Generation and deployment of context-specific machine learning models

By deriving and slicing the context-specific models from the basic machine learning model, the problem of insufficient storage and computing execution of large models on resource-constrained devices is solved, and efficient and low-cost machine learning model deployment and execution on client devices is achieved.

CN120380487APending Publication Date: 2025-07-25MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380086886.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-25
Filing Date
2023-12-21
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Large machine learning models have insufficient storage and computing resources when executed on resource-constrained client devices, resulting in high cost of cloud resources and latency and availability issues, and small models perform poorly in a wide range of contexts.

Method used

By deriving context-specific machine learning models from basic machine learning models, combining knowledge distillation and architectural search techniques, a context-specific model suitable for client device hardware is generated and sliced to appropriate sizes, and trained and deployed with cloud resources.

Benefits of technology

It realizes efficient execution of machine learning models on resource-constrained client devices, reduces the dependence of cloud resources, and improves the performance efficiency and accuracy of the model in a specific context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120380487A_ABST
    Figure CN120380487A_ABST
Patent Text Reader

Abstract

This document relates to automatic generation and deployment of machine learning models, such as neural networks. One example method involves obtaining a base machine learning model suitable for multiple contexts. The method further includes deriving, from the base machine learning model, a plurality of context-specific machine learning models adapted to different contexts of the plurality of contexts. The method also includes outputting a plurality of context-specific machine learning models for use in different contexts.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Machine learning models can be used in a wide range of applications. In some cases, machine learning models can be very large. For example, the GPT-3 model has approximately 175 billion parameters stored using 800 gigabytes. Large models such as BLOOM, GPT-3, ResNet-50, or NASNet Large tend to be more accurate than small models and can perform well in a wide range of contexts. However, there can be practical difficulties with the size of these models. Summary of the Invention

[0002] The present Summary of the Invention is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. The Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0003] This specification generally relates to techniques for the automatic generation of machine learning models. One example includes a method or technique that can be executed on a computing device. The method or technique can include: obtaining a plurality of context-specific machine learning models. Each context-specific machine learning model can be derived from a base machine learning model adapted to a plurality of contexts, and each context-specific machine learning model can be adapted to a different context among the plurality of contexts. The method or technique can further include: detecting a particular context of a particular device. The method or technique can further include: selecting a particular context-specific machine learning model from the plurality of context-specific machine learning models at least based on the particular context of the particular device; and providing the particular context-specific machine learning model to the particular device.

[0004] Another example includes a method or technique that can be executed on a computing device. The method or technique can include: obtaining a base machine learning model adapted to a plurality of contexts. The method or technique can further include: deriving a plurality of context-specific machine learning models adapted to different contexts among the plurality of contexts from the base machine learning model. The method or technique can further include: outputting the plurality of context-specific machine learning models for use in different contexts.

[0005] Another example includes a computing device that includes a hardware processing unit and storage resources. The storage resources store computer-readable instructions that, when executed by the hardware processing unit, cause the hardware processing unit to: receive a particular context-specific machine learning model adapted to a particular context. The particular context-specific machine learning model can be derived from a base machine learning model adapted to a plurality of contexts. The computer-readable instructions execute the particular context-specific machine learning model on the computing device when the computing device is in the particular context.

[0006] The examples listed above are intended to provide a quick reference to assist the reader and are not intended to define the scope of the concepts described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The detailed description is described with reference to the accompanying drawings. In the drawings, the leftmost digit of a reference numeral identifies the drawing in which the reference numeral first appears. Use of like reference numerals in different instances in the specification and the drawings may indicate like or identical items.

[0008] Figure 1 An example distillation scenario of knowledge from a base machine learning model to multiple context-specific machine learning models consistent with some embodiments of the present concept is shown.

[0009] Figure 2 An example deployment scenario of a context-specific machine learning model to a client device consistent with some embodiments of the present concept is shown.

[0010] Figure 3 An example teaching scenario for knowledge distillation from a base machine learning model to a context-specific machine learning model consistent with some embodiments of the present concept is shown.

[0011] Figure 4 An example search process for finding the architecture of a context-specific machine learning model consistent with some embodiments of the present concept is shown.

[0012] Figure 5A , Figure 5B , Figure 5C and Figure 5D Examples of variations that can be performed on a machine learning model during a search process consistent with some embodiments of the present concept are shown.

[0013] Figure 6 An example model generation workflow for generating a machine learning model consistent with some embodiments of the present concept is shown.

[0014] Figure 7A , Figure 7B and Figure 7C Scatter plots associated with successive iterations of a machine learning model search process consistent with some embodiments of the present concept are shown.

[0015] Figure 8 An example pruning scenario for knowledge distillation from a base machine learning model to a context-specific machine learning model consistent with some embodiments of the present concept is shown.

[0016] Figure 9 An example system consistent with some embodiments of the present concept is shown.

[0017] Figure 10 An example execution scenario of a compressed machine learning model consistent with some embodiments of the present concepts is shown.

[0018] Figure 11 An example graphical user interface consistent with some implementations of the present concepts is shown.

[0019] Figure 12 , Figure 13 and Figure 14 is a flow diagram of an example method or technique consistent with some implementations of the present concepts. DETAILED DESCRIPTION Overview

[0020] As mentioned above, large machine learning models tend to perform well in a variety of contexts. For example, consider GitHub Copilot, which starts with a pre-trained GPT-3-based model that has been tuned to generate code in a variety of different programming languages, such as Python, JavaScript, Perl, etc. The model is very large, so it is often run on cloud resources because the model is too large to run on most client devices. However, cloud resources tend to be quite expensive and may have latency and / or availability issues, especially during heavy use.

[0021] One approach for client-side execution of machine learning models involves training small models for a wide range of contexts. However, small models often do not perform as well as large models when trained directly on data that is trained for a wide range of contexts. For example, large models tend to learn conceptual abstractions that can help the model perform well in different contexts, but small models may not be able to learn such conceptual abstractions directly from training data.

[0022] The disclosed embodiments provide techniques for deriving context-specific machine learning models from base machine learning models. The context-specific machine learning models can be small enough so that they can be efficiently executed on client devices, which may have limited resources compared to cloud servers. In some cases, a context prediction scheme is employed to automatically detect the context of a particular client device, and the client device can then load and execute a corresponding context-specific machine learning model for the detected context.

[0023] Various schemes are provided for deriving context - specific machine learning models from a base machine learning model. The first scheme involves searching for a suitable architecture and then using the base machine learning model to teach different instances of the selected architecture using context - specific training data for different contexts. The second scheme involves using knowledge distillation to prune certain parameters from the base machine learning model or another model derived from the base model using the first scheme to obtain different context - specific machine learning models with fewer active parameters than the base machine learning model. Pruning can involve setting individual parameters of the model to zero that do not significantly contribute to the performance of the model in a particular context.

[0024] To further facilitate execution on resource - constrained client devices, each context - specific machine learning model can be compressed into corresponding slices. Each slice can have a corresponding size suitable for the hardware capabilities of the client device on which the model is to be executed. For example, each slice can include a parameter matrix for a specific layer of the machine learning model, where the parameters for that layer fit into the memory of the inference processing unit (e.g., a neural processing unit or “NPU”) of the client device. Machine Learning Background

[0025] There are various types of machine learning frameworks that can be trained to perform a given task, such as natural language understanding, natural language generation, detecting objects in images, generating images from text, etc. Support vector machines, decision trees, and neural networks are just a few examples of machine learning frameworks that have been used in a wide variety of applications. Some machine learning frameworks (such as neural networks) use operation layers or “nodes” connected together by one or more edges.

[0026] In a neural network, the nodes are connected to each other via one or more edges. A neural network can include an input layer, an output layer, and one or more intermediate layers. Each individual node in each layer can perform a specific operation on its inputs, such as a convolution operation, a vector operation, a matrix operation, a pooling operation, an activation function operation, an embedding operation, a decoding operation or an encoding operation, an attention operation, etc. Each operation can provide an output to a subsequent layer or, in some cases, to a previous layer. The input to a given node can be multiplied by the corresponding weight value of the edge between the input and the node. Additionally, the node can have individual bias values that are also used to produce the output. Various training processes can be applied to learn the edge weights and / or the bias values.

[0027] Neural networks and other machine learning models can be viewed as operating in two phases - training and inference. In the training phase, the model is used to make predictions on given training data, and the model parameters are updated based on whether those predictions are correct. In the inference phase, the trained model is used to process input data to perform a particular task, usually without further modification of the model parameters. Training can involve algorithms such as batch gradient descent, which performs the calculation of the error gradient to update the model parameters. In contrast, inference generally does not involve such calculations. Since inference processing does not necessarily involve training calculations, dedicated inference processing units, such as NPUs, can be built, which are particularly effective in performing inference operations and do not necessarily need to fully support model training.

[0028] For example, the inference processing unit can be implemented as a systolic array. In such an array, the input data can be partitioned and distributed to a set of parallel nodes, each parallel node performing the same operation on a subset of the input data and then passing their processing results to the next set of nodes in the array. For example, individual hardware nodes can perform multiplication and accumulation operations on specified data sizes very quickly, e.g., performing large convolution operations or matrix multiplication operations in relatively few processing cycles using dedicated circuitry and using relatively few memory transfer operations. For example, a large matrix or vector can be loaded into or output from memory in a single operation.

[0029] In contrast, implementing large convolutions or matrix multiplications on a conventional CPU often involves performing a long sequence of sequential operations. Individual instructions can retrieve parts of the matrix or vector from memory and process that part individually, and then subsequent operations can combine the intermediate results into the final output. Therefore, the efficiency of complex convolution operations or matrix operations on a general-purpose CPU is often much lower than when implemented on a dedicated inference processing unit. The CPU implementation of large convolution operations or matrix operations can have much higher latency, power consumption, and / or memory utilization than the same operations implemented with hardware-supported inference operations.

[0030] However, inference processing units often have certain practical limitations. For example, the SRAM capacity of an NPU can be around 500KB to a few megabytes. Therefore, it is usually not practical to execute a full-context based machine learning model on an NPU because the layers of a full-context based model may be too large to fit into the SRAM. In addition, full-context based machine learning models can include operations that are not supported by the NPU. Model Distillation

[0031] Figure 1Fig. 100 shows a distillation scenario for generating context-specific machine learning models. The base machine learning model 102 is processed using context-specific distillation 104 to obtain context-specific machine learning models 106(1), 106(2), and 106(3). The distillation can be performed using cloud resources (e.g., servers) in the cloud 108. Note that in some cases, the following discussion generally refers to any one of the context-specific machine learning models 106(1), 106(2), or 106(3) as the "context-specific machine learning model 106" without parentheses.

[0032] In many cases, the base machine learning model is a large model, e.g., having many parameters that are not suitable for execution on resource-constrained client devices. The base machine learning model can be adapted to multiple contexts. For example, as mentioned above, Copilot is an example of a base machine learning model that is adapted to code generation in multiple programming languages. As another example, a convolutional neural network (such as ResNet-50 or NASNet Large) can be adapted to identify objects for different contexts. For example, one context can involve identifying different animal species in wildlife photos, and another context can involve identifying organ damage in medical images.

[0033] Context-specific distillation 104 can involve transferring knowledge of a specific context from the large base machine learning model to a small context-specific machine learning model. For example, if the base machine learning model 102 is adapted to generate code in different programming languages, each context-specific machine learning model 106 can be adapted to generate code in a single programming language out of those programming languages. Similarly, if the base machine learning model 102 is adapted to identify thousands of object types in images, each context-specific machine learning model can identify a different subset of those object types. Model Deployment

[0034] Figure 2 Fig. 200 shows a deployment scenario. A request including context data 204 is received from a client device 202. Based on the context data, a selected context-specific machine learning model 206 is selected from the available context-specific machine learning models 106(1), 106(2), and 106(3). The selected context-specific machine learning model is deployed from the cloud 108 to the client device.

[0035] For example, if each context-specific machine learning model 106 is adapted to generate code in different programming languages, the context data 204 can include code snippets that can be used to determine what programming language the user of the client device 202 is currently coding in. If each context-specific machine learning model is adapted to identify different types of objects, the context data can indicate what application the user of the client device 202 is using, e.g., a medical imaging application versus a social media application where the user frequently posts and views images of wildlife. Teaching scenario

[0036] One way to implement knowledge distillation from a base machine learning model to a context-specific machine learning model involves using the base machine learning model as the teacher and the context-specific machine learning model as the student. Figure 3 A teaching scenario 300 for obtaining a context-specific machine learning model is shown. The base machine learning model 102 and the context-specific machine learning model 106 are evaluated on a context-specific training dataset 302. The parameters of the context-specific machine learning model 106 are adjusted using a standard loss 304 as well as a distillation loss 306.

[0037] The standard loss 304 reflects how accurately the context-specific machine learning model 106 predicts the value (e.g., label) of an example in the context-specific training dataset 302. The distillation loss 306 reflects how closely the output of the context-specific machine learning model matches the output of the base machine learning model 102. The distillation loss for any individual training example can include a term based on the difference between the output distribution of the base machine learning model for that training example and the output distribution of the context-specific machine learning model for that training example.

[0038] For example, assume the training example includes an image of a horse. If the base machine learning model 102 predicts that the image is a horse with a score of 0.7 and a zebra with a score of 0.3, the distillation loss for that training example increases as the output distribution of the context-specific machine learning model 106 further deviates from 0.7 horse, 0.3 zebra. Training according to the standard loss encourages the context-specific model to learn the correct labels from the training dataset, while training according to the distillation loss encourages the context-specific model to replicate the output distribution of the base machine learning model.

[0039] In some cases, multiple student models can be derived from a given base machine learning model. For example, the same student model architecture can be trained using the distillation loss from the same base machine learning model but with different context-specific datasets. In other words, the teaching scenario 300 can be performed as follows: first, a first context-specific machine learning model is trained using a first context-specific training dataset, second, a second context-specific machine learning model is trained using a second context-specific training dataset, and so on, where each context-specific model has the same architecture. However, in other cases, the student models can also have different architectures. Model architecture search

[0040] Generally, it may be useful for the context-specific machine learning model to be smaller than the base machine learning model. One way to obtain the student model architecture is to manually select the architecture of the student model. Another way is to perform neural architecture search using an evolutionary scheme, a reinforcement learning scheme, a Bayesian optimization scheme, a hill climbing scheme, a one-shot scheme, etc. The following discussion uses an evolutionary scheme as a specific example of how the disclosed concepts can be employed to automatically generate a machine learning model to be used as a context-specific student model. However, as discussed further below, the disclosed concepts can be easily incorporated into other schemes for generating machine learning models.

[0041] Figure 4 An example machine learning model evolutionary search process 400 for determining the architecture of a context-specific machine learning model 106 is shown. First, a parent model 410 is modified and trained to include a trained submodel 420. As described in more detail below, in some cases, the parent model is modified subject to certain constraints (such as hardware limitations). For example, the constraint can indicate that the modification includes inference operations supported by the target inference hardware architecture (e.g., a specific model of an NPU) or inference operations adapted within the memory (e.g., SRAM) constraints of the target inference hardware architecture. Additionally, in some cases, the parent model is modified by removing other inference operations that do not satisfy the constraints. The parent model can also be modified by adding or removing connections between individual inference operations.

[0042] Next, the trained sub-model 420 is pruned to obtain the pruned sub-model 430. For example, the trained sub-model can be pruned to remove individual sub-models that perform relatively poorly compared to one or more metrics. As discussed in more detail below, the metrics can relate to loss or accuracy, latency, power consumption, memory utilization, etc. In some cases, a distillation loss or "soft loss" value can be used to determine the metric, which is based on the difference between the output distribution produced by a given sub-model for a training example and the output distribution produced by the underlying machine learning model when evaluating the same training example.

[0043] After model pruning, the remaining trained sub-models are designated as the next generation of parent models 440. Further iterations of the model search can be performed by training and pruning further sub-models until a stopping condition is reached, at which point the final model can be selected from the available sub-models. Example Model Modifications

[0044] Figure 5A 、 Figure 5B 、 Figure 5C and Figure 5D illustrate example variations that can be performed to transform a parent model into a sub-model. Figures 5A to 5D These concepts are conveyed using a convolutional network architecture, but the concepts shown herein can be applied to other types of neural network architectures, such as transformer-based networks, long short-term ("LSTM") networks, etc.

[0045] Figure 5A Illustrated is a candidate inference operation 500, which includes three specific inference operations - CONV X, CONV Y, and CONV Z. In some cases, each candidate inference operation can be selected to meet hardware constraints. As an example of a hardware constraint, each candidate inference operation can have a specified size within the SRAM suitable for the target inference processing unit architecture. As another example of a hardware constraint, each candidate convolutional operation can have a specific input / output tensor size and / or kernel size supported by dedicated circuitry on the target inference processing unit architecture, or otherwise operate more efficiently on the target inference processing unit architecture than on a conventional CPU. In other words, a given inference processing unit implementing the target inference hardware architecture has dedicated circuitry for potentially performing a convolutional operation with those tensor and / or kernel sizes using, for example, a single machine instruction (e.g., opcode). The dedicated circuitry for implementing a given convolutional operation can have parallel hardware nodes, each of which can perform a part of the convolutional operation on a part of the input data to produce a part of the output data. Additional circuitry provided by the inference processing unit implementing the target hardware architecture can be used to combine and further process the output data.

[0046] The seed model 502 can be a model initially developed without considering the target inference architecture. For example, the seed model 502 can be a model developed manually or using automated techniques and is known to perform well for a specific task, such as a particular image processing operation (e.g., background segmentation, object recognition, etc.) or a natural language processing operation (e.g., natural language understanding, natural language generation). As shown, the seed model 502 includes three types of convolutional operations A, convolutional operation B, and convolutional operation C that are not necessarily supported by the target hardware architecture. In other words, convolutional operation A, convolutional operation B, and convolutional operation C may not be suitable for the SRAM of a given inference processing unit and / or may have different tensor and / or kernel sizes that are not supported in hardware. Note that the seed model 502 can be much smaller than the base machine learning model.

[0047] As described in more detail below, multiple iterations of a machine learning model search process can be performed starting from the seed model 502 as the parent model. In each iteration, one or more operations or connections between operations can be added or removed until a final model is generated. The final model can be suitable for execution on the target inference hardware architecture, e.g., because each operation fits the SRAM of the target inference hardware architecture and / or is supported in hardware by the target inference hardware architecture.

[0048] In the first model search iteration, the convolutional A operation 503 of the seed model 502 can be replaced to generate submodels 504, 508, and 512. Submodel 504 can be generated by replacing the convolutional A operation with a convolutional X operation 506, submodel 508 can be generated by replacing the convolutional A operation with a convolutional Y operation 510, and submodel 512 can be generated by replacing the convolutional A operation with a convolutional Z operation 514. As described above, the corresponding submodels can be trained and one or more submodels can be selected as the parent model for the next generation. The submodels can be trained on a context-specific training dataset using standard loss and / or distillation loss from a much larger base machine learning model.

[0049] For the purpose of example, assume that submodel 508 is selected as the parent model for the next generation and is re-designated as parent model 516 in Figure 5B This parent model can be modified by replacing the convolutional B operation 518 to generate submodels 520 and 526. Submodel 520 can be generated by replacing the convolutional B operation with a convolutional X operation 522 and a convolutional Z operation 524. Submodel 526 can be generated by replacing the convolutional B operation with a convolutional Y operation 528 and a convolutional Y operation 530. As described above, the corresponding submodels can be trained on a context-specific training dataset and one or more submodels can be selected as the next generation parent model.

[0050] For purposes of example, assume that sub-model 526 is selected as the parent model for the next generation and is re-designated as parent model 532 in Figure 5C . The parent model can be modified by replacing convolutional C operation 534 to generate sub-model 536 and sub-model 542. Sub-model 536 can be generated by replacing convolutional C operation with convolutional X operation 538 and ReLu operation 540. Sub-model 542 can be generated by replacing convolutional C operation with convolutional Z operation 544.

[0051] Now, assume that sub-model 542 is selected as the parent model for the next generation and is re-designated as parent model 546 in Figure 5D . The parent model can be modified by replacing convolutional A operation 548 to generate sub-model 550 and sub-model 554. Sub-model 550 can be generated by replacing convolutional A operation with convolutional Z operation 552. Sub-model 554 can be generated by replacing convolutional A operation with convolutional Y operation 556.

[0052] At this point, a stopping condition can be reached and a final model is selected from the models generated so far. For example, sub-model 550 can be selected as the final model, shown in bold text in Figure 5D . The final model can be output to be executed on an inference processing unit to perform the specific task that the model has been trained to do. Thus, referring back to Figure 5A , the seed model 502 has been transformed into a final model that can perform the same task as the seed model, but using inference operations that satisfy the hardware constraints associated with the target inference hardware architecture. Thus, the final model can provide similar functionality to the seed model while being suitable for execution on the target inference hardware architecture.

[0053] Although Figures 5A to 5D conveys how an evolutionary process can be used to search for convolutional architectures, similar schemes can be used for other architectures over time, such as transformer-based architectures. In the case of a transformer architecture, the search can consider different embedding sizes, different numbers / sizes of encoder layers and / or decoder layers, the number of attention heads, the size of the feed-forward or other layers in the model, etc. These characteristics of the transformer-based model can be modified to meet the corresponding memory limitations and / or hardware operations supported by a given target architecture. Example model generation workflow

[0054] Figure 6 shows an example model generation workflow 600 that can be used to search a machine learning model space. Hardware constraint memory 602 stores one or more hardware constraints, such as SRAM size or specific inference operations supported in the hardware, such as those described above in Figure 5Ato the convolution operations X, Y, and Z shown in FIGS. A-D. The parent model memory 604 stores a parent model that can be replaced over time with a new parent model, as described in more detail below. In some cases, one or more seed models are used to initialize the parent model memory, e.g., selected based on their performance (e.g., accuracy) at a particular task. Parent models of subsequent generations can be used to populate the parent model memory over time.

[0055] For each generation, one or more parent models 606 can be retrieved from the parent model store and input into the submodel generation 608. The submodel generation can modify the parent model consistent with the constraints in the hardware constraint store 602 to produce a submodel 610. The submodel can be trained at 612 to produce a trained submodel 614, e.g., using supervised learning, unsupervised learning, transfer learning, etc., with or without a distillation loss from a base model. The training can be based on a context-specific training dataset 302, which can include training examples processed by each submodel during training. In embodiments where a distillation loss is considered during model generation, the base machine learning model can also process each training example in the context-specific training dataset to determine the distillation loss.

[0056] The trained submodel can be executed at 616 to obtain a metric 618. For example, the metric can characterize the accuracy or loss (standard or distillation) of the trained submodel, the latency (e.g., execution time) of the trained submodel, the power consumption or memory utilization of the trained submodel, etc. The metric can be used to evaluate the trained submodel to identify a selected submodel 622. For example, in some cases, a submodel is selected based on a trade-off between two or more metrics, e.g., by selecting a submodel with a relatively low combined loss and a relatively low power consumption. A similar scheme can also be used to select a final model 624 when a stopping condition is reached.

[0057] The submodel generation 608 can involve replacement or addition operations and / or connections between operations, as described above with respect to Figures 5A to 5D shown. For example, the submodel generation can involve a random or deterministic scheme for selecting operations to add to the parent model, operations to remove from the submodel, and / or connections to add or remove between various operations. In some embodiments, the submodel generation is fully constrained by the hardware constraints. Evaluate and designate the submodel as a parent model

[0058] As described above, certain submodels are selected during evaluation 620 and added to the parent model memory 604 to be used as the parent model for subsequent generations. One scheme for deciding which submodels to add to the parent model memory involves using one or more metrics to predict which submodels are likely to produce offspring that will exhibit improvement relative to previously discovered models in subsequent iterations. Generally, the metrics can consider factors such as the standard and / or deviation loss or accuracy of a given submodel, the latency of a given submodel, the power consumption of a given submodel, the computational resource consumption of a given submodel (e.g., memory consumption). Submodels that exhibit characteristics such as relatively low loss or high accuracy, low latency, low power consumption, and / or low computational resource consumption can be favored for selection as the parent model in the next generation.

[0059] This document relates to Figure 7A illustrates a specific method for selecting submodels for the parent pool. The figure shows an example scatter plot 700 of various trained models. For each submodel that has completed training, the cost of that submodel can be calculated and plotted on the x-axis 702, where the cost can be defined based on latency, power consumption, computational resource consumption, etc. In some cases, the cost can be normalized to a number between 0 and 1, as Figure 7A shown. Additionally, the loss of that submodel can be calculated and plotted on the y-axis 704. Here, a combined loss function that takes into account both the standard loss and the distillation loss can be employed, such as a weighted sum of these two values. Once all the models for a given iteration are plotted, the lower convex hull 706 can be calculated from the plotted values. Note, however, that some schemes can use only the distillation loss or only the standard loss to select which submodels are chosen as the parent models for subsequent generations.

[0060] Additionally, note that the distillation loss can also be used to partially train the submodel and update the parameters of the submodel. In other words, distillation can be used for two different purposes - selecting which submodels to use as the parent models for the growth of the next generation models, and updating the model parameters during the training iterations. For example, the partial training of the submodel can be performed using only the distillation loss as the loss function or using a weighted combination of distillation and standard loss as the loss function.

[0061] The lower convex hull 706 can be used as the mechanism for deciding whether a given submodel is added to the parent model pool. For example, the submodels on the lower convex hull can be added to the parent model pool with a probability defined by the following specific algorithm. If m1 and m2 are two adjacent models on the hull, with costs c1 and c2 (c1 < c2), then the probability weight of m1 can be set in proportion to c2 - c1. The most accurate model according to the combined loss function, which has no following model on the curve, can be included in the parent model pool with a probability of 0.5. In Figure 7A the example, the most accurate model is model 708 because this model has the lowest combined loss.

[0062] Typically, the lower convex hull is a subset of the Pareto boundary, so another option is to select submodels on the Pareto boundary to include in the parent pool. Either option can provide good performance for selecting the submodels to add to the parent model pool. One way to view the lower convex hull and / or the Pareto boundary is as follows. Given a model on the lower convex hull or Pareto boundary, it cannot be improved with respect to one metric by moving to another model on the lower convex hull / Pareto boundary without degrading another metric.

[0063] Note that due to the randomness in forming the stochastic gradient, the same model may have different validation errors. Thus, the lower convex hull or Pareto boundary can be relaxed with a multiplicative bandwidth. Thus, submodels whose validation errors are within (1 + γ) times the lower convex hull validation error at the same computational cost can be considered to be on the lower convex hull and can be selected as parent models. Some implementations can be set to γ = 0.025. This option allows certain submodels that are close to but not strictly on the lower convex hull to still be designated as parent models.

[0064] Other options can also be used to allow submodels with positions within a predetermined vicinity of the lower convex hull to be selected as parent models. For example, some implementations can define a threshold distance from the lower convex hull and allow submodels within the threshold distance of the lower convex hull to be selected as parent models. This is just one of the various methods that can be used to select a subset of one or more submodels as parent models based on one or more metrics.

[0065] Figure 7A The models that are finished training are shown as black dots. For purposes of explanation, assume Figure 7A represents the state of scatter plot 700 after iteration N. One or more of the submodels on or near the lower convex hull 706 can be selected as parent models for the subsequent iteration N + 1, where additional operations are added from additional submodels as described above.

[0066] Figure 7B Scatter plot 700 in a subsequent state after iteration N + 1 is shown. Submodels trained during iteration N + 1 are shown as squares in Figure 7B . A new lower convex hull 710 can be calculated. The previous lower convex hull 706 is shown as a dashed line to show that the lower convex hull moves downward in iteration N + 1.

[0067] Similarly, one or more submodels in or near the lower convex hull 710 can be selected for the subsequent iteration N + 2. Submodels trained during iteration N + 2 are shown as triangles in Figure 7C . A new lower convex hull 712 can be calculated, and the previous lower convex hulls 706 and 710 are shown as dashed lines to show their positions relative to the lower convex hull 712.

[0068] View Figures 7A to 7C One way to view the scheme shown is a greedy scheme for finding cost-effective predictors. Note that this is a multi-objective approach that takes into account loss / accuracy and model performance with respect to latency, power consumption, or resource utilization. Alternative embodiments may use different metrics and / or additional metrics, such as multi-dimensional plots of three or more metrics, objective functions defined over one or more metrics, etc.

[0069] The above scheme typically uses a randomization scheme to grow the network. However, instead of a purely random scheme that may be computationally infeasible, the scheme is guided by favoring the selection of known good models as a basis for additional variants. As previously mentioned, training a model from scratch can be computationally very intensive. For example, the training dataset may include millions of training data items, and a given model may need to be trained over several training epochs before convergence. A training epoch can involve one forward propagation and one backward propagation operation of the entire model through each data item in the training dataset.

[0070] The above method provides various benefits over conventional methods for automatic model generation. Note that not every sub-model is used as the parent model for subsequent iterations. Instead, by using a subset of the sub-models that occur along the lower convex hull as the new parent models, the disclosed embodiments inherit the parent model structure of the known good models at each new iteration starting from the sub-model structure. This allows subsequent iterations to proceed without training models that occupy most of the search space far from the lower convex hull, and can save a significant amount of training time. Additionally, by using not only accuracy but also cost as the criterion for selecting which sub-models are used as the new parent models, the disclosed embodiments discourage the generation of new models that tend to have high latency or consume a large amount of power or computational resources.

[0071] Recall that prior art for the automatic generation of machine learning models typically does not consider distillation loss. In contrast, the techniques described herein can generate model architectures that satisfy specific hardware constraints while considering distillation loss with respect to a known base machine learning model. Additionally, the search can also consider other characteristics of the resulting model, such as latency, power consumption, and resource utilization.

[0072] In some cases, the same final architecture is used as the student model and trained on different context-specific training datasets to obtain different context-specific machine learning models. In other cases, the search can be performed using different context-specific training datasets with distillation loss as the evaluation metric. In this case, the resulting final architectures can vary. For example, different context-specific models not only have different weight and bias values but also have context-specific architectures found via neural architecture search. Simulation and Emulation

[0073] In some cases, the above techniques can be performed by using general-purpose hardware to simulate target hardware, without implementing a hardware execution model of the target inference hardware architecture and without emulating the target inference hardware architecture. Execution of the simulation model can involve approximating the functionality of a given target hardware architecture, without directly implementing the underlying inference operations supported by the target hardware architecture. Using simulation, the accuracy of a given model can still be predicted because general-purpose hardware can be used to simulate and execute operations that are mathematically or logically equivalent to those supported by the target inference hardware architecture.

[0074] For example, on a general-purpose CPU, it may take hundreds or thousands of operations and processing cycles to implement a convolution operation or a matrix operation, which can be performed by an NPU in just a few processing cycles using a single operation. However, because the operations are mathematically or logically equivalent (or at least approximately equivalent), the accuracy or loss of a given model can still be estimated on the CPU. Thus, a general-purpose CPU can be used to transform an initial seed model architecture with operations not supported by the target inference hardware architecture into a final model fully supported by the target inference hardware architecture. Even assuming that simulation on the CPU cannot estimate the performance of the final model with respect to latency, power consumption, or resource utilization without emulation, nevertheless, given that the final model can leverage the efficiency provided by the target inference hardware architecture by using hardware-supported operations and / or layers that fit the memory constraints of the target architecture, the final model will still exhibit a significant improvement over the seed model.

[0075] In additional embodiments, hardware emulation of the individual operations supported by the target inference hardware architecture can further guide the search. In hardware emulation, the CPU can implement operations that directly correspond to the inference operations of a given model. In other words, the CPU can replicate the target hardware architecture by mapping each inference operation in the given model to a corresponding CPU instruction set designated to emulate that inference operation. By using emulation, performance information about each model can be inferred. For example, the total latency, power consumption, and / or resource utilization of a given model can be estimated based on the emulation. This allows for a multi-objective search, where sub-models can be selected as parent models in the next generation based not only on accuracy or loss but also on the performance of each model. Alternative Search Techniques

[0076] The evolutionary search process was used above to convey the concepts described herein to show how the machine learning model space can be searched while taking into account the availability of hardware-supported inference operations. However, the above specific techniques can be easily extended to a variety of other schemes for the automatic generation of machine learning model architectures.

[0077] For example, consider a solution that uses reinforcement learning to find a new model architecture. In some embodiments, an exploration strategy can be provided that encourages searching for model architectures that satisfy hardware constraints such as SRAM limits or operations with hardware support. As another example, consider a method that uses Bayesian optimization to explore new machine learning models. In some embodiments, an acquisition function that takes into account hardware constraints can be defined when determining which models to explore. Similar solutions can be used for one-shot model generation, e.g., by defining a supernetwork with candidate inference operations that satisfy hardware constraints, training those operations together, and then pruning the supernetwork to select a particular path through the supernetwork as the final model. Pruning scenario

[0078] Another way to implement knowledge distillation from a base machine learning model to a context-specific machine learning model involves pruning the base machine learning model to produce a smaller context-specific machine learning model. Figure 8 A pruning scenario 800 for obtaining a context-specific machine learning model is shown. A base machine learning model 102 is used to process different context-specific training datasets 302(1), 302(2), and 302(3) to obtain pruned models 802(1), 802(2), and 802(3).

[0079] One way to prune a model involves magnitude pruning, where the model is executed on a given context-specific dataset and parameters with relatively low magnitudes are pruned from the model. Another way to prune a model involves gradient pruning, where parameters are pruned based on the error gradients of a given training dataset. Generally, pruning involves changing individual parameters (e.g., weights) to zero so that the parameters can be easily compressed and implemented using simple no-operations in inference hardware. Pruning can be done in a structured or unstructured manner. In unstructured pruning, weights are pruned individually. In structured pruning, an entire layer (e.g., a convolutional filter, an attention layer, etc.) can be pruned at once.

[0080] Note that each pruned model can have a different architecture from the initial model. In other words, different layers and / or connections can be removed from the individual models. For example, in the case of a convolutional model, a layer can be effectively removed by setting all parameters of a given convolutional filter to zero. In the case of a transformer model, the initial model can be pruned by removing one or more attention heads, encoder layers, or decoder layers.

[0081] In some cases, the teaching and pruning scenarios can be combined. First, a search is performed to identify a suitable architecture. Then, corresponding instances of the architecture are trained using different context-specific training datasets to obtain context-specific machine learning models. Further training on context-specific data can be performed to perform personalized pruning of the context-specific machine learning models. Example system

[0082] This embodiment can be executed on various devices in various scenarios. Figure 9 An example system 900 in which this embodiment can be employed is shown and discussed in more detail below.

[0083] As Figure 9 shown, system 900 includes client devices 910, server 920, client device 930, and client device 940 connected via one or more networks 950. Note that client devices can be embodied as mobile devices such as smart phones or tablets and fixed devices such as desktop computers, server devices, etc. Similarly, various types of computing devices can be used to implement the server. In some cases, Figure 9 any of the devices shown, but particularly the server, can be implemented in a data center, server farm, etc.

[0084] Figure 9 Certain components of the devices shown herein may be referred to by reference numerals in parentheses. For the purposes of the following description, parentheses (1) indicate the occurrence of a given component on client device 910, (2) indicate the occurrence of a given component on server 920, (3) indicate the occurrence on client device 930, and (4) indicate the occurrence on client device 940. Unless a specific instance of a given component is identified, this document will generally refer to the component without parentheses.

[0085] Generally, devices 910, device 920, device 930, and / or device 940 may have corresponding processing resources 901 and storage resources 902, which will be discussed in more detail below. The devices may also have various modules that use the processing resources and storage resources to perform the techniques discussed herein. The storage resources may include both persistent storage resources (such as magnetic drives or solid state drives) and volatile storage devices (such as one or more random access memory devices). In some cases, the modules are provided as executable instructions stored on a persistent storage device, loaded into a random access memory device, and read by the processing resources from the random access memory for execution.

[0086] The client device 910 may include a configuration module 911 that can interact with the model generation module 921 on the server 920. Generally speaking, the configuration module can provide certain configuration parameters to the model generation module. The model generation module uses these configuration parameters to perform model generation as discussed herein. The model generation module can derive context-specific models from a base machine learning model. The context detection module 922 can detect the context on each client device. The model providing module 923 can provide different context-specific machine learning models to each client device according to the context detected based on the corresponding context data received from each client device.

[0087] The client device 930 and the client device 940 may have corresponding instances of the context reporting module 903 and the model execution module 904. For example, the context reporting module can send context data to the server 920, and this context data can be processed by the context detection module 922 to predict the current context on a given client device. The model execution module 904 can execute the context-specific machine learning model returned by the server 920 after the context is detected. Context Detection

[0088] In some cases, the context detection module 922 can detect the context of a given client device based on the user's manual input. For example, the user can manually select C++ as the programming language. In other cases, context detection can be performed automatically. For example, certain programming languages tend to use certain operators more than other programming languages. For example, using symbols such as "@" and "$" may mean that the user is programming in Perl, while using many parentheses may mean that the user is programming in Lisp. Therefore, in some cases, the client device can send context data, such as a code snippet, to the server 920, and the server can detect the specific language used on the client device. In other cases, automatic context detection can be performed on the client device.

[0089] As another example, the user may be programming in Java and start doing some statistical work. In this case, context detection can look for the libraries used by the user (e.g., statistical libraries) and operations (e.g., many mathematical operations). At this time, a context-specific machine learning model for Java statistical code generation can be provided to the user. Later, the user can start doing some Java programming to interface with a database, and context detection can identify that the user is adopting SQL statements in certain function calls. Then, a context-specific machine learning model for Java database development can be provided to the user. Both context-specific machine learning models can be adapted to generate Java code, but for different programming scenarios.

[0090] In the case of image processing, a user can start by posting pictures of their pet online using a social media application. The user's post can include natural language text that can imply the pet context. For example, words like "puppy", "Rover", and "Siamese cat" can allow an automatic context detection algorithm to select a context-specific image processing model to identify the type of pet in the image. Later, the user can start using a medical application and use natural language text such as "liver" or "CT scan", and the automatic context detection algorithm can provide a context-specific image processing model for identifying objects in medical images in response to detecting that the user device has switched to another context. Compression and decompression

[0091] In some cases, the processing resources 901(3) and 901(4) can include a conventional CPU and an inference processing unit such as an NPU. In this case, the storage resources 902(3) and 902(4) can include storage resources for the CPU, such as solid state disk drives and / or main memory (RAM), and an inference processing unit memory, e.g., the internal memory of the NPU (SRAM).

[0092] Figure 10 An example execution scenario 1000 is shown, where the main memory 1002 of the client device stores a compressed model 1004. For example, the compressed model can be distributed by the server 920 to the client device. The model can be compressed by the server 920 in "slices" to obtain a compressed version, where each slice can include parameters of different layers, such as convolutional layers, attention layers, encoding layers, or decoding layers, etc. For example, each layer of a given model can be compressed on the server using ZIP or another compression algorithm. Recall that in some embodiments, the parameter matrix can be partially zeroed during pruning, and long runs of the same value (such as zeros) allow for efficient compression, e.g., a high compression ratio.

[0093] The CPU 1008 can retrieve the compressed slice 1006 from the main memory 1002 and perform decompression 1010 on the slice to obtain a decompressed slice 1012. The decompressed slice can be loaded into the SRAM 1014 of the NPU 1016. The NPU can use the processing circuitry 1018 to perform various operations on the decompressed slice, such as inference operation 1020 and inference operation 1022. This process can be repeated for each slice of the model until a final result is obtained. Example graphical interface

[0094] As described above, the configuration module 911 on the client device 910 can provide initial configuration parameters to the model generation module 921. The model generation module 921 can derive one or more context-specific machine learning models from a base machine learning model according to the configuration parameters provided by the configuration module.

[0095] Figure 11 An example configuration graphical user interface ("GUI") 1100 that can be presented on the client device 910 for the user to define these configuration parameters is shown. The base model element 1101 allows the user to specify what base model should be used to derive the context-specific model. In Figure 11 it, the user has selected the Copilot model, which is a large transformer-based model for generating programming code.

[0096] The derivation type element 1102 allows the user to specify what type of derivation to adopt. Here, the user has selected NAS, such as neural architecture search. Other options can include only pruning options and / or NAS plus pruning options. When the user selects the NAS option, the model generation module 921 can provide a default neural network structure to be used as a general seed model. Other options can include randomly generated models, where the module generation module selects a random model structure to be used as the seed model. Another option is for the user to navigate to an existing seed model that is known to provide relatively good performance for a specific task. In this case, the configuration module 911 can load the specified seed model into the model generation module to be used as the seed model.

[0097] The target architecture element 1103 allows the user to select the target inference hardware architecture to guide the search. In Figure 11 it, the user has selected the NPU model C, which can have a specific SRAM size and / or dedicated circuitry for performing specific operations of a specific inference hardware architecture. The model generation module can perform architecture search subject to the constraint that each layer fits into the SRAM and / or matches one of the inference operations supported by the circuitry of the NPU model C.

[0098] The metric 1 element 1104 allows the user to specify a first metric for evaluating the model, and the metric 2 element 1105 allows the user to specify a second metric. Here, the user has selected latency as the first metric and combined loss as the second metric. In other words, the user wishes to search the space of available model architectures that will exhibit relatively low latency while having a relatively low combined loss, where the combined loss is a function defined using both the standard loss and the derivation loss relative to the base machine learning model when executed on a context-specific training dataset.

[0099] Note that Figure 11The configuration parameters shown are merely exemplary, and various other embodiments are contemplated. For example, in some cases, the GUI may provide elements that allow a user to specify the location of a given context-specific training dataset for NAS or pruning-based derivation. As another example, the GUI may provide elements that allow a user to specify a budget for architecture search (e.g., specified GPU days) to be used as a stopping condition. As another example, the GUI may provide elements that allow a user to define the respective weights of a standard loss and a derivation loss on a context-specific training dataset. Thus, if the user weights the derivation loss relatively higher than the standard loss, the search will tend to preferentially find model architectures that approximate the performance of a base machine learning model. On the other hand, if the user weights the derivation loss relatively lower than the standard loss, the search will tend to preferentially find model architectures that perform well in matching labels from a context-specific training dataset.

[0100] In addition, note that some embodiments may provide one or more GUIs to show the progress of the model search. For example, some embodiments may generate a GUI that shows the variation of scatter plot 700 across different iterations of model growth in a manner similar to that shown in Figure 7A , Figure 7B and / or Figure 7C . Other embodiments may show a graphical representation of an individual model at the time of its generation. Method for providing a context-specific machine learning model

[0101] Figure 12 Illustrates an example method 1200 consistent with some embodiments of the present concept. Method 1200 may be implemented on many different types of devices, e.g., by one or more cloud servers, by a client device such as a laptop computer, a tablet computer, or a smart phone, or by a combination of one or more servers, client devices, etc.

[0102] Method 1200 begins at block 1202, where a plurality of context-specific machine learning models are obtained. As previously described, context-specific machine learning models may be derived from a large base machine learning model adapted to different contexts.

[0103] Method 1200 continues at block 1204, where the context of a particular device is detected. As previously described, in some cases, the context is manually identified by the user. In other cases, the context is automatically detected by using an automatic context prediction algorithm. For example, an SVM or a neural network classifier may classify programming code into different programming languages or different programming scenarios (e.g., a statistical program versus a database program).

[0104] Method 1200 continues at block 1206, where a particular context-specific machine learning model is selected. The particular context-specific machine learning model can be a model that is derived from a base machine learning model and adapted to a particular context using context-specific training data.

[0105] Method 1200 continues at block 1208, where the particular context-specific machine learning model is provided to a particular device. For example, the particular context-specific machine learning model can be sent from a cloud server to a client device for use while the client device remains in a particular context. In some cases, block 1208 can involve temporarily executing the particular context-specific machine learning model or the base machine learning model entirely on the cloud server for a period of time until the particular device is able to start locally executing the particular context-specific machine learning model. The method ensures that the functionality of the model remains available to the particular device via the cloud server after the context has changed (e.g., during the time the particular device downloads a particular context-specific machine learning model from the cloud). Method for generating a context-specific machine learning model

[0106] Figure 13 An example method 1300 consistent with some embodiments of the present concept is shown. Method 1300 can be implemented on many different types of devices, e.g., by one or more cloud servers, by a client device such as a laptop computer, a tablet computer, or a smart phone, or by a combination of one or more servers, client devices, etc.

[0107] Method 1300 begins at block 1302, where a base machine learning model is obtained. The base machine learning model can be a large model, such as BLOOM, GPT-3, ResNet-50, or NASNet Large, which is already adapted to multiple different contexts when received.

[0108] Method 1300 continues at block 1304, where multiple context-specific machine learning models are derived from a base machine learning model. As previously described, one way to derive context-specific machine learning models is by using the base machine learning model as a teacher, where the context-specific machine learning model and the base machine learning model are executed on corresponding context-specific training datasets to transfer knowledge from the base machine learning model to the context-specific machine learning model. Knowledge can be transferred by adjusting the parameters of the given context-specific machine learning model based on a loss function that considers the difference in the respective output distributions of the base machine learning model and the given context-specific machine learning model. Another way to derive context-specific machine learning models is to prune the parameters of the base machine learning model. The base machine learning model can be evaluated on the context-specific training dataset to determine which parameters to prune.

[0109] Method 1300 continues at block 1306, where the multiple context-specific machine learning models are output. For example, in some cases, the context-specific machine learning models are stored on a cloud server for subsequent distribution to individual client devices on a context-specific basis. Method for executing a context-specific machine learning model

[0110] Figure 14 An example method 1400 consistent with some embodiments of the present concept is shown. Method 1400 can be implemented on many different types of devices, for example, by one or more cloud servers, by client devices such as laptops, tablets, or smartphones, or by a combination of one or more servers, client devices, etc.

[0111] Method 1400 begins at block 1402, where a particular context of a computing device is detected. For example, context detection can be performed locally on the computing device or remotely on a server.

[0112] Method 1400 continues at block 1404, where a particular context-specific machine learning model adapted to the particular context is received. In some cases, the particular context-specific machine learning model is compressed into individual slices when it is received.

[0113] Method 1400 continues at block 1406, where the particular context-specific machine learning model is executed. In some cases, execution can involve using a CPU to decompress the individual slices of the model, loading the individual slices of the model into the memory of an inference processing unit, and performing hardware inference operations on data using the decompressed slices until a final result is obtained and output back to the CPU. Additional use case details

[0114] As described above, the disclosed techniques can be employed to derive context-specific machine learning models for a wide range of applications. In a code generation scenario, a user can input a description of a code function, such as a docstring, a comment, etc. The context-specific machine learning model can output code that performs the described function. In such a case, the base machine learning model can be a large model, such as Copilot that has been trained using docstrings or other functional descriptions of code for various programming languages. The context-specific training data for deriving a given context-specific model from the base model can include only code descriptions and corresponding code examples for a particular programming language.

[0115] Another example can be a natural language text generation scenario. Consider a BLOOM- or GPT-3-based model that answers user questions. A user may want to use the model to write a poem about some topic and then later ask detailed scientific questions. A context-specific machine learning model can be derived from BLOOM or GPT-3 using a context-specific training dataset of poem examples written by human users, accompanied by user descriptions of the poems. Another context-specific machine learning model can be derived from BLOOM or GPT-3 by using actual scientific questions and answers from scientists as the context-specific training dataset.

[0116] In the case of image processing, different labeled training datasets can be obtained. For example, a human user can label images of wild mammals with the correct species to obtain a first context-specific training dataset, and can label images of flowers with the correct species to obtain a second context-specific training dataset. A single large model such as ResNet-50 or NASNetLarge can be used as the base machine learning model, from which context-specific models for mammal recognition and context-specific models for flower recognition can be derived.

[0117] As another example, a text-to-image model such as Stable Diffusion can be used as the base machine learning model. Corresponding context-specific models can be derived from such a base machine learning model using text inputs and corresponding image pairs for different contexts. For example, a first context-specific training dataset can include images of artworks (e.g., paintings) and corresponding descriptions of the paintings. A second context-specific training dataset can include images of terrains (e.g., mountains, lakes, grasslands, forests) and corresponding descriptions of the terrains. Similar methods can be used for generating models for other types of media (such as audio and / or video). Technical Effects

[0118] As described above, modern inference hardware can greatly accelerate the efficiency of inference operations that can be performed on a given client device. However, many models are too large to be directly used on client devices. By starting with a large base machine learning model and deriving small context-specific models therefrom, it is feasible to implement inference processing on client hardware.

[0119] As described above, modern inference hardware architectures have limitations such as constrained memory sizes, or may only provide specific hardware instructions for implementing operations that tend to be done in neural networks, such as convolution operations or matrix operations. For example, an inference hardware architecture may provide instructions for performing convolution operations with specific input / output tensor and / or kernel sizes, vector operations or matrix operations with specific input or output tensor sizes, pooling operations, activation functions, etc. When developing a machine learning model using convolution operations or matrix operations supported by a given inference hardware architecture, the machine learning model can run very efficiently on a processing unit that supports that architecture.

[0120] By searching for models that satisfy the memory constraints of the inference hardware and / or include inference operations supported by the inference hardware architecture, new models that exhibit comparable accuracy to the base model in a particular context can be identified. Similarly, by pruning parameters from a large base machine learning model, context-specific machine learning models that satisfy the hardware constraints can be obtained.

[0121] Furthermore, by compressing a context-specific machine learning model into corresponding slices, it is reasonable to deliver the model on demand over the network when the context on a given client device changes. Thus, a given client device can swap a context-specific machine learning model in and out of memory as needed while consuming a reasonable amount of bandwidth to obtain the model over the network. Additionally, since the decompression layer is small enough to fit into the SRAM of the NPU, the decompression layer can be efficiently executed on the client device. Definitions

[0122] For the purposes of this document, the term "inference hardware architecture" refers to a set of operations provided by one or more inference processing units suitable for machine learning inference processing. For example, inference operations can be implemented in dedicated circuitry on a processing unit configured to use specific data sizes (e.g., input size, output size, kernel size, etc.). The term "inference operation" refers to an operation performed by a machine learning model to perform a task. For example, an inference operation can be performed by applying learned parameters obtained by training a machine learning model.

[0123] The term "base machine learning model" refers to a model that has been trained for a range of contexts. The term "context-specific machine learning model" refers to a model that is derived from a base machine learning model and adapted for a specific context. The term "context" refers to any type of use case for a machine learning model, such as a specific application scenario.

[0124] The term "learned parameters" refers to parameters learned by training a machine learning model (such as a neural network), such as edge weights and bias values. The term "operation" refers to a function that can be performed by one or more nodes. The term "model structure" refers to the overall architecture of the model, including the number of layers or nodes, the connectivity of the layers, and / or the types of operations performed by each layer. The term "neural network structure" refers to the model structure of a neural network. The term "trained model" refers to the model structure and the learned parameters of the model structure. Note that, for example, if the two models are trained on different training data or if there is an underlying stochastic process in the training process, then two trained models can share the same model structure and still have different learning parameters.

[0125] The term "parent model" refers to a model that is subsequently modified to obtain a "child model". A "seed model" is a type of parent model, e.g., a pre-existing model that is selected as a starting point for a search of a machine learning model search space. The term "final model" is used herein only to imply that a given model is designated for actual use in an application. In some cases, a final model output by a first search of a machine learning model search space can subsequently be used as a seed model to initiate a second search, thereby producing a second final model. Equipment Implementation

[0126] As mentioned above about Figure 9 As described, system 900 includes several devices, including client device 910, server 920, client device 930, and client device 940. As also mentioned above, according to the above and following descriptions, not all device implementations may be shown, and other device implementations should be obvious to those skilled in the art.

[0127] As used herein, the terms "device," "computer," "computing device," "client device," and / or "server device" may mean any type of device having a certain amount of hardware processing capability and / or hardware storage / memory capability. The processing capability may be provided by one or more hardware processors (e.g., hardware processing units / cores) that may execute data in the form of computer-readable instructions to provide functionality. The computer-readable instructions and / or data may be stored on a storage device, such as a storage device / memory and / or a data storage device. As used herein, the term "system" may refer to a single device, multiple devices, etc.

[0128] Storage resources can be internal or external to the respective devices with which they are associated. Storage resources can include any one or more of volatile memory or non-volatile memory, hard disk drives, flash devices, and / or optical storage devices (e.g., CDs, DVDs, etc.). As used herein, the term "computer-readable medium" can include signals. In contrast, the term "computer-readable storage medium" does not include signals. Computer-readable storage media include "computer-readable storage devices". Examples of computer-readable storage devices include volatile storage media such as RAM and non-volatile storage media such as hard disk drives, optical discs, and flash memory.

[0129] In some cases, a device is configured with a general-purpose hardware processor and storage resources. In other cases, a device can include a system-on-chip (SOC) type design. In an SOC design implementation, the functions provided by the device can be integrated on a single SOC or multiple coupled SOCs. One or more associated processors can be configured to coordinate with shared resources (such as memory, storage, etc.) and / or one or more dedicated resources (such as hardware blocks configured to perform certain specific functions). Thus, as used herein, the terms "processor", "hardware processor", or "hardware processing unit" can also refer to a central processing unit (CPU), a graphics processing unit (GPU), a controller, a microcontroller, a processor core, or other types of processing devices suitable for implementation in both conventional computing architectures and SOC designs.

[0130] Alternatively or additionally, the functions described herein can be performed at least in part by one or more hardware logic components. By way of example, and not limitation, illustrative types of hardware logic components that can be used include field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system-on-chip (SOC), complex programmable logic devices (CPLD), etc.

[0131] In some configurations, any module / code discussed herein can be implemented in software, hardware, and / or firmware. In any case, the module / code can be provided during the manufacture of the device, or by an intermediary of the device being prepared for sale to an end user. In other cases, the end user can install these modules / codes later, such as by downloading executable code and installing the executable code on the corresponding device.

[0132] Note also that a device can generally have input functionality and / or output functionality. For example, a computing device can have various input mechanisms such as a keyboard, a mouse, a touchpad, speech recognition, gesture recognition (e.g., using a depth camera such as a stereo or time-of-flight camera system, an infrared camera system, an RGB camera system or using an accelerometer / gyroscope, face recognition, etc.). The device can also have various output mechanisms such as a printer, a monitor, etc.

[0133] It should also be noted that the devices described herein can operate in an independent or collaborative manner to implement the described techniques. For example, the methods and functions described herein can be executed on a single computing device and / or distributed across multiple computing devices that communicate via a network 950. By way of non-limiting example, the network 950 can include one or more local area networks (LANs), wide area networks (WANs), the Internet, etc.

[0134] Various examples are described above. Additional examples are described below. One example includes a method that includes: obtaining a plurality of context-specific machine learning models, each context-specific machine learning model being derived from a base machine learning model adapted to a plurality of contexts and each context-specific machine learning model being adapted to a different context of the plurality of contexts; detecting a particular context of a particular device; selecting a particular context-specific machine learning model from the plurality of context-specific machine learning models based at least on the particular context of the particular device; and providing the particular context-specific machine learning model to the particular device.

[0135] Another example can include any of the above example and / or the following examples, wherein the method further includes: detecting that the particular device has switched to another context; and in response to detecting that the particular device has switched to another context: selecting another context-specific machine learning model from the plurality of context-specific machine learning models based at least on the other context; and providing the other context-specific machine learning model to the particular device.

[0136] Another example can include any of the above example and / or the following examples, wherein the method further includes: determining the particular context and the other context using an automatic context prediction algorithm based at least on context data received from the particular device.

[0137] Another example can include any of the above example and / or the following examples, wherein the base machine learning model is adapted to generate code in multiple programming languages, the particular context relates to a particular programming language, and the particular context-specific machine learning model is adapted to generate code in the particular programming language.

[0138] Another example may include any one of the above examples and / or the following examples, where the base machine learning model is adapted to identify multiple object types in an image, and a particular context-specific machine learning model is adapted to identify a subset of the multiple object types.

[0139] Another example may include any one of the above examples and / or the following examples, where the method further includes: compressing a particular context-specific machine learning model to obtain a compressed version, and sending the compressed version to a particular device via a network.

[0140] Another example may include any one of the above examples and / or the following examples, where the compressed version has corresponding slices corresponding to the respective layers of the particular context-specific machine learning model.

[0141] Another example is a method that includes: obtaining a base machine learning model adapted to multiple contexts; deriving multiple context-specific machine learning models adapted to different contexts among the multiple contexts from the base machine learning model; and outputting the multiple context-specific machine learning models for use in different contexts.

[0142] Another example may include any one of the above examples and / or the following examples, where the derivation includes: using the base machine learning model as a teacher and using the multiple context-specific machine learning models as students.

[0143] Another example may include any one of the above examples and / or the following examples, where the derivation includes: adjusting the parameters of a particular context-specific machine learning model so that the particular context-specific machine learning model is adapted to a particular context, and the adjustment is performed using a loss function based on the respective output distributions of the base machine learning model and the particular context-specific machine learning model when executed on particular context-specific training data for the particular context.

[0144] Another example may include any one of the above examples and / or the following examples, where the derivation includes: performing a search to identify an architecture shared by each of the multiple context-specific machine learning models.

[0145] Another example may include any one of the above examples and / or the following examples, where the search starts from a seed model architecture and iteratively selects new parent models from the Pareto frontier according to two or more criteria.

[0146] Another example may include any one of the above examples and / or the following examples, where the search is constrained based on hardware constraints for an inference processing unit.

[0147] Another example may include any of the above examples and / or the following examples, where the Pareto boundary includes a first criterion related to a loss function.

[0148] Another example may include any of the above examples and / or the following examples, where the derivation includes: pruning parameters from a base machine learning model.

[0149] Another example may include any of the above examples and / or the following examples, where the pruning is at least based on the magnitude or gradient of the parameters of the base machine learning model when training on special context-specific training data for a particular context.

[0150] Another example includes a computing device that includes: a hardware processing unit; and a storage resource storing computer-readable instructions that, when executed by the hardware processing unit, cause the hardware processing unit to: receive a special context-specific machine learning model adapted to a particular context, the special context-specific machine learning model being derived from a base machine learning model adapted to multiple contexts; and execute the special context-specific machine learning model on the computing device when the computing device is in the particular context.

[0151] Another example may include any of the above examples and / or the following examples, where the computer-readable instructions, when executed by the hardware processing unit, cause the hardware processing unit to: receive another context-specific machine learning model derived from the base machine learning model and adapted to another context; and execute the another context-specific machine learning model on the computing device when the computing device is in the another context.

[0152] Another example may include any of the above examples and / or the following examples, where the hardware processing unit includes a central processing unit, the computing device further includes an inference processing unit and an inference processing unit memory, and where the computer-readable instructions, when executed by the central processing unit, cause the central processing unit to: retrieve a compressed slice of the special context-specific machine learning model; decompress the slice; and load the decompressed slice into the inference processing unit memory for execution by the inference processing unit. Conclusion

[0153] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims need not be limited to the specific features or acts described above. Rather, the above specific features and acts are disclosed as example forms of implementing the claims, and other features and acts that would be recognized by those skilled in the art are intended to be within the scope of the claims.

Claims

1. A method, comprising: Obtaining a plurality of context-specific machine learning models, each context-specific machine learning model being derived from a base machine learning model adapted to a plurality of contexts, and each context-specific machine learning model being adapted to a different context among the plurality of contexts; Detecting a specific context of a specific device; Selecting a specific context-specific machine learning model from the plurality of context-specific machine learning models based at least on the specific context of the specific device; And Providing the specific context-specific machine learning model to the specific device.

2. The method according to claim 1, further comprising: Detecting that the specific device has switched to another context; And In response to detecting that the specific device has switched to the other context: Selecting another context-specific machine learning model from the plurality of context-specific machine learning models based at least on the other context; And Providing the other context-specific machine learning model to the specific device.

3. The method according to claim 2, further comprising: Determining the specific context and the other context using an automatic context prediction algorithm based at least on context data received from the specific device.

4. The method according to any one of claims 1 to 3, wherein the base machine learning model is adapted to generate code in multiple programming languages, the specific context relates to a specific programming language, and the specific context-specific machine learning model is adapted to generate code in the specific programming language.

5. The method according to any one of claims 1 to 3, wherein the base machine learning model is adapted to identify multiple object types in an image, and the specific context-specific machine learning model is adapted to identify a subset of the multiple object types.

6. The method according to claim 5, further comprising: Compressing the specific context-specific machine learning model to obtain a compressed version, and sending the compressed version to the specific device via a network.

7. According to the method of claim 6, the compressed version has corresponding slices corresponding to the respective layers of the specific context-specific machine learning model.

8. A method, comprising: Obtaining a base machine learning model adapted to a plurality of contexts; Deriving a plurality of context-specific machine learning models adapted to different contexts among the plurality of contexts from the base machine learning model; And Outputting the plurality of context-specific machine learning models for use in the different contexts.

9. The method according to claim 8, wherein the derivation comprises: Employing the base machine learning model as a teacher and the plurality of context-specific machine learning models as students.

10. The method according to claim 9, wherein the derivation comprises: Adjusting the parameters of a specific context-specific machine learning model to adapt the specific context-specific machine learning model to a specific context, The adjustment is performed using a loss function based on the output distributions of the base machine learning model and the context-specific machine learning model when executing on context-specific training data for the particular context.

11. The method according to claim 10, wherein the derivation comprises: Performing a search to identify an architecture shared by each of the context-specific machine learning models of the plurality of context-specific machine learning models.

12. The method according to claim 11, wherein the search starts from a seed model architecture and iteratively selects new parent models from the Pareto frontier according to two or more criteria.

13. The method according to claim 12, wherein the search is constrained based on hardware constraints for an inference processing unit.

14. The method according to any one of claims 12 or 13, wherein the Pareto frontier comprises a first criterion related to the loss function.

15. The method according to claim 14, wherein the derivation comprises: Pruning parameters from the base machine learning model.

16. The method according to claim 15, wherein the pruning is at least based on the magnitude or gradient of the parameters of the base machine learning model when training on context-specific training data for a particular context.

17. A computing device, comprising: A hardware processing unit; And A storage resource storing computer-readable instructions that, when executed by the hardware processing unit, cause the hardware processing unit to: Receive a context-specific machine learning model adapted to a particular context, the context-specific machine learning model being derived from a base machine learning model adapted to a plurality of contexts; And Execute the context-specific machine learning model on the computing device when the computing device is in the particular context.

18. The computing device according to claim 17, wherein the computer-readable instructions, when executed by the hardware processing unit, cause the hardware processing unit to: Receive another context-specific machine learning model derived from the base machine learning model and adapted to another context; and Execute the another context-specific machine learning model on the computing device when the computing device is in the another context.

19. The computing device according to claim 18, wherein the hardware processing unit comprises a central processing unit, and the computing device further comprises an inference processing unit and an inference processing unit memory, wherein the computer-readable instructions, when executed by the central processing unit, cause the central processing unit to: Retrieve a compressed slice of the context-specific machine learning model; Decompress the slice; and Load the decompressed slice into the inference processing unit memory for execution by the inference processing unit.

20. The computing device according to claim 19, wherein the compressed slice comprises parameters of respective layers of the context-specific machine learning model.