System and method for providing task and hardware architecture-specific machine learning models

The method of training a stacking model with knowledge distillation and neural architecture search addresses the inefficiencies of large-scale models by adapting them to specific tasks and hardware architectures, reducing computational and memory costs for efficient deployment.

CN120317397APending Publication Date: 2025-07-15ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510040638.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-12
Filing Date
2025-01-10
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing basic machine learning model is difficult to deploy efficiently in practical applications due to its many parameters and high computational storage costs. The existing knowledge transfer methods need to be recalculated for each target hardware architecture, resulting in high costs.

Method used

By training the superimposed model and fine-tuning it in a hardware architecture agnostic way, combining knowledge transfer and model architecture search, a task-specific machine learning model is generated, and a variety of alternative computing and knowledge distillation techniques of the superimposed model are used to reduce computing and storage requirements.

Benefits of technology

It realizes efficient deployment of machine learning models on different hardware architectures, reducing compute and storage costs, and improving model adaptability and deployment efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120317397A_ABST
    Figure CN120317397A_ABST
Patent Text Reader

Abstract

A system and method 300 for providing a task and hardware architecture-specific machine learning model is provided. A trained overlay model may be provided 320, which may include an overlay of a set of machine learning models, each machine learning model of the set of machine learning models may be extracted from the trained overlay model. A representation of a target hardware architecture may be received. The trained overlay model may be trimmed 330 for the application task in a hardware architecture agnostic manner. A machine learning model may be selected from the trimmed overlay model, where the selection may include for a target hardware architecture, the search is performed using a first function describing a first performance of a candidate machine learning model for the application task and a second function describing a second performance of the candidate machine learning model when executed on the target hardware architecture. The selected machine learning model may be provided as output for deployment on the target hardware architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The presently disclosed subject matter relates to a system and method for providing task-specific and hardware architecture-specific machine learning models. The presently disclosed subject matter also relates to a transient or non-transient computer-readable medium comprising data representing instructions that, when executed by a processor system, cause the processor system to perform one or more steps of the method. Background Art

[0002] A foundation model (FM) is a type of machine learning (ML) model that is relatively large in scale and is typically trained on large amounts of data via self-supervised learning or semi-supervised learning such that the model can be adapted for application to a wide range of application tasks or downstream tasks. In this sense, foundation models are quite general. This type of model has contributed to a major shift in how systems in artificial intelligence (AI) are built and how new technologies have been able to evolve, such as: language models, e.g., the Bidirectional Encoder Representations from Transformers (BERT) model from Google; foundation models in Generative Pretrained Transformer (GPT) models, e.g., the GPT-n models from OpenAI; chatbot technologies, such as Chat-GPT; and other user-interactive AI systems.

[0003] However, large-scale foundation models are not hardware-efficient. They suffer from a large number of parameters and very significant computational and memory costs, which makes them difficult to deploy in real-world settings.

[0004] Ways have been tried in the past to overcome this problem. To this end, several knowledge transfer methods have been introduced, by means of which knowledge can be transferred from a larger teacher network to a smaller student network. In this way, knowledge can be transferred from a foundation model to a smaller model that can have a smaller number of parameters and fewer computational and memory costs.

[0005] For example, in the paper "NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture Search" by Xu et al. that can be retrieved from https: / / arxiv.org / abs / 2105.14444, a task-agnostic knowledge transfer method using techniques in neural architecture search (NAS) is performed on the BERT model from Google to perform model compression, resulting in NAS-BERT. The compression algorithm employed is downstream task-agnostic, making the compressed model generally applicable to different downstream tasks. Then, NAS-BERT trains a large supernet on a design search space containing various architectures. During training, NAS-BERT gradually discards architectures based on its training loss, size, and latency on the target hardware. Finally, NAS-BERT outputs multiple models with different sizes and latencies to support different memory and latency constraints on the target device. Additionally, the training of NAS-BERT is conducted on standard self-supervised pre-training tasks and does not depend on specific downstream tasks. In this way, the compressed model can be used across various downstream tasks.

[0006] A drawback of NAS-BERT is that each step has to be re-run for each new target hardware architecture, which results in significant computational costs. Summary of the Invention

[0007] It would be desirable to be able to provide task- and hardware-architecture-specific machine learning models in a computationally more efficient manner.

[0008] According to a first aspect of the present invention, as defined in claim 1, there is provided a method for providing a task- and hardware-architecture-specific machine learning model. According to a further aspect of the present invention, there is provided a system as defined in claim 14. According to a further aspect of the present invention, there is provided a computer-readable medium as defined by claim 15.

[0009] The above measures may involve providing a trained superposition model. Such a superposition model (sometimes referred to as a one-shot model) includes a number of models stacked into one model, as detailed, for example, in Cai et al.'s "Once-for-All: Train One Network and Specialize It for Efficient Deployment", which can be retrieved from https: / / arxiv.org / pdf / 1908.09791. The trained superposition model may have been pre-trained using a general dataset and thus in a task-agnostic manner. In some examples, the trained superposition model may have been used as a student network in a knowledge transfer method to transfer knowledge from a base model to the superposition model. The superposition model may also have been trained from scratch, for example, if sufficient data and computing resources are available. The base model and / or the superposition model may include a neural network.

[0010] The above measures may also involve receiving a characterization of a target hardware architecture. Such a characterization of the target hardware architecture can describe aspects of the hardware architecture that can be unique in terms of computational efficiency or otherwise be characteristics of the target hardware architecture: for example, the number of floating-point operations (FLOP) that can be performed in one time unit, the number of multiply-accumulate (MAC) operations, the latency for executing a certain neural network, etc.

[0011] The above measures may also involve fine-tuning the trained superposition model in a hardware-architecture-agnostic manner for an application task. Although the superposition model may have been trained using a general dataset, here, the fine-tuning can include using a labeled dataset specific to the application task. For example, the labeled dataset can include a subset of the general dataset, or it can include the entire general dataset. The labeled dataset can include, for example, sensor signals as data, such as image data, radar data, and lidar data. The application task can include, for example, classification tasks (such as image perception) and / or regression tasks. The fine-tuning can be carried out, for example, using the sandwich rule and / or in-situ distillation. The fine-tuning can be hardware-architecture-agnostic because the characterization of the target hardware architecture may not be used in the fine-tuning, for example, to manipulate or otherwise control the fine-tuning.

[0012] The above measures may also involve selecting a machine learning model from the fine-tuned superposition model. The selection of the machine learning model can include performing a search for the target hardware architecture using a first function that describes the first performance of a candidate machine learning model for the application task and a second function that describes the second performance of the candidate machine learning model when executed on the target hardware architecture. This search can also be referred to as "model architecture search" because different candidate machine learning models can represent different model architectures. These measures can be further explained as follows.

[0013] The stacked model can be configured to provide alternative operations on data tensors. By selecting one operation or a subset of operations from the alternative operations for each data tensor, individual machine learning models can be extracted from the trained stacked model. For example, a neural network can include a sequence of operations (such as convolutions, max pooling, and / or activation functions) on data tensors. In a stacked model, several alternative operations can be applied to the same data tensor instead of one. Several alternative operations can together form a supernet, as in a hypernetwork. Such a supernet itself is known from, for example, see Figure 1 b in “DARTS: Differentiable Architecture Search” by Liu et al., which can be retrieved from https: / / arxiv.org / pdf / 1806.09055.pdf. Then, the stacked model can allow the extraction of machine learning models by selecting one operation or a subset of operations (and in some cases, not selecting any operation) from each set of alternative operations. For example, see the above-mentioned paper by Liu et al., where Figure 1 the model architectures shown in Figure 1 b form a subset of the supernet shown in

[0014] The above measures may also involve providing a selected machine learning model as output for deployment on a target hardware architecture. Since the machine learning model may already have been selected from an ensemble model after being fine-tuned for an application task and using a search based on both the application performance and the computing performance on the target hardware architecture, the resulting selected machine learning model can be considered a task-specific and hardware-architecture-specific machine learning model.

[0015] The above measures can be based on the following insight: When seeking to deploy a base model or a similar model that may have been trained on general training data in a real-world application, the application task is usually given (e.g., pedestrian detection), while the hardware architecture may change over time, for example, due to technological improvements in the computing architecture, or may vary at a given time due to different deployment targets (e.g., different products), etc. In other words, while the hardware architecture may often change, a given application task is unlikely to change over time, for example. Given this, the inventors have considered performing fine-tuning in a task-specific but hardware-architecture-agnostic manner, which is usually still relatively computationally intensive. As such, the fine-tuned ensemble model has not been fixed to a specific target hardware architecture but has only been fine-tuned for the application task at hand. By being an ensemble model, there are different machine learning models that can be extracted from the fine-tuned ensemble model, and these different machine learning models are each fine-tuned for the application task. Then, taking into account both the application performance and the computing performance on the target hardware architecture, adaptation to the target hardware architecture can be achieved by selecting from different machine learning models. When considering a new target hardware architecture, for example, due to the above-mentioned improvements in the computing architecture or different deployment targets, it may only be necessary to perform the search again, which is usually less computationally complex than fine-tuning.

[0016] In one embodiment, the search can be initialized based on the results of past searches. Given that the differences between the current and past searches in terms of, for example, the hardware architecture are likely to be relatively small, the solutions from past searches are likely to be candidates for the solutions of the current search. By using the past search as the initialization of the current search, the search efficiency of the current search can thus be improved.

[0017] In one embodiment, the search includes using the results of past searches to initialize an evolutionary search process. The results of past searches can be used as one or more seed models for the evolutionary search process. Generally speaking, an evolutionary search that depends on time can benefit from the initialization based on past searches at the initial time.

[0018] In one embodiment, the search includes using the results of past searches to initialize the distribution optimization over the model architecture. The results of past searches can include an optimized distribution over the model architecture, which can be used as a starting point for the distribution optimization over the model architecture in the search. The optimized distribution over the model architecture can be used as a starting point in one or more search methods that optimize the distribution over the model architecture in the stacked model. Given that the difference between the current and past distribution optimizations is likely to be relatively small in general, the optimized distribution as a result from the past search is typically used as a likely candidate for the distribution over the model architecture in the stacked model, which generally improves the search efficiency of the current distribution optimization obtained over the model architecture.

[0019] In one embodiment, the past search includes a task-agnostic and hardware-architecture-agnostic search, where the past search is applied to the trained stacked model using a task-independent function and a hardware-architecture-independent function. The task-independent function estimates the performance of the candidate machine learning models of the trained stacked model, and the hardware-architecture-independent function estimates the computational efficiency of the candidate machine learning models of the trained stacked model. The past search can be task-agnostic and / or hardware-agnostic, which means that no specific application task or target hardware characteristics are used in the search, and the search does not depend on it in any controlled way. Such a past model architecture can be generally applicable (e.g., for many application tasks) and can therefore be used as a generally possible candidate for the solution of the current search. Thus, this past search can be used to initialize all kinds of task-specific and / or hardware-architecture-specific searches. Using such a past, task- and hardware-architecture-agnostic search (the results of which are transferred to task- and hardware-specific searches) can accelerate task- and hardware-specific searches. To be able to achieve optimal or at least reasonable past search results, a task-independent function can be provided to characterize the performance of candidate model architectures on non-specific, "general" tasks. Thus, the task-independent function can be a function that describes how a candidate machine learning model will perform without considering a specific application task. The task-independent function can characterize the performance based on one or more generally applicable metrics related to model performance, such as the distillation loss used when training the stacked model, and the number of parameters of the candidate machine learning model, e.g., as a proxy for the representational ability of the candidate machine learning model. In addition to the task-independent function, a hardware-architecture-independent function can be used to characterize the performance of candidate model architectures on non-dedicated, "general" hardware architectures. The hardware-architecture-independent function can be a function that describes the computational efficiency of a candidate machine learning model independent of any specific hardware architecture. The hardware-architecture-independent function can characterize the computational efficiency of the candidate machine learning model based on one or more generally applicable metrics related to computational efficiency, such as the number of floating-point operations (FLOP), the number of multiply-accumulate (MAC) operations, the number of parameters of the candidate machine learning model (e.g., as a proxy for the memory traffic of the candidate machine learning model), and the latency on default hardware. The past search can be configured, for example, to search for a Pareto-optimal trade-off between a first task-independent function and a second hardware-architecture-independent function. This can then result in a Pareto-optimal model architecture.

[0020] In one embodiment, the method further includes using the selected machine learning model to initialize a subsequent search for another machine learning model for another application task and / or another target hardware architecture. Given that the differences between the subsequent search and the current search may generally be relatively small, for example, in terms of the hardware architecture, the results from the current model architecture can be used to initialize the subsequent search because it is likely to be a relatively suitable starting point for the subsequent search. The initialization of the subsequent search based on the current search can thus improve the search efficiency of the subsequent search. Although the application task in the subsequent search may be different from the application task in the current search, the target hardware architecture in the subsequent search can remain the same or similar to the target hardware architecture in the current search. Vice versa, although the target hardware architecture in the subsequent search may be different from the target hardware architecture in the current search, the application task in the subsequent search may remain the same or similar to the application task in the current search. Additionally, even if both the application task and the target hardware architecture change, these changes may be small enough such that the results from the current search can be used as a possible starting point for the subsequent search. In this way, the search can be iteratively performed in a computationally efficient manner.

[0021] In one embodiment, the method further includes training an ensemble model by using the ensemble model as the student model and using a base model as the teacher model to provide a trained ensemble model, the training including using a knowledge transfer method to transfer knowledge from the teacher model to the student model. In this way, valuable knowledge can be transferred from the base model to the smaller ensemble model, which has fewer computational and memory costs and potentially less training data compared to when the ensemble model has to be directly trained on extensive training data. The resulting trained ensemble model can then be generally applicable to, for example, many application tasks. The knowledge transfer method can include the use of knowledge distillation (KD) (such as task-agnostic knowledge distillation). As discussed by Hinton et al. in “Distilling the Knowledge in a Neural Network” (which can be retrieved from https: / / arxiv.org / abs / 1503.02531), knowledge distillation is a means of transferring knowledge from a large teacher network to a hardware-efficient student network, thereby providing a computationally and resource-efficient model from a large (pre)-trained model. KD can be employed before fine-tuning occurs; this enables the same student network to be reused for several application tasks, which results in the student network being widely applicable.

[0022] In one embodiment, the knowledge transfer method further includes using neural architecture search (NAS) to automate the design of the architecture of the neural network used in the knowledge transfer method, thereby avoiding the manual design of these architectures and accelerating the process of the knowledge transfer method.

[0023] In one embodiment, the method further includes fine-tuning the trained stacked model for additional application tasks to provide an additional fine-tuned stacked model, and selecting an additional machine learning model from the additional fine-tuned stacked model. In this way, multiple task-specific models can be extracted via fine-tuning from the trained, task-agnostic stacked model. In other words, the trained, task-agnostic stacked model can then serve as a basis for multiple task-specific stacked models, which in turn can each serve as a basis for multiple hardware-architecture-specific models. If desired, additional task-specific stacked models can be obtained only in the presence of additional application tasks. These multiple fine-tuning instances for application tasks can be performed simultaneously or sequentially with respect to the application tasks.

[0024] In one embodiment, the method further includes providing an additional task- and hardware-architecture-specific machine learning model for an additional target hardware architecture by again selecting a machine learning model from the fine-tuned stacked model. This selection of a machine learning model from the fine-tuned stacked model can be performed in a manner similar to that explained above, which involves a search. In doing so, additional hardware-specific models for the additional target hardware architecture can be derived from the fine-tuned task-specific stacked model. In this way, multiple hardware-architecture-specific models can be extracted via an additional search from the fine-tuned, task-specific stacked model (which is hardware-architecture-agnostic), and no additional fine-tuning process is required. Optionally, the results from the search for the target hardware architecture can be reused as an initialization for an additional search for an additional target hardware architecture. In the case where the target hardware architecture and the additional target hardware architecture relate to the same application task, this result from the model architecture may be a possible candidate for use as an initialization for the additional search. Note that the searches for different target hardware architectures can be performed simultaneously or sequentially.

[0025] In one embodiment, the stacked model is configured to provide alternative operations for data tensors, where individual machine learning models can be extracted from the trained stacked model by selecting one operation or a subset of operations from the alternative operations for each data tensor.

[0026] In another aspect of the present invention, a system is provided that includes one or more processors; and one or more storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operations of the method according to the embodiments discussed above.

[0027] In a further aspect of the present invention, there is provided a transient or non-transient computer-readable medium comprising data representing instructions which, when executed by a processor system, cause the processor system to perform one or more steps of the method according to the embodiments discussed above.

[0028] Those skilled in the art will appreciate that two or more of the above-described embodiments, implementations, and / or alternative aspects of the present invention may be combined in any useful manner.

[0029] Based on this specification, those skilled in the art can make modifications and variations to any device, system, network, computer-implemented method, and / or any computer-readable medium, which correspond to the modifications and variations of another such entity described. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Additional details, aspects, and embodiments will be described by way of example only with reference to the accompanying drawings. The elements in the figures are illustrated for simplicity and clarity and are not necessarily drawn to scale. In the figures, elements corresponding to those already described may have the same reference numerals. In the drawings,

[0031] Figure 1 and Figure 2 illustrates an example of a method for training and fine-tuning a base model;

[0032] Figure 3 illustrates a method for providing a task-specific machine learning model according to an embodiment;

[0033] Figures 4 - 6 illustrates a method for providing a task-specific and hardware-architecture-specific machine learning model according to an embodiment;

[0034] Figure 7a illustrates a computer-readable medium having a writable portion including a computer program; and

[0035] Figure 7b illustrates a representation of a processor system according to an embodiment.

[0036] LIST OF REFERENCE NUMERALS

[0037] The following list of reference numerals and abbreviations is provided for ease of interpretation of the drawings and should not be construed as limiting the claims.

[0038] 100 Method for training and fine-tuning a base model

[0039] 110 Training of the base model

[0040] 120A Perform KD / NAS for target hardware architecture A

[0041] 120B performs KD / NAS for target hardware architecture B

[0042] 130A, 130B task-agnostic search

[0043] 140A1, 140B1 are fine-tuned on application task 1

[0044] 140A2, 140B2 are fine-tuned on application task 2

[0045] 200 Method for training and fine-tuning a base model

[0046] 210 Training of the base model

[0047] 220 performs KD / NAS for a single-sample model

[0048] 230A performs task-agnostic search for target hardware architecture A

[0049] 230B performs task-agnostic search for target hardware architecture B

[0050] 240A1, 240B1 are fine-tuned on application task 1

[0051] 240A2, 240B2 are fine-tuned on application task 2

[0052] 300 Method for providing a machine learning model specific to a task and a hardware architecture

[0053] 310 Optionally train the base model

[0054] 320 Provide a trained stacked model

[0055] 330 Fine-tune the trained stacked model for an application task in a hardware-agnostic manner

[0056] 340 Search for a hardware architecture

[0057] 400 Method for providing a machine learning model specific to a task and a hardware architecture

[0058] 330.1 Fine-tune the trained stacked model 330.2 for application task 1 in a hardware-agnostic manner Fine-tune the trained stacked model for application task 2 in a hardware-agnostic manner 340.1, 340.2 Search for a hardware architecture

[0059] 500 Method for providing a machine learning model specific to a task and a hardware architecture

[0060] 340A Search for target hardware architecture A

[0061] 340B Search for target hardware architecture B

[0062] 350 Past task and hardware architecture agnostic search

[0063] 600 Method for providing a machine learning model specific to a task and a hardware architecture

[0064] 340.1A, 340.2A Search for target hardware architecture A

[0065] 340.1B, 340.2B Search for target hardware architecture B

[0066] 1000 Optical storage device

[0067] 1001 Memory card

[0068] 1020, 1021 Stored data

[0069] 1140 Processor system

[0070] 1110 Subsystem or component

[0071] 1120 Processing subsystem

[0072] 1122 Memory

[0073] 1124 Application-specific integrated circuit

[0074] 1126 Communication interface

[0075] 1130 Interconnection Detailed implementation manner

[0076] Although the presently disclosed subject matter admits of embodiments in many different forms, one or more specific embodiments are shown in the drawings and will be described in detail herein. It is to be understood that the present disclosure is to be considered as illustrative of the principles of the presently disclosed subject matter and is not intended to limit it to the specific embodiments shown and described.

[0077] Hereinafter, for the sake of understanding, the elements of the embodiments are described in operation. However, it will be clear that the various elements are arranged to perform the functions described as being performed by them.

[0078] Furthermore, the presently disclosed subject matter is not limited to the embodiments, but also includes every other combination of features described herein or recited in mutually different dependent claims.

[0079] Figure 1 An example of a method 100 for training and fine-tuning a base model is shown, which is similar in structure to NAS-BERT.

[0080] First, a base model is trained in step 110. In steps 120A and 120B, KD / NAS is performed on two target hardware architectures, which are identified by the characters "A" and "B" respectively. Hereinafter, steps belonging to the corresponding target hardware architecture can be identified by the corresponding character as a suffix. For example, step 120A may include performing KD / NAS on target hardware architecture A. Continuing with reference to KD / NAS, note that the knowledge distillation part is used as a means to transfer knowledge from a large teacher network (which is the base model in this example) to a one-shot model using neural architecture search (NAS). The one-shot model constitutes the student model. Thus, in steps 120A and 120B, task-agnostic but hardware-specific KD / NAS is performed for each target hardware architecture.

[0081] In steps 130A and 130B, task-agnostic search is performed for each target hardware architecture A and B respectively. Note that this task-agnostic search cannot optimize the architecture for the application task; the model architecture must be fine-tuned on the application task.

[0082] In steps 140A1-2 and 140B1-2, the obtained model architectures are fine-tuned on the application task. Here, two current application tasks are identified by the numbers "1" and "2". The aggregation of the foregoing steps of method 100 can be represented as NAS-BERT, which in turn can be regarded as a model compression version of BERT, where the compression algorithm is task-agnostic.

[0083] Method 100 has several drawbacks. First, computationally expensive KD / NAS - which requires hardware metrics such as the latency of candidate model architectures - needs to be run for each target hardware architecture. Second, the task-agnostic search cannot optimize the architecture for the application task.

[0084] Figure 2 An example of method 200 for training and fine-tuning a base model is shown, which is similar in structure to a method called AutoDistil, which is a known technique. AutoDistil represents another attempt to address the computational inefficiency problem of the base model. AutoDistil is discussed in the paper "AutoDistil: Few-shot Task-agnostic Neural Architecture Search for Distilling Large Language Models" by Xu et al., which can be retrieved from https: / / arxiv.org / pdf / 2201.12507, which is an example. In this paper, KD / NAS is used to compress a large model into a smaller student network, and several compressed student networks with varying computational costs are automatically distilled from the large model.

[0085] In Figure 2 it, an overview of a variant of AutoDistil is presented. First, a base model is trained in step 210. In step 220, KD / NAS is performed, which distills a single one-shot model from the base model. In steps 230A and 230B, task-agnostic search is performed for each target hardware architecture A and B, respectively. In steps 240A1-2 and 240B1-2, the obtained model architectures are fine-tuned on the application tasks. Two current application tasks are identified by the numbers "1" and "2".

[0086] This variant of AutoDistil 200 has several drawbacks. First, it cannot optimize the model architecture for the application tasks, thus wasting the optimization potential. Second, for each target hardware architecture, the search has to start from scratch, which results in significantly higher computational costs, especially in a larger search space.

[0087] Figure 3 An example of a method for providing a task- and hardware-architecture-specific machine learning model is shown.

[0088] In optional step 310, a base model can be trained. If sufficient data and computational resources are available, this training step of the model can be performed on large data. In most cases, it can be assumed that a (pre)-trained base model is available. An example of a pre-trained base model is described in the paper "EVA: Exploring the Limits of Masked Visual Representation Learning at Scale" by Fang et al., which can be retrieved from https: / / arxiv.org / abs / 2211.07636.

[0089] In step 320, a trained stacked model can be obtained. For example, during or before step 320, the stacked model may have been trained as a student model in a knowledge transfer method. In a specific example, the trained base model from step 310 can be the teacher model in such a knowledge transfer method. The knowledge transfer method employed can be knowledge distillation. The knowledge distillation can be task-agnostic knowledge distillation. Neural architecture search can be used, where the student network has a search space from this neural architecture search, thus constructing a KD / NAS method. In such a KD / NAS, the student network may be the aforementioned stacked model. The stacked model can represent a stack of many different neural networks. The stacked model can, for example, include different sequences of operations applied to a data tensor. Such operations can include, for example, convolution, max pooling, activation functions, etc. The stacked model can include alternative operations applied to the same tensor. For example, instead of only a single first operation, the stacked model can include several alternative first operations, such as a 3×3 convolution with 16 channels, a 5×5 convolution with 16 channels, and / or a 3×3 convolution with 32 channels. The different neural networks extractable from the stacked model can provide different trade-offs between computational resource requirements and achievable accuracy. Alternatively, the training of the stacked model may have been done "from scratch". In particular, if sufficient data and computational resources are available, the stacked model can be trained from scratch.

[0090] In step 330, the trained stacked model can be fine-tuned for an application task. For the fine-tuning 330 of the application task, a labeled dataset specific to the application task can be used as the input to the model 300. This labeled dataset specific to the application task can be relatively smaller than the general dataset on which the trained stacked model has been trained in step 310 and / or on which the base model has been trained in step 320. The fine-tuning of the trained stacked model for the application task using the labeled dataset specific to the application task can be performed in a task-specific and hardware-architecture-agnostic manner. For example, this can be done using the sandwich rule and / or in-situ distillation techniques, as in the paper by Yu et al., "BigNAS: Scaling Up Neural Architecture Search with Big Single-Stage Models", which can be retrieved from https: / / arxiv.org / abs / 2003.11142.

[0091] In step 340, a machine learning model can be selected from the fine-tuned stacked model. This selection can include performing a search. The search can be performed on a hardware architecture that has received a representation of the target hardware architecture as input. The receipt of the representation of the target hardware architecture can occur in an earlier step (e.g., one of steps 310 - 330). The search can use two functions that may have been received as input. The first function can be a function that describes the performance of candidate machine learning models of the fine-tuned stacked model for an application task. This function can be determined specifically for the application task at hand. For example, the metrics used in the function can be specifically selected to be able to quantify the performance of the application task at hand. The first function can, for example, include metrics such as classification accuracy, average precision, perplexity, etc., or a combination of several such metrics. The second function can be a function that describes the performance of candidate machine learning models of the fine-tuned stacked model when executed on the target hardware architecture. For this purpose, the second function can describe hardware metrics that can specifically characterize the performance on the target hardware architecture. For example, the second function can, for example, include metrics such as the latency, energy consumption, etc. of the hardware architecture measured on the target hardware, approximations of such metrics, or a combination of several such metrics. Then, these two functions can be used to perform a search on the fine-tuned stacked model, and the model architecture that best satisfies the combination of such functions can be returned. In a specific example, the search can be configured to search for a Pareto-optimal trade-off between the first function and the second function. This can result in a Pareto-optimal model architecture for the application task and the target hardware architecture, which can be returned by the search.

[0092] Compared with Figure 1 method 100, the step of providing the trained stacked model in step 320 of method 300 (which can be done, for example, by KD / NAS independent of the hardware architecture) can enable the computationally expensive steps of KD / NAS to be performed only once for various target hardware architectures and application tasks, rather than the requirement of method 100 for independent KD / NAS steps for each target hardware architecture.

[0093] Figure 4 Another example of a method 400 for providing a task- and hardware-architecture-specific machine learning model is shown.

[0094] In addition to method 300, method 400 may further include fine-tuning of the trained stacked model steps for multiple application tasks identified by the numbers "1" and "2" in steps 330.1 and 330.2, respectively. In steps 340.1 and 340.2, a machine learning model may be selected from the fine-tuned stacked model. Each selection of the machine learning model may include performing a search. The search may be performed for different target hardware architectures. The fine-tuning 330.1, 330.2 for application tasks 1 and 2 may be performed simultaneously with respect to application tasks 1 and 2, and / or the selection steps 340.1, 340.2 of the machine learning model for different target hardware architectures may be performed simultaneously with respect to the target hardware architectures. The fine-tuning steps 330.1, 330.2 may also occur sequentially (e.g., in a consecutive manner) with respect to application tasks 1 and 2. Similarly, the selection steps 340.1, 340.2 may also occur sequentially with respect to the target hardware architectures.

[0095] Figure 5 Another example of a method 500 for providing a machine learning model specific to a task and a hardware architecture is shown.

[0096] In addition to method 300, method 500 may further include searching for a fine-tuned stacked model for multiple target hardware architectures identified by the characters "A" and "B" in steps 340A and 340B, respectively. The searches 340A, 340B for target hardware architectures A and B may be performed simultaneously with respect to target hardware architectures A and B. The searches 340A, 340B may also occur sequentially (e.g., consecutively) with respect to target hardware architectures A and B.

[0097] In addition to method 300, method 500 may further include an additional step 350, which may include a past search. The past search in step 350 may be a search that is agnostic to the task and the hardware architecture, such that the results may be used for the search for both target hardware architectures A and B. The past model architecture may utilize a first function and a second function. The first function may be a task-independent function that may estimate the performance of a candidate machine learning model of the trained stacked model. The first function may be based on one or more generally applicable metrics related to model performance, such as the distillation loss used when training the trained stacked model, and the number of parameters of the candidate machine learning model (e.g., as a proxy for the representational power of the model architecture). The second function may be a hardware-independent function that may characterize the computational efficiency of the candidate model architecture. The second function may be based on one or more generally applicable metrics related to computational efficiency, such as the number of floating point operations (FLOP), the number of multiply-accumulate operations (MAC), the number of parameters of the candidate machine learning model (e.g., as a proxy for the memory traffic of the model architecture), and / or the latency on the default hardware. During the past search, the first and second functions may be used to perform a task- and hardware-agnostic evolutionary architecture search on the trained stacked model. In some embodiments, the past search may be configured to search for a Pareto-optimal trade-off between the first and second functions, and may thereby return the resulting Pareto-optimal model architecture.

[0098] The searches for target hardware architectures A and B in steps 340A and 340B may be initialized based on the results of the past search. For example, the search may be performed as an evolutionary search by using the results of the past model architecture as the seed model for the task- and hardware-specific evolutionary searches 340A, 340B. The searches in steps 340A and 340B may also be performed, for example, as a distribution optimization over the model architecture; the model architecture may use the optimized distribution as a starting point, where the optimized distribution may be obtained from the past search that may occur in step 350.

[0099] In some embodiments, the results of the past search in step 350 may include a set of Pareto-optimal model architectures with respect to the first and second functions, and may be used to initialize the task- and hardware-specific searches in steps 340A and 340B, where the set of Pareto-optimal model architectures may be used as the initial population.

[0100] Figure 6 An example of a method 600 for providing a task- and hardware-architecture-specific machine learning model is shown. Method 600 may combine features from models 400 and 500.

[0101] In addition to method 300, method 600 may further include fine-tuning the trained stacked model in steps 330.1 and 330.2 respectively, where the trained stacked model is the output of step 320 for multiple application tasks identified herein by the numbers "1" and "2". In steps 340.1A, 340.1B, 340.2A, and 340.2B, a search may be performed on multiple target hardware architectures, which are identified by the characters "A" and "B". The fine-tuning 330.1, 330.2 for application tasks 1 and 2 may be performed simultaneously with respect to application tasks 1 and 2, and / or the search 340.1A, 340.1B, 340.2A, and 340.2B for the target hardware architectures may be performed simultaneously with respect to hardware architectures A and B. The search 340.1A, 340.1B, 340.2A, 340.2B may also occur sequentially (e.g., in a consecutive manner) with respect to the target hardware architectures A and B. Additionally, the results of past searches in step 350 may initialize the searches in all of steps 340.1A, 340.1B, 340.2A, and 340.2B.

[0102] Any of the (one or more) methods described in this specification may be implemented on a computer as a computer-implemented method, as dedicated hardware, or as a combination of both. Also as Figure 7a illustrated, instructions (e.g., executable code) for a computer may be stored on computer-readable media 1000, 1001, for example, in the form of a series of machine-readable physical markings 1020, 1021 and / or as a series of elements having different electrical (e.g., magnetic) or optical properties or values. The computer-readable media 1000, 1001 may be a transient or non-transient medium. Examples of computer-readable media include memory devices, optical storage devices, integrated circuits, etc. By way of example, Figure 7a an optical storage device 1000 and a memory card 1001 are shown.

[0103] Figure 7bFIG. 1140 shows a processor system 1140 that may include or represent a system configured to perform methods as described elsewhere in this specification. The processor system may include one or more subsystems or components 1110. For example, a processing subsystem 1120 may be provided to execute computer program components to perform methods as described elsewhere in this specification. A memory 1122 may be provided to store programming code, data, etc. A communication subsystem 1126, such as a network interface, may allow communication with other entities. In some examples, an application specific integrated circuit 1124 may be provided to perform some or all of the processing related to methods as described elsewhere in this specification. The processing subsystem 1120, the memory 1122, the application specific IC 1124, and the communication subsystem 1126 may be interconnected via an interconnect 1130 (such as a bus). Although system 1140 is shown as including one of each described component, in various embodiments, the various components may be replicated. For example, the processing subsystem 1120 may include multiple microprocessors that are configured to independently perform methods as described in this specification, or are configured to perform steps or subroutines of methods described herein such that the multiple processors cooperate to implement the functions described in this specification. Additionally, in cases where system 1140 may be implemented in a cloud computing system, cloud server, and / or computing farm, the various hardware components may belong to separate physical systems. For example, the processing subsystem 1120 may include a first processor in a first server and a second processor in a second server.

[0104] In Figure 7bIn an alternative embodiment, the processor system 1140 may represent the target hardware architecture on which a selected machine learning model is deployed. In other words, the processor system may represent the deployment target, which may perform application tasks as described elsewhere in this specification. The processor system 1140 may be, for example, a device or an apparatus. The device or apparatus may include, for example, a sensor that may determine a measurement of the environment in the form of a sensor signal, which may be given by, for example, a digital image, such as a video, radar, lidar, ultrasound, thermal image of movement, or an audio signal. The device or apparatus may include, for example, a household appliance that includes a sensor for detecting the presence of an object in a washing machine, or a vehicle that includes a sensor for detecting the presence of an object in the vehicle environment. Application tasks may include classifying data from sensors, detecting the presence of objects in sensor data, and / or performing semantic segmentation on the data (e.g., with respect to traffic signs, road surfaces, pedestrians, and vehicles). Another application task may include determining one or more continuous values, such as performing regression analysis, e.g., with respect to distance, speed, acceleration, and / or tracking of items (e.g., objects) in the data. These examples of application tasks may be implemented on low-level features (such as edge or pixel nature in the case of image data). Other application tasks may include detecting anomalies in a technical system, calculating control signals for controlling a technical system, such as a computer-controlled machine, vehicle, household appliance (such as a washing machine), power tool, manufacturing machine, personal assistant, or access control system acting as a robotic system; or a system for transmitting information, such as a surveillance system or a medical system acting as a medical imaging system.

[0105] Examples, embodiments, or optional features — whether or not indicated as non-limiting — should not be construed as limiting the claimed invention.

[0106] It should be noted that the above embodiments illustrate rather than limit the present invention, and those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The use of the verb "comprise" and its conjugations does not exclude the presence of elements or steps other than those stated in the claim. The article "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. Expressions such as "at least one of..." when preceding a list of elements or a group of elements mean selecting all elements or any subset of elements from that list or group. For example, the expression "at least one of A, B, and C" should be understood to include only A, only B, only C, both A and B, both A and C, both B and C, or all of A, B, and C. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a device claim enumerating several components, several of these components may be embodied by the same item of hardware. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used advantageously.

Claims

1. A method (300, 400, 500, 600) for providing a task - specific and hardware - architecture - specific machine - learning model, the method (300, 400, 500, 600) comprising: - Providing (320) a trained stacked model, wherein the trained stacked model comprises a stack of a set of machine - learning models, and wherein individual machine - learning models can be extracted from the trained stacked model; - Receiving a characterization of the target hardware architecture; - Fine - tuning (330) the trained stacked model for an application task in a hardware - architecture - agnostic manner, wherein the stacked model has been trained using a general dataset, and wherein the fine - tuning (330) comprises using a labeled dataset specific to the application task; - Selecting a machine - learning model from the fine - tuned stacked model, wherein the selection of the machine - learning model comprises performing a search (340) for the target hardware architecture using a first function that describes a first performance of a candidate machine - learning model for the application task and a second function that describes a second performance of the candidate machine - learning model when executed on the target hardware architecture; - Providing the selected machine - learning model as an output for deployment on the target hardware architecture.

2. The method (300, 400, 500, 600) according to claim 1, further comprising initializing the search (340) based on the results of a past search (350).

3. The method (300, 400, 500, 600) according to claim 2, wherein each machine - learning model comprises a machine - learning model with a different model architecture, and wherein the search (340) comprises using the results of the past search (350) to initialize an evolutionary search process and / or an optimization of a distribution over different model architectures.

4. The method (300, 400, 500, 600) according to claim 2 or 3, wherein the past search (350) comprises a task - agnostic and hardware - architecture - agnostic search, wherein the past search (350) uses functions independent of the task and functions independent of the hardware architecture applied to the trained stacked model, the function independent of the task estimating the performance of candidate machine - learning models of the trained stacked model, and the function independent of the hardware architecture estimating the computational efficiency of the candidate machine - learning models.

5. The method (300, 400, 500, 600) according to claim 4, wherein the function independent of the task is based on values describing one or more of the following: the distillation loss used when training (310) the stacked model, and the number of parameters of the candidate machine - learning model, e.g., as a proxy for the representational ability of the candidate machine - learning model.

6. The method (300, 400, 500, 600) according to claim 4 or 5, wherein the function independent of the hardware architecture is based on values describing one or more of the following: the number of floating - point operations (FLOP), the number of multiply - accumulate (MAC) operations, the latency on a default hardware, and the number of parameters of the candidate machine - learning model, e.g., as a proxy for the memory traffic of the candidate machine - learning model.

7. The method (300, 400, 500, 600) according to any one of claims 1 to 6 further includes using the selected machine learning model to initialize a subsequent search (340A, 340B) for another machine learning model for another application task and / or another target hardware architecture.

8. The method (300, 400, 500, 600) according to any one of claims 1 to 7, wherein the set of machine learning models includes neural networks.

9. The method (300, 400, 500, 600) according to any one of claims 1 to 8 further includes training (310) an ensemble model to provide (320) a trained ensemble model by using the ensemble model as a student model and a base model as a teacher model, the training (310) including using a knowledge transfer method to transfer knowledge from the teacher model to the student model.

10. The method (300, 400, 500, 600) according to claim 9, wherein the knowledge transfer method includes using one or more of the following: knowledge distillation (KD), such as task-agnostic knowledge distillation; and neural architecture search (NAS).

11. The method (300, 400, 500, 600) according to any one of claims 1 to 10, further comprising: Fine-tuning (330.1, 330.2) the trained ensemble model for another application task to obtain another fine-tuned ensemble model; and selecting another machine learning model from the another fine-tuned ensemble model.

12. The method (300, 400, 500, 600) according to any one of claims 1 to 11 further includes providing another task- and hardware-architecture-specific machine learning model for another target hardware architecture by again selecting a machine learning model from the fine-tuned ensemble model.

13. The method (300, 400, 500, 600) according to any one of claims 1 to 12, wherein the search (340) is configured to search for a Pareto-optimal trade-off between a first function and a second function.

14. A system (1140), comprising: One or more processors; and one or more storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the steps of the method (300, 400, 500, 600) according to any one of claims 1 to 13.

15. A transient or non-transient computer-readable medium (1000) including data (1020) representing instructions that, when executed by a processor system (1140), cause the processor system to perform one or more steps of the method (300, 400, 500, 600) according to any one of claims 1 to 13.