System and method for providing task-specific and hardware architecture-specific machine learning model

The method of using a task-independent pre-trained superposed model and fine-tuning for application tasks, combined with hardware-agnostic selection, addresses the inefficiencies of existing models by creating efficient, adaptable machine learning models for various hardware architectures.

JP2025109702APending Publication Date: 2025-07-25ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025004140
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-12
Filing Date
2025-01-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing machine learning models, particularly large foundation models, are not hardware-efficient, leading to high computational and memory costs that hinder deployment in real-world settings, and existing methods for knowledge transfer incur significant computational costs when adapting to new hardware architectures.

Method used

A method involving a trained superposed model that is pre-trained in a task-independent manner, followed by fine-tuning for application tasks without considering hardware architecture, and selecting a machine learning model through a search process that balances application and hardware performance, using knowledge distillation and neural architecture search to optimize for both.

Benefits of technology

This approach enables the creation of task-specific and hardware-architecture-specific models efficiently, reducing computational overhead and allowing adaptation to new hardware architectures with minimal retraining, while maintaining performance on application tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025109702000001_ABST
    Figure 2025109702000001_ABST
Patent Text Reader

Abstract

To provide a system and method for providing a task-specific and hardware architecture-specific machine learning model.SOLUTION: The method includes a step 310 of training a basic model, a step 320 of acquiring a trained superposition model, a step 330 of fine-tuning the trained superposition model for an application task, and a step 340 of performing search to select a machine learning model from the fine-tuned superposition model.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The subject matter of the present disclosure relates to systems and methods for providing task-specific and hardware-architecture-specific machine learning models. The subject matter of the present disclosure further relates to a non-transitory or transitory computer-readable medium including data representing instructions for causing a processor system to perform one or more steps of a method when executed by the processor system.

Background Art

[0002] Background Art A foundation model (FM) is a type of machine learning (ML) model that is relatively large in size and is generally trained with a large amount of data by self-supervised learning or semi-supervised learning. As a result, the model can be configured to be applied to a wide range of application tasks or downstream tasks. In that sense, the foundation model is very general-purpose. This type of model has brought about a major transformation in the way artificial intelligence (AI) systems are built, the evolution of new types of technologies such as language models like Google's Bidirectional Encoder Representations from Transformers (BERT) model, Generative Pre-trained Transformers (GPT) models like OpenAI's GPT-

Number

[0003] However, large foundation models are not hardware-efficient. They incur a large number of parameters as well as very high computational and memory costs, making deployment in real-world settings difficult.

[0004] In an attempt to overcome this problem, past efforts have been made. For this purpose, several knowledge transfer methods have been introduced, by which knowledge can be transferred from a larger teacher network to a smaller student network. In this way, knowledge can be transferred from a base model to a smaller model that can have a smaller number of parameters as well as lower computational and memory costs.

[0005] For example, in the paper "NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture Search" by Xu et al., which can be obtained from https: / / arxiv.org / abs / 2105.14444, a task-agnostic knowledge transfer method using techniques in neural architecture search (NAS) is performed on the BERT model from Google for model compression, resulting in NAS-BERT. Since the compression algorithm used is independent of downstream tasks, the compressed model is generally applicable to different downstream tasks. And NAS-BERT trains a big supernet on a designed search space containing various architectures. During training, NAS-BERT gradually discards architectures based on their training losses, sizes, and latencies on the target hardware. Finally, NAS-BERT outputs multiple models with different sizes and latencies to support different memory and latency constraints for the target device. Furthermore, the training of NAS-BERT is performed with standard self-supervised pre-training tasks and is independent of specific downstream tasks. In this way, the compressed model can be used across various downstream tasks.

[0006] The drawback of NAS-BERT is that steps have to be re-executed for each new target hardware architecture, which incurs a significant computational cost.

Prior Art Documents

Non-Patent Documents

[0007]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0008] Summary of the Invention It is desirable to be able to provide task-specific and hardware-architecture-specific machine learning models in a more computationally efficient manner.

Means for Solving the Problems

[0009] According to a first aspect of the present invention, as defined by claim 1, a method for providing a task-specific and hardware-architecture-specific machine learning model is provided. According to a further aspect of the present invention, a system as defined by claim 14 is provided. According to a further aspect of the present invention, a computer-readable medium as defined by claim 15 is provided.

[0010] The above means may include providing a trained superposed model. Such a superposed model, sometimes referred to as a one-shot model, includes, for example, a superposition of multiple models as extended in "Once-for-All: Train One Network and Specialize It for Efficient Deployment" by Cai et al. (this paper can be obtained from https: / / arxiv.org / pdf / 1908.09791). The trained superposed model may be pre-trained using a general dataset and thus in a task-independent manner. In some examples, the trained superposed model may be used as a student network in a knowledge transfer method for transferring knowledge from a base model to the superposed model. The superposed model may also be trained from scratch, for example, when sufficient data and computing resources are available. The base model and / or the superposed model may include a neural network.

[0011] The above means may further include receiving a characteristic evaluation of the target hardware architecture. This characteristic evaluation of the target hardware architecture may describe aspects of the hardware architecture that may be specific to or otherwise characteristic of the target hardware architecture from the perspective of computational efficiency, such as the number of floating-point operations (FLOP) that can be performed per unit of time, the number of multiply-accumulate operations (MAC), and the latency for executing a specific neural network.

[0012] The above means may further include fine-tuning a trained superimposed model for an application task in a way that is independent of the hardware architecture. The superimposed model may be trained using a general dataset, but here, fine-tuning may include using a labeled dataset specific to the application task. For example, the labeled dataset may include a subset of the general dataset, or may include the entire general dataset. The labeled dataset may include sensor signals as data such as, for example, image data, radar data, and LiDAR data. The application task may include, for example, a classification task such as image perception, and / or a regression task. Fine-tuning may be performed using, for example, the sandwich rule and / or in-place distillation. Fine-tuning may be independent of the hardware architecture in that the evaluation of the characteristics of the target hardware architecture may not be used in the fine-tuning, for example, to manipulate or control the fine-tuning.

[0013] The above means may further include selecting a machine learning model from the fine-tuned superimposed model. The selection of the machine learning model may include performing a search using a first function that describes the first performance of a candidate machine learning model for the application task with respect to the target hardware architecture and a second function that describes the second performance of the candidate machine learning model when executed on the target hardware architecture. Since different candidate machine learning models may represent different model architectures, the search may also be referred to as "model architecture search". These means may be further described as follows.

[0014] The superimposed model can be configured to provide alternative operations on data tensors. Individual machine learning models can be extracted from the trained superimposed model by selecting one or a subset of the operations from each of the alternative operations on the data tensors. For example, a neural network may include a series of operations on data tensors such as convolution, max pooling, and / or activation functions. In the superimposed model, it may be that several alternative operations, rather than just one, are applied to the same data tensor. Some alternative operations may together form a supernet, such as a supernetwork. Such a supernet itself is known from the literature, and reference is made to the exemplary FIG. 1b of "DARTS: Differentiable Architecture Search" by Liu et al., which can be obtained from https: / / arxiv.org / pdf / 1806.09055.pdf. The superimposed model can then enable the extraction of a machine learning model by selecting one or a subset of the operations from each set of alternative operations, and in some cases, not selecting any operation. For an example, refer to the aforementioned paper by Liu et al. where the model architecture shown in FIG. 1d forms a subset of the supernet shown in FIG. 1b. Generally, the individual machine learning models extractable from the superimposed model may have different model architectures, for example, due to differences in the number of network layers or other model components, their relative arrangements (e.g., parallel, serial), the way the components are configured, arranged, and / or connected in the architecture topology (e.g., geometrically and / or logically in the flow of information between different components), the type of network layer or other model components, etc. It may be desirable to select a machine learning model that performs well not only on the local application task but also on the target hardware architecture.For that purpose, the selection of a machine learning model from a fine-tuned ensemble model may involve a search, in which the performance of a candidate machine learning model is evaluated based on a (first) function that describes how the candidate machine learning model performs with respect to an application task and a (second) function that describes the hardware performance of the candidate machine learning model when it is executed on a target hardware architecture. The first function may characterize this performance based on specific or other characteristic properties such as classification accuracy, average precision, complexity, and approximations thereof. The second function may characterize the hardware performance based on specific or other characteristic properties such as values corresponding to different hardware types, measured values such as latency on a sample hardware, energy consumption, memory characteristic evaluation, and / or approximations of either characteristic evaluation values and / or parameters. The search may then be configured to search for a trade-off between the first function and the second function, e.g., a Pareto optimal trade-off that results in a Pareto optimal model architecture.

[0015] The above means may further include providing, as an output, the selected machine learning model for deployment on a target hardware architecture. Since the machine learning model is fine-tuned for an application task and may have been selected from an ensemble model after using a search that may be based on both application performance and computational performance on the target hardware architecture, the resulting selected machine learning model may be considered a task-specific and hardware architecture-specific machine learning model.

[0016] When attempting to deploy the above-mentioned means, a base model, or a similar model that may be trained with general training data of an actual application, the application task is usually given, for example, pedestrian detection, but the hardware architecture may be based on the insight that it can change over time, for example, for technical improvements in the computing architecture, or at a given time for different existing deployment targets (e.g., different products). In other words, the hardware architecture may change frequently, but a given application task, for example, is less likely to change over time. Considering this, the inventors considered performing fine-tuning, which is typically still relatively computationally intensive, in a way that is task-specific but does not depend on the hardware architecture. Therefore, the fine-tuned superimposed model is not yet fixed to a specific target hardware architecture and is only fine-tuned to the in-hand application task. By being a superimposed model, there are different machine learning models extractable from the fine-tuned superimposed model, and the different machine learning models are each fine-tuned for the application task. Then, the adaptation to the target hardware architecture can be achieved by selecting from different machine learning models considering both the application performance and the computing performance on the target hardware architecture. For example, when a new hardware target architecture is considered due to the aforementioned improvements in the computing architecture or different deployment targets, only the search may need to be run again, which is typically less computationally complex than fine-tuning.

[0017] In one embodiment, the search may be initialized based on the results of past searches. For example, considering that the difference between the current search and past searches regarding the hardware architecture may generally be relatively small, the solutions from past searches can usually be potential candidates for the solution of the current search. Therefore, by using past searches for the initialization of the current search, the search efficiency of the current search can be improved.

[0018] In one embodiment, the search includes using the results of past searches to initialize an evolutionary search procedure. The results of past searches can be used as one or more seed models for the evolutionary search procedure. Generally, the evolutionary search, which is time-dependent, can benefit from an initialization based on past searches at the initial time.

[0019] In one embodiment, the search includes using the results of past searches to initialize distribution optimization on the model architecture. The results of past searches can include an optimized distribution on the model architecture that can be used as a starting point for distribution optimization on the model architecture in the search. The optimized distribution on the model architecture can be used as a starting point in one or more search methods for optimizing the distribution on the model architecture within the superimposed model. Considering that the difference between the current distribution optimization and the past distribution optimization is generally likely to be relatively small, and generally, thinking about improving the search efficiency of the current distribution optimization obtained as a result across the model architecture, the optimized distribution, which is the result of past searches, usually functions as a potential candidate for the distribution on the model architecture in the superimposed model.

[0020] In one embodiment, past searches include task-independent searches and hardware-architecture-independent searches. Past searches are applied to a trained stacked model using a task-independent function that estimates the performance of a candidate machine learning model of the trained stacked model and a hardware-architecture-independent function that estimates the computational efficiency of the candidate machine learning model of the trained stacked model. Past searches may be task-independent and / or hardware-independent, meaning that no specific application task or target hardware characteristics are used in the search and the search does not depend on them in any control method. Such past model architectures may, in a broad sense, be applicable to, for example, many application tasks and thus may function as generally possible candidates for solutions to current searches. Thus, this past search can be used to initialize all kinds of task-specific and / or hardware-architecture-specific searches. The use of such past, task-independent and hardware-architecture-independent searches, whose results are transferred to task-specific and hardware-specific searches, can speed up task-specific and hardware-specific searches. To be able to reach optimal or at least reasonable past search results, a task-independent function is provided that can characterize the performance of candidate model architectures for non-specific "general" tasks. Thus, the task-independent function may be a function that describes how a candidate machine learning model performs without considering a specific application task. The task-independent function can characterize this performance based on one or more generally applicable metrics related to the performance of the model, such as the distillation loss used during training of the stacked model, and, for example, the number of parameters of the candidate machine learning model as a proxy for the representational ability of the candidate machine learning model. In addition to the task-independent function, a hardware-architecture-independent function can be used to characterize the performance of candidate model architectures on non-specific "general" hardware architectures.Hardware architecture-independent functions may be functions that describe the computational efficiency of a candidate machine learning model without relying on any particular hardware architecture. The hardware architecture-independent functions may characterize the computational efficiency of a candidate machine learning model based on one or more universally applicable metrics related to computational efficiency, such as the number of floating point operations (FLOP), the number of multiply-accumulate operations (MAC), the number of parameters of the candidate machine learning model, for example, as a proxy for the memory traffic of the candidate machine learning model, and the latency on default hardware. Past searches may be configured to search for Pareto optimal trade-offs, for example, between a first task-independent function and a second hardware architecture-independent function. Thereby, a Pareto optimal model architecture may be obtained.

[0021] In one embodiment, the method further includes initializing a subsequent search for a further machine learning model for other application tasks and / or other target hardware architectures using the selected machine learning model. Considering that the difference between the subsequent search and the current search can generally be relatively small, for example, with respect to the hardware architecture, the results from the current model architecture are likely to be a relatively suitable starting point for the subsequent search and can thus be used to initialize the subsequent search. Thereby, initializing the subsequent search based on the current search can improve the search efficiency of the subsequent search. The application task in the subsequent search can be different from the application task in the current search, but the target hardware architecture in the subsequent search can remain the same or similar to the target hardware architecture in the current search. Conversely, the target hardware architecture in the subsequent search can be different from the target hardware architecture in the current search, but the application task in the subsequent search can remain the same or similar to the application task in the current search. Further, even when both the application task and the target hardware architecture are changed, the changes can be small enough such that the results of the current search can function as a potential starting point for the subsequent search. In this way, the search can be iteratively executed in a computationally efficient manner.

[0022] In one embodiment, the method further includes training the superimposed model to provide a trained superimposed model by using the superimposed model as the student model and the base model as the teacher model, and the training step includes transferring knowledge from the teacher model to the student model using a knowledge transfer method. In this way, useful knowledge can be transferred from the base model to a smaller superimposed model with potentially less training data, at a lower computational cost and memory cost than if the superimposed model had to be directly trained on extensive training data. The resulting trained superimposed model can be applicable, in a broad sense, to, for example, many application tasks. The knowledge transfer method may include the use of knowledge distillation (KD) such as task-agnostic knowledge distillation. Knowledge distillation, as described in "Distilling the Knowledge in a Neural Network" by Hinton et al., available from https: / / arxiv.org / abs / 1503.02531, is a means of transferring knowledge from a large-scale teacher network to a hardware-efficient student network, thereby providing a computationally and resource-efficient model from a large-scale (pre-)trained model. KD may be used before fine-tuning, which enables the reuse of the same student network for several application tasks and makes the student network widely applicable.

[0023] In one embodiment, the knowledge transfer method further includes the use of neural architecture search (NAS) to automate the design of the neural network architecture used in the knowledge transfer method, thereby avoiding the manual design of these architectures and speeding up the process of the knowledge transfer method.

[0024] In one embodiment, the method further includes tuning a trained stacked model for a further application task to provide a further fine-tuned stacked model, and selecting a further machine learning model from the further fine-tuned stacked model. In this way, a plurality of task-specific models can be extracted by tuning from a trained task-independent stacked model. In other words, the trained task-independent stacked model can function as a basis for a plurality of task-specific stacked models, and each of the plurality of task-specific stacked models can function as a basis for a plurality of hardware architecture-specific models. Optionally, additional task-specific stacked models can be obtained only if there are additional application tasks. These multiple instances of tuning for the application tasks may be executed simultaneously or sequentially with respect to the application tasks.

[0025] In one embodiment, the method further includes providing a further task-specific and hardware-architecture-specific machine learning model for a further target hardware architecture by reselecting a machine learning model from the fine-tuned superposition model. This selection of the machine learning model from the fine-tuned superposition model can be performed by a method similar to that described above, including a search. By doing so, a further hardware-specific model for a further target hardware architecture can be derived from the fine-tuned task-specific superposition model. In this way, a plurality of hardware-architecture-specific models can be extracted from the fine-tuned task-specific superposition model that is independent of the hardware architecture, via a further search, without requiring an additional fine-tuning procedure. Optionally, the result from the search for the target hardware architecture may be reused as an initialization for a further search for a further target hardware architecture. When the target hardware architecture and the further target hardware architecture are related to the same application task, this result from the model architecture may be a likely candidate as an initialization for a further search. Note that the searches for different target hardware architectures may be performed simultaneously or sequentially.

[0026] In one embodiment, the superposition model is configured to provide alternative operations of data tensors, and the individual machine learning models can be extracted from the superposition model trained by selecting one or a subset of the operations from each of the alternative operations of the data tensors.

[0027] In a further aspect of the invention, there is provided a system comprising one or more processors and one or more storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations for a method according to an embodiment as described above.

[0028] In a further aspect of the present invention, a transient or non-transient computer-readable medium is provided that includes data representing instructions, which, when executed by a processor system, cause the processor system to perform one or more steps of the methods according to the above-described embodiments.

[0029] Those skilled in the art will understand that two or more of the above-described embodiments, implementations, and / or optional aspects of the present invention may be combined in any useful way.

[0030] Modifications and variations of any device, system, network, computer-implemented method, and / or any computer-readable medium corresponding to other modifications and variations of such entities described can be carried out by those skilled in the art based on this specification.

[0031] Brief Description of the Drawings Further details, aspects, and embodiments are described by way of example only with reference to the drawings. The elements of the drawings are shown for simplicity and clarity and are not necessarily drawn to scale. In the drawings, elements corresponding to those already described may have the same reference numerals.

Brief Description of the Drawings

[0032]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7a

Figure 7b

[0033] Detailed description of embodiments The subject matter of the present disclosure is susceptible to embodiments in many different forms, but is shown in the drawings and will be described in detail herein in one or more specific embodiments. It should be understood, however, that the present disclosure is to be regarded as illustrative of the principles of the subject matter of the present disclosure and is not intended to be limited to the specific embodiments shown and described.

[0034] In the following, for the sake of understanding, the elements of the embodiments will be described during operation. However, it will be apparent that each element is configured to perform the functions described as being performed by them.

[0035] Furthermore, the presently disclosed subject matter is not limited to the embodiments only, but also includes any other combination of features described herein or in the mutually different dependent claims.

[0036] FIG. 1 shows an example of a method 100 for training and fine-tuning a base model similar to the structure of NAS-BERT.

[0037] First, in step 110, the base model is trained. In steps 120A and 120B, KD / NAS is executed for two target hardware architectures respectively identified by the characters "A" and "B". The steps related to each target hardware architecture may be identified by their respective characters as suffixes hereinafter. For example, step 120A may include the step of executing KD / NAS for target hardware architecture A. Continuing to refer to KD / NAS, it should be noted that the knowledge distillation part is adopted as a means of transferring knowledge from the large-scale teacher network, which is the base model in this example, to the one-shot model using neural architecture search (NAS). The one-shot model constitutes the student model. Therefore, in steps 120A and 120B, task-independent but hardware-specific KD / NAS is executed for each target hardware architecture.

[0038] In steps 130A and 130B, task-independent searches are separately executed for each of the target hardware architectures A and B. It should be noted that this task-independent search cannot optimize the architecture for the application task, and the model architecture must be fine-tuned on the application task.

[0039] In steps 140A1-2 and 140B1-2, the obtained model architectures are fine-tuned for the application task. Here, the two application tasks of this application are identified by the numbers "1" and "2". The aggregation of the foregoing steps as method 100 may be regarded as representing NAS-BERT, which can be considered as a model compression version of BERT with a task-independent compression algorithm.

[0040] Method 100 has several drawbacks. First, it is necessary to execute high-computation-cost KD / NAS, which requires hardware metrics such as the latency of the candidate model architecture, for each target hardware architecture. In addition to this, task-agnostic search cannot optimize the architecture of the application task.

[0041] FIG. 2 shows an example of method 200 for training and fine-tuning a base model, which is similar in structure to a method called AutoDistil that is a known technique. AutoDistil represents another attempt to address the problem of computational inefficiency of the base model. AutoDistil is discussed in the paper "AutoDistil: Few-shot Task-agnostic Neural Architecture Search for Distilling Large Language Models" by Xu et al., which can be obtained from https: / / arxiv.org / pdf / 2201.12507. In this paper, KD / NAS is used to compress a large model into a smaller student network and automatically distill several compressed student networks from the large model at various computational costs.

[0042] FIG. 2 shows an overview of a modified example of AutoDistil. First, the base model is trained in step 210. In step 220, KD / NAS is executed to distill a single one-shot model from the base model. In steps 230A and 230B, task-agnostic searches are separately executed for each target hardware architecture A and B. In steps 240A1-2 and 240B1-2, the obtained model architectures are fine-tuned for the application task. The two application tasks of this application are identified by the numbers "1" and "2".

[0043] This variant of AutoDistil200 has several drawbacks. First, it cannot optimize the model architecture for the application task, thereby wasting the potential for optimization. In addition to this, a search must be initiated from scratch for each target hardware architecture, which incurs a fairly high computational cost, especially in a larger search space.

[0044] Figure 3 shows an example of a method for providing a task-specific and hardware-architecture-specific machine learning model.

[0045] In optional step 310, the base model can be trained. If there are sufficient data and computational resources, this training step of the model can be performed on large data. In most cases, it can be assumed that a (pre-)trained base model is available. An example of a pre-trained base model can be found in the paper "EVA: Exploring the Limits of Masked Visual Representation Learning at Scale" by Fang et al., which can be obtained from https: / / arxiv.org / abs / 2211.07636.

[0046] In step 320, a trained stacked model can be obtained. For example, the stacked model may be trained as a student model in a knowledge transfer method during or before step 320. In a specific example, the trained base model from step 310 may be a teacher model in such a knowledge conversion method. The knowledge transfer method used can be knowledge distillation. The knowledge distillation may be task-independent knowledge distillation. Neural architecture search may be used, and the student network has a search space from this neural architecture search, thereby constructing a KD / NAS method. In such a KD / NAS, the student network may be the aforementioned stacked model. The stacked model may represent the stacking of many different neural networks. The stacked model may include, for example, different sequences of operations applied to a data tensor. Such operations may include, for example, convolution, max pooling, activation functions, etc. The stacked model may include alternative operations applied to the same tensor. For example, instead of only a single first operation, the stacked model may include

Number

Number

Number

Number

Number

Number

[0047] In step 330, the trained superposition model can be fine - tuned for the application task. In the fine - tuning 330 for the application task, a labeled dataset that may be specific to the application task can be used as the input to the model 300. This labeled dataset may be relatively smaller than the general dataset on which the trained superposition model can be trained in step 310 and / or on which the base model can be trained in step 320. The fine - tuning of the trained superposition model for the application task using a labeled dataset specific to the application task may be performed in a task - specific and hardware - architecture - independent manner. This may be done, for example, using sandwich rules and / or in - situ distillation techniques, as in the paper "BigNAS: Scaling Up Neural Architecture Search with Big Single - Stage Models" by Yu et al., which can be obtained from https: / / arxiv.org / abs / 2003.11142.

[0048] In step 340, a machine learning model can be selected from the fine-tuned superimposed model. The selection step may include the step of performing a search. The search may be performed for the hardware architecture that receives the characteristic evaluation of the target hardware architecture as input. The reception of the characteristic evaluation of the target hardware architecture may be performed in one of the previous steps, for example, one of steps 310 to 330. The search may use two functions that may have been received as input. The first function may be a function that describes the performance of the candidate machine learning model of the fine-tuned superimposed model for the application task. This function may be determined particularly for the application task at hand. For example, the metric used in the function may be particularly selected so as to quantify the performance of the application task at hand. The first function may include, for example, metrics such as classification accuracy, average precision, complexity, or some combination of such metrics. The second function may be a function that describes the performance of the candidate machine learning model of the fine-tuned superimposed mode when executed on the target hardware architecture. For that purpose, the second function may describe hardware metrics that can particularly characterize the performance on the target hardware architecture. For example, the second function may include metrics such as the latency, energy consumption, etc. of the hardware architecture measured on the target hardware, approximations of these metrics, or some combination of such metrics. Then, the search is performed on the fine-tuned superimposed model using the two functions, and may return the model architecture that best satisfies such a combination of functions. In a particular example, the search may be configured to search for the Pareto optimal trade-off between the first function and the second function. This may result in the Pareto optimal model architecture for the application task and the target hardware architecture that may be returned by the search.

[0049] Compared with the method 100 of FIG. 1, the step of providing a trained superimposed model in step 320 of method 300, which can be performed, for example, by hardware architecture-independent KD / NAS, instead of the requirements of method 100 for independent KD / NAS steps for each target hardware architecture, the computationally expensive steps of KD / NAS can be made to be executed only once for various target hardware architectures and application tasks.

[0050] FIG. 4 shows another example of method 400 for providing task-specific and hardware architecture-specific machine learning models.

[0051] In addition to method 300, method 400 may include fine-tuning of the trained superimposed model steps for a plurality of application tasks identified by the numbers "1" and "2" in steps 330.1 and 330.2, respectively. In steps 340.1 and 340.2, a machine learning model may be selected from the fine-tuned superimposed model. Each selection of the machine learning model may include performing a search. The search may be performed for different target hardware architectures. The fine-tuning 330.1, 330.2 for application tasks 1 and 2 may be performed simultaneously with respect to application tasks 1 and 2, and / or the selection steps 340.1, 340.2 of the machine learning model for different target hardware architectures may be performed simultaneously with respect to the target hardware architectures. The fine-tuning steps 330.1, 330.2 may also be performed sequentially, for example, continuously, with respect to application tasks 1 and 2. Similarly, the selection steps 340.1, 340.2 may also be performed sequentially with respect to the target hardware architectures.

[0052] FIG. 5 shows another example of a method 500 for providing task-specific and hardware-architecture-specific machine learning models.

[0053] In addition to method 300, method 500 may include, at steps 340A and 340B, a search for a fine-tuned superimposed model for a plurality of target hardware architectures identified by the letters "A" and "B", respectively. The searches 340A, 340B for target hardware architectures A and B may be performed simultaneously with respect to target hardware architectures A and B. The searches 340A, 340B may also be performed sequentially, e.g., continuously, with respect to target hardware architectures A and B.

[0054] In addition to method 300, method 500 may further include an additional step 350 that may include past searches. The past searches in step 350 may be task-independent searches and hardware-architecture-independent searches, such that the results may be used for searches for both target hardware architectures A and B. The past model architectures may utilize a first function and a second function. The first function may be a task-independent function that can estimate the performance of a candidate machine learning model of a trained stacked model. The first function may be based on one or more universally applicable metrics related to the performance of the model, such as, for example, the distillation loss used during training of the trained stacked model and the number of parameters of the candidate machine learning model, as a proxy for the representational ability of the model architecture. The second function may be a hardware-independent function that can characterize the computational efficiency of the candidate model architecture. The second function may be based on one or more universally applicable metrics related to computational efficiency, such as the number of floating-point operations (FLOP), the number of multiply-accumulate operations (MAC), the number of parameters of the candidate machine learning model as a proxy for the memory traffic of the model architecture, and / or the latency on default hardware. During the past searches, the first and second functions may be used to perform an evolutionary architecture search that is task-independent and hardware-independent for the trained stacked model. In some embodiments, the past searches may be configured to search for a Pareto-optimal trade-off between the first function and the second function, and thereby may return the resulting Pareto-optimal model architecture.

[0055] The search for target hardware architectures A and B in steps 340A and 340B may be initialized based on the results of past searches. The search may be performed as an evolutionary search, for example, by using the results of past model architectures as seed models for task-specific and hardware-specific evolutionary searches 340A, 340B. The searches in steps 340A and 340B may also be performed as, for example, distribution optimizations on the model architecture, and the model architecture may use the optimized distribution as a starting point, and the optimized distribution may result from past searches that may be performed in step 350.

[0056] In some embodiments, the results of the past searches in step 350 may include a set of Pareto-optimal model architectures for the first and second functions, may be used to initialize task-specific and hardware-specific searches in steps 340A and 340B, and the set of Pareto-optimal model architectures may be used as an initial population.

[0057] FIG. 6 shows an example of a method 600 for providing task-specific and hardware-architecture-specific machine learning models. Method 600 may combine features from models 400 and 500.

[0058] Method 600 may include, in addition to method 300, fine-tuning of the trained superimposed model, which is the output of step 320 for a plurality of application tasks identified by the numbers "1" and "2" in steps 330.1 and 330.2 respectively here. In steps 340.1A, 340.1B, 340.2A and 340.2, searches may be performed for a plurality of target hardware architectures identified by the letters "A" and "B". The fine-tuning 330.1, 330.2 for application tasks 1 and 2 may be performed simultaneously with respect to application tasks 1 and 2, and / or the searches 340.1A, 340.1B, 340.2A and 340.2B for the target hardware architectures may be performed simultaneously with respect to hardware architectures A and B. The searches 340.1A, 340.1B, 340.2A, 340.2B may also be performed sequentially, for example continuously, with respect to the target hardware architectures A and B. Further, the results of past searches in step 350 may be used to initialize the searches in all of steps 340.1A, 340.1B, 340.2A and 340.2B.

[0059] Any of the methods described herein may be implemented on a computer as a computer-implemented method, as dedicated hardware, or as a combination of both. As also shown in FIG. 7a, instructions for a computer, such as executable code, may be stored in a computer-readable medium 1000, 1001 in the form of, for example, a series of machine-readable physical marks 1020, 1021 and / or as a series of elements having different electrical, for example magnetic, or optical characteristics or values. The computer-readable media 1000, 1001 may be either a temporary or a non-temporary medium. Examples of computer-readable media include memory devices, optical storage devices, integrated circuits, etc. By way of example, FIG. 7a shows an optical storage device 1000 and a memory card 1001.

[0060] FIG. 7b shows a processor system 1140 that may comprise or represent a system configured to execute a method as described elsewhere herein. The processor system may comprise one or more subsystems or components 1110. For example, a processing subsystem 1120 may be provided for executing computer program components for executing a method as described elsewhere herein. A memory 1122 may be provided for storing programming code, data, etc. A communication subsystem 1126, such as a network interface, may enable communication with other entities. In some examples, an application specific integrated circuit 1124 may be provided for executing some or all of the processing related to the methods described elsewhere herein. The processing subsystem 1120, the memory 1122, the application specific IC 1124, and the communication subsystem 1126 may be interconnected with each other via an interconnection 1130, such as a bus. System 1140 is shown as including one of each of the components described, but various components may be replicated in various embodiments. For example, the processing subsystem 1120 may include a plurality of microprocessors configured to independently execute the methods described herein, or configured to execute steps or subroutines of the methods described herein, and the plurality of processors may cooperate to achieve the functions described herein. Further, where system 1140 may be implemented in a cloud computing system, cloud server and / or computing farm, various hardware components may belong to separate physical systems. For example, the processing subsystem 1120 may include a first processor in a first server and a second processor in a second server.

[0061] In an alternative embodiment of FIG. 7b, the processor system 1140 may represent the target hardware architecture on which the selected machine learning model is deployed. In other words, the processor system may represent a deployment target on which application tasks may be executed as described elsewhere herein. The processor system 1140 may be, for example, a device or apparatus. The device or apparatus may, for example, comprise a sensor capable of identifying measurements of the environment in the form of sensor signals, where the sensor signals may be provided, for example, by digital images such as video, radar, LiDAR, ultrasonic, thermographic motion images, or audio signals. The device or apparatus may, for example, be a household device comprising a sensor for detecting the presence of an object inside a washing machine, or a vehicle comprising a sensor for detecting the presence of an object in the environment of a vehicle. The application tasks may include classifying data from the sensor, detecting the presence of an object in the sensor data, and / or performing semantic segmentation on the data, for example, with respect to traffic signs, road surfaces, pedestrians, and vehicles. Other application tasks may include determining one or more continuous values, for example, performing a regression analysis on items in the data, such as the distance, speed, acceleration, and / or tracking of an object. Examples of these application tasks may be performed on low-level features such as edges or pixel attributes in the case of image data. Other application tasks may include, for example, detecting anomalies in technical systems such as computer-controlled machines as robotic systems, vehicles, household devices such as washing machines, power tools, manufacturing machines, personal assistants, or access control systems, or systems for transmitting information, such as monitoring systems or medical systems as medical imaging systems, and calculating control signals for controlling the technical systems.

[0062] Features of the examples, embodiments, or alternatives, whether or not limiting, should not be understood as limiting the claimed invention.

[0063] It should be noted that the above embodiments are illustrative rather than limiting the present invention, and those skilled in the art can design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claims. The use of the verb "comprise" and its conjugations does not exclude the presence of elements or steps other than those recited in the claims. The article "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. Expressions such as "at least one" when preceding a list or group of elements represent the selection of all or any subset of elements from the list or group. For example, the expression "at least one of A, B, and C" should be understood to include only A, only B, only C, both A and B, both A and C, both B and C, or all of A, B, and C. The present invention can be implemented by means of hardware including several distinct elements and, appropriately programmed computers. In device claims listing several means, some of these means may be embodied by the same item of hardware. The mere fact that certain means are recited in mutually different dependent claims does not indicate that a combination of these means cannot be used advantageously.

Explanation of Reference Signs

[0064] List of Reference Signs The following list of signs and abbreviations is provided to facilitate the interpretation of the drawings and should not be construed as limiting the claims. 100 Method for training and fine-tuning a base model 110 Training of a base model 120A Execution of KD / NAS for target hardware architecture A 120B Execution of KD / NAS for target hardware architecture B 130A, 130B Task-independent search Fine-tuning of Application Task 1 for 140A1, 140B1 Fine-tuning of Application Task 2 for 140A2, 140B2 Method for training and fine-tuning a base model Training of the base model Execution of KD / NAS for the one-shot model Execution of task-independent search for target hardware architecture A Execution of task-independent search for target hardware architecture B Fine-tuning of Application Task 1 for 240A1, 240B1 Fine-tuning of Application Task 2 for 240A2, 240B2 Method for providing a task-specific and hardware-architecture-specific machine learning model Optional training of the base model Provision of a trained stacked model Fine-tuning of the trained stacked model for application tasks in a hardware-independent manner Search for the hardware architecture Method for providing a task-specific and hardware-architecture-specific machine learning model Fine-tuning of the trained stacked model for Application Task 1 in a hardware-independent manner Fine-tuning of the trained stacked model for Application Task 2 in a hardware-independent manner Search for the hardware architecture Method for providing a task-specific and hardware-architecture-specific machine learning model Search for target hardware architecture A Search for Target Hardware Architecture B of 340B 350 Past, Task-Independent Search and Hardware Architecture-Independent Search 600 Method for Providing Task-Specific and Hardware Architecture-Specific Machine Learning Models Search for Target Hardware Architecture A of 340.1A, 340.2A Search for Target Hardware Architecture B of 340.1B, 340.2B 1000 Optical Memory Device 1001 Memory Card 1020, 1021 Memory Data 1140 Processor System 1110 Subsystem or Component 1120 Processing Subsystem 1122 Memory 1124 Application-Specific Integrated Circuit 1126 Communication Interface 1130 Interconnection

Claims

1. A method (300, 400, 500, 600) for providing a machine learning model specific to a task and a hardware architecture, comprising: The method (300, 400, 500, 600) comprises: providing a trained stacked model (320), wherein the trained stacked model comprises a stack of a set of machine learning models, and individual machine learning models are extractable from the trained stacked model; receiving an evaluation of the characteristics of a target hardware architecture; fine-tuning (330) the trained stacked model for an application task by a method independent of the hardware architecture, wherein the stacked model is trained using a general dataset, and the step of fine-tuning (330) comprises using a labeled dataset specific to the application task; selecting a machine learning model from the fine-tuned stacked model, wherein the step of selecting the machine learning model comprises performing a search (340) using a first function that describes a first performance of a candidate machine learning model for the application task and a second function that describes a second performance of the candidate machine learning model when executed on the target hardware architecture, for the target hardware architecture; providing the selected machine learning model as an output for deployment on a target hardware architecture; A method (300, 400, 500, 600) comprising the above steps.

2. The method (300, 400, 500, 600) according to claim 1, further comprising initializing the search (340) based on the results of a past search (350).

3. The individual machine learning models include machine learning models having different model architectures, The search (340) according to claim 2, wherein the search (340) comprises initializing an evolutionary search procedure and / or distribution optimization across the different model architectures using the results of the past search (350).

4. The past search (350) includes a task-independent search and a hardware-architecture-independent search, The past search (350) is applied to the trained stacked model using a task-independent function for estimating the performance of a candidate machine learning model of the trained stacked model and a hardware-architecture-independent function for estimating the computational efficiency of the candidate machine learning model, according to the method (300, 400, 500, 600) of claim 2 or 3. **Claim 5** The task-independent function is based on a distillation loss used during the step (310) of training the stacked model and one or more of the values describing the number of parameters of the candidate machine learning model, for example, as a proxy for the representational ability of the candidate machine learning model, according to the method (300, 400, 500, 600) of claim 4. **Claim 6** The hardware-architecture-independent function is based on the number of floating-point operations (FLOP), the number of multiply-accumulate operations (MAC), the latency on default hardware, and one or more of the values describing the number of parameters of the candidate machine learning model, for example, as a proxy for the memory traffic of the candidate machine learning model, according to the method (300, 400, 500, 600) of claim 4 or 5. **Claim 7** The method (300, 400, 500, 600) according to any one of claims 1 to 6 further includes the step (340A, 340B) of initializing a subsequent search for a further machine learning model for other application tasks and / or other target hardware architectures using the selected machine learning model. **Claim 8** The set of machine learning models includes a neural network, according to the method (300, 400, 500, 600) of any one of claims 1 to 7. **Claim 9** The method (300, 400, 500, 600) according to any one of claims 1 to 8 further includes the step (310) of training the stacked model to provide the trained stacked model (320) by using the stacked model as a student model and using the base model as a teacher model. The step (310) of training includes the step of transferring knowledge from the teacher model to the student model using a knowledge transfer method. **Claim 10** The method of knowledge transfer as claimed in claim 9, wherein the method (300, 400, 500, 600) includes using one or more of knowledge distillation (KD) such as task-independent knowledge distillation and neural architecture search (NAS).

11. The step (330.1, 330.2) of fine-tuning the trained stacked model for a further application task to obtain a further fine-tuned stacked model, The step of selecting a further machine learning model from the further fine-tuned stacked model, The method (300, 400, 500, 600) according to any one of claims 1 to 10, further comprising.

12. The method (300, 400, 500, 600) according to any one of claims 1 to 11, further comprising the step of providing a further task-specific and hardware architecture-specific machine learning model for a further target hardware architecture by reselecting a machine learning model from the fine-tuned stacked model.

13. The method (300, 400, 500, 600) according to any one of claims 1 to 12, wherein the search (340) is configured to search for a Pareto optimal trade-off between the first function and the second function.

14. One or more processors, One or more storage devices storing instructions for causing the one or more processors to perform the steps for the method (300, 400, 500, 600) according to any one of claims 1 to 13 when executed by the one or more processors, A system (1140) comprising.

15. A temporary or non-temporary computer-readable medium (1000) comprising data (1020) representing instructions for causing a processor system (1140) to perform one or more steps of the method (300, 400, 500, 600) according to any one of claims 1 to 13 when executed by the processor system.