ARTIFICIAL INTELLIGENCE (AI) FOR THE HARDWARE / SOFTWARE CO-DESIGN OF ACCELERATORS AND MACHINE LEARNING MODELS

By iteratively designing optimized hardware and software configurations using a machine learning regression process with active learning, the system addresses the inefficiencies of standard hardware processors for AI workloads, achieving low latency, high throughput, and accurate model inference.

DE102024115801A1Pending Publication Date: 2025-06-05HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024115801
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-06-06
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Modern AI workloads process large datasets, which is not ideal for standard or generically configured hardware processors. Existing hardware processors, even when configured as accelerators, often do not provide an optimal configuration for specific functions or workloads, requiring manual and resource-intensive adaptation.

Method used

The system iteratively determines hardware and software configurations by jointly designing these elements to form an optimized system configuration. This process searches the configuration spaces of hardware and software to determine software and hardware parameters, using a machine learning regression process with active learning to simulate and optimize device configurations for maximum model accuracy and minimized hardware costs.

Benefits of technology

The approach achieves optimized hardware and software configurations that provide low latency, high throughput, and accurate inference of machine learning models, while meeting physical and performance constraints, thereby improving the efficiency and effectiveness of AI workloads.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Systems and methods are provided for the iterative co-design of hardware and software elements of a device configuration in order to optimize it. This process can predict how software and hardware parameters in the device configuration will behave by using a machine learning process that simulates how the device configuration will behave. The corresponding inputs (e.g., software / hardware parameters) and outputs (e.g., model accuracy score for software parameters and hardware cost estimate for hardware parameters) determined from the machine learning process can be used for various purposes, including training a machine learning (ML) model to select the optimized device configuration given the various constraints.
Need to check novelty before this filing date? Find Prior Art

Description

background

[0001] Processors are used to execute machine-readable instructions that cause the processor to perform various actions on the computer system. In some examples, the actions implement machine learning models and other high-performance computing operations. The characteristics of the processor can determine how quickly, accurately, and efficiently the processor processes these actions. For example, general-purpose graphics processing units (GPUs) can run models that other processors can run, but other processors may perform the same tasks more slowly and less accurately. Indeed, the right hardware configuration for the processing task can improve the overall efficiency of the system. Brief description of the drawings

[0002] The present disclosure will be described in detail in accordance with one or more various embodiments with reference to the following figures. The figures are for illustrative purposes only and represent only typical or exemplary embodiments. Fig. Figure 1 shows a hardware and software co-design system corresponding to some examples of the system. Fig. Figure 2 shows a decision tree ensemble, its mapping to a hardware device, and a prediction step in accordance with some examples of the system. Fig. Figure 3 illustrates a process for generating hardware and software optimization configurations using the hardware and software co-design system, in accordance with some examples of the system. Fig. Figure 4 shows a process optimization using the hardware and software co-design system in accordance with some examples of the system. Fig. Figure 5 shows a machine learning regression model consistent with some examples of the system. Fig. Figure 6 shows a function for expected improvement detection in accordance with some examples of the system. Fig. Figure 7 shows the pseudocode of computer-readable instructions for level two co-design using a machine learning regression process with active learning according to some examples of the system. Fig. Figure 8 shows optimization metrics according to some examples. Fig. Figure 9 shows an example computer component that can be used to implement hardware and software co-design in accordance with some examples of the system. Fig. 10 is an example of a computer component that may be used to implement various features of the embodiments described in the present disclosure.

[0003] The figures are not exhaustive and do not limit the present disclosure to the precise form disclosed. Detailed description

[0004] Modern artificial intelligence (AI) workloads process large datasets, which is not ideal for standard or generically configured hardware processors. Some hardware processors are specifically designed as accelerators for a specific function or workload. Even if the processor is configured as an accelerator for a specific purpose, the processor / accelerator can handle a wide range of workloads but rarely provides an optimal configuration for specific functions or workloads. Some administrative users can customize hardware processors for ML workloads or other software-based processes, although creating these configurations is often a manual process and resource-intensive.Developing new hardware for specific applications and implementing software that leverages the hardware's unique features are interdependent problems that should be addressed jointly to achieve maximum improvements. Optimizing the design of both hardware and software components presents many challenges and may require frequent communication between hardware and software experts.

[0005] Examples of the disclosure iteratively determine hardware and software configurations by jointly designing these elements to form an optimized system configuration. This process can determine software parameters and hardware components and is applicable to general search spaces (e.g., the finite hardware and software configuration space to be searched) and optimization objective variables (e.g., the features of a software / hardware pair to be maximized in the model output). For example, the system may first search the search spaces of hardware and software configurations to determine software parameters, model parameters, and / or hardware parameters.Software or model parameters (used interchangeably) include values ​​associated with a software application, including a type of ML model, a data size, a software program size, and other measurable characteristics of the software that can be stored as software parameters. Hardware parameters include values ​​associated with a hardware device, such as the type of hardware device, the size of memory or other device component, the speed of the processor, and other measurable characteristics of the hardware that can be stored as hardware parameters.

[0006] After determining the parameters, the system can determine a first device configuration associated with the parameters. For example, the latency, range, and throughput of the first device configuration can be determined / estimated using a closed-form hardware cost model, where the metrics associated with the first device configuration are measured in a simulated environment that implements the device configuration with hardware and software parameters. The machine learning regression process model can be implemented in a simulated or virtual environment using the software and hardware parameters determined from the configuration sample and used to generate metrics associated with the first device configuration.

[0007] In some examples, the machine learning process is Gaussian process regression with active learning, although other forms of machine learning processes can be implemented without detracting from the disclosure. Active learning refers to a type of machine learning in which the learning algorithm can interactively query a source of information to assign specific values ​​to new data points (e.g., using the Botorch® or GpyTorch® packages). The machine learning process can iteratively and sequentially simulate different device configurations with different hardware and software parameters.

[0008] The metrics predicted from the simulation can include a model accuracy score and a hardware cost estimate for each device configuration. A model accuracy score includes a value corresponding to a relative degree of accuracy of the software output (e.g., ML model, software / model parameters) when combined with the particular hardware processor configuration (e.g., hardware parameters). A better model accuracy score can be maximized compared to other model accuracy scores. On the hardware side, the hardware cost estimate includes values ​​to be maximized or minimized according to the hardware parameters of a device configuration, including latency (minimize), area (minimize), throughput (maximize), and other hardware cost values.

[0009] The model accuracy score and the hardware cost estimate score can be generated simultaneously. For example, the machine learning process can apply the device configuration in the simulated environment using the appropriate sets of hardware and software parameters. These parameters can be applied simultaneously and for a single configuration pair. Once the sets of hardware and software parameters are simulated, the process can wait / delay execution of the next iteration of the process until the output is determined. In some examples, the process can wait to determine the result of the previous simulation, which can help select the next set of hardware and software parameters.The selection of parameters in the two parameter spaces (hardware and software) is performed simultaneously and jointly, which is ensured by the expected improvement capture function.

[0010] The simulated output can be used in various ways. For example, the simulated output can be used to train an ML model, which can apply weights and biases that can tune and optimize any values ​​associated with the simulated output. The output of the ML model can determine a new device configuration that maximizes output for the corresponding hardware and software parameters. For example, the new device configuration can maximize a model accuracy score while minimizing an estimated hardware cost of the new device configuration. In some examples, the output can be used to train an ML model during additional training stages of device configurations, or can predict a device configuration that maximizes output for corresponding hardware / software parameters.

[0011] Technical improvements are achieved throughout the process. On the hardware side, the optimized solution can offer low latency and high throughput, while meeting physical constraints (e.g., the area available on a silicon chip) and power limitations. On the software or application side, effective model implementations can leverage the individually chosen hardware for the software implementation to deliver accurate and fast inference of the ML model (e.g., ensembles of decision trees). In some examples, an active learning Gaussian process regression model is used to decide where in the search space to test the next available implementation of hardware and software to ensure they work together.

[0012] In some examples, the method can also improve the machine learning model. For example, the system can help identify fewer decision trees in decision ensembles while mapping more decision trees to each hardware configuration. Joint optimization can also reduce the accuracy of features for software parameters derived in each device configuration by performing feature binning without sacrificing model accuracy while optimizing hardware performance and enabling large throughput improvements.

[0013] Fig. 1 shows a hardware and software co-design system in accordance with some examples of the system. In example 100, the hardware and software co-design system 102 is shown in communication with the network 140 and a set of hardware devices 130. The hardware and software co-design system 102 includes a processor 104, a memory 105, and a machine-readable medium 106 for storing computer-readable instructions for performing various operations described herein. The set of hardware devices 130 may include second and separate processors, memory, and software components (not shown) for implementing the device configuration determined by the hardware and software co-design system 102.

[0014] Processor 104 may be a general-purpose or special-purpose processing engine, such as a microprocessor, a controller, or other control logic. Processor 104 may be connected to a bus, although any communication medium may be used to facilitate interaction with other components of hardware and software co-design system 102 or to communicate externally.

[0015] Memory 105 may include random access memory (RAM) or other dynamic memory for storing information and instructions to be executed by processor 104. Memory 105 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 104. Memory 105 may also include read-only memory ("ROM") or other static storage device connected to a bus to store static information and instructions for processor 104.

[0016] Machine-readable media 106 may include one or more interfaces, circuits, and modules for implementing the functions described herein. Machine-readable media 106 may contain one or more sequences of one or more instructions executed by processor 104. Such instructions embodied on machine-readable media 106 may enable hardware and software co-design system 102 to perform features or functions of the technology described herein. The interfaces, circuits, and modules of machine-readable media 106 may include, for example, data processing module 108, device configuration module 110, simulation module 112, artificial intelligence (AI) module 114, and user interface module 116.Various data stores may also be managed by the hardware and software co-design system 102, including the hardware parameter data store 120, the software parameter data store 122, the model parameter data store 124, and the hardware cost estimation data store 126.

[0017] The data processing module 108 is configured to receive a set of hardware parameters and a set of software parameters to configure a device consistent with individual parameters of the set of hardware parameters and individual parameters of the set of software parameters. As explained herein, software or model parameters (used interchangeably) include values ​​associated with a software application, including a type of ML model, a data size, a software program size, and other measurable characteristics of the software that can be stored as software parameters. Hardware parameters include values ​​associated with a hardware device, such as the type of hardware device, the size of memory or other device component, the speed of the processor, and other measurable characteristics of the hardware that can be stored as hardware parameters.

[0018] In some examples, once the set of hardware parameters and a set of software parameters for configuring a device are received as a data set and selected by the hardware and software co-design system 102, the data processing module 108 may not further interact with the parameters and instead rely on various device configurations generated by the device configuration module 110.

[0019] The device configuration module 110 may determine a first device configuration for the device using a first set of hardware parameters from the set of hardware parameters and a first set of software parameters from the set of software parameters. The device configuration module 110 may also determine a second device configuration using the first software model accuracy score and the first hardware cost estimate score from the ML model (implemented by the AI ​​module 114).

[0020] A device configuration may consider hardware and software constraints. For example, on the hardware side, the device configuration module 110 may consider optimizing features to achieve low latency and high throughput and to adhere to the physical constraints (e.g., the area available on a silicon chip) and the performance limitations of the hardware. On the software or application side, the device configuration module 110 may consider optimizing features to utilize the hardware and provide accurate and fast inference of the model. The system may require frequent communication between hardware and software components of the simulated device configuration.

[0021] The device configuration module 110 is also configured to provide the device configuration to the simulation module 112 to evaluate the corresponding hardware and software parameters of the device configuration in the simulated or virtual environment. The simulations can be executed sequentially so that the results of a first simulation can be used in the second simulation.

[0022] To determine the performance of the configuration, various evaluation methods can be employed, including using a machine learning model such as a Gaussian process regression model, and the next set of hardware and software parameters can be selected using active learning methods. Other types of evaluation methods can be implemented without detracting from the essence of the disclosure, including a multi-arm bandit model, global optimization, Pareto optimum, or other probabilistic methods. For example, the machine learning process can determine the output of a particular device configuration, and active learning can select the next hardware and software parameters to be tested using the Gaussian process regression model (e.g., to determine model accuracy and hardware cost values).

[0023] The system also identifies the next parameters corresponding to a specific device configuration to be investigated using active learning. The next parameters and configurations can be fed back into the machine learning process to sequentially determine the optimization values ​​of the various configuration settings. The process can stop determining the device configurations when the output corresponding to the hardware parameters and the software parameters exceeds a predetermined threshold (e.g., the threshold, which is a specific optimization value, is determined, or a certain amount of improvement between the hardware and software configurations increases less than a threshold).

[0024] The machine learning process can attempt to optimize hardware and software metrics for an optimized device configuration. A normalized set of measurements "y," corresponding to a parameter sample "x," is written as follows: y=P(X)=z(∑iz(Pi(X))),where z(X)=X−μ^(X)σ^(X).

[0025] Using this formula, the system can attempt to optimize various hardware and software metrics. For example, using three hardware metrics and one software metric, the system can minimize the model's hardware area, hardware latency, and root mean square error (RMSE) metric while maximizing hardware throughput and model accuracy. The metrics can be combined into a single scale, and the system can independently calculate scores (z-scores) for each metric for the sum of the normalized metrics. The system can define the objective function f(x) = y as a real-valued single-output function of a vector of hardware and software parameters.

[0026] In some examples, the device configuration module 110 (with the simulation module 112) may find the best compromise for the hardware capable of executing the corresponding software, e.g., a large set of optimized models. The system may explore models and specific hardware together to determine the best compromise.

[0027] The simulation module 112 is configured to determine a simulated output from the device configuration determined by the device configuration module 110. For example, the simulation module 112 may create concurrent systems in a virtual environment where the instruction set architecture (ISA), microarchitecture, and memory interact with the programming model and the communication system. Some examples may include components of a virtualized processor, memory, disk, network, and software model. The simulator may remain unchanged while accepting various hardware and software parameters as input and determining event types or values ​​as output. During a simulation, the simulation module 112 may, for example, determine events related to computations, communications (e.g., MPI events), sleep, or memory reads, as well as the corresponding values ​​associated with such events.

[0028] In some examples, the simulation module 112 may apply the first set of hardware parameters and the first set of software parameters to the ML model (implemented by the artificial intelligence (AI) module 114). The output of the model may help determine when and how much time is spent executing processes for these events, in addition to other metrics discussed herein.

[0029] In some examples, simulation module 112 may receive a parameter provided in a parallel simulation environment based on MPI events. This process may provide a high level of performance and the ability to consider large systems. The model may determine the expected performance of device configurations ranging from in-memory processing to conventional processors connected via conventional network interfaces and executing MPI.

[0030] In some examples, the simulation module 112 may perform the simulation on an external system or with a toolkit (e.g., Structural Simulation Toolkit (SST)). As an external system, the simulation may virtualize the software and hardware components of the device configuration, with the ISA, microarchitecture, and memory interacting with the programming model and communication system while measuring the effects of the configuration. The measurements may be returned to the simulation module 112. In this sense, the simulation module 112 may enable a modular design, allowing the measurement of a single system parameter without changing the components of the simulation module 112, or provide a parallel simulation environment based on MPI.

[0031] The artificial intelligence (AI) module 114 is configured to determine the output of the first device configuration by applying the first set of hardware parameters and the first set of software parameters to the ML model. The output of the ML model simultaneously generates a first software model accuracy score and a first hardware cost estimate for the first device configuration consistent with the first set of hardware parameters and the first set of software parameters.

[0032] In some examples, the AI ​​module 114 implements tree-based model inference for the co-design process. The inference can determine the trade-offs between software and hardware metrics, such as software / model accuracy and hardware area, latency, and throughput. The ML model can correspond to a linear regression model or a multi-objective Gaussian process (GP) regression model to merge hardware and software search spaces. In this process, the joint space can be determined using active learning capture functions (e.g., the Expected Improvement (El) criterion).

[0033] In some examples, the ML model is a Gradient Boosted Decision Tree Ensemble. For example, if an input point x = (x 1 , ... , x K), where K is the number of input dimensions or features, and a function f (x) = y, where y is a real number or a class label, the learning problem is to construct, given data (X, y = f (x ∈ X)), a predictor ^fθ that is able to estimate f(x') for new points x' not contained in X.

[0034] In some examples, a binary decision tree ensemble performs the inference. The ensemble of binary decision trees is T = {T 1 , ... , T P}, where P is the number of trees in the ensemble. Inference is performed by comparing the input feature values ​​x 1 , ... , x K with thresholds in the nodes of each tree.

[0035] In some examples, each node compares a single threshold with a single feature. The features can appear on multiple nodes, and the trees do not need to be balanced. All subpredictions of the tree can consist of individual paths taken from each input point, and the leaves reached by each tree in the ensemble are combined in a subsequent reduction step. This can yield the final regression or classification prediction y^ = f^ θ (x'), where θ consists of the thresholds and parameters of the tree ensemble.

[0036] The AI ​​module 114 may implement a training process for the ML model. The training may consist of training the tree ensemble model using the gradient-boosted decision tree algorithm, with each new tree iteratively fitted to the partial residuals of the ensemble. The process may include many parameters for configuring and constraining the resulting tree ensemble and the training process, such as the maximum number of trees, the learning rate, and the implementation of various subsampling methods.

[0037] In some examples, the ML model may perform feature binning. For example, feature binning may divide continuous or other numeric features into distinct groups. This can allow the ML model to highlight important trends in the data by focusing feature binning on tree splits in the decision tree ensemble.

[0038] In some examples, the performance of tree ensemble models can be measured by the percentage of accurate predictions for classification and root mean squared error (RMSE) for regression tasks.

[0039] The AI ​​module 114 may compare the accuracy values ​​to thresholds or previous accuracy values. The comparison may include, for example, comparing the second device configuration to a second software model accuracy score. In another example, the comparison may compare a second hardware cost estimate to a first software model accuracy score and the first hardware cost estimate of the first device configuration. If the second values ​​are greater than the first values, the second device configuration may be communicated to the interface (via the user interface module 116).

[0040] The user interface module 116 is configured to provide a device configuration for a display. The device configuration may, for example, use a set of hardware parameters from the set of hardware parameters and a set of software parameters from the set of hardware parameters that optimize the determined metrics. The device configuration may, for example, include a device configuration that is better than other device configurations with respect to one or more of the following characteristics: low latency, high throughput, hardware configuration that fits within the chip area, compliance with power constraints, accurate and fast inference of the ML model, or other consistently described metrics.

[0041] The hardware parameter data store 120 includes values ​​associated with a hardware device, including the type of hardware device, the size of memory or other device component, the processor speed, and other measurable characteristics of the hardware that can be stored as hardware parameters. An illustrative hardware device may include an analog content-addressable memory (aCAM), and the corresponding hardware parameters for the aCAM device may include a number and type of cells in the aCAM, various organizational structures of the aCAM cells such as the height and width of groups of aCAM cells, various organizational structures of the network-on-chip (NoC) communication structure between aCAM cells such as tree branching and depth, various memory buffer sizes such as buffers at each level of the NoC, and various pipeline parameters such as pipeline depth and bubble sizes.

[0042] The software parameter data store 122 includes values ​​associated with a software application, including an ML model, a data size, a software program size, software execution time, memory and power consumption, and other measurable characteristics of the software that can be stored as software parameters.

[0043] The model parameter data store 124 includes values ​​associated with the ML model, including an algorithm type (e.g., an ensemble of decision trees or a neural network), a machine learning process (e.g., a Bayesian optimization process with Gaussian process regression or another general Bayesian approach), or other model parameter values, such as, for ensembles of decision trees, the tree depth and number of leaves, the number of features per tree, the number of trees, the learning rate of the gradient boosting algorithm, various subsampling parameters applied to the dataset, such as subsampling of columns (features), subsampling of rows (or data points) per tree, and feature binning or precision.

[0044] The hardware cost estimation data store 126 includes the cost and constraint functions of the optimization problem that the Gaussian process uses for optimization, such as hardware latency, area, and throughput.

[0045] Hardware device 130 includes various devices that can be configured according to the simulated device configuration. For example, the device may include a graphics processor, which is a common machine learning accelerator and well-suited for parallelism. Some devices can implement fast non-uniform memory access, overcoming various memory limitations that may arise when executing ML models, processes, and fast non-uniform data access.

[0046] In some examples, hardware device 130 includes an aCAM architecture (e.g., for Tree Ensemble Inference). An aCAM is a type of CAM that compares analog inputs to stored intervals and returns a match if all inputs fall within the intervals. An ensemble can be mapped to an aCAM array by placing the input features in the aCAM columns and the root-leaf paths of each tree in the aCAM rows. Thresholds can then be programmed in each aCAM cell, and traversing all trees to find each selected leaf, stored in a separate RAM, can be performed in a single aCAM matching operation for each input point. The aCAM architecture can implement parallel in-memory processes and solve problems with irregular memory access patterns.

[0047] Fig. Figure 2 shows a decision tree ensemble, its association with a hardware device, and a prediction step in accordance with some examples of the system. In example 200, the decision tree ensemble 210, the Fig. to a hardware device (e.g., an aCAM array) and the ensemble prediction 230. The Fig. to a hardware device (e.g., an aCAM array) illustrates an aCAM-based architecture consisting of cores with aCAM arrays with a specific number of rows and columns arranged in stacks and queues. The cores can be interconnected via a configurable H-Tree Network on Chip (NoC). Ensemble prediction 230 illustrates one way in which the values ​​of leaves can be reduced in the inference process. In this phase of exploratory architecture design, the parameters can be freely varied and trade-offs between hardware performance metrics such as area, latency, and throughput can be explored.

[0048] Fig. Figure 3 illustrates a process for generating hardware and software optimization configurations using the hardware and software co-design system according to some examples of the system. In example 300, software parameters 310, hardware parameters 320, model parameters 330, and a hardware cost estimate 340 are provided for optimization 350, which is used to generate output 360 corresponding to the hardware and model optimization configurations.

[0049] Software parameters 310 include software values ​​associated with a software application, including a type of ML model, a data size, a size of a software program, and other measurable characteristics of the software that can be stored as software parameters.

[0050] Hardware parameters 320 include values ​​associated with a hardware device, including the type of hardware device, the size of memory or other device component, the speed of the processor, and other measurable characteristics of the hardware that can be stored as hardware parameters.

[0051] The model parameters 330 include values ​​that define the machine learning model selection (e.g., hyperparameters such as topology or size of a neural network), weights, biases, values ​​that affect the speed and quality of the learning process, or values ​​that affect the algorithm (e.g., learning rate or dataset size). In some examples, an output of the model may correspond to model parameters, e.g., a model accuracy score or other relative degree of accuracy in the output of the software (e.g., ML model, software / model parameters).

[0052] In some examples, hardware cost estimation 340 includes a closed-loop hardware cost model in which metrics associated with various device configurations are measured in a simulated environment that implements the device configuration with hardware and software parameters. In some examples, the hardware cost estimate corresponds to values ​​to be maximized or minimized according to the hardware parameters of a device configuration, including latency (minimize), area (minimize), throughput (maximize), and other hardware cost values.

[0053] Optimization 350 includes determining predicted metrics from the simulation using software parameters 310, hardware parameters 320, model parameters 330, and hardware cost estimate 340. These values ​​may be optimized relative to a threshold. In some examples, optimization may determine a model accuracy score and a hardware cost estimate for each hardware / software / model configuration.

[0054] In some examples, the optimization of the model accuracy score and the hardware cost estimate can occur simultaneously. For example, the machine learning process can apply the device configuration in the simulated environment using the appropriate sets of hardware and software parameters. These parameters can be applied simultaneously and for a single configuration pair. Once the sets of hardware and software parameters are simulated, the process can wait / delay execution of the next iteration of the process until the output is determined. In some examples, the process can wait to determine the result of the previous simulation, which can help select the next set of hardware and software parameters.The selection of parameters in the two parameter spaces (hardware and software) is performed simultaneously and jointly, which is ensured by the expected improvement capture function.

[0055] Output 360 includes the output of the hardware / software / model configurations generated in the simulated environment. The output can be used in a variety of ways. For example, the simulated output can be used to train an ML model, which can apply weights and biases that can tune and optimize any values ​​associated with the simulated output. The output of the ML model can determine a new device configuration that maximizes output for the corresponding hardware and software parameters. For example, the new device configuration can maximize a model accuracy score while minimizing an estimated hardware cost of the new device configuration.In some examples, the output can be used to train an ML model during additional training levels of device configurations or to predict a device configuration that maximizes output for corresponding hardware / software parameters. These and other uses of the initial configurations are illustrative and are not intended to limit the disclosure.

[0056] Fig. Figure 4 shows a process optimization using the hardware and software co-design system in accordance with some examples of the system. The process 400 can be optimized using the Fig. 1 illustrated hardware and software co-design system 102.

[0057] In block 402, a dataset selection is made. A dataset may include any collection of data in any industry or domain. The data may correspond to a specific time interval (e.g., time series data) or another purpose (e.g., training data, validation data, or test data).

[0058] In block 404, the software parameters are determined. The software parameters may correspond to the dataset selection (block 402). During the training process, a model may be trained that adapts to the training dataset according to the software parameters (e.g., weights of the connections between the model's nodes).

[0059] In block 406, the hardware constraints are determined. In some examples, the hardware constraints correspond to the physical constraints of the device (e.g., a specific size, speed, or other value constraint as a physical constraint).

[0060] In block 408, the hardware parameters are determined. The hardware parameters may correspond to the record selection (block 402) that executes instructions to execute the software or other modules / engines. The hardware parameters may include values ​​associated with a hardware device, including the type of hardware device, the size of memory or other device component, the processor speed, and other measurable characteristics of the hardware that may be stored as hardware parameters.

[0061] An illustrative example is the selection of data sets (block 402) used by the system to determine software parameters (block 404) and hardware parameters (block 408). The hardware parameters 408 may be limited by hardware constraints (block 406). The process may sample a data set through data set selection (block 402) to generate software parameters (block 404) and hardware parameters (block 408), as well as N corresponding device configurations using these selected parameters and in accordance with any hardware constraints (block 406).

[0062] At block 420, process regression and active learning may be initiated. For example, process regression and active learning may include the machine learning process of measuring the metrics associated with each device configuration in a simulated environment. During process regression and active learning, predicted metrics may be determined from the simulation, including a model accuracy score and an estimated hardware cost for each device configuration.

[0063] In block 430, a next or second set of software parameters may be selected through the active learning process, and in some examples, a next set of hardware parameters may be selected through the active learning process at the same time in block 440. In other words, the software parameters and the hardware parameters may be selected and tested simultaneously in a combined process. In some examples, block 430 and block 440 may not be executed concurrently.

[0064] In some examples, the next set of software parameters (block 430) and hardware parameters (block 440) may be selected in a sequential order by the active learning process, such that the first device configuration is measured using a first set of software parameters and a first set of hardware parameters, and then a next or second configuration is simulated. In other words, the first set of software parameters may be selected, then a second set of software parameters may be selected sequentially. Similarly, the first set of hardware parameters may be selected, then a second set of hardware parameters may be selected sequentially.

[0065] In block 432, the process may use the next or second set of software parameters to train the ML model with active learning, and in some examples, the process may simultaneously use the next or second set of hardware parameters to train the ML model with active learning in block 442. The process regression and active learning may determine the corresponding model accuracy (block 434) and hardware cost (block 444) of the device configuration corresponding to the simulated software parameters and hardware parameters for the next or second configuration of the software parameters (block 430) and hardware parameters (block 440). In some examples, the input (e.g., software parameters and hardware parameters) and the output (e.g.,Model accuracy rating value for software parameters and hardware cost estimate value for hardware parameters) determined from the process regression and active learning are provided for training the ML model (block 432).

[0066] In block 434, a value for model accuracy is determined, and in some examples, a value for hardware cost is determined concurrently in block 444. For example, if a new set of parameters is provided to the trained ML model, the trained ML model may generate a value for model accuracy (block 434) and also generate the hardware cost (block 444) using a cost estimate of the determined hardware configuration / parameters (block 442) used in the simulated environment to execute the ML model.

[0067] A better model accuracy score can be maximized compared to other model accuracy scores, allowing the software parameters to be optimized and the process to select the most efficient / accurate software configuration. On the hardware side, the hardware cost estimate includes values ​​to be maximized or minimized according to the hardware parameters of a device configuration, including latency (minimize), area (minimize), throughput (maximize), and other hardware cost values. A model accuracy score includes a value that corresponds to a relative degree of accuracy of the software output (e.g., ML model, software / model parameters) when combined with the specific hardware processor configuration (e.g., hardware parameters). A better model accuracy score can be maximized compared to other model accuracy scores.

[0068] Fig. 5 shows a machine learning regression model according to some examples of the system. The machine learning regression model may include, for example, a Gaussian process regression model or other machine learning regression models without detracting from the essence of the disclosure. Equation 500 may be calculated using the Fig. 1 illustrated hardware and software co-design system 102.

[0069] In some examples, the system can determine the initial experiments X by sampling a joint search space of hardware and software parameters. The sample points can be evaluated by first training a decision tree model on one or all of the target datasets and then estimating the hardware range, latency, and throughput using a closed-form hardware cost model. The target datasets can be selected depending on the optimization objective. The machine learning prior model and the posterior model based on the measurements are given in Equation 500.

[0070] In some examples, the machine learning process can blend hardware and software parameters from different search spaces in the same model without altering the underlying fitting algorithm. Using the normalization strategy described in Equation 500, the system can also blend metrics from different hardware and software optimization levels. This approach is independent of the parameter spaces and processes used to determine the metrics and can be applied to hardware or software search spaces regardless of differentiability.

[0071] The normalization strategy described in Equation 500 is an example of normalizing metrics from different domains (such as hardware and software). An example of a software metric is the accuracy of a tree-based ML model, while an example of a hardware metric is the latency of a machine learning model accelerator. The normalization in Equation 500 may involve removing the mean µ̂(X) from X and dividing it by the standard deviation σ̂(X). This may correspond to shifting the mean to "0" and modulating the standard deviation to "1." After different metrics have been normalized in this way, the metrics can be combined / aggregated / blended by ensuring that they have the same or substantially similar weight.

[0072] In some examples, equation 500 describes the prior and posterior of the Gaussian process (GP). The prior f(x) ∼ N(µ 0 ,σ 0 ) can correspond to a normal distribution with mean µ 0 and standard deviation σ 0 are, while the posterior f̂ θ (x) is the prior f(x) weighted by the probability given the parameters, samples, and measurements. This equation can be used for modeling the space, i.e., for learning to predict a metric (e.g., tree-based ML model accuracy or hardware accelerator latency) given a set of parameters.

[0073] Fig. Figure 6 shows a function for capturing the expected improvement (El) in accordance with some examples of the system. Equation 600 can be used with the Fig. 1 illustrated hardware and software co-design system 102.

[0074] In some examples, equation 600 can be implemented to describe the expected improvement (EI) criterion used by the active learning optimization model. At each iteration, the process can determine a new set of input parameters that can improve metrics (e.g., the accuracy of the tree-based ML model and hardware latency). The active learning model can select the set of inputs that maximizes the EI.

[0075] In some examples, the EI criterion consists of two terms. The first weights the normal cumulative distribution function Φ by the difference between the best observed point and the GP mean. This indicates how far the point falls outside the distribution and favors points that are good but lie in an unexplored area. The second term weights the normal probability distribution function ϕ by the variance of the Gaussian process, so that large variance is favored (the explored space is larger).

[0076] In some examples, both distributions are calculated based on the Fig. normalized to the equation shown in Figure 5.

[0077] In some examples, the EI detection function is executed to calculate a next point to be measured among the candidate points. The candidate points can be sampled uniformly in the common space.

[0078] In some examples, Sobol' sequences can be used to draw a uniform sample in a high dimension, thus generating quasi-random sequences with low discrepancy. Sobol' sequences can be performed in base two. Using a base of two can help form successively finer uniform partitions of the unit interval and then reorder the coordinates in each dimension.

[0079] These candidate points do not need to be measured and can be used to guide the optimization. The use of the El criterion to guide an optimization model as part of the active learning process. The EI criterion is given in equation 600. The nearest point x N+1 after N experiments, argmax x E [I (y, x)] is given.

[0080] Fig. 7 shows an example of pseudocode of computer-readable instructions for a level two co-design using a machine learning regression process with active learning according to some examples of the system. The level two co-design may include a joint hardware and software optimization process in which hardware and software are optimized simultaneously. The level one co-design may include a hardware-related optimization process that includes software optimization, hardware evaluation, and hardware optimization, each of which is optimized separately. The level zero co-design may include an independent optimization process in which software and hardware are optimized separately. The computer-readable instructions 700 may be configured with the Fig. 1 illustrated hardware and software co-design system 102.

[0081] The instructions can illustrate part of the overall process for level two co-construction using the linear regression process with active learning. A similar approach can be used for level one co-construction using a linear regression process for each individual room.

[0082] This example illustrates the shared and separate room co-design settings. The co-design approach allowed the investigation of two aspects of the interaction between hardware and software parameters.

[0083] In some examples of separate hardware and software development, the process may train and optimize decision tree models and later optimize the specific hardware for each model to investigate the impact of hardware configuration on model performance (Level 1: Hardware-Aware Optimization). Since the application domain is fixed in this case, we can also optimize a generalized hardware implementation capable of running all trained ML models, averaging hardware performance across all models. In co-design work targeting already developed architectures, such as GPUs, the hardware configuration typically does not limit the application domain.However, since this work deals with a completely new architecture, we have the freedom to explore generalization or specialization based on performance requirements and software domains of interest.

[0084] We then investigate joint hardware and software co-design, where model and hardware parameters are merged into a single, larger search space, and we optimize both sets of parameters in the same optimization loop (stage two: true hardware / software co-design). This joint investigation allows us to uncover trade-offs between hardware and software performance that could be investigated in the separate co-design setting, since the software-side optimizer is unable to directly influence the parameter choice of the second-stage optimizer. We did not investigate hardware generalization in this framework because it is not practical to optimize all models with the same set of software parameters, even though the same set of hardware parameters can generate hardware that supports all models.Generalizing the hardware for co-design of the shared space would require a more complex approach than the one presented in this work.

[0085] The computer-readable instructions 700 may describe a linear regression process (e.g., a Bayesian optimization algorithm using GP regression models) implemented in the GPyTorch package and the El detection function from the BoTorch package, both of which run natively on hardware devices (e.g., GPUs or other processors). The pseudocode of the computer-readable instructions for second-level co-design using an active learning machine learning regression process is described here in conjunction with the pseudocode lines presented in the example.

[0086] In lines 1 and 2, the software space and the hardware space are each introduced with parameters. For example, the software area S is introduced with K SParameters and the hardware area with K H parameters initiated.

[0087] In line 3, the software space and the hardware space are combined to generate an input for the Bayesian optimization function.

[0088] In line 4, the first function is defined to accept four inputs, including the software and hardware space defined in lines 1 and 2, the combined software and hardware space, and a variable "experiments." The first function can be consistent with a Bayesian optimization function.

[0089] In line 5, the first function can sample and measure the seed points.

[0090] In line 6, the first function can define a "while" clause. For example, as long as the variable "i" is less than or equal to the variable "experiments," lines 7-11 are executed.

[0091] In line 7, the first function can fit the Gaussian process model to (X, y).

[0092] In line 8, the first function can select the next experiment based on the expected improvement (EI) criterion.

[0093] In line 9, the first function can update the data.

[0094] In line 10, the first function can move on to the next experiment.

[0095] In line 11, the first function can terminate the “while” clause.

[0096] In line 12, the first function can return values ​​and data associated with the identified variables.

[0097] The first function can end in line 13.

[0098] In line 14, the second function is defined. The second function can define the power of points in a given sample space, including N x (K S + K H ).

[0099] In line 15, the second function can perform various operations related to the range, latency, and throughput of the defined parameters.

[0100] In line 16, the second function can return a normalized value.

[0101] The second function can end in line 17.

[0102] In line 18 the third function is defined, which is called in the first function in line 8.

[0103] In line 19, the third function can return a maximum value from a range of values. The third function can strike a balance between exploitation and exploration.

[0104] The third function can end in line 20.

[0105] In line 21 the fourth function is defined, which is called in the second function in line 16.

[0106] In line 22, the fourth function can determine a normalized X sample space with a sample mean and variance estimates. The value can be returned.

[0107] In line 23 the fourth function can end.

[0108] In line 24 the fifth function is defined, which is called in the first function in line 5.

[0109] In line 25, the fifth function returns a value for matrix multiplication that is uniformly connected to the common sample space.

[0110] In line 26 the fifth function can end.

[0111] Fig. Figure 8 shows optimization metrics corresponding to some system examples. Example 800 provides optimization metrics corresponding to various hardware and software parameters. The ML model can indicate the optimized device configuration by highlighting the combination of hardware and software parameters with a black border around the best compromise for those parameters. In other words, the hardware and software parameters could not be improved for any of the four metrics without "costing" the other three parameters.

[0112] In some examples, the display shows a progressive color scale that translates into dashed boxes on the graph labeled "A," "B," and "C." As experiments are performed, values ​​comparing the model's performance with each perspective are plotted on the graph. When approximately fifty experiments have been performed, the approximate area on each graph is labeled "A." When approximately one hundred experiments are performed, the approximate area on each graph is labeled "B." When approximately two hundred experiments are performed, the approximate area on each map is labeled "C." In summary, the color (and the converted ABC labels) show that the optimization process can end at points that exploit the structure in the parameter space while defining parameters to optimize each one for the overall benefit of the device configuration.

[0113] For example, example 800 may provide different views of the same data set. For example, in block 810, the model may receive a single data set three times, and the optimizer results are provided from three different perspectives, including area ratio (top), latency ratio (middle), and throughput ratio (bottom). In block 820, the model may receive a single data set three times, and the optimizer results are provided from three different perspectives, including area ratio (top), latency ratio (middle), and throughput ratio (bottom). In block 830, the model may receive a single data set three times, and the optimization results are provided from three different perspectives, including area ratio (top), latency ratio (middle), and throughput ratio (bottom).

[0114] The three data sets (shown in blocks 810, 820, 830) can be analyzed from different perspectives. For the area ratio (top), the model can determine how the area ratio affects model performance. For the latency ratio (middle), the model can determine how the latency ratio affects model performance. For the throughput ratio (bottom), the model can determine how the throughput ratio affects model performance. In some examples, area, latency, and throughput can be associated with hardware metrics that are measured simultaneously with software metrics.

[0115] In some examples, values ​​are measured in specific areas of the graph as the system runs more experiments. This can lead to trade-offs in model performance and the measured perspective. For example, as the hardware area becomes smaller, model accuracy may be slightly reduced while simultaneously reducing latency. Model accuracy can be maximized relative to the other values.

[0116] It should be noted that the terms "optimize," "optimal," and the like, as used herein, may be used to mean making or achieving performance as effective or perfect as possible. However, as one skilled in the art reading this document will recognize, perfection cannot always be achieved. Accordingly, these terms may also mean making or achieving performance as good or effective as possible or practical under the circumstances, or making or achieving performance better than that achievable with other settings or parameters.

[0117] Fig. 9 shows an example of a computing component that can be used to implement burst preloading for estimating available bandwidth in accordance with various embodiments. As shown in Fig. 9, the computing component 900 may be, for example, a server computer, a controller, or other similar computing component that can process data. In the example implementation of Fig. 9, the computer component 900 comprises a hardware processor 902 and a machine-readable storage medium 904.

[0118] The hardware processor 902 may be one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for fetching and executing instructions stored in the machine-readable storage medium 904. The hardware processor 902 may fetch, decode, and execute instructions, such as instructions 906-912, to control processes or operations for burst preloading to estimate available bandwidth. Alternatively, or in addition to fetching and executing instructions, the hardware processor 902 may include one or more electronic circuits comprising electronic components for performing the functionality of one or more instructions, such as a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other electronic circuitry.

[0119] A machine-readable storage medium, such as machine-readable storage medium 904, may be any electronic, magnetic, optical, or other physical storage device that contains or stores executable instructions. For example, machine-readable storage medium 904 may be random access memory (RAM), non-volatile RAM (NVRAM), electrically erasable programmable read-only memory (EEPROM), a storage device, an optical disk, and the like. In some embodiments, machine-readable storage medium 904 may be a non-transitory storage medium, where the term "non-transitory" does not include the transitory transmission signals. As described in detail below, machine-readable storage medium 904 may be encoded with executable instructions, e.g., instructions 906-912.

[0120] The hardware processor 902 may execute instruction 906 to obtain a set of hardware parameters and a set of software parameters for configuring a device. For example, the hardware processor 902 may scan a search space of hardware and software configurations to determine software parameters and / or hardware parameters. The software parameters include values ​​associated with a software application, including a type of ML model, a size of data, a size of a software program, and other measurable characteristics of the software that may be stored as software parameters. The hardware parameters include values ​​associated with a hardware device, including the type of hardware device, the size of memory or other device component, the speed of the processor, and other measurable characteristics of the hardware that may be stored as hardware parameters.

[0121] Hardware processor 902 may execute instruction 908 to determine a first device configuration for the device. The first device configuration may be determined from the set of hardware parameters and the set of software parameters. In an illustrative example, the latency, range, and throughput of the first device configuration may be determined / estimated using a closed-form hardware cost model, where the metrics associated with the first device configuration are measured in a simulated environment implementing the device configuration with hardware and software parameters.

[0122] The hardware processor 902 may execute instruction 910 to apply the first set of hardware parameters and the first set of software parameters to a machine learning process. The model may be a machine learning regression method. The machine learning process may be implemented in a simulated or virtual environment using the software and hardware parameters determined from the configuration sample and used to generate metrics associated with the first device configuration. In some examples, the machine learning process is Gaussian process regression with active learning, although other forms of machine learning processes may be implemented without departing from the disclosure.

[0123] Hardware processor 902 may execute instruction 912 to sequentially apply a second set of hardware parameters and a second set of software parameters to the machine learning process to generate a second output. The machine learning process may iteratively and sequentially simulate various device configurations with different hardware and software parameters.

[0124] The metrics predicted from the simulation can include a model accuracy score and a hardware cost estimate for each device configuration. A model accuracy score includes a value corresponding to a relative degree of accuracy of the software output (e.g., ML model, software / model parameters) when combined with the particular hardware processor configuration (e.g., hardware parameters). A better model accuracy score can be maximized compared to other model accuracy scores. On the hardware side, the hardware cost estimate includes values ​​to be maximized or minimized according to the hardware parameters of a device configuration, including latency (minimize), area (minimize), throughput (maximize), and other hardware cost values.

[0125] The model accuracy score and the hardware cost estimate score can be generated simultaneously. For example, the machine learning process can apply the device configuration in the simulated environment using the corresponding sets of hardware and software parameters. These parameters can be applied simultaneously and for a single configuration pair.

[0126] Once the sets of hardware and software parameters are simulated, the process can wait / delay execution of the next iteration of the process until the output is determined. In some examples, the process may wait to determine the result of the previous simulation, which can help in selecting the next set of hardware and software parameters. The selection of parameters in the two parameter spaces (hardware and software) is performed simultaneously and jointly, which is ensured by the expected improvement capture function. The simulated output can be used in several ways. For example, the simulated output can be used to train an ML model, which can apply weights and biases to tune and optimize any values ​​associated with the simulated output.

[0127] In some examples, the output of the ML model may determine a new device configuration that maximizes output for the corresponding hardware and software parameters. For example, the new device configuration may maximize a model accuracy score while minimizing an estimated hardware cost of the new device configuration. In some examples, the output may be used to train an ML model during additional training stages of device configurations or may predict a device configuration that maximizes output for corresponding hardware / software parameters.

[0128] Fig.10 shows a block diagram of an example computer system 1000 in which various embodiments described herein may be implemented. Computer system 1000 includes a bus 1002 or other communication mechanism for conveying information, and one or more hardware processors 1004 connected to bus 1002 for processing information. Hardware processor(s) 1004 may be, for example, one or more general-purpose microprocessors.

[0129] Computer system 1000 also includes main memory 1006, such as random access memory (RAM), a cache, and / or other dynamic storage devices connected to bus 1002, for storing information and instructions to be executed by processor 1004. Main memory 1006 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 1004. When such instructions are stored in storage media accessible to processor 1004, computer system 1000 becomes a special-purpose machine adapted to perform the operations specified in the instructions.

[0130] Computer system 1000 also includes a read-only memory (ROM) 1008 or other static storage device connected to bus 1002 to store static information and instructions for processor 1004. A storage device 1010, such as a magnetic disk, an optical disk, or a USB stick (flash drive), etc., is provided and connected to bus 1002 to store information and instructions.

[0131] Computer system 1000 may be connected to a display 1012, such as a liquid crystal display (LCD) (or a touch screen), via bus 1002 to display information to a computer user. An input device 1014, including alphanumeric and other keys, is coupled to bus 1002 to communicate information and command selections to processor 1004. Another type of user input device is cursor control 1016, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to processor 1004 and controlling cursor movement on display 1012. In some embodiments, the same direction information and command selections as with cursor control may be implemented via receiving touches on a touchscreen without a cursor.

[0132] Computer system 1000 may include a user interface module for implementing a graphical user interface, which may be stored on a mass storage device as executable software code executed by the computing device(s). This and other modules may include, for example, components such as software components, object-oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.

[0133] In general, the term "component", "engine", "system", "database", "data store", and the like, as used herein, may refer to logic embodied in hardware or firmware, or to a collection of software instructions that may have entry and exit points and be written in a programming language such as Java, C, or C++. A software component may be compiled and linked into an executable program, installed in a dynamic link library, or written in an interpreted programming language such as BASIC, Perl, or Python. It is understood that software components may be called by other components or by themselves, and / or may be called in response to detected events or interrupts. Software components configured to run on computing devices may be embodied on a computer-readable medium, such as a hard disk.a compact disc, digital video disc, flash drive, magnetic disk, or other tangible medium, or as a digital download (and may be originally stored in a compressed or installable format that requires installation, decompression, or decryption before execution). Such software code may be stored partially or entirely in a memory of the executing computing device so that it can be executed by the computing device. Software instructions may be embedded in firmware, such as an EPROM. In addition, the hardware components may consist of interconnected logic units, such as gates and flip-flops, and / or programmable units, such as programmable gate arrays or processors.

[0134] Computer system 1000 may implement the techniques described herein using custom hard-wired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, in combination with the computer system, causes or programs computer system 1000 to be a special-purpose machine. According to one embodiment, the techniques described herein are performed by computer system 1000 in response to processor(s) 1004 executing one or more sequences of one or more instructions contained in main memory 1006. Such instructions may be read into main memory 1006 from another storage medium, such as storage device 1010. Execution of the instruction sequences contained in main memory 1006 causes processor(s) 1004 to perform the process steps described herein.In alternative embodiments, hard-wired circuits may be used instead of or in combination with software instructions.

[0135] The term "non-volatile media" and similar terms as used herein refer to any media that stores data and / or instructions that cause a machine to operate in a particular manner. Such non-volatile media may include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 1010. Volatile media includes dynamic memory, such as main memory 1006. Common forms of non-volatile media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or other magnetic data storage media, CD-ROMs, other optical data storage media, physical media with hole patterns, RAM, PROM and EPROM, FLASH EPROM, NVRAM, other memory chips or cartridges, and networked versions thereof.

[0136] Non-transitory media are distinct from transmission media but can be used in conjunction with them. Transmission media are involved in the transfer of information between non-transitory media. Examples of transmission media include coaxial, copper, and fiber optic cables, including the wires that make up bus 1002. Transmission media can also take the form of sound or light waves, such as those generated in radio and infrared data communications.

[0137] Computer system 1000 also includes an interface 1018 connected to bus 1002. Interface 1018 establishes a two-way data communication connection to one or more network connections connected to one or more local area networks. For example, interface 1018 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for establishing a data communication connection to a corresponding type of telephone line. As another example, interface 1018 may be a Local Area Network (LAN) card for establishing a data communication connection to a compatible LAN (or a WAN component for communicating with a WAN). Wireless connections may also be implemented.In each of these implementations, the interface 1018 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams containing various types of information.

[0138] A network connection typically enables data communication over one or more networks to other data devices. For example, a network connection may establish a connection over a local area network to a host computer or to data devices operated by an Internet service provider (ISP). The ISP, in turn, provides data communication services over the worldwide packet data communication network, now commonly referred to as the "Internet." Both the local area network and the Internet use electrical, electromagnetic, or optical signals that carry digital data streams. The signals over the various networks and the signals on the network connection and across interface 1018 that carry the digital data to and from computer system 1000 are examples of transmission media.

[0139] Computer system 1000 can send messages and receive data, including program code, over the network(s), the network connection, and interface 1018. In the Internet example, a server could transmit requested code for an application program over the Internet, the ISP, the local network, and interface 1018.

[0140] The received code may be executed by processor 1004 upon receipt and / or stored in storage device 1010 or other non-volatile memory for later execution.

[0141] Each of the processes, methods, and algorithms described in the preceding sections may be embodied in, and fully or partially automated by, code components executed by one or more computer systems or computer processors comprising computer hardware. The one or more computer systems or computer processors may also operate to support the performance of the corresponding operations in a cloud computing environment or as software as a service (SaaS). The processes and algorithms may be partially or fully implemented in application-specific circuitry. The various features and methods described above may be used independently of one another or combined in various ways.Various combinations and subcombinations are intended to be within the scope of this disclosure, and certain method or process blocks may be omitted in some implementations. The methods and processes described herein are also not limited to any particular order, and the associated blocks or states may be performed in other suitable orders, in parallel, or otherwise. Blocks or states may be added to or removed from the disclosed examples. The execution of certain operations or processes may be distributed among computer systems or computer processors that are not located solely on a single machine, but are distributed across a number of machines.

[0142] A circuit may be implemented in any form of hardware, software, or a combination thereof. For example, one or more processors, controllers, ASICs, PLAs, PALs, CPLDs, FPGAs, logic components, software routines, or other mechanisms may be implemented to form a circuit. In implementation, the various circuits described herein may be implemented as discrete circuits, or the described functions and features may be partitioned, in part or in whole, among one or more circuits. Although various features or functional elements are individually described or claimed as separate circuits, those features and functions may be shared by one or more common circuits, and such description is not intended to assume or imply that separate circuits are required to implement those features or functions.If a circuit is implemented in whole or in part with software, that software may be implemented to operate with a computer or processing system capable of performing the functionality described with respect thereto, such as computer system 1000.

[0143] As used herein, the term "or" can be interpreted in both an inclusive and an exclusive sense. Furthermore, descriptions of resources, acts, or structures in the singular should not be construed to exclude the plural. Conditional terms such as "may," "could," "might," or "may," unless expressly stated otherwise or understood by context, are generally intended to convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not.

[0144] Unless expressly stated otherwise, the terms and expressions used in this document, as well as variations thereof, are not to be interpreted as limiting, but as open-ended. Adjectives such as "conventional," "traditional," "normal," "standard," "known," and terms of similar import are not to be construed as limiting the subject matter described to a particular period of time or to a subject matter available at a particular time, but should be understood to include conventional, traditional, normal, or standard technologies that may be available or known now or at any time in the future.The presence of broader words and phrases such as “one or more,” “at least,” “but not limited to,” or similar phrases in some cases should not be construed as meaning that the narrower case is intended or required in the absence of such broader phrases.

Claims

[1] A process comprising: Receiving a set of hardware parameters and a set of software parameters for configuring a device; determining a first device configuration for the device using a first set of hardware parameters from the set of hardware parameters and a first set of software parameters from the set of software parameters; Applying the first set of hardware parameters and the first set of software parameters to a machine learning process, wherein a first output from the machine learning process comprises a first software model accuracy score for the first set of hardware parameters from the set of hardware parameters and a first hardware cost estimate for the first set of software parameters from the set of software parameters, wherein the first output from the machine learning process simultaneously determines the first software model accuracy score and the first hardware cost estimate for the first device configuration; and sequentially applying a second set of hardware parameters and a second set of software parameters to the machine learning process to generate a second output from the machine learning process. [2] The method of claim 1, further comprising: Training a machine learning (ML) model during a first training level using the first device configuration based on the first set of hardware parameters, the first set of software parameters, and the first output from applying the first set of hardware parameters and the first set of software parameters to the machine learning process; Training the ML model during a second training level using a second device configuration, the second set of hardware parameters, the second set of software parameters, and the second output; and Use the trained ML model to predict a third device configuration that maximizes performance for the corresponding hardware and software parameters. [3] The method of claim 1, wherein the machine learning process is a Bayesian optimization process with a Gaussian process regression. [4] The method of claim 1, wherein the first output comprises a latency, a range, and a throughput of the first device configuration measured in a simulated environment implementing the first device configuration with the first hardware parameters and the first software parameters. [5] The method of claim 1, wherein the first output is generated using a closed-form hardware cost model of the machine learning process. [6] The method of claim 1, wherein the first set of hardware parameters, the first set of software parameters, the second set of hardware parameters, and the second set of software parameters are selected using an active learning process. [7] The method of claim 1, wherein the first set of hardware parameters and the first set of software parameters are returned to the machine learning process to sequentially determine optimization values ​​for different configuration settings. [8] The method of claim 1, further comprising: Stop the determination of device configurations when the output corresponding to the hardware parameters and the software parameters exceeds a specified threshold. [9] A computer system with: a memory; and one or more processors configured to execute machine-readable instructions stored in memory to cause the processor to: to receive a set of hardware parameters and a set of software parameters for configuring a device; determine a first device configuration for the device using a first set of hardware parameters from the set of hardware parameters and a first set of software parameters from the set of software parameters; apply the first set of hardware parameters and the first set of software parameters to a machine learning process, wherein the first output from the machine learning process comprises a first software model accuracy rating value for the first set of hardware parameters from the set of hardware parameters and a first hardware cost estimation value for the first set of software parameters from the set of software parameters, wherein the first output from the machine learning process simultaneously determines the first software model accuracy score and the first hardware cost estimate for the first device configuration; and sequentially apply a second set of hardware parameters and a second set of software parameters to the machine learning process to produce a second output from the machine learning process. [10] The computer system of claim 9, wherein the processor further comprises: train a machine learning (ML) model during a first training level using the first device configuration based on the first set of hardware parameters, the first set of software parameters, and the first output from applying the first set of hardware parameters and the first set of software parameters to the machine learning process; train the ML model during a second training level using a second device configuration, the second set of hardware parameters, the second set of software parameters, and the second output; and use the trained ML model to predict a third device configuration that maximizes performance for the corresponding hardware and software parameters. [11] The computer system of claim 9, wherein the machine learning process is a Bayesian optimization process with Gaussian process regression. [12] The computer system of claim 9, wherein the first output comprises a latency, a range, and a throughput of the first device configuration measured in a simulated environment implementing the first device configuration with the first hardware parameters and the first software parameters. [13] The computer system of claim 9, wherein the first output is generated using a closed-form hardware cost model of the machine learning process. [14] The computer system of claim 9, wherein the first set of hardware parameters, the first set of software parameters, the second set of hardware parameters, and the second set of software parameters are selected using an active learning process. [15] The computer system of claim 9, wherein the first set of hardware parameters and the first set of software parameters are returned to the machine learning process to sequentially determine optimization values ​​for different configuration settings. [16] The computer system of claim 9, wherein the processor further comprises: to stop the determination of device configurations when the output corresponding to the hardware parameters and the software parameters exceeds a specified threshold. [17] A non-transitory, computer-readable storage medium storing a plurality of instructions executable by a processor, the plurality of instructions, when executed by the processor, causing the processor to: to obtain a set of hardware parameters and a set of software parameters for configuring a device; determine a first device configuration for the device using a first set of hardware parameters from the set of hardware parameters and a first set of software parameters from the set of software parameters; apply the first set of hardware parameters and the first set of software parameters to a machine learning process, wherein the first output from the machine learning process comprises a first software model accuracy rating value for the first set of hardware parameters from the set of hardware parameters and a first hardware cost estimation value for the first set of software parameters from the set of software parameters, wherein the first output from the machine learning process simultaneously determines the first software model accuracy score and the first hardware cost estimate for the first device configuration; and sequentially apply a second set of hardware parameters and a second set of software parameters to the machine learning process to produce a second output from the machine learning process. [18] The non-transitory computer-readable storage medium of claim 17, further comprising: Training a machine learning (ML) model during a first training level using the first device configuration based on the first set of hardware parameters, the first set of software parameters, and the first output from applying the first set of hardware parameters and the first set of software parameters to the machine learning process; Training the ML model during a second training level using a second device configuration, the second set of hardware parameters, the second set of software parameters, and the second output; and Use the trained ML model to predict a third device configuration that maximizes performance for the corresponding hardware and software parameters. [19] The non-transitory computer-readable storage medium of claim 17, wherein the machine learning process is a Bayesian optimization process with a Gaussian process regression. [20] The non-transitory computer-readable storage medium of claim 17, wherein the first output comprises a latency, a range, and a throughput of the first device configuration measured in a simulated environment implementing the first device configuration with the first hardware parameters and the first software parameters.