Artificial intelligence (AI) for hardware / software collaborative design of accelerator and machine learning models

Through the hardware and software collaborative design system, the hardware and software configuration of the processor is optimized by the machine learning regression process, the inefficiency problem in the existing technology is solved and efficient and automated AI workload processing is achieved.

CN120069120APending Publication Date: 2025-05-30HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410668969.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-05-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively optimize processor hardware and software configuration, resulting in inefficient efficiency when processing AI workloads of large data sets, and manual adaptation process resources intensive and inefficient.

Method used

The hardware and software configuration is iteratively determined through the collaborative design of the system by hardware and software, forming an optimized system configuration. The system optimizes model accuracy and hardware cost through machine learning regression process, combining hardware and software parameters, and realizes automation and efficiency of device configuration.

Benefits of technology

It realizes improving processor efficiency and accuracy when handling AI workloads, reducing inefficiency problems of resource consumption and manual adaptation, and optimizing the coordinated performance of hardware and software.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069120A_ABST
    Figure CN120069120A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to artificial intelligence (AI) for hardware / software co-design of accelerators and machine learning models. Systems and methods are provided for iteratively co-designing hardware and software elements of a device configuration to optimize it. The process may use a machine learning process that simulates how the device configuration is to execute to predict how the software and hardware parameters are to execute in the device configuration. Corresponding inputs (e.g., software / hardware parameters) and outputs (e.g., model accuracy estimates for software parameters and hardware cost estimates for hardware parameters) determined from a machine learning process may be used for various purposes, including being used to train a machine learning (ML) model to select an optimized device configuration in consideration of various constraints.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Processors are used to execute machine-readable instructions that cause the processors to perform various actions on a computer system. In some examples, the actions implement machine learning models and other high-performance computing operations. The characteristics of a processor can determine the speed, accuracy, and efficiency at which the processor performs these actions. For example, a general-purpose graphics processing unit (GPU) can execute models that other processors can also execute, however, other processors may be slower and may execute the same task inaccurately. In fact, the correct hardware configuration for processing a task can improve the overall efficiency of the system. Brief Description of the Drawings

[0002] The present disclosure will be described in detail with reference to the following drawings. The drawings are provided for illustrative purposes only and depict typical or example embodiments.

[0003] Figure 1 Illustrates a hardware and software co-design system according to some examples of the system.

[0004] Figure 2 Is a decision tree ensemble, its mapping to hardware devices, and prediction steps according to some examples of the system.

[0005] Figure 3 Illustrates a process for using a hardware and software co-design system to generate hardware and software optimized configurations according to some examples of the system.

[0006] Figure 4 Illustrates process optimization using a hardware and software co-design system according to some examples of the system.

[0007] Figure 5 Illustrates a machine learning regression model according to some examples of the system.

[0008] Figure 6 Illustrates an expected improvement acquisition function according to some examples of the system.

[0009] Figure 7 Illustrates the pseudocode of computer-readable instructions for using a machine learning regression process with active learning for secondary co-design according to some examples of the system.

[0010] Figure 8 Illustrates optimization metrics according to some examples.

[0011] Figure 9 Illustrates example computing components that can be used to implement hardware and software co-design according to some examples of the system.

[0012] Figure 10is an example computing component that can be used to implement various features of the embodiments described in this disclosure.

[0013] The figures are not exhaustive and do not limit the disclosure to the exact forms disclosed. Detailed Description

[0014] Modern artificial intelligence (AI) workloads process large datasets, which is not ideal in standard or general-configured hardware processors. Some hardware processors are specifically built as accelerators for accelerating specific functions or workloads. Even when a processor is configured as an accelerator for a specific purpose, the processor / accelerator can handle a wide range of workloads, but rarely achieves the optimal configuration for a specific function or workload. Some administrative users can adapt the hardware processor for ML workloads or other software-based processing, but this is typically a manual process and creating these configurations requires intensive resources. Designing new hardware for specific applications and implementing software that leverages the hardware characteristics are interdependent issues that should be addressed together to maximize improvements. Optimizing the design of both hardware and software components presents many challenges and may require frequent communication between hardware and software experts, and can be time-consuming and inefficient if handled separately.

[0015] Examples of this disclosure iteratively determine hardware and software configurations by co-designing these elements to form an optimized system configuration. The process can determine software parameters and hardware components and can be applied to a general search space (e.g., a finite hardware and software configuration space to be searched) and an optimization target variable (e.g., a characteristic of a software / hardware pair to be maximized in a model result). For example, the system can first sample the search space of hardware and software configurations to determine software parameters, model parameters, and / or hardware parameters. Software or model parameters (used interchangeably) include values associated with a software application, including the type of ML model, the size of the data, the size of the software program, and other measurable characteristics of the software that can be stored as software parameters. Hardware parameters include values associated with a hardware device, including the type of hardware device, the size of the memory or other device components, the speed of the processor, and other measurable characteristics of the hardware that can be stored as hardware parameters.

[0016] When determining the parameters, the system can determine a first device configuration associated with the parameters. For example, the latency, area, and throughput of the first device configuration can be determined / estimated using a closed-form hardware cost model, where metrics associated with the first device configuration are measured in a simulated environment in which the device configuration is implemented using the hardware and software parameters. A model of a machine learning regression process can be implemented in a simulated or virtual environment using the software parameters and hardware parameters determined from the configuration samples and used to generate metrics associated with the first device configuration.

[0017] In some examples, the machine learning process is Gaussian process regression with active learning, but other forms of machine learning processes can be implemented without departing from the present disclosure. Active learning refers to a type of machine learning where the learning algorithm can interactively query an information source to label new data points with specific values (e.g., using or packages). The machine learning process can iteratively and sequentially simulate various device configurations using different hardware and software parameters.

[0018] Metrics predicted from the simulations can include a model accuracy evaluation value and a hardware cost estimate for each device configuration. The model accuracy evaluation value includes a value corresponding to the amount of relative accuracy in the output of the software (e.g., ML model, software / model parameters) when combined with a specific configuration of the hardware processor (e.g., hardware parameters). A better model accuracy evaluation value can be maximized compared to other model accuracy evaluation values. On the hardware side, the hardware cost estimate includes values that are maximized or minimized corresponding to the hardware parameters of the device configuration, including latency (minimized), area (minimized), throughput (maximized), and other hardware cost values.

[0019] The model accuracy evaluation value and the hardware cost estimate can be generated simultaneously. For example, the machine learning process can apply a device configuration in a simulation environment using a corresponding set of hardware and software parameters. These parameters can be applied simultaneously to a single configuration pair. Once the set of hardware and software parameters has been simulated, the process can wait / delay the execution of the next iteration of the process until the output is determined. In some examples, the process can wait to determine the output from a previous simulation, which can help guide the selection of the next set of hardware and software parameters. The parameter selection in the two parameter spaces (hardware and software) is performed simultaneously in a joint manner, guaranteed by the expected improvement acquisition function.

[0020] The simulation output can be used in various ways. For example, the simulation output can be used to train an ML model, which can apply weights and biases that can tune and optimize any value associated with the simulation output. The output from the ML model can determine a new device configuration that maximizes the output corresponding to the hardware and software parameters. For example, the new device configuration can maximize the model accuracy evaluation value and simultaneously minimize the hardware cost estimate for the new device configuration. In some examples, the output can be used to train the ML model during additional levels of training of the device configuration, or can predict a device configuration that maximizes the output for the corresponding hardware / software parameters.

[0021] Provide technical improvements throughout the process. On the hardware side, optimized solutions can deliver low latency, high throughput, and comply with physical limitations (e.g., the area available on a silicon chip) and power constraints. On the software or application side, effective model implementations can utilize separately selected hardware for the software implementation to deliver accurate and fast inference of ML models (e.g., decision tree ensembles). In some examples, an active learning Gaussian process regression model is used to determine where to explore in the search space to test the next available implementation of the hardware and software working together.

[0022] In some examples, the process can also improve machine learning models. For example, the system can help identify fewer decision trees in a decision ensemble while mapping more decision trees to each hardware device configuration. Joint optimization can also reduce the feature precision for software parameters derived in each device configuration by performing feature binning without loss of model accuracy, while optimizing hardware performance and achieving large throughput improvements.

[0023] Figure 1 Illustrated is a hardware and software co - design system according to some examples of the system. In Example 100, illustrated is a hardware and software co - design system 102 communicating with a network 140 and a set of hardware devices 130. The hardware and software co - design system 102 includes a processor 104, a memory 105, and a machine - readable medium 106 for storing computer - readable instructions for performing the various operations discussed herein. The set of hardware devices 130 may include a second separate processor, memory, and software components (not shown) for implementing the device configurations determined by the hardware and software co - design system 102.

[0024] The processor 104 may include a general - purpose or a special - purpose processing engine, such as, for example, a microprocessor, a controller, or other control logic. The processor 104 may be connected to a bus, but any communication medium may be used to facilitate interaction with other components of the hardware and software co - design system 102 or external communication.

[0025] The memory 105 may include random access memory (RAM) or other dynamic memory for storing information and instructions to be executed by the processor 104. The memory 105 may also be used to store temporary variables or other intermediate information during the execution of instructions by the processor 104. The memory 105 may also include read - only memory (“ROM”) or other static storage devices coupled to the bus for storing static information and instructions for the processor 104.

[0026] The machine-readable medium 106 may include one or more interfaces, circuits, and modules for implementing the functions discussed herein. The machine-readable medium 106 may carry one or more sequences of instructions for execution by the one or more instruction processors 104. Such instructions embodied on the machine-readable medium 106 may cause the hardware and software co-design system 102 to perform the features or functions of the disclosed technologies discussed herein. For example, the interfaces, circuits, and modules of the machine-readable medium 106 may include, for example, a data processing module 108, a device configuration module 110, an analog module 112, an artificial intelligence (AI) module 114, and a user interface module 116. Various data storage devices may also be maintained by the hardware and software co-design system 102, including a hardware parameter data storage device 120, a software parameter data storage device 122, a model parameter data storage device 124, and a hardware cost estimate data storage device 126.

[0027] The data processing module 108 is configured to receive a set of hardware parameters and a set of software parameters that are used to configure a device that conforms to each parameter in the set of hardware parameters and each parameter in the set of software parameters. As discussed herein, software or model parameters (used interchangeably) include values associated with a software application, including the type of ML model, the size of the data, the size of the software program, and other measurable characteristics of the software that can be stored as software parameters. Hardware parameters include values associated with a hardware device, including the type of hardware device, the size of the memory or other device components, the speed of the processor, and other measurable characteristics of the hardware that can be stored as hardware parameters.

[0028] In some examples, once the set of hardware parameters and the set of software parameters for configuring the device are received as a data set and selected by the hardware and software co-design system 102, the data processing module 108 may not interact with the parameters further but may rely on various device configurations generated by the device configuration module 110.

[0029] The device configuration module 110 may use a first set of hardware parameters from the set of hardware parameters and a first set of software parameters from the set of software parameters to determine a first device configuration for the device. The device configuration module 110 may also use a first software model accuracy evaluation value and a first hardware cost estimate value from an ML model (implemented by the AI module 114) to determine a second device configuration.

[0030] In device configuration, hardware and software constraints can be considered. For example, on the hardware side, the device configuration module 110 can consider optimization features to deliver low latency, high throughput, and comply with the physical limitations of the hardware (e.g., the available area on a silicon chip) and power constraints. On the software or application side, the device configuration module 110 can consider optimization features to utilize the hardware to deliver accurate and fast inference of the model. The system may require frequent communication between the hardware and software components of the simulated device configuration.

[0031] The device configuration module 110 is also configured to provide the device configuration to the simulation module 112 to evaluate the corresponding hardware parameters and software parameters of the device configuration in a simulated or virtual environment. The simulations can be executed sequentially such that the output from a first simulation can be used in a second simulation.

[0032] Various evaluation techniques can be implemented to determine how the configuration will perform, including using machine learning models such as Gaussian process regression models, and the next set of hardware parameters and software parameters can be selected using active learning methods. Other types of evaluation techniques can be implemented without departing from the essence of the present disclosure, including multi-armed bandit models, global optimization, Pareto optimization, or other probabilistic processes. For example, a machine learning process can determine the output from a particular device configuration and active learning can use a Gaussian process regression model to select the next hardware parameters and software parameters to test (e.g., to determine model accuracy and hardware cost values).

[0033] The system also identifies the next parameters corresponding to a particular device configuration for sampling using active learning. The next parameters and configuration can be provided back to the machine learning process to sequentially determine the optimized values for different configuration settings. When the output corresponding to the hardware parameters and software parameters exceeds a predetermined threshold (e.g., a threshold is determined as a particular optimized value, or a certain improvement amount between the hardware and software configurations increases less than the threshold), the process can stop determining the device configuration.

[0034] The machine learning process can seek to optimize the hardware and software metrics for the optimized device configuration. An array of normalized measurement values “y” corresponding to the parameter sample “X” can be written as:

[0035] where

[0036] Using this formula, the system can seek to optimize various hardware and software metrics. For example, using three hardware metrics and one software metric, the system can minimize the hardware area, hardware latency, and root mean square error (RMSE) performance metric of the model, while maximizing the hardware throughput and model accuracy. The metrics can be combined in a single scalarization, and the system can independently calculate the score (Z-score) for each metric against the normalized metrics. The system can define the objective function f(x) = y as a single-output real-valued function of the hardware and software parameter vectors.

[0037] In some examples, the device configuration module 110 (with the simulation module 112) can find the best trade-off of hardware that can run the corresponding software, e.g., a large set of optimization models. The system can explore the model and the dedicated hardware together to determine the best trade-off.

[0038] The simulation module 112 is configured to determine a simulation output based on the device configuration determined by the device configuration module 110. For example, the simulation module 112 can generate a concurrent system in a virtual environment where the instruction set architecture (ISA), microarchitecture, and memory interact with the programming model and communication system. Some examples can incorporate components of virtualized processors, memories, disks, networks, and software models. The simulator can remain invariant when it accepts various hardware parameters and software parameters as inputs and determines an event type or value as an output. For example, during simulation, the simulation module 112 can determine events associated with computation, communication (e.g., Message Passing Interface (MPI) events), sleep, or memory reads and the corresponding values associated with such events.

[0039] In some examples, the simulation module 112 can apply a first set of hardware parameters and a first set of software parameters to an ML model (implemented by the artificial intelligence (AI) module 114). In addition to the other metrics discussed herein, the output of the model can help determine when and how much time to spend executing the processes for these events.

[0040] In some examples, the simulation module 112 can receive parameters provided in a parallel simulation environment based on MPI events. This process can provide a high level of performance and the ability to view large systems. The model can determine the predicted output for a range of device configurations from processing in memory to a conventional processor connected via a conventional network interface and running MPI.

[0041] In some examples, the simulation module 112 may implement the simulation on an external system or using a toolkit (e.g., the Structural Simulation Toolkit (SST)). As an external system, the simulation can virtualize the software and hardware components of the device configuration, where the ISA, microarchitecture, and memory interact with the programming model and communication system while the impact of the configuration is measured. The measurements can be provided back to the simulation module 112. In this sense, the simulation module 112 can implement a modular design that enables measuring individual system parameters without changing the components of the simulation module 112, or providing an MPI-based parallel simulation environment.

[0042] The artificial intelligence (AI) module 114 is configured to determine an output from a first device configuration by applying a first set of hardware parameters and a first set of software parameters to an ML model. The output from the ML model simultaneously generates a first software model accuracy assessment value and a first hardware cost estimate value for the first device configuration that conforms to the first set of hardware parameters and the first set of software parameters.

[0043] In some examples, the AI module 114 implements tree-based model inference for the co-design process. The inference can determine trade-offs between software and hardware metrics, such as the trade-off between software / model accuracy and hardware area, latency and throughput. The ML model can correspond to a linear regression model or a multi-objective Gaussian process (GP) regression model to incorporate the hardware and software search spaces. The process can use an active learning acquisition function (e.g., the Expected Improvement (EI) criterion) to identify the joint space.

[0044] In some examples, the ML model is a gradient-boosted decision tree ensemble. For example, given an input point x = (x 1 ,..., x K ), where K is the number of input dimensions or features, and a function f(x) = y, where y is a real number or class label, the learning problem involves using data (X, y = f(x ∈ X)) to build a predictor ^fθ that can estimate f(x′) for new points x′ not in X.

[0045] In some examples, a binary decision tree ensemble performs the inference. The binary decision tree ensemble is where P is the number of trees in the ensemble. The inference is performed by comparing the input feature values x 1 ,..., x K with the thresholds in the nodes of each tree.

[0046] In some examples, each node compares a single threshold with a single feature. The features can appear on multiple nodes and the tree can be unbalanced. Any part of the tree prediction can consist of a single path taken by each input point, and the leaves reached by each tree in the ensemble are combined in subsequent reduction steps. This can result in a final regression or classification prediction ŷ = f̂θ(x′), where θ includes the thresholds and parameters of the tree ensemble.

[0047] The AI module 114 can implement the training process of the ML model. The training can include using a gradient-boosted decision tree algorithm to train the tree ensemble model, where each new tree is iteratively fit to the partial residuals of the ensemble. The process can expose many parameters for configuring and constraining the resulting tree ensemble and the training process, such as the maximum number of trees, the learning rate, and implement different subsampling methods.

[0048] In some examples, the ML model can implement feature binning. For example, feature binning can divide continuous or other numerical features into different groups. This can allow the ML model to emphasize important trends in the data by concentrating the feature binning on the tree splits in the decision tree ensemble.

[0049] In some examples, the performance of the tree ensemble model can be measured by the percentage of accurate predictions for classification and the root mean square error (RMSE) for regression tasks.

[0050] The AI module 114 can compare the accuracy value with a threshold or a previous accuracy value. The comparison can include, for example, comparing a second device configuration with a second software model accuracy assessment value. In another example, the comparison can include comparing a second hardware cost estimate with a first software model accuracy estimate and a first hardware cost estimate of a first device configuration. When the second value is greater than the first value, the second device configuration can be provided to the interface (via the user interface module 116).

[0051] The user interface module 116 is configured to provide a device configuration to a display. For example, the device configuration can use a set of hardware parameters from a hardware parameter set and a set of software parameters from a hardware parameter set that optimizes the determined metric. The device configuration can include, for example, a device configuration that is superior to other device configurations in one or more aspects such as low latency, high throughput, a hardware configuration that can fit within a chip area, maintaining power constraints, delivering accurate and fast inference of the ML model, or other metrics described throughout the text.

[0052] The hardware parameter data storage device 120 includes values associated with hardware devices, including the type of hardware device, the size of memory or other device components, the speed of the processor, and other measurable characteristics of the hardware that can be stored as hardware parameters. Exemplary hardware devices may include analog content addressable memories (aCAMs), and the corresponding hardware parameters for an aCAM device may include the number and type of cells in the aCAM, the various organizational structures of the aCAM cells (such as the height and width of a group of aCAM cells), the various organizational structures of the on-chip network (NoC) communication structure between aCAM cells (such as tree branches and depth), various memory buffer sizes (such as buffers at each level of the NoC), and various pipeline parameters, such as pipeline depth and bubble size.

[0053] The software parameter data storage device 122 includes values associated with software applications, including the type of ML model, the size of the data, the size of the software program, the software execution time and memory, and power consumption, and other measurable characteristics of the software that can be stored as software parameters.

[0054] The model parameter data storage device 124 includes values associated with ML models, including algorithm types (e.g., decision trees or neural network ensembles), machine learning processes (e.g., Bayesian optimization processes with Gaussian process regression, or other general Bayesian methods), or other model parameter values such as for decision tree ensembles, tree depth and number of leaves, number of features per tree, number of trees, learning rate of gradient boosting algorithms, various subsampling parameters for operating on the dataset (such as column (feature) subsampling, row (or data point) subsampling for each tree), and feature binning or precision.

[0055] The hardware cost estimation data storage device 126 includes the cost and constraint functions of the optimization problem that the Gaussian process is used to optimize, such as hardware latency, area, and throughput.

[0056] The hardware device 130 includes various devices that can be configured according to an analog device configuration. For example, the device may include a GPU, which can be a general-purpose accelerator for machine learning functions and may be well-suited for parallelism. Some devices can implement fast non-uniform memory access and overcome various memory constraints that may be encountered when running ML models, processes, and fast non-uniform data access.

[0057] In some examples, the hardware device 130 includes an aCAM architecture (e.g., for tree ensemble inference). An aCAM is a type of CAM that compares analog inputs to stored intervals and returns a match when all inputs are within the intervals. By setting the input features on the aCAM columns and each root-to-leaf path of each tree on the aCAM rows, the ensemble can be mapped to an aCAM array. Thresholds can then be programmed in each aCAM cell, and all the trees are traversed to find each selected leaf stored in a separate RAM that can be executed in a single aCAM match operation for each input point. The aCAM architecture can implement parallel memory and overcome the problem of irregular memory access patterns.

[0058] Figure 2 is a decision tree ensemble, its mapping to a hardware device, and a prediction step according to some examples of the system. In Example 200, a decision tree ensemble 210, a mapping 220 to a hardware device (e.g., an aCAM array), and an ensemble prediction 230 are provided. The mapping 220 to the hardware device (e.g., an aCAM array) illustrates an aCAM-based architecture, which consists of a core with an aCAM array having a given number of rows and columns arranged in stacks and queues. The cores can be interconnected by a configurable on-chip H-tree Network-on-Chip (NoC). The ensemble prediction 230 illustrates one way to prune leaf values during the inference process. During this exploratory architecture design phase, the parameters are freely changed and the trade-offs between hardware performance metrics such as area, latency, and throughput are explored.

[0059] Figure 3 illustrates a process for generating hardware and software optimized configurations using a hardware and software co-design system according to some examples of the system. In Example 300, software parameters 310, hardware parameters 320, model parameters 330, and a hardware cost estimate 340 are provided to an optimizer 350, which is used to generate an output 360 corresponding to the hardware and model optimized configurations.

[0060] The software parameters 310 include software values associated with a software application, including the type of ML model, the size of the data, the size of the software program, and other measurable characteristics of the software that can be stored as software parameters.

[0061] The hardware parameters 320 include values associated with a hardware device, including the type of the hardware device, the size of the memory or other device components, the speed of the processor, and other measurable characteristics of the hardware that can be stored as hardware parameters.

[0062] The model parameters 330 include values that define the selection of a machine learning model (e.g., hyperparameters such as the topology or size of a neural network), weights, biases, values that affect the speed and quality of the learning process, or values that affect the algorithm (e.g., learning rate or size of a data sample set). In some examples, the output of the model can correspond to a model parameter, such as a model accuracy assessment value or other relatively accurate metric in the software output (e.g., ML model, software / model parameters).

[0063] In some examples, the hardware cost estimate 340 includes a closed - form hardware cost model where metrics associated with various device configurations are measured in a simulated environment using the hardware and software parameters to implement the device configuration. In some examples, the hardware cost estimate corresponds to a value that is maximized or minimized corresponding to the hardware parameters of the device configuration, including latency (minimized), area (minimized), throughput (maximized), and other hardware cost values.

[0064] Optimization 350 includes determining a predictive metric from simulations using the software parameters 310, hardware parameters 320, model parameters 330, and hardware cost estimate 340. These values can be optimized relative to a threshold. In some examples, the optimization can determine a model accuracy assessment value and a hardware cost estimate value for each hardware / software / model configuration.

[0065] In some examples, the optimization of the model accuracy assessment value and the hardware cost estimate value can be generated simultaneously. For example, a machine learning process can apply a device configuration in a simulated environment using a corresponding set of hardware and software parameters. These parameters can be applied simultaneously to a single configuration pair. Once the set of hardware and software parameters has been simulated, the process can wait / delay the execution of the next iteration of the process until the output is determined. In some examples, the process can wait to determine the output from a previous simulation, and the output from the previous simulation can help guide the selection of the next set of hardware and software parameters. The parameter selection in the two parameter spaces (hardware and software) is performed simultaneously in a joint manner, guaranteed by an expected improvement acquisition function.

[0066] Output 360 includes outputs of the hardware / software / model configuration generated in a simulation environment. The output can be used in various ways. For example, the simulation output can be used to train an ML model, and the ML model can apply weights and biases that can tune and optimize any values associated with the simulation output. The output from the ML model can determine a new device configuration that maximizes the output corresponding to the hardware parameters and software parameters. For example, the new device configuration can maximize the model accuracy evaluation value and at the same time minimize the hardware cost estimate value for the new device configuration. In some examples, the output can be used to train an ML model during additional levels of training of the device configuration, or can predict a device configuration that maximizes the output for the corresponding hardware / software parameters. These and other uses of the output configuration are provided for illustrative purposes and should not limit the present disclosure.

[0067] Figure 4 Illustrates process optimization using a hardware and software co-design system according to some examples of the system. Process 400 can be performed using Figure 1 the hardware and software co-design system 102 shown.

[0068] At block 402, a dataset selection is determined. The dataset can include any data set in any type of industry or scope. The data can correspond to a specific time interval (e.g., time series data) or other purposes (e.g., training data, validation data, or test data).

[0069] At block 404, software parameters are determined. The software parameters can correspond to the dataset selection (block 402). The model can be trained during the training process to fit the training dataset according to the software parameters (e.g., the weights of the connections between the nodes of the model).

[0070] At block 406, hardware constraints are determined. In some examples, the hardware constraints correspond to the physical limitations of the device (e.g., specific dimensions, speed, or other value limitations as physical constraints).

[0071] At block 408, hardware parameters are determined. The hardware parameters can correspond to the dataset selection (block 402) for executing instructions to run the software or other modules / engines. The hardware parameters can include values associated with the hardware device, including the type of the hardware device, the size of the memory or other device components, the speed of the processor, and other measurable characteristics of the hardware that can be stored as hardware parameters.

[0072] As an illustrative example, the system uses dataset selection (block 402) to determine software parameters (block 404) and hardware parameters (block 408). The hardware parameters 408 can be restricted by hardware constraints (block 406). The process can sample the dataset (block 402) through dataset selection to generate software parameters (block 404) and hardware parameters (block 408), and use these selected parameters and generate N corresponding device configurations according to any hardware constraints (block 406).

[0073] At block 420, process regression and active learning can be initiated. For example, process regression and active learning can include machine learning processes where metrics associated with each device configuration are measured in a simulation environment. During process regression and active learning, predicted metrics from the simulation can be determined, including model accuracy evaluation values and hardware cost estimates for each device configuration.

[0074] At block 430, the next or second set of software parameters can be selected by the active learning process, and in some examples, simultaneously at block 440, the next set of hardware parameters can be selected by the active learning process. In other words, the software parameters and the hardware parameters can be selected and tested simultaneously in a combined process. In some examples, blocks 430 and 440 can be implemented non - simultaneously.

[0075] In some examples, the next set of software parameters (block 430) and the set of hardware parameters (block 440) can be selected sequentially by the active learning process, such that the first device configuration is measured using the first set of software parameters and the first set of hardware parameters, and then the next or second configuration is simulated. In other words, the first set of software parameters can be selected, and then sequentially, the second set of software parameters can be selected. Similarly, the first set of hardware parameters can be selected, and then sequentially, the second set of hardware parameters can be selected.

[0076] At block 432, the process can use the next or second set of software parameters to train the ML model using active learning, and in some examples, simultaneously, at block 442, the process can use the next or second set of hardware parameters to train the ML model using active learning. Process regression and active learning can determine the corresponding model accuracy (block 434) and hardware cost (block 444) of the device configuration corresponding to the simulated software parameters and hardware parameters for the next or second configuration of each software parameter (block 430) and hardware parameter (block 440). In some examples, the inputs (e.g., software parameters and hardware parameters) and outputs (e.g., model accuracy evaluation values for software parameters and hardware cost estimates for hardware parameters) determined from process regression and active learning can be provided to train the ML model (block 432).

[0077] At block 434, a model accuracy value is determined and, in some examples, at block 444, a hardware cost value is determined concurrently. For example, when a new set of parameters is provided to a trained ML model, the trained ML model can generate a model accuracy value (block 434) and also generate a hardware cost (block 444) using a cost estimate of a particular hardware configuration / parameters (block 442) used to execute the ML model in a simulation environment.

[0078] Compared to other model accuracy evaluation values, a better model accuracy evaluation value can be maximized, which can optimize software parameters, and the process can select the most efficient / accurate software configuration. On the hardware side, the hardware cost estimate includes values that are maximized or minimized corresponding to the hardware parameters of the device configuration, including latency (minimized), area (minimized), throughput (maximized), and other hardware cost values. The model accuracy evaluation value includes a value corresponding to a relative accuracy metric in the output of the software (e.g., an ML model, software / model parameters) when combined with a particular configuration (e.g., hardware parameters) of the hardware processor. Compared to other model accuracy evaluation values, a better model accuracy evaluation value can be maximized.

[0079] Figure 5 Illustrated is a machine learning regression model according to some examples of the system. For example, without departing from the essence of the present disclosure, the machine learning regression model can include a Gaussian process regression model or other machine learning regression models. Equation 500 can be executed using the Figure 1 hardware and software co - design system 102 shown.

[0080] In some examples, the system can determine an initial experiment X by sampling a joint search space of hardware and software parameters. The sampling points can be evaluated by first training a decision tree model on one or all of the target datasets and then using a closed - form hardware cost model to estimate hardware area, latency, and throughput. The target datasets can be selected according to the optimization objective. The previous machine learning model and the posterior adjustment of the measurements are written as Equation 500.

[0081] In some examples, the machine learning process can mix hardware and software parameters from different search spaces in the same model without modifying the underlying fitting algorithm. Using the normalization strategy described in Equation 500, the system can also mix metrics from different hardware and software optimization phases. The method can be agnostic to the parameter space and process used to obtain the metrics and can be applied to the hardware or software search space regardless of differentiability.

[0082] The normalization strategy described in Equation 500 is an example of normalizing metrics from different domains such as hardware and software. An example of a software metric is the accuracy of a tree-based ML model, while an example of a hardware metric is the latency of a machine learning model accelerator. The normalization in Equation 500 can remove the mean from X and divide by the standard deviation This can correspond to shifting the mean to "0" and modulating the standard deviation to "1". After different metrics are normalized in this way, the metrics can be combined / aggregated / blended by ensuring that these metrics have the same or substantially similar weights.

[0083] In some examples, Equation 500 describes the Gaussian process (GP) model prior and posterior. The prior f(x) ∼ N(μ 0 , σ 0 ) can correspond to a normal distribution, where the mean μ 0 and the standard deviation σ 0 , and the posterior is the prior f(x) weighted by the likelihood of the given parameter samples and measurements. This equation can be used to model the space, or in other words, learn how to predict a metric (e.g., tree-based ML model accuracy or hardware accelerator latency) given a set of parameters.

[0084] Figure 6 Illustrates the expected improvement (EI) acquisition function according to some examples of the system. Equation 600 can be executed using Figure 1 the hardware and software co-design system 102 shown.

[0085] In some examples, Equation 600 can be implemented to describe the expected improvement (EI) criterion used by the active learning optimization model. At each iteration, the process can determine a new set of input parameters that can improve a metric (e.g., such as tree-based ML model accuracy and hardware latency). The active learning model can select the input set that maximizes the EI.

[0086] In some examples, the EI criterion consists of two terms. First, the normal cumulative distribution function Φ is weighted by the difference between the best observed point and the GP mean. This tells how far a point is from the distribution, boosting points that are good but in regions that have not yet been explored. The second term weights the normal probability distribution function by the Gaussian process variance, such that large variances are boosted (the more space has been explored).

[0087] In some examples, both distributions are normalized based on the equation illustrated in Figure 5 .

[0088] In some examples, an EI acquisition function is executed to calculate the next point for measurement among candidate points. The candidate points may be sampled uniformly in the joint space.

[0089] In some examples, a Sobol’ sequence may be used to sample uniformly in high dimensions to generate a quasi-random low-discrepancy sequence. The Sobol’ sequence may be listed in base 2. Using base 2 can help form successively finer uniform partitions of the unit interval and then reorder the coordinates in each dimension.

[0090] These candidate points may not be measured and may be used to guide the optimization. The EI criterion is used to guide the optimization model as part of an active learning process. The EI criterion is written as Equation 600. The next point x after N experiments N+1 is written as argmax x E[I(y,x)].

[0091] Figure 7 Illustrated is example pseudocode of computer-readable instructions for performing second-level co-design using a machine learning regression process with active learning according to some examples of the system. Second-level co-design may include a joint optimization process of hardware and software, where the hardware and software are optimized simultaneously and concurrently. First-level co-design may include a hardware-aware optimization process, which includes software optimization, hardware evaluation, and hardware optimization, each optimized individually. Zero-level co-design may include independent optimization processes where the software and hardware are optimized separately. The computer-readable instructions 700 may be executed using Figure 1 the hardware and software co-design system 102.

[0092] The instructions may illustrate a part of the overall process of performing second-level co-design using a linear regression process with active learning. Similar methods may be used for first-level co-design, using a linear regression process for each individual space.

[0093] In this example, a joint and individual space co-design setup is illustrated. The co-design method enables exploration of both aspects of the interaction between hardware and software parameters.

[0094] In some examples of separate hardware and software co-design, the process can train and optimize a decision tree model and then optimize dedicated hardware for each model, exploring the impact of hardware configuration on model performance (Level 1: Hardware-aware optimization). Since the application domain is fixed in this setting, we can also optimize a general-purpose hardware implementation that can run all the trained ML models, where the hardware performance is averaged across all models. In co-design efforts targeting developed architectures such as GPUs, the hardware configuration usually does not limit the application domain, but since this paper deals with a new architecture developed from scratch, we have the freedom to explore generalization or specialization based on the performance requirements and the software domain of interest.

[0095] We then explore joint hardware and software co-design, where model and hardware parameters are combined in a single larger search space, and we optimize both sets of parameters in the same optimization loop (Level 2: True hardware / software co-design). Since the optimizer on the software side cannot directly influence the parameter selection of the optimizer in the second stage, this joint exploration enables the revelation of the trade-off between hardware and software performance that can be explored in separate co-design settings. We do not explore hardware generalization in this setting because although the same set of hardware parameters can generate hardware that supports all models, it does not make sense to optimize all models using the same set of software parameters. Hardware generalization for co-design of the joint space requires a more complex approach than the method presented in this paper.

[0096] Computer-readable instructions 700 can describe a linear regression process implemented in the GPyTorch package (e.g., a Bayesian optimization algorithm using a GP regression model) and an EI acquisition function from the BoTorch package, both of which run natively on a hardware device (e.g., a GPU or other processor). This paper describes the pseudocode of computer-readable instructions for secondary co-design using a machine learning regression process with active learning in conjunction with the pseudocode lines illustrated in the examples.

[0097] At lines 1 and 2, the software space and the hardware space are initiated using parameters. For example, the software space S is initiated using the Ks parameter and the hardware space is initiated using the K H parameter.

[0098] At line 3, the software space and the hardware space are combined to generate the input to the Bayesian optimization function.

[0099] At line 4, the first function is defined to accept four inputs, including the software space and the hardware space defined at lines 1 and 2, the combined software space and hardware space, and the "experiment" variable. The first function can correspond to the Bayesian optimization function.

[0100] At line 5, the first function can sample and measure the seed points.

[0101] At line 6, the first function can define a "while" clause. For example, when the variable "i" is less than or equal to the "experiments" variable, lines 7 - 11 are executed.

[0102] At line 7, the first function can fit a Gaussian process model to (X, y).

[0103] At line 8, the first function can use the expected improvement (EI) criterion to select the next experiment.

[0104] At line 9, the first function can update the data.

[0105] At line 10, the first function can iterate to the next experiment.

[0106] At line 11, the first function can end the "while" clause.

[0107] At line 12, the first function can return the values and data associated with the determined variables.

[0108] At line 13, the first function can end.

[0109] At line 14, the second function is defined. The second function can define the performance of points in a specific sample space, including N x (K s +K H )).

[0110] At line 15, the second function can perform various operations associated with the area, waiting time, and throughput of the defined parameters.

[0111] At line 16, the second function can return a normalized value.

[0112] At line 17, the second function can end.

[0113] At line 18, the third function is defined and is called in the first function at line 8.

[0114] At line 19, the third function can return the maximum value in the value set. The third function can balance exploration and exploitation.

[0115] At line 20, the third function can end.

[0116] At line 21, the fourth function is defined and is called in the second function at line 16.

[0117] At line 22, the fourth function can use sample mean and variance estimates to determine the normalized X sample space. Values can be returned.

[0118] At line 23, the fourth function can end.

[0119] At line 24, the fifth function is defined and the fifth function is called in the first function at line 5.

[0120] At line 25, the fifth function uniformly returns the matrix multiplication value associated with the sample joint space.

[0121] At line 26, the fifth function can end.

[0122] Figure 8 Illustrated is an optimization metric according to some examples of the system. In example 800, optimization metrics corresponding to various hardware and software parameters are provided. The ML model can display the optimized device configuration by highlighting the combination of hardware and software parameters with a black contour of the best trade-off around these parameters. In other words, the hardware and software parameters cannot be improved on any one of the four metrics without a "cost" on the other three metrics.

[0123] In some examples, the display shows a progressive color scale that is converted in the figure to dashed boxes labeled "A", "B", and "C". When the experiment is executed, the values comparing the model performance with each perspective are mapped on the chart. When approximately 50 experiments are executed, the approximate area is identified as "A" on each chart. When approximately 100 experiments are executed, the approximate area is identified as "B" on each chart. When approximately 200 experiments are executed, the approximate area is identified as "C" on each chart. In summary, the color (and the converted ABC labels) shows that the optimization process can be completed at points exploring the structure in the parameter space and, considering the overall benefit to the device configuration, concurrently define the parameters for optimizing each other.

[0124] For example, example 800 can provide different views of the same data set. For example, at block 810, the model can receive a single data set three times and the results of the optimizer are provided from three different perspectives, including area ratio (top), latency ratio (middle), and throughput ratio (bottom). At block 820, the model can receive a single data set three times and the results of the optimizer are provided from three different perspectives, including area ratio (top), latency ratio (middle), and throughput ratio (bottom). At block 830, the model can receive a single data set three times and the results of the optimizer are provided from three different perspectives, including area ratio (top), latency ratio (middle), and throughput ratio (bottom).

[0125] Three data sets (shown at boxes 810, 820, 830) can be analyzed from different perspectives. For the area ratio (top), the model can determine how the area ratio relates to model performance. For the wait time ratio (middle), the model can determine how the wait time ratio relates to model performance. For the throughput ratio (bottom), the model can determine how the throughput ratio relates to model performance. In some examples, area, wait time, and throughput can be associated with hardware metrics that are measured concurrently with software metrics.

[0126] In some examples, as the system performs more experiments, values are measured at a particular area of the graph. This can identify trade - offs made regarding model performance and the measured perspectives. For example, as the hardware area becomes smaller, the model accuracy may decrease slightly, while also reducing the wait time. The model accuracy can be maximized relative to other values.

[0127] It should be noted that the terms "optimize", "optimal", etc., as used herein, can be used to denote manufacturing or achieving performance that is as efficient or perfect as possible. However, as will be recognized by one of ordinary skill in the art upon reading this document, perfection cannot always be achieved. Thus, these terms can also encompass making performance or achieving performance as good or efficient or practical as possible in a given environment, or making performance or achieving performance better than can be achieved with other settings or parameters.

[0128] Figure 9 Illustrated are example computing components that can be used to implement burst pre - loading for available bandwidth estimation in accordance with various embodiments. Now refer Figure 9 , the computing component 900 can be, for example, a server computer, a controller, or any other similar computing component capable of processing data. In Figure 9 an example implementation of, the computing component 900 includes a hardware processor 902 and a machine - readable storage medium 904.

[0129] The hardware processor 902 can be one or more central processing units (CPUs), semiconductor - based microprocessors, and / or other hardware devices suitable for retrieving and executing instructions stored in the machine - readable storage medium 904. The hardware processor 902 can extract, decode, and execute instructions such as instructions 906 - 912 to control the process or operation for burst pre - loading for available bandwidth estimation. As an alternative or addition to retrieving and executing instructions, the hardware processor 902 can include one or more electronic circuits, where one or more electronic circuits include electronic components for performing the functions of one or more instructions, such as a field - programmable gate array (FPGA), an application - specific integrated circuit (ASIC), or other electronic circuits.

[0130] A machine-readable storage medium, such as machine-readable storage medium 904, can be any electronic, magnetic, optical, or other physical storage device that contains or stores executable instructions. Thus, machine-readable storage medium 904 can be, for example, random access memory (RAM), non-volatile RAM (NVRAM), electrically erasable programmable read-only memory (EEPROM), a storage device, an optical disc, etc. In some embodiments, machine-readable storage medium 904 can be a non-transitory storage medium, where the term "non-transitory" does not include transitory propagated signals. As described in detail below, machine-readable storage medium 904 can be encoded with executable instructions (e.g., instructions 906 - 912).

[0131] Hardware processor 902 can execute instruction 906 to receive a set of hardware parameters and a set of software parameters for configuring the device. For example, hardware processor 902 can sample a search space of hardware and software configurations to determine software parameters and / or hardware parameters. Software parameters include values associated with software applications, including the type of ML model, the size of the data, the size of the software program, and other measurable characteristics of the software that can be stored as software parameters. Hardware parameters include values associated with hardware devices, including the type of hardware device, the size of the memory or other device components, the speed of the processor, and other measurable characteristics of the hardware that can be stored as hardware parameters.

[0132] Hardware processor 902 can execute instruction 908 to determine a first device configuration for the device. The first device configuration can be determined based on the set of hardware parameters and the set of software parameters. In an illustrative example, the latency, area, and throughput of the first device configuration can be determined / estimated using a closed-form hardware cost model, where metrics associated with the first device configuration are measured in a simulated environment in which the device configuration is implemented using the hardware and software parameters.

[0133] Hardware processor 902 can execute instruction 910 to apply the first set of hardware parameters and the first set of software parameters to a machine learning process. The model can be a machine learning regression process. The ML process can be implemented in a simulated or virtual environment using software parameters and hardware parameters determined from configuration samples, and is used to generate metrics associated with the first device configuration. In some examples, the machine learning process is Gaussian process regression with active learning, but other forms of machine learning processes can be implemented without departing from the present disclosure.

[0134] Hardware processor 902 can execute instruction 912 to sequentially apply a second set of hardware parameters and a second set of software parameters to the machine learning process to generate a second output. The machine learning process can iteratively and sequentially simulate various device configurations using different hardware and software parameters.

[0135] Metrics from simulation predictions can include a model accuracy evaluation value and a hardware cost estimate for each device configuration. The model accuracy evaluation value includes a value corresponding to a relative accuracy metric in the output of the software (e.g., an ML model, software / model parameters) when combined with a specific configuration of the hardware processor (e.g., hardware parameters). A better model accuracy evaluation value can be maximized compared to other model accuracy evaluation values. On the hardware side, the hardware cost estimate includes values that are maximized or minimized corresponding to the hardware parameters of the device configuration, including latency (minimized), area (minimized), throughput (maximized), and other hardware cost values.

[0136] The model accuracy evaluation value and the hardware cost estimate can be generated simultaneously. For example, a machine learning process can apply a device configuration in a simulation environment using a corresponding set of hardware parameters and a set of software parameters. These parameters can be applied simultaneously to a single configuration pair.

[0137] Once the hardware parameters and the set of software parameters are simulated, the process can wait / delay the execution of the next iteration of the process until the output is determined. In some examples, the process can wait to determine the output from a previous simulation, and the output from the previous simulation can help guide the selection of the next set of hardware parameters and software parameters. The parameter selection in both parameter spaces (hardware and software) is performed simultaneously in a joint manner, guaranteed by an expected improvement acquisition function. The simulation output can be used in various ways. For example, the simulation output can be used to train an ML model, and the ML model can apply weights and biases that can tune and optimize any value associated with the simulation output.

[0138] In some examples, the output from the ML model can determine a new device configuration that maximizes the output corresponding to the hardware parameters and software parameters. For example, the new device configuration can maximize the model accuracy evaluation value and simultaneously minimize the hardware cost estimate for the new device configuration. In some examples, the output can be used to train the ML model during additional levels of training of the device configuration, or can predict a device configuration that maximizes the output for the corresponding hardware / software parameters.

[0139] Figure 10 A block diagram of an example computer system 1000 in which various embodiments described herein can be implemented is depicted. Computer system 1000 includes a bus 1002 or other communication mechanism for conveying information, and one or more hardware processors 1004 coupled to bus 1002 for processing information. The (multiple) hardware processors 1004 can be, for example, one or more general-purpose microprocessors.

[0140] The computer system 1000 also includes a main memory 1006, such as a random access memory (RAM), a cache, and / or other dynamic storage devices, which is coupled to the bus 1002 to store information and instructions to be executed by the processor 1004. The main memory 1006 can also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by the processor 1004. When such instructions are stored in a storage medium accessible by the processor 1004, the computer system 1000 becomes a special-purpose machine customized to perform the operations specified in the instructions.

[0141] The computer system 1000 also includes a read-only memory (ROM) 1008 or other static storage devices coupled to the bus 1002 for storing static information and instructions for the processor 1004. A storage device 1010, such as a magnetic disk, an optical disk, or a USB thumb drive (flash drive), can be provided and coupled to the bus 1002 to store information and instructions.

[0142] The computer system 1000 can be coupled via the bus 1002 to a display 1012, such as a liquid crystal display (LCD) (or a touch screen), for displaying information to a computer user. An input device 1014, including alphanumeric keys and other keys, is coupled to the bus 1002 for transmitting information and command selections to the processor 1004. Another type of user input device is a cursor control 1016, such as a mouse, a trackball, or cursor direction keys, for transmitting direction information and command selections to the processor 1004 and for controlling the movement of a cursor on the display 1012. In some embodiments, the same direction information and command selections as for the cursor control can be implemented by receiving touches on a touch screen without a cursor.

[0143] The computing system 1000 can include a user interface module for implementing a GUI, which can be stored in a mass storage device as executable software code executed by the (one or more) computing devices. For example, the module and other modules can include components such as software components, object-oriented software components, class components, and task components; processes; functions; attributes; programs; subroutines; program code segments; drivers; firmware; microcode; circuitry; data; databases; data structures; tables; arrays; and variables.

[0144] Generally, terms such as "component", "engine", "system", "database", "data storage device", etc. as used herein can refer to logic embodied in hardware or firmware, or can refer to a collection of software instructions written in a programming language (such as, for example, Java, C, or C++) that may have entry and exit points. Software components can be compiled and linked into an executable program installed in a dynamic link library, or can be written using an interpreted programming language such as, for example, BASIC, Perl, or Python. It should be understood that software components can be called from other components or from themselves, and / or can be called in response to detected events or interrupts. Software components configured to execute on a computing device can be provided on a computer-readable medium, such as a compact disc, digital video disc, flash drive, magnetic disk, or any other tangible medium, or as a digital download (and can initially be stored in a compressed or installable format that requires installation, decompression, or decryption before execution). Such software code can be stored, in whole or in part, on the memory device of the executing computing device for execution by the computing device. Software instructions can be embedded in firmware such as an EPROM. It should also be understood that hardware components can include connected logic units, such as gates and flip-flops, and / or can include programmable units, such as programmable gate arrays or processors.

[0145] Computer system 1000 can implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that is combined with or programmed into the computer system such that the computer system 1000 becomes a special-purpose machine. According to one embodiment, the techniques herein are performed by computer system 1000 in response to one or more sequences of one or more instructions contained in main memory 1006 being executed by (one or more) processors 1004. Such instructions can be read into main memory 1006 from another storage medium, such as storage device 1010. Execution of the instruction sequences contained in main memory 1006 causes (one or more) processors 1004 to perform the process steps described herein. In an alternative embodiment, hardwired circuitry can be used in place of or in combination with software instructions.

[0146] As used herein, the term "non-transitory medium" and like terms refer to any medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such non-transitory media can include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 1010. Volatile media includes dynamic memory, such as main memory 1006. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid state drive devices, magnetic tape, or any other magnetic data storage media, CD-ROM, any other optical data storage media, any physical media with hole patterns, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge, and networked versions thereof.

[0147] Non-transitory media is different from transmission media, but can be used in combination with transmission media. Transmission media participates in transferring information between non-transitory media. For example, transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that comprise bus 1002. Transmission media can also take the form of acoustic or light waves, such as acoustic or light waves generated during radio wave and infrared data communications.

[0148] Computer system 1000 also includes interface 1018 coupled to bus 1002. Interface 1018 provides two-way data communication coupled to one or more network links connected to one or more local networks. For example, interface 1018 can be an Integrated Services Digital Network (ISDN) card, cable modem, satellite modem, or a modem for providing a data communication connection to a corresponding type of telephone line. As another example, interface 1018 can be a Local Area Network (LAN) card to provide a data communication connection to a compatible LAN (or a WAN component communicating with a WAN). A wireless link can also be implemented. In any such implementation, interface 1018 transmits and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.

[0149] Network links typically provide data communication to other data devices via one or more networks. For example, a network link can provide a connection to a host computer or to a data device operated by an Internet Service Provider (ISP) via a local network. The ISP in turn provides data communication services via the global packet data communication network now commonly referred to as the "Internet". Both the local network and the Internet use electrical, electromagnetic, or optical signals that carry digital data streams. Signals through various networks and signals on network links and signals through interface 1018 are example forms of transmission media that carry digital data to and from computer system 1000.

[0150] The computer system 1000 can send messages and receive data, including program code, via (multiple) networks, network links, and interfaces 1018. In the Internet example, the server can transmit the code requested for the application via the Internet, an ISP, a local network, and interface 1018.

[0151] The received code can be executed by the processor 1004 when it is received and / or stored in the storage device 1010 or other non-volatile storage means for later execution.

[0152] Each of the processes, methods, and algorithms described in the foregoing sections can be embodied in code components executed by one or more computer systems or computer processors including computer hardware and be fully or partially automated. One or more computer systems or computer processors can also operate to support the execution of related operations in a "cloud computing" environment or as "software as a service" (SAAS). The processes and algorithms can be implemented partially or fully in dedicated circuitry. The various features and processes described above can be used independently of each other or can be combined in various ways. Different combinations and sub-combinations are intended to fall within the scope of the present disclosure, and certain method or process blocks can be omitted in some implementations. The methods and processes described herein are also not limited to any particular sequence, and the blocks or states associated therewith can be executed in other appropriate sequences or can be executed in parallel or in some other manner. Blocks or states can be added to or removed from the disclosed example embodiments. The execution of certain operations or processes can be distributed among computer systems or computer processors, not only residing within a single machine but also deployed across multiple machines.

[0153] As used herein, a circuit can be implemented using any form of hardware, software, or a combination thereof. For example, one or more processors, controllers, ASICs, PLAs, PALs, CPLDs, FPGAs, logic components, software routines, or other mechanisms can be implemented to constitute a circuit. In an implementation, the various circuits described herein can be implemented as discrete circuits, or the described functions and features can be shared partially or fully among one or more circuits. Even though various features or functional elements can be described or claimed separately as discrete circuits, these features and functions can be shared among one or more common circuits, and such a description should not require or imply the need for separate circuits to implement such features or functions. In cases where a circuit is implemented wholly or partially using software, such software can be implemented to operate using a computing or processing system (such as the computer system 1000) capable of executing the functions described thereof.

[0154] As used herein, the term "or" may be interpreted in an inclusive or exclusive sense. Additionally, the description of a resource, operation, or structure in the singular form should not be construed as excluding the plural. Unless otherwise expressly stated or otherwise understood in the context in which it is used, conditional language such as "can", "could", "might", "may", etc., generally, among other things, is intended to convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not include certain features, elements, and / or steps.

[0155] Unless otherwise expressly stated, the terms and phrases used in this document and their variants should be construed as open-ended rather than limiting. Adjectives such as "conventional", "traditional", "normal", "standard", "known", and terms of similar import should not be construed as limiting the items described to those available at a given time period or given time, but should be understood to encompass conventional, traditional, normal, or standard techniques available or known at any time, present or future. In some instances, the presence of broadening words and phrases such as "one or more", "at least", "but not limited to", or other similar phrases should not be construed as meaning that a narrower case is intended or required in instances where such broadening phrases may be absent.

Claims

1. A method comprising: receiving a hardware parameter set and a software parameter set for configuring a device; determining a first device configuration for the device using a first hardware parameter set from the hardware parameter sets and a first software parameter set from the software parameter sets; applying the first hardware parameter set and the first software parameter set to a machine learning process, wherein a first output from the machine learning process includes a first software model accuracy estimate for the first hardware parameter set from the hardware parameter sets and a first hardware cost estimate for the first software parameter set from the software parameter sets, wherein the first output from the machine learning process simultaneously determines the first software model accuracy estimate and the first hardware cost estimate for the first device configuration; as well as A second set of hardware parameters and a second set of software parameters are sequentially applied to the machine learning process to generate a second output from the machine learning process.

2. The method according to claim 1, further comprising: training a machine learning (ML) model using the first device configuration during a first level of training based on the first hardware parameter set, the first software parameter set, and applying the first hardware parameter set and the first software parameter set to the first output of the machine learning process; training the ML model during a second level of training using a second device configuration, the second set of hardware parameters, the second set of software parameters, and the second output; as well as The trained ML model is used to predict a third device configuration that maximizes an output corresponding to the hardware parameters and the software parameters.

3. The method of claim 1, wherein the machine learning process is a Bayesian optimization process with Gaussian process regression.

4. The method of claim 1, wherein the first output comprises latency, area, and throughput of the first device configuration measured in a simulation environment implementing the first device configuration using the first hardware parameters and the first software parameters.

5. The method of claim 1, wherein the first output is generated using a closed hardware cost model of the machine learning process. 6 . The method of claim 1 , wherein the first hardware parameter set, the first software parameter set, the second hardware parameter set, and the second software parameter set are selected using an active learning process.

7. The method of claim 1, wherein the first set of hardware parameters and the first set of software parameters are provided back to the machine learning process to sequentially determine optimized values ​​for different configuration settings.

8. The method according to claim 1, further comprising: When the output corresponding to the hardware parameter and the software parameter exceeds a predetermined threshold, determining the device configuration is stopped.

9. A computer system comprising: Memory; as well as one or more processors configured to execute machine-readable instructions stored in the memory, such that the processors are configured to: receiving a hardware parameter set and a software parameter set for configuring a device; determining a first device configuration for the device using a first hardware parameter set from the hardware parameter sets and a first software parameter set from the software parameter sets; applying the first hardware parameter set and the first software parameter set to a machine learning process, wherein a first output from the machine learning process includes a first software model accuracy estimate for the first hardware parameter set from the hardware parameter sets and a first hardware cost estimate for the first software parameter set from the software parameter sets, wherein the first output from the machine learning process simultaneously determines the first software model accuracy estimate and the first hardware cost estimate for the first device configuration; as well as A second set of hardware parameters and a second set of software parameters are sequentially applied to the machine learning process to generate a second output from the machine learning process.

10. The computer system of claim 9, wherein the processor is further configured to: training a machine learning (ML) model using the first device configuration during a first level of training based on the first hardware parameter set, the first software parameter set, and applying the first hardware parameter set and the first software parameter set to the first output of the machine learning process; training the ML model during a second level of training using a second device configuration, the second set of hardware parameters, the second set of software parameters, and the second output; as well as The trained ML model is used to predict a third device configuration that maximizes an output corresponding to the hardware parameters and the software parameters.

11. The computer system of claim 9, wherein the machine learning process is a Bayesian optimization process with Gaussian process regression.

12. The computer system of claim 9, wherein the first output comprises latency, area, and throughput of the first device configuration measured in a simulation environment implementing the first device configuration using the first hardware parameters and the first software parameters.

13. The computer system of claim 9, wherein the first output is generated using a closed hardware cost model of the machine learning process.

14. The computer system of claim 9, wherein the first hardware parameter set, the first software parameter set, the second hardware parameter set, and the second software parameter set are selected using an active learning process.

15. The computer system of claim 9, wherein the first set of hardware parameters and the first set of software parameters are provided back to the machine learning process to sequentially determine optimized values ​​for different configuration settings.

16. The computer system of claim 9, wherein the processor is further configured to: When the output corresponding to the hardware parameter and the software parameter exceeds a predetermined threshold, determining the device configuration is stopped.

17. A non-transitory computer-readable storage medium storing a plurality of instructions executable by a processor, the plurality of instructions, when executed by the processor, causing the processor to: receiving a hardware parameter set and a software parameter set for configuring a device; determining a first device configuration for the device using a first hardware parameter set from the hardware parameter sets and a first software parameter set from the software parameter sets; applying the first hardware parameter set and the first software parameter set to a machine learning process, wherein a first output from the machine learning process includes a first software model accuracy estimate for the first hardware parameter set from the hardware parameter sets and a first hardware cost estimate for the first software parameter set from the software parameter sets, wherein the first output from the machine learning process simultaneously determines the first software model accuracy estimate and the first hardware cost estimate for the first device configuration; as well as A second set of hardware parameters and a second set of software parameters are sequentially applied to the machine learning process to generate a second output from the machine learning process.

18. The non-transitory computer-readable storage medium of claim 17, further comprising: training a machine learning (ML) model using the first device configuration during a first level of training based on the first hardware parameter set, the first software parameter set, and applying the first hardware parameter set and the first software parameter set to the first output of the machine learning process; training the ML model during a second level of training using a second device configuration, the second set of hardware parameters, the second set of software parameters, and the second output; as well as The trained ML model is used to predict a third device configuration that maximizes an output corresponding to the hardware parameters and the software parameters.

19. The non-transitory computer-readable storage medium of claim 17, wherein the machine learning process is a Bayesian optimization process with Gaussian process regression.

20. The non-transitory computer-readable storage medium of claim 17, wherein the first output comprises latency, area, and throughput of the first device configuration measured in a simulation environment implementing the first device configuration using the first hardware parameters and the first software parameters.