A Hardware-Software Coordination Method and Device for Deep Learning Accelerators

Through the deep learning accelerator software and hardware collaboration method, through multiple iterations, the problem of low model deployment efficiency in the existing technology is solved, and the efficient utilization of accelerator hardware resources and the improvement of model computing efficiency is achieved.

CN119294274BActive Publication Date: 2025-06-20ZHEJIANG LAB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411832029.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-06-20
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

The prior art cannot effectively improve the running speed and performance of the model in terms of model deployment of deep learning accelerators, especially in the synchronous operation of multi-models and complex AI algorithms, which is difficult to meet the computing needs of modern AI workloads.

Method used

The deep learning accelerator software and hardware collaboration method is adopted to obtain the model parameters of the target model and accelerator area constraints, and the initial hardware parameter configuration sample is determined, and the candidate model operation information and hardware parameter configuration are optimized through multiple iterations of the software optimizer and the hardware optimizer until the preset iteration conditions are met, and the target model operation mode and target hardware parameter configuration are obtained.

Benefits of technology

The optimal accelerator specification parameters for a given model set are realized, which improves the accelerator hardware resource utilization and model computing efficiency, and ensures that the target device has the best performance and efficiency when executing the computing tasks of the target model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119294274B_ABST
    Figure CN119294274B_ABST
Patent Text Reader

Abstract

This specification discloses a software-hardware co - optimization method and device for a deep learning accelerator. In this method, after a software optimizer configures samples based on initial hardware parameters and determines the running information of a candidate model, a hardware optimizer determines the initial hardware parameter configuration samples for the next iteration based on the task efficiency characterization value corresponding to the candidate model running information. If it is monitored that the deviation between the task efficiency characterization values corresponding to the initial hardware parameter configuration samples in the previous and next rounds after reaching the preset number of rounds is less than the preset deviation, then the initial hardware parameter configuration samples obtained when meeting the preset iteration conditions are used as the target hardware parameter configuration, and the candidate model running mode corresponding to the candidate model running information determined by the software optimizer based on the target hardware parameter configuration is used as the target model running mode. Through multiple rounds of iteration, the software and hardware are continuously co - configured to achieve the optimal accelerator specification parameters for a given model set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the fields of computer technology and artificial intelligence, and particularly to a method and device for software-hardware cooperation of a deep learning accelerator. Background Art

[0002] A deep learning accelerator is a type of hardware specifically designed to accelerate deep learning computations. With the rapid development of deep learning technology, the amount of computation required to process large amounts of data and complex models is also increasing continuously. Therefore, deep learning accelerators play an important role in improving computational performance and efficiency.

[0003] A deep learning accelerator is a type of AI chip used to accelerate the computational process in deep learning and artificial intelligence tasks.

[0004] Currently, in the design of artificial intelligence (AI) chips, a software-hardware integration design method is usually adopted to make full use of the computational characteristics and memory access characteristics of AI models to achieve more efficient computation and better performance.

[0005] In the design of AI chips, it is necessary to carefully configure hardware resources such as computing, storage, and communication, and also ensure that when an AI model runs on the chip, it has good performance and fast computing speed. Therefore, when deploying a model currently, corresponding deployment solutions need to be designed at the hardware and software levels. However, the current model deployment solution strategy cannot effectively improve the running speed and performance of the model.

[0006] In addition, with multiple models running synchronously, the success of AI algorithms in a single task (such as hand tracking, depth estimation, speech recognition) has led to the emergence of multi-model AI workloads. For the efficient design of AI accelerators, in order to meet the growing computational needs and changing application requirements of modern AI workloads, an efficient dedicated chip design method is required.

[0007] Facing the above problems, how to deploy the model more accurately is an urgent problem to be solved. Summary of the Invention

[0008] Embodiments of this specification provide a method and device for software-hardware cooperation of a deep learning accelerator to partially solve the problems existing in the above-mentioned prior art.

[0009] Embodiments of this specification adopt the following technical solutions:

[0010] A method for software-hardware cooperation of a deep learning accelerator provided in this specification includes:

[0011] Obtain the model parameters of the target model and the accelerator area constraint, where the accelerator area constraint is used to characterize the hardware configuration budget allocated by the accelerator to the target model;

[0012] Determine an initial hardware parameter configuration sample according to the hardware parameter configuration included in the accelerator area constraint;

[0013] Input the hardware parameters corresponding to the initial hardware parameter configuration sample and the model parameters into a preset software optimizer, so that the software optimizer determines each initial model running information in the given hardware running environment corresponding to the initial hardware parameter configuration sample, and determines candidate model running information based on the each initial model running information;

[0014] Input the candidate model running information into a preset analysis model, so as to predict, through the analysis model, the task efficiency characterization value obtained when performing the computing task of the target model in the hardware running environment corresponding to the candidate model running information according to the candidate model running mode corresponding to the candidate model running information. According to the task efficiency characterization value corresponding to the candidate model running mode, determine the initial hardware parameter configuration sample for the next iteration under the condition of the given candidate model running mode, and send the initial hardware parameter configuration sample for the next iteration to the software optimizer, so that the software optimizer determines the candidate model running mode corresponding to the next iteration according to the initial hardware parameter configuration sample for the next iteration and the model parameters, until the preset iteration condition is met, to obtain the target model running mode and the target hardware parameter configuration;

[0015] Configure the target hardware running environment corresponding to the target hardware parameter configuration on the target device, and deploy the target model on the target device based on the target hardware running environment, so as to execute the computing task corresponding to the target model on the target device according to the target model running mode.

[0016] Optionally, determining candidate model running information based on the each initial model running information specifically includes:

[0017] Sample from the each initial model running information to obtain sampled model running information;

[0018] Input the sampled model running information and the initial hardware parameter configuration sample into a preset evaluator, so as to predict, through the evaluator, the task efficiency characterization value obtained when performing the computing task of the target model in the hardware running environment corresponding to the initial hardware parameter configuration sample according to the sampled model running information, and use it as the sampled characterization value;

[0019] Input the sampling characterization value, the sampling model operation information, and the initial hardware parameter configuration sample into the software optimizer, so as to fit the sampling characterization value and the sampling model operation information through the software optimizer in the hardware operation environment corresponding to the initial hardware parameter configuration sample, and obtain a first optimization model;

[0020] Based on the first optimization model and the initial model operation information of each, determine the candidate model operation information.

[0021] Optionally, according to the task efficiency characterization value corresponding to the candidate model operation mode, determine the initial hardware parameter configuration sample for the next iteration under the condition of given the candidate model operation mode, specifically including:

[0022] Input the task efficiency characterization value corresponding to the candidate model operation mode, the candidate model operation mode, and the initial hardware parameter configuration sample into a preset hardware optimizer, so as to determine the initial hardware parameter configuration sample for the next iteration through the hardware optimizer under the condition of given the candidate model operation mode.

[0023] Optionally, determine the initial hardware parameter configuration sample for the next iteration through the hardware optimizer under the condition of given the candidate model operation mode, specifically including:

[0024] Based on the candidate model operation mode, fit the task efficiency characterization value corresponding to the candidate model operation mode and the initial hardware parameter configuration sample through the hardware optimizer to obtain a second optimization model;

[0025] Based on the hardware parameter configurations required when the target model executes tasks, determine each basic hardware parameter configuration sample under the condition of given the candidate model operation mode through the second optimization model;

[0026] Input each basic hardware parameter configuration sample into a preset evaluator, so as to predict the task efficiency characterization value obtained when executing the arithmetic task of the target model in the hardware operation environment corresponding to each basic hardware parameter configuration sample through the evaluator, and use it as the task efficiency characterization value corresponding to each basic hardware parameter configuration sample;

[0027] Determine the initial hardware parameter configuration sample for the next iteration according to the task efficiency characterization value corresponding to each basic hardware parameter configuration sample.

[0028] Optionally, until the preset iteration condition is met, obtain the target model operation mode and the target hardware parameter configuration, specifically including;

[0029] If it is monitored that the deviation between the task efficiency characterization values corresponding to the initial hardware parameter configuration samples in the previous and next rounds after reaching the preset number of rounds is less than the preset deviation, it is determined that the preset iteration condition is met, and the initial hardware parameter configuration sample obtained when the preset iteration condition is met is used as the target hardware parameter configuration, and the candidate model operation mode determined by the software optimizer based on the target hardware parameter configuration is used as the target model operation mode.

[0030] Optionally, the target model includes at least one model for performing arithmetic tasks.

[0031] A deep learning accelerator software and hardware co - operating device provided in this specification includes:

[0032] An acquisition module, configured to acquire model parameters of a target model and an accelerator area constraint, where the accelerator area constraint is used to characterize the hardware configuration budget allocated by the accelerator to the target model;

[0033] A first determination module, configured to determine an initial hardware parameter configuration sample according to the hardware parameter configuration included in the accelerator area constraint;

[0034] A second determination module, configured to input the hardware parameters corresponding to the initial hardware parameter configuration sample and the model parameters into a preset software optimizer, so that the software optimizer determines each initial model operation information in the hardware operation environment corresponding to the initial hardware parameter configuration policy, and determines candidate model operation information based on the each initial model operation information;

[0035] An analysis module, configured to input the candidate model operation information into a preset analysis model, so as to predict, through the analysis model, the task efficiency characterization value obtained when performing the arithmetic task of the target model in the hardware operation environment corresponding to the candidate model operation information corresponding to the candidate model operation information, determine the initial hardware parameter configuration sample for the next iteration under the condition of the candidate model operation mode corresponding to the candidate model operation mode, and send the initial hardware parameter configuration sample for the next iteration to the software optimizer, so that the software optimizer determines the candidate model operation mode corresponding to the next iteration according to the initial hardware parameter configuration sample for the next iteration and the model parameters, until the preset iteration condition is met, and obtain the target model operation mode and the target hardware parameter configuration;

[0036] An operating module, configured to configure a target hardware operating environment corresponding to the target hardware parameter configuration sample on a target device, and based on the target hardware operating environment, deploy the target model on the target device, so as to execute an arithmetic task corresponding to the target model on the target device according to the running mode of the target model.

[0037] Optionally, the second determination module is specifically configured to sample from the initial model running information to obtain sampled model running information; input the sampled model running information and the initial hardware parameter configuration sample into a preset evaluator, so as to predict, through the evaluator, a task efficiency characterization value obtained when executing the arithmetic task of the target model in the hardware operating environment corresponding to the initial hardware parameter configuration sample according to the sampled model running information, as a sampled characterization value; input the sampled characterization value, the sampled model running information, and the initial hardware parameter configuration sample into the software optimizer, so as to fit the sampled characterization value and the sampled model running information through the software optimizer in the hardware operating environment corresponding to the initial hardware parameter configuration sample to obtain a first optimized model; and determine candidate model running information based on the initial model running information through the first optimized model.

[0038] A computer-readable storage medium provided in this specification, where the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned deep learning accelerator software and hardware cooperation method is implemented.

[0039] An electronic device provided in this specification includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned deep learning accelerator software and hardware cooperation method is implemented.

[0040] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects:

[0041] In the embodiments of this specification, in each round of iteration, through a software optimizer, based on the initial hardware parameter configuration samples of each round and based on the operation information of each initial model, after determining the candidate model operation information, through a hardware optimizer, based on the task efficiency characterization value corresponding to the candidate model operation mode corresponding to the candidate model operation mode, under the condition of a given candidate model operation mode, determine the initial hardware parameter configuration samples for the next round of iteration. If it is monitored that the deviation between the task efficiency characterization values corresponding to the initial hardware parameter configuration samples of the previous and next rounds after reaching the preset number of rounds is less than the preset deviation, then use the initial hardware parameter configuration samples obtained when meeting the preset iteration conditions as the target hardware parameter configuration, and use the candidate model operation mode determined by the software optimizer based on the target hardware parameter configuration as the target model operation mode. Furthermore, configure the target hardware operating environment on the target device according to the target hardware parameter configuration, and deploy the target model on the target device according to the target model operation mode to achieve the optimal accelerator specification parameters for a given model set.

[0042] In this method, through multiple rounds of iteration, the model and accelerator hardware are continuously co-designed to obtain the best model deployment method and the corresponding hardware parameter configuration, so as to achieve the optimal accelerator specification parameters for a given model set, and to execute arithmetic tasks on the accelerator more efficiently, greatly improving the utilization rate of accelerator hardware resources and the arithmetic efficiency of the model. Brief Description of the Drawings

[0043] The drawings described herein are used to provide a further understanding of this specification and form a part of this specification. The illustrative embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:

[0044] Figure 1 It is a schematic flowchart of a method for collaborative design of software and hardware of a deep learning accelerator provided by an embodiment of this specification;

[0045] Figure 2 It is a schematic diagram of the system architecture of a method for collaborative design of software and hardware of a deep learning accelerator provided by an embodiment of this specification;

[0046] Figure 3 It is a schematic diagram of the structure of a device for collaborative design of software and hardware of a deep learning accelerator provided by an embodiment of this specification;

[0047] Figure 4 It is a schematic diagram of the structure of an electronic device provided by an embodiment of this specification. Detailed Embodiments

[0048] To make the objectives, technical solutions, and advantages of this specification clearer, the following will clearly and completely describe the technical solutions of this specification in combination with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.

[0049] The following will, with reference to the drawings, elaborate on the technical solutions provided by each embodiment of this specification.

[0050] Figure 1 The following is a schematic flowchart of a software and hardware co - operation method for a deep - learning accelerator provided by an embodiment of this specification, including:

[0051] S100: Obtain the model parameters of the target model and the accelerator area constraint, where the accelerator area constraint is used to characterize the hardware configuration budget allocated by the accelerator to the target model.

[0052] A deep - learning accelerator is an artificial - intelligence chip used to accelerate the computing process in deep - learning and artificial - intelligence tasks. By allocating the hardware resources of the artificial - intelligence chip, the computing efficiency can be improved.

[0053] In the design of artificial - intelligence (AI) chips, a software - hardware integration design method is usually adopted to make full use of the computing characteristics and memory - access characteristics of AI models, so as to achieve more efficient computing and better performance.

[0054] During the design process, it is necessary to carefully configure hardware resources such as computing, storage, and communication, and also ensure that after the AI model is deployed on the chip using this design method, it has good model performance and fast computing speed.

[0055] Currently, using a spatial search algorithm to conduct co - exploration in software and hardware involves two aspects: design - space search and mapping - space search. Design - space search targets the hardware aspect, and mapping - space search mainly focuses on how to better map the computing method of the AI model algorithm to the hardware resources to execute the computing task. The spatial search algorithm can find the optimal design scheme for the performance of the AI chip within a reasonable time. For example, random search, grid search, and evolutionary search, etc. Although relatively reasonable AI - chip design schemes can be obtained, they can only obtain relatively reasonable design schemes when designing AI chips that meet the requirements, which has certain limitations.

[0056] In view of the above problems, although the prior art can also deploy the model on the deep learning accelerator, it still causes waste of hardware resources of the deep learning accelerator and low model operation efficiency. To solve the above problems, in the embodiments of this specification, in each round of iteration, through a software optimizer, based on the initial hardware parameter configuration samples in each round and the initial model operation information, after determining the candidate model operation information, through a hardware optimizer, based on the task efficiency characterization value corresponding to the candidate model operation mode corresponding to the candidate model operation information, under the condition of a given candidate model operation mode, determine the initial hardware parameter configuration samples for the next round of iteration. If it is monitored that the deviation between the task efficiency characterization values corresponding to the initial hardware parameter configuration samples in the previous and next rounds after reaching the preset number of rounds is less than the preset deviation, then use the initial hardware parameter configuration samples obtained when meeting the preset iteration conditions as the target hardware parameter configuration, and use the candidate model operation mode determined by the software optimizer based on the target hardware parameter configuration as the target model operation mode. Furthermore, configure the target hardware operating environment on the target device according to the target hardware parameter configuration, and deploy the target model on the target device according to the target model operation mode, so that the target device can execute the operation tasks corresponding to the target model more efficiently.

[0057] In this method, through multiple rounds of iteration, the model and the accelerator hardware are continuously co-designed to obtain the best model deployment method and the corresponding hardware parameter configuration, so as to achieve the optimal accelerator specification parameters for a given model set, and execute the operation tasks on the deep learning accelerator more efficiently, greatly improving the utilization rate of accelerator hardware resources and the operation efficiency of the model.

[0058] Next, the target device needs to first obtain the model parameters of the target model and the accelerator area constraint, where the accelerator area constraint is used to characterize the hardware configuration budget allocated by the accelerator to the target model.

[0059] In the embodiments of this specification, after the target device obtains the model parameters of the target model, it will determine the subsequent initial hardware parameter configuration samples according to the model parameters of the target model and the hardware parameter configuration included in the accelerator area constraint. The target device mentioned in this specification can be a terminal device such as a desktop computer or a laptop computer, or a server, or a dedicated device specifically used for training and deploying models.

[0060] Among them, the accelerator area constraint mentioned here is used to characterize the hardware configuration budget allocated by the accelerator to the target model, and the hardware configuration budget includes: Processing Element (PE): responsible for performing basic arithmetic operations in the deep learning model, such as convolution and rectangular array methods, etc.; on-chip network bandwidth: used to represent the efficiency and speed of data transmission between components inside the accelerator; AI chip computing precision: used to represent the computing precision when the accelerator performs operations; on-chip cache size: used to represent the size of the buffer for storing data inside the accelerator, etc.

[0061] Each part included in the above-mentioned hardware configuration budget can actually be reflected by specific information. For example, the on-chip network bandwidth and AI chip computing precision, etc., can be reflected by specific data or information.

[0062] It should be noted that when the target device deploys the target model, the target model includes at least one model for performing computing tasks. That is to say, the target device can perform co-design in terms of software and hardware for multiple models to achieve the application scenario where one chip supports multiple different models simultaneously.

[0063] In this specification, the model can have various forms. For example, it can be a large language model and a generative AI model. The large language model can perform reasoning based on the user input text to output the content required by the user. The generative AI model mainly synthesizes the content required by the user according to the existing database based on the user's needs. For another example, the model can also be an image recognition model, which can recognize the target objects contained in the input image.

[0064] Therefore, the accelerator area constraint in this specification can actually provide various alternative hardware parameter configurations in the accelerator, and these hardware parameter configurations actually determine which hardware resources can be used in the subsequent configuration samples.

[0065] S102: Determine an initial hardware parameter configuration sample according to the hardware parameter configuration included in the accelerator area constraint.

[0066] In the embodiment of this specification, after the terminal device obtains the hardware parameter configurations required for the target model to perform tasks, it can select computing components, storage components, and communication components based on the model parameters of the target model on the basis that the target device can perform the computing tasks of the target model, and then obtain a preliminary hardware parameter configuration sample, that is, an initial hardware parameter configuration sample.

[0067] Among them, the initial hardware parameter configuration sample mentioned here is mainly used to reflect how much hardware resources in the hardware configuration budget of the accelerator area constraint are used to execute the computing tasks of the target model. For example, in an initial hardware parameter configuration sample, it is specified that 200 PE execution units can be selected, the on-chip network bandwidth is 100MB, the AI chip computing precision is int8, and the on-chip cache size is 180KB.

[0068] By selecting the hardware parameter configuration included in the accelerator area constraint, the requirement that the target model can execute computing tasks is met, and subsequent hardware parameter adjustments are made on the initial hardware parameter configuration sample, so as to obtain the target hardware parameter configuration, that is, the hardware resources of the accelerator when the computing efficiency of the target model can reach the best.

[0069] S104: Input the hardware parameters corresponding to the initial hardware parameter configuration sample and the model parameters into a preset software optimizer, so that the software optimizer determines each initial model running information in the given hardware running environment corresponding to the initial hardware parameter configuration sample, and determines candidate model running information based on the each initial model running information.

[0070] In this specification, before the target device inputs the obtained initial hardware parameter configuration sample and the model parameters of the target model into the preset software optimizer, various mapping relationships between the computing tasks executed by the target model and the hardware resources can be obtained according to the initial hardware parameter configuration sample, that is, the hardware resources. That is, various initial model running information. Furthermore, the software optimizer obtains candidate model running information according to various initial model runs and the initial hardware parameter configuration sample, and through the evaluation results after the target model executes computing tasks according to various model running information.

[0071] Among them, the initial model running information mentioned here is different from the initial hardware parameter configuration sample mentioned above. The initial model running information is mainly reflected in the execution method of the computing tasks executed by the model at the software level. That is, the target device can obtain various initial model running methods according to various mapping relationships between the loop operations included in the computing tasks of the target model and the hardware resources.

[0072] It should be noted that there are various mapping relationships between the loop operations included in the computing tasks executed by the target model and the hardware resources. Here, the mapping relationship means that each layer of the loop operations included in the computing tasks executed by the target model can be grouped for operation. For example, a loop contains a total of 10 layers of loops. Each layer of the loop can be configured with hardware resources, or the 10 layers of loops can be divided into 5 groups, and each group is configured with accelerator hardware resources, etc.

[0073] According to different mapping methods, multiple implementation methods of model operation tasks can be obtained. Executing the operation tasks of the target model according to each implementation method will result in different operation efficiencies. By continuously adjusting these implementation methods under the conditions of the initial hardware parameter configuration samples, candidate model operation information is obtained, that is, the candidate model operation methods corresponding to the candidate model operation information.

[0074] Specifically, the target device first samples various initial model operation information. After obtaining the sampled model operation information, it inputs the sampled model operation information and the initial hardware parameter configuration samples into a preset evaluator. Through the preset evaluator, the corresponding task efficiency evaluation index after the target model executes the operation tasks in the hardware environment corresponding to the initial hardware parameter configuration samples according to the sampled model operation information is obtained as the sampled characterization value. Furthermore, the sampled characterization value, the sampled model operation information, and the initial hardware parameter configuration samples are input into a software optimizer. Under the hardware operation environment corresponding to the initial hardware parameter configuration samples, the sampled characterization value and the sampled model operation information are fitted to obtain a first optimized model. By using the method of taking the extreme value of the first optimized model, the basic candidate model operation information is obtained.

[0075] For example, after the software optimizer fits the relationship between the sampled characterization value and the sampled model operation method, a curve is obtained. Appropriate basic candidate model operation information can be obtained from the extreme value or the neighborhood of the extreme value on this curve.

[0076] Furthermore, the basic candidate model operation information and the initial hardware parameter configuration samples are input into an analysis model to predict the task efficiency characterization value obtained when the target model executes the operation tasks in the hardware operation environment corresponding to the initial hardware parameter configuration samples according to the basic candidate model operation information, and the task efficiency characterization value corresponding to each basic candidate model operation information is obtained, so that the software optimizer can obtain the candidate model operation information according to the task efficiency characterization value corresponding to each basic candidate model operation information.

[0077] It should be noted that the task efficiency characterization value and the sampled characterization value mentioned here are both indicators used to evaluate the model performance. For example, the time consumption and energy consumption, etc. The time consumption is used to represent the time required for the model to complete the operation task and can reflect the execution efficiency of the operation task. The energy consumption represents the power used by the model to complete the operation task. Among them, the task efficiency characterization value is the model performance evaluation of the analysis model for each basic candidate model operation information obtained by the software optimizer, and the sampled characterization value refers to the evaluation value of the model performance corresponding to the sampled model operation information required for the software optimizer to perform fitting.

[0078] In addition, there can be multiple preset evaluators. For example, a Maestro Evaluator (Maestro) or a Time Loop Evaluator (TimeLoop) can be used to perform the evaluation task.

[0079] Secondly, there are multiple methods for the target device to sample various initial model operation information. For example, the Latin Hypercube method can be used for uniform sampling to ensure that various initial model operation information can be uniformly sampled, further ensuring that when the target device adjusts the operation information in terms of software, it is more comprehensive and can obtain the optimal target model operation method corresponding to the optimal target model operation information.

[0080] S106: Input the candidate model operation information into a preset analysis model to predict, through the analysis model, the task efficiency characterization value obtained when performing the operation task of the target model in the hardware operation environment corresponding to the initial hardware parameter configuration sample according to the candidate model operation method corresponding to the candidate model operation information. According to the task efficiency characterization value corresponding to the candidate model operation method, determine the initial hardware parameter configuration sample for the next iteration under the condition of giving the candidate model operation method, and send the initial hardware parameter configuration sample for the next iteration to the software optimizer, so that the software optimizer determines the candidate model operation method corresponding to the next iteration according to the initial hardware parameter configuration sample for the next iteration and the model parameters, until the preset iteration condition is met, to obtain the target model operation method and the target hardware parameter configuration.

[0081] In this specification, the target device inputs the candidate model operation method determined by the software optimizer, the task efficiency characterization value corresponding to the candidate model operation method, and the initial hardware parameter configuration sample into a preset hardware optimizer to obtain the initial hardware parameter configuration sample for the next iteration under the condition of giving the candidate model operation method, and through the cyclic iteration between the software optimizer and the hardware optimizer, obtains the target hardware parameter configuration that meets the preset iteration condition, and the target model operation method determined based on the target hardware parameter configuration.

[0082] Specifically, the hardware optimizer fits the task efficiency characterization value corresponding to the candidate model operation mode and the initial hardware parameter configuration sample according to the candidate model operation mode to obtain a second optimization model. Furthermore, under the condition of a given candidate model operation mode, the second optimization model is used to adjust the initial hardware parameter configuration sample to obtain each basic hardware parameter configuration sample. Then, each basic hardware parameter configuration sample is input into a preset evaluator to predict the task efficiency characterization value of the target model when performing an arithmetic task in the hardware operation environment corresponding to each basic hardware parameter configuration sample, and the initial hardware parameter configuration sample for the next iteration is obtained according to the task efficiency characterization value corresponding to each basic hardware parameter configuration sample.

[0083] It should be noted that there are various methods for the hardware optimizer to fit the task efficiency characterization value corresponding to the candidate model operation mode and the initial hardware parameter configuration sample according to the candidate model operation mode. It can be to use Bayesian optimization based on heterogeneous variance evolution (Hogere Europese Beroepen Opleiding, HEBO) to fit this process.

[0084] S108: Configure the target hardware operation environment corresponding to the target hardware parameter configuration on the target device, and deploy the target model on the target device based on the target hardware operation environment to execute the arithmetic task corresponding to the target model on the target device according to the target model operation mode.

[0085] In this specification, with the goal of making the deviation between the task efficiency characterization values of the initial hardware parameter configuration samples in the two rounds before and after reaching the preset number of rounds less than the preset deviation, the initial hardware parameter configuration samples are adjusted, and based on the adjusted initial hardware parameter configuration samples, the selection and adjustment of the model operation mode are performed to obtain the target model operation mode and the corresponding target hardware parameter configuration samples.

[0086] After obtaining the target model operation mode and the target hardware parameter configuration through the above method, the target device is configured with the target hardware operation environment according to the target hardware parameter configuration sample, and the target model is executed on the target hardware according to the target model operation mode. Then, in the embodiments of this specification, for the finally obtained target model operation mode and the target hardware parameter configuration, it can be considered the best target hardware parameters and model deployment method obtained after the collaborative design of hardware and software, enabling the target device to execute the arithmetic task corresponding to the target model more efficiently.

[0087] As can be seen from the above method, based on the initial hardware parameter configuration sample, the model running mode of the target model is selected, and then, through the selected candidate model running mode, the initial hardware parameter configuration sample used in the next round of iteration is adjusted to further optimize the hardware parameter configuration. Moreover, each round of adjustment is based on the previous round of adjustment. This not only ensures that the more rounds of adjustment are carried out, the better the effect of the finally obtained initial hardware parameter configuration sample, but also selects and adjusts the model running mode of the target model on this basis, so as to achieve the co-design of hardware and software, ensuring a higher matching degree when the target model runs on the target device, and greatly improving the model running performance and the utilization rate of accelerator hardware resources.

[0088] Figure 2 It is a schematic diagram of the co-system architecture of software and hardware for a deep learning accelerator provided by an embodiment of this specification.

[0089] As Figure 2 shown, during the process of co-configuring the software and hardware of the target model and the target device, it includes mapping space search and design space search. The mapping space search is to select and adjust the model running information based on the initial hardware parameter configuration sample, while the design space search is to adjust the initial hardware parameter configuration sample based on the candidate model running mode corresponding to the candidate model running information obtained after one round of selection and adjustment of the model running information in the mapping space search, and use the adjusted initial hardware parameter configuration sample as the initial hardware configuration for the next round to input into the software optimizer of the mapping space search. Furthermore, based on the adjusted initial hardware parameter configuration sample, the mapping space search adjusts the model running mode, that is, the software optimizer determines each initial model running information according to the input hardware configuration, and inputs the determined initial model running information and the input hardware configuration into the analysis model to obtain the task efficiency characterization values corresponding to each initial model running information based on the determined hardware configuration. Thus, the software optimizer adjusts the model running information according to the task efficiency characterization values corresponding to each basic candidate model running information.

[0090] Through this method, in each round of mapping space search, based on the adjusted initial hardware parameter configuration sample, the candidate model operation mode corresponding to the candidate model operation information is obtained. The design space search is to adjust the initial hardware parameter configuration sample based on the candidate model operation mode, and each round of adjustment is carried out based on the previous round of adjustment. By analogy, after multiple iterations, the target hardware parameter configuration that meets the preset iteration conditions and the target model operation mode determined based on the target hardware parameter configuration sample can be obtained, that is, the best deployment mode of the target model on the target device and the allocation of hardware resources corresponding to the best operation efficiency of the target model.

[0091] By targeting the hardware operating environment configured according to the target hardware parameter configuration on the target device, the target model is deployed using the corresponding target model operation mode, thereby realizing the operation task of the target model, greatly improving the model operation efficiency and the utilization rate of accelerator hardware resources.

[0092] The above is a deep learning accelerator software and hardware co - operation method provided by the embodiments of this specification. Based on the same idea, this specification also provides corresponding devices, storage media, and electronic devices.

[0093] Figure 3 It is a schematic structural diagram of a deep learning accelerator software and hardware co - operation device provided by the embodiments of this specification. The device includes:

[0094] An acquisition module 301, configured to acquire the model parameters of the target model and the accelerator area constraint, where the accelerator area constraint is used to characterize the hardware configuration budget allocated by the accelerator to the target model;

[0095] A first determination module 302, configured to determine an initial hardware parameter configuration sample according to the hardware parameter configuration included in the accelerator area constraint;

[0096] A second determination module 303, configured to input the hardware parameters corresponding to the initial hardware parameter configuration sample and the model parameters into a preset software optimizer, so that the software optimizer determines each initial model operation information in the given hardware operating environment corresponding to the initial hardware parameter configuration sample, and determines candidate model operation information based on the each initial model operation information;

[0097] An analysis module 304 is configured to input candidate model running information into a preset analysis model, so as to, through the analysis model, predict a task efficiency characterization value obtained when performing an operation task of the target model in a hardware operating environment corresponding to the initial hardware parameter configuration mode according to a candidate model running mode corresponding to the candidate model running information. According to the task efficiency characterization value corresponding to the candidate model running mode, an initial hardware parameter configuration sample for the next iteration is determined under the condition of giving the candidate model running mode, and the initial hardware parameter configuration sample for the next iteration is sent to the software optimizer, so that the software optimizer determines a candidate model running mode corresponding to the next iteration according to the initial hardware parameter configuration sample for the next iteration and the model parameters, until a preset iteration condition is met, and a target model running mode and a target hardware parameter configuration are obtained.

[0098] An operation module 305 is configured to configure a target hardware operating environment corresponding to the target hardware parameter configuration on a target device, and deploy the target model on the target device based on the target hardware operating environment, so as to perform an operation task corresponding to the target model on the target device according to the target model running mode.

[0099] Optionally, the second determination module 303 is specifically configured to: sample from the initial model running information to obtain sampled model running information; input the sampled model running information and the initial hardware parameter configuration sample into a preset evaluator, so as to, through the evaluator, predict a task efficiency characterization value obtained when performing an operation task of the target model in a hardware operating environment corresponding to the initial hardware parameter configuration sample according to the sampled model running information, and use it as a sampled characterization value; input the sampled characterization value, the sampled model running information, and the initial hardware parameter configuration sample into the software optimizer, so as to, in a hardware operating environment corresponding to the initial hardware parameter configuration sample, fit the sampled characterization value and the sampled model running information through the software optimizer to obtain a first optimization model; and determine candidate model running information based on the initial model running information through the first optimization model.

[0100] Optionally, the analysis module 304 is specifically configured to input the task efficiency characterization value corresponding to the candidate model running mode, the candidate model running mode, and the initial hardware parameter configuration sample into a preset hardware optimizer, so as to, through the hardware optimizer, determine an initial hardware parameter configuration sample for the next iteration under the condition of giving the candidate model running mode.

[0101] Optionally, the analysis module 304 is specifically configured to, based on the candidate model operation mode, fit the task efficiency characterization value corresponding to the candidate model operation mode and the initial hardware parameter configuration sample through the hardware optimizer to obtain a second optimization model;

[0102] Based on each hardware parameter option required when the target model executes a task, through the second optimization model, determine each basic hardware parameter configuration sample under the condition of giving the candidate model operation mode;

[0103] Input each of the basic hardware parameter configuration samples into a preset evaluator, so as to, through the evaluator, predict the task efficiency characterization value obtained when executing the arithmetic task of the target model in the hardware operation environment corresponding to each basic hardware parameter configuration sample, and use it as the task efficiency characterization value corresponding to each basic hardware parameter configuration sample;

[0104] Determine the initial hardware parameter configuration sample for the next round of iteration according to the task efficiency characterization value corresponding to each basic hardware parameter configuration sample.

[0105] Optionally, the analysis module 304 is specifically configured to, if it is monitored that the deviation between the task efficiency characterization values corresponding to the initial hardware parameter configuration samples in the previous and next rounds after reaching the preset number of rounds is less than the preset deviation, determine that the preset iteration condition is met, and use the initial hardware parameter configuration sample obtained when the preset iteration condition is met as the target hardware parameter configuration, and use the candidate model operation mode determined by the software optimizer based on the target hardware parameter configuration as the target model operation mode.

[0106] Optionally, the target model includes at least one model for executing arithmetic tasks.

[0107] This specification also provides a computer-readable storage medium, and the storage medium stores a computer program, and when the computer program is executed by a processor, it can be used to execute the above Figure 1 A deep learning accelerator software and hardware cooperation method provided.

[0108] Based on Figure 1 A deep learning accelerator software and hardware cooperation method shown, an embodiment of this specification also provides Figure 4 A schematic structural diagram of an electronic device shown. As Figure 4 , at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, there may also be other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 A deep learning accelerator software and hardware cooperation method described.

[0109] Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or logic devices.

[0110] In the 1990s, it was obvious to distinguish whether an improvement to a technology was an improvement in hardware (e.g., improvement to circuit structures such as diodes, transistors, switches, etc.) or an improvement in software (improvement to method flows). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structures by programming the improved method flows into the hardware circuits. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. The designer can program by himself / herself to "integrate" a digital system on a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be clear that by simply making a little logical programming of the method flow with the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0111] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0112] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0113] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0114] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0115] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 one or more of the blocks for implementing the specified functions.

[0116] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 one or more of the blocks.

[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 one or more of the blocks.

[0118] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0119] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0120] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0121] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0122] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0123] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0124] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the corresponding description in the method embodiment.

[0125] The above is only the embodiment of this specification and is not used to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A deep learning accelerator software and hardware collaboration method, characterized in that: include: Acquire model parameters of the target model and an accelerator area constraint, where the accelerator area constraint is used to characterize a hardware configuration budget allocated by the accelerator to the target model; Determining an initial hardware parameter configuration sample according to the hardware parameter configuration included in the accelerator area constraint; Input the hardware parameters corresponding to the initial hardware parameter configuration sample and the model parameters into a preset software optimizer, so that the software optimizer determines each initial model operation information under a given hardware operation environment corresponding to the initial hardware parameter configuration sample, and determines the candidate model operation information based on the each initial model operation information, wherein sampling is performed from each initial model operation information to obtain the sampled model operation information; input the sampled model operation information and the initial hardware parameter configuration sample into a preset evaluator, so as to predict, through the evaluator, the task efficiency characterization value obtained when the computing task of the target model is executed according to the sampled model operation information under the hardware operation environment corresponding to the initial hardware parameter configuration sample, as a sampled characterization value; input the sampled characterization value, the sampled model operation information and the initial hardware parameter configuration sample into the software optimizer, so as to fit the sampled characterization value and the sampled model operation information through the software optimizer under the hardware operation environment corresponding to the initial hardware parameter configuration sample to obtain a first optimization model; determine the candidate model operation information based on the each initial model operation information through the first optimization model; The candidate model operation information is input into a preset analysis model, so as to predict, through the analysis model, the task efficiency characterization value obtained when the operation task of the target model is executed in the hardware operation environment corresponding to the initial hardware parameter configuration sample according to the candidate model operation mode corresponding to the candidate model operation information; according to the task efficiency characterization value corresponding to the candidate model operation mode, the initial hardware parameter configuration sample of the next iteration is determined under the condition of the given candidate model operation mode; the initial hardware parameter configuration sample of the next iteration is sent to the software optimizer, so that the software optimizer determines the candidate model operation mode corresponding to the next iteration according to the initial hardware parameter configuration sample of the next iteration and the model parameters, until the preset iteration conditions are met, and the target model operation mode and the target hardware parameter configuration are obtained; A target hardware operating environment corresponding to the target hardware parameter configuration is configured on the target device, and based on the target hardware operating environment, the target model is deployed on the target device to execute a computing task corresponding to the target model on the target device according to the target model operating mode.

2. The method according to claim 1, characterized in that According to the task efficiency characterization value corresponding to the candidate model operation mode, determining the initial hardware parameter configuration sample for the next iteration under the condition of the given candidate model operation mode, specifically including: The task efficiency characterization value corresponding to the candidate model operation mode, the candidate model operation mode and the initial hardware parameter configuration sample are input into a preset hardware optimizer, so as to determine the initial hardware parameter configuration sample for the next iteration under the condition of the given candidate model operation mode through the hardware optimizer.

3. The method according to claim 2, characterized in that Determining, by the hardware optimizer, an initial hardware parameter configuration sample for the next iteration under the condition of the given operation mode of the candidate model, specifically including: Based on the candidate model operation mode, the hardware optimizer is used to fit the task efficiency characterization value corresponding to the candidate model operation mode and the initial hardware parameter configuration sample to obtain a second optimization model; Based on the hardware parameter configurations required by the target model to perform the task, determine, through the second optimization model, various basic hardware parameter configuration samples under the condition of the given candidate model operation mode; Inputting each of the basic hardware parameter configuration samples into a preset evaluator, so as to predict, through the evaluator, a task efficiency characterization value obtained when executing the computing task of the target model under the hardware operating environment corresponding to each of the basic hardware parameter configuration samples, as the task efficiency characterization value corresponding to each of the basic hardware parameter configuration samples; According to the task efficiency characterization value corresponding to each basic hardware parameter configuration sample, the initial hardware parameter configuration sample for the next iteration is determined.

4. The method according to claim 1, characterized in that Until the preset iteration conditions are met, the target model operation mode and target hardware parameter configuration are obtained, including: If it is monitored that the deviation between the task efficiency characterization values ​​corresponding to the initial hardware parameter configuration samples corresponding to the two rounds before and after the preset round is less than the preset deviation, it is determined that the preset iteration condition is met, and the initial hardware parameter configuration sample obtained when the preset iteration condition is met is used as the target hardware parameter configuration, and the candidate model operation mode determined by the software optimizer based on the target hardware parameter configuration is used as the target model operation mode.

5. The method according to any one of claims 1 to 4, characterized in that The target model includes at least one model for performing a computing task.

6. A deep learning accelerator hardware and software collaborative device, characterized in that: include: An acquisition module, used for acquiring model parameters of a target model and an accelerator area constraint, wherein the accelerator area constraint is used for characterizing a hardware configuration budget allocated by the accelerator to the target model; A first determination module, configured to determine an initial hardware parameter configuration sample according to the hardware parameter configuration included in the accelerator area constraint; A second determination module is used to input the hardware parameters corresponding to the initial hardware parameter configuration sample and the model parameters into a preset software optimizer, so that the software optimizer determines each initial model operation information under a given hardware operation environment corresponding to the initial hardware parameter configuration sample, and determines the candidate model operation information based on the each initial model operation information, wherein sampling is performed from each initial model operation information to obtain the sampled model operation information; the sampled model operation information and the initial hardware parameter configuration sample are input into a preset evaluator, so as to predict, through the evaluator, the task efficiency characterization value obtained when the computing task of the target model is executed according to the sampled model operation information under the hardware operation environment corresponding to the initial hardware parameter configuration sample, as the sampled characterization value; the sampled characterization value, the sampled model operation information and the initial hardware parameter configuration sample are input into the software optimizer, so as to fit the sampled characterization value and the sampled model operation information through the software optimizer under the hardware operation environment corresponding to the initial hardware parameter configuration sample to obtain a first optimization model; and the candidate model operation information is determined based on the each initial model operation information through the first optimization model; An analysis module is used to input the candidate model operation information into a preset analysis model, so as to predict, through the analysis model, a task efficiency characterization value obtained when the operation task of the target model is executed in the hardware operation environment corresponding to the initial hardware parameter configuration sample according to the candidate model operation mode corresponding to the candidate model operation information, and determine, based on the task efficiency characterization value corresponding to the candidate model operation mode, the initial hardware parameter configuration sample of the next iteration under the condition of the given candidate model operation mode, so as to send the initial hardware parameter configuration sample of the next iteration to the software optimizer, so that the software optimizer determines the candidate model operation mode corresponding to the next iteration according to the initial hardware parameter configuration sample of the next iteration and the model parameters, until the preset iteration conditions are met, and the target model operation mode and the target hardware parameter configuration are obtained; An operation module is used to configure a target hardware operating environment corresponding to the target hardware parameter configuration on a target device, and deploy the target model on the target device based on the target hardware operating environment, so as to execute the computing tasks corresponding to the target model on the target device according to the target model operation mode.

7. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Systems and methods for agile and explainable optimization of efficient hardware / software codesigns for domain-specific computing systems using bottleneck analysis

    US20240134769A1