Chip optimization method and device, electronic device and storage medium
By selecting chip architecture parameters and optimization strategies in deep neural network computing tasks, and calculating and selecting the optimization strategy with the least time-consuming time on the device side, the problem of low chip resource utilization in the existing technology is solved and the efficiency of computing tasks is improved.
Patent Information
- Application Number
- CN202311570185.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-11-22
AI Technical Summary
When optimizing deep neural network computing tasks, it is difficult for the prior art to effectively utilize chip resources, resulting in insufficient computing efficiency and performance.
By obtaining multiple fusion operators in the artificial intelligence model, selecting the architecture parameters of the chip and multiple optimization strategies applied to the target fusion operator, calculating the device-side time-consuming of each optimization strategy, and selecting the optimization strategy with the least time-consuming as the target optimization strategy.
Given architectural parameters, the device-side time-consuming of the target fusion operator is optimized, and the utilization rate of chip resources and the execution efficiency of computing tasks are improved.
Smart Images

Figure CN117556756B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a chip optimization method and device, an electronic device, and a storage medium. Background Art
[0002] In the field of Artificial Intelligence (AI), Deep Neural Network (DNN) has become the foundation of the most advanced technology and the core driving force of many applications. Operator fusion is a method to improve the execution efficiency of deep learning computing tasks; before executing the computing task, multiple operators that meet certain conditions or rules in the neural network can be fused to form a fused operator. By fusing multiple operators, memory reuse can be achieved and the utilization rate of computing resources such as central processing unit (CPU), graphics processing unit (GPU), and general-purpose graphics processing unit (GPGPU) can be improved. Summary of the invention
[0003] At least one embodiment of the present disclosure provides a chip optimization method, which includes: obtaining an artificial intelligence model, wherein the artificial intelligence model includes multiple fusion operators, and the multiple fusion operators include a target fusion operator; selecting a first architecture parameter of the chip and multiple optimization strategies applied to the target fusion operator; based on the first architecture parameter, respectively calculating the device-side time consumption of the target fusion operator when applying each of the multiple optimization strategies; in response to the target fusion operator having the smallest device-side time consumption when applying the first optimization strategy among the multiple optimization strategies, selecting the first optimization strategy as the target optimization strategy of the target fusion operator under the first architecture parameter.
[0004] For example, the chip optimization method provided by at least one embodiment of the present disclosure also includes: merging the device-side time consumption and the host-side time consumption of the target fusion operator when applying the target optimization strategy to obtain the total time consumption of the target fusion operator when applying the target optimization strategy.
[0005] For example, the chip optimization method provided by at least one embodiment of the present disclosure also includes: decomposing the artificial intelligence model into the multiple fusion operators; selecting the target optimization strategy corresponding to each of the multiple fusion operators under the first architecture parameters; based on the first architecture parameters, respectively calculating the total time consumed by the multiple fusion operators when applying the corresponding target optimization strategies; merging the total time consumed by the multiple fusion operators when applying the corresponding target optimization strategies to obtain the total time consumed by the artificial intelligence model under the first architecture parameters.
[0006] For example, in the chip optimization method provided by at least one embodiment of the present disclosure, the target fusion operator includes multiple single operators, the multiple optimization strategies include the 1st optimization strategy to the Nth optimization strategy, where N is a positive integer, and the device-side time consumption of the target fusion operator when applying each of the multiple optimization strategies based on the first architecture parameters is calculated separately, including: decomposing the target fusion operator into the multiple single operators; based on the first architecture parameters, calculating the time consumption of the multiple single operators when applying the i-th optimization strategy among the multiple optimization strategies, where i=1,2,…,N; merging the multiple time consumptions of the multiple single operators when applying the i-th optimization strategy to obtain the device-side time consumption of the target fusion operator when applying the i-th optimization strategy.
[0007] For example, in the chip optimization method provided by at least one embodiment of the present disclosure, each of the multiple single operators corresponds to an operator strategy, and the operator strategy of each single operator is associated with an architecture parameter of the chip, and the time consumption of the multiple single operators in applying the i-th optimization strategy among the multiple optimization strategies based on the first architecture parameter includes: based on the operator strategy of each single operator corresponding to the first architecture parameter, calculating the time consumption of each single operator in applying the i-th optimization strategy.
[0008] For example, in the chip optimization method provided by at least one embodiment of the present disclosure, the multiple optimization strategies include the first optimization strategy to the Nth optimization strategy, where N is a positive integer, and in response to the target fusion operator having the smallest device-side time consumption when applying the first optimization strategy among the multiple optimization strategies, selecting the first optimization strategy as the target optimization strategy of the target fusion operator under the first architecture parameters includes: setting a preset time consumption of the target fusion operator, where the preset time consumption is initialized to infinity; comparing the device-side time consumption of the target fusion operator when applying the i-th optimization strategy among the multiple optimization strategies with the preset time consumption in the traversal order of the first optimization strategy to the Nth optimization strategy, and using the smaller value therebetween to update the preset time consumption, where i=1,2,…,N; in response to the completion of the traversal of the first optimization strategy to the Nth optimization strategy, selecting the first optimization strategy corresponding to the preset time consumption as the target optimization strategy of the target fusion operator under the first architecture parameters.
[0009] For example, the chip optimization method provided by at least one embodiment of the present disclosure also includes: selecting multiple architecture parameters of the chip, wherein the multiple architecture parameters include the first architecture parameter; calculating the total time consumed by the artificial intelligence model when applying each of the multiple architecture parameters; in response to the artificial intelligence model having the smallest total time consumed when applying the second architecture parameter of the multiple architecture parameters, selecting the second architecture parameter as the target architecture parameter of the chip.
[0010] For example, the chip optimization method provided by at least one embodiment of the present disclosure further includes: configuring the chip using the target architecture parameters.
[0011] For example, in the chip optimization method provided by at least one embodiment of the present disclosure, the multiple architecture parameters include the 1st architecture parameter to the Mth architecture parameter, where M is a positive integer, and the total time consumed by the artificial intelligence model when applying each of the multiple architecture parameters includes: decomposing the artificial intelligence model into the multiple fusion operators; selecting the target optimization strategy corresponding to each of the multiple fusion operators under the kth architecture parameter among the multiple architecture parameters, where k=1, 2,…, M; based on the kth architecture parameter, respectively calculating the total time consumed by the multiple fusion operators when applying the corresponding target optimization strategies; merging the total time consumed by the multiple fusion operators when applying the corresponding target optimization strategies to obtain the total time consumed by the artificial intelligence model under the kth architecture parameter.
[0012] For example, in the chip optimization method provided by at least one embodiment of the present disclosure, the target fusion operator includes multiple single operators, each of the multiple single operators corresponds to an operator strategy, and for each of the multiple single operators, the multiple architecture parameters correspond to multiple operator strategies respectively.
[0013] At least one embodiment of the present disclosure also provides a chip optimization device, which includes: an acquisition module, configured to acquire an artificial intelligence model, wherein the artificial intelligence model includes multiple fusion operators, and the multiple fusion operators include a target fusion operator; a selection module, configured to select a first architecture parameter of the chip and multiple optimization strategies applied to the target fusion operator; a calculation module, configured to calculate the device-side time consumption of the target fusion operator when applying each of the multiple optimization strategies based on the first architecture parameter, wherein the selection module is further configured to select the first optimization strategy as the target optimization strategy of the target fusion operator under the first architecture parameter in response to the target fusion operator having the smallest device-side time consumption when applying the first optimization strategy among the multiple optimization strategies.
[0014] At least one embodiment of the present disclosure further provides an electronic device. The electronic device includes: a processor; a memory including one or more computer program modules; wherein the one or more computer program modules are stored in the memory and configured to be executed by the processor, and the one or more computer program modules are used to implement the chip optimization method provided in any embodiment of the present disclosure.
[0015] At least one embodiment of the present disclosure further provides a storage medium storing non-transitory computer-readable instructions, which, when executed by a computer, implements the chip optimization method provided by any embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, but are not intended to limit the present disclosure.
[0017] Figure 1 A schematic diagram of a method for executing computing tasks on an artificial intelligence chip;
[0018] Figure 2 An exemplary flow chart of a chip optimization method provided for at least one embodiment of the present disclosure;
[0019] Figure 3 Another exemplary flow chart of a chip optimization method provided for at least one embodiment of the present disclosure;
[0020] Figure 4 A schematic diagram of an example of a chip optimization method provided by at least one embodiment of the present disclosure;
[0021] Figure 5 Another exemplary flow chart of a chip optimization method provided for at least one embodiment of the present disclosure;
[0022] Figure 6 A schematic diagram of another example of a chip optimization method provided by at least one embodiment of the present disclosure;
[0023] Figure 7 A schematic block diagram of a chip optimization device provided for at least one embodiment of the present disclosure;
[0024] Figure 8 A schematic block diagram of an electronic device provided for at least one embodiment of the present disclosure;
[0025] Fig. 9 A schematic block diagram of another electronic device provided for at least one embodiment of the present disclosure; and
[0026] Fig.10 A schematic diagram of a storage medium provided for at least one embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0028] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, similar words such as "one", "one" or "the" do not indicate quantity restrictions, but indicate that there is at least one. Similar words such as "include" or "comprise" mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Similar words such as "connect" or "connected" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0029] The present disclosure is described below through several specific embodiments. In order to keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and known components may be omitted. When any component of the embodiments of the present disclosure appears in more than one figure, the component is represented by the same or similar reference numeral in each figure.
[0030] Artificial Neural Networks (ANNs), also referred to as neural networks, are an algorithmic mathematical model that imitates the behavioral characteristics of animal neural networks and performs distributed parallel information processing. This network relies on the complexity of the system to adjust the interconnected relationships between a large number of internal nodes to achieve the purpose of processing information. Neural networks can include recurrent neural networks (RNNs), long / short-term memory (LSTM), deep belief networks (DBNs), convolutional neural networks (CNNs), etc. Regardless of the type of artificial neural network, their common characteristics are large-scale parallel processing, distributed storage, elastic topology, high redundancy and nonlinear operations, etc., and have capabilities in terms of computing speed, associative ability, adaptability, fault tolerance and self-organization. These characteristics and capabilities constitute the technical basis for artificial neural networks to simulate intelligent activities and have been used in various technical fields. For example, artificial neural networks can be used in data compression, image processing, video coding, signal processing, etc.
[0031] A deep learning computing task consists of multiple computing units, each of which can be called an operator (Operator, Op). In a neural network model, an operator corresponds to the computing logic of a layer or node. For example, a convolution computing task in a convolution layer can be an operator; a weight sum computing task in a fully-connected layer (FC layer) can also be an operator. When implementing a neural network, an operator can be implemented in software (for example, through a computer program) or in hardware (for example, through a circuit).
[0032] Deep learning computing tasks need to be deployed on different artificial intelligence chips, such as graphics processing units (GPUs), general-purpose graphics processing units (GPGPUs), etc. The following uses GPUs / GPGPUs as examples for illustration, but the present disclosure is not limited to GPUs / GPGPUs (the following description of GPUs / GPGPUs is also applicable to other types of processors).
[0033] Figure 1 A schematic diagram of how an artificial intelligence chip performs computing tasks.
[0034] For example, Figure 1 As shown in the figure, the program architecture of the processor on the artificial intelligence chip (such as GPU or GPGPU) can include two parts: (1) the host program running on the CPU, and (2) the device program running on the processor. When the processor is used to perform a computing task, the CPU calls the kernel function and provides the kernel function and input data to the device side, and then the thread block in the grid of the device side executes the kernel function to process the input data; the CPU then controls the processor to copy the processing results back to the host side, thereby completing the computing task.
[0035] For example, a chip can have multiple architecture parameters (also called hardware architecture parameters). Taking GPGPU as an example, the hardware architecture parameters of GPGPU mainly include the following categories:
[0036] 1) Core parameters: including the number of GPGPU cores, core frequency, instruction set, and computing unit. These parameters affect the computing performance of GPGPU. The number of cores refers to the number of basic units in GPGPU that can execute instructions in parallel; the more cores there are, the more data and tasks GPGPU can process simultaneously. The core frequency refers to the number of instruction cycles that each core in GPGPU can execute per second, generally expressed in GHz; the higher the core frequency, the faster the single core of GPGPU operates. The instruction set refers to the instruction type and format supported by each core in GPGPU, generally expressed in ISA (Instruction Set Architecture); the instruction set determines the operations and functions that GPGPU can perform. The computing unit refers to the hardware module that performs specific types of operations inside each core in GPGPU, such as integer computing unit (INT32), floating point computing unit (FP32, FP64), tensor computing unit (or tensor core, Tensor Core), etc.; the type and number of computing units affect the support and performance of GPGPU for different data types and precisions.
[0037] 2) Storage parameters: including GPGPU memory capacity, memory type, memory bandwidth, cache size, shared memory size, etc. These parameters affect the storage efficiency of GPGPU. Memory capacity refers to the size of the memory space used to store data and programs in GPGPU, generally expressed in GB; the larger the memory capacity, the larger the data and tasks that GPGPU can handle. Memory type refers to the type and specification of memory chips used in GPGPU, such as DDR, GDDR, HBM, etc.; the memory type determines the memory access method and characteristics of GPGPU. Memory bandwidth refers to the total amount of data that can be read and written from memory per second in GPGPU, generally expressed in GB / s; the larger the memory bandwidth, the higher the memory access efficiency of GPGPU. Cache size refers to the size of the high-speed buffer used to accelerate memory access in GPGPU, generally at different levels such as L1 and L2; the larger the cache size, the more commonly used data GPGPU can cache, reducing memory access latency. Shared memory size refers to the size of the memory space used to share data between different cores in GPGPU; the larger the shared memory size, the more collaborative computing between cores GPGPU can support.
[0038] 3) Power consumption parameters: including the operating voltage, operating current, operating temperature, power consumption, etc. of GPGPU. These parameters affect the stability and life of GPGPU. The operating voltage refers to the voltage value required for each unit in GPGPU to work normally; the higher the operating voltage, the faster the GPGPU runs, but it also generates more heat. The operating current refers to the current value required for each unit in GPGPU to work normally; the larger the operating current, the higher the operating power of GPGPU, but it also generates more heat. The operating temperature refers to the temperature value when each unit in GPGPU works normally; the higher the operating temperature, the greater the heat dissipation demand of GPGPU, but it also affects the stability and life of GPGPU. Power consumption refers to the electric energy value consumed by each unit in GPGPU when working normally; the greater the power consumption, the higher the energy consumption cost of GPGPU, but it also provides higher performance.
[0039] For example, for the hardware architecture parameters of a chip, the number of cores and core frequency determine the peak floating-point computing capability of the chip. The memory capacity and memory bandwidth determine the memory access capability of the chip. If storage is too fast and calculation is too slow, the calculation will become the bottleneck of the operator, and the storage unit will often be idle, and vice versa. Only when the data access speed matches the calculation speed can the ultimate performance of the chip be fully utilized.
[0040] For example, for chip power consumption and performance, generally speaking, power consumption and performance are proportional, that is, the higher the parameters (such as the number of cores, core frequency, memory bandwidth, etc.), the higher the performance, but the higher the power consumption. Therefore, when designing and using chips, it is necessary to weigh the balance between power consumption and performance according to different scenarios and requirements.
[0041] Before a chip leaves the factory, it needs to be evaluated in many ways, including pre-silicon and post-silicon performance simulation, architecture verification, and AI model performance. In order to evaluate the chip, a large number of high-performance operators need to be developed and the corresponding operators need to be adapted on the simulation platform or chip to run the evaluation model. Due to the large number of operators and the complex development and adaptation processes, it takes a long time to run the operators on the hardware simulation platform, which reduces the efficiency of chip evaluation and verification.
[0042] At least one embodiment of the present disclosure provides a chip optimization method, which includes: obtaining an artificial intelligence model, wherein the artificial intelligence model includes multiple fusion operators, and the multiple fusion operators include a target fusion operator; selecting a first architecture parameter of the chip and multiple optimization strategies applied to the target fusion operator; based on the first architecture parameter, respectively calculating the device-side time consumption of the target fusion operator when applying each of the multiple optimization strategies; in response to the target fusion operator having the smallest device-side time consumption when applying the first optimization strategy among the multiple optimization strategies, selecting the first optimization strategy as the target optimization strategy of the target fusion operator under the first architecture parameter.
[0043] At least one embodiment of the present disclosure further provides a chip optimization device, an electronic device, and a storage medium for implementing the chip optimization method of the above embodiment.
[0044] The method, device, electronic device and storage medium provided by at least one embodiment of the present disclosure can obtain the optimal optimization strategy corresponding to the fusion operator in the artificial intelligence model when the architectural parameters of the chip are fixed, so as to guide operator development, operator fusion and software configuration, etc., thereby quickly evaluating the expected performance of the chip and improving the evaluation and verification efficiency of the chip.
[0045] Hereinafter, at least one embodiment of the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that the same reference numerals in different drawings will be used to refer to the same elements described.
[0046] Figure 2 An exemplary flow chart of a chip optimization method provided for at least one embodiment of the present disclosure.
[0047] For example, Figure 2 As shown, at least one embodiment of the present disclosure provides a chip optimization method, which may include the following steps S110 to S140.
[0048] Step S110: Obtain an artificial intelligence model.
[0049] For example, in step S110, the artificial intelligence model includes multiple fusion operators, and the multiple fusion operators include a target fusion operator. For example, before executing a deep learning computing task, multiple single operators that meet certain conditions or rules in the artificial intelligence model can be fused based on a specific fusion strategy to form a fused operator.
[0050] Step S120: selecting first architecture parameters of the chip and multiple optimization strategies applied to the target fusion operator.
[0051] For example, in step S120, one of the multiple groups of hardware parameters of the chip is selected as the first architecture parameter. For example, as described above, the architecture parameters of the chip may include core parameters, storage parameters, and power consumption parameters, etc., which are used to characterize the chip's memory or cache size, supported instructions, the time consumption of each instruction, computing power (how fast each instruction is executed), and bandwidth (how fast data is accessed).
[0052] For example, multiple different optimization strategies can be selected for the fusion operator. For example, in the same artificial intelligence model, the name and function of the operator are the same, but if different optimization strategies are used, the specific implementation in the operator may be different. For example, operators have different writing methods, and different optimization strategies write different operators (that is, different operators implemented at the software level), so the decomposed hardware instructions are also different. For example, when the first architecture parameter of the chip is fixed, the time consumption of each hardware instruction is fixed, but the performance of the operator is different under different optimization strategies.
[0053] Step S130: Based on the first architecture parameter, respectively calculating the device-side time consumed by the target fusion operator when applying each of the multiple optimization strategies.
[0054] For example, in step S130, when the first architecture parameter is fixed, multiple different optimization strategies of the target fusion operator are traversed, and the device-side time consumption of the target fusion operator when applying each optimization strategy is calculated respectively. For example, for multiple single operators included in the target fusion operator, each single operator corresponds to an operator strategy (such as a tiling strategy), and the operator strategy is a strategy related to the operator itself. For example, the device-side time consumption of the target fusion operator is related to the operator strategy in addition to the optimization strategy. For example, different architecture parameters correspond to different operator strategies; when the first architecture parameter of the chip is fixed, the corresponding operator strategy is also fixed.
[0055] Step S140: In response to the target fusion operator having the smallest device-side time consumption when applying a first optimization strategy among multiple optimization strategies, selecting the first optimization strategy as the target optimization strategy of the target fusion operator under the first architecture parameters.
[0056] For example, in step S140, after traversing multiple different optimization strategies, if the target fusion operator consumes the least time on the device side when applying the first optimization strategy, the first optimization strategy is selected as the target optimization strategy of the target fusion operator under the first architecture parameters. For example, a small time consumption on the device side indicates a fast instruction execution speed (correspondingly, good chip performance), so the target optimization strategy obtained in step S140 is the optimal optimization strategy of the target fusion operator under the first architecture parameters.
[0057] In some examples, such as Figure 2 As shown, the chip optimization method provided by at least one embodiment of the present disclosure may further include step S150.
[0058] Step S150: The device-side time consumption and the host-side time consumption of the target fusion operator when applying the target optimization strategy are combined to obtain the total time consumption of the target fusion operator when applying the target optimization strategy.
[0059] For example, in step S150, the total time consumption of the target fusion operator is obtained by combining the device-side time consumption and the host-side time consumption. For example, the host-side time consumption is related to factors such as CPU performance, and an empirical value can be taken as the host-side time consumption. For example, the combination of the device-side time consumption and the host-side time consumption is not a simple addition; since the device-side and the host-side can be concurrent, the total time consumption of the target fusion operator obtained by combining may be less than the sum of the device-side time consumption and the host-side time consumption.
[0060] It should be noted that, in addition to artificial intelligence models, the chip optimization method provided in at least one embodiment of the present disclosure is also applicable to other models containing multiple operators; accordingly, in addition to artificial intelligence chips, the chip optimization method provided in at least one embodiment of the present disclosure is also applicable to other types of chips; the specific selection can be based on actual needs, and the embodiments of the present disclosure are not limited to this.
[0061] It should be noted that the calculation of the total time consumption of the target fusion operator is not limited to the method of step S150. The type of chip is different, and the corresponding total time consumption calculation method is also different; for example, for a specific type of chip, the device-side time consumption of the target fusion operator is the total time consumption of the target fusion operator; the specific selection can be made according to the actual situation such as the type of chip, and the embodiments of the present disclosure are not limited to this.
[0062] It should be noted that Figure 2Steps S110 to S150 are used to select the target optimization strategy of the target fusion operator from the multiple fusion operators and calculate the corresponding total time consumption; for each of the multiple fusion operators, Figure 2 Steps S110 to S150 calculate the corresponding target optimization strategy (ie, the optimal optimization strategy) and the corresponding total time consumption.
[0063] Figure 3 Another exemplary flow chart of a chip optimization method provided for at least one embodiment of the present disclosure.
[0064] For example, Figure 3 As shown, the chip optimization method provided by at least one embodiment of the present disclosure may also include the following steps S210 to S240. For example, the artificial intelligence model in step S210 is Figure 2 The artificial intelligence model obtained in step S110; in step S220, the target optimization strategy of each fusion operator is obtained by Figure 2 The total time consumption of each fusion operator in step S230 is obtained by Figure 2 Calculated in step S150.
[0065] Step S210: Decompose the artificial intelligence model into multiple fusion operators.
[0066] For example, in step S210, multiple single operators are fused based on the fusion strategy to obtain a fusion operator, and multiple fusion operators are also fused based on the fusion strategy to obtain a fusion operator artificial intelligence model. For example, the artificial intelligence model can be disassembled into multiple fusion operators based on the fusion strategy, and Figure 2 The steps traverse multiple fusion operators, and each fusion operator traverses multiple optimization strategies.
[0067] Step S220: selecting a target optimization strategy corresponding to each fusion operator among the multiple fusion operators under the first architecture parameters.
[0068] For example, in step S220, when the first architecture parameters are fixed, each fusion operator is used as a target fusion operator, using Figure 2 Steps S120 to S140 traverse multiple optimization strategies of each fusion operator, and select a target optimization strategy corresponding to each fusion operator.
[0069] Step S230: Based on the first architecture parameter, the total time consumed by the multiple fusion operators when applying the corresponding target optimization strategies is calculated respectively.
[0070] For example, in step S230, after obtaining the target optimization strategy corresponding to each fusion operator, based on the device-side time consumed by each fusion operator in applying the corresponding target optimization strategy calculated in step S220, use Figure 2 In step S150, the device-side time consumption and the host-side time consumption of each fusion operator when applying the corresponding target optimization strategy are combined to obtain the total time consumption of each fusion operator when applying the corresponding target optimization strategy.
[0071] Step S240: Combine the total time consumed by multiple fusion operators when applying the corresponding target optimization strategies to obtain the total time consumed by the artificial intelligence model under the first architecture parameters.
[0072] For example, in step S240, based on the total time consumed by each fusion operator in applying the corresponding target optimization strategy calculated in step S230, the total time consumed by multiple fusion operators in applying the corresponding target optimization strategies are merged to obtain the total time consumed by the artificial intelligence model under the first architecture parameters.
[0073] In some examples, the target fusion operator includes multiple single operators, and the multiple optimization strategies include the first optimization strategy to the Nth optimization strategy, where N is a positive integer. For example, Figure 2 Step S130 may further include steps S1301 to S1303:
[0074] Step S1301: Decompose the target fusion operator into multiple single operators;
[0075] Step S1302: based on the first architecture parameter, calculating the time consumed by multiple single operators when applying the i-th optimization strategy among multiple optimization strategies, where i=1, 2, ..., N;
[0076] Step S1303: Merge multiple time consumptions of multiple single operators when applying the i-th optimization strategy to obtain the device-side time consumption of the target fusion operator when applying the i-th optimization strategy.
[0077] For example, each of the multiple single operators corresponds to an operator strategy, and the operator strategy of each single operator is associated with the architecture parameter of the chip. For example, step S1302 may further include: based on the operator strategy of each single operator corresponding to the first architecture parameter, calculating the time consumed by each single operator when applying the i-th optimization strategy.
[0078] For example, Figure 2 Step S140 may further include steps S1401 to S1403:
[0079] Step S1401: setting a preset time consumption of a target fusion operator, wherein the preset time consumption is initialized to infinity;
[0080] Step S1402: in the traversal order from the 1st optimization strategy to the Nth optimization strategy, compare the device-side time consumption of the target fusion operator when applying the i-th optimization strategy among the multiple optimization strategies with the preset time consumption, and use the smaller value to update the preset time consumption, where i = 1, 2, ..., N;
[0081] Step S1403: In response to the traversal of the first optimization strategy to the Nth optimization strategy being completed, a first optimization strategy corresponding to a preset time consumption is selected as a target optimization strategy of the target fusion operator under the first architecture parameters.
[0082] Figure 4 A schematic diagram of an example of a chip optimization method provided by at least one embodiment of the present disclosure. For example, Figure 4 for Figure 2 and Figure 3 A specific example of the chip optimization method shown.
[0083] For example, Figure 4 As shown, in process 1, first execute Figure 2 Step S110, obtaining an artificial intelligence model; then executing Figure 3 In step S210, based on the fusion strategy of the operators in the artificial intelligence model, the artificial intelligence model is decomposed into multiple fusion operators; further, executing Figure 3 In steps S220 to S240, when the first architecture parameter of the chip is fixed, multiple fusion operators are traversed, the target optimization strategy corresponding to each fusion operator under the first architecture parameter is selected, the total time consumed by the multiple fusion operators when applying the corresponding target optimization strategy is calculated respectively, and the multiple total time consumptions of the multiple fusion operators are merged to obtain the total time consumption t of the artificial intelligence model under the first architecture parameter.
[0084] For example, Figure 4 As shown, process 2 is a further refinement of the execution process of steps S220 to S240 in process 1. For example, in process 2, the total time consumption t of the artificial intelligence model is first initialized, for example, it is initialized to 0 (ie, t=0); then, the target fusion operator is traversed and selected from multiple fusion operators, and the target optimization strategy and time consumption of each fusion operator are calculated.
[0085] For example, in process 2, for the selected target fusion operator, first execute Figure 2 Step S120, selecting multiple optimization strategies applied to the target fusion operator; then executing Figure 2In steps S130 to S140, when the first architecture parameters are fixed, multiple optimization strategies are traversed to calculate the device-side time consumption of the target fusion operator when each of the multiple optimization strategies is applied. If the device-side time consumption of the target fusion operator when applying the first optimization strategy is the smallest, the first optimization strategy is selected as the target optimization strategy of the target fusion operator, and the optimal device-side time consumption t1 when using the target optimization strategy is obtained; further, execution Figure 2 In step S150, the device-side time consumption and the host-side time consumption of the target fusion operator when applying the target optimization strategy are merged to obtain the total time consumption t2 of the target fusion operator when applying the target optimization strategy; further, the obtained t2 is merged with the current total time consumption t of the artificial intelligence model to update the total time consumption t of the artificial intelligence model (i.e., t=t+t2); further, it is determined whether there are any fusion operators that have not been traversed. If there are any fusion operators that have not been traversed, the next fusion operator is taken as the new target fusion operator to perform traversal calculation according to the above steps. If all fusion operators have been traversed, the updated total time consumption t of the artificial intelligence model is the final total time consumption t of the artificial intelligence model.
[0086] For example, Figure 4 As shown, process 3 is a further refinement of the execution process of steps S130 to S140 in process 2. For example, in process 3, firstly execute step S1401 to set the preset time t1 of the target fusion operator, for example, set the preset time t1 to infinity (ie, t1 = +inf); then execute step S1301 to decompose the target fusion operator into multiple single operators based on the fusion strategy to traverse multiple optimization strategies of the target fusion operator. For example, the multiple optimization strategies include the 1st optimization strategy to the Nth optimization strategy (where N is a positive integer).
[0087] For example, in process 3, the i-th optimization strategy is selected, and step S1302 is first executed to calculate the time consumption of each single operator when applying the i-th optimization strategy based on the first architecture parameter (here i = 1, 2, ..., N). Specifically, based on the operator strategy of each single operator corresponding to the first architecture parameter, the time consumption of each single operator when applying the i-th optimization strategy is calculated; then step S1303 is executed to merge the multiple time consumptions of multiple single operators when applying the i-th optimization strategy to obtain the device-side time consumption t3 of the target fusion operator when applying the i-th optimization strategy; further, step S1402 is executed to merge the target fusion operator at The device-side time consumption t3 when applying the i-th optimization strategy is compared with the current preset time consumption t1, and the smaller value is used to update the preset time consumption t1 (i.e., t1=min(t1,t3)); further, it is determined whether there are any optimization strategies that have not been traversed. If there are any optimization strategies that have not been traversed, the next optimization strategy is traversed and calculated according to the above steps. If all N optimization strategies have been traversed, step S1403 is executed to select the first optimization strategy corresponding to the current preset time consumption t1 as the target optimization strategy of the target fusion operator, and the current preset time consumption t1 is the optimal device-side time consumption t1 of the target fusion operator.
[0088] It should be noted that when executing step S1303, merging multiple time consumptions of multiple single operators is not a simple addition, and the device-side time consumption of the merged target fusion operator may be less than the sum of the multiple time consumptions of the multiple single operators.
[0089] The chip optimization method provided by at least one embodiment of the present disclosure can obtain the optimal optimization strategy corresponding to the fusion operator in the artificial intelligence model when the chip architecture parameters are fixed, so as to guide operator development, operator fusion and software configuration, etc., thereby quickly evaluating the expected performance of the chip and improving the chip evaluation and verification efficiency.
[0090] For example, Figure 2 to Figure 4 The example method shown is to evaluate the performance of multiple fusion operators of an artificial intelligence model under different optimization strategies when the chip's architectural parameters are fixed. In other examples, for a given artificial intelligence model, the performance of the chip under multiple architectural parameters can be further evaluated to obtain the optimal architectural parameters, so as to configure the chip using the optimal architectural parameters.
[0091] Figure 5 Another exemplary flow chart of a chip optimization method provided for at least one embodiment of the present disclosure.
[0092] For example, Figure 5 As shown, the chip optimization method provided by at least one embodiment of the present disclosure may further include the following steps S310 to S340.
[0093] Step S310: Select multiple architecture parameters of the chip.
[0094] For example, in step S310, the plurality of architecture parameters include Figure 2 to Figure 4 The first architecture parameter selected in the example shown; for a given artificial intelligence model, each architecture parameter of the chip can be used as the first architecture parameter, using Figure 2 to Figure 4 The example method shown performs a traversal calculation.
[0095] Step S320: Calculate the total time consumed by the artificial intelligence model when applying each of the multiple architecture parameters.
[0096] For example, in step S320, for each architecture parameter, Figure 2 to Figure 4 The example method shown calculates the total time taken by the artificial intelligence model to apply each architectural parameter.
[0097] For example, each fusion operator of an artificial intelligence model includes multiple single operators, each of which has an operator strategy. For each single operator, multiple architecture parameters of the chip correspond to multiple operator strategies. For example, when applying the same optimization strategy, the chip architecture parameters are different, and the time consumption of each single operator is different.
[0098] Step S330: In response to the artificial intelligence model having the smallest total time consumption when applying the second architecture parameter among the multiple architecture parameters, the second architecture parameter is selected as the target architecture parameter of the chip.
[0099] For example, in step S330, after the traversal of multiple different architecture parameters is completed, if the total time consumption of the artificial intelligence model when applying the second architecture parameter is the smallest, the second architecture parameter is selected as the target architecture parameter of the chip. For example, under the target architecture parameter, the total time consumption of the artificial intelligence model is the smallest and the performance of the chip is the best.
[0100] In some examples, such as Figure 5 As shown, the chip optimization method provided by at least one embodiment of the present disclosure may further include step S340: configuring the chip using target architecture parameters.
[0101] For example, when the hardware architecture of the chip has not yet been formed, the hardware resources can be configured by setting the chip's architecture parameters, that is, how much time each instruction takes can also be configured. For example, in step S340, the chip can be configured using the target architecture parameters to optimize the chip's performance under a given artificial intelligence model.
[0102] In some examples, the plurality of architecture parameters include the 1st architecture parameter to the Mth architecture parameter, where M is a positive integer. For example, Figure 5Step S320 may further include steps S3201 to S3204:
[0103] Step S3201: Decomposing the artificial intelligence model into multiple fusion operators;
[0104] Step S3202: selecting a target optimization strategy corresponding to each fusion operator among the multiple fusion operators under the kth architecture parameter among the multiple architecture parameters, where k=1, 2, ..., M;
[0105] Step S3203: based on the kth architecture parameter, respectively calculating the total time consumed by multiple fusion operators when applying the corresponding target optimization strategies;
[0106] Step S3204: Combine the total time consumed by multiple fusion operators when applying the corresponding target optimization strategies to obtain the total time consumed by the artificial intelligence model under the kth architecture parameter.
[0107] Figure 6 A schematic diagram of another example of a chip optimization method provided by at least one embodiment of the present disclosure. For example, Figure 6 for Figure 5 A specific example of the chip optimization method shown.
[0108] For example, Figure 6 As shown, in process 4, first an artificial intelligence model is selected; then based on the given artificial intelligence model, execution Figure 5 In step S310, multiple architecture parameters of the chip are selected. Specifically, multiple architecture parameters of the chip can be obtained by collecting a list of different architecture parameter combinations. For example, the multiple architecture parameters include the first architecture parameter to the Mth architecture parameter (where M is a positive integer).
[0109] For example, further, in process 4, the traversal calculation is performed in the order of the 1st architecture parameter to the Mth architecture parameter. For example, first traverse and take the kth architecture parameter (here k = 1, 2, ..., M), and then execute Figure 5 Step S320, according to Figure 2 to Figure 4 The example method shown calculates the total time consumed by the artificial intelligence model when applying the kth architectural parameter, and records the calculated total time consumed.
[0110] For example, in step S320, in order to calculate the total time consumed by the artificial intelligence model when applying the kth architectural parameter, specifically, first execute step S3201 to decompose the artificial intelligence model into multiple fusion operators; then execute step S3202 to select the target optimization strategy corresponding to each fusion operator under the kth architectural parameter among the multiple architectural parameters; further, execute step S3203 to calculate the total time consumed by multiple fusion operators when applying the corresponding target optimization strategies based on the kth architectural parameter; further, execute step S3204 to merge the total time consumed by multiple fusion operators when applying the corresponding target optimization strategies to obtain the total time consumed by the artificial intelligence model under the kth architectural parameter.
[0111] For example, further, in process 4, it is determined whether there are still untraversed architecture parameters. If there are still untraversed architecture parameters, the next architecture parameter is traversed and calculated according to the above steps. If all M architecture parameters have been traversed, step S330 is executed to compare the total time consumption of the artificial intelligence model under the M architecture parameters. If the total time consumption of the artificial intelligence model when applying the second architecture parameter is the smallest, the second architecture parameter is selected as the target architecture parameter of the chip. For example, under the target architecture parameter, the total time consumption of the artificial intelligence model is the smallest and the performance of the chip is optimal.
[0112] The chip optimization method provided by at least one embodiment of the present disclosure can, for a given artificial intelligence model, obtain the optimal target architecture parameters that minimize the total time consumption of the artificial intelligence model by traversing multiple architecture parameters of the chip, and use the target architecture parameters to configure the chip to achieve optimal performance, thereby quickly evaluating and guiding chip design.
[0113] Figure 7 A schematic block diagram of a chip optimization device provided for at least one embodiment of the present disclosure.
[0114] For example, Figure 7 As shown, the chip optimization device 200 includes an acquisition module 210, a selection module 220 and a calculation module 230.
[0115] For example, the acquisition module 210 is configured to acquire an artificial intelligence model; in some examples, the artificial intelligence model includes multiple fusion operators, and the multiple fusion operators include a target fusion operator. That is, the acquisition module 210 can be configured to perform, for example Figure 2 Step S110 is shown.
[0116] For example, the selection module 220 is configured to select a first architecture parameter of the chip and a plurality of optimization strategies applied to the target fusion operator. That is, the selection module 220 may be configured to perform, for example Figure 2 Step S120 is shown.
[0117] For example, the calculation module 230 is configured to calculate the device-side time consumption of the target fusion operator when applying each of the multiple optimization strategies based on the first architecture parameter. That is, the calculation module 230 can be configured to perform, for example Figure 2 Step S130 is shown.
[0118] For example, the selection module 220 is further configured to, in response to the target fusion operator having the smallest device-side time consumption when applying the first optimization strategy among the multiple optimization strategies, select the first optimization strategy as the target optimization strategy of the target fusion operator under the first architecture parameter. That is, the selection module 220 can also be configured to perform, for example Figure 2 Step S140 is shown.
[0119] In some examples, such as Figure 7 As shown, the chip optimization device 200 may further include a merging module 240. For example, the merging module 240 is configured to merge the device-side time consumption and the host-side time consumption of the target fusion operator when applying the target optimization strategy to obtain the total time consumption of the target fusion operator when applying the target optimization strategy.
[0120] In some examples, such as Figure 7 As shown, the chip optimization device 200 may also include a disassembly module 250. For example, the disassembly module 250 is configured to disassemble the artificial intelligence model into multiple fusion operators; the selection module 220 is also configured to select the target optimization strategy corresponding to each fusion operator in the multiple fusion operators under the first architecture parameters; the merging module 240 is also configured to merge the total time consumed by the multiple fusion operators when applying the corresponding target optimization strategies to obtain the total time consumed by the artificial intelligence model under the first architecture parameters.
[0121] For example, the target fusion operator includes multiple single operators, and the multiple optimization strategies include the 1st optimization strategy to the Nth optimization strategy, where N is a positive integer. For example, the disassembly module 250 is also configured to disassemble the target fusion operator into multiple single operators; the calculation module 230 is also configured to calculate the time consumption of multiple single operators when applying the i-th optimization strategy among multiple optimization strategies based on the first architecture parameter, where i=1, 2, ..., N; the merging module 240 is also configured to merge the multiple time consumptions of multiple single operators when applying the i-th optimization strategy to obtain the device-side time consumption of the target fusion operator when applying the i-th optimization strategy.
[0122] For example, each of the multiple single operators corresponds to an operator strategy, and the operator strategy of each single operator is associated with the architecture parameters of the chip. For example, the calculation module 230 is further configured to: based on the operator strategy of each single operator corresponding to the first architecture parameter, calculate the time consumed by each single operator when applying the i-th optimization strategy.
[0123] For example, the selection module 220 is also configured to: set a preset time consumption of the target fusion operator, wherein the preset time consumption is initialized to infinity; in accordance with the traversal order from the 1st optimization strategy to the Nth optimization strategy, compare the device-side time consumption of the target fusion operator when applying the i-th optimization strategy among multiple optimization strategies with the preset time consumption, and use the smaller value therein to update the preset time consumption, where i = 1, 2, ..., N; in response to the completion of the traversal from the 1st optimization strategy to the Nth optimization strategy, select the first optimization strategy corresponding to the preset time consumption as the target optimization strategy of the target fusion operator under the first architecture parameters.
[0124] In some examples, the selection module 220 is further configured to select multiple architecture parameters of the chip, for example, the multiple architecture parameters include a first architecture parameter; the calculation module 230 is further configured to calculate the total time consumed by the artificial intelligence model when applying each of the multiple architecture parameters; the selection module 220 is further configured to select the second architecture parameter as the target architecture parameter of the chip in response to the artificial intelligence model having the smallest total time consumed when applying the second architecture parameter among the multiple architecture parameters.
[0125] For example, Figure 7 As shown, the chip optimization device 200 may further include a configuration module 260. For example, the configuration module 260 is configured to configure the chip using target architecture parameters.
[0126] For example, the multiple architecture parameters include the 1st architecture parameter to the Mth architecture parameter, where M is a positive integer. For example, the disassembly module 250 is also configured to disassemble the artificial intelligence model into multiple fusion operators; the selection module 220 is also configured to select the target optimization strategy corresponding to each fusion operator in the multiple fusion operators under the kth architecture parameter in the multiple architecture parameters, where k = 1, 2, ..., M; the merging module 240 is also configured to merge the total time consumed by the multiple fusion operators when applying the corresponding target optimization strategies to obtain the total time consumed by the artificial intelligence model under the kth architecture parameter.
[0127] For example, the target fusion operator includes multiple single operators, and each of the multiple single operators corresponds to an operator strategy; for each of the multiple single operators, multiple architecture parameters correspond to the multiple operator strategies respectively.
[0128] Since in the above description, for example Figure 2 to Figure 6 In the process of the chip optimization method shown in FIG. 1 , the details of the operation of the chip optimization device 200 have been introduced, so for the sake of brevity, they will not be repeated here. For relevant details, please refer to the above Figure 2 to Figure 6 Description.
[0129] It should be noted that Figure 7 The modules in the chip optimization device 200 can be configured as software, hardware, firmware, or any combination of the above items to perform specific functions. For example, these modules can correspond to dedicated integrated circuits, pure software codes, or modules that combine software and hardware. Figure 7 The described device may be a PC computer, a tablet device, a personal digital assistant, a smart phone, a web application, or other device capable of executing program instructions, but is not limited thereto.
[0130] In addition, although the chip optimization device 200 is described above as being divided into modules for performing corresponding processing respectively, it is clear to those skilled in the art that the processing performed by each module can also be performed without any specific module division in the device or without clear demarcation between the modules. Figure 7 The described chip optimization device 200 is not limited to including the modules described above, but some other modules (for example, a reading module, a control module, etc.) may be added as needed, or the above modules may also be combined.
[0131] At least one embodiment of the present disclosure further provides an electronic device, the electronic device comprising a processor and a memory; the memory comprising one or more computer program modules; the one or more computer program modules are stored in the memory and configured to be executed by the processor, and the one or more computer program modules include a method for implementing the chip optimization method provided by the embodiment of the present disclosure described above. For example, the processor may be a single-core processor or a multi-core processor.
[0132] Figure 8 A schematic block diagram of an electronic device provided for at least one embodiment of the present disclosure.
[0133] For example, Figure 8 As shown, the electronic device 300 includes a processor 310 and a memory 320. For example, the memory 320 is used to store non-transitory computer-readable instructions (e.g., one or more computer program modules). The processor 310 is used to run non-transitory computer-readable instructions, and when the non-transitory computer-readable instructions are run by the processor 310, one or more steps of the chip optimization method described above can be executed. The memory 320 and the processor 310 can be interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0134] For example, the processor 310 may be a central processing unit (CPU), a graphics processing unit (GPU), a general purpose graphics processing unit (GPGPU), a digital signal processor (DSP), or other forms of processing units with chip optimization capabilities and / or program execution capabilities, such as a field programmable gate array (FPGA), etc.; for example, the central processing unit (CPU) may be an X86, RISC-V, or ARM architecture, etc. The processor 310 may be a general-purpose processor or a dedicated processor, and may control other components in the electronic device 300 to perform desired functions.
[0135] For example, the memory 320 may include any combination of one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules may be stored on the computer-readable storage medium, and the processor 310 may run one or more computer program modules to implement various functions of the electronic device 300. Various applications and various data, as well as various data used and / or generated by the application, etc. may also be stored in the computer-readable storage medium.
[0136] It should be noted that in the embodiments of the present disclosure, the specific functions and technical effects of the electronic device 300 can refer to the above description of the chip optimization method provided in at least one embodiment of the present disclosure, and will not be repeated here.
[0137] Fig. 9 A schematic block diagram of another electronic device provided for at least one embodiment of the present disclosure.
[0138] For example, Fig. 9 As shown, the electronic device 400 is suitable for implementing the chip optimization method provided by the embodiment of the present disclosure. It should be noted that, Fig. 9 The electronic device 400 shown is only an example and does not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0139] For example, Fig. 9As shown, the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 41, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 42 or a program loaded from a storage device 48 into a random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 41, the ROM 42, and the RAM 43 are connected to each other through a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44. Generally, the following devices can be connected to the I / O interface 45: an input device 46 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 47 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 48 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 49. The communication device 49 can allow the electronic device 400 to communicate with other electronic devices wirelessly or by wire to exchange data.
[0140] Although Fig. 9 The electronic device 400 is shown to have various devices, but it should be understood that it is not required to implement or possess all the devices shown, and the electronic device 400 may alternatively implement or possess more or fewer devices.
[0141] For the detailed description and technical effects of the electronic device 400, please refer to the above description of the chip optimization method, which will not be repeated here.
[0142] Fig.10 A schematic diagram of a storage medium provided for at least one embodiment of the present disclosure.
[0143] For example, Fig.10 As shown, the storage medium 500 stores non-transitory computer-readable instructions 510. For example, when the non-transitory computer-readable instructions 510 are executed by a computer, one or more steps in the chip optimization method described above are performed.
[0144] For example, the storage medium 500 can be applied to Figure 8 For example, the storage medium 500 may be the memory 320 in the electronic device 300. For example, the description of the storage medium 500 may refer to Figure 8 The corresponding description of the memory 320 in the electronic device 300 is shown and will not be repeated here.
[0145] There are a few points to note about this disclosure:
[0146] (1) In the drawings of the embodiments of the present disclosure, only the structures related to the embodiments of the present disclosure are involved, and other structures can refer to the general design.
[0147] (2) In the absence of conflict, features in the same embodiment or in different embodiments of the present disclosure may be combined with each other.
[0148] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present disclosure, which should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.
Claims
1. A chip optimization method, used to optimize the performance of an artificial intelligence model corresponding to the chip before the chip leaves the factory, wherein: The chip optimization method comprises: Acquire the artificial intelligence model, wherein the artificial intelligence model includes a plurality of fusion operators, and the plurality of fusion operators include a target fusion operator; Selecting a first architecture parameter of the chip and a plurality of optimization strategies applied to the target fusion operator; Based on the first architecture parameter, respectively calculating the device-side time consumed by the target fusion operator when applying each of the multiple optimization strategies; In response to the target fusion operator having the smallest device-side time consumption when applying a first optimization strategy among the multiple optimization strategies, selecting the first optimization strategy as the target optimization strategy of the target fusion operator under the first architecture parameters, The target fusion operator includes multiple single operators, and the multiple optimization strategies include the first optimization strategy to the Nth optimization strategy, where N is a positive integer. The calculating, based on the first architecture parameter, respectively the device-side time consumption of the target fusion operator when applying each of the multiple optimization strategies includes: Decomposing the target fusion operator into the multiple single operators; Based on the first architecture parameter, calculating the time consumed by the multiple single operators when applying the i-th optimization strategy among the multiple optimization strategies, where i=1, 2, ..., N; The multiple time consumptions of the multiple single operators when applying the i-th optimization strategy are merged to obtain the device-side time consumption of the target fusion operator when applying the i-th optimization strategy.
2. The chip optimization method according to claim 1, further comprising: The device-side time consumption and the host-side time consumption of the target fusion operator when applying the target optimization strategy are combined to obtain the total time consumption of the target fusion operator when applying the target optimization strategy.
3. The chip optimization method according to claim 1, further comprising: Decomposing the artificial intelligence model into the multiple fusion operators; Selecting a target optimization strategy corresponding to each fusion operator of the multiple fusion operators under the first architecture parameters; Based on the first architecture parameter, respectively calculating the total time consumed by the multiple fusion operators when applying the corresponding target optimization strategies; The total time consumed by the multiple fusion operators when applying the corresponding target optimization strategies is combined to obtain the total time consumed by the artificial intelligence model under the first architecture parameters.
4. The chip optimization method according to claim 1, wherein: Each of the multiple single operators corresponds to an operator strategy, and the operator strategy of each single operator is associated with an architecture parameter of the chip, The calculating, based on the first architecture parameter, the time consumption of the multiple single operators when applying the i-th optimization strategy among the multiple optimization strategies includes: Based on the operator strategy of each single operator corresponding to the first architecture parameter, the time consumption of each single operator when applying the i-th optimization strategy is calculated.
5. The chip optimization method according to claim 1, wherein: The multiple optimization strategies include the first optimization strategy to the Nth optimization strategy, wherein N is a positive integer. In response to the target fusion operator having the smallest device-side time consumption when applying a first optimization strategy among the multiple optimization strategies, selecting the first optimization strategy as the target optimization strategy of the target fusion operator under the first architecture parameters includes: Setting a preset time consumption of the target fusion operator, wherein the preset time consumption is initialized to infinity; According to the traversal order of the first optimization strategy to the Nth optimization strategy, the device-side time consumption of the target fusion operator when applying the i-th optimization strategy among the multiple optimization strategies is compared with the preset time consumption, and the smaller value thereof is used to update the preset time consumption, where i=1, 2, ..., N; In response to the traversal of the first optimization strategy to the Nth optimization strategy being completed, the first optimization strategy corresponding to the preset time consumption is selected as the target optimization strategy of the target fusion operator under the first architecture parameters.
6. The chip optimization method according to claim 1, further comprising: Selecting a plurality of architecture parameters of the chip, wherein the plurality of architecture parameters include the first architecture parameter; Calculating the total time consumed by the artificial intelligence model when applying each of the plurality of architecture parameters; In response to the artificial intelligence model taking the least total time when applying a second architecture parameter among the multiple architecture parameters, the second architecture parameter is selected as the target architecture parameter of the chip.
7. The chip optimization method according to claim 6, further comprising: The chip is configured using the target architecture parameters.
8. The chip optimization method according to claim 6, wherein: The plurality of architecture parameters include the 1st architecture parameter to the Mth architecture parameter, wherein M is a positive integer, The calculating the total time consumed by the artificial intelligence model when applying each of the plurality of architecture parameters comprises: Decomposing the artificial intelligence model into the multiple fusion operators; Selecting a target optimization strategy corresponding to each fusion operator of the multiple fusion operators under a k-th architecture parameter of the multiple architecture parameters, where k=1, 2, ..., M; Based on the kth architecture parameter, respectively calculating the total time consumed by the multiple fusion operators when applying the corresponding target optimization strategies; The total time consumed by the multiple fusion operators when applying the corresponding target optimization strategies is combined to obtain the total time consumed by the artificial intelligence model under the kth architecture parameter.
9. The chip optimization method according to claim 6, wherein: The target fusion operator includes multiple single operators, each of the multiple single operators corresponds to an operator strategy, For each single operator of the multiple single operators, the multiple architecture parameters correspond to multiple operator strategies respectively.
10. A chip optimization device, used to optimize the performance of an artificial intelligence model corresponding to the chip before the chip leaves the factory, wherein: The chip optimization device comprises: An acquisition module configured to acquire the artificial intelligence model, wherein the artificial intelligence model includes a plurality of fusion operators, and the plurality of fusion operators include a target fusion operator; A selection module configured to select a first architecture parameter of the chip and a plurality of optimization strategies applied to the target fusion operator; a calculation module configured to calculate, based on the first architecture parameter, the device-side time consumption of the target fusion operator when applying each of the multiple optimization strategies, The selection module is further configured to, in response to the target fusion operator having the smallest device-side time consumption when applying the first optimization strategy among the multiple optimization strategies, select the first optimization strategy as the target optimization strategy of the target fusion operator under the first architecture parameter, The target fusion operator includes multiple single operators, and the multiple optimization strategies include the first optimization strategy to the Nth optimization strategy, where N is a positive integer. The computing module is further configured as: Decomposing the target fusion operator into the multiple single operators; Based on the first architecture parameter, calculating the time consumed by the multiple single operators when applying the i-th optimization strategy among the multiple optimization strategies, where i=1, 2, ..., N; The multiple time consumptions of the multiple single operators when applying the i-th optimization strategy are merged to obtain the device-side time consumption of the target fusion operator when applying the i-th optimization strategy.
11. An electronic device, comprising: processor; a memory including one or more computer program modules; The one or more computer program modules are stored in the memory and configured to be executed by the processor, and the one or more computer program modules are used to implement the chip optimization method according to any one of claims 1 to 9.
12. A storage medium storing non-transitory computer-readable instructions, which, when executed by a computer, implements the chip optimization method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method and device for optimizing vectorization in compiling process and electronic equipment
CN112947932A
Server, client, system, method, device and medium for strategy optimization
CN115225497A
Operator fusion method and device, electronic equipment and computer readable medium
CN116933841A