Data Processing Method, Apparatus, and Storage Medium
By obtaining the calculation diagram and processor information of the neural network model, determining the intermediate representation and optimizing the processor, the problem that existing compilers are difficult to support new chips and neural network models is solved, and efficient compilation and processor optimization are achieved.
Patent Information
- Application Number
- CN202210500937.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-09
AI Technical Summary
Existing neural network compilers are difficult to efficiently support brain-like chips and flexible neural network models, such as SNN and HNN.
By obtaining the calculation diagram of the neural network model and the processor information, the first intermediate representation (including the operator supported by the processor), and then the second intermediate representation (including the time scheduling information and spatial mapping information of the operator), the target instructions are finally determined and the processor is optimized.
It realizes efficient compilation and deployment of new neural network models onto new chips, and guides processor optimization, improving the flexibility and efficiency of the compilation process.
Smart Images

Figure CN114970847B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a data processing method, apparatus, and storage medium. Background Art
[0002] With the continuous development of artificial intelligence technology, many new hardware platforms and neural network models have emerged. The new hardware platforms (i.e., chips) are represented by rapidly iterating and developing graphics processing units (GPUs), multi-core chips, many-core chips, deep learning accelerators, neuromorphic chips, brain-like computing chips, etc. The new neural network models are represented by artificial neural networks, spiking neural networks (SNNs), hybrid neural networks (HNNs), etc.
[0003] In order to enable neural network models to run efficiently on hardware platforms, many neural network compilation processes and methods have been proposed, such as TVM (tensor virtual machine), to deploy deep learning models under different neural network programming frameworks to hardware platforms in the most efficient way possible. However, current neural network compilers still cannot provide general and efficient support for many-core chips such as brain-like chips, and flexible neural network models such as SNNs and HNNs. Summary of the Invention
[0004] In view of this, the present disclosure provides a data processing method, apparatus, and storage medium.
[0005] According to one aspect of the present disclosure, a data processing method is provided. The method includes: obtaining a computational graph of a neural network model, where the computational graph represents the structure of the neural network model;
[0006] obtaining information about a processor;
[0007] determining a first intermediate representation according to the computational graph and the information about the processor, where the first intermediate representation includes operators supported by the processor;
[0008] determining a second intermediate representation according to the first intermediate representation and the information about the processor, where the second intermediate representation includes time scheduling information and spatial mapping information of the operators on the processor;
[0009] determining a target instruction corresponding to the neural network model according to the second intermediate representation and the information about the processor, where the target instruction is used to be executed by the processor;
[0010] Optimize the processor according to at least one of the computation graph, the first intermediate representation, the second intermediate representation, the target instruction, and the information of the processor.
[0011] According to the embodiments of the present application, by obtaining the computation graph of the neural network model and the information of the processor, the compilation of the neural network model can be realized, so that the new neural network model can be efficiently deployed on the new chip. At the same time, during the compilation process, the processor can be guided to be optimized to achieve the collaborative optimization of the compilation process and the processor, making the compilation process flexible and efficient. At the same time, by determining the first intermediate representation, the second intermediate representation, and finally determining the target instruction during the compilation process, the compilation process can be made hierarchical, with different levels targeting different compilation requirements, making the compilation more targeted and achieving better compilation effects.
[0012] In a possible implementation manner, the neural network model includes at least one of an artificial neural network model ANN, a spiking neural network model SNN, and a hybrid neural network model HNN, and the processor includes at least one of a graphics processing unit GPU, a neuromorphic chip, a brain-like computing chip, and a many-core neural network acceleration chip.
[0013] Thus, the compilation of the new neural network model can be realized, so that the new neural network model can be deployed on the new chip to meet the needs of the new neural network model and the new chip.
[0014] In a possible implementation manner, the information of the processor includes the operator information supported by the processor, and the operator information includes at least one of the operator type and the data precision of the operator. Determining the first intermediate representation according to the computation graph and the information of the processor includes:
[0015] Determine the first intermediate representation according to the computation graph and the operator information supported by the processor.
[0016] According to the embodiments of the present application, by using the operator information supported by the processor, the operators in the computation graph can be converted into the operators supported by the processor to determine the first intermediate representation, so that the operator information supported by the processor can be considered during the compilation optimization process, hierarchical compilation can be realized, and the compilation process is more in line with the specific hardware characteristics, more targeted, and the compilation optimization effect is better.
[0017] In a possible implementation manner, determining the first intermediate representation according to the computation graph and the operator information supported by the processor includes:
[0018] Perform operator type conversion and / or operator quantization on the operators in the computational graph according to the operator information supported by the processor, and determine the first intermediate representation.
[0019] According to the embodiments of the present application, by determining the first intermediate representation according to the computational graph and the operator information supported by the processor, the neural network model can be compiled for the operator types, data precisions, and requirements for the computing precision of the processor supported by the processor, making the compilation process more targeted and having a better compilation effect.
[0020] In a possible implementation, the information of the processor includes the resource information of the processor, and the resource information includes at least one of memory resource information, computing resource information, and routing resource information. Determining the second intermediate representation according to the first intermediate representation and the information of the processor includes:
[0021] Determine the second intermediate representation according to the first intermediate representation and the resource information.
[0022] According to the embodiments of the present application, by using the resource information of the processor, it is possible to determine the second intermediate representation with spatio-temporal mapping information. Considering the resource information of the processor during the compilation optimization process can achieve hierarchical compilation, and make the compilation process more in line with the specific hardware characteristics, more targeted, and have a better compilation optimization effect.
[0023] In a possible implementation, the information of the processor includes the behavioral characteristic information of the processor, and the behavioral characteristic information includes the operations executable by the processor. Determining the target instruction corresponding to the neural network model according to the second intermediate representation and the information of the processor includes:
[0024] Determine the target instruction according to the second intermediate representation and the behavioral characteristic information.
[0025] According to the embodiments of the present application, by using the behavioral characteristic information of the processor, it is possible to determine the target instruction. Considering the behavioral characteristic information of the processor during this process can achieve hierarchical compilation, and make the compilation process more in line with the specific hardware characteristics, more targeted, and have a better compilation optimization effect.
[0026] In a possible implementation, determining the target instruction according to the second intermediate representation and the behavioral characteristic information includes:
[0027] Perform at least one of memory address allocation, routing path planning, computing pipeline optimization, and register allocation on the second intermediate representation according to the second intermediate representation and the behavioral characteristic information, and determine the target instruction.
[0028] According to the embodiments of the present application, by performing memory address allocation, routing path planning, computing pipeline optimization, and register allocation, the compilation effect can be better, and the performance of the obtained target instructions on the processor is better.
[0029] In a possible implementation manner, optimizing the processor according to at least one of the computation graph, the first intermediate representation, the second intermediate representation, the target instruction, and the information of the processor includes:
[0030] Determine whether the processor meets the operator function support requirements for the neural network model according to the computation graph and the operator information supported by the processor;
[0031] When the operator function support requirements for the neural network model are not met, optimize the processor.
[0032] According to the embodiments of the present application, it is possible to optimize the processor from the perspective of hardware operator design, so as to more specifically achieve the collaborative optimization of the processor and the compilation process.
[0033] In a possible implementation manner, optimizing the processor according to at least one of the computation graph, the first intermediate representation, the second intermediate representation, the target instruction, and the information of the processor includes:
[0034] Determine whether the processor meets the computing power support requirements for the neural network model according to the first intermediate representation and the resource information;
[0035] When the computing power support requirements for the neural network model are not met, optimize the processor.
[0036] According to the embodiments of the present application, it is possible to optimize the processor from the perspective of processor resources, so as to more specifically achieve the collaborative optimization of the processor and the compilation process.
[0037] In a possible implementation manner, optimizing the processor according to at least one of the computation graph, the first intermediate representation, the second intermediate representation, the target instruction, and the information of the processor includes:
[0038] Determine the accuracy of the algorithm of the neural network model when executed on the processor according to the second intermediate representation and the behavior characteristic information;
[0039] Optimize the processor according to the accuracy.
[0040] According to the embodiments of the present application, it is possible to optimize the processor from the perspective of the execution of the processor, so that the cooperative optimization of the processor and the compilation process can be more targeted.
[0041] In a possible implementation manner, the information of the processor includes the clock information of the processor. Optimizing the processor according to at least one of the computational graph, the first intermediate representation, the second intermediate representation, and the target instruction, and the information of the processor includes:
[0042] Determine the time performance of the algorithm of the neural network model executed on the processor according to the clock information and the target instruction;
[0043] Optimize the processor according to the time performance.
[0044] According to the embodiments of the present application, it is possible to optimize the processor from the perspective of clock accuracy, so that the optimization of the processor can be more targeted.
[0045] According to another aspect of the present disclosure, a data processing device is provided. The device includes:
[0046] A first acquisition module, configured to acquire a computational graph of a neural network model, where the computational graph represents the structure of the neural network model;
[0047] A second acquisition module, configured to acquire information of the processor;
[0048] A first determination module, configured to determine a first intermediate representation according to the computational graph and the information of the processor, where the first intermediate representation includes operators supported by the processor;
[0049] A second determination module, configured to determine a second intermediate representation according to the first intermediate representation and the information of the processor, where the second intermediate representation includes time scheduling information and spatial mapping information of the operator on the processor;
[0050] A third determination module, configured to determine a target instruction corresponding to the neural network model according to the second intermediate representation and the information of the processor, where the target instruction is used to be executed by the processor;
[0051] An optimization module, configured to optimize the processor according to at least one of the computational graph, the first intermediate representation, the second intermediate representation, and the target instruction, and the information of the processor.
[0052] In a possible implementation, the neural network model includes at least one of an artificial neural network model ANN, a spiking neural network model SNN, and a hybrid neural network model HNN, and the processor includes at least one of a graphics processing unit GPU, a neuromorphic chip, a brain-inspired computing chip, and a many-core neural network acceleration chip.
[0053] In a possible implementation, the information of the processor includes the operator information supported by the processor, and the operator information includes at least one of the operator type and the data precision of the operator. The first determination module is configured to:
[0054] Determine the first intermediate representation according to the computational graph and the operator information supported by the processor.
[0055] In a possible implementation, determining the first intermediate representation according to the computational graph and the operator information supported by the processor includes:
[0056] Perform operator type conversion and / or operator quantization on the operators in the computational graph according to the operator information supported by the processor, and determine the first intermediate representation.
[0057] In a possible implementation, the information of the processor includes the resource information of the processor, and the resource information includes at least one of memory resource information, computing resource information, and routing resource information. The second determination module is configured to:
[0058] Determine the second intermediate representation according to the first intermediate representation and the resource information.
[0059] In a possible implementation, the information of the processor includes the behavioral characteristic information of the processor, and the behavioral characteristic information includes the operations executable by the processor. The third determination module is configured to:
[0060] Determine the target instruction according to the second intermediate representation and the behavioral characteristic information.
[0061] In a possible implementation, determining the target instruction according to the second intermediate representation and the behavioral characteristic information includes:
[0062] Perform at least one of memory address allocation, routing path planning, computing pipeline optimization, and register allocation on the second intermediate representation according to the second intermediate representation and the behavioral characteristic information, and determine the target instruction.
[0063] In a possible implementation, the optimization module is configured to:
[0064] Determine whether the processor meets the operator function support requirements for the neural network model according to the computational graph and the operator information supported by the processor;
[0065] Optimize the processor when the operator function support requirement for the neural network model is not met.
[0066] In a possible implementation, an optimization module is configured to:
[0067] Determine whether the processor meets the computing power support requirement for the neural network model according to the first intermediate representation and the resource information;
[0068] Optimize the processor when the computing power support requirement for the neural network model is not met.
[0069] In a possible implementation, an optimization module is configured to:
[0070] Determine the accuracy of the algorithm of the neural network model when executed on the processor according to the second intermediate representation and the behavior characteristic information;
[0071] Optimize the processor according to the accuracy.
[0072] In a possible implementation, the information of the processor includes the clock information of the processor, and the optimization module is configured to:
[0073] Determine the time performance of the algorithm of the neural network model when executed on the processor according to the clock information and the target instruction;
[0074] Optimize the processor according to the time performance.
[0075] According to another aspect of the present disclosure, a data processing device is provided, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to implement the above method when executing the instructions stored in the memory.
[0076] According to another aspect of the present disclosure, a non - volatile computer - readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0077] According to another aspect of the present disclosure, a computer program product is provided, including computer - readable code, or a non - volatile computer - readable storage medium carrying the computer - readable code, when the computer - readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0078] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. Description of the Drawings
[0079] The drawings included in and constituting a part of the specification illustrate exemplary embodiments, features, and aspects of the present disclosure together with the specification, and are used to explain the principles of the present disclosure.
[0080] Figure 1 A schematic diagram showing an application scenario according to an embodiment of the present application.
[0081] Figure 2 A flowchart showing a data processing method according to an embodiment of the present application.
[0082] Figure 3 A structural diagram showing a data processing apparatus according to an embodiment of the present application.
[0083] Figure 4 A block diagram showing a data processing apparatus 1900 according to an embodiment of the present application. Detailed Description of the Invention
[0084] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the drawings. Like reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0085] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.
[0086] In addition, for a better description of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0087] With the continuous development of artificial intelligence technology, many new types of chips and neural network models have emerged. In order to enable neural network models to run efficiently on chips with increasing design complexity, neural network compilers cannot provide general and efficient support for new types of chips such as brain-like chips, nor for flexible neural network models such as SNN and HNN, and cannot verify the design of the chips.
[0088] In view of this, the present application provides a data processing method. According to the method of the embodiments of the present application, a computational graph of a new neural network model and information of a new chip (i.e., a processor) can be obtained. Through a hierarchical compilation process such as operator function conversion, spatio-temporal resource mapping, and hardware code generation, the compilation of the neural network model can be realized, enabling these new neural network models to be efficiently deployed on the new chip. At the same time, during the compilation process, the processor can also be guided to perform optimization to achieve collaborative optimization between the compilation process and the chip, making the compilation process flexible and efficient.
[0089] Figure 1 FIG. shows a schematic diagram of an application scenario according to an embodiment of the present application. The data processing method of the embodiments of the present application can be used in the process of the data processing system 100 compiling a neural network model for deployment on the processor 200 (i.e., the chip). As Figure 1 shown, the data processing system 100 of the embodiments of the present application may include a storage module 101, a processor emulation module 102, and a compilation module 103. Among them, the storage module 101 can be used to store the computational graph of the neural network model and the intermediate representation during the compilation process; the processor emulation module 102 can be used to store the information of the processor, and can also be used to perform emulation based on at least one of the computational graph of the neural network model, the intermediate representation, and the target instruction and the information of the processor to determine the information for optimizing the processor; the compilation module 103 can be used to determine the target instruction corresponding to the neural network model according to the computational graph (or the intermediate representation) and the information of the processor (as shown in the figure), and this target instruction can be executed by the processor 200, so that the neural network model can be deployed on the processor 200.
[0090] In the above process, the neural network model can include an artificial neural network (ANN), a spiking neural network model SNN, a hybrid neural network model HNN, a convolutional neural network (CNN), a recurrent neural network (RNN), etc., and the present application does not limit this. The processor can include a graphics processing unit GPU, a neuromorphic chip, a brain-inspired computing chip, a central processing unit (CPU), a many-core neural network acceleration chip, etc., and the present application also does not limit this.
[0091] The following will Figure 1 be used as a basis to introduce the data processing method of the embodiments of the present application in detail.
[0092] See Figure 2, showing a flowchart of a data processing method according to an embodiment of the present application. This data processing method can be used in the above-mentioned data processing system 100, such as Figure 2 As shown, the method includes:
[0093] Step S201, obtaining the computation graph of the neural network model.
[0094] Among them, the computation graph (computation graph) can represent the structure of the neural network model. In a possible implementation, the neural network model can include at least one of an artificial neural network model ANN, a spiking neural network model SNN, and a hybrid neural network model HNN.
[0095] The computation graph of the neural network model can be obtained, for example, from Figure 1 the storage module 101 shown. For example, the basic operation unit of the computation graph can be an operator, and the basic data structure can be a tensor. The computation graph can be composed of nodes and edges (nodes represent operators, and edges represent tensors), and the structure of the neural network model is represented by a data flow graph.
[0096] In a possible implementation, the process of compiling the neural network model can be divided into multiple levels according to different compilation requirements, and different intermediate representations (intermediate representation, IR) can be obtained at different levels. See the following for details.
[0097] The present application does not limit the type of the neural network model.
[0098] Step S202, obtaining information about the processor.
[0099] In a possible implementation, the processor includes at least one of a graphics processing unit GPU, a neuromorphic chip, a brain-inspired computing chip, and a many-core neural network acceleration chip. The present application does not limit the type of the processor.
[0100] Among them, the information about the processor can be obtained, for example, from Figure 1 the processor emulation module 102 shown. The information about the processor can include operator information supported by the processor, resource information of the processor, etc. The method of obtaining the information about the processor can be implemented based on the prior art. During the compilation process, according to the compilation requirements at different levels, the information about the processor obtained at different levels can be different. See the following for details.
[0101] Step S203, determining a first intermediate representation according to the computation graph and the information about the processor.
[0102] Among them, the first intermediate representation includes operators supported by the processor. For the detailed process of determining the first intermediate representation according to the computational graph and the information of the processor, refer to the following.
[0103] In a possible implementation, the information of the processor may include operator information supported by the processor, and the operator information may include at least one of operator type and data precision of the operator. For example, the information of the processor may indicate which types and data precisions of operators the processor supports, and so on. This step S203 includes:
[0104] Determine the first intermediate representation according to the computational graph and the operator information supported by the processor.
[0105] Among them, based on the operator information supported by the processor (such as including the operator type and data precision supported by the processor), the operators in the computational graph can be converted into operators supported by the processor to determine the first intermediate representation. In a possible implementation, the first intermediate representation obtained in this way can be understood as the computational graph obtained after converting the operators in the computational graph into operators supported by the processor.
[0106] According to the embodiments of the present application, by using the operator information supported by the processor, it is possible to convert the operators in the computational graph into operators supported by the processor to determine the first intermediate representation, so that the operator information supported by the processor can be considered during the compilation optimization process, hierarchical compilation can be achieved, and the compilation process is more in line with the specific hardware characteristics, more targeted, and the effect of compilation optimization is better.
[0107] In a possible implementation, determining the first intermediate representation according to the computational graph and the operator information supported by the processor includes:
[0108] Perform operator type conversion and / or operator quantization on the operators in the computational graph according to the operator information supported by the processor to determine the first intermediate representation.
[0109] For example, the operator with function A in the computational graph can be operator quantized to make its data precision conform to the data precision of the operator with function A supported by the processor. Also, the operator with function A in the computational graph can be operator type converted into operator types B and C supported by the processor, where operators B and C can equivalently or approximately implement function A.
[0110] Among them, the process of performing operator type conversion and / or operator quantization (i.e., converting floating-point operations to fixed-point operations) on the operators in the computational graph can be implemented based on existing technologies. In a possible implementation, compilation optimization can also be performed during this process. The ways of performing compilation optimization can include, for example, operator fusion, redundant node elimination, dimension alignment, etc.
[0111] According to an embodiment of the present application, by determining a first intermediate representation based on a computational graph and operator information supported by a processor, the neural network model can be compiled according to the operator types supported by the processor, data precision, and requirements for the computing precision of the processor, etc., making the compilation process more targeted and achieving better compilation results.
[0112] Step S204: Determine a second intermediate representation according to the first intermediate representation and the information of the processor.
[0113] Among them, the second intermediate representation may include time scheduling information and spatial mapping information of the operator on the processor. The time scheduling information may represent the execution order of the node tasks corresponding to each operator on the processor, and the spatial mapping information may represent the node tasks corresponding to each operator and the computing units (i.e., cores) for executing the node tasks. For the detailed process of determining the second intermediate representation according to the first intermediate representation and the information of the processor, refer to the following.
[0114] In a possible implementation manner, the information of the processor may include resource information of the processor, and the resource information may include at least one of memory resource information, computing resource information, and routing resource information. This step S204 includes:
[0115] Determine a second intermediate representation according to the first intermediate representation and the resource information.
[0116] Among them, the memory resource information may, for example, include the size of the processor memory and the memory configuration information of each partition. The memory configuration information of each partition may, for example, include the memory partition allocation method and the memory capacity size configured for each partition. The computing resource information may, for example, include parallelism information, synchronization and asynchronous information, pipeline information, etc. when processing data. The pipeline information may, for example, include the number of pipeline stages; the routing resource model may, for example, include data transmission path information between each computing unit, etc. Based on the resource information of the processor, each node task in the first intermediate representation can be mapped to the corresponding time step and corresponding computing unit of the processor to determine the second intermediate representation.
[0117] According to an embodiment of the present application, by using the resource information of the processor, it is possible to realize the determination of the second intermediate representation with spatio-temporal mapping information. Considering the resource information of the processor during the compilation optimization process can achieve hierarchical compilation, and make the compilation process more in line with the specific hardware characteristics, more targeted, and achieve better compilation optimization results.
[0118] Step S205: Determine the target instruction corresponding to the neural network model according to the second intermediate representation and the information of the processor.
[0119] Among them, the target instruction is used to be executed by the processor. The target instruction may include code corresponding to a target language executable by the processor. For the detailed process of determining the target instruction according to the second intermediate representation and the information of the processor, please refer to the following.
[0120] In a possible implementation, the information of the processor may include the behavioral characteristic information of the processor. The behavioral characteristic information includes operations executable by the processor. This step S205 includes:
[0121] Determine the target instruction according to the second intermediate representation and the behavioral characteristic information.
[0122] Among them, the behavioral characteristic information of the processor may be relevant information during the operation of the processor (such as memory access logic, on-chip network routing policy, data storage method, operating frequency during operation, etc.) obtained based on the operator information supported by the above-mentioned processor and the resource information of the processor. Determining the target instruction according to the second intermediate representation and the behavioral characteristic information may include compiling the second intermediate representation based on the behavioral characteristic information, and converting the second intermediate representation into a target instruction executable by hardware. The process of converting the second intermediate representation into a target instruction may be implemented based on the prior art, and this application does not limit it.
[0123] According to the embodiments of the present application, by using the behavioral characteristic information of the processor, it is possible to determine the target instruction. During this process, considering the behavioral characteristic information of the processor, hierarchical compilation can be achieved, and the compilation process is more in line with the specific hardware characteristics, more targeted, and the effect of compilation optimization is better.
[0124] In a possible implementation, determining the target instruction according to the second intermediate representation and the behavioral characteristic information further includes:
[0125] Perform at least one of memory address allocation, routing path planning, calculation pipeline optimization, and register allocation on the second intermediate representation according to the second intermediate representation and the behavioral characteristic information, and determine the target instruction.
[0126] Among them, the methods of performing memory address allocation, routing path planning, calculation pipeline optimization, and register allocation may be implemented based on the prior art. For example, the optimal memory placement can be solved according to integer linear programming, and memory optimization can also be performed by using methods such as minimum memory allocation based on topological order. This application does not limit it.
[0127] According to the embodiments of the present application, by performing memory address allocation, routing path planning, calculation pipeline optimization, and register allocation, the compilation effect can be better, and the performance of the obtained target instruction on the processor is better.
[0128] During the compilation of the neural network model, the processor can also be guided to perform iterative optimization, as described below.
[0129] Step S206: Optimize the processor according to at least one of the computational graph, the first intermediate representation, the second intermediate representation, the target instruction, and the information of the processor.
[0130] Among them, according to at least one of the computational graph, the first intermediate representation, the second intermediate representation, the target instruction, and the information of the processor, the execution result of at least one of the computational graph, the first intermediate representation, the second intermediate representation, and the target instruction on the processor can be simulated to determine the simulation result, that is, the information executed by the processor. The processor can be optimized according to this simulation result.
[0131] In a possible implementation, the above compilation process can also be iteratively optimized according to the above simulation result, and one or more steps in steps S203 - S205 are re-executed until a predetermined iteration stop condition is met, and a better compilation result (i.e., intermediate representation and / or target instruction) is obtained. Thus, the compilation process and the processor can be iteratively and collaboratively optimized until a predetermined iteration stop condition is met, and the optimized compilation result and processor are obtained respectively. During this process, the process of determining the simulation result in step S206 can be executed after one or more steps in steps S203 - S205. For example, when determining the first intermediate representation, steps S203 and S206 can be iteratively executed until a predetermined convergence condition is met, and the optimized first intermediate representation is obtained, and then step S204 is executed. The process of determining the second intermediate representation and the target instruction is the same.
[0132] According to the embodiments of the present application, by obtaining the computational graph of the neural network model and the information of the processor, the compilation of the neural network model is realized, so that the new neural network model can be efficiently deployed on the new chip. At the same time, during the compilation process, the processor can also be guided to be optimized to realize the collaborative optimization of the compilation process and the processor, making the compilation process flexible and efficient. At the same time, by determining the first intermediate representation, the second intermediate representation, and finally the target instruction during the compilation process, the compilation process can be made hierarchical, and different levels are aimed at different compilation requirements, making the compilation more targeted and obtaining better compilation effects.
[0133] Referring to the above hierarchical compilation process, the process of optimizing the processor in step S206 can also be hierarchical, that is, the performance of the processor can be simulated at different levels, so as to guide the processor to perform different degrees and different emphases of optimization, so as to realize the collaborative optimization of the compilation process and the processor at different levels.
[0134] In a possible implementation, step S206 may include:
[0135] Determine whether the processor meets the operator function support requirements for the neural network model according to the computational graph and the operator information supported by the processor; optimize the processor when the operator function support requirements for the neural network model are not met.
[0136] Among them, it is possible to verify the degree of demand support of the operators supported by the processor for each operator in the computational graph. For example, according to the computational graph and the operator information supported by the processor, simulate the execution process on the processor, so as to determine the coverage rate of the functions of the operators supported by the processor for the requirements of each operator in the computational graph. The coverage rate can be obtained, for example, by calculating the ratio of the number of operators in the computational graph that can be converted into the operators supported by the processor to the total number of operators in the computational graph, or obtained by other means. When the operator function support requirements for the neural network model are not met, for example, the processor can be optimized by designing new operators in the processor.
[0137] According to the embodiments of the present application, it is possible to optimize the processor from the perspective of hardware operator design, so as to more specifically achieve the collaborative optimization of the processor and the compilation process.
[0138] In a possible implementation, step S206 may include:
[0139] Determine whether the processor meets the computing power support requirements for the neural network model according to the first intermediate representation and the resource information; optimize the processor when the computing power support requirements for the neural network model are not met.
[0140] Among them, determine whether the memory resources, computing resources, and routing resources of the processor can support the needs of the neural network model algorithm respectively. For example, according to the first intermediate representation and the resource information, simulate the execution process on the processor, so as to perform performance evaluation on the processor from different dimensions such as memory resources, computing efficiency, and routing volume, so as to determine whether the processor meets the computing power support requirements for the neural network model, so as to specifically optimize the processor. The method for determining whether the processor meets the computing power support requirements for the neural network model can be implemented based on the prior art. The method for optimizing the processor can be, for example, optimizing the memory resources, routing design, etc. of the processor. This process can be implemented based on the prior art, and the present application does not limit this.
[0141] According to the embodiments of the present application, it is possible to optimize the processor from the perspective of processor resources, so as to more specifically achieve the collaborative optimization of the processor and the compilation process.
[0142] In a possible implementation, step S206 may include:
[0143] Determine the accuracy of the algorithm of the neural network model when executed on the processor according to the second intermediate representation and the behavior characteristic information; optimize the processor according to the accuracy.
[0144] Among them, the accuracy of the algorithm of the neural network model when executed on the processor may include the latency, power consumption, throughput, and the correctness of the obtained prediction results when the model is executed on the processor. For example, according to the second intermediate representation and the behavior characteristic information, the execution process on the processor can be simulated to determine the accuracy, so that the memory resources and routing design of the processor can be further optimized according to the accuracy, and this process can be implemented based on the prior art.
[0145] According to the embodiments of the present application, it is possible to optimize the processor from the perspective of the execution of the processor, so as to more specifically achieve the collaborative optimization of the processor and the compilation process.
[0146] In a possible implementation, the information of the processor may include the clock information of the processor, and step S206 may include:
[0147] Determine the time performance of the algorithm of the neural network model when executed on the processor according to the clock information and the target instruction; optimize the processor according to the time performance.
[0148] For example, based on the clock information, the actual execution process of the target instruction on the processor can be simulated to determine the time performance of the algorithm of the neural network model when executed on the processor, and this process can be implemented based on the prior art. Among them, the time performance may include the execution time of the algorithm on the processor, the throughput rate, etc., and the throughput rate may represent the number of instructions that can be executed concurrently per unit time when the processor executes the target instruction. Further optimization of the memory resources and routing design of the processor can be carried out based on the determined time performance.
[0149] According to the embodiments of the present application, it is possible to optimize the processor from the perspective of clock accuracy, so as to more specifically achieve the optimization of the processor.
[0150] In a possible implementation, in order to more quickly and comprehensively verify the compilation strategy and processor design, different test cases can also be stored for each level of storage according to the needs of testing the compilation process and the needs of testing the hardware. For example, the test cases can be stored in the storage module 101.
[0151] Among them, the test cases may include cases for verifying the execution accuracy requirements of the algorithm on the processor to test the compilation accuracy during the compilation process; the test cases may also include cases for comprehensively testing the operations executable by the processor and the processor behavior, such as cases for functional coverage testing (indicating the integrity of the test), stress testing (indicating the processor performance), and boundary testing (indicating the error rate of the neural network model in boundary scenarios).
[0152] Figure 3 Shows a structural diagram of a data processing device according to an embodiment of the present application. As Figure 3 shown, the device includes:
[0153] A first acquisition module 301, configured to acquire a computation graph of a neural network model, where the computation graph represents the structure of the neural network model;
[0154] A second acquisition module 302, configured to acquire information about the processor;
[0155] A first determination module 303, configured to determine a first intermediate representation according to the computation graph and the information about the processor, where the first intermediate representation includes operators supported by the processor;
[0156] A second determination module 304, configured to determine a second intermediate representation according to the first intermediate representation and the information about the processor, where the second intermediate representation includes time scheduling information and spatial mapping information of the operators on the processor;
[0157] A third determination module 305, configured to determine a target instruction corresponding to the neural network model according to the second intermediate representation and the information about the processor, where the target instruction is used to be executed by the processor;
[0158] An optimization module 306, configured to optimize the processor according to at least one of the computation graph, the first intermediate representation, the second intermediate representation, the target instruction, and the information about the processor.
[0159] According to an embodiment of the present application, by acquiring the computation graph of the neural network model and the information about the processor, the compilation of the neural network model is realized, so that the new neural network model can be efficiently deployed on the new chip. At the same time, during the compilation process, the processor can also be guided to be optimized to realize the collaborative optimization of the compilation process and the processor, making the compilation process flexible and efficient. At the same time, by determining the first intermediate representation, the second intermediate representation, and finally determining the target instruction during the compilation process, the compilation process can be made hierarchical, and different levels are targeted at different compilation requirements, making the compilation more targeted and obtaining a better compilation effect.
[0160] In a possible implementation, the neural network model includes at least one of an artificial neural network model ANN, a spiking neural network model SNN, and a hybrid neural network model HNN, and the processor includes at least one of a graphics processing unit GPU, a neuromorphic chip, a brain-inspired computing chip, and a many-core neural network acceleration chip.
[0161] Thereby, the compilation of the new neural network model can be realized, so that the new neural network model can be deployed on the new chip to meet the needs of the new neural network model and the new chip.
[0162] In a possible implementation, the information of the processor includes the operator information supported by the processor, and the operator information includes at least one of the operator type and the data precision of the operator. The first determination module 303 is configured to:
[0163] Determine the first intermediate representation according to the computational graph and the operator information supported by the processor.
[0164] According to the embodiments of the present application, by using the operator information supported by the processor, the operators in the computational graph can be converted into the operators supported by the processor to determine the first intermediate representation, so that the operator information supported by the processor can be considered in the process of compilation optimization, hierarchical compilation can be realized, and the compilation process is more in line with the specific hardware characteristics, more targeted, and the effect of compilation optimization is better.
[0165] In a possible implementation, determining the first intermediate representation according to the computational graph and the operator information supported by the processor includes:
[0166] Perform operator type conversion and / or operator quantization on the operators in the computational graph according to the operator information supported by the processor to determine the first intermediate representation.
[0167] According to the embodiments of the present application, by determining the first intermediate representation according to the computational graph and the operator information supported by the processor, the neural network model can be compiled according to the operator type, data precision supported by the processor, and the requirements for the computing precision of the processor, so that the compilation process is more targeted and the compilation effect is better.
[0168] In a possible implementation, the information of the processor includes the resource information of the processor, and the resource information includes at least one of memory resource information, computing resource information, and routing resource information. The second determination module 304 is configured to:
[0169] Determine the second intermediate representation according to the first intermediate representation and the resource information.
[0170] According to the embodiments of the present application, by utilizing the resource information of the processor, it is possible to determine a second intermediate representation with spatio-temporal mapping information. Considering the resource information of the processor during the compilation optimization process can achieve hierarchical compilation, make the compilation process more in line with specific hardware characteristics, more targeted, and have better compilation optimization effects.
[0171] In a possible implementation manner, the information of the processor includes the behavioral characteristic information of the processor. The behavioral characteristic information includes the operations executable by the processor. The third determination module 305 is configured to:
[0172] Determine the target instruction according to the first intermediate representation and the behavioral characteristic information.
[0173] According to the embodiments of the present application, by utilizing the behavioral characteristic information of the processor, it is possible to determine the target instruction. Considering the behavioral characteristic information of the processor during this process can achieve hierarchical compilation, make the compilation process more in line with specific hardware characteristics, more targeted, and have better compilation optimization effects.
[0174] In a possible implementation manner, determining the target instruction according to the second intermediate representation and the behavioral characteristic information includes:
[0175] Perform at least one of memory address allocation, routing path planning, calculation pipeline optimization, and register allocation on the second intermediate representation according to the second intermediate representation and the behavioral characteristic information to determine the target instruction.
[0176] According to the embodiments of the present application, by performing memory address allocation, routing path planning, calculation pipeline optimization, and register allocation, the compilation effect can be better, and the obtained target instruction performs better on the processor.
[0177] In a possible implementation manner, the optimization module 306 is configured to:
[0178] Determine whether the processor meets the operator function support requirements for the neural network model according to the computational graph and the operator information supported by the processor;
[0179] Optimize the processor when the operator function support requirements for the neural network model are not met.
[0180] According to the embodiments of the present application, it is possible to optimize the processor from the perspective of hardware operator design, so as to more specifically achieve the collaborative optimization of the processor and the compilation process.
[0181] In a possible implementation manner, the optimization module 306 is configured to:
[0182] Determine whether the processor meets the computing power support requirement for the neural network model according to the first intermediate representation and the resource information;
[0183] When the computing power support requirement for the neural network model is not met, optimize the processor.
[0184] According to the embodiments of the present application, it is possible to optimize the processor from the perspective of processor resources, so that the collaborative optimization of the processor and the compilation process can be realized more pertinently.
[0185] In a possible implementation manner, the optimization module 306 is configured to:
[0186] Determine the accuracy of the algorithm of the neural network model when executed on the processor according to the second intermediate representation and the behavior feature information;
[0187] Optimize the processor according to the accuracy.
[0188] According to the embodiments of the present application, it is possible to optimize the processor from the perspective of processor execution, so that the collaborative optimization of the processor and the compilation process can be realized more pertinently.
[0189] In a possible implementation manner, the information of the processor includes the clock information of the processor, and the optimization module 306 is configured to:
[0190] Determine the time performance of the algorithm of the neural network model when executed on the processor according to the clock information and the target instruction;
[0191] Optimize the processor according to the time performance.
[0192] According to the embodiments of the present application, it is possible to optimize the processor from the perspective of clock accuracy, so that the optimization of the processor can be realized more pertinently.
[0193] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be elaborated here.
[0194] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0195] An embodiment of the present disclosure also provides an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0196] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above method.
[0197] Figure 4 The block diagram of a data processing device 1900 according to an embodiment of the present application is shown. For example, the device 1900 may be provided as a server or a terminal device. Referring to Figure 4 , the device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to implement the above method.
[0198] The device 1900 may further include a power component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output (I / O) interface 1958. The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows ServerTM, MacOS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
[0199] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions. The above computer program instructions can be executed by the processing component 1922 of the device 1900 to complete the above method.
[0200] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0201] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0202] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0203] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0204] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0205] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0206] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0207] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0208] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A data processing method, characterized in that, the method includes: obtaining a computational graph of a neural network model, where the computational graph represents the structure of the neural network model; obtaining information of a processor, where the information of the processor includes operator information supported by the processor, resource information of the processor, and behavioral characteristic information of the processor; determining a first intermediate representation according to the computational graph and the operator information supported by the processor, where the first intermediate representation includes operators supported by the processor; determining a second intermediate representation according to the first intermediate representation and the resource information, where the second intermediate representation includes time scheduling information and spatial mapping information of the operators on the processor; determining a target instruction corresponding to the neural network model according to the second intermediate representation and the behavioral characteristic information, where the target instruction is used to be executed by the processor; optimizing the processor according to at least one of the computational graph, the first intermediate representation, the second intermediate representation, the target instruction, and the information of the processor.
2. The method according to claim 1, characterized in that, the operator information includes at least one of operator type and data precision of the operator.
3. The method according to claim 1, characterized in that, the determining the first intermediate representation according to the computational graph and the operator information supported by the processor includes: performing operator type conversion and / or operator quantization on the operators in the computational graph according to the operator information supported by the processor to determine the first intermediate representation.
4. The method according to claim 1, characterized in that, the resource information includes at least one of memory resource information, computing resource information, and routing resource information.
5. The method according to claim 1, characterized in that, the behavioral characteristic information includes operations executable by the processor.
6. The method according to claim 1, characterized in that, the determining the target instruction according to the second intermediate representation and the behavioral characteristic information includes: performing at least one of memory address allocation, routing path planning, computing pipeline optimization, and register allocation on the second intermediate representation according to the second intermediate representation and the behavioral characteristic information to determine the target instruction.
7. The method according to claim 1, characterized in that, the optimizing the processor according to at least one of the computational graph, the first intermediate representation, the second intermediate representation, the target instruction, and the information of the processor includes: determining whether the processor meets the operator function support requirement for the neural network model according to the computational graph and the operator information supported by the processor; optimizing the processor when the operator function support requirement for the neural network model is not met.
8. The method according to claim 1, characterized in that, the optimizing the processor according to at least one of the computational graph, the first intermediate representation, the second intermediate representation, the target instruction, and the information of the processor includes: Determine whether the processor meets the computing power support requirement for the neural network model according to the first intermediate representation and the resource information; Optimize the processor in case the computing power support requirement for the neural network model is not met.
9. The method according to claim 1, wherein, the optimizing the processor according to at least one of the computation graph, the first intermediate representation, the second intermediate representation, the target instruction, and the information of the processor includes: Determine the accuracy of the algorithm of the neural network model executed on the processor according to the second intermediate representation and the behavior characteristic information; Optimize the processor according to the accuracy.
10. The method according to any one of claims 1-9, wherein, the information of the processor includes the clock information of the processor, and the optimizing the processor according to at least one of the computation graph, the first intermediate representation, the second intermediate representation, the target instruction, and the information of the processor includes: Determine the time performance of the algorithm of the neural network model executed on the processor according to the clock information and the target instruction; Optimize the processor according to the time performance.
11. The method according to claim 1, wherein, the neural network model includes at least one of an artificial neural network model ANN, a spiking neural network model SNN, and a hybrid neural network model HNN, and the processor includes at least one of a graphics processing unit GPU, a neuromorphic chip, a brain-like computing chip, and a many-core neural network acceleration chip.
12. A data processing device, wherein, the device includes: A first acquisition module, configured to acquire a computation graph of a neural network model, where the computation graph represents the structure of the neural network model; A second acquisition module, configured to acquire information of a processor, where the information of the processor includes operator information supported by the processor, resource information of the processor, and behavior characteristic information of the processor; A first determination module, configured to determine a first intermediate representation according to the computation graph and the operator information supported by the processor, where the first intermediate representation includes the operators supported by the processor; A second determination module, configured to determine a second intermediate representation according to the first intermediate representation and the resource information, where the second intermediate representation includes time scheduling information and spatial mapping information of the operators on the processor; A third determination module, configured to determine a target instruction corresponding to the neural network model according to the second intermediate representation and the behavior characteristic information, where the target instruction is used to be executed by the processor; An optimization module, configured to optimize the processor according to at least one of the computation graph, the first intermediate representation, the second intermediate representation, the target instruction, and the information of the processor.
13. A data processing device, wherein, includes: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method according to any one of claims 1 to 11 when executing the instructions stored in the memory.
14. A non-volatile computer-readable storage medium, on which computer program instructions are stored, characterized in that when the computer program instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Operation method and device, electronic equipment and storage medium
CN112947935A
Neural network processing method and apparatus, computer device and storage medium
WO2021057746A1