A compilation method, device, computer device and storage medium for a neural network

By analyzing the configuration information of the operation statement generated by the intermediate program, adapting to the instruction packaging relationship of the AI acceleration chip, the problem that neural networks in the prior art cannot fully utilize the computing capabilities of the AI acceleration chip are improved, and inference performance is improved.

CN114356340BActive Publication Date: 2025-07-22SHANGHAI POWERTENSORS INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111652713.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-07-22
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

The existing compilation framework cannot effectively compile neural networks onto AI acceleration chips, and cannot fully utilize their complex computing capabilities, resulting in insufficient inference performance.

Method used

By analyzing the internal data structure of the intermediate program, generating configuration information of the operation statement, and using the instruction packaging relationship of the data processing chip, generating machine instructions to adapt to the operation type of the AI acceleration chip.

Benefits of technology

It fully utilizes the computing power of AI acceleration chips and improves the inference performance of neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114356340B_ABST
    Figure CN114356340B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, an apparatus, a computer device and a storage medium for compiling a neural network. The method includes: parsing an intermediate program to obtain an internal data structure of the intermediate program; the internal data structure includes objects and the association relationships between the objects; the objects include: operation statements, and variables and variable definition statements corresponding to the operation statements; wherein, the intermediate program is a program written in a preset intermediate language converted from an original program of a target neural network written in a target high-level language by using conversion relationship information between the preset intermediate language and the target high-level language; generating configuration information of the operation statements based on the internal data structure; generating machine instructions of the data processing chip when running the target neural network based on the instruction encapsulation relationship of the operation statements for the data processing chip and the configuration information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, an apparatus, a computer device, and a storage medium for compiling a neural network. Background Art

[0002] Artificial Intelligence (AI) acceleration chips are an effective way to improve the inference performance of neural networks. They implement common computations in neural networks at the hardware level and enhance the inference process of the network. In order to infer a neural network on an AI acceleration chip, a compiler is required to compile the original neural network into instructions for the AI acceleration card chip.

[0003] However, common compilation frameworks are mainly oriented towards general hardware accelerators, that is, they are more suitable for machine instructions that can be decomposed into simple arithmetic operations, and are not suitable for machine instructions that directly adopt arithmetic operations such as convolution operations. Therefore, they are not suitable for compiling neural networks on data processing chips such as AI acceleration chips. Summary of the Invention

[0004] Embodiments of the present disclosure at least provide a method, an apparatus, a computer device, and a storage medium for compiling a neural network.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for compiling a neural network, including: parsing an intermediate program to obtain an internal data structure of the intermediate program; the internal data structure includes objects and the association relationships between the objects; the objects include: operation statements, and variables and variable definition statements corresponding to the operation statements; wherein, the intermediate program is a program written in a preset intermediate language by converting an original program of a target neural network written in a target high-level language by using conversion relationship information between the preset intermediate language and the target high-level language; generating configuration information for the operation statements based on the internal data structure; generating machine instructions for the data processing chip when running the target neural network based on the instruction encapsulation relationship of the operation statements for the data processing chip and the configuration information.

[0006] In an optional implementation manner, the configuration information of the operation statements includes: the target storage address of the variable corresponding to the operation statement; generating the configuration information for each operation statement based on the internal data structure includes: for a target operation statement in the operation statements, allocating a corresponding target storage address for the variable corresponding to the target operation statement based on the variable definition statement of the variable corresponding to the target operation statement; the target storage address includes a storage address in the memory and / or a storage address in the cache.

[0007] In an alternative embodiment, before generating the configuration information of each of the operation statements based on the internal data structure, the method further includes: determining, from the operation statements, operation statements to be deleted that repeatedly access a preset storage space based on the type of the operation statements and the variables corresponding to the operation statements; and deleting the operation statements to be deleted from the operation statements to obtain the target operation statements.

[0008] In an alternative embodiment, the types of the operation statements include variable storage and variable reading. Determining, from the operation statements, operation statements to be deleted that repeatedly access a preset storage space based on the type of the operation statements and the variables corresponding to the operation statements includes: for each first operation statement of which the type is variable storage, determining whether there is a second operation statement of which the type is variable reading and the variable corresponding to the second operation statement is the same as the variable corresponding to the first operation statement, where the first operation statement and the second operation statement belong to adjacent different network layers; and if so, determining the first operation statement and the corresponding second operation statement as the operation statements to be deleted.

[0009] In an alternative embodiment, the variables include input variables, output variables, and intermediate variables respectively corresponding to each network layer. Allocating corresponding target storage addresses for the variables corresponding to the target operation statements based on the variable definition statements of the variables corresponding to the target operation statements includes: in response to the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an input variable, allocating storage addresses in the memory and the cache for the variable corresponding to the target operation statement; in response to the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an output variable, allocating storage addresses in the memory and the cache for the variable corresponding to the target operation statement; and in response to the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an intermediate variable, allocating a storage address in the cache for the variable corresponding to the target operation statement.

[0010] In an alternative embodiment, before generating the machine instructions of the data processing chip when running the target neural network based on the instruction encapsulation relationship of the operation statements with respect to the data processing chip and the configuration information, the method further includes: for each operation statement, adding wait information for the operation statement according to the association relationship between the operation statement and other operation statements, where the wait information is used to indicate the execution order of each operation instruction.

[0011] In an alternative embodiment, after adding waiting information to each operation statement according to the association relationship between each operation statement and other operation statements, the method further includes: generating first debugging information based on the operation statement and the waiting information added to the operation statement; the first debugging information is used to debug the original program of the target neural network.

[0012] In an alternative embodiment, after generating the configuration information of the operation statement based on the internal data structure, the method further includes: generating second debugging information based on the operation statement and the configuration information corresponding to the operation statement; the second debugging information is used to debug the original program of the target neural network.

[0013] In an alternative embodiment, the compilation method further includes: in response to the encapsulation operation of multiple machine instructions of multiple data processing chips, generating operation statements corresponding to the encapsulation operation; establishing a conversion relationship between the operation statements and the target high-level language.

[0014] In a second aspect, an embodiment of the present disclosure further provides a neural network compilation device, including: a parsing module, configured to parse an intermediate program to obtain an internal data structure of the intermediate program; the internal data structure includes objects and the association relationship between the objects; the objects include: operation statements, and variables and variable definition statements corresponding to the operation statements; wherein, the intermediate program is a program written in a preset intermediate language converted from an original program of a target neural network written in a target high-level language by using conversion relationship information between the preset intermediate language and the target high-level language; a first generation module, configured to generate configuration information of the operation statement based on the internal data structure; a second generation module, configured to generate machine instructions of the data processing chip when running the target neural network based on the instruction encapsulation relationship of the operation statement on the data processing chip and the configuration information.

[0015] In an alternative embodiment, the configuration information of the operation statement includes: the target storage address of the variable corresponding to the operation statement; when generating the configuration information of each operation statement based on the internal data structure, the first generation module is configured to: for a target operation statement in the operation statement, allocate a corresponding target storage address for the variable corresponding to the target operation statement based on the variable definition statement of the variable corresponding to the target operation statement; the target storage address includes a storage address in the memory and / or a storage address in the cache.

[0016] In an alternative embodiment, before generating the configuration information for each of the operation statements based on the internal data structure, the first generation module is further configured to: determine, from the operation statements, the operation statements to be deleted that repeatedly access a preset storage space based on the type of the operation statements and the variables corresponding to the operation statements; and delete the operation statements to be deleted from the operation statements to obtain the target operation statements.

[0017] In an alternative embodiment, the types of the operation statements include: variable storage and variable reading; when determining, from the operation statements, the operation statements to be deleted that repeatedly access a preset storage space based on the type of the operation statements and the variables corresponding to the operation statements, the first generation module is configured to: for each first operation statement of which the type is variable storage, determine whether there is a second operation statement of which the type is variable reading and the variable corresponding to the second operation statement is the same as the variable corresponding to the first operation statement; wherein the first operation statement and the second operation statement belong to different adjacent network layers; if so, determine the first operation statement and the corresponding second operation statement as the operation statements to be deleted.

[0018] In an alternative embodiment, the variables include: input variables, output variables, and intermediate variables respectively corresponding to each network layer; when allocating corresponding target storage addresses for the variables corresponding to the target operation statements based on the variable definition statements of the variables corresponding to the target operation statements, the first generation module is configured to: in response to the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an input variable, allocate storage addresses in the memory and the cache for the variable corresponding to the target operation statement; in response to the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an output variable, allocate storage addresses in the memory and the cache for the variable corresponding to the target operation statement; in response to the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an intermediate variable, allocate a storage address in the cache for the variable corresponding to the target operation statement.

[0019] In an alternative embodiment, before generating the machine instructions for the data processing chip when running the target neural network based on the instruction encapsulation relationship of the operation statements for the data processing chip and the configuration information, the second generation module is further configured to: for each operation statement, add waiting information for the operation statement according to the association relationship between the operation statement and other operation statements; the waiting information is used to indicate the execution order of each operation instruction.

[0020] In an alternative embodiment, after adding waiting information to each operation statement according to the association relationship between each operation statement and other operation statements, the second generation module is further configured to: generate first debugging information based on the operation statement and the added waiting information for the operation statement; the first debugging information is used to debug the original program of the target neural network.

[0021] In an alternative embodiment, after generating the configuration information of the operation statement based on the internal data structure, the first generation module is further configured to: generate second debugging information based on the operation statement and the configuration information corresponding to the operation statement; the second debugging information is used to debug the original program of the target neural network.

[0022] In an alternative embodiment, the compilation device further includes a processing module, configured to: generate an operation statement corresponding to the encapsulation operation in response to the encapsulation operation of multiple machine instructions of the multiple data processing chips; establish a conversion relationship between the operation statement and the target high-level language.

[0023] In a third aspect, an alternative implementation of the present disclosure further provides a computer device, a processor, and a memory. The memory stores machine-readable instructions executable by the processor. The processor is configured to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the machine-readable instructions execute the steps in the first aspect or any possible implementation manner in the first aspect when the machine-readable instructions are executed by the processor.

[0024] In a fourth aspect, an alternative implementation of the present disclosure further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run, it executes the steps in the first aspect or any possible implementation manner in the first aspect.

[0025] The neural network compilation method provided by the embodiments of the present disclosure can further generate the configuration information of the operation statement by using the internal data structure obtained by parsing the intermediate program, and then generate machine instructions by using the instruction encapsulation relationship between the operation statement and the data processing chip and the configuration information. In this way, since the instruction encapsulation relationship for encapsulating the instructions of the data processing chip by the preset intermediate language can be established in advance, the corresponding operation statement can be constructed by using the intermediate language according to the operation types actually supported by the data processing chip. In this way, in the process of converting the original program written in the target high-level language into machine instructions, it is not necessary to refine it into basic machine instructions such as addition, subtraction, and multiplication, but to convert it into machine instructions corresponding to the operation types supported by the data processing chip, so as to make more full use of the computing power of the data processing chip.

[0026] To make the above objectives, features, and advantages of the present disclosure more apparent and understandable, the following provides preferred embodiments in conjunction with the accompanying drawings and describes them in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for use in the embodiments will be briefly introduced below. These accompanying drawings are incorporated into the specification and constitute a part of this specification. These drawings show embodiments that conform to the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following accompanying drawings only show certain embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related accompanying drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 Shows a flowchart of a method for compiling a neural network provided by an embodiment of the present disclosure;

[0029] Figure 2 Shows a schematic diagram of a compiler provided by an embodiment of the present disclosure;

[0030] Figure 3 Shows an example diagram of determining an operation statement to be deleted provided by an embodiment of the present disclosure;

[0031] Figure 4 Shows a flowchart corresponding to a specific embodiment during compilation provided by an embodiment of the present disclosure;

[0032] Figure 5 Shows a schematic diagram of a neural network compilation device provided by an embodiment of the present disclosure;

[0033] Figure 6 Shows a schematic diagram of a computer device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. Usually, the components of the embodiments of the present disclosure described and shown here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the present disclosure to be protected, but only represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.

[0035] Research has found that among data processing chips, AI acceleration chips are an effective way to improve the reasoning performance of neural networks. In order to reason a network on an AI acceleration chip, the compiler needs to compile the original network into instructions for the AI acceleration card chip. According to the hardware design of the AI acceleration chip, dozens of register parameters need to be configured each time a specific calculation is performed, and the synchronization of multiple calculations when executed in parallel also needs to be considered. The compiler's direct generation of machine instructions does not have debuggability and scalability for large networks, so an intermediate language is needed to connect the front-end and back-end of the compiler: the front-end is responsible for parsing the network and implementing operators to generate intermediate language; the back-end focuses on optimizing and generating machine instructions. The intermediate languages corresponding to the currently commonly used compilation frameworks are mainly aimed at general hardware accelerators, which usually use basic machine instructions to perform data processing tasks. Basic machine instructions can be instructions corresponding to basic calculations such as addition, subtraction, and multiplication. As the types of operators in neural networks continue to increase, in order to adapt to the increase in operator types, more and more AI acceleration chips have implemented more complex operations in hardware, such as convolution operations and pooling operations. The current neural network compilation method can only compile neural networks into basic machine instructions, and therefore cannot fully utilize the computing power of AI acceleration chips, and is not suitable for compiling neural networks on AI acceleration chips.

[0036] The defects existing in the above solutions are the results obtained by the inventor after practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the present disclosure for the above problems below should be the contributions made by the inventor to the present disclosure during the disclosure process.

[0037] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0038] For ease of understanding this embodiment, a neural network compilation method disclosed in this disclosure embodiment will be introduced in detail first. The execution subject of the neural network compilation method provided in this disclosure embodiment is generally a computer device with certain computing capabilities. Such a computer device includes, for example: a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the neural network compilation method may be implemented by a processor invoking computer-readable instructions stored in a memory.

[0039] The neural network compilation method provided in this disclosure embodiment will be described below.

[0040] See Figure 1 As shown, it is a flowchart of a neural network compilation method provided in this disclosure embodiment. The method includes steps S101 to S103, where:

[0041] S101: Parse the intermediate program to obtain the internal data structure of the intermediate program; the internal data structure includes objects and the association relationships between the objects; the objects include: operation statements, and variables and variable definition statements corresponding to the operation statements; where the intermediate program is a program written in a preset intermediate language converted from the original program of the target neural network written in a target high-level language by using the conversion relationship information between the preset intermediate language and the target high-level language;

[0042] S102: Generate configuration information for the operation statements based on the internal data structure;

[0043] S103: Generate machine instructions for the data processing chip when running the target neural network based on the instruction encapsulation relationship of the data processing chip for the operation statements and the configuration information.

[0044] After converting the original program of the target neural network into an intermediate program written in a preset intermediate language in the embodiments of the present disclosure, the internal data structure obtained by parsing the intermediate program can be used to further generate the configuration information of the operation statements, and then the operation statements, the instruction encapsulation relationship, and the configuration information of the data processing chip are used to generate machine instructions. In this way, since the instruction encapsulation relationship for encapsulating the instructions of the data processing chip by using the preset intermediate language can be established in advance, the corresponding operation statements can be constructed by using the intermediate language according to the operation types actually supported by the data processing chip. In this way, in the process of converting the original program written in the target high-level language into machine instructions, it is not necessary to refine it into basic machine instructions such as addition, subtraction, and multiplication, but to convert it into machine instructions corresponding to the operation types supported by the data processing chip, so as to make more full use of the computing power of the data processing chip.

[0045] In addition, in the embodiments of the present disclosure, the data processing tasks of the target neural network may include inference and / or training tasks.

[0046] The above S101 to S103 will be described in detail below.

[0047] Regarding the above S101, first, the preset intermediate language, the target high-level language, the target neural network, and the original program of the target neural network included in this step will be described.

[0048] First, regarding the target neural network, the target neural network may be different according to different actual application scenarios. Exemplarily, in the application scenario of image recognition, for example, in the application scenario of object recognition and classification, the corresponding target neural network may include a convolutional neural network; in the application scenario of speech recognition, for example, in the recognition of audio-to-text conversion, the corresponding target neural network may include a recurrent neural network.

[0049] After determining the target neural network, the original program of the target neural network can also be determined accordingly. Exemplarily, before deploying the target neural network determined in different application scenarios on a specific hardware device, by adjusting the parameters of the determined target neural network, etc., the target neural network for deployment on the hardware device can be obtained, which is actually the original program of the target neural network.

[0050] Among them, when writing the original program of the target neural network, for example, a target high-level language can be used to develop the neural network, and after training the neural network, it is converted into an inference network for representation, such as the Convolutional Architecture for Fast Feature Embedding (Caffe) or other inference networks. In a specific implementation, the target high-level language can include computer programming languages, such as the C language, the Python language, etc.; the original program of the target neural network written using the high-level language described above includes, for example, the inference network described above.

[0051] For the preset intermediate language, exemplarily, the intermediate language can include, for example, the Groot Intermediate Language (GIL). For the intermediate language GIL, the compiler can be divided into a front end and a back end, where the front end is used to process the parsing of the target neural network without paying attention to the requirements of the hardware device to be deployed; the back end is used to optimize the operation statements and generate corresponding machine instructions using the operation statements. Exemplarily, see Figure 2 As shown, it is a schematic diagram of a compiler provided by an embodiment of the present disclosure. Among them, the intermediate language GIL divides the compiler into a compiler front end and a compiler back end. In the case of inputting the target neural network into the compiler, actually the target neural network is input into the compiler front end; through the processing of the intermediate language GIL, machine instructions can be output by the compiler back end. In this way, the target neural network can be converted into machine instructions that can be executed by the hardware device through the compiler.

[0052] In addition, the intermediate language GIL also has good debuggability. During debugging, the parsing result of the front end for the target neural network can be directly analyzed and optimized, such as deleting operation statements, etc. For details, please refer to the following description and will not be elaborated here.

[0053] Exemplarily, for the original program of the target neural network written using the target high-level language C, the original program of the target neural network can be converted into an intermediate program written using the preset intermediate language by using the conversion relationship between the target high-level language C and the preset intermediate language GIL. In a possible case, since in this step, GIL actually completes the work of the compiler front end, the obtained intermediate program can be represented as a frontend.gil program, where "frontend" represents the front end and ".gil" represents being written using the intermediate language GIL.

[0054] After determining the intermediate language corresponding to the target neural network, the intermediate program can also be parsed to obtain the internal data structure of the intermediate program. The internal data structure of the intermediate program is used to represent the abstract syntax structure of the program code corresponding to the intermediate program. Among them, the internal data structure of the intermediate program can be represented by an internal abstract syntax tree (Abstract Syntax Tree, AST) for example, or can also be represented by other data structures. Taking the internal abstract syntax tree as an example, the internal abstract syntax tree is an abstract representation of the source code syntax structure, presenting the syntax structure of the programming language in a tree-like form, that is, the syntax structure of the intermediate program frontend.gil program can be presented in a tree-like form. Among them, each node on the tree represents a structure in the source code. If the internal abstract syntax tree is optimized, it can specifically include processing of replacing or deleting nodes on the tree. The specific processing logic of replacement or deletion is related to the syntax structure corresponding to the node. For example, when storing and reading data, optimization can be performed by deleting relevant statements, which can correspond to deleting the nodes corresponding to these statements on the internal abstract syntax tree. Finally, machine instructions are generated based on the optimized abstract syntax tree. The principle of such an optimization method will be described and explained in the following text.

[0055] Specifically, in the case of obtaining the internal data structure, the internal data structure includes objects and information on the association relationships of the objects. Among them, the objects include: operation statements, and variables and variable definition statements corresponding to the operation statements. Specifically, an object is a concept of encapsulation, an entity under the concept of a class; multiple objects can belong to the same class for example. The operation statements can include specific arithmetic operations, or operations of specific variable storage and variable reading types for example. Among them, the variables can include specific data to be stored, and for different data, the data types, etc. are also different, then variable definition statements can be used to determine information such as the corresponding data types.

[0056] For the above S102, in the case of determining the internal data structure, configuration information for the operation statements can also be generated accordingly.

[0057] In a specific implementation, the configuration information of the operation statements includes, for example: the target storage address of the variable corresponding to the target operation statement, and a waiting message added to the operation statement.

[0058] Among them, during the execution of the target operation statement, some variables may be generated or used; before the execution of the target operation statement, it is necessary to pre-allocate a target storage address for the variable corresponding to the target operation statement; during the execution of the target operation statement, if there is a situation where a variable needs to be stored, the variable can be stored in the storage space indicated by the target storage address based on the pre-allocated target storage address.

[0059] A wait message refers to a description of the specific information of other operation statements that need to be executed before executing a certain operation statement. Based on this wait message, the execution order of different operation statements can be defined, so as to ensure that the operation statements can be executed smoothly according to the logic of the program.

[0060] Since there may be redundant statements in the internal data structure, such as operation statements for repeated memory access, before the step of determining the configuration information of the operation statements, for example, the operation statements to be operated that can be deleted can be determined first to obtain the target operation statements, and the configuration information corresponding to the target operation statements can be generated. In this way, unnecessary read and read operations can be reduced, and the obtained configuration information is more concise.

[0061] Specifically, for example, the following method can be used to determine the target operation statements: based on the type of the operation statements and the variables corresponding to the operation statements, determine the operation statements to be deleted for repeated memory access to a preset storage space from the operation statements; delete the operation statements to be deleted from the operation statements to obtain the target operation statements.

[0062] Among them, the types of operation statements include: variable storage and variable reading. In a specific implementation, when determining the operation statements to be deleted for repeated memory access to a preset storage space from the operation statements, for example, for each first operation statement of the type of variable storage, based on the variable corresponding to the first operation statement, it is determined whether there is a second operation statement of the type of variable reading and with the same variable as the variable corresponding to the first operation statement; where the first operation statement and the second operation statement belong to adjacent different network layers; if so, the first operation statement and the corresponding second operation statement are determined as the operation statements to be deleted.

[0063] Exemplarily, see Figure 3As shown, it is an example diagram for determining an operation statement to be deleted provided by an embodiment of the present disclosure. In this example diagram, two network layers are represented at the upper position, and the operation statements corresponding to the two network layers are described at the lower position respectively. In this example, there are two adjacent network layers: network layer L1 and network layer L2. In this process, since the hardware computing unit of the AI acceleration chip can only access the cache and cannot directly access the memory, for network layer L1, its corresponding output data is the input data of network layer L2, and the data transmitted from network layer L1 to network layer L2 is, for example, variable A. In network layer L1, variable A can be read from the cache and calculated, and since the value corresponding to variable A may change, the calculated variable A is denoted as A'. After obtaining A', it can be stored in the memory. In network layer L2, A' is read from the memory and placed in the cache for calculation. That is, for network layer L1 and network layer L2, when specifically calculating variable A, the operation can be performed according to the process indicated by the solid arrow. Among them, for the convenience of explanation, the operation statement corresponding to variable storage is labeled as F1, and the operation statement corresponding to variable reading is labeled as F2.

[0064] For the internal data structure parsed by the intermediate program, if there is enough space in the cache to store data, then data can be directly read from the cache by the next network layer for further calculation of A' without storing it in the memory. That is, according to the above example, for the operation statement of variable storage type, such as F1, if this operation statement F1 is used as the first operation statement, then based on the variable corresponding to the first operation statement, that is, A, it can be determined whether there is a second operation statement of variable reading type with the same variable A, that is, operation statement F2. And for operation statement F1 and operation statement F2, since they belong to two different network layers respectively, the determined first operation statement F1 and second operation statement F2 can both be used as operation statements to be deleted. Figure 3 In the figure, the operation statements F1 and F2 to be deleted that can be deleted are boxed in the form of a dashed box.

[0065] In the above example, according to the logic before the deletion operation: when data is placed in the cache, it will be stored in the memory again; when reading, it will be read from the memory to the cache and then read from the cache for processing. So the step of storing in the memory can actually be omitted. Therefore, the statements to be deleted are the statements related to storing data from the cache to the memory and reading it back from the memory to the cache.

[0066] In the case of determining the operation statements to be deleted, the operation statements to be deleted can be deleted from the operation statements to obtain the target operation statements.

[0067] In a possible case, when determining the target operation statement, a corresponding target storage address can also be allocated for the variable corresponding to the target operation statement. Specifically, for example, the following method can be adopted: for the target operation statement in the operation statements, based on the variable definition statement of the variable corresponding to the target operation statement, allocate a corresponding target storage address for the variable corresponding to the target operation statement; the target storage address includes a storage address in the memory and / or a storage address in the cache. Wherein, the configuration information of the operation statement includes: the target storage address of the variable corresponding to the operation statement.

[0068] In a specific implementation, for the target operation statement described above, using the variable definition statement corresponding to the variable, a suitable target storage address can be determined for different variables. Specifically, for the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an input variable, allocate a storage address in the memory and in the cache for the variable corresponding to the target operation statement. For the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an output variable, allocate a storage address in the memory and in the cache for the variable corresponding to the target operation statement.

[0069] Exemplarily, for the input data and output data of each network layer, due to the large quantity, it is more suitable to use the memory for storage. For the intermediate data generated during the calculation in each network layer, etc., in order to ensure the calculation efficiency, the cache can be used for storage.

[0070] In a possible case, since the input data and / or output data will also participate in the operation of the internal reference data, in this case, for example, a storage address in the cache can also be allocated for the input data and / or output data. Since the statement for allocating the storage address for the input data and / or output data is not a statement that can be deleted and is also Figure 3 a statement included in the storage and reading operations, it will not affect the efficiency due to the address allocation statement. However, storing and reading data in the cache can improve the calculation efficiency.

[0071] That is, for the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an intermediate variable, allocate a storage address in the cache for the variable corresponding to the target operation statement.

[0072] In another embodiment of the present disclosure, for each operation statement, for example, waiting information can also be added to each operation statement according to the association relationship between each operation statement and other operation statements; the waiting information is used to indicate the execution order of each operation instruction.

[0073] Among them, for multiple operation statements, in the prior art, the manual calculation and waiting technology is mainly adopted to determine the execution order of each operation statement. That is, the execution order corresponding to each operation statement is related to multiple operation statements. In this way, if there is an addition or deletion of an operation statement, it is necessary to re-adjust the order of each operation statement, which is very inconvenient for the case of a large number of operation statements.

[0074] In the embodiments of the present disclosure, by using the correlation relationship information between operation statements, the connection between one operation statement and multiple operation statements can be reduced, and there is only a connection with the operation statements having the correlation relationship information with this operation instruction. In this way, when adding or deleting an operation statement, it will not cause multiple operation statements to be adjusted accordingly due to the change of one operation statement. While improving flexibility, it can also effectively avoid the execution order error caused by re-arranging multiple operation statements.

[0075] Exemplarily, when adding waiting information, the waiting information can, for example, indicate the execution order when two adjacent operation statements are executed. For example, for operation instruction P1 and operation instruction P2, if operation instruction P1 and operation instruction P2 are executed sequentially, then waiting information W1 can be added to operation instruction P1, and the waiting information W1 can indicate that after operation instruction P1 is completed, operation instruction P2 is continued to be executed.

[0076] In addition, for each operation statement, after adding waiting information according to the correlation relationship between each operation statement and other operation statements, for example, first debugging information can also be generated based on the operation statement and the waiting information added to the operation statement; the first debugging information is used to debug the original program of the target neural network.

[0077] Here, the principle of debugging the original program is similar to the process of debugging a program through a program debugging software. After adding waiting information to the operation statement, the original program is debugged again. In principle, by adding waiting information, the error in the execution order of multiple operation statements can be avoided, but it is still necessary to verify through debugging to determine whether there will be an error when the original program is executed according to the order after adding the waiting information, that is, whether there will be an incorrect operation during manual addition, resulting in incorrect waiting information.

[0078] Specifically, when adding waiting information, errors may occur during manual addition, resulting in incorrect waiting information being added. For example, when determining the waiting information for each operation statement based on the association relationship between each operation statement and other operation statements, the waiting information corresponding to other operation statements that do not have an association relationship may be incorrectly set for some operation statements. Therefore, using the operation statements and the added waiting information for the operation statements, first debugging information can be generated. The first debugging information may include, for example, the principle of adding waiting information described in words or characters, enabling the staff to debug the original program based on the first debugging information and the execution logic required by the program. The first debugging information can be used to assist in determining whether there is waiting information added incorrectly as described above, so as to debug the original program of the target neural network, thereby reducing the occurrence of errors.

[0079] In this way, the configuration information of the operation statement can be generated according to the internal data structure.

[0080] In another embodiment of the present disclosure, after generating the configuration information of the operation statement, second debugging information can also be generated based on the operation statement and the configuration information corresponding to the operation statement; the second debugging information is used to debug the original program of the target neural network.

[0081] Here, similar to the generation of the first debugging information above, using the obtained second debugging information, the obtained configuration information can also be further checked accordingly to reduce the occurrence of errors. When the second debugging information is used to debug the original program here, it realizes the debugging of the original program by checking the configuration information. The above first debugging information is to determine whether there is waiting information added incorrectly to realize the debugging of the original program.

[0082] Regarding the above S103, in the case of determining the configuration information, using the operation statement, the instruction encapsulation relationship of the artificial intelligence data processing chip, and the configuration information, machine instructions for executing the inference task of the target neural network by using the data processing chip can be generated. Here, the data processing chip specifically includes an AI acceleration chip, and the configuration information has a specific impact on the operation statement. For example, the variable allocation corresponding to the target operation statement assigns the target storage address for sensing, etc. For details, refer to the description in the above S102. And the operation statement can be used to generate machine instructions. For details, refer to the following description. Therefore, the configuration information will also affect the finally generated machine instructions by affecting the operation statement.

[0083] When generating machine instructions using configuration information, if the configuration information includes the target storage address of a variable, the generated machine instructions carry the target storage address, so that when the AI acceleration chip executes the corresponding machine instructions, it can obtain the variable from the target storage address.

[0084] If the configuration information includes a waiting message, the generated machine instructions can carry the corresponding waiting message to indicate the execution order of different machine instructions for the AI acceleration chip.

[0085] In addition, the waiting message can also be used to determine the position of each machine instruction in the instruction stream composed of different machine instructions. The order of each machine instruction in the instruction stream represents the execution order of the machine instructions.

[0086] When generating the configuration information of the operation statement based on the internal data structure, the generation of the target storage address and the waiting message is also carried out on the basis of the internal data structure. After generating the configuration information, the corresponding target storage address and waiting message are added to the corresponding operation statement on the basis of the internal data structure, and the final internal data structure is formed. Then, the machine instructions are generated using the final internal data structure.

[0087] In a specific implementation, in the case of determining the configuration information, the instruction encapsulation relationship between the operation statement and the AI acceleration chip can be used to generate machine instructions for each operation statement in sequence. The machine instructions generated here are also the machine instructions when the AI acceleration chip actually executes the inference task of the target neural network when deployed on the hardware device. Since the machine instructions that the acceleration chip can execute are different when different AI acceleration chips are selected, for different AI acceleration chips, the instruction encapsulation relationship between the operation statement and the artificial intelligence AI acceleration chip is also different.

[0088] Here, the instruction encapsulation relationship refers to the conversion relationship between the operation statement and the chip instructions of the AI acceleration chip. The instruction encapsulation relationship between the operation statement and the AI acceleration chip can be used to convert the operation statement into the instructions of the AI acceleration chip.

[0089] In another embodiment of the present disclosure, for example, it is also possible to generate an operation statement corresponding to the encapsulation operation in response to the encapsulation operation of multiple machine instructions of multiple AI acceleration chips; and establish a conversion relationship between the operation statement and the target high-level language.

[0090] Exemplarily, when compiling a target neural network, different high-level languages can be used, for example. In this way, the target neural network has a corresponding relationship with the high-level language. For an AI acceleration chip, there is also a corresponding relationship between the encapsulation operation of multiple machine instructions and the corresponding operation statements. In addition, when the target neural network is deployed on the AI acceleration chip, there is also a corresponding deployment relationship when the AI acceleration chip is deployed on the target neural network, for example, which data type of the target neural network can be deployed on the AI acceleration chip. In this way, using the corresponding relationships described above, the conversion relationship between the operation statements and the target high-level language can be determined. When determining a new target neural network to be deployed using the target high-level language, the corresponding operation statements can be directly determined using this conversion relationship, which is more efficient.

[0091] In another embodiment of the present disclosure, a specific embodiment for compiling a neural network is also provided. In the example, the data processing chip specifically includes an AI acceleration chip. Refer to Figure 4 As shown, it is a flowchart corresponding to a specific embodiment for compilation provided by an embodiment of the present disclosure; where.

[0092] S401: Determine the target neural network and the AI acceleration chip;

[0093] In this embodiment, the target neural network can include, for example, a convolutional neural network; the target high-level language for writing the convolutional neural network is C language; the AI acceleration chip can be selected according to actual needs and is not limited here.

[0094] S402: Determine the intermediate program corresponding to the target neural network;

[0095] In this embodiment, by transforming the original program written in C language into an intermediate language GIL, the intermediate program corresponding to the target neural network can be obtained, denoted as the CNN_frontend.gil program.

[0096] S403: Parse the intermediate program to obtain the internal data structure of the intermediate program;

[0097] In this embodiment, the internal data structure of the intermediate program is represented by an internal abstract syntax tree.

[0098] S404: Determine the operation statements to be deleted that can be deleted in the internal data structure, and delete the operation statements to be deleted from the operation statements in the internal data structure to obtain the target operation statements;

[0099] In this embodiment, by determining the types of operation statements and corresponding variables in the internal data structure, for example, the operation statements to be deleted can be identified, and the deletion process for the operation statements to be deleted can be performed accordingly. In this way, the obtained target operation statements are more concise compared to the operation statements in the initially obtained internal data structure, and the data volume during subsequent operations can be reduced.

[0100] S405: Allocate corresponding target storage addresses for the variables corresponding to the target operation statements;

[0101] In this embodiment, the target storage addresses that can be allocated for variables include: allocating memory and storage addresses in the cache for input variables; allocating memory and storage addresses in the cache for output variables; and allocating storage addresses in the cache for intermediate variables.

[0102] S406: Add waiting information to each operation statement;

[0103] In this embodiment, by using the association relationship between each operation statement and other operation statements, waiting information can be added to each operation statement.

[0104] In addition, for the above steps S404 - S406, only one possible order is provided in this embodiment. Specifically, it can be deleted and the order of steps can be adjusted according to the actual situation, and no limitation is made here. In addition, in the above steps S404 - S406, corresponding debugging information can also be generated after each step to further confirm whether the operation steps are accurate.

[0105] S407: Use the target storage addresses and waiting messages as the configuration information of the operation statements, and generate machine instructions for the AI acceleration chip to run the target neural network by using the configuration information.

[0106] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order that constitutes any limitation to the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0107] Based on the same inventive concept, an apparatus for compiling a neural network corresponding to the method for compiling a neural network is also provided in the embodiments of the present disclosure. Since the principle of solving problems by the apparatus in the embodiments of the present disclosure is similar to the above method for compiling a neural network in the embodiments of the present disclosure, the implementation of the apparatus can refer to the implementation of the method, and the repeated parts will not be described again.

[0108] Refer to Figure 5As shown in the figure, it is a schematic diagram of a neural network compilation device provided by an embodiment of the present disclosure. The device includes: a parsing module 51, a first generation module 52, and a second generation module 53. Among them,

[0109] The parsing module 51 is configured to parse the intermediate program to obtain the internal data structure of the intermediate program. The internal data structure includes objects and the association relationships between the objects. The objects include: operation statements, and variables and variable definition statements corresponding to the operation statements. Among them, the intermediate program is a program written in a preset intermediate language by converting the original program of the target neural network written in a target high-level language by using the conversion relationship information between the preset intermediate language and the target high-level language.

[0110] The first generation module 52 is configured to generate configuration information of the operation statements based on the internal data structure.

[0111] The second generation module 53 is configured to generate machine instructions for the data processing chip when running the target neural network based on the instruction encapsulation relationship of the artificial intelligence (AI) acceleration chip for the operation statements and the configuration information.

[0112] In an optional implementation manner, the configuration information of the operation statements includes: the target storage address of the variable corresponding to the operation statement. When generating the configuration information of each operation statement based on the internal data structure, the first generation module 52 is configured to: for the target operation statement in the operation statements, based on the variable definition statement of the variable corresponding to the target operation statement, allocate a corresponding target storage address for the variable corresponding to the target operation statement. The target storage address includes a storage address in the memory and / or a storage address in the cache.

[0113] In an optional implementation manner, before generating the configuration information of each operation statement based on the internal data structure, the first generation module 52 is further configured to: determine, from the operation statements, the operation statements to be deleted that repeatedly access the preset storage space based on the type of the operation statements and the variables corresponding to the operation statements; delete the operation statements to be deleted from the operation statements to obtain the target operation statements.

[0114] In an alternative embodiment, the types of the operation statements include variable storage and variable reading. When determining, from the operation statements, the operation statements to be deleted that repeatedly access the preset storage space based on the types of the operation statements and the variables corresponding to the operation statements, the first generation module 52 is configured to: for each first operation statement of which the type is variable storage, determine, based on the variable corresponding to the first operation statement, whether there is a second operation statement of which the variable corresponding to the second operation statement is the same variable as the variable corresponding to the first operation statement and the type is variable reading; wherein the first operation statement and the second operation statement belong to different adjacent network layers; if so, determine the first operation statement and the corresponding second operation statement as the operation statements to be deleted.

[0115] In an alternative embodiment, the variables include input variables, output variables, and intermediate variables respectively corresponding to each network layer. When allocating corresponding target storage addresses for the variables corresponding to the target operation statement based on the variable definition statements of the variables corresponding to the target operation statement, the first generation module 52 is configured to: in response to the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an input variable, allocate storage addresses in the memory and the cache for the variable corresponding to the target operation statement; in response to the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an output variable, allocate storage addresses in the memory and the cache for the variable corresponding to the target operation statement; in response to the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an intermediate variable, allocate storage addresses in the cache for the variable corresponding to the target operation statement.

[0116] In an alternative embodiment, before generating the machine instructions for the data processing chip to run the target neural network based on the instruction encapsulation relationship of the operation statements with respect to the artificial intelligence (AI) acceleration chip and the configuration information, the second generation module 53 is further configured to: for each operation statement, add waiting information for the operation statement according to the association relationship between the operation statement and other operation statements; the waiting information is used to indicate the execution order of each operation instruction.

[0117] In an alternative embodiment, after adding, for each operation statement, waiting information for the operation statement according to the association relationship between the operation statement and other operation statements, the second generation module 53 is further configured to: generate first debugging information based on the operation statements and the waiting information added for the operation statements; the first debugging information is used to debug the original program of the target neural network.

[0118] In an alternative embodiment, after generating the configuration information of the operation statement based on the internal data structure, the first generation module 52 is further configured to: generate second debugging information based on the operation statement and the configuration information corresponding to the operation statement; the second debugging information is used to debug the original program of the target neural network.

[0119] In an alternative embodiment, the compilation device further includes a processing module 54, configured to: generate an operation statement corresponding to the encapsulation operation in response to the encapsulation operation of multiple machine instructions of the multiple AI acceleration chips; establish a conversion relationship between the operation statement and the target high-level language.

[0120] For the description of the processing flow of each module in the device and the interaction flow between modules, reference may be made to the relevant descriptions in the above method embodiments, which will not be elaborated here.

[0121] The embodiments of the present disclosure also provide a computer device, as Figure 6 shown, which is a schematic structural diagram of the computer device provided by the embodiments of the present disclosure, including:

[0122] a processor 10 and a memory 20; the memory 20 stores machine-readable instructions executable by the processor 10, and the processor 10 is configured to execute the machine-readable instructions stored in the memory 20. When the machine-readable instructions are executed by the processor 10, the processor 10 performs the following steps:

[0123] Parse the intermediate program to obtain the internal data structure of the intermediate program; the internal data structure includes objects and the association relationships between the objects; the objects include: operation statements, and variables and variable definition statements corresponding to the operation statements; wherein, the intermediate program is a program written in a preset intermediate language by converting the original program of the target neural network written in the target high-level language by using the conversion relationship information between the preset intermediate language and the target high-level language; generate the configuration information of the operation statement based on the internal data structure; generate the machine instructions of the data processing chip when running the target neural network based on the instruction encapsulation relationship of the operation statement for the data processing chip and the configuration information.

[0124] The above-mentioned memory 20 includes an internal memory 210 and an external memory 220; here, the internal memory 210 is also called the main memory, which is used to temporarily store the operation data in the processor 10 and the data exchanged with the external memory 220 such as the hard disk. The processor 10 exchanges data with the external memory 220 through the internal memory 210.

[0125] The specific execution process of the above instructions can refer to the steps of the neural network compilation method described in the embodiments of the present disclosure, which will not be elaborated here.

[0126] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the neural network compilation method described in the above method embodiments. Wherein, the storage medium may be a volatile or non-volatile computer-readable storage medium.

[0127] The embodiments of the present disclosure also provide a computer program product, which carries program codes. The instructions included in the program codes can be used to execute the steps of the neural network compilation method described in the above method embodiments. For details, please refer to the above method embodiments and will not be elaborated here.

[0128] Wherein, the above computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0129] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here. In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces. The indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.

[0130] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0131] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit exists physically alone, or two or more units are integrated into one unit.

[0132] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0133] Finally, it should be noted that: the above-mentioned embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A compilation method for a neural network, characterized in that, Including: Parsing an intermediate program to obtain an internal data structure of the intermediate program; The internal data structure includes objects and the association relationships between the objects; The objects include: operation statements, and variables and variable definition statements corresponding to the operation statements; Among them, the intermediate program is a program written in a preset intermediate language obtained by converting an original program of a target neural network written in a target high-level language by using conversion relationship information between the preset intermediate language and the target high-level language; Generating configuration information of the operation statement based on the internal data structure; the configuration information of the operation statement includes: the target storage address of the variable corresponding to the operation statement; Generating machine instructions of the data processing chip when running the target neural network based on the instruction encapsulation relationship of the data processing chip for the operation statement and the configuration information; The generating the configuration information of the operation statement based on the internal data structure includes: For a target operation statement in the operation statements, based on the variable definition statement of the variable corresponding to the target operation statement, allocating a corresponding target storage address for the variable corresponding to the target operation statement; the target storage address includes a storage address in memory and / or a storage address in a cache; Before generating the configuration information of the operation statement based on the internal data structure, it further includes: Determining, from the operation statements, a to-be-deleted operation statement that repeatedly accesses a preset storage space based on the type of the operation statement and the variable corresponding to the operation statement; the types of the operation statements include: variable storage and variable reading; Deleting the to-be-deleted operation statement from the operation statements to obtain the target operation statement; The determining, from the operation statements, a to-be-deleted operation statement that repeatedly accesses a preset storage space based on the type of the operation statement and the variable corresponding to the operation statement includes: For each first operation statement of which the type is variable storage, based on the variable corresponding to the first operation statement, determining whether there is a second operation statement of which the type is variable reading and the variable corresponding to the second operation statement is the same as the variable corresponding to the first operation statement; wherein, the first operation statement and the second operation statement belong to adjacent different network layers; If so, determining the first operation statement and the corresponding second operation statement as the to-be-deleted operation statement.

2. The compilation method according to claim 1, wherein The variables include: input variables, output variables, and intermediate variables respectively corresponding to each network layer; The allocating a corresponding target storage address for the variable corresponding to the target operation statement based on the variable definition statement of the variable corresponding to the target operation statement includes: In response to the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an input variable, allocating a storage address in memory and a storage address in a cache for the variable corresponding to the target operation statement; In response to the variable definition statement of the variable corresponding to the target operation statement indicating that the variable corresponding to the target operation statement is an output variable, allocating a storage address in memory and a storage address in a cache for the variable corresponding to the target operation statement; In response to the variable definition statement corresponding to the variable of the target operation statement indicating that the variable corresponding to the target operation statement is an intermediate variable, allocate a storage address in the cache for the variable corresponding to the target operation statement.

3. The compilation method according to claim 1 or 2, characterized in that, Before generating the machine instructions for the data processing chip to run the target neural network based on the instruction encapsulation relationship of the operation statements for the data processing chip and the configuration information, it further includes: For each operation statement, add wait information for the each operation statement according to the association relationship between the each operation statement and other operation statements; the wait information is used to indicate the execution order of the each operation statement.

4. The compilation method according to claim 3, wherein After adding the wait information for each operation statement according to the association relationship between each operation statement and other operation statements, it further includes: Generate first debugging information based on the operation statements and the wait information added for the operation statements. The first debugging information is used to debug the original program of the target neural network.

5. The compilation method according to claim 1 or 2, characterized in that After generating the configuration information of the operation statements based on the internal data structure, it further includes: Generate second debugging information based on the operation statements and the configuration information corresponding to the operation statements. The second debugging information is used to debug the original program of the target neural network.

6. The compilation method according to claim 1 or 2, characterized in that It further includes: In response to the encapsulation operation of multiple machine instructions for multiple data processing chips, generate operation statements corresponding to the encapsulation operation. Establish the conversion relationship between the operation statements and the target high-level language.

7. A compilation device for a neural network, characterized in that, It includes: A parsing module, configured to parse the intermediate program to obtain the internal data structure of the intermediate program. The internal data structure includes objects and the association relationships between the objects; the objects include: operation statements, and variables and variable definition statements corresponding to the operation statements; wherein, the intermediate program is a program written in a preset intermediate language converted from the original program of the target neural network written in the target high-level language by using the conversion relationship information between the preset intermediate language and the target high-level language. A first generation module, configured to generate the configuration information of the operation statements based on the internal data structure; the configuration information of the operation statements includes: the target storage address of the variable corresponding to the operation statement. A second generation module, configured to generate the machine instructions for the data processing chip to run the target neural network based on the instruction encapsulation relationship of the operation statements for the data processing chip and the configuration information. When generating the configuration information of the operation statements based on the internal data structure, the first generation module is configured to: For the target operation statement in the operation statements, based on the variable definition statement corresponding to the variable of the target operation statement, allocate the corresponding target storage address for the variable corresponding to the target operation statement; the target storage address includes the storage address in the memory and / or the storage address in the cache. Before generating the configuration information of the operation statements based on the internal data structure, the first generation module is further configured to: Determine, from the operation statements, the operation statements to be deleted that repeatedly access a preset storage space based on the type of the operation statements and the variables corresponding to the operation statements; the types of the operation statements include variable storage and variable reading; delete the operation statements to be deleted from the operation statements to obtain the target operation statements; When the first generation module determines, from the operation statements, the operation statements to be deleted that repeatedly access a preset storage space based on the type of the operation statements and the variables corresponding to the operation statements, it is configured to: For each first operation statement of the variable storage type, determine whether there is a second operation statement of the variable reading type whose corresponding variable is the same as that of the first operation statement, where the first operation statement and the second operation statement belong to different adjacent network layers; if so, determine the first operation statement and the corresponding second operation statement as the operation statements to be deleted.

8. A computer device, characterized in that, Including: A processor and a memory, the memory stores machine-readable instructions executable by the processor, the processor is configured to execute the machine-readable instructions stored in the memory, and when the machine-readable instructions are executed by the processor, the processor executes the steps of the neural network compilation method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by a computer device, the computer device executes the steps of the neural network compilation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Neural network model compiling method and device, equipment and storage medium

    CN111860816A