Deep learning operator optimization and operator fusion calculation instruction generation method and device

By fusion and hierarchical processing of the target node set in the deep learning computing graph, fusion operators and calculation instructions are generated, the problem of low optimization efficiency of computing graphs in the prior art is solved, and higher operator execution efficiency is achieved.

CN119990220APending Publication Date: 2025-05-13太初(无锡)电子科技有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411955113.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, the calculation graph optimization is performed through the developers' past experience and fixed optimization templates, and the optimization efficiency is low and the optimization effect is poor.

Method used

By obtaining the deep learning calculation graph, traverse each node to obtain the target node set, fuse the operators of the target node according to the association relationship, generate fusion operators, and generate calculation instructions through hierarchical processing and preset register allocation.

Benefits of technology

The overhead of vacuoles when a single operator is sent and the interoperators access between multiple caches is eliminated. The calculation tasks of multiple target operators are realized through one instruction, which improves the execution efficiency of the operator.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990220A_ABST
    Figure CN119990220A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electronics, and discloses a deep learning operator optimization and fusion operator calculation instruction generation method and device.The deep learning operator optimization method comprises the steps that all nodes in a deep learning calculation graph are traversed to obtain at least one target node set, the target node set comprises multiple target nodes, and the target nodes in the target node set are extracted; the operator corresponding to each target node is a target operator of a target type, and different target operators have an association relationship; and fusing the operators corresponding to different target nodes in the target node set to obtain a fused operator. The target nodes which have the incidence relation in the computational graph and the corresponding operators are the target type operators are used as a set, the target operators corresponding to the target nodes in the same set are fused, the fused operators can be fully optimized in the aspect of specific instruction implementation, and the execution efficiency of the operators is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic technology, and in particular to a method and device for generating computing instructions for deep learning operator optimization and fusion operators. Background Art

[0002] A computational graph is a description of a model. It abstracts the computational process into multiple computational nodes. Each node includes a description of the input, output, and computational process. The same computational process in the original model will be abstracted into the same computational node, i.e., the operator node. The computational tasks of each operator node in the computational graph are executed sequentially to achieve model construction and optimization.

[0003] However, there are some redundant nodes in the calculation graph. The existence of these redundant nodes will slow down the entire calculation process, resulting in low calculation efficiency of the calculation graph. Therefore, it is very important to optimize the operator nodes in the calculation graph. In related technologies, the calculation graph is generally optimized based on the past experience of developers and fixed optimization templates, which has low optimization efficiency and poor optimization effect. Summary of the invention

[0004] In view of this, the present invention provides a method and device for generating computing instructions for deep learning operator optimization and fusion operators, so as to solve the problem of low optimization efficiency and poor optimization effect when optimizing calculation graphs based on developers' past experience and fixed optimization templates.

[0005] In a first aspect, the present invention provides a deep learning operator optimization method, the method comprising: obtaining a deep learning calculation graph, the deep learning calculation graph comprising nodes corresponding to different operators respectively; traversing each node in the deep learning calculation graph to obtain at least one target node set, the target node set comprising multiple target nodes, the operator corresponding to each target node being a target operator of a target type, and there being an association relationship between different target operators; fusing the operators corresponding to different target nodes in the target node set to obtain a fused operator.

[0006] The deep learning operator optimization method provided by the present invention traverses each node in the deep learning computation graph to obtain at least one set of target nodes. The set of target nodes includes multiple target nodes, and the operator corresponding to each target node is a target operator of a target type. There is an association relationship between different target operators; the operators corresponding to different target nodes in the set of target nodes are fused to obtain a fused operator. The method provided by the present invention takes the target nodes in the computation graph that have an association relationship and whose corresponding operators are target type operators as a set, and fuses the target operators corresponding to the target nodes in the same set, eliminating the bubbles when sending a single operator and the overhead of accessing and storing between multiple levels of caches for operators. The fused operator can execute the computing tasks corresponding to multiple target operators through one instruction, without the need to send multiple instructions. The fused operator can be fully optimized in the specific instruction implementation, improving the execution efficiency of the operator.

[0007] In an alternative embodiment, the method further includes: determining a fused node according to the fused operator; using the fused node to update the deep learning computation graph to obtain an updated computation graph.

[0008] In a second aspect, the present invention provides a method for generating a computation instruction for a fused operator. The method includes: obtaining the fused operator and the identifiers of multiple preset registers. The fused operator is obtained through the deep learning operator optimization method of the first aspect or its corresponding embodiment, and the fused operator includes multiple target operators; hierarchically processing the target nodes corresponding to different target operators based on the association relationship between different target operators in the fused operator to determine N sets of target nodes corresponding to different levels respectively. The depths of different target nodes within the same level are the same; let level n = 1, allocate preset registers to each target node in the nth level through the identifiers of multiple preset registers to obtain the preset registers corresponding to each target node in the nth level. The input of each target node in the first level is an external input; let level n = n + 1, identify at least one target preset register that will be released after the computation of different target nodes in the nth level is completed, and allocate registers to different target nodes in the (n + 1)th level based on the identifier of the target preset register and the identifiers of multiple preset registers to obtain the preset registers corresponding to each target node in the (n + 1)th level; if n < N, return to the step of letting level n = n + 1 until n = N to determine the preset registers of each target node; determine the computation instructions for the corresponding target operators based on the preset registers of each target node and the association relationship between different target operators. The computation instructions are used to control the corresponding preset registers to execute the computing tasks of the target operators.

[0009] The computing instruction generation method of the fusion operator provided by the present invention performs hierarchical processing on each target operator in the fusion operator, determines a set of target nodes corresponding to N different levels, and then matches the preset registers corresponding to the target operators level by level. When matching the preset registers, the memory resources released by the operators corresponding to the target nodes in the upper level after the calculation is completed are taken into consideration, and the released preset registers are also added to the matching until the register resources are fully utilized and reasonably bound to the operation of each operator. Finally, the preset registers of each target node and the association relationship between different target operators determine the computing instructions of the corresponding target operators, and the obtained computing instructions can achieve the maximum instruction pipeline efficiency when the number of preset registers is limited, thereby avoiding the performance loss of registers.

[0010] In an optional embodiment, multiple target operators include at least one first operator, the first operator is used to characterize an operator that needs to use external input to perform calculations, and the calculation instructions of the corresponding target operator are determined based on the preset registers of each target node and the association relationship between different target operators. The calculation instructions are used to control the corresponding preset registers to perform the calculation tasks of the target operators. The steps include: determining the external input data information required by the first operator and the preset register of the first operator; generating a first instruction, a second instruction and a third instruction of the second operator, the first instruction is used to control the preset register corresponding to the first operator to read external input data from the target storage location based on the external input data information, and store the read external input data in the cache, the second instruction is used to control the preset register corresponding to the first operator to perform the calculation task of the first operator based on the external input data to obtain the calculation result of the first operator, and the third instruction is used to send the calculation result of the first operator to the preset location for storage; the first instruction, the second instruction and the third instruction of the first operator are packaged to obtain the calculation instruction of the first operator.

[0011] The method provided by this optional implementation mode is to read all the required external input data of the first operator through the first control instruction for the first operator that needs to use external input to perform calculations, and then the second instruction controls the register to perform all the calculation tasks of the first operator based on the read external input data and obtain the calculation results. The third instruction outputs the calculation results of each calculation task. Each calculation task needs to go through the reading of input data, task calculation and result output, and so on to complete multiple calculation tasks in the first operator. Through unified input data reading, task calculation and calculation result output, instruction pipeline masking is maximized, data hazards are eliminated, redundant access between the preset register and the target storage location of the operator is eliminated, and redundant memory access hardware characteristics are eliminated, so that the register is maximized.

[0012] In an optional embodiment, multiple target operators include at least one second operator, and the second operator does not have a lower-level target operator. The calculation instructions of the corresponding target operator are determined based on the preset registers of each target node and the association relationship between different target operators. The calculation instructions are used to control the corresponding preset registers to perform the calculation tasks of the target operator. The steps include: generating a first instruction, a second instruction, and a third instruction of the second operator, the first instruction is used to control the preset register corresponding to the second operator to read the target input data, the second instruction is used to control the preset register corresponding to the second operator to perform the calculation task of the second operator based on the target input data to obtain the output result of the second operator, and the third instruction is used to write the output result of the second operator to the target storage location; the first instruction, the second instruction, and the third instruction of the second operator are packaged to obtain the calculation instructions of the second operator.

[0013] In a third aspect, the present invention provides a deep learning operator optimization device, which includes: a first acquisition module, used to acquire a deep learning calculation graph, the deep learning calculation graph includes node traversal modules corresponding to different operators, used to traverse each node in the deep learning calculation graph to obtain at least one target node set, the target node set includes multiple target nodes, the operator corresponding to each target node is a target operator of a target type, and there is an association relationship between different target operators; a fusion module, used to fuse the operators corresponding to different target nodes in the target node set to obtain a fused operator.

[0014] Fourth aspect, the present invention provides a computing instruction generation device for a fusion operator, the device comprising: a second acquisition module, configured to acquire a fusion operator and identifiers of a plurality of preset registers, the fusion operator being obtained by the deep learning operator optimization method of the first aspect or its corresponding implementation, and the fusion operator including a plurality of target operators; a processing module, configured to perform hierarchical processing on target nodes corresponding to different target operators respectively based on the association relationship between different target operators in the fusion operator, to determine target node sets corresponding to N different levels, and the depths of different target nodes within the same level are the same; a first allocation module, configured to set level n = 1, and allocate preset registers to each target node in the nth level through the identifiers of the plurality of preset registers, to obtain the preset registers corresponding to each target node in the nth level, and the input of each target node in the first level is an external input; a second allocation module, configured to set level n = n + 1, identify at least one target preset register that will be released after the calculation of different target nodes in the nth level is completed, and allocate registers to different target nodes in the (n + 1)th level based on the identifiers of the target preset registers and the identifiers of the plurality of preset registers, to obtain the preset registers corresponding to each target node in the (n + 1)th level; a first determination module, configured to, if n < N, return to the step of setting level n = n + 1 until n = N, to determine the preset registers of each target node; a second determination module, configured to determine the calculation instructions for the corresponding target operators based on the preset registers of each target node and the association relationship between different target operators, and the calculation instructions are used to control the corresponding preset registers to execute the calculation tasks of the target operators.

[0015] Fifth aspect, the present invention provides a computer device, comprising: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the deep learning operator optimization method of the first aspect or its corresponding implementation, or execute the computing instruction generation method for the fusion operator of the second aspect or any of its corresponding implementations.

[0016] Sixth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the deep learning operator optimization method of the first aspect or its corresponding implementation, or execute the computing instruction generation method for the fusion operator of the second aspect or any of its corresponding implementations.

[0017] Seventh aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the deep learning operator optimization method of the first aspect or its corresponding implementation, or execute the computing instruction generation method for the fusion operator of the second aspect or any of its corresponding implementations. Description of the Drawings

[0018] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 is a flowchart of a deep learning operator optimization method according to an embodiment of the present invention;

[0020] Figure 2 is a flow chart of a method for generating a computing instruction of a fusion operator according to an embodiment of the present invention;

[0021] Figure 3 is a flow chart of a method for generating computing instructions for another fusion operator according to an embodiment of the present invention;

[0022] Figure 4 is a flowchart of a method for generating a computing instruction of another fusion operator according to an embodiment of the present invention;

[0023] Figure 5 is a structural block diagram of a deep learning operator optimization device according to an embodiment of the present invention;

[0024] Figure 6 is a structural block diagram of a computing instruction generating device for a fusion operator according to an embodiment of the present invention;

[0025] Figure 7 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0027] There are some redundant nodes in the calculation graph. The existence of these redundant nodes will slow down the entire calculation process, resulting in low calculation efficiency of the calculation graph. Therefore, it is very important to optimize the operator nodes in the calculation graph. In related technologies, the calculation graph is generally optimized based on the past experience of developers and fixed optimization templates, which has low optimization efficiency and poor optimization effect.

[0028] In view of this, a deep learning operator optimization method provided in an embodiment of the present application can be applied to a server to achieve the optimization of deep learning operators. The method provided by the present invention traverses each node in a deep learning calculation graph to obtain at least one target node set, the target node set includes multiple target nodes, and the operator corresponding to each target node is a target operator of a target type, and there is an association relationship between different target operators; the operators corresponding to different target nodes in the target node set are merged to obtain a fused operator. The method provided by the present invention treats the target nodes in the calculation graph that have an association relationship and the corresponding operators are target type operators as a set, and merges the target operators corresponding to the target nodes in the same set, eliminating the bubbles when sending a single operator and the overhead of accessing between operators in multi-level caches. The fused operator can realize the execution of the corresponding computing tasks of multiple target operators through one instruction, without sending multiple instructions. The fused operator can be fully optimized in the specific instruction implementation, which improves the execution efficiency of the operator.

[0029] According to an embodiment of the present invention, an embodiment of a deep learning operator optimization method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in an order different from that shown here.

[0030] In this embodiment, a deep learning operator optimization method is provided, which can be used for the above-mentioned server. Figure 1 is a flowchart of a deep learning operator optimization method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0031] Step S101, obtaining a deep learning calculation graph, where the deep learning calculation graph includes nodes corresponding to different operators.

[0032] Exemplarily, a deep learning calculation graph is a graph structure for describing and organizing the operation of a neural network model. The calculation graph consists of nodes and edges, where nodes represent operations (such as addition, multiplication, activation functions, etc.), and edges represent data flows (i.e., input and output). Through the calculation graph, the dependencies and calculation processes of various operations in the model can be clearly understood, thereby achieving effective training and reasoning. The deep learning algorithm consists of individual computing units, which are operators (OP). In the network model, the operator corresponds to the calculation logic in the layer. In an embodiment of the present application, a plurality of nodes are included in the deep learning calculation graph, and there is a dependency relationship between nodes with a connection relationship, and nodes and operators correspond one to one.

[0033] Step S102, traverse each node in the deep learning calculation graph to obtain at least one target node set, the target node set includes multiple target nodes, the operator corresponding to each target node is a target operator of the target type, and there is an association relationship between different target operators.

[0034] Exemplarily, the target operator of the target type may include, but is not limited to, point-by-point calculation operators and reduction operators. Different nodes in the same target node set correspond to the same operator type and have a dependency relationship. In the embodiment of the present application, based on the greedy algorithm with the maximum number of registers and the maximum on-chip cache space as constraints, as many nodes of point-by-point calculation operators as possible are included in the target node set.

[0035] Step S103, fusing the operators corresponding to different target nodes in the target node set to obtain a fused operator.

[0036] Exemplarily, different target nodes in the same target node set are fused into one operator for calculation. The fusion method is not limited in the embodiment of the present application, as long as it is reasonable. In the embodiment of the present application, it is assumed that operator 1 and operator 2 are two operators with a dependency relationship, and the calculation of operator 2 depends on operator 1. Operator 1 and operator 2 are both point-by-point calculation operators. The host side sends a calculation instruction of operator 1 to the device side, and the device side calculates operator 1 based on the received instruction, and then feeds back the calculation result of the operator to the host side. The host side receives the calculation result of operator 1, and then sends a calculation instruction of operator 2 to the device side. The device side reads the calculation result of operator 1 from the host side based on the received instruction, and calculates operator 2 based on the calculation result of operator 1 and feeds back the calculation result of operator 2 to the host side. Based on this calculation method, it can be found that there are several redundant steps between the calculation of operator 1 and operator 2, namely, "feeding back the calculation result of the operator to the host side", "the host side receives the calculation result of operator 1, and then sends the calculation instruction of operator 2 to the device side", and "the device side reads the calculation result of operator 1 from the host side based on the received instruction". If the host side directly sends instructions to the device side to calculate operator 1 and operator 2, the device side directly calculates operator 1, and after the calculation is completed, the calculation of operator 2 is performed based on the calculation result of operator 1. The calculation result of operator 1 does not need to be written back to the host side, but only needs to reside in the cache of the device side and wait for operator 2 to use it. On this basis, the implementation of operator 1 and operator 2 can be optimized in a targeted manner. By merging operator 1 and operator 2 for calculation, the cavitation when the host side sends a single operator to the device side and the overhead of accessing operators between multi-level caches are eliminated. In addition, after replacing with one operator, the specific instruction implementation can be fully optimized.

[0037] The deep learning operator optimization method provided in this embodiment traverses each node in the deep learning calculation graph to obtain at least one target node set, the target node set includes multiple target nodes, the operator corresponding to each target node is a target operator of the target type, and there is an association relationship between different target operators; the operators corresponding to different target nodes in the target node set are merged to obtain a fused operator. The method provided by the present invention takes the target nodes in the calculation graph that have an association relationship and the corresponding operators are target type operators as a set, and merges the target operators corresponding to the target nodes in the same set, eliminating the bubbles when sending a single operator and the overhead of access between operators in multi-level caches. The fused operator can realize the execution of the corresponding computing tasks of multiple target operators through one instruction, without sending multiple instructions. The fused operator can be fully optimized in the specific instruction implementation, which improves the execution efficiency of the operator.

[0038] In some optional implementations, the method further includes: determining a fusion node according to a fusion operator; and updating the deep learning computation graph using the fusion node to obtain an updated computation graph.

[0039] Exemplarily, in an embodiment of the present application, corresponding fusion nodes are generated based on the nodes of the fusion operator corresponding to different target operators, and the fusion nodes are used to replace the nodes of different target operators in the original deep learning calculation graph to obtain an updated calculation graph.

[0040] In this embodiment, a method for generating a calculation instruction of a fusion operator is provided, which can be used for the above-mentioned server. Figure 2 is a flow chart of a method for generating a calculation instruction of a fusion operator according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0041] Step S201, obtaining identifiers of a fusion operator and multiple preset registers.

[0042] Exemplarily, the fusion operator uses the deep learning operator optimization method of the above embodiment, and the fusion operator includes multiple target operators. The preset registers may include but are not limited to single instruction multiple data (SIMD) registers. The specific number of SIMD registers is not limited in the embodiment of the present application, and those skilled in the art can determine it according to needs.

[0043] Step S202, based on the association relationship between different target operators in the fusion operator, the target nodes corresponding to different target operators are hierarchically processed to determine the target node sets corresponding to N different levels, and the depths of different target nodes in the same level are the same.

[0044] Exemplarily, the association relationships between different target operators can determine the depths of each target node, and the predecessor nodes and successor nodes of each target node are the nodes that have a dependency relationship with the corresponding target node.

[0045] Step S203: Let the level n = 1, and allocate preset registers to each target node in the nth level through the identifiers of multiple preset registers, to obtain the preset registers corresponding to each target node in the nth level. The input of each target node in the first level is an external input.

[0046] Exemplarily, for the first level, the input of each target node comes from external data, and its calculation process does not depend on other target nodes. Therefore, first allocate registers to the target nodes in the first level to determine the SIMD registers of each target node in the first level.

[0047] Step S204: Let the level n = n + 1, identify at least one target preset register that will be released after the calculation of different target nodes in the nth level, and based on the identifier of the target preset register and the identifiers of multiple preset registers, allocate registers to different target nodes in the (n + 1)th level, to obtain the corresponding preset registers of each target node in the (n + 1)th level.

[0048] Exemplarily, in the embodiments of the present application, the calculations of each target node in the (n + 1)th level depend on the output results of the corresponding nodes in the previous level. Identify that there are idle register resources after the previous level's calculation is completed, and include the idle register resources in the matching to reuse the register space. For example, a fusion operator A includes three target operators op1, op2, and op3. The input of op1 is x, the input of op2 is y, op3 depends on the results of op1 and op2, and outputs z after calculation. If the registers corresponding to op1 and op2 are not used again after the data input and calculation are completed, then the calculation and output of op3 can reuse the register space of op1, replace it with in-place calculation, and the register resources corresponding to op2 will be released after the calculation is completed.

[0049] Step S205: If n < N, return to the step of letting the level n = n + 1 until n = N, to determine the preset registers of each target node.

[0050] Exemplarily, determine the preset registers of each target node through layer-by-layer matching.

[0051] Step S206: Based on the preset registers of each target node and the association relationships between different target operators, determine the calculation instructions for the corresponding target operators. The calculation instructions are used to control the corresponding preset registers to execute the calculation tasks of the target operators.

[0052] Exemplarily, the computing instructions of the target operator may include but are not limited to SIMD instructions, where SIMD refers to single instruction multiple data streams, that is, one instruction controls multiple parallel processing elements and performs calculations on a set of independent data in parallel. If the input data is 64 in length, then compared to the original implementation of looping 64 times and calculating one each time, it can be optimized by looping 4 times, using one SIMD instruction each time to calculate 16 numbers in parallel. Subsequently, by executing the computing instructions of each target operator, the calculation of the fused operator can be realized.

[0053] The calculation instruction generation method of the fusion operator provided in this embodiment is to perform hierarchical processing on each target operator in the fusion operator, determine the target node sets corresponding to N different levels, and then match the preset registers corresponding to the target operators step by step. When matching the preset registers, the memory resources released by the operators corresponding to the target nodes in the upper level after the calculation is completed are considered, and the released preset registers are also added to the matching until the register resources are fully utilized and reasonably bound to the operation of each operator. Finally, the preset registers of each target node and the association relationship between different target operators determine the calculation instructions of the corresponding target operator, and the calculation instructions can be realized in the case of a limited number of preset registers. The efficiency of the instruction pipeline is maximized and the performance loss of the register is avoided. At the same time, compared with manually writing the calculation instructions of each target operator in the fusion operator, the method provided by this application effectively improves the generation efficiency of the calculation instructions.

[0054] In this embodiment, a method for generating a calculation instruction of a fusion operator is provided, which can be used for the above-mentioned server. Figure 3 is a flow chart of a method for generating a calculation instruction of a fusion operator according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:

[0055] Step S301, obtaining the identifiers of a fusion operator and multiple preset registers, the fusion operator uses the deep learning operator optimization method of the above embodiment, and the fusion operator includes multiple target operators. For details, please refer to Figure 2 Step S201 of the illustrated embodiment will not be described in detail here.

[0056] Step S302: hierarchically process the target nodes corresponding to different target operators based on the association relationship between different target operators in the fusion operator, and determine the target node sets corresponding to N different levels, and the depths of different target nodes in the same level are the same. Figure 2 Step S202 of the illustrated embodiment will not be described in detail here.

[0057] Step S303: Set level n = 1. Allocate preset registers to each target node in the nth level through the identifiers of multiple preset registers, obtaining the corresponding preset registers for each target node in the nth level. The input of each target node in the first level is an external input. For details, please refer to Figure 2 Step S203 of the embodiment shown, which will not be elaborated here.

[0058] Step S304: Set level n = n + 1. Identify at least one target preset register that different target nodes in the nth level will release after calculation. Based on the identifiers of the target preset registers and the identifiers of multiple preset registers, allocate registers to different target nodes in the (n + 1)th level, obtaining the corresponding preset registers for each target node in the (n + 1)th level. For details, please refer to Figure 2 Step S204 of the embodiment shown, which will not be elaborated here.

[0059] Step S305: If n < N, return to the step of setting level n = n + 1 until n = N, and determine the preset registers of each target node. For details, please refer to Figure 2 Step S205 of the embodiment shown, which will not be elaborated here.

[0060] Step S306: Determine the calculation instructions for the corresponding target operators based on the preset registers of each target node and the association relationship between different target operators. The calculation instructions are used to control the corresponding preset registers to execute the calculation tasks of the target operators.

[0061] Specifically, the above Step S306 includes:

[0062] Step S3061: Determine the external input data information required by the first operator and the preset register of the first operator.

[0063] Exemplarily, in the embodiments of the present application, for the first operator that requires external input, based on the target node corresponding to the first operator, the preset register of the first operator can be determined. There is a corresponding relationship between the target node and the preset register, which will not be elaborated here.

[0064] Step S3062: Generate the first instruction, the second instruction, and the third instruction for the second operator. The first instruction is used to control the preset register corresponding to the first operator to read the external input data from the target storage location based on the external input data information and store the read external input data in the cache. The second instruction is used to control the preset register corresponding to the first operator to execute the calculation task of the first operator based on the external input data, obtaining the calculation result of the first operator. The third instruction is used to send the calculation result of the first operator to the preset location for storage.

[0065] Exemplarily, the target storage location may be a fast memory on the device side. The preset registers for executing the first instruction, the second instruction, and the third instruction may be different SIMD memories. In the embodiment of the present application, taking the fusion operator op1, op2, and op3 structure as an example, the inputs of op1 and op2 are x and y, and the output of op3 is, the first part of the backend code first traverses the inputs x and y of the SIMD block operator provided after the frontend replacement, and generates using INPUT1 = PlaceHolder<x,0> and using INPUT2 = PlaceHolder<y,1> , the 0 and 1 in the brackets are the SIMD register subscripts to which they are bound, and then op1, op2, and op3 are traversed. op1 and op2 will generate instantiated operator template classes using OP1 = Op1<INPUT1,0> and using OP2=Op2<INPUT2,1> , op3 is using OP3=Op3<OP1,OP2,0> , where we can see that op1 reuses register 0, op2 reuses register 1, and op3 reuses register 0, and then instantiate the SIMD block template class using BLOCKTYPE = SimdBlock<OP3,INPUT1,INPUT2> , which marks the dependencies between the operators of the entire fusion block, and finally generates a block object BLOCKTYPEblock and generates a calling code block.run(x,y,z).

[0066] Step S3063: Pack the first instruction, the second instruction, and the third instruction of the first operator to obtain a calculation instruction of the first operator.

[0067] For example, the embodiments of the present application do not limit the packaging method, and those skilled in the art can determine it according to needs.

[0068] In this embodiment, a method for generating a calculation instruction of a fusion operator is provided, which can be used for the above-mentioned server. Figure 4 is a flow chart of a method for generating a calculation instruction of a fusion operator according to an embodiment of the present invention. Figure 4 As shown, the process includes the following steps:

[0069] Step S401, obtaining the identifiers of a fusion operator and multiple preset registers, the fusion operator uses the deep learning operator optimization method of the above embodiment, and the fusion operator includes multiple target operators. For details, please refer to Figure 3 Step S301 of the illustrated embodiment will not be described in detail here.

[0070] Step S402: Hierarchically process the target nodes corresponding to different target operators based on the association relationships between different target operators in the fusion operator, and determine the target node sets corresponding to N different levels. The depths of different target nodes within the same level are the same. For details, please refer to Figure 3 Step S302 of the embodiment shown, which will not be elaborated here.

[0071] Step S403: Let level n = 1, and allocate preset registers to each target node in the nth level through the identifiers of multiple preset registers, to obtain the preset registers corresponding to each target node in the nth level. The input of each target node in the first level is an external input. For details, please refer to Figure 3 Step S303 of the embodiment shown, which will not be elaborated here.

[0072] Step S404: Let level n = n + 1, identify at least one target preset register that will be released after the calculation of different target nodes in the nth level is completed, and allocate registers to different target nodes in the (n + 1)th level based on the identifier of the target preset register and the identifiers of multiple preset registers, to obtain the preset registers corresponding to each target node in the (n + 1)th level. For details, please refer to Figure 2 Step S304 of the embodiment shown, which will not be elaborated here.

[0073] Step S405: If n < N, return to the step of letting level n = n + 1 until n = N, and determine the preset registers of each target node. For details, please refer to Figure 3 Step S305 of the embodiment shown, which will not be elaborated here.

[0074] Step S406: Determine the calculation instructions for the corresponding target operators based on the preset registers of each target node and the association relationships between different target operators. The calculation instructions are used to control the corresponding preset registers to execute the calculation tasks of the target operators.

[0075] Specifically, the above Step S406 includes:

[0076] Step S4061: Generate the first instruction, the second instruction, and the third instruction for the second operator. The first instruction is used to control the preset register corresponding to the second operator to read the target input data, the second instruction is used to control the preset register corresponding to the second operator to execute the calculation task of the second operator based on the target input data to obtain the output result of the second operator, and the third instruction is used to write the output result of the second operator to the target storage location.

[0077] Exemplarily, the target input data may be the output of the target operator on which the second operator depends, or may include external input data. In an embodiment of the present application, the preset registers for executing the first instruction, the second instruction, and the third instruction of the second operator may be different SIMD memories. When it is necessary to read the external input, the preset register includes a fast memory on the device side. When the target input data is the output of the target operator on which the second operator depends, the preset register may involve the cache space of the register, because the output of the target operator on which the second operator depends will be stored in the cache space of the corresponding preset register.

[0078] Step S4062: Pack the first instruction, the second instruction, and the third instruction of the second operator to obtain a calculation instruction of the second operator.

[0079] For example, the embodiments of the present application do not limit the packaging method, and those skilled in the art can determine it according to needs.

[0080] The following is an explanation of the deep learning operator optimization and fusion operator computing instruction generation method provided in this application through a specific embodiment.

[0081] Example:

[0082] (1) Operator fusion refers to replacing multiple consecutive operators with one, aiming to optimize and eliminate bubbles when the host sends a single operator to the device, and the overhead of accessing operators between multi-level caches. In addition, after replacing with one operator, the specific instruction implementation can be fully optimized. For example, the original host needs to receive a signal from the device that the result of operator 1 has been transmitted, and from the device side, it will be found that there is idle time between the calculation of operator 1 and operator 2. If the host only sends one operator once, there will be no such redundancy. The calculation result of operator 1 does not need to be written back to the host side, but only needs to reside in the device side cache and wait for operator 2 to use it. On this basis, the implementation of operator 1 and operator 2 can be optimized in a targeted manner.

[0083] SIMD instruction pipeline arrangement mainly refers to 1. Manually expand the loop to increase pipeline masking; 2. Manually adjust the calculation order to eliminate data hazards SIMD refers to single instruction multiple data flow, that is, one instruction controls multiple parallel processing elements and calculates a set of independent data in parallel. If the input data is 64 in length, then compared with the original implementation of looping 64 times and calculating one each time, it can be optimized by looping 4 times, and one SIMD instruction each time to calculate 16 numbers in parallel. The pipeline masking in the above 1 refers to the overlap of SIMD instructions, thereby covering up some time-consuming phenomena. SIMD instructions are all asynchronously issued, that is, in the absence of data dependency, the next instruction does not need to wait for the previous instruction to be calculated before it can be issued. Each instruction can be regarded as five stages: instruction fetch (IF), decode (ID), execute / effective address (EX), memory access (MEM), and write back (WB). That is, under ideal conditions (the adjacent SIMD instruction sequence is long enough), each instruction only exposes the instruction fetch stage, and the other stages overlap and cover each other. On average, each instruction only needs one instruction cycle to complete. Therefore, in the previous example, the implementation of looping 4 times and executing one SIMD instruction each time can be optimized to an unrolled loop, looping only once, and executing 4 adjacent SIMD instructions each time, so that its calculation process is fully covered by the pipeline. Eliminating data hazards in the above 2 refers to eliminating data dependencies. SIMD calculations need to read data from the on-chip cache space to the SIMD registers, and after the calculation is completed, it needs to be written back to the on-chip cache, which means that 3 instructions are required. The second SIMD calculation instruction needs to rely on the first SIMD read instruction to complete, that is, the second instruction fetch stage depends on the write-back stage of the first instruction to complete, which blocks the instruction pipeline mentioned above. In order to eliminate dependencies and fully pipeline, the SIMD read parts of all operators can be manually arranged at the beginning of the loop, because they have no dependencies on each other, and then SIMD calculations are unified, and finally SIMD writes to the on-chip cache are unified. For example, the original op1 SIMD read -> op1 SIMD calculation -> op1 SIMD write -> op2 SIMD read -> op2 SIMD calculation -> op2 SIMD write has become op1 SIMD read -> op2 SIMD read -> op1 SIMD calculation -> op2 SIMD calculation -> op1 SIMD write -> op2 SIMD write. The advantage of this is that the op2 SIMD read, which is dependent on the op1 SIMD calculation, has been completed, and its asynchronous delivery is not affected.

[0084] (2) Fusion operator code generation algorithm considering the number of registers.

[0085] The design implements a fusion operator code generation algorithm that takes into account the number of vector registers. When performing operator fusion and SIMD instruction pipeline optimization, a large number of SIMD registers need to be used. Since the number of SIMD registers is limited, excessive SIMD register requirements will lead to overflow. Therefore, the optimization goal of the present invention is to maximize the operator fusion depth and instruction pipeline efficiency under the SIMD register number limit to avoid performance loss.

[0086] This algorithm is applied to deep learning computational graphs. A computational graph is a description of a model that abstracts the computational process into computational nodes. Each node includes a description of the input, output, and computational process. The same computational process in the original model will be abstracted into the same computational node, also known as the operator node. The same operator information is extracted to facilitate subsequent optimization.

[0087] After obtaining the calculation graph, the operator fusion will first be performed in the compiler frontend according to the supported fusion operator template. The compiler frontend here refers to the optimization operation based on the high-level graph structure that does not involve the generation of specific instructions. The fusion process here refers to traversing the calculation graph, determining the target operator set and replacing it. The fusion template refers to a matching rule used to determine the target operator set. For example, a fusion template is to fuse as many pointwise operators (pointwiseop) as possible. In this way, whenever a pointwise operator is encountered during the traversal process, the matching logic will be entered. This pointwise operator is the entry operator. At this time, based on the greedy algorithm with the maximum number of SIMD registers and the maximum on-chip cache space as constraints, as many pointwise operators as possible are included in the target operator set. This step is recorded as step 1. Then, the target operator structure is analyzed for liveness, and SIMD registers used for its input, output or possible intermediate results are allocated to each operator. The register space that is no longer used is reused. For example, the input of an operator A is x and y, and the output is z. The inputs x and y are not used in the subsequent calculation process. Then the output z can reuse the register space of input x and replace it with in-situ calculation. The register resources corresponding to y will be released after the calculation is completed. This step of the analysis process is recorded as step 2. After step 2, because some registers are released and the available register resources for reuse increase, more target operators may be obtained by executing step 1 again. Steps 1 and 2 are repeated until the register resources are fully utilized and reasonably bound to each calculation operation. At this time, the fusion operation for this entry operator is completed, and the fused target operator is skipped to continue the original traversal process until the entire calculation graph is traversed. There are more than one fusion templates in this traversal process. There are priorities between fusion templates. For each entry operator, the fusion template with the highest priority will be executed first. The calculation graph traversal is repeated until the calculation graph is stable and no longer replaced. Each target operator set will be replaced by a corresponding block operator, which encapsulates the overall input and output, operator attributes and other information.

[0088] After the compiler front-end fusion replacement, the compiler back-end fusion operator code generation will be performed. The compiler back-end here refers to the low-level instruction generation for each operator that does not involve high-level graph structures. The innovation of this patent in this part is to design and implement a SIMD block template based on intermediate result retention, and optimize the target operator that has been marked for fusion in the above front-end at the code implementation level. The SIMD block template parameter is the structure of the target operator, that is, the type and dependency of the calculation, and the template content, that is, the part shared by all target structures, is a loop body, which contains three steps: reading, calculating, and writing back, that is, the operators in the target structure perform these three steps uniformly in the order of their calculation. For example, if the target structure is three operators op1, op2, and op3, the input of op1 is x, the input of op2 is y, and op3 depends on the results of op1 and op2, and the output after calculation is z, and xyz all refer to the on-chip cache, then the template will first uniformly read the input data in each loop iteration, that is, read the input data of op1op2 from x and y, and then uniformly execute the SIMD calculation instructions of op1op2op3 in sequence, and finally execute the SIMD write-back instruction separately to write the result of op3 back to z. The benefits of doing this are firstly the SIMD instruction pipeline arrangement optimization mentioned in the DSL part of the first subsection above, which maximizes the use of instruction pipeline masking and eliminates data hazards. Secondly, it eliminates redundant access between SIMD registers and on-chip caches of operators, that is, in the above example, op1 and op2 do not need to be SIMD written back to the on-chip cache after calculation, and op3 does not need to be read from the on-chip cache. The calculation results of op1 and op2 reside in the SIMD registers to which they are bound, eliminating redundant memory access hardware features and maximizing utilization.

[0089] 3) Use template metaprogramming technology to encapsulate the SDAAC code template.

[0090] In addition, this patent innovatively applies template metaprogramming technology to SIMD block code generation in the backend of the compiler, with the aim of implementing static analysis to improve compile-time performance, while decoupling the codes of each module to facilitate subsequent expansion and maintenance. Specifically, there are three parts of implementation in the backend. The first part is the instruction generation logic code, which analyzes the front-end code to generate back-end calls; the second part is the SIMD block template code, which has a member function run and a data member simdvalue. The input and output of the member function are the input and output of the entire SIMD block. The function of this method is to loop the calculation of the unified reading, calculation, and writing mentioned above, and the specific relationship between the operators is described by the type of the member simdvalue, and the type of the member is provided by the template parameters of the SIMD block template. The third part is the specific operator class code, which is nested level by level to provide the template parameters of the SIMD block template. The operator class is implemented using CRTP (Curiously Recurring Template Pattern), which is derived from an operator base class and passed to the base class as a template parameter. The main method rewritten is the run method, which is used to provide specific calculation code; at the same time, the operator base class also accepts other operator classes as template parameters to express the dependencies between operators. The second and third parts are on-chip codes, that is, device-side codes, and the first part generates calls to the second and third parts of the code.

[0091] Taking the op1, op2, and op3 structures mentioned above as an example, op1 and op2 are inputs, and op3 is output. The first part of the backend code first traverses the input x and y of the SIMD block operator provided by the frontend after replacement, and generates using INPUT1 = PlaceHolder<x,0> and using INPUT2 = PlaceHolder<y,1> , the 0 and 1 in the brackets are the SIMD register subscripts to which they are bound, and then traverse op1, opop3, op1, op2 will generate an instantiated operator template class using OP1 = Op1<INPUT1,0> and using OP2=Op2<INPUT2,1> , op3 is using OP3=Op3<OP1,OP2,0> , where we can see that op1 reuses register 0, op2 reuses register 1, and op3 reuses register 0, and then instantiate the SIMD block template class using BLOCKTYPE = SimdBlock<OP3,INPUT1,INPUT2> , which marks the dependency between the operators of the entire fusion block, and finally generates a block object BLOCKTYPE block and generates a calling code block.run(x,y,z). The implementation and related content of the above template classes such as PlaceHolder and Op1Op1 are the third part of the code, and SimdBlock is the second part of the code.

[0092] From the above examples, we can see that the advantages of applying template metaprogramming are: 1. All dependencies are based on static analysis, eliminating the redundant overhead caused by the use of virtual functions; 2. The template code and operator implementation are decoupled, with high maintainability and extensibility. Subsequent expansion to support more fusion operators only requires adding a derived subclass to rewrite the run method, without modifying the second part of the template code; 3. Instruction generation and template code are decoupled. If the instruction generation logic needs to be modified, only the first part needs to be modified, without modifying the specific operator or template implementation.

[0093] The solution provided in this application greatly reduces the programming difficulty of users. Users do not need to write complex SIMD vectorized code or deeply understand the underlying hardware optimization strategy. They only need to use the familiar Python language to write code according to the Single Program Multiple Data (SPMD) paradigm. Secondly, the compiler automatically performs operator fusion and SIMD instruction pipeline optimization, which reduces the number of data read and write times, fully utilizes the parallel computing capabilities of the hardware, improves computing performance, and solves the memory access performance bottleneck in the traditional computing process. In addition, through the effective allocation and reuse of SIMD registers, the register overflow problem in the case of deep fusion is avoided, the SIMD register resources are maximized, and the utilization rate of hardware resources is improved. Finally, template metaprogramming technology is used to optimize compile-time performance, decouple code generation and operator implementation, and the code structure is clear and easy to expand and maintain.

[0094] In this embodiment, a deep learning operator optimization device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0095] This embodiment provides a deep learning operator optimization device, such as Figure 5 As shown, including:

[0096] A first acquisition module 501 is used to acquire a deep learning calculation graph, where the deep learning calculation graph includes nodes corresponding to different operators respectively;

[0097] A traversal module 502 is used to traverse each node in the deep learning calculation graph to obtain at least one target node set, where the target node set includes multiple target nodes, and the operator corresponding to each target node is a target operator of a target type, and there is an association relationship between different target operators;

[0098] A fusion module 503 is configured to fuse the operators respectively corresponding to different target nodes in the target node set to obtain a fused operator.

[0099] In some alternative embodiments, the apparatus further includes:

[0100] A third determination module, configured to determine a fused node according to the fused operator;

[0101] An update module, configured to update the deep learning computation graph by using the fused node to obtain an updated computation graph.

[0102] This embodiment provides a computing instruction generation apparatus for a fused operator, as Figure 6 shown, including:

[0103] A second acquisition module 601, configured to acquire a fused operator and identifiers of a plurality of preset registers. The fused operator is obtained by the deep learning operator optimization method of the foregoing embodiment, and the fused operator includes a plurality of target operators;

[0104] A processing module 602, configured to hierarchically process the target nodes respectively corresponding to different target operators based on the association relationship between different target operators in the fused operator, to determine target node sets respectively corresponding to N different levels, and the depths of different target nodes within the same level are the same;

[0105] A first allocation module 603, configured to set level n = 1, and allocate preset registers to each target node in the nth level through the identifiers of the plurality of preset registers, to obtain the preset registers corresponding to each target node in the nth level. The input of each target node in the first level is an external input;

[0106] A second allocation module 604, configured to set level n = n + 1, identify at least one target preset register that will be released after the computation of different target nodes in the nth level is completed, and allocate registers to different target nodes in the (n + 1)th level based on the identifiers of the target preset registers and the identifiers of the plurality of preset registers, to obtain the preset registers corresponding to each target node in the (n + 1)th level;

[0107] A first determination module 605, configured to, if n < N, return to the step of setting level n = n + 1 until n = N, and determine the preset registers of each target node;

[0108] A second determination module 606, configured to determine computation instructions corresponding to the target operators based on the preset registers of each target node and the association relationship between different target operators. The computation instructions are used to control the corresponding preset registers to execute the computation tasks of the target operators.

[0109] In some optional implementations, the plurality of target operators include at least one first operator, the first operator being used to characterize an operator that needs to utilize external input to perform calculations, and the second determination module 606 includes:

[0110] A first determination submodule, used to determine external input data information required by the first operator and a preset register of the first operator;

[0111] A first instruction generation submodule is used to generate a first instruction, a second instruction and a third instruction of a second operator, wherein the first instruction is used to control a preset register corresponding to the first operator to read external input data from a target storage location based on external input data information, and store the read external input data in a cache, the second instruction is used to control the preset register corresponding to the first operator to perform a calculation task of the first operator based on the external input data to obtain a calculation result of the first operator, and the third instruction is used to send the calculation result of the first operator to a preset location for storage;

[0112] The second determination submodule is used to package the first instruction, the second instruction and the third instruction of the first operator to obtain the calculation instruction of the first operator.

[0113] In some optional embodiments, multiple target operators include at least one second operator, and the second operator does not have a lower-level target operator. The step of determining the calculation instructions of the corresponding target operator based on the preset registers of each target node and the association relationship between different target operators includes: generating a first instruction, a second instruction, and a third instruction of the second operator, the first instruction is used to control the preset register corresponding to the second operator to read the target input data, the second instruction is used to control the preset register corresponding to the second operator to perform the calculation task of the second operator based on the target input data to obtain the output result of the second operator, and the third instruction is used to write the output result of the second operator to the target storage location; the first instruction, the second instruction, and the third instruction of the second operator are packaged to obtain the calculation instructions of the second operator.

[0114] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0115] The deep learning operator optimization device and the computing instruction generation device for the fusion operator in this embodiment are presented in the form of functional units, where the units refer to ASIC (Application Specific Integrated Circuit) circuits, processors and memories that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0116] The embodiment of the present invention also provides a computer device having the above Figure 5The deep learning operator optimization device shown, or having Figure 6 The computing instruction generating device of the fusion operator shown.

[0117] See also Figure 7 , Figure 7 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 7 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 7 A processor 10 is taken as an example.

[0118] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0119] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0120] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0121] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0122] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0123] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0124] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0125] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A deep learning operator optimization method, characterized in that: The method includes: Obtain a deep learning computation graph, where the deep learning computation graph includes nodes corresponding to different operators respectively; Traverse each node in the deep learning computation graph to obtain at least one set of target nodes. The set of target nodes includes multiple target nodes, and the operators corresponding to each target node are target operators of a target type, and there is an association relationship between different target operators; Fuse the operators corresponding to different target nodes within the set of target nodes to obtain a fused operator.

2. The method according to claim 1, characterized in that The method further includes: Determine a fused node according to the fused operator; Use the fused node to update the deep learning computation graph to obtain an updated computation graph.

3. A method for generating computing instructions for a fusion operator, characterized in that: The method includes: Obtain a fused operator and the identifiers of multiple preset registers. The fused operator is obtained by the deep learning operator optimization method described in claim 1 or 2, and the fused operator contains multiple target operators; Perform hierarchical processing on the target nodes corresponding to different target operators based on the association relationship between different target operators in the fused operator to determine target node sets corresponding to N different levels. The depths of different target nodes within the same level are the same; Let level n = 1, and allocate preset registers to each target node in the nth level through the identifiers of multiple preset registers to obtain the preset registers corresponding to each target node in the nth level. The input of each target node in the first level is an external input; Let level n = n + 1, identify at least one target preset register that will be released after the computation of different target nodes in the nth level is completed, and allocate registers to different target nodes in the (n + 1)th level based on the identifier of the target preset register and the identifiers of the multiple preset registers to obtain the preset registers corresponding to each target node in the (n + 1)th level; If n < N, return to the step of letting level n = n + 1 until n = N to determine the preset registers of each target node; Determine the computation instructions for the corresponding target operators based on the preset registers of each target node and the association relationship between different target operators. The computation instructions are used to control the corresponding preset registers to execute the computation tasks of the target operators.

4. The method according to claim 3, characterized in that The multiple target operators include at least one first operator, where the first operator is used to represent an operator that needs to use external input to perform computations. The step of determining the computation instructions for the corresponding target operators based on the preset registers of each target node and the association relationship between different target operators includes: Determine the external input data information required by the first operator and the preset register of the first operator; Generate a first instruction, a second instruction, and a third instruction for the second operator. The first instruction is used to control the preset register corresponding to the first operator to read external input data from a target storage location based on the external input data information and store the read external input data in a cache. The second instruction is used to control the preset register corresponding to the first operator to perform the computation task of the first operator based on the external input data to obtain the computation result of the first operator. The third instruction is used to send the computation result of the first operator to a preset location for storage; Pack the first instruction, second instruction, and third instruction of the first operator to obtain the calculation instruction of the first operator.

5. The method according to claim 3, characterized in that: The multiple target operators include at least one second operator that has no subordinate target operator. The step of determining the calculation instruction of the corresponding target operator based on the preset registers of each target node and the association relationship between different target operators includes: Generate a first instruction, a second instruction, and a third instruction for the second operator. The first instruction is used to control the preset register corresponding to the second operator to read the target input data. The second instruction is used to control the preset register corresponding to the second operator to perform the calculation task of the second operator based on the target input data to obtain the output result of the second operator. The third instruction is used to write the output result of the second operator to the target storage location. Pack the first instruction, second instruction, and third instruction of the second operator to obtain the calculation instruction of the second operator.

6. A deep learning operator optimization device, characterized in that: The device includes: A first acquisition module for acquiring a deep learning computation graph, where the deep learning computation graph includes nodes corresponding to different operators. A traversal module for traversing each node in the deep learning computation graph to obtain at least one target node set. The target node set includes multiple target nodes, and the operator corresponding to each target node is a target operator of a target type, and there is an association relationship between different target operators. A fusion module for fusing the operators corresponding to different target nodes in the target node set to obtain a fused operator.

7. A computing instruction generating device for a fusion operator, characterized in that: The device includes: A second acquisition module for acquiring a fused operator and the identifiers of multiple preset registers. The fused operator is obtained by the deep learning operator optimization method described in claim 1 or 2, and the fused operator includes multiple target operators. A processing module for hierarchically processing the target nodes corresponding to different target operators in the fused operator based on the association relationship between different target operators in the fused operator to determine target node sets corresponding to N different levels, and the depths of different target nodes within the same level are the same. A first allocation module for setting level n = 1, and allocating preset registers to each target node in the nth level through the identifiers of multiple preset registers to obtain the preset registers corresponding to each target node in the nth level. The input of each target node in the first level is an external input. A second allocation module for setting level n = n + 1, identifying at least one target preset register that will be released after the calculation of different target nodes in the nth level is completed, and allocating registers to different target nodes in the (n + 1)th level based on the identifier of the target preset register and the identifiers of the multiple preset registers to obtain the preset registers corresponding to each target node in the (n + 1)th level. A first determination module for, if n < N, returning to the step of setting level n = n + 1 until n = N to determine the preset registers of each target node. The second determination module is used to determine the calculation instructions of the corresponding target operator based on the preset registers of each target node and the association relationship between different target operators, and the calculation instructions are used to control the corresponding preset registers to execute the calculation tasks of the target operator.

8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the deep learning operator optimization method described in claim 1 or 2, or executes the computing instruction generation method for the fusion operator described in any one of claims 3 to 5 by executing the computer instructions.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the deep learning operator optimization method described in claim 1 or 2, or the calculation instruction generation method for the fusion operator described in any one of claims 3 to 5.

10. A computer program product, characterized in that It includes computer instructions, which are used to enable a computer to execute the deep learning operator optimization method described in claim 1 or 2, or the calculation instruction generation method for the fusion operator described in any one of claims 3 to 5.

Citation Information

Cited By

  • Data integration method and device based on multi-dimensional weight mechanism, equipment and medium

    CN120277130A

  • Data transmission method and device, electronic equipment and storage medium

    CN121940447A