A Compilation and Execution Method, Device, Equipment and Medium for a Dynamic Neural Network

By obtaining the dynamic fusion fragments of the dynamic neural network during the compilation stage and converting them into static fusion fragments, the problem that the dynamic neural network cannot be optimized in the compilation stage is solved, and higher execution efficiency and performance are achieved.

CN119806541BActive Publication Date: 2025-05-27SHANGHAI ENFLAME INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510300600.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-05-27
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

Because dynamic neural networks cannot obtain exact shape information during the compilation stage, they cannot perform compilation cache optimization, resulting in significant performance and execution efficiency.

Method used

During the compilation stage, dynamic fusion fragments are obtained by obtaining each operator node of the dynamic neural network, and the optimal fusion slice is calculated based on the chip parameters of the multi-level memory chip. The dynamic fusion fragment is converted into a static fusion fragment, forming a fusion unit, tensor handling is performed in the shared storage layer, and the fusion unit is dynamically called during the execution stage.

Benefits of technology

By optimizing in the compilation stage, dynamic shape-related computing nodes are eliminated, the execution efficiency of dynamic neural networks is improved, and performance is improved through static scheduling planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119806541B_ABST
    Figure CN119806541B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, equipment and medium for compiling and executing a dynamic neural network. The method includes: calculating the optimal fusion slice corresponding to each dynamic fusion segment according to chip parameters during the compilation stage; obtaining the static fusion segment matching the corresponding dynamic fusion segment according to the optimal fusion slice, and forming a fusion unit with each dynamic fusion segment and the matching static fusion segment; obtaining the loop scalar control logic of the dynamic neural network during the execution stage, and dynamically calling the fusion unit based on the loop scalar control logic. Without relying on shape information, the optimal fusion slice suitable for the hardware cache size of each dynamic fusion segment is obtained, and the dynamic fusion segment is converted into a static fusion segment based on the optimal fusion slice to eliminate the calculation nodes related to the dynamic shape. During the execution stage, dynamic scheduling and planning can be performed on the static fusion segment, greatly improving the execution efficiency of the dynamic neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to computer technology, and in particular, to a method, device, equipment, and medium for compiling and executing a dynamic neural network. Background Art

[0002] Compared with a static neural network, the input shape of a dynamic neural network is uncertain. With the application of neural networks in scenarios such as large language models and image processing, more and more dynamic neural networks with variable shapes are being applied in various fields.

[0003] However, since a dynamic neural network cannot obtain exact shape information during the compilation stage, it is impossible to perform compilation-time cache optimization. All data is obtained from the external memory space of the chip, which significantly reduces the performance and execution efficiency of the dynamic neural network. Summary of the Invention

[0004] The embodiments of the present invention provide a method for compiling and executing a dynamic neural network to achieve compilation optimization of the dynamic neural network during the compilation stage.

[0005] In a first aspect, the embodiments of the present invention provide a method for compiling and executing a dynamic neural network, including: obtaining dynamic fusion segments according to each operator node of the dynamic neural network during the compilation stage;

[0006] Obtaining the chip parameters of the multi-level storage chip carrying the dynamic neural network, and calculating the optimal fusion slices corresponding to each of the dynamic fusion segments according to the chip parameters;

[0007] Obtaining the static fusion segments matching the corresponding dynamic fusion segments according to the optimal fusion slices, and forming a fusion unit with each of the dynamic fusion segments and the matching static fusion segments, wherein the fusion unit performs tensor transfer in the shared storage layer of the multi-level storage chip;

[0008] Obtaining the loop scalar control logic of the dynamic neural network during the execution stage, and dynamically calling the fusion unit based on the loop scalar control logic.

[0009] In a second aspect, the embodiments of the present invention provide a device for compiling and executing a dynamic neural network, including: a dynamic fusion segment obtaining module, configured to obtain dynamic fusion segments according to each operator node of the dynamic neural network during the compilation stage;

[0010] An optimal fusion slice obtaining unit, configured to obtain the chip parameters of the multi-level storage chip carrying the dynamic neural network, and calculate the optimal fusion slices corresponding to each of the dynamic fusion segments according to the chip parameters;

[0011] The fusion unit acquisition module is used to obtain the static fusion fragments corresponding to the dynamic fusion fragments matched according to the optimal fusion slices, and form a fusion unit with each of the dynamic fusion fragments and the matched static fusion fragments, wherein the fusion unit performs tensor transfer in the shared storage layer of the multi-level storage chip;

[0012] The fusion unit call module is used to obtain the loop scalar control logic of the dynamic neural network during the execution phase, and dynamically call the fusion unit based on the loop scalar control logic.

[0013] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method as described above when executing the program.

[0014] In a fourth aspect, an embodiment of the present invention provides a storage medium storing computer-executable instructions, on which a computer program is stored, wherein the program implements the method as described above when executed by a processor.

[0015] In the present invention, dynamic fusion fragments are obtained during the compilation phase, the optimal fusion slices suitable for the hardware cache size of each dynamic fusion fragment are obtained without relying on shape information, and the dynamic fusion fragments are converted into static fusion fragments based on the optimal fusion slices to eliminate the calculation nodes related to the dynamic shape, so as to realize the optimization of the dynamic neural network during the compilation phase. During the execution phase, dynamic scheduling planning can be performed on the static fusion fragments, so as to convert the internal processing dynamics at the original operator level into the dynamic processing of the scheduling logic, greatly improving the execution efficiency of the dynamic neural network. Description of the Drawings

[0016] Figure 1 is a flowchart of a method for compiling and executing a dynamic neural network provided in Embodiment 1 of the present invention;

[0017] Figure 2 is a schematic structural diagram of a dynamic neural network provided in Embodiment 1 of the present invention;

[0018] Figure 3 is a schematic structural diagram of the compiled dynamic neural network provided in Embodiment 1 of the present invention;

[0019] Figure 4 is a flowchart of a method for compiling and executing a dynamic neural network provided in Embodiment 2 of the present invention;

[0020] Figure 5 is a schematic structural diagram of a device for compiling and executing a dynamic neural network provided in Embodiment 3 of the present invention;

[0021] Figure 6 It is a schematic structural diagram of a computer device provided in Embodiment 4 of the present invention. Specific implementation manners

[0022] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the accompanying drawings.

[0023] Embodiment 1

[0024] Figure 1 It is a flowchart of a method for compiling and executing a dynamic neural network provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of compiling and executing a dynamic neural network. This method can be executed by a device for compiling and executing a dynamic neural network, and this device can be implemented in the form of hardware and / or software. As Figure 1 shown, the method includes:

[0025] Step S101, obtaining a dynamic fusion segment according to each operator node of the dynamic neural network during the compilation phase.

[0026] Optionally, obtaining a dynamic fusion segment according to each operator node of the dynamic neural network during the compilation phase includes: screening out fusion operator nodes from each operator node of the dynamic neural network during the compilation phase; obtaining a dynamic fusion segment according to the fusion operator nodes, where the dynamic fusion segment includes at least two fusion operator nodes with a dependency relationship.

[0027] Specifically, a dynamic neural network is a directed acyclic graph composed of several operator nodes and data dependencies, and an operator node includes an input, a calculation process, and an output. Among them, the calculation process is usually fixed for the type of the operator node itself. For example, convolution, pooling, addition, and subtraction. At the same time, the attributes of the operator node itself are also determined during the operator phase. For example, the stride of convolution, the size of the convolution kernel, etc. These information are also determined for the dynamic neural network, and the only thing that is uncertain is only some specific dimensions of some inputs. For example, the width and height dimensions of convolution, or the specific batch size. And for a dynamic neural network in a specific field, which dimensions are dynamic dimensions are determined during the compilation phase of the network, and all operator nodes in the network basically follow them. Therefore, the dynamic dimensions of the inputs of all operator nodes in the entire dynamic neural network can also be determined during the compilation period. As Figure 2The following is a schematic structural diagram of the dynamic neural network in this embodiment. In this dynamic neural network, there are seven dynamic operator nodes, such as dynamic OP1, dynamic OP2... dynamic OP7, and the input shapes of each operator node are uncertain, so the data is obtained from the external memory space of the chip. Of course, this is only an example in this embodiment, and the number of operator nodes included in the dynamic neural network is not limited. In addition, in this embodiment, in order to optimize the dynamic neural network during the compilation stage, the fusion operator nodes are selected from each operator node of the dynamic neural network, and the dynamic fusion segment is obtained based on the fusion operator nodes.

[0028] Optionally, selecting the fusion operator nodes from each operator node of the dynamic neural network during the compilation stage includes: obtaining the weight transfer types of each operator node in the dynamic neural network, where the weight transfer types include weight transfer sparse type and weight transfer dense type; obtaining the dynamic dimensions and reduction calculation dimensions of each operator node in the dynamic neural network; and taking the operator nodes in the dynamic neural network with the weight transfer type of weight transfer sparse type and the dynamic dimensions not coinciding with the reduction calculation dimensions as the fusion operator nodes.

[0029] Optionally, obtaining the weight transfer types of each operator node in the dynamic neural network includes: obtaining the weight size and chip bandwidth of each operator node in the dynamic neural network, and obtaining the weight transfer overhead according to the ratio of the weight size to the chip bandwidth; obtaining the data transfer overhead and data calculation overhead of each operator node in the dynamic neural network, and obtaining the overhead threshold according to the sum of the data transfer overhead and the data calculation overhead; and determining whether the weight transfer overhead of the operator node is less than the overhead threshold. If so, determining that the weight transfer type of the operator node is the weight transfer sparse type, otherwise, determining that the weight transfer type of the operator node is the weight transfer dense type.

[0030] Specifically, in this embodiment, when analyzing each operator node in the dynamic neural network one by one, the screening is specifically carried out from two aspects: function evaluation and performance evaluation, and the operator nodes that meet the requirements are used as the fused operator nodes. On the one hand, when performing performance evaluation, it is specifically to judge the weight transfer type of the operator node, and the operator node with the weight transfer type of sparse weight transfer is used as the candidate operator node. Among them, the weight transfer type includes sparse weight transfer and dense weight transfer. When determining the weight transfer type of the operator node, it is specifically to obtain the weight size of the operator node, the chip bandwidth, and the data transfer overhead and data calculation overhead. Among them, the weight size and the chip bandwidth are known quantities and can be directly obtained. The data transfer overhead and data calculation overhead need to be calculated by a certain calculation method based on the operation characteristics of the operator node. The specific calculation method is not the focus of this application, so it will not be elaborated in this embodiment. In this embodiment, the weight transfer overhead will be obtained according to the ratio of the weight size and the chip bandwidth, and the overhead threshold will be obtained according to the data transfer overhead and the data calculation overhead, so as to determine the weight transfer type of the operator node based on the relationship between the weight transfer overhead of the operator node and the overhead threshold. For example, when the weight transfer overhead is less than the overhead threshold, it is determined that the operator node is of the sparse weight transfer type; when the weight transfer overhead is greater than the overhead threshold, it is determined that the operator node is of the dense weight transfer type. Of course, only examples are given in this embodiment, and the specific determination method of the weight transfer type of the operator node is not limited. For example, through performance evaluation, six operator nodes such as dynamic OP1, dynamic OP2, dynamic OP3, dynamic OP4, dynamic OP5, and dynamic OP6 are obtained as six candidate operator nodes that meet the performance requirements. On the other hand, after performance evaluation, in this embodiment, the candidate operator nodes will also be subjected to function evaluation, specifically to judge whether the dynamic dimension and the reduction calculation dimension of the operator node coincide, and the candidate operator nodes that do not coincide are used as the fused operator nodes. For example, through function evaluation, five fused operator nodes such as dynamic OP2, dynamic OP3, dynamic OP4, dynamic OP5, and dynamic OP6 that meet the function requirements are obtained. Among them, the dynamic dimension of the operator node is a known quantity as it has been determined during the compilation stage of the network, and the reduction calculation dimension of the operator node can be obtained by looking up the reduction dimension table. Among them, the reduction dimension table records the dimensions with reduction calculations corresponding to different types of operator nodes, as shown in Table 1 below as an example of the reduction dimension table:

[0031] Table 1

[0032]

[0033] Among them, due to space limitations, only three types of operator nodes are taken as examples in Table 1 for illustration. In actual applications, the specific number of operator types recorded in the reduction dimension table in this embodiment is not limited.

[0034] Optionally, obtaining a dynamic fusion segment according to the fusion operator nodes includes: obtaining the input-output ratio of each fusion operator node; searching and fusing the fusion operator nodes with input-output ratios within the first numerical range and having a dependency relationship by using a greedy algorithm to obtain at least one dynamic fusion segment.

[0035] Specifically, in this embodiment, after obtaining the fusion operator nodes, such as dynamic OP2, dynamic OP3, dynamic OP4, dynamic OP5, and dynamic OP6, the input-output ratio of each fusion operator node will be obtained, and the fusion operator nodes with input-output ratios within the first numerical range and having a dependency relationship will be fused to obtain a dynamic fusion segment. During the fusion process, a greedy algorithm will be used to search for dynamic fusion segments in the dynamic neural network, and the number of fusion operator nodes included in a dynamic fusion segment is at least 2. Although the principle of "fusing as much as possible" is followed during the search process, in actual applications, the fusion is usually interrupted due to functional limitations or performance limitations. Therefore, the dynamic fusion segment is not infinitely large.

[0036] Among them, in terms of performance limitations, although the specific shape information (i.e., the absolute size) of each fused operator node cannot be obtained by the compiler, the multiple relationship (i.e., the relative size) is determined when multiple fused operator nodes are connected in series. For example, fused operator node a is an element-wise operation operator, the size of its input dynamic dimension is x, the output of a is given to fused operator node b, b is a convolutional operator with a stride of 2, the input of b is equal to the output x of a, and the output of b becomes 1 / 2x. That is to say, although the exact value of the dynamic dimension of the output of operator b is unknown, it must be half of the input x. Such multiple relationships can be deduced and determined during compilation. That is, although the actual memory size of a dynamic shape network cannot be obtained during compilation, the memory size trend curve can be obtained. Therefore, in this embodiment, specifically, fused operator nodes with consistent memory trends (i.e., the input-output ratio is within the first numerical range) are fused, where the first numerical range can be 1 - 1.3, that is, fused operator nodes with similar or even the same input-output ratio are fused. In addition, in terms of functional limitations, during the fusion process using the greedy algorithm, if the dynamic dimension of the newly fused dynamic fusion segment does not overlap with the reduction calculation dimension, they can be fused together, but when the dynamic dimension overlaps with the reduction calculation dimension, the fusion will be interrupted. For example, the dynamic dimensions of the dynamic fusion segment obtained after fusing dynamic OP4, dynamic OP5, and dynamic OP6 do not overlap with the reduction calculation dimension, but when dynamic OP3 is incorporated and the dynamic dimension of the newly obtained dynamic fusion segment through amplification overlaps with the reduction calculation dimension, the fusion will be interrupted between dynamic OP4 and dynamic OP3 at this time. Thus, it can be determined that dynamic OP4, dynamic OP5, and dynamic OP6 form a dynamic fusion segment 1, and a new anchor point is set as dynamic OP3 to perform greedy search upward again, obtaining a dynamic fusion segment 2 containing dynamic OP2 and dynamic OP3. Of course, in this embodiment, only two dynamic fusion segments are taken as examples for illustration, and the specific number of the obtained dynamic fusion segments is not limited.

[0037] Step S102: Obtain the chip parameters of the multi-level storage chip carrying the dynamic neural network, and calculate the optimal fusion slices corresponding to each dynamic fusion segment according to the chip parameters.

[0038] Optionally, the chip parameters include the chip cache size. Calculating the optimal fusion slice corresponding to each dynamic fusion segment based on the chip parameters includes: setting an initial slice value for the dynamic dimension of each dynamic fusion segment, and calculating the memory cache size occupied by the dynamic fusion segment at the initial slice value; determining whether the memory cache size is greater than the chip cache size, if so, reducing the initial slice value to obtain an updated slice value, otherwise, increasing the initial slice value to obtain an updated slice value; calculating the difference between the updated slice value and the chip cache size, and when the difference is within the second numerical range, taking the updated slice value as the optimal fusion slice.

[0039] Optionally, after calculating the optimal fusion slice corresponding to each dynamic fusion segment according to the chip parameters, it further includes: aligning the dimension size of the optimal fusion slice to an even number; calculating the data volume size corresponding to the optimally fusion slice after even alignment, and aligning the data volume size to an integer multiple with reference to the chip transfer bandwidth.

[0040] Specifically, in this embodiment, the multi-level storage chip can carry the operation of the dynamic neural network, and the chip parameters of the chip are known. For example, the chip cache size and chip bandwidth of each level are known. Since the dynamic dimensions of each fusion operator node are known, the dynamic dimensions of the obtained dynamic fusion segments after fusion are also known. Therefore, in this embodiment, an initial slice value can be set for each dynamic fusion segment for trial calculation to obtain the memory cache size of the dynamic fusion segment at the given initial slice value, so as to determine whether the current slice can be completely cached on the chip. If the memory cache size is smaller than the chip cache size, it means it can be placed, and at this time, an attempt is made to increase the initial slice value to obtain an updated slice value; if the memory cache size is greater than the chip cache size, it means it cannot be placed, and at this time, an attempt is made to reduce the initial slice value to obtain an updated slice value. Repeat the above operations until the maximum value of the dynamic dimension of the dynamic fusion segment that can be cached by the current chip is found, and this value is the limit that the chip can just cache. When specifically calculating, the difference between the updated slice value and the chip cache size is obtained, and when the difference is within the second numerical range, the updated slice value is taken as the optimal fusion slice. For example, the second numerical range can be 0 - 0.5. Of course, in this embodiment, only an example is given, and the specific content of the second numerical range is not limited. For example, when there is only one dynamic dimension in the dynamic fusion segment 1, the obtained optimal fusion slice is 320. Correspondingly, when there are two dynamic dimensions, the optimal fusion slices corresponding to each dynamic dimension are 320 and 340, that is, two optimal fusion slices. Of course, in this embodiment, only an example is given, and the specific values of the optimal fusion slices corresponding to each dynamic fusion segment are not limited.

[0041] It should be noted that, in this embodiment, after obtaining the optimal fusion slices corresponding to each dynamic fusion segment, alignment processing is also performed on the optimal fusion slices to ensure the full utilization of hardware resources such as chip computing power and bandwidth. Among them, the alignment processing is specifically carried out from two perspectives. The first perspective is to align the dimension size of the optimal fusion slice to an even number, that is, to align it to a value friendly to the chip computing unit, such as a power of 2. For example, when the optimal fusion slice is 323, it will be aligned to 320. The second perspective is to align the data volume size corresponding to the optimal fusion slice in integer multiples with reference to the chip transfer bandwidth. For example, for dynamic fusion segment 1, the original is 367*?*360, where "?" represents the position of the dynamic dimension. Then, after the optimal fusion slice is determined, it is substituted into the position of the dynamic dimension to obtain the data volume "367×320×360", and it is necessary to ensure that the product result is an integer multiple of 512 bits. If not, the optimal fusion slice needs to be further aligned and adjusted. Of course, only examples are given in this embodiment, and the specific alignment method of the optimal fusion slice is not limited.

[0042] Step S103, obtain the static fusion segment matching the corresponding dynamic fusion segment according to the optimal fusion slice, and form a fusion unit with each dynamic fusion segment and the matching static fusion segment.

[0043] Optionally, obtaining the static fusion segment matching the corresponding dynamic fusion segment according to the optimal fusion slice includes: filling each aligned optimal fusion slice onto the dynamic dimension of the corresponding dynamic fusion segment to obtain the static fusion segment matching the dynamic fusion segment; performing constant folding and algebraic simplification on each static fusion segment to obtain the static fusion segment optimized for the ordinary layer; performing static operator fusion and redundant operator node elimination on the static fusion segment optimized for the ordinary layer to obtain the static fusion segment optimized for the cache layer.

[0044] Specifically, in this embodiment, after finding the aligned optimal fusion slices corresponding to each dynamic fusion segment, the optimal fusion slices after alignment can be filled onto the dynamic dimension of the corresponding dynamic fusion segment to obtain the static fusion segment matching the dynamic fusion segment. And after obtaining the static fusion segment, a series of internal optimizations can be performed, such as ordinary layer optimizations like constant folding and algebraic simplification, and cache layer optimizations like static operator fusion and redundant operator node elimination. Since the specific optimization process of the static neural network is not the focus of this application, it will not be elaborated in this embodiment. At the same time, for the operator nodes, since the determined size information has been obtained, the operator nodes can be compiled to select the most suitable data stream and computing stream for the current size, rather than the dynamic operator that needs to apply various input shapes and thus uses a more general data stream and computing stream, thereby improving the operation performance in various specific scenarios.

[0045] It should be noted that a prerequisite for converting a dynamic fusion segment into a static fusion segment in this embodiment is that all dynamic dimensions in the dynamic fusion segment need to be segmented, and the obtained segmentation size and static values are used to replace the dynamic dimensions, so as to achieve the transformation from dynamic to static. If any one dimension is still dynamic, it cannot be converted into a static fusion segment. In addition, the original dynamic fusion segment in this embodiment will be retained, and the converted static fusion segment and the corresponding dynamic fusion segment are formed into a fusion unit. The fusion unit performs tensor transfer in the shared storage layer of the multi-level storage chip, that is, the data in the fusion operator node located in the fusion unit is obtained from the internal space of the chip, without the need to frequently transfer from the outside of the chip, thereby improving the overall performance of the dynamic neural network. Among them, as Figure 3 shown in the structural schematic diagram of the compiled dynamic neural network, and for each dynamic neural network, a corresponding loop scalar control logic is pre-compiled. During subsequent operation, when each fusion unit is executed, this loop scalar control logic will be called once to determine the calling method for each fusion unit. Of course, in this embodiment, only the case where the dynamic neural network contains two fusion units is taken as an example for illustration, and the specific number of fusion units in the dynamic neural network is not limited in this embodiment.

[0046] Step S104, obtain the loop scalar control logic of the dynamic neural network during the execution phase, and dynamically call the fusion unit based on the loop scalar control logic.

[0047] Optionally, dynamically calling the fusion unit based on the loop scalar control logic includes: obtaining the best fusion slice corresponding to the currently to-be-executed fusion unit and the actual input size of the dynamic neural network; inputting the actual input size and the best fusion slice into the loop scalar control logic to obtain the calling method for the fusion unit, and dynamically calling the currently to-be-executed fusion unit based on the calling method.

[0048] Optionally, the actual input size and the best fusion slice input loop scalar control logic are used to obtain the calling method for the fusion unit, including: when the actual input size is smaller than the best fusion slice, it is determined that the calling method of the loop scalar control logic for the fusion unit is to only call the dynamic fusion segment in the fusion unit once; when the actual input size is not less than the best fusion slice and is divisible, the integer value obtained by dividing the actual input size by the best fusion slice is obtained, and it is determined that the calling method of the loop scalar control logic for the fusion unit is to loop and call the static fusion segments in the fusion unit with the integer value; when the actual input size is greater than the best fusion slice and is not divisible, the integer value and the remainder value obtained by dividing the actual input size by the best fusion slice are obtained, and it is determined that the calling method of the loop scalar control logic for the fusion unit is to loop and call the static fusion segments in the fusion unit with the integer value, and to call the dynamic fusion segments in the fusion unit with the remainder value.

[0049] Specifically, in this embodiment, the actual input size of the dynamic neural network can be determined during the execution phase. At this time, the specific calling method for each fusion unit for the current input size, that is, how many times the dynamic fusion segment and the static fusion segment need to be called, will be calculated based on the loop scalar control logic. Among them, the loop scalar control logic in this embodiment is mainly applied to the following two application scenarios: Scenario one is that when there is a dynamic dimension that needs to be looped and expanded N times, it can be ensured that the first (N - 1) loops enter the static graph calculation. If the Nth loop cannot be divisible by the slice size, it needs to be calculated by the dynamic graph; Scenario two is that when multiple dynamic dimensions need to be sliced and expanded, for example, the convolution width / height is a dynamic dimension, and the width / height is expanded M / N times respectively, then it can be ensured that (M - 1)*(N - 1) loops enter the static graph. If the remaining part cannot be divisible, it needs to be calculated by the dynamic graph. Assuming that the best static slice size is H_step = 64 / W_step = 64 and the dynamic shape is H = 128 / W = 128, the static graph can be hit and looped 2 times ((128 / 64)*(128 / 64)); when the dynamic shape is H = 32 / W = 128, the static graph cannot be hit. In summary, it can be known that the larger the dynamic shape, the more parts of the loop count that hit the static graph.

[0050] In a specific implementation, when running to the fusion unit 1, the loop scalar control logic is called, and the actual input size of the current dynamic neural network and the optimal fusion slice A are both input into the loop scalar control logic, and the loop scalar control logic determines the call times of the dynamic fusion segment 1 and the static fusion segment 1. For example, when the actual input size is smaller than the optimal fusion slice, it indicates that the current data processing volume is relatively small, so the fusion unit 1 does not require any segmentation or loop and only needs to call the dynamic fusion segment once; when the actual input size is not less than the optimal fusion slice and is divisible, the integer value m obtained by dividing the actual input size by the optimal fusion slice is calculated, and the static fusion segment 1 is called m times; when the actual input size is greater than the optimal fusion slice and is not divisible, the integer value p and the remainder value q obtained by dividing the actual input size by the optimal fusion slice are calculated, and the static fusion segment 1 is called p times and the dynamic fusion segment 1 is called q times. Usually, the number of times the dynamic fusion segment 1 is called is very small, and most of the calls are to the static fusion segment. And when the dynamic fusion segment 1 and the static fusion segment 1 are called simultaneously, according to the actually calculated runtime calculation logic, after calling the static fusion segment 1 and the dynamic fusion segment 1, the results after the calculation of each fusion segment are concatenated and then written back to the original output memory space, so as to complete the calculation and output of the entire fusion result. Through the above steps, the dynamic nature of a dynamic neural network is transformed from complete internal operator processing to dynamic scheduling of layers plus static operators, and on the host side, the dynamic scheduling usually has a scalar calculation overhead that is almost zero, thus realizing the transformation of a dynamic network into a static slice network, enabling the performance of layer operators to be fully optimized and ensuring the performance of the entire dynamic neural network. Of course, in this embodiment, only the call to the fusion unit 1 is used as an example for illustration, and the method of calling other fusion units is roughly the same, and will not be elaborated in this embodiment.

[0051] It is worth mentioning that in this embodiment, it does not rely on shape information, can automatically find the optimal fusion slice suitable for the hardware cache size, and provides the possibility of fusion for the dynamic network; according to the determined optimal fusion slice of the searched size, the dynamic shape of the neural network is converted into a static shape, and the calculation nodes related to the dynamic shape can be eliminated; further optimization processing of the converted static fusion slice can make full use of various optimization means in the static network, including algebraic optimization, static operator fusion, and constant folding, greatly improving the performance of the fusion slice; the originally dynamic operators in the network are transformed into dynamic scheduling plus static operators, and the extreme performance of the static operators can be obtained with extremely small scheduling overhead.

[0052] In this embodiment, dynamic fusion segments are obtained during the compilation phase. Without relying on shape information, the optimal fusion slices suitable for the hardware cache size of each dynamic fusion segment are obtained, and the dynamic fusion segments are converted into static fusion segments based on the optimal fusion slices to eliminate the computational nodes related to dynamic shapes, thereby achieving the optimization of dynamic neural networks during the compilation phase. During the execution phase, dynamic scheduling planning can be performed on the static fusion segments, thus converting the internal processing dynamics at the original operator level into dynamic processing of the scheduling logic, greatly improving the execution efficiency of dynamic neural networks.

[0053] Embodiment 2

[0054] Figure 4 FIG. is a flowchart of a method for compiling and executing a dynamic neural network provided in Embodiment 2 of the present invention. Based on the above embodiment, after dynamically calling the fusion unit based on the loop scalar control logic, the method further includes: detecting the call results of each fusion unit of the dynamic neural network. As Figure 4 shown, the method includes:

[0055] Step S201, obtaining dynamic fusion segments according to each operator node of the dynamic neural network during the compilation phase.

[0056] Optionally, obtaining dynamic fusion segments according to each operator node of the dynamic neural network during the compilation phase includes: screening out fusion operator nodes from each operator node of the dynamic neural network during the compilation phase; obtaining dynamic fusion segments according to the fusion operator nodes, where the dynamic fusion segments include at least two fusion operator nodes with dependency relationships.

[0057] Optionally, screening out fusion operator nodes from each operator node of the dynamic neural network during the compilation phase includes: obtaining the weight transfer types of each operator node in the dynamic neural network, where the weight transfer types include weight transfer sparse type and weight transfer dense type; obtaining the dynamic dimensions and reduction calculation dimensions of each operator node in the dynamic neural network; using the operator nodes in the dynamic neural network with the weight transfer type of weight transfer sparse type and the dynamic dimensions not coinciding with the reduction calculation dimensions as fusion operator nodes.

[0058] Optionally, obtaining the weight transfer types of each operator node in the dynamic neural network includes: obtaining the weight size and chip bandwidth of each operator node in the dynamic neural network, and obtaining the weight transfer overhead according to the ratio of the weight size to the chip bandwidth; obtaining the data transfer overhead and data calculation overhead of each operator node in the dynamic neural network, and obtaining an overhead threshold according to the sum of the data transfer overhead and the data calculation overhead; determining whether the weight transfer overhead of the operator node is less than the overhead threshold. If so, determining that the weight transfer type of the operator node is weight transfer sparse type, otherwise, determining that the weight transfer type of the operator node is weight transfer dense type.

[0059] Optionally, obtain dynamic fusion segments according to the fusion operator nodes, including: obtaining the input-output ratios of each fusion operator node; searching and fusing the fusion operator nodes with input-output ratios within the first numerical range and having a dependency relationship by using a greedy algorithm to obtain at least one dynamic fusion segment.

[0060] Step S202, obtain the chip parameters of the multi-level storage chip carrying the dynamic neural network, and calculate the optimal fusion slices corresponding to each dynamic fusion segment according to the chip parameters.

[0061] Optionally, the chip parameters include the chip cache size. Calculating the optimal fusion slices corresponding to each dynamic fusion segment according to the chip parameters includes: setting an initial slice value for the dynamic dimension of each dynamic fusion segment, and calculating the memory cache size occupied by the dynamic fusion segment under the initial slice value; determining whether the memory cache size is greater than the chip cache size, if so, reducing the initial slice value to obtain an updated slice value, otherwise, increasing the initial slice value to obtain an updated slice value; calculating the difference between the updated slice value and the chip cache size, and when the difference is within the second numerical range, taking the updated slice value as the optimal fusion slice.

[0062] Optionally, after calculating the optimal fusion slices corresponding to each dynamic fusion segment according to the chip parameters, it further includes: aligning the dimension sizes of the optimal fusion slices to be even; calculating the data volume size corresponding to the optimally fusion slices after even alignment, and aligning the data volume size to an integer multiple with reference to the chip transfer bandwidth.

[0063] Step S203, obtain the static fusion segments matching the corresponding dynamic fusion segments according to the optimal fusion slices, and form fusion units with each dynamic fusion segment and the matching static fusion segments.

[0064] Optionally, obtaining the static fusion segments matching the corresponding dynamic fusion segments according to the optimal fusion slices includes: filling each aligned optimal fusion slice onto the dynamic dimension of the corresponding dynamic fusion segment to obtain the static fusion segments matching the dynamic fusion segments; performing constant folding and algebraic simplification on each static fusion segment to obtain the static fusion segments optimized for ordinary layers; performing static operator fusion and eliminating redundant operator nodes on the static fusion segments optimized for ordinary layers to obtain the static fusion segments optimized for cache layers.

[0065] Step S204, obtain the loop scalar control logic of the dynamic neural network during the execution phase, and dynamically call the fusion unit based on the loop scalar control logic.

[0066] Optionally, the fusion unit is dynamically called based on the loop scalar control logic, including: obtaining the optimal fusion slice corresponding to the currently to-be-executed fusion unit and the actual input size of the dynamic neural network; inputting the actual input size and the optimal fusion slice into the loop scalar control logic to obtain the calling method for the fusion unit, and dynamically calling the currently to-be-executed fusion unit based on the calling method.

[0067] Optionally, inputting the actual input size and the optimal fusion slice into the loop scalar control logic to obtain the calling method for the fusion unit, including: when the actual input size is less than the optimal fusion slice, determining that the calling method of the loop scalar control logic for the fusion unit is to only call the dynamic fusion segment in the fusion unit once; when the actual input size is not less than the optimal fusion slice and is divisible, obtaining the integer value of the actual input size divided by the optimal fusion slice, and determining that the calling method of the loop scalar control logic for the fusion unit is to loop and call the static fusion segments in the fusion units with the integer value; when the actual input size is greater than the optimal fusion slice and is not divisible, obtaining the integer value and the remainder value of the actual input size divided by the optimal fusion slice, and determining that the calling method of the loop scalar control logic for the fusion unit is to loop and call the static fusion segments in the fusion units with the integer value, and call the dynamic fusion segment in the fusion unit with the remainder value.

[0068] Step S205, detecting the call results of each fusion unit of the dynamic neural network.

[0069] Specifically, in this embodiment, after the dynamic neural network finishes execution, the call results of each fusion unit will be detected. For example, it is detected whether the fusion unit calls the dynamic fusion segment and the static fusion segment according to the calculated number of calls, or whether there are obvious errors such as garbled calculation outputs in the fusion unit. When it is determined that the fusion unit is not called according to the calculated number of calls, it indicates that there may be a fault in the call logic part of the loop scalar control logic. At this time, an alarm prompt message will be generated to prompt the supervisor to maintain the loop scalar control logic in time; when the calculation output of the fusion unit is garbled, it indicates that there may be a fault in the calculation part of the fusion unit. At this time, an alarm prompt will be generated to prompt the supervisor to maintain the abnormal fusion unit in time. Of course, only examples are given in this embodiment, and the specific content of the detection is not limited.

[0070] It should be noted that in this embodiment, a detection report will also be generated after the detection. The detection report includes detection items, detection results, and maintenance suggestions, and the detection report will be presented to the supervisor in a visual form, so as to quickly find abnormal points according to the detection report and quickly maintain the abnormal points according to the maintenance suggestions, thereby improving the repair efficiency of the loop scalar control logic or the fusion unit.

[0071] In this embodiment, dynamic fusion segments are obtained during the compilation stage. Without relying on shape information, the best fusion slices suitable for the hardware cache size of each dynamic fusion segment are obtained, and based on the best fusion slices, the dynamic fusion segments are converted into static fusion segments to eliminate the calculation nodes related to dynamic shapes, thereby realizing the optimization of dynamic neural networks during the compilation stage. During the execution stage, dynamic scheduling planning can be performed on the static fusion segments, thus converting the internal processing of the original operator level for dynamics into the dynamic processing of the scheduling logic, greatly improving the execution efficiency of dynamic neural networks.

[0072] Embodiment III

[0073] Figure 5 It is a schematic structural diagram of a compilation and execution device for a dynamic neural network provided in Embodiment III of the present invention. As Figure 5 shown, the device includes: a dynamic fusion segment acquisition module 310, an optimal fusion slice acquisition unit 320, a fusion unit acquisition module 330, and a fusion unit call module 340.

[0074] Among them, the dynamic fusion segment acquisition module 310 is used to obtain dynamic fusion segments according to the operator nodes of the dynamic neural network during the compilation stage;

[0075] The optimal fusion slice acquisition unit 320 is used to obtain the chip parameters of the multi-level storage chip carrying the dynamic neural network, and calculate the optimal fusion slices corresponding to each dynamic fusion segment according to the chip parameters;

[0076] The fusion unit acquisition module 330 is used to obtain the static fusion segments matching the corresponding dynamic fusion segments according to the optimal fusion slices, and form fusion units with each dynamic fusion segment and the matching static fusion segments. Among them, the fusion units perform tensor transfer in the shared storage layer of the multi-level storage chip;

[0077] The fusion unit call module 340 is used to obtain the loop scalar control logic of the dynamic neural network during the execution stage, and dynamically call the fusion units based on the loop scalar control logic.

[0078] Optionally, the dynamic fusion segment acquisition module includes: a fusion operator node screening unit, which is used to screen out fusion operator nodes from the operator nodes of the dynamic neural network during the compilation stage;

[0079] A dynamic fusion segment acquisition unit, which is used to obtain dynamic fusion segments according to the fusion operator nodes, where at least two fusion operator nodes with dependency relationships are included in the dynamic fusion segments.

[0080] Optionally, a fused operator node screening unit is configured to obtain the weight transfer types of operator nodes in a dynamic neural network, where the weight transfer types include sparse weight transfer type and dense weight transfer type;

[0081] Obtain the dynamic dimensions and reduction calculation dimensions of operator nodes in the dynamic neural network;

[0082] Use the operator nodes in the dynamic neural network with a sparse weight transfer type and non - overlapping dynamic dimensions and reduction calculation dimensions as fused operator nodes.

[0083] Optionally, a fused operator node screening unit is configured to obtain the weight sizes and chip bandwidths of operator nodes in the dynamic neural network, and obtain the weight transfer overhead according to the ratio of the weight size to the chip bandwidth;

[0084] Obtain the data transfer overhead and data calculation overhead of operator nodes in the dynamic neural network, and obtain an overhead threshold according to the sum of the data transfer overhead and the data calculation overhead;

[0085] Determine whether the weight transfer overhead of the operator node is less than the overhead threshold. If so, determine that the weight transfer type of the operator node is the sparse weight transfer type; otherwise, determine that the weight transfer type of the operator node is the dense weight transfer type.

[0086] Optionally, a dynamic fusion segment obtaining unit is configured to obtain the input - output ratios of each fused operator node;

[0087] Search for fusion of the fused operator nodes with input - output ratios within a first numerical range and having a dependency relationship using a greedy algorithm to obtain at least one dynamic fusion segment.

[0088] Optionally, the chip parameters include the chip cache size. A best - fusion slice obtaining unit is configured to set an initial slice value for the dynamic dimensions of each dynamic fusion segment, and calculate the memory cache size occupied by the dynamic fusion segment under the initial slice value;

[0089] Determine whether the memory cache size is greater than the chip cache size. If so, decrease the initial slice value to obtain an updated slice value; otherwise, increase the initial slice value to obtain an updated slice value;

[0090] Calculate the difference between the updated slice value and the chip cache size. When the difference is within a second numerical range, use the updated slice value as the best - fusion slice.

[0091] Optionally, the device further includes a best - fusion slice alignment module configured to align the dimension sizes of the best - fusion slice to be even;

[0092] Calculate the data volume size corresponding to the best fusion slice after even alignment, and align the data volume size to an integer multiple with reference to the chip transfer bandwidth.

[0093] Optionally, the fusion unit acquisition module includes a static fusion segment acquisition unit, which is used to fill each aligned best fusion slice into the dynamic dimension of the corresponding dynamic fusion segment to obtain a static fusion segment that matches the dynamic fusion segment.

[0094] Perform constant folding and algebraic simplification on each static fusion segment to obtain the static fusion segment optimized for the ordinary layer.

[0095] Perform static operator fusion and redundant operator node elimination on the static fusion segment optimized for the ordinary layer to obtain the static fusion segment optimized for the cache layer.

[0096] Optionally, the fusion unit invocation module is used to obtain the best fusion slice corresponding to the currently to-be-executed fusion unit and the actual input size of the dynamic neural network.

[0097] Input the actual input size and the best fusion slice into the loop scalar control logic to obtain the invocation method for the fusion unit, and dynamically invoke the currently to-be-executed fusion unit based on the invocation method.

[0098] Optionally, the fusion unit invocation module is further used to, when the actual input size is less than the best fusion slice, determine that the invocation method of the loop scalar control logic for the fusion unit is to only invoke the dynamic fusion segment in the fusion unit once.

[0099] When the actual input size is not less than the best fusion slice and is divisible, obtain the integer value of the actual input size divided by the best fusion slice, and determine that the invocation method of the loop scalar control logic for the fusion unit is to loop and invoke the static fusion segments in the fusion units with the integer value.

[0100] When the actual input size is greater than the best fusion slice and is not divisible, obtain the integer value and the remainder value of the actual input size divided by the best fusion slice, and determine that the invocation method of the loop scalar control logic for the fusion unit is to loop and invoke the static fusion segments in the fusion units with the integer value, and to invoke the dynamic fusion segment in the fusion unit with the remainder value.

[0101] The dynamic neural network compilation and execution device provided by the embodiments of the present invention can execute the dynamic neural network compilation and execution method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0102] Embodiment 4

[0103] Figure 6 It is a schematic structural diagram of a computer device provided for Embodiment 4 of the present invention, asFigure 6 As shown, the computer device includes a processor 610, a memory 620, an input device 630, and an output device 640; the number of processors 610 in the computer device can be one or more, Figure 6 and one processor 610 is taken as an example herein; the processor 610, the memory 620, the input device 630, and the output device 640 in the computer device can be connected through a bus or other means, Figure 6 and connection through a bus is taken as an example herein.

[0104] The memory 620, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the compilation optimization method of the computational graph in the embodiments of the present invention. The processor 610 executes various functional applications and data processing of the computer device by running the software programs, instructions, and modules stored in the memory 620, that is, implements the above-mentioned compilation and execution method of the dynamic neural network.

[0105] The memory 620 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory 620 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 620 may further include a memory remotely set relative to the processor 610, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0106] The input device 630 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device. The output device 640 may include a display device such as a display screen.

[0107] Embodiment Five

[0108] Embodiment Five of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a compilation and execution method of a dynamic neural network when executed by a computer processor, including: obtaining dynamic fusion segments according to each operator node of the dynamic neural network during the compilation stage;

[0109] obtaining chip parameters of a multi-level storage chip carrying the dynamic neural network, and calculating the optimal fusion slices corresponding to each dynamic fusion segment according to the chip parameters;

[0110] Obtain the static fusion segment corresponding to the dynamic fusion segment matching the best fusion slice, and form a fusion unit with each dynamic fusion segment and the matching static fusion segment, where the fusion unit performs tensor transfer in the shared storage layer of the multi-level storage chip;

[0111] Obtain the loop scalar control logic of the dynamic neural network during the execution phase, and dynamically call the fusion unit based on the loop scalar control logic.

[0112] Of course, the storage medium containing computer-executable instructions provided by the embodiments of the present invention, the computer-executable instructions are not limited to the method operations as above, and can also execute the relevant operations in the compilation and execution method of the dynamic neural network provided by any embodiment of the present invention.

[0113] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disc of a computer, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.

[0114] It should be noted that in the embodiments of the above compilation optimization device of the loop calculation graph, the included units and modules are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.

[0115] Note that the above is only the preferred embodiment of the present invention and the applied technical principle. Those skilled in the art will understand that the present invention is not limited to the specific embodiments here, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A method for compiling and executing a dynamic neural network, characterized in that: include: In the compilation phase, dynamic fusion fragments are obtained according to each operator node of the dynamic neural network; Obtaining chip parameters of a multi-level storage chip carrying the dynamic neural network, and calculating the best fusion slice corresponding to each of the dynamic fusion segments according to the chip parameters; According to the optimal fusion slice, a static fusion segment matching the corresponding dynamic fusion segment is obtained, and each of the dynamic fusion segments and the matching static fusion segment constitutes a fusion unit, wherein the fusion unit performs tensor transport in the shared storage layer of the multi-level storage chip; The loop scalar control logic of the dynamic neural network is obtained during the execution phase, and the fusion unit is dynamically called based on the loop scalar control logic.

2. The method according to claim 1, characterized in that The step of obtaining the dynamic fusion fragment according to each operator node of the dynamic neural network in the compilation stage includes: In the compilation phase, a fusion operator node is selected from each operator node of the dynamic neural network; The dynamic fusion segment is obtained according to the fusion operator node, wherein the dynamic fusion segment includes at least two of the fusion operator nodes having a dependent relationship.

3. The method according to claim 2, characterized in that The step of selecting a fusion operator node from each operator node of the dynamic neural network during the compilation phase includes: Obtaining a weight transfer type of each operator node in the dynamic neural network, wherein the weight transfer type includes a sparse weight transfer type and an intensive weight transfer type; Obtaining the dynamic dimension and the reduced calculation dimension of each operator node in the dynamic neural network; The weight transfer type in the dynamic neural network is a weight transfer sparse type, and the operator node whose dynamic dimension does not overlap with the reduced calculation dimension is used as the fusion operator node.

4. The method according to claim 3, characterized in that The obtaining of the weight transfer type of each operator node in the dynamic neural network includes: Obtaining the weight size and chip bandwidth of each operator node in the dynamic neural network, and obtaining the weight transport overhead according to the ratio of the weight size to the chip bandwidth; Obtaining data handling overhead and data calculation overhead of each operator node in the dynamic neural network, and obtaining an overhead threshold according to the sum of the data handling overhead and the data calculation overhead; Determine whether the weight transfer overhead of the operator node is less than the overhead threshold; if so, determine that the weight transfer type of the operator node is the weight transfer sparse type; otherwise, determine that the weight transfer type of the operator node is the weight transfer intensive type.

5. The method according to claim 2, characterized in that: The acquiring the dynamic fusion segment according to the fusion operator node includes: Obtaining the input-output ratio of each fusion operator node; The fusion operator nodes whose input-output ratios are within a first numerical range and have a dependency relationship are searched and fused using a greedy algorithm to obtain at least one of the dynamic fusion fragments.

6. The method according to claim 1, characterized in that The chip parameters include a chip cache size, and calculating the best fusion slice corresponding to each of the dynamic fusion segments according to the chip parameters includes: Setting an initial slice value for the dynamic dimension of each of the dynamic fusion segments, and calculating the memory cache size occupied by the dynamic fusion segment under the initial slice value; Determine whether the memory cache size is greater than the chip cache size, if so, reduce the initial slice value to obtain an updated slice value, otherwise, increase the initial slice value to obtain an updated slice value; The difference between the updated slice value and the chip cache size is calculated, and when the difference is within a second numerical range, the updated slice value is used as the optimal fused slice.

7. The method according to claim 1, characterized in that After calculating the best fusion slice corresponding to each of the dynamic fusion segments according to the chip parameters, the method further includes: Aligning the dimensions of the best fused slices to an even number; The data size corresponding to the optimal fused slice after even-number alignment is calculated, and the data size is aligned as an integer multiple with reference to the chip transport bandwidth.

8. The method according to claim 7, characterized in that The step of obtaining a static fusion segment that matches a corresponding dynamic fusion segment according to the optimal fusion slice includes: Filling the aligned best fused slices into the dynamic dimension of the corresponding dynamic fused segment to obtain the static fused segment matching the dynamic fused segment; Performing constant folding and algebraic simplification on each of the static fusion fragments to obtain the static fusion fragment after ordinary layer optimization; Static operator fusion and redundant operator node elimination are performed on the static fusion fragment after the ordinary layer is optimized to obtain the static fusion fragment after the cache layer is optimized.

9. The method according to claim 1, characterized in that: The dynamically calling the fusion unit based on the loop scalar control logic includes: Obtaining the best fusion slice corresponding to the fusion unit to be executed currently and the actual input size of the dynamic neural network; The actual input size and the optimal fused slice are input into the loop scalar control logic to obtain a calling mode for the fusion unit, and the fusion unit to be executed currently is dynamically called based on the calling mode.

10. The method according to claim 9, characterized in that The step of inputting the actual input size and the optimal fused slice into the loop scalar control logic to obtain a calling mode for the fusion unit includes: When the actual input size is smaller than the optimal fused slice, determining that the calling mode of the loop scalar control logic for the fusion unit is to call the dynamic fusion segment in the fusion unit only once; When the actual input size is not less than the best fused slice and is divisible by an integer, obtaining an integer value of the actual input size divided by the best fused slice, and determining that the calling mode of the loop scalar control logic for the fusion unit is to loop call the static fusion segment in the fusion unit of the integer value; When the actual input size is larger than the optimal fusion slice and is not divisible, obtain the integer value and remainder value of the actual input size divided by the optimal fusion slice, and determine that the calling mode of the loop scalar control logic for the fusion unit is to cyclically call the static fusion fragment in the fusion unit of the integer value and call the dynamic fusion fragment in the fusion unit of the remainder value.

11. A compiling and executing device for a dynamic neural network, characterized in that: include: A dynamic fusion fragment acquisition module is used to acquire dynamic fusion fragments according to each operator node of the dynamic neural network during the compilation stage; An optimal fusion slice acquisition unit, used to acquire chip parameters of a multi-level storage chip carrying the dynamic neural network, and calculate the optimal fusion slice corresponding to each of the dynamic fusion segments according to the chip parameters; A fusion unit acquisition module, used for acquiring a static fusion segment matching the corresponding dynamic fusion segment according to the optimal fusion slice, and forming a fusion unit with each of the dynamic fusion segments and the matching static fusion segment, wherein the fusion unit performs tensor transport in the shared storage layer of the multi-level storage chip; The fusion unit calling module is used to obtain the loop scalar control logic of the dynamic neural network during the execution phase, and dynamically call the fusion unit based on the loop scalar control logic.

12. A computer device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 10 is implemented.

13. A computer executable instruction storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Compiling method, compiler, neural network accelerator, chip and electronic equipment

    CN117492766A

  • Operator fusion method and device, electronic equipment and storage medium

    CN118394475A