A method, device, electronic device and medium for on-chip resource scheduling of neural network
By segmenting neural networks based on memory trends and optimizing cache batch sizes, the method addresses inefficiencies in traditional resource scheduling, enhancing cache utilization and reducing redundant data movement for improved performance.
Patent Information
- Application Number
- CN202510353255.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-25
AI Technical Summary
In the on-chip resource scheduling method of traditional neural networks, the cache utilization rate is low and the weight is repeated and the number of times is often transferred, resulting in wasted cache space and chip computing power.
By drawing the change trend chart of the neural network operator execution order and tensor memory usage, each network fragment is obtained in segments, and the target cache batch size of each fragment is calculated, the data processing process is optimized, and the repeated handling and system overhead are reduced.
It improves the utilization rate of on-chip cache, reduces the duplicate handling of weighted data and system fixed overhead, and improves the operation efficiency of neural networks.
Smart Images

Figure CN119862153B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of on-chip scheduling of neural networks, and in particular, to a method, device, electronic device and medium for on-chip resource scheduling of neural networks. Background Art
[0002] Although there are many types of neural networks, their data size scale or change period usually follows a certain change rule. For example, the data size scale changes from large to small or from small to large, and the change period changes reciprocally from large to small. Among them, neural networks with a "pyramid" shape of data scale are particularly common. For example, in a typical resnet network (residual network), the data has the largest size at the beginning of the network, and as the number of network layers deepens, the size gradually decreases, and the overall data scale is like a "pyramid".
[0003] The logic of large-batch resource scheduling for traditional neural networks is as Figure 1 shown. When optimizing the on-chip cache of such neural networks, they usually make a single determination of the batch based on the position where the network data scale is the largest, that is, the peak memory of the entire network. For example, in an AI (Artificial Intelligence) chip with a determined cache for a resnet network with a batch size of 32, the peak memory usage of the network can be cached in batches of size 4. From top to bottom of the entire network, the batch size for cache optimization will be 4, and it will loop 8 times. Although such cache optimization methods can reduce the interaction between the entire network and the external cache, there are still many problems: 1) Since the memory peak is used to determine the slice size, the number of slices of the entire network is too large, and too many slice times will cause a large amount of weight data to be repeatedly moved, and the fixed overhead of the system will increase exponentially, which instead erodes the benefits of the on-chip cache to a certain extent. 2) For the slice size determined by the memory peak, in the initial stage of the network, due to the "pyramid" memory trend, the memory is at the peak at this time, so the cache utilization rate is also very high. However, as the network deepens layer by layer, when the size of the tensor gradually decreases to half, one-fourth, or even one-eighth of the peak, this results in the chip cache utilization rate in the latter part of the network also being reduced to half, one-fourth, or one-eighth of the peak, causing serious cache waste. 3) In addition to cache waste, in the latter half of the network, due to the too small tensor, the calculation becomes a bottleneck for weight transfer, and the calculation is in an idle state, resulting in the idle and waiting of the chip computing resources. Summary of the Invention
[0004] The present invention provides a method, device, electronic device and medium for on-chip resource scheduling of neural networks to solve the problems of low on-chip cache utilization rate, many weight repeated transfer times, and waste of cache space and chip computing power existing in the batch scheduling method that directly determines the entire network using the memory peak.
[0005] According to one aspect of the present invention, there is provided a method for on-chip resource scheduling of a neural network, including:
[0006] Based on the execution order of operators and the memory usage of operator tensors of the current neural network, draw a trend chart and segment the trend chart to obtain each network segment;
[0007] Calculate the target cache batch size of each network segment;
[0008] Perform on-chip data stitching on the last loop output data of the previous network segment and the remaining input of the subsequent network segment in the current adjacent network segments to obtain the target input data for the operator corresponding to the subsequent network segment, so as to perform data processing on the target input data based on the operator corresponding to the subsequent network segment.
[0009] According to another aspect of the present invention, there is provided an on-chip resource scheduling device for a neural network, including:
[0010] A trend chart drawing and segmentation module, configured to draw a trend chart based on the execution order of operators and the memory usage of operator tensors of the current neural network, and segment the trend chart to obtain each network segment;
[0011] A batch size calculation module, configured to calculate the target cache batch size of each network segment;
[0012] An on-chip data stitching module, configured to perform on-chip data stitching on the last loop output data of the previous network segment and the remaining input of the subsequent network segment in the current adjacent network segments to obtain the target input data for the operator corresponding to the subsequent network segment, so as to perform data processing on the target input data based on the operator corresponding to the subsequent network segment.
[0013] According to another aspect of the present invention, there is provided an electronic device, where the electronic device includes:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for on-chip resource scheduling of a neural network according to any embodiment of the present invention.
[0017] According to another aspect of the present invention, there is provided a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the method for on-chip resource scheduling of a neural network according to any embodiment of the present invention when executed.
[0018] In the technical solution of the embodiment of the present invention, by drawing a trend graph based on the execution order of operators and the memory usage of operator tensors in the current neural network, and segmenting the trend graph to obtain each network segment, the target cache batch size of each network segment is calculated. Then, the last loop output data of the previous network segment and the remaining input of the subsequent network segment in the current adjacent network segments are spliced within the chip to obtain the target input data for the operator corresponding to the subsequent network segment, so as to perform data processing on the target input data based on the operator corresponding to the subsequent network segment. This solution calculates the cache batch size for each network segment, breaks the drawbacks of using the memory peak value to determine the batch size of the entire network, and by dynamically changing the size of the batch in the cache, significantly reduces the number of loop slices in the middle and end segments of the network, so as to optimize the utilization rate of neural network memory resources. And through the optimization of the operator execution order, it reduces the repeated transfer of weight data and the fixed overhead of the system, solves the problems of low utilization rate of on-chip cache, many repeated transfers of weights, and waste of cache space and chip computing power existing in the batch scheduling method of directly determining the entire network using the memory peak value, can make full use of the on-chip cache space, reduce the repeated transfer of weight data and the fixed overhead of the system, and greatly improve the optimized operation efficiency of the on-chip cache.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Brief Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 It is a logical schematic diagram of large-batch resource scheduling for a traditional neural network provided by the present invention;
[0022] Figure 2 It is a flowchart of a method for on-chip resource scheduling of a neural network provided in Embodiment 1 of the present invention;
[0023] Figure 3 It is a flowchart of a method for on-chip resource scheduling of a neural network provided in Embodiment 2 of the present invention;
[0024] Figure 4 It is a schematic diagram for dividing network segments based on a trend graph provided in Embodiment 2 of the present invention;
[0025] Figure 5 Schematic diagram of an on-chip resource scheduling logic of a neural network provided in the second embodiment of the present invention;
[0026] Figure 6 Real-time operation graph of an operator provided in the second embodiment of the present invention;
[0027] Figure 7 Operator operation graph during traditional batch scheduling of an existing entire neural network provided in the second embodiment of the present invention;
[0028] Figure 8 Schematic structural diagram of an on-chip resource scheduling device of a neural network provided in the third embodiment of the present invention;
[0029] Figure 9 Schematic structural diagram of an electronic device that can be used to implement the embodiments of the present invention. Detailed implementation manners
[0030] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0031] It should be noted that the terms "previous", "next", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0032] Embodiment 1
[0033] Figure 2The flowchart of a method for on-chip resource scheduling of a neural network provided by Embodiment 1 of the present invention. This embodiment is applicable to the situation of optimizing on-chip resource scheduling of a neural network. This method can be executed by an on-chip resource scheduling device of the neural network. The on-chip resource scheduling device of the neural network can be implemented in the form of hardware and / or software, and the on-chip resource scheduling device of the neural network can be configured in an electronic device. The electronic device can include, but is not limited to, a computer or a server, etc. As Figure 2 shown, the method includes:
[0034] Step 110: Based on the execution order of operators and the memory usage of operator tensors in the current neural network, draw a trend chart, and segment the trend chart to obtain each network segment.
[0035] Among them, the execution order of operators can be the execution order of operators in the neural network. The memory usage of operator tensors can be used to describe the memory occupation of operator tensors. The trend chart can be used to describe the change trend of the tensor memory occupation of the current neural network with the execution order of operators. The network segment can be the segmentation result obtained by segmenting the current neural network based on the trend chart.
[0036] In the embodiment of the present invention, the memory usage of the entire network can be analyzed during the compilation stage of the current neural network to obtain the execution order of operators and the memory usage of operator tensors in the current neural network. Then, based on the execution order of operators and the memory usage of operator tensors in the current neural network, a trend chart is drawn, and several operators with relatively stable memory usage (the measurement standard for relatively stable memory usage can be set by oneself) are divided into one segment to achieve the segmentation of the entire network of the current neural network and obtain each network segment.
[0037] Step 120: Calculate the target cache batch size of each network segment.
[0038] Among them, the target cache batch size can be the batch size cached by the network segment.
[0039] In the embodiment of the present invention, for each network segment, the target cache batch size can be calculated based on the corresponding memory peak value and the on-chip cache capacity of the chip on which the current neural network is deployed, so as to realize dynamically changing the size of the batch in the cache according to the tensor memory size in the network. Especially for neural networks such as "inverted pyramid", as the network deepens, the tensor size of a single batch gradually becomes smaller, and the cache batch size gradually increases, which can make full use of the on-chip cache space and greatly reduce the number of loop slices in the middle and tail segments of the network.
[0040] Step 130: Concatenate the last loop output data of the previous network segment in the current adjacent network segments with the remaining input of the subsequent network segment inside the chip to obtain the target input data for the operator corresponding to the subsequent network segment, so as to perform data processing on the target input data based on the operator corresponding to the subsequent network segment.
[0041] Among them, the current adjacent network segments can be composed of two currently adjacent network segments. The network segment with a higher operator execution order in the current adjacent network segments is the previous network segment, and the network segment with a lower operator execution order is the subsequent network segment. The in-chip data concatenation can be a process of combining the last loop output data of the previous network segment with the remaining input of the subsequent network segment. The target input data can be the data completely input to the subsequent network segment in the current adjacent network segments.
[0042] In the embodiment of the present invention, the currently adjacent network segments for which the operator execution order is scheduled can be determined, and then the previous network segment and the subsequent network segment in the currently adjacent network segments can be determined, so that the last loop output data of the previous network segment is concatenated with the remaining input of the subsequent network segment inside the chip (without caching the last loop output data of the previous network segment outside the chip), and the data concatenation result is used as the target input data for the operator corresponding to the subsequent network segment. Further, the operator corresponding to the subsequent network segment performs data processing on the target input data, reducing the repeated transfer of weight data and the fixed overhead of the system, and greatly improving the optimized operation efficiency of the in-chip cache.
[0043] The technical solution of the embodiment of the present invention is to draw a trend graph based on the operator execution order and the operator tensor memory usage of the current neural network, segment the trend graph to obtain each network segment, calculate the target cache batch size of each network segment, and then concatenate the last loop output data of the previous network segment in the current adjacent network segments with the remaining input of the subsequent network segment inside the chip to obtain the target input data for the operator corresponding to the subsequent network segment, so as to perform data processing on the target input data based on the operator corresponding to the subsequent network segment. This solution calculates the cache batch size for each network segment, breaking the drawbacks of using the memory peak value to determine the batch size of the entire network. By dynamically changing the size of the batch in the cache, the number of loop slices in the middle and end segments of the network is greatly reduced to optimize the utilization rate of the neural network memory resources. Through the optimization of the operator execution order, the repeated transfer of weight data and the fixed overhead of the system are reduced, solving the problems of low in-chip cache utilization rate, large number of repeated transfers of weights, and waste of cache space and chip computing power existing in the batch scheduling method of directly determining the entire network using the memory peak value. It can make full use of the in-chip cache space, reduce the repeated transfer of weight data and the fixed overhead of the system, and greatly improve the optimized operation efficiency of the in-chip cache.
[0044] Embodiment 2
[0045] Figure 3 The figure is a flowchart of a method for on-chip resource scheduling of a neural network provided in Embodiment 2 of the present invention. This embodiment is a specific implementation based on the above embodiment, and provides a specific and optional implementation manner for calculating the target cache batch size of each network segment. As Figure 3 shown, the method includes:
[0046] Step 210: Based on the execution order of operators and the memory usage of operator tensors in the current neural network, draw a trend chart, and segment the trend chart to obtain each network segment.
[0047] In an optional embodiment of the present invention, segmenting the trend chart to obtain each network segment may include: determining an operator to be compared for the current tensor memory usage, and determining the similarity of the operator tensor memory usage of the operator to be compared for the current tensor memory usage; segmenting the trend chart according to the similarity of the operator tensor memory usage and the lower limit threshold of the number of operators in a single network segment to obtain the current network segment.
[0048] Wherein, the operator to be compared for the current tensor memory usage may be the operator that needs to compare the tensor memory usage currently. The similarity of the operator tensor memory usage may be used to describe the similarity degree of the tensor memory usage between different operators. The lower limit threshold of the number of operators may be the minimum value of the number of operators in a single network segment set in advance.
[0049] In the embodiment of the present invention, the operator that needs to compare the tensor memory usage currently may be determined first, that is, the operator to be compared for the current tensor memory usage is determined first, and then based on the evaluation rule of the similarity of the operator tensor memory usage set in advance, the similarity of the operator tensor memory usage of the operator to be compared for the current tensor memory usage is calculated. Then, the trend chart is segmented by using the similarity of the operator tensor memory usage and the lower limit threshold of the number of operators in a single network segment, so that the memory usage of the operators in the current network segment divided based on the trend chart is relatively stable.
[0050] In an optional embodiment of the present invention, determining the similarity of the operator tensor memory usage of the operator to be compared for the current tensor memory usage may include: based on the operator tensor memory usage of the operator to be compared for the current tensor memory usage, determining the first memory usage similarity between the starting operator and the ending operator in the operator to be compared for the current tensor memory usage, or the second memory usage similarity between non-bifurcation point operators; using the first memory usage similarity, or the second memory usage similarity as the similarity of the operator tensor memory usage.
[0051] Among them, the start-end operator can be the operator with the earliest execution order among the operators to be compared for the current tensor memory usage. The end-end operator can be the operator with the latest execution order among the operators to be compared for the current tensor memory usage. The first memory usage similarity can be the similarity of the memory usage between the start-end operator and the end-end operator among the operators to be compared for the current tensor memory usage. The non-two-endpoint operator can be an operator that is not at the endpoint position at the same time among the operators to be compared for the current tensor memory usage. The second memory usage similarity can be the similarity of the memory usage of the non-two-endpoint operator.
[0052] In an embodiment of the present invention, based on the change trend graph, the operator tensor memory usage of the start-end operator and the end-end operator among the operators to be compared for the current tensor memory usage can be determined. Furthermore, the ratio of the operator tensor memory usage of the two endpoint operators can be used as the first memory usage similarity between the two endpoint operators. It is also possible to calculate the ratio of the operator tensor memory usage of any two non-two-endpoint operators among the operators to be compared for the current tensor memory usage to obtain the second memory usage similarity between the non-two-endpoint operators. Further, the first memory usage similarity or the second memory usage similarity is used as the operator tensor memory usage similarity.
[0053] In an alternative embodiment of the present invention, segmenting the change trend graph according to the operator tensor memory usage similarity and the lower limit threshold of the number of operators in a single network segment to obtain the current network segment may include: when the number of operators to be compared for the current tensor memory usage is less than the lower limit threshold of the number of operators in a single network segment, or the operator tensor memory usage similarity meets the similarity discrimination condition, then update the operators to be compared for the current tensor memory usage based on the operator with the next execution order after the end-end operator among the operators to be compared for the current tensor memory usage; and return to execute the operation of determining the operator tensor memory usage similarity based on the operators to be compared for the current tensor memory usage until the number of operators to be compared for the current tensor memory usage is greater than the lower limit threshold of the number of operators in a single network segment and the operator tensor memory usage similarity does not meet the similarity discrimination condition, then segment the change trend graph using the operators to be compared for the current tensor memory usage to obtain the current network segment.
[0054] Among them, the similarity discrimination condition can be a condition for judging the memory usage stability of the operators to be compared for the current tensor memory usage. Optionally, the similarity discrimination condition may include that the first memory usage similarity is greater than 70% (which can be set by oneself), or the ratio of the maximum operator tensor memory usage to the minimum operator tensor memory usage among the non-two-endpoint operators is greater than 50% (which can be set by oneself).
[0055] In an embodiment of the present invention, if the number of operators to be compared for the current tensor memory usage is less than the lower threshold of the number of operators within a single network segment, or the similarity of the operator tensor memory usage meets the similarity discrimination condition, it indicates that the operators to be compared for the current tensor memory usage can be further expanded. Thus, the next operator in the execution order of the end operator among the operators to be compared for the current tensor memory usage is determined, and the newly determined next operator in the execution order is added to the operators to be compared for the current tensor memory usage. Then, the operation of determining the similarity of the operator tensor memory usage based on the operators to be compared for the current tensor memory usage is returned and executed until the number of operators to be compared for the current tensor memory usage is greater than the lower threshold of the number of operators within a single network segment, and the similarity of the operator tensor memory usage does not meet the similarity discrimination condition. At this time, it can be considered that the operators with relatively stable memory usage have been determined. Therefore, the operator that causes the number of operators to be compared for the current tensor memory usage to be greater than the lower threshold of the number of operators within a single network segment and the similarity of the operator tensor memory usage does not meet the similarity discrimination condition (i.e., the operator that was last added to the operators to be compared for the current tensor memory usage) is used as the demarcation point, and the network segment corresponding to the operators to be compared for the current tensor memory usage before adding this operator is used as the current network segment, and this operator is used as the initial operator of the next network segment, and the next network segment is divided according to the division rules of the network segment.
[0056] Step 220: Obtain the peak memory within a single batch segment of each network segment.
[0057] Among them, the peak memory within a single batch segment can be the peak memory of the network segment when the batch size is 1.
[0058] In an embodiment of the present invention, if the obtained peak memory of the network segment is not the peak memory of a single batch, it is necessary to calculate the peak memory of the network segment when the number of batches is equal to 1, so as to obtain the peak memory within a single batch segment of each network segment.
[0059] Step 230: Calculate the target cache batch size of each network segment based on the on-chip cache capacity of the target chip and the peak memory within a single batch segment of each network segment.
[0060] Among them, the on-chip cache capacity can be used to describe the size of the on-chip cache of the chip.
[0061] In an embodiment of the present invention, the ratio of the on-chip cache capacity of the target chip to the peak memory within a single batch segment of each network segment can be calculated, and the calculated ratio is used as the target cache batch size of the corresponding network segment.
[0062] In an alternative embodiment of the present invention, after calculating the target cache batch size of each network segment, the following steps may further be included: determining, according to the target cache batch size of each network segment, the network segments to be fused that meet the network segment fusion condition; fusing the network segments to be fused to obtain a fused processing network segment, and updating the network segments according to the fused processing network segment.
[0063] Among them, the network segment fusion condition may be a condition for determining whether two adjacent network segments can be used as one network segment for cache batch adjustment. The network segments to be fused may be two adjacent network segments that meet the network segment fusion condition. The fused processing network segment may be a network segment formed by the network segments to be fused.
[0064] In an embodiment of the present invention, the target cache batch size of each network segment may be obtained first, and then the difference between the target cache batch sizes of two adjacent network segments may be calculated. If the difference between the target cache batch sizes of two adjacent network segments is not greater than a preset difference, the two adjacent network segments with the difference between the target cache batch sizes not greater than the preset difference are used as the network segments to be fused that meet the network segment fusion condition, so as to fuse the network segments to be fused to obtain a fused processing network segment, that is, the network segments to be fused are used as a complete network segment, and the network segments that have been determined are updated based on the fused processing network segment.
[0065] Step 240: Perform on-chip data splicing on the last loop output data of the previous network segment and the remaining input of the subsequent network segment in the current adjacent network segments to obtain the target input data for the operator corresponding to the subsequent network segment, so as to perform data processing on the target input data based on the operator corresponding to the subsequent network segment.
[0066] In an alternative embodiment of the present invention, after obtaining the target input data for the operator corresponding to the subsequent network segment, the following steps may further be included: after the operator corresponding to the subsequent network segment completes data processing on the target input data, updating the current adjacent network segments, and returning to execute the operation of performing on-chip data splicing on the loop output data of the previous network segment and the remaining input of the subsequent network segment in the current adjacent network segments to obtain the target input data for the operator corresponding to the subsequent network segment until the target input data for the operator corresponding to the full-scale network segment is determined.
[0067] In an embodiment of the present invention, after data processing is performed on target input data based on an operator corresponding to a subsequent network segment, the subsequent network segment and the next adjacent network segment of the subsequent network segment are used as new current adjacent network segments, and then the operation of performing in-chip data splicing on the loop output data of the previous network segment in the current adjacent network segment and the remaining input of the subsequent network segment is returned to obtain the target input data of the operator corresponding to the subsequent network segment, until the target input data of the operators corresponding to all network segments is determined, so as to implement data processing of all operators on the corresponding target input data.
[0068] In a specific example, first analyze the tensor memory when the number of batches (batch) of the current neural network is equal to 1, and draw a trend graph of the change in tensor memory occupancy of the current neural network with the operator execution order according to the operator execution order and the tensor memory usage of the operator. According to the trend graph, divide the network into segments. Usually, several operators with relatively stable memory usage are divided into one segment. The division of network segments satisfies the following conditions: 1. Inside the same network segment, the memory usage of each operator should be close, or the operators at both ends of the start and end of the network segment have relatively close memory usage, so as to ensure a high cache utilization rate for memory caching within the entire network segment, because within the network segment, the size of the slice is determined by the memory peak operator within the network segment. 2. A certain number of operators (at least more than 3) should be included in the network segment to ensure the overall benefit of in-chip caching and balance the amplitude of operator memory changes and the number of operators. If the memory usage fluctuates greatly and changes rapidly, dividing the network segments only according to the memory change trend will result in too few operators within each fusion segment, making it impossible to perform effective in-chip caching. The balance should be made according to the tensor memory change trend and the number of operators, and each operator should be divided into the corresponding cache segment. A network segment is equivalent to a fusion segment, and the computing operators within a fusion segment are cached in the chip as a whole, and the computing operators within the fusion segment are continuously executed with data dependencies.
[0069] Figure 4 FIG. is a schematic diagram of network segment division based on a trend graph provided in Embodiment II of the present invention. As Figure 4 shown, the horizontal axis is the operator execution order (op) of the operators in the network, and the vertical axis is the operator tensor memory usage of each operator. The memory usage gradually decreases as the operators are executed. Through this solution, the entire network can be divided into 4 segments at the place where the memory decreases significantly, and separate cache analysis is performed within each segment.
[0070] Figure 5 FIG. is a schematic diagram of the in-chip resource scheduling logic of a neural network provided in Embodiment II of the present invention. As Figure 5For a given network segment (abbreviated as Stage in Figure 5 ), with a fixed-chip-specification cache, the optimal batch size for the cache, i.e., the target cache batch size, can be determined. For different network segments, the optimal batch size for each network segment is also different because their memory usage is usually different. Based on the optimization at the scheduling level, that is, when scheduling two adjacent network segments, the output data of the last loop of the previous network segment is not moved out of the on-chip cache of the chip. The output data of the last loop of the previous network segment is used as part of the input of the subsequent network segment, and the remaining input of the subsequent network segment is continuously moved into the on-chip cache, and the splicing operation of the input is directly completed on the on-chip cache.
[0071] The actual operation graph of the operators in this solution can be seen in Figure 6 . Figure 6 The white dots in Figure 7 represent the tensors of the batches, which are used to reflect the propagation behavior of data in the network. The yellow dots represent the splicing behavior, which is used to merge the previous serial data into a tensor. The dashed arrows represent the operator execution order, not the data dependency relationship. The solid arrows represent both the operator execution order and the data dependency, and the two are unified. By comparing with the operator operation graph during the traditional batch scheduling of the existing entire neural network shown in Figure 7 , it can be intuitively seen that in this solution, the output data of the previous network segment is directly retained in the on-chip cache without being moved out, and is directly cached in the on-chip cache with the input of the subsequent network segment, which can effectively reduce the redundant movement of data. Among them, Figure 7 describes the processing method of the inference stack when the batch size is 16, that is, all batches form a tensor and pass through all layers of the neural network layer by layer, Figure 7 the white dots represent the tensors of the batches, and the solid arrows represent the operator execution order. By optimizing the execution order of Figure 7 , disassembling the batches, assembling them into an appropriate size and then executing them in order (there can be multiple evolved optimization versions), the caching ability of the L2 buf (L2 cache area) layer can be improved, and the locality of the L2 stored data can be fully utilized.
[0072] This solution no longer focuses solely on the memory peak of the entire neural network, but rather on the trend graph of the memory usage of the entire network. It segments the network based on the points where the memory usage jumps in the trend graph. Each segment after segmentation is regarded as a network segment, such that the memory usage within the network segment is similar, while the memory usage between network segments varies significantly. This ensures that the memory usage of each network segment is similar, and it prevents the phenomenon where the cache utilization rate of the remaining operators in the entire network is low due to the high peak values of individual operators in the network. For each network segment after segmentation, the batch size that can be cached is calculated as the target cache batch size for that network segment. Instead of making a single division of the on-chip cache area based on the memory peak of the entire network, the entire network is divided into several network segments with similar memory usage according to the memory usage trend, and each network segment is cached separately. The batch cache size within a network segment is only related to the memory peak within the current network segment and is no longer tied to the memory peak of the entire network, ensuring the on-chip cache utilization rate within each network segment.
[0073] During the operation of the neural network, a new method is added. By associating the data dependencies of two consecutive network segments, the output data of the previous network segment is directly retained in the on-chip memory without being transferred externally, and is directly cached in the on-chip memory with the input of the subsequent network segment. In this way, the data interaction between adjacent segments outside the chip can be significantly reduced, eliminating the redundant transfer step of this result data from on-chip -> off-chip -> on-chip, improving performance. Moreover, this scheduling method for the execution order of operators does not depend on the network type and operator type, and only needs to focus on the memory usage of the operators, having complete generalization performance.
[0074] The technical solution of the embodiment of the present invention is as follows: By drawing a trend graph based on the execution order of operators and the memory usage of operator tensors in the current neural network, and segmenting the trend graph to obtain each network segment, the in-memory peak value within a single batch segment of each network segment can be obtained. Then, based on the on-chip cache capacity of the target chip and the in-memory peak value within a single batch segment of each network segment, the target cache batch size of each network segment is calculated. Further, the last loop output data of the previous network segment and the remaining input of the subsequent network segment in the current adjacent network segments are spliced in the on-chip data to obtain the target input data for the operator corresponding to the subsequent network segment, so as to perform data processing on the target input data based on the operator corresponding to the subsequent network segment. This solution calculates the cache batch size for each network segment, breaking the drawbacks of using the in-memory peak value to determine the batch size of the entire network. By dynamically changing the size of the batches in the cache, the number of loop slices in the middle and end segments of the network is significantly reduced, optimizing the utilization rate of neural network memory resources. And through the optimization of the operator execution order, the repeated transfer of weight data and the fixed overhead of the system are reduced, solving the problems of low on-chip cache utilization rate, large number of repeated weight transfers, and waste of cache space and chip computing power in the batch scheduling method that directly determines the entire network using the in-memory peak value. It can make full use of the on-chip cache space, reduce the repeated transfer of weight data and the fixed overhead of the system, and greatly improve the optimized operation efficiency of the on-chip cache.
[0075] Embodiment III
[0076] Figure 8 It is a schematic structural diagram of a neural network on-chip resource scheduling device provided by Embodiment III of the present invention. As Figure 8 shown, the device includes:
[0077] A trend graph drawing and segmentation module 310, configured to draw a trend graph based on the execution order of operators and the memory usage of operator tensors in the current neural network, and segment the trend graph to obtain each network segment.
[0078] A batch size calculation module 320, configured to calculate the target cache batch size of each network segment.
[0079] An on-chip data splicing module 330, configured to splice the last loop output data of the previous network segment and the remaining input of the subsequent network segment in the current adjacent network segments in the on-chip data to obtain the target input data for the operator corresponding to the subsequent network segment, so as to perform data processing on the target input data based on the operator corresponding to the subsequent network segment.
[0080] The technical solution of the embodiment of the present invention is to draw a trend graph based on the execution order of operators and the memory usage of operator tensors in the current neural network, segment the trend graph to obtain each network segment, calculate the target cache batch size of each network segment, and then splice the last loop output data of the previous network segment and the remaining input of the next network segment in the current adjacent network segments in-chip to obtain the target input data for the operator corresponding to the next network segment, so as to process the target input data based on the operator corresponding to the next network segment. This solution calculates the cache batch size for each network segment, breaks the drawbacks of using the memory peak value to determine the batch size of the entire network, and significantly reduces the number of loop slices in the middle and end segments of the network by dynamically changing the size of the batch in the cache, so as to optimize the utilization rate of the memory resources of the neural network. Through the optimization of the operator execution order, the repeated transfer of weight data and the fixed overhead of the system are reduced, and the problems of low in-chip cache utilization rate, many repeated transfers of weights, and waste of cache space and chip computing power existing in the batch scheduling method of directly determining the entire network using the memory peak value are solved. It can make full use of the in-chip cache space, reduce the repeated transfer of weight data and the fixed overhead of the system, and greatly improve the optimized operation efficiency of the in-chip cache.
[0081] Optionally, the trend graph drawing and segmentation module 310 is specifically configured to determine the operator to be compared with the current tensor memory usage and determine the similarity of the operator tensor memory usage of the operator to be compared with the current tensor memory usage; segment the trend graph according to the similarity of the operator tensor memory usage and the lower limit threshold of the number of operators in a single network segment to obtain the current network segment.
[0082] Optionally, the trend graph drawing and segmentation module 310 is specifically configured to determine the first memory usage similarity between the start-end operator and the end-end operator in the operator to be compared with the current tensor memory usage, or the second memory usage similarity between non-double-endpoint operators based on the operator tensor memory usage of the operator to be compared with the current tensor memory usage; use the first memory usage similarity or the second memory usage similarity as the similarity of the operator tensor memory usage.
[0083] Optionally, the trend graph plotting and segmentation module 310 is specifically configured to, when the number of operators to be compared in the current tensor memory usage is less than the lower limit threshold of the number of operators within a single network segment, or the similarity of the operator tensor memory usage satisfies the similarity discrimination condition, update the operator to be compared in the current tensor memory usage based on the operator with the next execution order of the end operator in the operator to be compared in the current tensor memory usage; and return the operation of determining the similarity of the operator tensor memory usage based on the operator tensor memory usage of the operator to be compared in the current tensor memory usage, until the number of operators to be compared in the current tensor memory usage is greater than the lower limit threshold of the number of operators within a single network segment, and the similarity of the operator tensor memory usage does not satisfy the similarity discrimination condition, then segment the change trend graph using the operator to be compared in the current tensor memory usage to obtain the current network segment.
[0084] Optionally, the batch size calculation module 320 is specifically configured to obtain the in-batch memory peak value of each network segment; calculate the target cache batch size of each network segment based on the on-chip cache capacity of the target chip and the in-batch memory peak value of each network segment.
[0085] Optionally, the neural network on-chip resource scheduling device further includes a network segment update module, configured to determine the network segments to be fused that meet the network segment fusion conditions according to the target cache batch size of each network segment; fuse the network segments to be fused to obtain a fused processing network segment, and update the network segments according to the fused processing network segment.
[0086] Optionally, the on-chip data splicing module 330 is specifically configured to, after the operator of the subsequent network segment completes data processing on the target input data, update the current adjacent network segment, and return the operation of splicing the loop output data of the previous network segment and the remaining input of the subsequent network segment in the current adjacent network segment in the on-chip to obtain the target input data of the operator of the subsequent network segment, until the target input data of the operator corresponding to the full network segment is determined.
[0087] The neural network on-chip resource scheduling device provided by the embodiments of the present invention can execute the neural network on-chip resource scheduling method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0088] Embodiment 4
[0089] Figure 9The figure shows a schematic structural diagram of an electronic device that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0090] As Figure 9 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as ROM 12, RAM 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the ROM 12 or the computer program loaded from the storage unit 18 into the RAM 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. The I / O interface 15 is also connected to the bus 14. Among them, the ROM 12 is a read-only memory, the RAM 13 is a random access memory, and the I / O interface 15 is an input / output interface.
[0091] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0092] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the on-chip resource scheduling method for neural networks.
[0093] In some embodiments, the method for on-chip resource scheduling of a neural network can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the above-described method for on-chip resource scheduling of a neural network can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the method for on-chip resource scheduling of a neural network in any other suitable manner (e.g., by means of firmware).
[0094] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0095] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0096] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0097] For providing interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used for providing interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0098] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0099] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of traditional physical hosts and VPS servers, such as high management difficulty and weak business scalability.
[0100] An embodiment of the present application also discloses a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the neural network on-chip resource scheduling method provided in any embodiment of the present application. This program product and the neural network on-chip resource scheduling method disclosed in each embodiment of the present application belong to the same inventive concept, and thus will not be elaborated herein.
[0101] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. No limitation is made herein.
[0102] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for on-chip resource scheduling of a neural network, characterized in that Including: Based on the execution order of operators and the memory usage of operator tensors in the current neural network, draw a trend graph and segment the trend graph to obtain each network segment; Calculate the target cache batch size for each of the network segments; Perform in-chip data splicing on the last loop output data of the previous network segment and the remaining input of the subsequent network segment among the current adjacent network segments to obtain the target input data for the operator corresponding to the subsequent network segment, so as to perform data processing on the target input data based on the operator corresponding to the subsequent network segment; Segment the trend graph to obtain each network segment, including: Determine the operator to be compared for the current tensor memory usage and determine the similarity of the operator tensor memory usage of the operator to be compared for the current tensor memory usage; Segment the trend graph according to the similarity of the operator tensor memory usage and the lower limit threshold of the number of operators in a single network segment to obtain the current network segment.
2. The method according to claim 1, characterized in that, Determine the similarity of the operator tensor memory usage of the operator to be compared for the current tensor memory usage, including: Based on the operator tensor memory usage of the operator to be compared for the current tensor memory usage, determine the first memory usage similarity between the start-end operator and the end-end operator in the operator to be compared for the current tensor memory usage, or the second memory usage similarity between non-two-endpoint operators; Use the first memory usage similarity, or the second memory usage similarity as the similarity of the operator tensor memory usage.
3. The method according to claim 2, wherein Segment the trend graph according to the similarity of the operator tensor memory usage and the lower limit threshold of the number of operators in a single network segment to obtain the current network segment, including: If the number of the operators to be compared for the current tensor memory usage is less than the lower limit threshold of the number of operators in a single network segment, or the similarity of the operator tensor memory usage meets the similarity discrimination condition, then update the operator to be compared for the current tensor memory usage based on the operator with the next execution order after the end-end operator in the operator to be compared for the current tensor memory usage; And return to execute the operation of determining the similarity of the operator tensor memory usage based on the operator tensor memory usage of the operator to be compared for the current tensor memory usage until the number of the operators to be compared for the current tensor memory usage is greater than the lower limit threshold of the number of operators in a single network segment, and the similarity of the operator tensor memory usage does not meet the similarity discrimination condition, then segment the trend graph using the operator to be compared for the current tensor memory usage to obtain the current network segment.
4. The method according to claim 1, characterized in that Calculate the target cache batch size for each of the network segments, including: Obtain the in-memory peak value within a single batch segment for each of the network segments; Calculate the target cache batch size for each of the network segments based on the on-chip cache capacity of the target chip and the in-memory peak value within a single batch segment for each of the network segments.
5. The method according to claim 1, wherein After calculating the target cache batch size for each of the network segments, it further includes: Determine the network segments to be fused that meet the network segment fusion condition according to the target cache batch size for each of the network segments; Fuse the to-be-fused network segments to obtain a fused processing network segment, and update the network segments according to the fused processing network segment.
6. The method according to claim 1 or 5, characterized in that, After obtaining the target input data of the operator corresponding to the subsequent network segment, it further includes: After the operator corresponding to the subsequent network segment finishes processing the target input data, update the current adjacent network segment, and return to execute the operation of performing in-chip data splicing on the loop output data of the previous network segment and the remaining input of the subsequent network segment in the current adjacent network segment to obtain the target input data of the operator corresponding to the subsequent network segment, until the target input data of the operator corresponding to the full network segment is determined.
7. A neural network on-chip resource scheduling device, characterized in that, It includes: A trend graph drawing and segmentation module, configured to draw a change trend graph based on the operator execution order and operator tensor memory usage of the current neural network, and segment the change trend graph to obtain each network segment; A batch size calculation module, configured to calculate the target cache batch size of each of the network segments; An in-chip data splicing module, configured to perform in-chip data splicing on the last loop output data of the previous network segment and the remaining input of the subsequent network segment in the current adjacent network segment to obtain the target input data of the operator corresponding to the subsequent network segment, so as to process the target input data based on the operator corresponding to the subsequent network segment; The function of the trend graph drawing and segmentation module segmenting the change trend graph to obtain each network segment is specifically configured to determine the operator to be compared with the current tensor memory usage, and determine the operator tensor memory usage similarity of the operator to be compared with the current tensor memory usage; segment the change trend graph according to the operator tensor memory usage similarity and the lower limit threshold of the number of operators in a single network segment to obtain the current network segment.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the neural network in-chip resource scheduling method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to implement the neural network in-chip resource scheduling method according to any one of claims 1-6 when executed.
Citation Information
Patent Citations
Intermediate representation method and device for neural network model calculation
CN114186687A
Dynamic neural network compiling method and device, electronic equipment and storage medium
CN115658331A