Scheduling method, scheduling device, electronic device, and storage medium

By determining the data copying and transfer mode and effective data row filling among multiple computing units, the problem of data duplication in dynamic shape computation graphs is solved, thereby improving the computational efficiency of multi-layer convolutional neural networks.

CN117827386BActive Publication Date: 2026-08-04BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING YOUZHUJU NETWORK TECH CO LTD
Filing Date
2022-09-27
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing multi-core compilation and scheduling schemes cannot effectively handle computation graphs with dynamic shapes, resulting in data duplication and computational redundancy in data parallelism, which affects the computational efficiency of multi-layer convolutional neural networks.

Method used

By determining the data replication and transmission mode among multiple computing units, the filling and transmission of effective data rows can be achieved, optimizing the resource allocation of computing units, reducing redundant calculations, and improving computing power utilization.

Benefits of technology

It achieves balanced utilization of computing power in computing units, reduces redundant data calculations, improves the utilization rate of computing power, and enhances the computational efficiency of multi-layer convolutional neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117827386B_ABST
    Figure CN117827386B_ABST
Patent Text Reader

Abstract

A scheduling method, a scheduling device, an electronic device, and a storage medium. The scheduling method comprises: a plurality of calculation units respectively performing first convolution calculation on a plurality of corresponding data groups to obtain a plurality of corresponding first calculation result groups, wherein the plurality of first calculation result groups are used to constitute a first convolution layer obtained by first convolution calculation; determining a data replication transmission mode corresponding to the plurality of first calculation result groups on the plurality of calculation units according to a configuration rule of a second convolution layer obtained by performing second convolution calculation on the first convolution layer by the plurality of calculation units; and a first calculation unit in the plurality of calculation units that needs to perform effective data row padding, based on the corresponding data replication transmission mode, obtains a first intermediate data row required for the first calculation unit to perform padding in the second convolution calculation process from the first calculation result group on the second calculation unit. The scheduling method can effectively reduce the repeated calculation of data and improve the utilization rate of chip computing power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to a scheduling method, scheduling device, electronic device, and storage medium. Background Technology

[0002] Artificial Intelligence (AI) chips are chips specifically designed for performing neural network operations, specifically to accelerate the execution of neural networks. With the development of AI, the number of parameters in algorithm models is increasing dramatically, leading to a growing demand for computing power.

[0003] Because traditional hardware architectures (such as central processing units, CPUs) consider the balance between different business needs during the architecture design phase, their computing power for AI applications is limited. For the sake of high computing power, current AI chips, in addition to using general-purpose graphics processing units (GPUs), also widely employ domain-specific ASICs (DSAs) with homogeneous multi-core architectures. Summary of the Invention

[0004] This disclosure provides at least one embodiment of a scheduling method for a multi-layer convolutional neural network. The scheduling method includes: multiple computing units performing a first convolution calculation on corresponding multiple data groups to obtain corresponding multiple first calculation result groups, wherein the multiple first calculation result groups are used to constitute a first convolutional layer obtained by the first convolution calculation, and the multiple computing units include a first computing unit and a second computing unit; determining a data copying and transmission mode corresponding to the multiple first calculation result groups on the multiple computing units according to the configuration rules of the second convolutional layer obtained by the multiple computing units after performing a second convolution calculation on the first convolutional layer on the multiple computing units; and a first computing unit that needs to fill valid data rows in the multiple computing units, based on the corresponding data copying and transmission mode, obtaining a first intermediate data row required by the first computing unit to fill during the second convolution calculation process from the first calculation result group on the second computing unit.

[0005] At least one embodiment of this disclosure also provides a scheduling device, which includes: a computation control module configured to enable multiple computation units to perform a first convolution calculation on corresponding multiple data groups to obtain corresponding multiple first computation result groups, wherein the multiple first computation result groups are used to constitute a first convolutional layer obtained by the first convolution calculation, and the multiple computation units include a first computation unit and a second computation unit; an allocation scheduling module configured to determine a data copying and transmission mode corresponding to the multiple first computation result groups on the multiple computation units according to the configuration rules of the second convolutional layer obtained by the multiple computation units after performing a second convolution calculation on the first convolutional layer on the multiple computation units; and a data transmission module configured to enable a first computation unit that needs to fill valid data rows in the multiple computation units to obtain a first intermediate data row required by the first computation result group on the second computation unit for filling during the second convolution calculation, based on the corresponding data copying and transmission mode.

[0006] At least one embodiment of this disclosure also provides an electronic device, including the scheduling device provided in any embodiment of this disclosure.

[0007] At least one embodiment of this disclosure also provides an electronic device, the electronic device comprising: a processor; a memory including at least one computer program module; wherein the at least one computer program module is stored in the memory and configured to be executed by the processor, the at least one computer program module being used to implement the scheduling method described in any embodiment of this disclosure.

[0008] At least one embodiment of this disclosure also provides a storage medium storing non-transitory computer-readable instructions that, when executed by a computer, implement the scheduling method described in any embodiment of this disclosure. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.

[0010] Figure 1A This is a schematic diagram illustrating model parallelism.

[0011] Figure 1B This is a schematic diagram illustrating data parallelism.

[0012] Figure 2 This is a schematic diagram of a data-parallel convolution computation process represented on a computation graph;

[0013] Figure 3A flowchart illustrating a scheduling method provided in at least one embodiment of this disclosure;

[0014] Figure 4 for Figure 3 A schematic flowchart of step S20;

[0015] Figure 5 for Figure 3 A schematic flowchart of step S30;

[0016] Figures 6A-6D A schematic diagram illustrating a data copying and transfer process provided for at least one embodiment of this disclosure;

[0017] Figure 7 A schematic block diagram of a scheduling device provided for at least one embodiment of this disclosure;

[0018] Figure 8 A schematic block diagram of an electronic device provided for at least one embodiment of this disclosure;

[0019] Figure 9 A schematic block diagram of another electronic device provided for at least one embodiment of this disclosure;

[0020] Figure 10 A schematic block diagram of another electronic device provided for at least one embodiment of this disclosure; and

[0021] Figure 11 This is a schematic diagram of a storage medium provided for at least one embodiment of the present disclosure. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0023] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as “comprising” or “including” mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as “connected” or “linked” are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as “upper,” “lower,” “left,” and “right” are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes.

[0024] The present disclosure will now be described through several specific embodiments. To keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and known components may be omitted. When any component of an embodiment of the present disclosure appears in more than one drawing, that component is represented by the same or similar reference numerals in each drawing.

[0025] The core idea of ​​Domain Specific ASICs (DSA) is to use dedicated hardware for dedicated tasks. For example, a DSA is designed to meet the needs of an application within a specific domain, rather than a fixed application. Therefore, DSA achieves a trade-off between flexibility and specialization. A DSA accelerator refers to a chip that uses multiple interconnected computing units (or Process Element (PE) or processing cores) to expand computing power and accelerate computation. These multiple processing units are interconnected via communication lines, such as buses or on-chip networks. For this type of multi-PE interconnected AI accelerator, effectively scheduling hardware resources to efficiently complete the inference of deep neural network models is a major challenge faced by AI compilers.

[0026] Currently, there are two main implementation methods for multi-core compilation scheduling schemes.

[0027] One approach is model parallelism, where different devices are responsible for computations on different parts of the computation graph. For example, in some domain-specific accelerators, the operators in the computation graph are statically assigned to different physical processors (PEs) within the chip. Then, input data is fed into the PE containing the first operator. Once the first PE has completed its computation, the data is transferred to the next PE, and so on, until all operators have been computed. In model parallelism, each PE executes only the operators in its local memory. Each PE must wait for the preceding operator to finish before it can begin its computation; this is a pipelined approach to serialize computations.

[0028] For example, such as Figure 1A As shown, PE0, PE1, and PE2 each load three different operators, which are convolution operators corresponding to different layers of the neural network. For example, PE0 has convolution operator A, PE1 has convolution operator B, and PE2 has convolution operator C. During computation, the server first loads the input data into PE0, where convolution operator A resides. After PE0 completes the first convolution calculation on the input data, it passes the result of the first convolution calculation to PE1. On PE1, the second convolution calculation corresponding to convolution operator B is executed, and so on, until convolution operator C on PE2 is completed. Finally, the result obtained on PE2 is returned to the host.

[0029] However, model parallelism requires static compilation, which necessitates distributing some operators in the computation graph across a fixed number of physical instances (PEs). Therefore, it cannot handle computation graphs with dynamic shapes. Dynamic shapes refer to tensors whose shapes depend on specific operations and cannot be obtained through pre-calculation; that is, a dynamic computation graph is built upon each step of the computation. Therefore, model parallelism that only includes the computation graph portion cannot handle dynamic computation graphs.

[0030] Another approach is data parallelism, where each device has a complete computation graph. For example, in some GPUs, each operator in the computation graph is first loaded sequentially into the GPU device, and each operator is placed in parallel on multiple PEs. Then, the input data is split into multiple parts and loaded onto multiple PEs respectively. The multiple PEs process the input data in parallel until the computation of all operators is completed. Finally, the computation results from the multiple PEs are aggregated and returned to the host.

[0031] For example, such as Figure 1BAs shown, PE0, PE1, and PE2 each have three different operators loaded, such as convolution operators for different layers of the neural network. For example, PE0, PE1, and PE2 each have convolution operators A, B, and C. During computation, the server first splits the input data into multiple parts, for example, evenly splitting it into three parts. Then, the three parts of input data are loaded onto PE0, PE1, and PE2 respectively. PE0, PE1, and PE2 perform multiple convolution calculations on the input data in parallel, sequentially using the corresponding convolution operators A, B, and C, until the computation is complete. Finally, the system aggregates the final results from PE0, PE1, and PE2 and returns them to the host.

[0032] There is currently no fixed and feasible model for implementing data parallelism; scheduling optimization is required based on different hardware characteristics. In actual model deployment, some models have a large amount of operator data, while others have a small amount of operator data. Therefore, different deployment schemes for AI chips are needed based on the model to obtain the optimal inference performance.

[0033] In current data parallelism methods, for input data with a large amount of data, the PE cannot hold all the input data required for computation at once. Therefore, it is usually necessary to split the input data into multiple groups, and then place the data of multiple groups on different PEs for parallel computation by multiple PEs. After the computation is completed, the computation results on multiple PEs are aggregated to obtain the final output result corresponding to the input data.

[0034] However, in the inference computation process of multi-layer neural networks, considering the padding operation of convolution computation, directly splitting large images or large amounts of data into multiple groups can easily introduce the problem of repeated data computation.

[0035] For example, an exemplary computation graph contains the following three convolution operators (conv):

[0036] %5: [(1 224 224 64), F, FP32] = conv (%0: [(1 224 224 64), F, FP32], 1%: [(64 3 3 64), W, FP32], kh=3, kw=3, pad_h_top=1, pad_h_bottom=1, pad_w_left=1, pad_w_right=1, stride_h=1, stride_w=1);

[0037] %10: [(1 224 224 64), F, FP32] = conv (%5: [(1 224 224 64), F, FP32], 6%: [(64 3 3 64), W, FP32], kh=3, kw=3, pad_h_top=1, pad_h_bottom=1, pad_w_left=1, pad_w_right=1, stride_h=1, stride_w=1);

[0038] %15: [(1 224 224 64), F, FP32] = conv (%10: [(1 224 224 64), F, FP32], 11%: [(64 3 3 64), W, FP32], kh=3, kw=3, pad_h_top=1, pad_h_bottom=1, pad_w_left=1, pad_w_right=1, stride_h=1, stride_w=1).

[0039] %0, %5, %10, and %15 represent tensor data. For example, the input image for the first convolutional layer of a neural network can be represented as [B, H, W, C], meaning the dimensions of the input data are B×H×W×C. For example, %0 indicates that the size of the input image for the first convolutional layer of the neural network is [1, 224, 224, 64], that is, the input image has 64 channels, and the image size (height H × width W) of each channel is 224×224, for a total of 1 batch. For example, the size of the convolution kernel used by each convolution operator conv can be represented as kh×kw, that is, kh=3, kw=3 indicates that the size of the convolution kernel is 3×3. For example, the padding operation of each convolution operator (conv) is represented by `pad`. For instance, `pad_h_top = 1` means padding one row (e.g., 0) at the top of the image along the height (h); `pad_h_bottom = 1` means padding one row (e.g., 0) at the bottom of the image along the height (h); `pad_w_left = 1` means padding one column (e.g., 0) at the left side of the image along the width (w); and `pad_w_right = 1` means padding one column (e.g., 0) at the right side of the image along the width (w). Similarly, the stride of the convolution kernel in each convolution operator (conv) is represented by `stride`. For instance, `stride_h = 1` means the kernel moves one pixel at a time along the height (h); and `stride_w = 1` means the kernel moves one pixel at a time along the width (w). This calculation uses floating-point arithmetic with a precision of FP32.

[0040] For example, Figure 2 A schematic diagram of a data-parallel convolution computation process is shown, represented on a computation graph containing the three convolution operators described above. For example, when the memory of a single computational object (PE) cannot hold all the data required for computation at once, the input data needs to be split into multiple data groups based on the PE's memory and quantity. For example, in... Figure 2 The input data has 3 PEs and the size of the input data is 224×224. The input data is split into 3 data groups and loaded into the 3 PEs respectively for convolution calculation.

[0041] The selection of input data splitting and the size and range of the resulting data groups depends on the output data. For example, in Figure 2 In the computation graph shown, the output data %15 is 224×224. By dividing the data into three roughly evenly partitioned parts along the height H dimension of the output data, the heights of the three output data groups in the three rounds (round1, round2, and round3) corresponding to the three PEs are 75 rows, 75 rows, and 74 rows, respectively. When partitioning the input data, each output data group requires a depth-first traversal of the entire data group to recursively derive the range of each input data group.

[0042] It should be noted that, Figure 2 The data shown in the computation graph are all "valid data". For simplification, "valid data" here refers to the meaningful and effective data obtained from each convolution calculation. In actual calculation, because padding is required after each convolution calculation to maintain the image size, the area formed by the data at the "edge" of each split data group and the padded "0" for convolution with the convolution kernel is an invalid region. The result obtained from convolution in this region is invalid data and needs to be discarded. That is, the output data obtained from convolution of complete input data is valid data. Therefore, in the split data groups, only the data from the valid regions obtained from convolution is valid data. For clarity and conciseness, the data shown in the computation graph in the accompanying drawings of this application are all valid data.

[0043] For example, Figure 2The output data %15 in round1 has 75 valid rows, i.e., [0, 74]. Since the padding operation is: pad_h_top = 1, pad_h_bottom = 1, the intermediate data %10 obtained by recursively tracing forward from %15 needs 76 valid rows, i.e., [0, 75]. The intermediate data %5 obtained by recursively tracing forward from intermediate data %10 needs 77 valid rows, i.e., [0, 76], and so on. Therefore, the final input data %0 needs 78 valid rows. Thus, the range of data read from the input data %0 in round1 along the height direction is [0, 77] (row numbers start from 0), a total of 78 rows.

[0044] Similarly, Figure 2 Round 2 shows that the output data %15 has 75 valid rows, i.e., [75, 149]. Since the padding operation is: pad_h_top = 1, pad_h_bottom = 1, the intermediate data %10 obtained by recursively tracing forward from %15 needs 77 valid rows, i.e., [74, 150]. The intermediate data %5 obtained by recursively tracing forward from intermediate data %10 needs 79 valid rows, i.e., [73, 151], and so on. Therefore, the final input data %0 needs 81 valid rows. Thus, the range of data read from the input data %0 in the height direction in round 2 is [72, 152], a total of 81 rows.

[0045] Similarly, Figure 2 Round 3 shows that the output data %15 has 74 valid rows, i.e., [150, 223]. Since the padding operation is: pad_h_top = 1, pad_h_bottom = 1, the intermediate data %10 obtained by recursively tracing forward from %15 needs 75 valid rows, i.e., [149, 223]. The intermediate data %5 obtained by recursively tracing forward from intermediate data %10 needs 76 valid rows, i.e., [148, 223], and so on. Therefore, the final input data %0 needs 77 valid rows. Thus, the range of data read from the input data %0 in the height direction in round 3 is [147, 223], a total of 77 rows.

[0046] from Figure 2The computation graph clearly shows that the repeated computation regions in rounds 1 and 2 are [72, 77] rows, meaning 6 rows are repeated. Similarly, the repeated computation regions in rounds 2 and 3 are [147, 152] rows, also repeating 6 rows. This indicates that the deeper the convolutional network and the more times it iterates forward, the more effective data is needed at the edges of data groups caused by padding in the previous layer. In other words, the regions requiring repeated computation are larger, leading to significant computational redundancy.

[0047] This disclosure provides at least one embodiment of a scheduling method for a multi-layer convolutional neural network. The scheduling method includes: multiple computing units performing first convolution calculations on corresponding multiple data groups to obtain corresponding multiple first calculation result groups, wherein the multiple first calculation result groups constitute a first convolutional layer obtained from the first convolution calculation; the multiple computing units include a first computing unit and a second computing unit; determining a data copying and transmission mode corresponding to the multiple first calculation result groups on the multiple computing units based on the configuration rules of the second convolutional layer obtained by the multiple computing units performing a second convolution calculation on the first convolutional layer; and a first computing unit that needs to fill valid data rows, based on the corresponding data copying and transmission mode, obtaining the first intermediate data rows required for filling during the second convolution calculation from the first calculation result groups on the second computing unit. This scheduling method, by copying and transmitting valid data from one computing unit to another, can achieve balanced utilization of computing power, reduce redundant data calculations, and improve the utilization rate of computing power.

[0048] At least one embodiment of this disclosure also provides a scheduling device, electronic device, and storage medium. Similarly, by copying and transmitting valid data from one computing unit to another, this scheduling device, electronic device, and storage medium can achieve balanced utilization of computing power, reduce redundant data computation, and improve the utilization rate of computing power.

[0049] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the same reference numerals in different drawings will be used to refer to the same elements described.

[0050] Figure 3 This is a flowchart illustrating a scheduling method provided in at least one embodiment of the present disclosure. Figure 3 As shown, the scheduling method includes steps S10 to S30.

[0051] Step S10: Multiple computing units perform first convolution calculations on corresponding multiple data groups to obtain corresponding multiple first calculation result groups, wherein the multiple first calculation result groups are used to form the first convolution layer obtained by the first convolution calculation, and the multiple computing units include a first computing unit and a second computing unit.

[0052] Step S20: Based on the configuration rules of the second convolutional layer obtained after the second convolutional calculation of the first convolutional layer by multiple computing units in multiple computing units, determine the data copying and transmission mode corresponding to the multiple first calculation result groups on multiple computing units.

[0053] Step S30: The first computing unit that needs to fill the effective data row among the multiple computing units obtains the first intermediate data row required for filling in the second convolution calculation process from the first calculation result group on the second computing unit based on the corresponding data copy and transfer mode.

[0054] For example, this scheduling method can be used in computing devices such as AI chips that perform convolution operations on multi-layer convolutional neural networks, general-purpose graphics processing units (GPGPUs), etc. For example, the AI ​​chip can be an AI chip using a DSA accelerator or an AI chip with multiple PE units. For example, it can be used in data-parallelized distributed training. The embodiments of this disclosure do not limit this.

[0055] It should be noted that, in at least one embodiment of this disclosure, "first convolution calculation" can refer to a convolution calculation performed on the complete input image, or it can refer to a convolution calculation performed on the current input data group in the computing unit. "First convolution calculation" can be any convolution calculation in a multi-layer convolution calculation, rather than a convolution calculation performed only on the first layer. For example, "first convolution calculation" can be a convolution calculation performed on the first layer, or it can be a convolution calculation performed on the second layer, the third layer, and so on.

[0056] In at least one embodiment of this disclosure, "second convolution calculation" refers to the next convolution calculation corresponding to "first convolution calculation," that is, a convolution calculation performed using the "first convolutional layer" obtained after the first convolution calculation as input. The "first convolutional layer" is the convolutional layer composed of the calculation results obtained from the "first convolution calculation," where "first convolutional layer" refers to the actual convolutional layer obtained by performing the first convolution calculation on the complete input image without splitting. For example, the first convolutional layer can be the input data for the second convolution calculation. Similarly, "second convolutional layer" is the convolutional layer composed of the calculation results obtained from the "second convolution calculation," where "second convolutional layer" refers to the actual convolutional layer obtained by performing the second convolution calculation on the first convolutional layer composed of the data obtained without splitting.

[0057] The following will combine Figures 6A-6D The computational graph shown details the scheduling method of this disclosure according to embodiments. In the following exemplary description, three PEs (...) Figures 6A-6C The following example is used for illustration, but the embodiments of this disclosure are not limited thereto. The computing device may include, for example, 2 PEs, 4 PEs (…). Figure 6D Or more PEs.

[0058] It should be noted that, Figures 6A-6D The elongated region shown is an abstract representation of the range of at least a portion of the valid data in a certain dimension (e.g., height H or width W) of the input image for each convolution calculation. The diagonal region, grid line region, and diamond line region are abstract representations of data groups on the calculation units (PE0, PE1, PE2, etc.), and their respective positions on the elongated region are merely schematic representations of the corresponding positions of their respective data groups on the input image, and do not constitute a limitation on the embodiments of this disclosure.

[0059] For example, in step S10, the "multiple data groups" corresponding to the multiple computing units can be multiple original input data groups obtained by splitting the original input matrix (these original input data groups are input into the computing device), or they can be multiple input data groups targeted by any convolution calculation in the computing device. For example, the multiple data groups can be Figure 6A The multiple data sets 10, 20, and 30 on PE0, PE1, and PE2 that perform convolution calculations on %0 can also be multiple data sets that perform convolution calculations on %5. For example, the data set for each computing unit to perform convolution calculations on %5 includes a first calculation result set and an intermediate data row obtained from another computing unit. That is, in at least one embodiment of this disclosure, the first calculation result set on the first computing unit and the first intermediate data row obtained from the second computing unit constitute the data set required for the first computing unit to perform the second convolution calculation.

[0060] For example, in step S10, multiple computing units perform first convolution calculations on corresponding data groups to obtain corresponding first calculation result groups. For example, the multiple first calculation result groups can be used to construct a first convolutional layer obtained from the first convolution calculation. For example, as... Figure 6A As shown, multiple computing units PE0, PE1, and PE2 perform first convolution calculations on corresponding multiple data groups 10, 20, and 30 to obtain multiple first calculation result groups 11, 21, and 31. These multiple first calculation result groups 11, 21, and 31 can be used to form the first convolution layer obtained from the first convolution calculation.

[0061] For example, in the embodiments of this disclosure, the multiple first calculation result groups in multiple computing units can directly constitute the first convolutional layer obtained by the first convolution calculation, or they can constitute a part of the first convolutional layer obtained by the first convolution calculation. For example, when the input image is small, multiple computing units can carry all the input data at once, and the multiple first calculation result groups in multiple computing units can be combined to form the first convolutional layer; when the input image is large, multiple computing units cannot carry all the input data at once, and it is necessary to divide the input data into multiple loading computing units, split the input data loaded each time into multiple groups, and put them into multiple computing units respectively, then the multiple first calculation result groups in multiple computing units only constitute a part of the first convolutional layer corresponding to the input data loaded in that time.

[0062] For example, in step 10, the data rows in the multiple first calculation result groups are continuous without overlap in the first convolutional layer. For example, as Figure 6A As shown, the data rows in the multiple first calculation result groups 11, 21, and 31 are continuous without overlap in the first convolutional layer. That is, the data in the multiple first calculation result groups 11, 21, and 31 obtained by performing the first convolution calculation on %0 on the multiple calculation units PE0, PE1, and PE2 have no overlapping parts.

[0063] For example, in embodiments of this disclosure, a "data row" can represent a row of data or a column of data in an input image. For example, in Figure 6A In the data group shown that is split along the height H dimension, a "data row" represents a row of data in the height direction of the input image. Of course, a "data row" can also represent a column of data in the width direction of the input image. The embodiments of this disclosure do not limit this.

[0064] For example, in step S20, based on the configuration rules of the second convolutional layer obtained after performing a second convolution calculation on the first convolutional layer by multiple computing units, the data replication and transmission mode corresponding to the multiple first calculation result groups on the multiple computing units is determined. It should be noted that "performing a second convolution calculation on the first convolutional layer" here does not refer to the actual convolution calculation in the computing unit, but rather to the fact that, based on the rules of convolution calculation, the size of the second convolutional layer obtained after the second convolution calculation can be determined from the size of the first convolutional layer. Therefore, "the obtained second convolutional layer" does not refer to the specific calculation result of the data in the second convolutional layer. For example, the splitting mode of the second convolutional layer can be determined based on its size, that is, the configuration rules of the second convolutional layer in the multiple computing units can be determined.

[0065] Figure 4 for Figure 3 A schematic flowchart of step S20. For example, in some examples, such as Figure 4As shown, step S20 may further include steps S21 to S23.

[0066] Step S21: Determine the configuration rules of the second convolutional layer in the multiple computing units based on the memory size of the memory on the multiple computing units;

[0067] Step S22: Based on the configuration rules, determine the distribution of the second convolutional layer across multiple computational units;

[0068] Step S23: Determine the data replication and transmission mode corresponding to the first calculation result group on multiple computing units based on the distribution.

[0069] For example, in step S21, the memory (e.g., RAM) sizes on multiple computing units are obtained, and the configuration rules for the second convolutional layer in the multiple computing units are determined based on the memory size of each computing unit. For example, the second convolutional layer is split into multiple data segments, and the size of each data segment must not exceed the available capacity of the memory allocated to its corresponding computing unit; this is a prerequisite for determining the configuration rules.

[0070] For example, in step S22, based on the determined configuration rules, it is determined how to split the second convolutional layer, that is, to determine the distribution of the multiple data groups split into the second convolutional layer on multiple computing units. For example, corresponding to the distribution of the multiple first calculation result groups on multiple computing units and the determined configuration rules, the distribution of the multiple data groups split into the second convolutional layer on the corresponding multiple computing units can be determined.

[0071] For example, in step S23, based on the distribution of the second convolutional layer across multiple computing units as determined by the configuration rules, a data copying and transmission mode corresponding to the first computation result group across the multiple computing units is determined. For example, the data copying and transmission mode may include a unidirectional data transmission mode and a bidirectional data transmission mode. For example, a unidirectional data transmission mode means that a computing unit can only obtain copied data from another unit, while a bidirectional data transmission mode means that a computing unit can not only obtain copied data from another computing unit but also copy and transmit its own data to that other computing unit. For example, multiple computing units may be configured to have the same data copying and transmission mode, or multiple computing units may be configured to have different data copying and transmission modes. The data copying and transmission mode is not limited to the above two modes and may also be other feasible implementations; the embodiments of this disclosure do not limit this.

[0072] For example, in one example, the configuration rule could be: the distribution of the data from the second convolutional layer across multiple computational units is the same as the data range of the first computational result group in those multiple computational units. For example, as... Figure 6AAs shown, multiple computing units PE0, PE1, and PE2 perform a first convolution calculation on the data group %0 to obtain a first convolutional layer composed of multiple first calculation result groups 11, 21, and 31. The size of the second convolutional layer obtained by performing a second convolution calculation on the first convolutional layer can be calculated from this. Therefore, it is proposed to split the data of the second convolutional layer and proceed according to the following... Figure 6A The distribution of %10 is allocated to multiple computing units PE0, PE1, and PE2. Based on this configuration rule, the data replication and transfer mode for each computing unit when performing convolution calculations on %5 can be determined. For example, this configuration rule can determine that the data replication and transfer mode for multiple computing units PE0, PE1, and PE2 is a bidirectional data transfer mode.

[0073] For example, in another example, the configuration rule could be: the distribution of the data from the second convolutional layer across a subset of computational units is the same as the data range of the first computational result group within that subset of computational units. For example, as... Figure 6B As shown, after knowing the size of the second convolutional layer, the data of the second convolutional layer is to be processed as follows: Figure 6B The data on %10 is distributed across multiple computing units PE0, PE1, and PE2. In other words, the data of the second convolutional layer is split into three parts, with the data range allocated to PE1 being equal in size to the data range of the first computation result group on PE1. Based on this configuration rule, the data copying and transmission mode for each computing unit performing convolution calculations on %5 can be determined. For example, the data copying and transmission mode for multiple computing units PE0, PE1, and PE2 can be determined to be a unidirectional data transmission mode.

[0074] In the embodiments of this disclosure, the configuration rules can be set not only according to the memory size of the computing unit, but also flexibly according to factors such as the computing power allocation of multiple computing units and data transmission overhead. The embodiments of this disclosure do not limit the specific content of the configuration rules.

[0075] For example, in step S30, the first computing unit among multiple computing units that needs to be filled with valid data rows is first determined. For example, it could be that all multiple computing units need to be filled with valid data rows, or it could be that only some of the multiple computing units need to be filled with valid data rows. For example, in... Figure 6A In the above, PE0, PE1, and PE2 are all the first calculation units that need to be filled with valid data rows; Figure 6B In the diagram, PE1 and PE2 are the first calculation units that need to be filled with valid data rows; Figure 6C In this context, only PE1 is the first calculation unit that needs to be filled with valid data rows.

[0076] For example, in step S30, after determining the first computing unit that needs to perform valid data row filling and the data copying and transmission mode corresponding to the first computing unit, the first computing unit obtains at least one first intermediate row data from the first calculation result group of another computing unit. For example, the first computing unit can obtain two first intermediate row data from the first calculation result groups of two different computing units, both of which are required for filling during the second convolution calculation process of the first computing unit. Figure 6A In this process, PE0, as the first calculation unit, obtains one or more rows of first intermediate row data from the first calculation result group 21 of PE1; PE1, as the first calculation unit, obtains one or more rows of first intermediate row data from the first calculation result group 11 of PE0, and obtains one or more rows of first intermediate row data from the first calculation result group 31 of PE2; PE2, as the first calculation unit, obtains one or more rows of first intermediate row data from the first calculation result group 21 of PE1.

[0077] Figure 5 for Figure 3 A schematic flowchart of step S30. For example, in some examples, such as Figure 5 As shown, step S30 may further include steps S31 and S32.

[0078] Step S31: Based on the relationship between the second calculation result group obtained after the second convolution calculation of the first calculation unit and the first calculation result group obtained after the first convolution calculation of the first calculation unit, determine the size of the first intermediate data row;

[0079] Step S32: Based on the data copy transmission mode and the size of the first intermediate data row, obtain the first intermediate data row from the first calculation result group on the second computing unit.

[0080] For example, in step S31, the size of the first intermediate data row obtained by the first computing unit from another computing unit is determined by the relationship between the second computing result group and the first computing result group. The relationship between the second computing result group and the first computing result group is determined by factors such as the distribution of the second convolutional layer on the computing unit, the size of the convolutional kernel, and the padding operation.

[0081] For example, in step S32, the first intermediate data line can be obtained from the first calculation result group on the second computing unit through inter-core transfer between computing units. For example, by obtaining the address of the first intermediate data line in the first calculation result group of the second computing unit, the first intermediate data line is read and copied to the first computing unit, and the copied first intermediate data line is written to the corresponding address of the first computing unit. For example, in... Figure 6AIn this example, PE1 has 74,150 rows of valid data (%0). These 77 rows correspond to storage addresses address_0 to address_75. After the first convolution calculation, the first calculation result group on PE1 consists of rows [75,149] in the first convolutional layer, totaling 75 rows. These rows can be allocated to addresses address_1 to address_74. The first intermediate data row obtained from the first calculation result group 11 of PE0 is written to address_0, and the first intermediate data row obtained from the first calculation result group 31 of PE2 is written to address_75. The embodiments of this disclosure do not limit the specific implementation of the data copying and transmission process between computing units.

[0082] For example, the scheduling method further includes step S40 (not shown in the figure): splitting the original input matrix to obtain multiple original input data groups, and transmitting the multiple original input data groups to multiple computing units respectively for the first convolution calculation.

[0083] For example, the splitting pattern of the original input matrix can be determined based on the size of the first convolutional layer obtained after performing the first convolution calculation on the original input matrix, the number of computation units, and the memory size. For instance, the splitting pattern could be to divide the original input matrix into multiple groups of original input data evenly, or to split the original input matrix into multiple groups of original input data according to the memory size of the computation units, so that each group is placed into a separate computation unit. For example, the original input data groups can be used as the data groups targeted for the first convolution calculation.

[0084] For example, the scheduling method further includes step S50 (not shown in the figure): in response to the second convolutional layer being the output layer in a multi-layer convolutional neural network, multiple second computation result groups from multiple computation units are aggregated to obtain at least a portion of the computational output of the multi-layer convolutional neural network. For example, when the second convolutional layer is the last convolutional layer after the last convolution operator has been computed, the second convolutional layer is the output result. That is, the second convolutional layer composed of the second computation result groups from multiple computation units is the computational output or at least a portion of the computational output of the multi-layer convolutional neural network.

[0085] In embodiments of this disclosure, the plurality of computing units includes at least two computing units, such as a first computing unit and a second computing unit, and the multi-layer convolutional neural network includes at least two convolutional layers. For example, Figures 6A-6C The example shown includes three computational units, PE0, PE1, and PE2, containing the following two convolution operators (conv):

[0086] %5: [(1 224 224 64), F, FP32] = conv (%0: [(1 224 224 64), F, FP32], 1%: [(64 3 3 64), W, FP32], kh=3, kw=3, pad_h_top=1, pad_h_bottom=1, pad_w_left=1, pad_w_right=1, stride_h=1, stride_w=1);

[0087] %10: [(1 224 224 64), F, FP32] = conv (%5: [(1 224 224 64), F, FP32], 6%: [(64 3 3 64), W, FP32], kh=3, kw=3, pad_h_top=1, pad_h_bottom=1, pad_w_left=1, pad_w_right=1, stride_h=1, stride_w=1);

[0088] For example, such as Figure 6A As shown, the chip uses three computing units PE0, PE1, and PE2 to perform parallel convolution calculations on the 224×224 input data. For example, in step S40, the original input matrix is ​​roughly divided into three original input data groups 10, 20, and 30 in terms of the height H dimension of the input data, and loaded onto the three computing units PE0, PE1, and PE2 respectively, as shown in %0. For example, original input data group 10 includes [0, 75], with a total of 76 rows of data; original input data group 20 includes [74, 150], with a total of 77 rows of data; and original input data group 30 includes [149, 223], with a total of 75 rows of data. For example, in step S10, multiple computing units PE0, PE1, and PE2 perform first convolution calculations on the corresponding three original input data groups 10, 20, and 30 respectively, to obtain first calculation result groups 11 ([0,74] rows), 21 ([75,149] rows), and 31 ([150,223] rows) as shown in %5. These first calculation result groups constitute the first convolutional layer corresponding to the first convolution calculation performed on the original input matrix.

[0089] For example, in step S20, based on the configuration rule that the distribution of the data of the second convolutional layer across multiple computing units is the same as the data range of the first calculation result group in the multiple computing units, it is determined that the three computing units PE0, PE1, and PE2 all have the same data copying and transmission mode, i.e., a bidirectional data transmission mode. Then, in step S30, it is determined that the three computing units PE0, PE1, and PE2 are all computing units that need to perform data copying and transmission, and it is determined that the first intermediate data row required by each computing unit is one row. Then, each computing unit obtains the required first intermediate data row from another computing unit to fill in the second convolution calculation.

[0090] For example, PE0 is the first computing unit and PE1 is the second computing unit. As shown by arrow 1, the first computing unit PE0 obtains the first intermediate data row required for filling during the second convolution calculation from the first calculation result group 21 of the second computing unit PE1

[75] . As shown by arrow 2, the second computing unit PE1 obtains the second intermediate data row required for filling during the second convolution calculation from the first calculation result group 11 of the first computing unit PE0

[74] . The first intermediate data row and the second intermediate data row obtained by the first computing unit PE0 and the second computing unit PE1 from each other are the same size, which is 1 row of data.

[0091] For example, PE1 is the first computing unit and PE2 is the second computing unit. As shown by arrow 3, the first computing unit PE1 obtains the first intermediate data row required for filling during the second convolution calculation from the first calculation result group 31 of the second computing unit PE2

[150] . As shown by arrow 4, the second computing unit PE2 obtains the second intermediate data row required for filling during the second convolution calculation from the first calculation result group 21 of the first computing unit PE1

[149] . The first intermediate data row and the second intermediate data row obtained by the first computing unit PE1 and the second computing unit PE2 from each other are the same size, which is 1 row of data.

[0092] For example, the first calculation result group [0,74] on PE0 and the first intermediate data row

[75] obtained from PE1 constitute the data group [0,75] required for PE0 to perform the second convolution calculation.

[0093] For example, the first calculation result group [75,149] on PE1 and the second intermediate data row

[74] and the first intermediate data row

[150] obtained from PE0 and PE2 respectively constitute the data group [74,150] required for PE1 to perform the second convolution calculation.

[0094] For example, the first calculation result [150,223] on PE2 and the second intermediate data row

[149] obtained from PE2 constitute the data set [149,223] required for PE2 to perform the second convolution calculation.

[0095] Three computational units PE0, PE1, and PE2 perform a second convolution calculation on the data sets [0,75], [74,150], and [149,223] after data copying and transmission, resulting in the second calculation result set [0,74], [75,149], and [150,223] as shown in %10. Finally, step S50 is executed to combine the second calculation result sets [0,74], [75,149], and [150,223] on PE0, PE1, and PE2, which is the calculation result of the output layer.

[0096] Therefore, the scheduling method provided in at least one embodiment of this disclosure reduces the amount of computation of repetitive data in convolutional neural networks and improves the computational efficiency of multi-layer convolutional neural networks by using the copying and transmission of effective data between computing units and using effective data as filling rows.

[0097] In such Figure 6B In another example shown, in step S20, it can be determined that the three computing units PE0, PE1, and PE2 all have the same unidirectional data transmission pattern. In step S30, computing units PE1 and PE2 are determined to be the computing units that need to perform data copying and transmission, and it is determined that the first intermediate data line required by each computing unit is two lines. That is, PE1 obtains two lines of valid data from PE0 for filling, and PE2 obtains two lines of valid data from PE1 for filling. For example, with Figure 6B In comparison, Figure 6C In this configuration, only computing unit PE1 can be designated as the computing unit requiring data copying and transmission. That is, PE1 retrieves two rows of valid data from PE0 and PE2 respectively for filling. Figure 6B and Figure 6C The scheduling method shown in the example can not only reduce redundant data calculations and improve computational efficiency, but also reduce the number of data transfers between computing units and reduce the overhead required for data transmission.

[0098] For example, the computing device may also include more computing units, and the convolutional network may also include more convolution operators; the embodiments of this disclosure are not limited in this regard. For example, Figure 6D The example shown includes four computational units, PE0, PE1, PE2, and PE3, containing the following three convolution operators (conv):

[0099] %5: [(1 224 224 64), F, FP32] = conv (%0: [(1 224 224 64), F, FP32], 1%: [(64 3 3 64), W, FP32], kh=3, kw=3, pad_h_top=1, pad_h_bottom=1, pad_w_left=1, pad_w_right=1, stride_h=1, stride_w=1);

[0100] %10: [(1 224 224 64), F, FP32] = conv (%5: [(1 224 224 64), F, FP32], 6%: [(64 3 3 64), W, FP32], kh=3, kw=3, pad_h_top=1, pad_h_bottom=1, pad_w_left=1, pad_w_right=1, stride_h=1, stride_w=1);

[0101] %15: [(1 224 224 64), F, FP32] = conv (%10: [(1 224 224 64), F, FP32], 11%: [(64 3 3 64), W, FP32], kh=3, kw=3, pad_h_top=1, pad_h_bottom=1, pad_w_left=1, pad_w_right=1, stride_h=1, stride_w=1).

[0102] For example, in Figure 6D As shown, the multiple computing units PE0, PE1, PE2, and PE3 can have different data replication and transmission modes. For example, the size of the first intermediate data line obtained by each computing unit from another computing unit can be different, and the embodiments of this disclosure do not limit this. For example, PE2 obtains 1 line of first intermediate data from PE1, while obtaining 2 lines of first intermediate data from PE3. Figure 6D The steps of the scheduling method in the example shown can be found in the documentation. Figure 6A The description will not be repeated here. (Using...) Figure 6D The scheduling method shown in the example can not only reduce redundant data computation and improve the utilization of computing power, but also achieve a balanced allocation of memory, computing power and data transmission overhead among different computing units based on the overall performance of the system.

[0103] Figure 7 This is a schematic block diagram of a scheduling device provided for some embodiments of this disclosure. Figure 7As shown, the scheduling device 100 includes a calculation control module 110, an allocation scheduling module 120, and a data transmission module 130. These components can be interconnected via a bus and / or other forms of connection mechanisms (not shown).

[0104] The calculation control module 110 is configured to enable multiple calculation units to perform first convolution calculations on corresponding multiple data groups to obtain corresponding multiple first calculation result groups, wherein the multiple first calculation result groups are used to form a first convolution layer obtained by the first convolution calculation, and the multiple calculation units include a first calculation unit and a second calculation unit.

[0105] The allocation and scheduling module 120 is configured to determine the data copying and transmission mode corresponding to the multiple first calculation result groups on the multiple computing units according to the configuration rules of the second convolutional layer obtained after the second convolutional calculation of the first convolutional layer by the multiple computing units in the multiple computing units.

[0106] The data transmission module 130 is configured to enable the first computing unit, which requires effective data row filling among multiple computing units, to obtain the first intermediate data row required for filling during the second convolution calculation process from the first calculation result group on the second computing unit based on the corresponding data copy transmission mode.

[0107] It should be noted that in the embodiments of this disclosure, each module of the scheduling device 100 corresponds to each step of the aforementioned scheduling method. For the specific functions of the scheduling device 100, please refer to the relevant description of the scheduling method above, which will not be repeated here. Figure 7 The components and structure of the scheduling device 100 shown are exemplary and not limiting; the scheduling device 100 may also include other components and structures as needed.

[0108] Figure 8 This is a schematic block diagram of an electronic device provided for some embodiments of this disclosure. For example... Figure 8 As shown, the electronic device 200 includes a scheduling device 210, which can be any scheduling device provided in any embodiment of this disclosure, such as the aforementioned scheduling device 100. The electronic device 200 can be any device with computing capabilities, such as a server, terminal device, personal computer, etc., and the embodiments of this disclosure do not limit this.

[0109] Figure 9 This is a schematic block diagram of another electronic device provided for some embodiments of this disclosure. For example... Figure 9As shown, the electronic device 300 includes a processor 310 and a memory 320, and can be used to implement a client or server. The memory 320 stores computer-executable instructions (e.g., at least one or more computer program modules) non-transitoryly. The processor 310 executes the computer-executable instructions, which, when run by the processor 310, can perform one or more steps of the convolution operation method described above, thereby implementing the convolution operation method described above. The memory 320 and the processor 310 can be interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0110] For example, processor 310 may be a central processing unit (CPU), a graphics processing unit (GPU), or other form of processing unit with data processing and / or program execution capabilities. For example, the central processing unit (CPU) may be an x86 or ARM architecture. Processor 310 may be a general-purpose processor or a special-purpose processor, capable of controlling other components in electronic device 300 to perform desired functions.

[0111] For example, memory 320 may include any combination of at least one (e.g., one or more) computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. At least one (e.g., one or more) computer program modules may be stored on the computer-readable storage medium, and processor 310 may run at least one (e.g., one or more) computer program modules to implement various functions of electronic device 300. Various application programs and various data, as well as various data used and / or generated by the application programs, may also be stored in the computer-readable storage medium.

[0112] It should be noted that, in the embodiments of this disclosure, the specific functions and technical effects of the electronic device 300 can be referred to the description of the scheduling method above, and will not be repeated here.

[0113] Figure 10This is a schematic block diagram of another electronic device provided in some embodiments of this disclosure. The electronic device 400 is, for example, suitable for implementing the scheduling method provided in the embodiments of this disclosure. The electronic device 400 may be a terminal device, etc., and can be used to implement a client or server. The electronic device 400 may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, etc., as well as fixed terminals such as digital TVs, desktop computers, smart home devices, etc. It should be noted that... Figure 10 The illustrated electronic device 400 is merely an example and does not impose any limitation on the functionality and scope of use of the embodiments of this disclosure.

[0114] like Figure 10 As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 410, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 420 or a program loaded from storage device 480 into random access memory (RAM) 430. The RAM 430 also stores various programs and data required for the operation of electronic device 400. The processing device 410, ROM 420, and RAM 430 are interconnected via bus 440. Input / output (I / O) interface 450 is also connected to bus 440.

[0115] Typically, the following devices can be connected to I / O interface 450: input devices 460 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 470 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 480 including, for example, magnetic tapes, hard disks, etc.; and communication devices 490. Communication device 490 allows electronic device 400 to communicate wirelessly or wiredly with other electronic devices to exchange data. Although Figure 10 An electronic device 400 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown, and the electronic device 400 may alternatively implement or have more or fewer devices.

[0116] For example, according to embodiments of this disclosure, the scheduling method described above can be implemented as a computer software program. For instance, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program including program code for executing the scheduling method described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 490, or installed from a storage device 480, or installed from a ROM 420. When the computer program is executed by the processing device 410, the functions defined in the scheduling method provided by embodiments of this disclosure can be implemented.

[0117] At least one embodiment of this disclosure also provides a storage medium. Using this storage medium, the utilization rate of matrix operation units can be improved, the computing power of the matrix operation units can be effectively utilized, the convolution operation time can be shortened, the operation efficiency can be improved, and data transmission time can be saved.

[0118] Figure 11 This is a schematic diagram of a storage medium provided for some embodiments of this disclosure. For example, such as... Figure 11 As shown, the storage medium 500 can be a non-transitory computer-readable storage medium storing non-transitory computer-readable instructions 510. When the non-transitory computer-readable instructions 510 are executed by the processor, the scheduling method described in the embodiments of this disclosure can be implemented. For example, when the non-transitory computer-readable instructions 510 are executed by the processor, one or more steps in the scheduling method described above can be performed.

[0119] For example, the storage medium 500 can be used in the aforementioned electronic device. For instance, the storage medium 500 may include the memory 320 in the electronic device 300.

[0120] For example, the storage medium may include a memory card for a smartphone, a storage component for a tablet computer, a hard disk for a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, or any combination of the above storage media, or other suitable storage media.

[0121] For example, the description of storage medium 500 can be found in the description of memory in the embodiments of the electronic device, and will not be repeated here. The specific functions and technical effects of storage medium 500 can be found in the description of the scheduling method above, and will not be repeated here.

[0122] In the above text, referring to Figures 1 to 12... Figure 11This disclosure describes a scheduling method, scheduling device, electronic device, and storage medium provided in embodiments of this disclosure. The scheduling method provided in these embodiments can reduce the computational load of repetitive data in convolutional neural networks and improve the computational efficiency of multi-layer convolutional neural networks by using the copying and transmission of valid data between computing units as filler rows.

[0123] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, capable of transmitting, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0124] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0125] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0126] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0128] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0129] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0130] According to one or more embodiments of this disclosure, Example 1 provides a scheduling method for a multi-layer convolutional neural network, comprising: multiple computing units performing a first convolution calculation on corresponding multiple data groups to obtain corresponding multiple first calculation result groups, wherein the multiple first calculation result groups are used to constitute a first convolutional layer obtained by the first convolution calculation, and the multiple computing units include a first computing unit and a second computing unit; determining a data copying and transmission mode corresponding to the multiple first calculation result groups on the multiple computing units according to the configuration rules of the second convolutional layer obtained by the multiple computing units after performing a second convolution calculation on the first convolutional layer on the multiple computing units; and a first computing unit that needs to perform effective data row filling on the multiple computing units, based on the corresponding data copying and transmission mode, obtaining a first intermediate data row required by the first computing unit for filling during the second convolution calculation process from the first calculation result group on the second computing unit.

[0131] Example 2: According to the scheduling method described in Example 1, the data rows in the plurality of first calculation result groups are continuous without overlap in the first convolutional layer.

[0132] Example 3: According to the scheduling method described in Example 1, the first calculation result group on the first computing unit and the first intermediate data row obtained from the second computing unit constitute the data group required for the first computing unit to perform the second convolution calculation.

[0133] Example 4: According to the scheduling method described in Example 1, the multiple data sets used by the multiple computing units in performing the first convolution calculation are multiple original input data sets.

[0134] Example 5: The scheduling method according to Example 4 further includes:

[0135] The original input matrix is ​​split to obtain multiple original input data groups, and the multiple original input data groups are respectively transmitted to the multiple computing units to perform the first convolution calculation.

[0136] Example 6. According to the scheduling method described in Example 1, wherein determining the data replication and transmission mode corresponding to the plurality of first calculation result groups on the plurality of computing units based on the configuration rules of the second convolutional layer obtained by the plurality of computing units performing a second convolution calculation on the first convolutional layer in the plurality of computing units includes:

[0137] Based on the memory size of the memory on the plurality of computing units, the configuration rules of the second convolutional layer in the plurality of computing units are determined;

[0138] Based on the configuration rules, the distribution of the second convolutional layer across the plurality of computing units is determined;

[0139] Based on the distribution, the data replication and transmission mode corresponding to the first calculation result group on the plurality of computing units is determined.

[0140] Example 7. According to the scheduling method described in Example 1, the step of obtaining the first intermediate data row required for filling by the first computing unit during the second convolution calculation process from the first calculation result group on the second computing unit includes:

[0141] Based on the relationship between the second calculation result group obtained after the second convolution calculation performed by the first calculation unit and the first calculation result group obtained after the first convolution calculation performed by the first calculation unit, the size of the first intermediate data row is determined.

[0142] Based on the data copying and transmission mode and the size of the first intermediate data row, the first intermediate data row is obtained from the first calculation result group on the second computing unit.

[0143] Example 8: According to the scheduling method described in Example 7, the size of the second calculation result group obtained by the first computing unit is the same as the size of the first calculation result group on the first computing unit.

[0144] Example 9. The scheduling method according to Example 7 further includes:

[0145] In response to the second convolutional layer being the output layer in the multi-layer convolutional neural network, multiple second computation result groups from the plurality of computational units are aggregated to obtain at least a portion of the computational output of the multi-layer convolutional neural network.

[0146] Example 10. The scheduling method according to any one of Examples 1-9 above further includes:

[0147] Based on the corresponding data copying and transmission mode, the second computing unit obtains the second intermediate data row required for filling during the second convolution calculation process from the first calculation result group on the first computing unit.

[0148] Example 11: According to the scheduling method described in Example 10, the first computing unit and the second computing unit obtain the same size of the first intermediate data line and the second intermediate data line from each other.

[0149] According to one or more embodiments of this disclosure, Example 12 provides a scheduling apparatus for a multi-layer convolutional neural network, comprising:

[0150] The calculation control module is configured to enable multiple calculation units to perform a first convolution calculation on corresponding multiple data groups to obtain corresponding multiple first calculation result groups, wherein the multiple first calculation result groups are used to form a first convolution layer obtained by the first convolution calculation, and the multiple calculation units include a first calculation unit and a second calculation unit.

[0151] The allocation and scheduling module is configured to determine the data copying and transmission mode corresponding to the multiple first calculation result groups on the multiple computing units according to the configuration rules of the second convolutional layer obtained after the multiple computing units perform a second convolution calculation on the first convolutional layer in the multiple computing units;

[0152] The data transmission module is configured such that the first computing unit, which requires effective data row filling among multiple computing units, obtains the first intermediate data row required for filling during the second convolution calculation process from the first calculation result group on the second computing unit based on the corresponding data copy transmission mode.

[0153] Example 13: The scheduling apparatus according to Example 12, wherein the data rows in the plurality of first calculation result groups are continuous without overlap in the first convolutional layer.

[0154] Example 14: The scheduling apparatus according to Example 12, wherein the first calculation result group on the first computing unit and the first intermediate data row obtained from the second computing unit constitute the data group required for the first computing unit to perform the second convolution calculation.

[0155] Example 15: The scheduling device according to Example 12, wherein the plurality of data groups used by the plurality of computing units in performing the first convolution calculation are plurality of original input data groups.

[0156] Example 16. The scheduling device according to Example 15 further includes:

[0157] The data splitting module is configured to split the original input matrix to obtain multiple original input data groups, and transmit the multiple original input data groups to the multiple computing units respectively for the first convolution calculation.

[0158] Example 17. The scheduling device according to Example 12, wherein determining the data copying and transmission mode corresponding to the plurality of first calculation result groups on the plurality of computing units according to the configuration rules of the second convolutional layer obtained by the plurality of computing units performing a second convolution calculation on the first convolutional layer in the plurality of computing units includes:

[0159] Based on the memory size of the memory on the plurality of computing units, the configuration rules of the second convolutional layer in the plurality of computing units are determined;

[0160] Based on the configuration rules, the distribution of the second convolutional layer across the plurality of computing units is determined;

[0161] Based on the distribution, the data replication and transmission mode corresponding to the first calculation result group on the plurality of computing units is determined.

[0162] Example 18. The scheduling apparatus according to Example 12, wherein obtaining the first intermediate data row required for filling by the first computing unit during the second convolution calculation process from the first calculation result group on the second computing unit includes:

[0163] Based on the relationship between the second calculation result group obtained after the second convolution calculation performed by the first calculation unit and the first calculation result group obtained after the first convolution calculation performed by the first calculation unit, the size of the first intermediate data row is determined.

[0164] Based on the data copying and transmission mode and the size of the first intermediate data row, the first intermediate data row is obtained from the first calculation result group on the second computing unit.

[0165] Example 19. The scheduling apparatus according to Example 18, wherein the size of the second calculation result group obtained by the first computing unit is the same as the size of the first calculation result group on the first computing unit.

[0166] Example 20: The scheduling device according to Example 18 further includes:

[0167] The data output module is configured to, in response to the second convolutional layer being the output layer in the multi-layer convolutional neural network, aggregate multiple second computation result groups from the multiple computation units to obtain at least a portion of the computation output of the multi-layer convolutional neural network.

[0168] Example 21. The scheduling apparatus according to any one of Examples 12-20, wherein the data transmission module is further configured as follows:

[0169] The second computing unit, based on the corresponding data copying and transmission mode, obtains the second intermediate data row required for filling during the second convolution calculation process from the first calculation result group on the first computing unit.

[0170] Example 22: The scheduling apparatus according to Example 10, wherein the first computing unit and the second computing unit obtain the same size of the first intermediate data line and the second intermediate data line from each other.

[0171] According to one or more embodiments of this disclosure, Example 23 provides an electronic device including the scheduling device described in any one of Examples 12-22 above.

[0172] According to one or more embodiments of this disclosure, Example 24 provides an electronic device including: a processor; a memory including at least one computer program module; wherein the at least one computer program module is stored in the memory and configured to be executed by the processor, the at least one computer program module being used to implement the scheduling method described in any one of Examples 1-11 above.

[0173] According to one or more embodiments of this disclosure, Example 25 provides a storage medium storing non-transitory computer-readable instructions that, when executed by a computer, implement the scheduling method described in any one of Examples 1-11 above.

[0174] Although the present disclosure has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to the embodiments of the present disclosure, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present disclosure are within the scope of protection claimed by the present disclosure.

[0175] The following points should be noted regarding this disclosure:

[0176] (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure. Other structures can be referred to the general design.

[0177] (2) For clarity, the thickness of layers or regions in the drawings used to describe embodiments of the present disclosure is enlarged or reduced, i.e., these drawings are not drawn to actual scale.

[0178] (3) Where there is no conflict, the embodiments of this disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0179] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. The scope of protection of this disclosure should be determined by the scope of protection of the claims.

Claims

1. A scheduling method for multi-layer convolutional neural networks, comprising: Multiple computing units within the chip perform a first convolution calculation on corresponding multiple data groups to obtain corresponding multiple first calculation result groups, wherein the multiple first calculation result groups are used to form a first convolution layer obtained by the first convolution calculation, and the multiple computing units include a first computing unit and a second computing unit. According to the configuration rules of the second convolutional layer in the plurality of computing units, the data copying and transmission mode corresponding to the plurality of first calculation result groups on the plurality of computing units is determined, wherein the second convolutional layer is obtained by the plurality of computing units performing a second convolution calculation on the first convolutional layer; The first computing unit, which requires effective data row filling among multiple computing units, obtains the first intermediate data row required for filling during the second convolution calculation process from the first calculation result group on the second computing unit, based on the corresponding data copying and transmission mode. The step of obtaining the first intermediate data row required for filling during the second convolution calculation process by the first calculation result group on the second calculation unit includes: Based on the relationship between the second calculation result group obtained after the second convolution calculation performed by the first calculation unit and the first calculation result group obtained after the first convolution calculation performed by the first calculation unit, the size of the first intermediate data row is determined. Based on the data copying and transmission mode and the size of the first intermediate data row, the first intermediate data row is obtained from the first calculation result group on the second computing unit.

2. The scheduling method of claim 1, wherein, The data rows in the plurality of first calculation result groups are continuous without overlap in the first convolutional layer.

3. The scheduling method of claim 1, wherein, The first calculation result group on the first computing unit and the first intermediate data row obtained from the second computing unit constitute the data group required for the first computing unit to perform the second convolution calculation.

4. The scheduling method according to claim 1, wherein, The multiple data sets used by the multiple computing units in performing the first convolution calculation are multiple original input data sets.

5. The scheduling method according to claim 4 further includes: The original input matrix is ​​split to obtain multiple original input data groups, and the multiple original input data groups are respectively transmitted to the multiple computing units to perform the first convolution calculation.

6. The scheduling method according to claim 1, wherein, The step of determining the data copying and transmission mode corresponding to the plurality of first computation result groups on the plurality of computation units according to the configuration rules of the second convolutional layer in the plurality of computation units includes: Based on the memory size of the memory on the plurality of computing units, the configuration rules of the second convolutional layer in the plurality of computing units are determined; Based on the configuration rules, the distribution of the second convolutional layer across the plurality of computing units is determined; Based on the distribution, the data replication and transmission mode corresponding to the first calculation result group on the plurality of computing units is determined.

7. The scheduling method according to claim 6, wherein, The size of the second calculation result group obtained by the first calculation unit is the same as the size of the first calculation result group on the first calculation unit.

8. The scheduling method according to claim 6 further includes: In response to the second convolutional layer being the output layer in the multi-layer convolutional neural network, multiple second computation result groups from the plurality of computational units are aggregated to obtain at least a portion of the computational output of the multi-layer convolutional neural network.

9. The scheduling method according to any one of claims 1-8, further comprising: Based on the corresponding data copying and transmission mode, the second computing unit obtains the second intermediate data row required for filling during the second convolution calculation process from the first calculation result group on the first computing unit.

10. The scheduling method according to claim 9, wherein, The first calculation unit and the second calculation unit obtain the same size of the first intermediate data row and the second intermediate data row from each other.

11. A scheduling device for a multi-layer convolutional neural network, comprising: The calculation control module is configured to enable multiple calculation units within the chip to perform a first convolution calculation on corresponding multiple data groups to obtain corresponding multiple first calculation result groups, wherein the multiple first calculation result groups are used to form a first convolution layer obtained by the first convolution calculation, and the multiple calculation units include a first calculation unit and a second calculation unit. The allocation and scheduling module is configured to determine the data copying and transmission mode corresponding to the multiple first calculation result groups on the multiple computing units according to the configuration rules of the second convolutional layer in the multiple computing units, wherein the second convolutional layer is obtained by the multiple computing units performing a second convolution calculation on the first convolutional layer; The data transmission module is configured such that, based on the corresponding data copying and transmission mode, the first computing unit, which requires effective data row filling among multiple computing units, obtains the first intermediate data row required for filling during the second convolution calculation process from the first calculation result group on the second computing unit. The step of obtaining the first intermediate data row required for filling during the second convolution calculation process by the first calculation result group on the second calculation unit includes: Based on the relationship between the second calculation result group obtained after the second convolution calculation performed by the first calculation unit and the first calculation result group obtained after the first convolution calculation performed by the first calculation unit, the size of the first intermediate data row is determined. Based on the data copying and transmission mode and the size of the first intermediate data row, the first intermediate data row is obtained from the first calculation result group on the second computing unit.

12. An electronic device comprising the scheduling device of claim 11.

13. An electronic device, comprising: processor; The memory includes at least one computer program module; The at least one computer program module is stored in the memory and configured to be executed by the processor, the at least one computer program module being used to implement the scheduling method according to any one of claims 1-10.

14. A storage medium storing non-transitory computer-readable instructions that, when executed by a computer, implement the scheduling method according to any one of claims 1-10.