Scheduling method, scheduling device, electronic device, and storage medium
The scheduling method optimizes AI chip performance by balancing data distribution and reducing redundant calculations through data duplication and transmission modes, addressing inefficiencies in multi-core AI chip architectures.
Patent Information
- Application Number
- JP2025517927
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-27
- Filing Date
- 2023-09-18
- Publication Date
- 2025-10-22
- Estimated Expiration
- 2043-09-18
AI Technical Summary
Current AI chip architectures face challenges in efficiently scheduling multi-core computations for deep neural networks, particularly in handling dynamic computation graphs and reducing redundant data calculations in multi-layer convolutional neural networks.
A scheduling method and device that involves data duplication and transmission modes among computing units to optimize the utilization of computing power, reducing redundant calculations by balancing the distribution and padding of data across multiple computing units.
The method improves computing efficiency by minimizing redundant calculations and optimizing resource utilization in AI chips, particularly in multi-layer convolutional neural networks.
Smart Images

Figure 2025534993000001_ABST
Abstract
Description
[Technical Field]
[0001] This application claims priority to Chinese Patent Application No. 202211188739.0, filed on September 27, 2022, the entire contents of which are hereby incorporated by reference into this application.
[0002] SUMMARY OF THE INVENTION An embodiment of the present disclosure relates to a scheduling method, a scheduling device, an electronic device, and a storage medium. [Background technology]
[0003] Artificial intelligence (AI) chips are chips specially designed to perform neural network calculations and accelerate the execution of neural networks. With the development of artificial intelligence (AI), the number of parameters in algorithm models has increased dramatically, resulting in an ever-increasing need for computing power.
[0004] Traditional hardware architectures (e.g., central processing units (CPUs)) are designed to balance different service needs, limiting the computing power they can provide for AI applications. Considering high computing power, current AI chips not only use general-purpose graphics processing units (GPUs), but also widely use domain-specific accelerators (DSAs) with homogeneous multi-core architectures. Summary of the Invention [Means for solving the problem]
[0005] At least one embodiment of the present disclosure provides a scheduling method for a multi-layer convolutional neural network, the scheduling method including: a plurality of computing units performing first convolutional calculations on corresponding plurality of data sets to obtain corresponding first computation result sets, the plurality of first computation result sets being for constituting a first convolutional layer obtained by the first convolutional calculations, the plurality of computing units including a first computing unit and a second computing unit; determining data duplication and transmission modes corresponding to the plurality of first computation result sets in the plurality of computing units according to an arrangement rule of the plurality of computing units of a second convolutional layer obtained by the plurality of computing units performing second convolutional calculations on the first convolutional layer; and obtaining a first intermediate data row required for padding in the second convolutional calculation process by the first computing unit from the first computation result set in the second computing unit based on the data duplication and transmission mode corresponding to a first computing unit among the plurality of computing units that is to pad a valid data row.
[0006] At least one embodiment of the present disclosure further provides a scheduling device, the scheduling device including: a computation control module configured to cause a plurality of computing units to perform first convolutional calculations on corresponding plurality of data sets, respectively, to obtain corresponding first computation result sets, the plurality of first computation result sets being for constituting a first convolutional layer obtained by the first convolutional calculation, the plurality of computing units including a first computing unit and a second computing unit; an allocation scheduling module configured to determine data duplication and transmission modes corresponding to the plurality of first computation result sets in the plurality of computing units according to an arrangement rule in the plurality of computing units of a second convolutional layer obtained by the plurality of computing units performing second convolutional calculations on the first convolutional layer; and a data transmission module configured to obtain a first intermediate data row required for padding in the second convolutional calculation process by the first computing unit from the first computation result set in the second computing unit, based on the data duplication and transmission mode corresponding to a first computing unit among the plurality of computing units that should pad a valid data row.
[0007] At least one embodiment of the present disclosure further provides an electronic device, comprising the scheduling device according to any one embodiment of the present disclosure.
[0008] At least one embodiment of the present disclosure further provides an electronic device, the electronic device comprising a processor and a memory including at least one computer program module, the at least one computer program module being stored in the memory and configured to be executed by the processor, the at least one computer program module being for implementing the scheduling method described in any one embodiment of the present disclosure.
[0009] At least one embodiment of the present disclosure further provides a storage medium, on which non-transitory computer-readable instructions are stored, and which, when executed by a computer, realizes the scheduling method described in any one embodiment of the present disclosure.
[0010] In order to more clearly describe the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments are briefly described below. Obviously, the drawings described below only relate to some embodiments of the present disclosure, and do not limit the present disclosure. [Brief explanation of the drawings]
[0011] [Figure 1A] Figure 1A is a schematic diagram of the model parallelism. [Figure 1B] Figure 1B is a schematic diagram of data parallelism. [Figure 2] Figure 2 is a schematic diagram of the convolution calculation process based on data parallelism shown in the calculation graph. [Figure 3] FIG. 3 is a schematic flow chart of a scheduling method in accordance with at least one embodiment of the present disclosure. [Figure 4] FIG. 4 is a schematic flowchart of step S20 in FIG. [Figure 5] FIG. 5 is a schematic flowchart of step S30 in FIG. [Figure 6A] FIG. 6A is a schematic diagram of a data replication and transmission process in accordance with at least one embodiment of the present disclosure. [Figure 6B] FIG. 6B is a schematic diagram of a data replication and transmission process in accordance with at least one embodiment of the present disclosure. [Figure 6C] FIG. 6C is a schematic diagram of a data replication and transmission process in accordance with at least one embodiment of the present disclosure. [Figure 6D] FIG. 6D is a schematic diagram of a data replication and transmission process in accordance with at least one embodiment of the present disclosure. [Figure 7] FIG. 7 is a schematic block diagram of a scheduling device in accordance with at least one embodiment of the present disclosure. [Figure 8] FIG. 8 is a schematic block diagram of an electronic device in accordance with at least one embodiment of the present disclosure. [Figure 9] FIG. 9 is a schematic block diagram of another electronic device in accordance with at least one embodiment of the present disclosure. [Figure 10] FIG. 10 is a schematic block diagram of another electronic device in accordance with at least one embodiment of the present disclosure. [Figure 11] FIG. 11 is a schematic diagram of a storage medium in accordance with at least one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0012] In order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, but not all of the embodiments. Based on the described embodiments of the present disclosure, any other embodiments that a person skilled in the art can obtain without inventive efforts fall within the scope of protection of the present disclosure.
[0013] Unless otherwise defined, technical or scientific terms used in this disclosure should have the ordinary meaning understood by those skilled in the art to which this disclosure belongs. The terms "first," "second," and similar terms used in this disclosure do not denote any order, number, or importance, but are merely used to distinguish different components. Similar terms such as "comprise" or "include" refer to the elements or components described before the term covering the elements or components listed thereafter and their equivalents, without excluding other elements or components. Similar terms such as "connect" or "coupled" are not limited to physical or mechanical connections, but also include electrical connections, whether direct or indirect. Terms such as "top," "bottom," "left," "right," and the like only refer to relative positions, and if the absolute positions of the described objects change, the relative positions may change accordingly.
[0014] The present disclosure will be described below by means of several specific embodiments. In order to maintain the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and known components (elements) may be omitted. When any one component (element) of the embodiments of the present disclosure appears in more than one drawing, the component (element) is indicated by the same or similar symbol in each drawing.
[0015] The core concept of domain-specific accelerators (DSAs, or Domain Specific ASICs) is to use dedicated hardware to perform specialized tasks. For example, DSAs satisfy applications in a single domain rather than a fixed set of applications. Therefore, DSAs can balance flexibility and specialization. A DSA accelerator refers to an ASIC that interconnects multiple computing units (also called processing units (PEs, Process Elements) or processing cores) on a single chip to expand computing power and accelerate calculations. The multiple processing units are connected to each other, for example, via communication lines, which may be buses, on-chip networks, etc. A major challenge facing AI compilers is how to effectively schedule hardware resources in such an AI accelerator with multiple interconnected PEs to efficiently complete inference of deep neural network models.
[0016] Currently, there are two ways to implement multi-core compilation scheduling:
[0017] One approach is model parallelism, where different devices are responsible for computing different parts of the computation graph. For example, in some domain-specific accelerators, each operator in the computation graph is statically assigned to a different PE within the chip. Then, input data is input to the PE where the first operator is located. After the PE where the first operator is located completes its calculation, the calculated data is sent to the next PE, and so on until all operators are calculated. In model parallelism, each PE only executes operators in its local internal memory; each PE cannot continue computing until the calculation of the previous operator is completed. This is a pipelined approach to sequentially execute operations.
[0018] For example, as shown in FIG. 1A, three different operators are loaded into PE0, PE1, and PE2, respectively. For example, the three operators are convolution operators corresponding to different layers of a neural network. For example, PE0 has convolution operator A, PE1 has convolution operator B, and PE2 has convolution operator C. During calculation, the server first loads input data into PE0, where convolution operator A is located. After PE0 completes the first convolution calculation on the input data, it transmits the result of the first convolution calculation to PE1, and then PE1 performs the second convolution calculation corresponding to convolution operator B. By analogy, after PE2 completes the execution of convolution operator C, the server returns the final result obtained by PE2 to the host.
[0019] However, model parallelism requires static compilation, and some operators in the computation graph must be assigned to a fixed number of PEs, so it cannot process computation graphs that include dynamic shapes. Dynamic shapes mean that the shape of a tensor depends on the specific operation and cannot be calculated in advance; that is, the dynamic computation graph is constructed through the calculation of each step. Therefore, model parallelism, which only includes the computation graph portion, cannot process dynamic computation graphs.
[0020] The other method is data parallelism, which means that each device has a complete computation graph. For example, in some GPUs, each operator in the computation graph is first loaded sequentially into the GPU device, and each operator is placed in parallel on multiple PEs. Then, the input data is decomposed into multiple parts and loaded into multiple PEs respectively. The multiple PEs process the input data in parallel. After all the operators are calculated, the computation results of the multiple PEs are aggregated and sent back to the host.
[0021] For example, as shown in FIG. 1B, PE0, PE1, and PE2 are each loaded with three different operators, each of which is a convolution operator corresponding to a different layer of a neural network. For example, PE0, PE1, and PE2 each have convolution operators A, B, and C. During calculation, the server first decomposes the input data into multiple parts, for example, uniformly decomposing it into three parts of input data, and then loads the three parts of input data into PE0, PE1, and PE2, respectively. PE0, PE1, and PE2 then perform multiple convolution calculations corresponding to convolution operators A, B, and C on the input data in parallel and sequentially. After the calculation is completed, the system collects the final results from PE0, PE1, and PE2 and sends them back to the host.
[0022] Data parallelism does not yet have a consistent implementation mode, and scheduling optimization is required for different hardware characteristics. When deploying actual models, some models have large operator data volumes, while others have small operator data volumes. Therefore, it is necessary to implement different deployment schemes for AI chips based on the model to achieve optimal inference performance.
[0023] In current data parallel methods, when the amount of input data is large, it is not possible to store all of the input data required for calculation inside the PE at once. Therefore, the input data is generally divided into multiple groups, and then each group of data is placed in a different PE, and the multiple PEs perform parallel calculations. After the calculation is completed, the calculation results from the multiple PEs are collected to obtain the final output result corresponding to the input data.
[0024] However, in the inference calculation process of a multi-layer neural network, if a large picture or a large amount of data is directly decomposed into multiple sets in consideration of the padding operation of the convolution calculation, it is easy to introduce the problem of repeated calculation of data.
[0025] For example, one exemplary computation graph is: %5:[(1 224 224 64),F,FP32]=conv(%0:[(1 224 224 64),F,FP32], 1%:[(64 3 3 64),W,FP32],kh=3,kw=3,pad_h_top=1,pad_h_bottom=1,pad_w_left=1,pad_w_right=1,stride_h=1,stride_w=1), %10:[(1 224 224 64),F,FP32]=conv(%5:[(1 224 224 64),F,FP32], 6%:[(64 3 3 64),W,FP32],kh=3,kw=3,pad_h_top=1,pad_h_bottom=1,pad_w_left=1,pad_w_right=1,stride_h=1,stride_w=1), and It contains three convolution operators (conv): %15:[(1 224 224 64),F,FP32]=conv(%10:[(1 224 224 64),F,FP32], 11%:[(64 3 3 64),W,FP32],kh=3,kw=3,pad_h_top=1,pad_h_bottom=1,pad_w_left=1,pad_w_right=1,stride_h=1,stride_w=1).
[0026] %0, %5, %10, and %15 indicate tensor data. For example, the input image of the first convolution layer of a neural network may be represented as [B, H, W, C], i.e., the dimensions of the input data are B × H × W × C. For example, %0 indicates that the size of the input image of the first convolution layer of a neural network is [1, 224, 224, 64], i.e., the input image has 64 channels, and the image dimensions (height H × width W) of each channel are 224 × 224, totaling one batch. For example, the size of the convolution kernel used by each convolution operator conv may be represented as kh × kw, i.e., kh = 3 and kw = 3 indicate that the size of the convolution kernel is 3 × 3. For example, the padding operation of each convolution operator conv is indicated by pad, where, for example, pad_h_top=1 indicates padding one row at the top of the picture in the direction of the height h of the picture, for example, padding with 0, pad_h_bottom=1 indicates padding one row at the bottom of the picture in the direction of the height h of the picture, for example, padding with 0, pad_w_left=1 indicates padding one column at the left side of the picture in the direction of the width w of the picture, for example, padding with 0, and pad_w_right=1 indicates padding one column at the right side of the picture in the direction of the width w of the picture, for example, padding with 0. For example, the sliding step width of the convolution kernel in each convolution operator conv is represented by stride, where stride_h=1 indicates that the convolution kernel moves by one pixel point each time it moves along the height h of the picture, and stride_w=1 indicates that the convolution kernel moves by one pixel point each time it moves along the width w of the picture. The calculation uses floating-point calculation and has a precision of FP32.
[0027] For example, Figure 2 is a schematic diagram showing a convolution calculation process using data parallelism shown in a calculation graph, which includes the three convolution operators described above. For example, if all the data required for calculation cannot be stored in the internal memory of a single PE at once, it is necessary to decompose the input data into multiple data sets based on the internal memory and number of PEs. For example, in Figure 2, there are three PEs and the input data size is 224 x 224, so the input data is decomposed into three data sets and loaded into three PEs respectively to perform the convolution calculation.
[0028] The decomposition of input data and the size and range of the decomposed dataset may be selected based on the output data. For example, in the computation graph shown in FIG. 2, if the size of the output data 15 is 224 × 224 and the data is selected to be decomposed approximately uniformly into three parts in the height dimension H of the output data, the (height) sizes of the three parts of the output dataset in three rounds, round 1, round 2, and round 3, corresponding to three PEs, are 75 rows, 75 rows, and 74 rows, respectively. When dividing the input data, the output data of each part must be obtained by deep traversing the entire dataset first, performing forward recursion to obtain the range of the input data for each part.
[0029] It should be noted that all data displayed in the computation graph shown in FIG. 2 is "valid data." For simplicity, "valid data" here refers to the actual valid data among the data obtained each time a convolution calculation is performed. During actual calculation, a padding operation is required to maintain the image size unchanged after the convolution calculation. Therefore, the area for performing the convolution kernel and convolution calculation, which consists of data at the "edges" of each decomposed data set and padding "0," is an invalid area. The results obtained by performing the convolution calculation in this area are invalid data and must be truncated. That is, since the output data obtained by performing a convolution calculation on the complete input data is valid data, only the data obtained by performing a convolution calculation on the data in the valid area is valid data when corresponding to the decomposed data set. For clarity of illustration and simplicity of explanation, all data shown in the computation graph in the drawings of this application is valid data.
[0030] For example, in Figure 2, round1 shows that the output data %15 has 75 rows, i.e., [0,74] rows of valid data. Because the padding operations are pad_h_top=1 and pad_h_bottom=1, the intermediate data %10 obtained by recursing forward from %15 must have 76 rows, i.e., [0,75] rows of valid data, and the intermediate data %5 obtained by recursing forward from intermediate data %10 must have 77 rows, i.e., [0,76] rows of valid data. By analogy, the final input data %0 must have 78 rows of valid data. Therefore, the range of read data in the height direction of the input data %0 obtained in round1 is [0,77] (row numbers start from 0), for a total of 78 rows.
[0031] Similarly, in Figure 2, round2 shows that the output data %15 has 75 rows, or [75,149] rows of valid data. Because the padding operations are pad_h_top=1 and pad_h_bottom=1, the intermediate data %10 obtained by recursing forward from %15 must have 77 rows, or [74,150] rows of valid data, and the intermediate data %5 obtained by recursing forward from intermediate data %10 must have 79 rows, or [73,151] rows of valid data. By analogy, the final input data %0 must have 81 rows of valid data. Therefore, the range of read data in the height direction for the input data %0 obtained in round2 is [72,152], for a total of 81 rows.
[0032] Similarly, in Figure 2, round3 shows that the output data %15 has 74 rows, i.e., [150, 223] rows of valid data. Because the padding operations (pad_h_top=1, pad_h_bottom=1) are used, the intermediate data %10 obtained by recursing forward from %15 must have 75 rows, i.e., [149, 223] rows of valid data. The intermediate data %5 obtained by recursing forward from intermediate data %10 must have 76 rows, i.e., [148, 223] rows of valid data. By analogy, the final input data %0 obtained must have 77 rows of valid data. Therefore, the range of read data in the height direction for the input data %0 obtained in round3 is [147, 223], for a total of 77 rows.
[0033] It is clear from the computation graph in Figure 2 that the iterative calculation area for round1 and round2 is [72,77] rows, i.e., 6 rows are repeated, and the iterative calculation area for round2 and round3 is [147,152] rows, i.e., 6 rows are repeated. As can be seen from this, the deeper the convolutional network, the more forward recursions there are, and the more valid data is required for the edge of the dataset due to padding on the previous layer, i.e., the larger the area to be iteratively calculated, which results in a large amount of redundant overhead of computing power.
[0034] At least one embodiment of the present disclosure provides a scheduling method for a multi-layer convolutional neural network, the scheduling method including: a plurality of computing units each performing a first convolutional calculation on a corresponding plurality of data sets to obtain a corresponding plurality of first calculation result sets, the plurality of first calculation result sets being for constituting a first convolutional layer obtained by the first convolutional calculation, the plurality of computing units including a first computing unit and a second computing unit; determining data duplication and transmission modes corresponding to the plurality of first calculation result sets in the plurality of computing units according to an arrangement rule for the plurality of computing units of a second convolutional layer obtained by the plurality of computing units performing a second convolutional calculation on the first convolutional layer; and obtaining a first intermediate data row required for padding in the second convolutional calculation process by the first computing unit from the first calculation result set in the second computing unit based on the data duplication and transmission mode corresponding to the first computing unit among the plurality of computing units that is to pad a valid data row. The scheduling method can achieve balanced utilization of the computing power of a computing unit by duplicating and transmitting valid data in the computing unit to another computing unit, thereby reducing repeated calculation of data and improving utilization of computing power.
[0035] At least one embodiment of the present disclosure further provides a scheduling device, an electronic device, and a storage medium, which can similarly duplicate and transmit valid data in a computing unit to another computing unit, thereby realizing balanced utilization of the computing capabilities of the computing units, thereby reducing repeated calculation of data and improving utilization of the computing capabilities.
[0036] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings, in which it should be noted that the same reference numerals in different drawings represent the same elements.
[0037] 3 is a schematic flowchart of a scheduling method according to at least one embodiment of the present disclosure. As shown in FIG. 3, the scheduling method includes steps S10 to S30.
[0038] Step S10: A plurality of calculation units perform first convolution calculations on corresponding plurality of data sets to obtain corresponding plurality of first calculation result sets, and the plurality of first calculation result sets are for constituting a first convolution layer obtained by the first convolution calculations, and the plurality of calculation units include a first calculation unit and a second calculation unit; Step S20: Determine data replication and transmission modes corresponding to the first calculation result sets in the plurality of calculation units according to the allocation rules of the plurality of calculation units in the second convolution layer obtained by the plurality of calculation units performing second convolution calculations on the first convolution layer; Step S30: According to a data duplication transmission mode corresponding to a first computing unit among the plurality of computing units that is to pad valid data rows, obtain a first intermediate data row required for padding in the second convolution calculation process by the first computing unit from the first calculation result set in the second computing unit.
[0039] For example, the scheduling method may be used in a computing device such as an AI chip that performs convolutional operations on a multi-layer convolutional neural network, a general-purpose graphics processing unit (GPGPU), etc. For example, the AI chip may be an AI chip that uses a DSA accelerator or an AI chip that has multiple PE units, and may be used for, for example, data-parallel distributed training, and the embodiments of the present disclosure are not limited thereto.
[0040] It should be noted that in at least one embodiment of the present disclosure, a "first convolution calculation" may refer to a single convolution calculation performed on a complete input image, or may refer to a single convolution calculation performed on a current input data set in a computation unit. The "first convolution calculation" may not refer to a convolution calculation performed only on the first layer of convolution, but may refer to any one of multiple layers of convolution calculations. For example, the "first convolution calculation" may refer to a single convolution calculation performed on the first layer of convolution, or may refer to a single convolution calculation performed on the second layer, third layer, etc.
[0041] In at least one embodiment of the present disclosure, the "second convolution calculation" refers to the next convolution calculation corresponding to the "first convolution calculation," i.e., a single convolution calculation performed using the "first convolution layer" obtained after the first convolution calculation as input. The "first convolution layer" refers to a convolution layer formed from the calculation result obtained by the "first convolution calculation," where the "first convolution layer" refers to an actual convolution layer obtained by performing the first convolution calculation on a complete, undecomposed input image. For example, the first convolution layer may be input data for the second convolution calculation. Similarly, the "second convolution layer" refers to a convolution layer formed from the calculation result obtained by the "second convolution calculation," where the "second convolution layer" refers to an actual convolution layer obtained by performing the second convolution calculation on the first convolution layer formed from the obtained undecomposed data.
[0042] Hereinafter, a scheduling method according to an embodiment of the present disclosure will be described in detail with reference to the computation graphs shown in Figures 6A to 6D. In the following exemplary description, three PEs (Figures 6A to 6C) will be used as an example, but the embodiment of the present disclosure is not limited thereto, and the computing device may include, for example, two PEs, four PEs (Figure 6D), or more PEs.
[0043] 6A to 6D abstractly represent the range of at least a portion of valid data in a certain dimension (e.g., height H or width W) of the input image to be convolved each time. The diagonal line region, grid line region, and diamond line region abstractly represent data sets in the respective computation units (PE0, PE1, PE2, etc.), and the positions in these strip-like regions merely schematically represent the corresponding positions of these data sets in the input image, and do not limit the embodiments of the present disclosure.
[0044] For example, in step S10, the "multiple datasets" corresponding to the multiple computation units may be multiple initial input datasets obtained by decomposing an initial input matrix (these initial input datasets are input to the computation device), or multiple input datasets that are objects of a single convolution computation in the computation device. For example, the multiple datasets may be multiple datasets 10, 20, and 30 obtained by performing a convolution computation on %0 on PE0, PE1, and PE2 in FIG. 6A , or multiple datasets obtained by performing a convolution computation on %5. For example, the dataset obtained by each computation unit performing a convolution computation on %5 includes a first computation result set and intermediate data rows obtained from another computation unit. That is, in at least one embodiment of the present disclosure, the first computation result set in the first computation unit and the first intermediate data rows obtained from the second computation unit constitute a dataset required for the first computation unit to perform a second convolution computation.
[0045] For example, in step S10, the plurality of calculation units each perform a first convolution calculation on a corresponding plurality of data sets to obtain a corresponding plurality of first calculation result sets. For example, the plurality of first calculation result sets may be for constituting a first convolution layer obtained by the first convolution calculation. For example, as shown in FIG. 6A , the plurality of calculation units PE0, PE1, and PE2 each perform a first convolution calculation on a corresponding plurality of data sets 10, 20, and 30 to obtain a plurality of first calculation result sets 11, 21, and 31, and the plurality of first calculation result sets 11, 21, and 31 may be for constituting the first convolution layer obtained by the first convolution calculation.
[0046] For example, in the embodiments of the present disclosure, the multiple sets of first calculation results in the multiple calculation units may directly constitute the first convolution layer obtained by the first convolution calculation, or may constitute a part of the first convolution layer obtained by the first convolution calculation. For example, when the input image is relatively small, the multiple calculation units can carry all the input data at once, and the multiple sets of first calculation results in the multiple calculation units can be collected to form the first convolution layer. When the input image is relatively large, the multiple calculation units cannot carry all the input data at once, and the input data needs to be loaded into the calculation units in multiple batches. Each time, the loaded input data is divided into multiple sets and input into the multiple calculation units respectively. The multiple sets of first calculation results in the multiple calculation units form only a part of the first convolution layer corresponding to the currently loaded input data.
[0047] For example, in step 10, the data rows in the plurality of first calculation result sets are continuous without overlapping in the first convolutional layer. For example, as shown in FIG. 6A , the data rows in the plurality of first calculation result sets 11, 21, and 31 are continuous without overlapping in the first convolutional layer, that is, there are no overlapping portions in the data in the plurality of first calculation result sets 11, 21, and 31 obtained by performing the first convolutional calculation on %0 in the plurality of calculation units PE0, PE1, and PE2.
[0048] For example, in the embodiments of the present disclosure, a "data row" may refer to one row of data or one column of data in an input image. For example, in a dataset decomposed in the height dimension H shown in FIG. 6A, a "data row" may refer to one row of data in the height direction of the input image, and of course, a "data row" may refer to one column of data in the width direction of the input image, and the embodiments of the present disclosure are not limited thereto.
[0049] For example, in step S20, a data replication and transmission mode corresponding to the plurality of first calculation result sets in the plurality of calculation units is determined according to an arrangement rule for the plurality of calculation units of the second convolutional layer obtained by the plurality of calculation units performing the second convolutional calculation on the first convolutional layer. It should be noted that "performing the second convolutional calculation on the first convolutional layer" herein does not refer to the actual convolutional calculation in the calculation unit, but refers to the rule based on the convolutional calculation. The size of the second convolutional layer to be obtained by performing the second convolutional calculation can be determined based on the size of the first convolutional layer. Therefore, the "second convolutional layer to be obtained" does not mean that the specific calculation result of the data in the second convolutional layer needs to be obtained. For example, the decomposition mode for the second convolutional layer may be determined based on the size of the second convolutional layer, i.e., the arrangement rule for the plurality of calculation units of the second convolutional layer may be determined.
[0050] Fig. 4 is a schematic flowchart of step S20 in Fig. 3. For example, in some examples, as shown in Fig. 4, step S20 may further include steps S21 to S23.
[0051] Step S21: Determine an allocation rule for the plurality of computing units of the second convolution layer according to the internal memory sizes of the plurality of computing units; Step S22: Determine the distribution status of the plurality of calculation units of the second convolution layer according to the placement rule; Step S23: Determine a data duplication transmission mode corresponding to the first set of calculation results in the plurality of calculation units according to the distribution situation.
[0052] For example, in step S21, the memory (e.g., internal memory) sizes of the multiple computing units are obtained, and an allocation rule for the multiple computing units of the second convolutional layer is determined based on the internal memory size of each computing unit. For example, the second convolutional layer is decomposed into multiple pieces of data, and the size of each piece of data must not exceed the available capacity of the internal memory allocated to its corresponding computing unit, which is the premise for determining the allocation rule.
[0053] For example, in step S22, it determines how to decompose the second convolutional layer according to the determined arrangement rule, i.e., determines the distribution of the multiple data sets decomposed from the second convolutional layer among the multiple computing units. For example, it may determine the distribution of the multiple data sets decomposed from the second convolutional layer among the corresponding multiple computing units according to the distribution of the multiple first calculation result sets among the multiple computing units and the determined arrangement rule.
[0054] For example, in step S23, a data duplication and transmission mode corresponding to the first computation result set among the plurality of computing units is determined according to the distribution state among the plurality of computing units of the second convolutional layer determined according to the allocation rule. For example, the data duplication and transmission mode may include a unidirectional data transmission mode and a bidirectional data transmission mode. For example, the unidirectional data transmission mode indicates that a computing unit can only receive duplicated data from another computing unit, while the bidirectional data transmission mode indicates that a computing unit can not only receive duplicated data from another computing unit but also duplicate and transmit its own data to the other computing unit. For example, the plurality of computing units may be configured to have the same data duplication and transmission mode, or the plurality of computing units may be configured to have different data duplication and transmission modes. The data duplication and transmission mode is not limited to the above two modes and may be other feasible implementation methods, which are not limited to the embodiments of the present disclosure.
[0055] For example, in one example, the allocation rule may be such that the distribution of data of the second convolutional layer among the multiple computing units is the same as the data range of the first computation result set among the multiple computing units. For example, as shown in FIG. 6A , the multiple computing units PE0, PE1, and PE2 perform a first convolutional calculation on the data set %0 to obtain a first convolutional layer consisting of multiple first computation result sets 11, 21, and 31. The size of the second convolutional layer obtained by performing a second convolutional calculation on the first convolutional layer can then be calculated and obtained. Therefore, the data of the second convolutional layer is decomposed and allocated to the multiple computing units PE0, PE1, and PE2 according to the distribution pattern of %10 in FIG. 6A . The data duplication and transmission mode for each computing unit when performing convolutional calculation on %5 can be determined according to the allocation rule. For example, the data duplication and transmission mode for the multiple computing units PE0, PE1, and PE2 can be determined to be a bidirectional data transmission mode according to the allocation rule.
[0056] For example, in another example, the allocation rule may be such that the distribution of the data of the second convolutional layer among some of the computing units is the same as the data range of the first calculation result set among the computing units. For example, as shown in FIG. 6B, after obtaining the size of the second convolutional layer, the data of the second convolutional layer is allocated to the computing units PE0, PE1, and PE2 according to the distribution scheme of %10 in FIG. 6B, that is, the data of the second convolutional layer is decomposed into three parts, and the data range allocated to PE1 is equal to the size of the data range of the first calculation result set in PE1. Thus, according to the allocation rule, the data duplication and transmission mode when each computing unit performs convolution calculation on %5 can be determined, for example, the data duplication and transmission modes of the computing units PE0, PE1, and PE2 are all determined to be unidirectional data transmission modes.
[0057] In the embodiments of the present disclosure, the allocation rule can be set not only based on the size of the internal memory of the computing unit, but also flexibly based on factors such as the allocation of computing power of multiple computing units and data transmission overhead, and the embodiments of the present disclosure do not limit the specific content of the allocation rule.
[0058] For example, in step S30, a first computing unit among the plurality of computing units that should pad valid data rows is determined. For example, all of the plurality of computing units may need to pad valid data, or only some of the plurality of computing units may need to pad valid data rows. For example, in FIG. 6A, PE0, PE1, and PE2 are all first computing units that should pad valid data rows; in FIG. 6B, PE1 and PE2 are first computing units that should pad valid data rows; and in FIG. 6C, only PE1 is the first computing unit that should pad valid data rows.
[0059] For example, in step S30, after determining the first computing unit to pad valid data rows and the corresponding data duplication transmission mode, the first computing unit obtains at least one first intermediate row data from the first computation result set of another computing unit. For example, the first computing unit may obtain two first intermediate row data from the first computation result sets of two different computing units, and both of these first intermediate row data are required for padding in the second convolution calculation process by the first computing unit. For example, in FIG. 6A , PE0 as the first computing unit obtains one or more rows of first intermediate row data from the first computation result set 21 of PE1, PE1 as the first computing unit obtains one or more rows of first intermediate row data from the first computation result set 11 of PE0 and obtains one or more rows of first intermediate row data from the first computation result set 31 of PE2, and PE2 as the first computing unit obtains one or more rows of first intermediate row data from the first computation result set 21 of PE1.
[0060] Fig. 5 is a schematic flowchart of step S30 in Fig. 3. For example, in some examples, as shown in Fig. 5, step S30 may further include steps S31 and S32.
[0061] Step S31: Determine a size of a first intermediate data row according to a relationship between a second calculation result set obtained by the first calculation unit performing the second convolution calculation and a first calculation result set obtained by the first calculation unit performing the first convolution calculation; Step S32: Obtain a first intermediate data row from the first calculation result set in the second calculation unit according to the data duplication transmission mode and the size of the first intermediate data row.
[0062] For example, in step S31, the size of the first intermediate data row obtained by the first computing unit from another computing unit is determined by the relationship between the second computing result set and the first computing result set, and the relationship between the second computing result set and the first computing result set is determined by factors such as the distribution situation in the computing units of the second convolution layer, the size of the convolution kernel, and the padding operation.
[0063] For example, in step S32, the first intermediate data row may be obtained from the first calculation result set in the second calculation unit through kernel-to-kernel transmission between the calculation units, for example, by obtaining the address of the first intermediate data row in the first calculation result set of the second calculation unit, reading the first intermediate data row, and copying it to the first calculation unit, and then writing the copied and obtained first intermediate data row to the corresponding address of the first calculation unit. 6A , PE1 has valid data of [74,150] rows of %0, and the storage addresses corresponding to the 77 rows of data are address_0 to address_75. Then, after the first convolution calculation, the first calculation result set in PE1 is [75,149] rows in the first convolution layer, which is 75 rows of data in total, and these rows are assigned to addresses address_1 to address_74, respectively. The first intermediate data row obtained from the first calculation result set 11 of PE0 may be written to address_0, and the first intermediate data row obtained from the first calculation result set 31 of PE2 may be written to address_75. The embodiments of the present disclosure do not limit the specific implementation manner of the data replication and transmission process between calculation units.
[0064] For example, the scheduling method further includes step S40 (not shown) of decomposing an initial input matrix to obtain a plurality of initial input data sets, and transmitting the plurality of initial input data sets to a plurality of calculation units respectively to perform a first convolution calculation.
[0065] For example, the decomposition mode of the initial input matrix may be determined based on the size of the first convolution layer obtained by performing the first convolution calculation on the initial input matrix, the number of calculation units, the size of the internal memory, etc. For example, the decomposition mode may be to uniformly decompose the initial input matrix into multiple initial input data sets, or to decompose the initial input matrix into multiple initial input data sets based on the size of the internal memory of the calculation unit, and input the data into multiple calculation units, respectively. For example, the initial input data set may be a data set that is an object for performing the first convolution calculation.
[0066] For example, the scheduling method further includes step S50 (not shown) of acquiring at least a portion of the calculation output of the multilayer convolutional neural network by collecting a plurality of second calculation result sets of the plurality of calculation units in response to the second convolutional layer being an output layer in the multilayer convolutional neural network. For example, after the calculation of the last convolution operator is completed, if the second convolutional layer is the final convolutional layer, the second convolutional layer is the output result. That is, the second convolutional layer consisting of the second calculation result sets of the plurality of calculation units is the calculation output or at least a portion of the calculation output of the multilayer convolutional neural network.
[0067] In an embodiment of the present disclosure, the plurality of computing units includes at least two computing units, such as a first computing unit and a second computing unit, and the multi-layer convolutional neural network includes at least two convolutional layers. For example, the example shown in Figures 6A to 6C includes three computing units PE0, PE1, and PE2, %5:[(1 224 224 64),F,FP32]=conv(%0:[(1 224 224 64),F,FP32], 1%:[(64 3 3 64),W,FP32],kh=3,kw=3,pad_h_top=1,pad_h_bottom=1,pad_w_left=1,pad_w_right=1,stride_h=1,stride_w=1), and It contains two convolution operators (conv): %10:[(1 224 224 64),F,FP32]=conv(%5:[(1 224 224 64),F,FP32], 6%:[(64 3 3 64),W,FP32],kh=3,kw=3,pad_h_top=1,pad_h_bottom=1,pad_w_left=1,pad_w_right=1,stride_h=1,stride_w=1). 6A, the chip uses three computation units PE0, PE1, and PE2 to perform data-parallel convolution calculations on 224 × 224 input data. For example, step S40 is performed to uniformly decompose the initial input matrix into approximately three initial input data sets 10, 20, and 30 in the height dimension H of the input data, and load them into the three computation units PE0, PE1, and PE2, respectively, as shown in %0. For example, initial input data set 10 is [0, 75], containing a total of 76 rows of data; initial input data set 20 is [74, 150], containing a total of 77 rows of data; and initial input data set 30 is [149, 223], containing a total of 75 rows of data. For example, step S10 is executed, and the multiple calculation units PE0, PE1, and PE2 perform first convolution calculations on the three corresponding initial input data sets 10, 20, and 30, respectively, to obtain first calculation result sets 11 (row [0, 74]), 21 (row [75, 149]), and 31 (row [150, 223]) shown in %5, and the first calculation result sets constitute a first convolution layer obtained by performing the first convolution calculation on the initial input matrix.
[0068] For example, step S20 is performed to determine that the three computing units PE0, PE1, and PE2 all have the same data duplication and transmission mode, i.e., bidirectional data transmission mode, according to an arrangement rule that ensures that the distribution of data in the multiple computing units of the second convolution layer is the same as the data range of the first calculation result set in the multiple computing units. Next, step S30 is performed to determine that the three computing units PE0, PE1, and PE2 are all computing units that should perform data duplication and transmission, and that the first intermediate data row required for each computing unit is one row. Next, each computing unit obtains the first intermediate data row required for padding when performing the second convolution calculation from another computing unit.
[0069] For example, PE0 is the first calculation unit, and PE1 is the second calculation unit. As shown by arrow 1, the first calculation unit PE0 obtains the first intermediate data row
[75] required for padding in the second convolution calculation process by PE0 from the first calculation result set 21 of the second calculation unit PE1. As shown by arrow 2, the second calculation unit PE1 obtains the second intermediate data row
[74] required for padding in the second convolution calculation process by PE1 from the first calculation result set 11 of the first calculation unit PE0. The first and second intermediate data rows obtained by the first calculation unit PE0 and the second calculation unit PE1 from each other have the same size, and each is one row of data.
[0070] For example, PE1 is the first calculation unit, and PE2 is the second calculation unit. As shown by arrow 3, the first calculation unit PE1 obtains the first intermediate data row
[0150] required for padding in the second convolution calculation process by PE1 from the first calculation result set 31 in the second calculation unit PE2. As shown by arrow 4, the second calculation unit PE2 obtains the second intermediate data row
[0149] required for padding in the second convolution calculation process by PE2 from the first calculation result set 21 of the first calculation unit PE1. The first and second intermediate data rows obtained by the first calculation unit PE1 and the second calculation unit PE2 from each other have the same size and are both one row of data.
[0071] For example, the first calculation result set [0,74] in PE0 and the first intermediate data row
[75] obtained from PE1 constitute the data set [0,75] required for the second convolution calculation by PE0.
[0072] For example, the first calculation result set [75,149] in PE1 and the second intermediate data row
[74] and the first intermediate data row
[0150] obtained from PE0 and PE2, respectively, constitute the data set [74,150] required for the second convolution calculation by PE1.
[0073] For example, the first calculation result [150,223] in PE2 and the second intermediate data row
[0149] obtained from PE2 constitute the data set [149,223] required for the second convolution calculation by PE2.
[0074] The three calculation units PE0, PE1, and PE2 perform second convolution calculations on the data sets [0,75], [74,150], and [149,223] after data duplication and transmission, and obtain the second calculation result sets [0,74], [75,149], and [150,223] shown in %10. Finally, step S50 is executed, and the second calculation result sets [0,74], [75,149], and [150,223] in PE0, PE1, and PE2 are collected to form the calculation results of the output layer.
[0075] Therefore, the scheduling method according to at least one embodiment of the present disclosure can reduce the amount of calculation of repeated data in the convolutional neural network by copying and transmitting valid data between computing units and using the valid data as padding rows, thereby improving the calculation efficiency of the multi-layer convolutional neural network.
[0076] In another example shown in FIG. 6B, in step S20, it may be determined that three computing units PE0, PE1, and PE2 all have the same unidirectional data transmission mode. In step S30, it is determined that computing units PE1 and PE2 are the computing units that should perform data duplication transmission, and that the first intermediate data row required for each computing unit is two rows. That is, PE1 obtains two rows of valid data from PE0 for padding, and PE2 obtains two rows of valid data from PE1 for padding. For example, compared to FIG. 6B, in FIG. 6C, only computing unit PE1 may be the computing unit that should perform data duplication transmission. That is, PE1 obtains two rows of valid data from PE0 and PE2 for padding. The scheduling methods in the examples shown in FIGS. 6B and 6C can not only reduce repeated data calculations and improve computation efficiency, but also reduce the number of data transmissions between computing units and reduce the overhead required for data transmission.
[0077] For example, the computing device may further include more computing units, and the convolutional network may include more convolution operators, and the embodiments of the present disclosure are not limited thereto. For example, the example shown in FIG. 6D includes four computing units PE0, PE1, PE2, and PE3, %5:[(1 224 224 64),F,FP32]=conv(%0:[(1 224 224 64),F,FP32], 1%:[(64 3 3 64),W,FP32],kh=3,kw=3,pad_h_top=1,pad_h_bottom=1,pad_w_left=1,pad_w_right=1,stride_h=1,stride_w=1), %10:[(1 224 224 64),F,FP32]=conv(%5:[(1 224 224 64),F,FP32], 6%:[(64 3 3 64),W,FP32],kh=3,kw=3,pad_h_top=1,pad_h_bottom=1,pad_w_left=1,pad_w_right=1,stride_h=1,stride_w=1), and It contains three convolution operators (conv): %15:[(1 224 224 64),F,FP32]=conv(%10:[(1 224 224 64),F,FP32], 11%:[(64 3 3 64),W,FP32],kh=3,kw=3,pad_h_top=1,pad_h_bottom=1,pad_w_left=1,pad_w_right=1,stride_h=1,stride_w=1).
[0078] For example, as shown in FIG. 6D , multiple computing units PE0, PE1, PE2, and PE3 may have different data duplication and transmission modes. For example, the size of the first intermediate data row acquired by each computing unit from another computing unit may be different, and this is not a limitation of the embodiments of the present disclosure. For example, the first intermediate data row acquired by PE2 from PE1 is one row, while the first intermediate data row acquired from PE3 is two rows. The steps of the scheduling method in the example shown in FIG. 6D may refer to the description of FIG. 6A , and detailed descriptions will be omitted here. The scheduling method in the example shown in FIG. 6D can not only reduce repeated data calculations and improve the utilization rate of computing power, but also achieve a balanced allocation of internal memory, computing power, and data transmission overhead among different computing units in consideration of the overall performance of the system.
[0079] 7 is a schematic block diagram of a scheduling device according to some embodiments of the present disclosure. As shown in FIG. 7, the scheduling device 100 includes a calculation control module 110, an allocation scheduling module 120, and a data transmission module 130. These components may be interconnected via a bus and / or other type of connection mechanism (not shown).
[0080] The calculation control module 110 is configured to cause a plurality of calculation units to perform a first convolution calculation on a corresponding plurality of data sets, respectively, to obtain a corresponding plurality of first calculation result sets, the plurality of first calculation result sets being for constituting a first convolution layer obtained by the first convolution calculation, the plurality of calculation units including a first calculation unit and a second calculation unit; The allocation scheduling module 120 is configured to determine data replication transmission modes corresponding to the first calculation result sets in the plurality of calculation units according to allocation rules in the plurality of calculation units of the second convolution layer obtained by the plurality of calculation units performing second convolution calculations on the first convolution layer; The data transmission module 130 is configured to cause a first computing unit among the plurality of computing units, which is to pad a valid data row, to obtain a first intermediate data row required for padding in the second convolution calculation process by the first computing unit from the first calculation result set in the second computing unit based on a corresponding data duplication transmission mode.
[0081] It should be noted that in the embodiment of the present disclosure, each module of the scheduling device 100 corresponds to each step of the above scheduling method, and the specific functions of the scheduling device 100 may refer to the relevant description of the above scheduling method, and detailed description will be omitted here. The components and structure of the scheduling device 100 shown in Figure 7 are exemplary and not limiting, and the scheduling device 100 may further include other components and structures as necessary.
[0082] 8 is a schematic block diagram of an electronic device according to some embodiments of the present disclosure. As shown in FIG. 8, the electronic device 200 includes a scheduling device 210. The scheduling device 210 may be a scheduling device according to any one of the embodiments of the present disclosure, such as the scheduling device 100. The electronic device 200 may be any device with a computing function, such as a server, a terminal device, a personal computer, etc., and the embodiments of the present disclosure are not limited thereto.
[0083] FIG. 9 is a schematic block diagram of another electronic device according to some embodiments of the present disclosure. As shown in FIG. 9, the electronic device 300 includes a processor 310 and a memory 320 and may be used to implement a client terminal or a server. The memory 320 is used to non-instantaneously store computer-executable instructions (e.g., at least one computer program module). The processor 310 is used to execute the computer-executable instructions, which, when executed by the processor 310, may perform one or more steps in the convolution method described above and further realize the convolution method described above. The memory 320 and the processor 310 may be interconnected via a bus system and / or other types of connection mechanisms (not shown).
[0084] For example, the processor 310 may be a central processing unit (CPU), a graphics processing unit (GPU), or other type of processing unit having data processing and / or program execution capabilities. For example, the central processing unit (CPU) may have an X86 or ARM architecture, etc. The processor 310 may be a general-purpose processor or a special-purpose processor and may control other components in the electronic device 300 to perform desired functions.
[0085] For example, the memory 320 may include any combination of at least one (e.g., one or more) computer program products, which may include various types of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. At least one (e.g., one or more) computer program modules may be stored in the computer-readable storage medium, and the processor 310 may execute at least one (e.g., one or more) computer program modules to implement various functions of the electronic device 300. The computer-readable storage medium may also store various application programs, various data, and various data used and / or generated by the application programs.
[0086] It should be noted that in the embodiments of the present disclosure, the specific functions and technical effects of the electronic device 300 may refer to the above description of the scheduling method, and detailed description thereof will be omitted here.
[0087] FIG. 10 is a schematic block diagram of another electronic device according to some embodiments of the present disclosure. The electronic device 400 is suitable for implementing, for example, a scheduling method according to embodiments of the present disclosure. The electronic device 400 may be a terminal device or the like, and may be used to implement a client terminal or a server. The electronic device 400 may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), in-car terminals (e.g., in-car navigation terminals), and wearable electronic devices, as well as fixed terminals such as digital TVs, desktop computers, and smart home devices. It should be noted that the electronic device 400 shown in FIG. 10 is merely an example and does not limit the functionality and scope of use of embodiments of the present disclosure.
[0088] 10, the electronic device 400 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 410, which can perform various appropriate operations and processes based on programs stored in a read-only memory (ROM) 420 or programs loaded from a storage device 480 into a random access memory (RAM) 430. The RAM 430 further stores various programs and data necessary for operation by the electronic device 400. The processing unit 410, the ROM 420, and the RAM 430 are connected to each other via a bus 440. An input / output (I / O) interface 450 is also connected to the bus 440.
[0089] Generally, input devices 460 including, for example, a touch panel, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc., output devices 470 including, for example, a liquid crystal display (LCD), loudspeaker, vibrator, etc., storage devices 480 including, for example, a magnetic tape, hard disk, etc., and communication devices 490 may be connected to the I / O interface 450. The communication devices 490 may allow the electronic device 400 to exchange data with other electronic devices through wireless or wired communication. While FIG. 10 shows the electronic device 400 with various devices, it should be understood that it is not required to embody or include all of the illustrated devices, and the electronic device 400 may alternatively embody or include more or fewer devices.
[0090] For example, according to an embodiment of the present disclosure, the scheduling method may be implemented as a computer software program. For example, an embodiment of the present disclosure may include a computer program product including a computer program embodied in a non-transitory computer-readable medium, the computer program including program code for executing the scheduling method. In such an embodiment, the computer program may be downloaded and installed from a network via communication device 490, or may be installed from storage device 480, or may be installed from ROM 420. When the computer program is executed by processing device 410, it may implement limited functionality in the scheduling method according to an embodiment of the present disclosure.
[0091] At least one embodiment of the present disclosure further provides a storage medium, which can be used to improve the utilization rate of the matrix operation unit, effectively utilize the calculation capability of the matrix operation unit, shorten the time for convolution calculation, improve calculation efficiency, and save data transmission time.
[0092] 11 is a schematic diagram of a storage medium according to some embodiments of the present disclosure. For example, as shown in FIG. 11, storage medium 500 may be a non-transitory computer-readable storage medium, storing non-transitory computer-readable instructions 510. When non-transitory computer-readable instructions 510 are executed by a processor, the scheduling method described in the embodiments of the present disclosure can be realized, for example, when non-transitory computer-readable instructions 510 are executed by a processor, one or more steps of the scheduling method described above can be performed.
[0093] For example, the storage medium 500 may be applied to the electronic device, for example, the storage medium 500 may comprise the memory 320 in the electronic device 300 .
[0094] For example, the storage medium may include a memory card in a smartphone, a storage member in a tablet computer, a hard disk in a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a flash memory, or any combination of the above storage media, or may be other applicable storage media.
[0095] For example, the description of the storage medium 500 may refer to the description of the memory in the embodiment of the electronic device, and the overlapping parts will not be described again. The specific functions and technical effects of the storage medium 500 may refer to the description of the scheduling method above, and the detailed description will not be repeated here.
[0096] The scheduling method, scheduling device, electronic device, and storage medium according to the embodiments of the present disclosure have been described above with reference to Figures 1 to 11. The scheduling method according to the embodiments of the present disclosure can use valid data as padding rows by replicating and transmitting valid data between computing units, thereby reducing the amount of calculation of repeated data in the convolutional neural network and improving the calculation efficiency of the multilayer convolutional neural network.
[0097] It should be noted that in the context of the present disclosure, a computer-readable medium may be a tangible medium that can contain or store a program used by or in combination with an instruction execution system, apparatus, or device. A computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of both. A computer-readable storage medium may be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. Further specific examples of a computer-readable storage medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program, which may be used by or in combination with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a propagating data signal, either in baseband or as part of a carrier wave, that includes computer-readable program code. Such propagating data signals may take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may transmit, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code included in the computer-readable medium may be transmitted via any suitable medium, including, but not limited to, wire, optical cable, RF (radio frequency), etc., or any suitable combination of the above.
[0098] In some embodiments, client terminals and servers may communicate using any now known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and may interconnect with any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any now known or later developed networks.
[0099] The computer-readable medium may be included in the electronic device, or may exist separately but not assembled into the electronic device.
[0100] Computer program code for carrying out the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may run entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. When a remote computer is involved, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0101] The flowcharts and block diagrams in the figures illustrate possible system architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, program segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function. It should be noted that in some permutation implementations, the functions depicted in the blocks may be implemented in a different order than depicted in the figures. For example, two successively shown blocks may actually be executed essentially simultaneously, or they may be executed in the reverse order, depending on the functionality involved. It should be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0102] The units according to the embodiments of the present disclosure may be implemented in software or hardware, and the names of the units may not limit the units themselves.
[0103] The functionality described herein above may be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc.
[0104] According to one or more embodiments of the present disclosure, Example 1 provides a scheduling method for a multi-layer convolutional neural network, including: a plurality of computing units performing first convolutional calculations on corresponding plurality of data sets, respectively, to obtain corresponding first calculation result sets, the plurality of first calculation result sets being for constituting a first convolutional layer obtained by the first convolutional calculations, the plurality of computing units including a first computing unit and a second computing unit; determining data duplication and transmission modes corresponding to the plurality of first calculation result sets in the plurality of computing units according to an arrangement rule of the plurality of computing units of a second convolutional layer obtained by the plurality of computing units performing second convolutional calculations on the first convolutional layer; and obtaining a first intermediate data row required for padding in the second convolutional calculation process by the first computing unit from the first calculation result set in the second computing unit based on the data duplication and transmission mode corresponding to a first computing unit among the plurality of computing units that is to pad a valid data row.
[0105] In Example 2, based on the scheduling method described in Example 1, data rows in the plurality of first calculation result sets are consecutive without overlapping in the first convolutional layer.
[0106] In Example 3, based on the scheduling method described in Example 1, the first calculation result set in the first calculation unit and the first intermediate data row obtained from the second calculation unit form a data set required for the first calculation unit to perform the second convolution calculation.
[0107] In Example 4, based on the scheduling method described in Example 1, the corresponding plurality of data sets used by the plurality of computing units when performing the first convolution calculation are a plurality of initial input data sets.
[0108] In Example 5, based on the scheduling method described in Example 4, The method further includes decomposing an initial input matrix to obtain the plurality of initial input data sets, and transmitting the plurality of initial input data sets to the plurality of calculation units respectively to perform the first convolution calculation.
[0109] In Example 6, based on the scheduling method described in Example 1, determining data replication and transmission modes corresponding to the plurality of first calculation result sets in the plurality of calculation units according to an allocation rule in the plurality of calculation units of a second convolution layer obtained by the plurality of calculation units performing second convolution calculations on the first convolution layer, Determining the placement rule in the plurality of computing units of the second convolutional layer based on the size of internal memory in the plurality of computing units; Determining a distribution state of the plurality of computing units of the second convolutional layer according to the placement rule; and determining the data duplication transmission mode corresponding to the first set of calculation results in the plurality of calculation units according to the distribution status.
[0110] In Example 7, based on the scheduling method described in Example 1, obtaining a first intermediate data row required for padding in the second convolution calculation process by the first calculation unit from the first calculation result set of the second calculation unit includes: determining a size of the first intermediate data row according to a relationship between a second calculation result set obtained by the first calculation unit performing the second convolution calculation and a first calculation result set obtained by the first calculation unit performing the first convolution calculation; and obtaining the first intermediate data row from the first calculation result set in the second calculation unit according to the data duplication transmission mode and the size of the first intermediate data row.
[0111] In Example 8, based on the scheduling method described in Example 7, the size of the second calculation result set that the first computing unit is to obtain is the same as the size of the first calculation result set in the first computing unit.
[0112] In Example 9, based on the scheduling method described in Example 7, In response to the second convolutional layer being an output layer in the multilayer convolutional neural network, the method further includes acquiring at least a portion of a calculation output of the multilayer convolutional neural network by collecting a plurality of second calculation result sets in the plurality of calculation units.
[0113] In Example 10, based on the scheduling method described in any one of Examples 1 to 9 above, The method further includes obtaining a second intermediate data row required for padding in the second convolution calculation process by the second calculation unit from the first calculation result set in the first calculation unit based on the data duplication transmission mode corresponding to the second calculation unit.
[0114] In Example 11, based on the scheduling method described in Example 10, the first and second intermediate data rows acquired by the first and second computing units from each other have the same size.
[0115] According to one or more embodiments of the present disclosure, Example 12 provides a scheduling apparatus for a multi-layer convolutional neural network, comprising: a calculation control module configured to cause a plurality of calculation units to perform a first convolution calculation on a corresponding plurality of data sets to obtain a corresponding plurality of first calculation result sets, the plurality of first calculation result sets being for constituting a first convolution layer obtained by the first convolution calculation, the plurality of calculation units including a first calculation unit and a second calculation unit; an allocation scheduling module configured to determine data replication transmission modes corresponding to the plurality of first calculation result sets in the plurality of calculation units according to allocation rules in the plurality of calculation units of a second convolution layer obtained when the plurality of calculation units perform second convolution calculations on the first convolution layer; and a data transmission module configured to cause a first calculation unit among the plurality of calculation units, which is to pad a valid data row, to obtain a first intermediate data row required for padding in the second convolution calculation process by the first calculation unit from a first calculation result set in the second calculation unit based on the corresponding data duplication transmission mode.
[0116] In Example 13, based on the scheduling device described in Example 12, data rows in the plurality of first calculation result sets are consecutive without overlapping in the first convolutional layer.
[0117] In Example 14, based on the scheduling device described in Example 12, the first calculation result set in the first calculation unit and the first intermediate data row obtained from the second calculation unit form a data set required for the first calculation unit to perform the second convolution calculation.
[0118] In Example 15, based on the scheduling device of Example 12, the corresponding plurality of data sets used by the plurality of calculation units when performing the first convolution calculation are a plurality of initial input data sets.
[0119] In Example 16, based on the scheduling device described in Example 15, The method further includes a data decomposition module configured to decompose an initial input matrix to obtain the plurality of initial input data sets, and transmit the plurality of initial input data sets to the plurality of calculation units, respectively, to perform the first convolution calculation.
[0120] In Example 17, based on the scheduling device described in Example 12, determining data replication and transmission modes corresponding to the plurality of first calculation result sets in the plurality of calculation units according to an arrangement rule in the plurality of calculation units of a second convolution layer obtained by the plurality of calculation units performing second convolution calculations on the first convolution layer, Determining the placement rule in the plurality of computing units of the second convolutional layer based on the size of internal memory in the plurality of computing units; Determining a distribution state of the plurality of computing units of the second convolutional layer according to the placement rule; and determining the data duplication transmission mode corresponding to the first set of calculation results in the plurality of calculation units according to the distribution status.
[0121] In Example 18, based on the scheduling device described in Example 12, obtaining a first intermediate data row required for padding in the second convolution calculation process by the first calculation unit from the first calculation result set in the second calculation unit includes: determining a size of the first intermediate data row according to a relationship between a second calculation result set obtained by the first calculation unit performing the second convolution calculation and a first calculation result set obtained by the first calculation unit performing the first convolution calculation; and obtaining the first intermediate data row from the first calculation result set in the second calculation unit according to the data duplication transmission mode and the size of the first intermediate data row.
[0122] In Example 19, based on the scheduling device described in Example 18, the size of the second calculation result set that the first computing unit is to acquire is the same as the size of the first calculation result set in the first computing unit.
[0123] In Example 20, based on the scheduling device described in Example 18, The multilayer convolutional neural network further includes a data output module configured, in response to the second convolutional layer being an output layer in the multilayer convolutional neural network, to acquire a calculation output of at least a portion of the multilayer convolutional neural network by collecting a plurality of second calculation result sets in the plurality of calculation units.
[0124] In Example 21, based on the scheduling device according to any one of Examples 12 to 20, the data transmission module further comprises: The second calculation unit is configured to obtain, based on the corresponding data duplication transmission mode, a second intermediate data row required for padding in the second convolution calculation process by the second calculation unit from the first calculation result set in the first calculation unit.
[0125] In Example 22, based on the scheduling device of Example 10, the first and second intermediate data rows acquired by the first and second computing units from each other have the same size.
[0126] According to one or more embodiments of the present disclosure, Example 23 provides an electronic device, comprising the scheduling device according to any one of Examples 12 to 22 above.
[0127] According to one or more embodiments of the present disclosure, Example 24 provides an electronic device, comprising: a processor; and a memory including at least one computer program module, the at least one computer program module being stored in the memory and configured to be executed by the processor, the at least one computer program module for implementing the scheduling method described in any one of Examples 1 to 11 above.
[0128] According to one or more embodiments of the present disclosure, Example 25 provides a storage medium having non-transitory computer-readable instructions stored thereon, the non-transitory computer-readable instructions, when executed by a computer, implementing the scheduling method described in any one of Examples 1 to 11 above.
[0129] Although the present disclosure has been described in detail above through general descriptions and specific embodiments, it is obvious to those skilled in the art that some modifications or improvements can be made based on the examples of the present disclosure, and therefore, any modifications or improvements made without departing from the gist of the present disclosure shall fall within the scope of the claims of the present disclosure.
[0130] This disclosure requires some explanation in the following respects.
[0131] (1) The drawings of the embodiments of the present disclosure only relate to the structures related to the embodiments of the present disclosure, and other structures may refer to conventional designs.
[0132] (2) For clarity, in the drawings illustrating the embodiments of the present disclosure, the thicknesses of layers or regions are exaggerated or reduced, i.e., the drawings are not drawn to actual scale.
[0133] (3) Unless there is a conflict, the embodiments and features of the embodiments of the present disclosure can be combined with each other to obtain new embodiments.
[0134] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto, and should be consistent with the scope of protection of the claims.
Claims
1. 1. A scheduling method for a multi-layer convolutional neural network, comprising: A plurality of calculation units perform a first convolution calculation on a corresponding plurality of data sets to obtain a corresponding plurality of first calculation result sets, the plurality of first calculation result sets being for constituting a first convolution layer obtained by the first convolution calculation, and the plurality of calculation units include a first calculation unit and a second calculation unit; Determine data duplication and transmission modes corresponding to the first calculation result sets in the plurality of calculation units according to an arrangement rule in the plurality of calculation units of a second convolution layer obtained by the plurality of calculation units performing second convolution calculations on the first convolution layer; and obtaining a first intermediate data row required for padding in the second convolution calculation process by the first calculation unit from a first calculation result set in the second calculation unit based on the data duplication transmission mode corresponding to a first calculation unit that is to pad a valid data row among the plurality of calculation units.
2. The scheduling method of claim 1 , wherein data rows in the plurality of first calculation result sets are consecutive without overlapping in the first convolutional layer.
3. 3. The scheduling method according to claim 1, wherein the first calculation result set in the first calculation unit and the first intermediate data row obtained from the second calculation unit constitute a data set required for the first calculation unit to perform the second convolution calculation.
4. 4. The scheduling method according to claim 1, wherein the plurality of data sets used by the plurality of calculation units when performing the first convolution calculation are a plurality of initial input data sets.
5. 5. The scheduling method of claim 4, further comprising: decomposing an initial input matrix to obtain the plurality of initial input data sets; and transmitting the plurality of initial input data sets to the plurality of calculation units, respectively, to perform the first convolution calculation.
6. Determining data replication and transmission modes corresponding to the plurality of first calculation result sets in the plurality of calculation units according to an arrangement rule in the plurality of calculation units of a second convolution layer obtained by the plurality of calculation units performing second convolution calculations on the first convolution layer includes: Determining the placement rule in the plurality of computing units of the second convolutional layer based on the size of internal memory of the plurality of computing units; Determining a distribution state of the plurality of computing units of the second convolutional layer according to the placement rule; The scheduling method according to any one of claims 1 to 5, further comprising: determining the data duplication transmission mode corresponding to the first calculation result set in the plurality of calculation units according to the distribution situation.
7. The step of obtaining a first intermediate data row required for padding in the second convolution calculation process by the first calculation unit from the first calculation result set of the second calculation unit includes: determining a size of the first intermediate data row according to a relationship between a second calculation result set obtained by the first calculation unit performing the second convolution calculation and a first calculation result set obtained by the first calculation unit performing the first convolution calculation; and obtaining the first intermediate data row from the first calculation result set in the second calculation unit based on the data duplication transmission mode and the size of the first intermediate data row.
8. The scheduling method according to claim 7 , wherein the size of the second calculation result set that the first calculation unit is to acquire is the same as the size of the first calculation result set in the first calculation unit.
9. 8. The scheduling method of claim 7, further comprising, in response to the second convolutional layer being an output layer in the multilayer convolutional neural network, acquiring at least a portion of a computation output of the multilayer convolutional neural network by collecting a plurality of second computation result sets in the plurality of computation units.
10. 10. The scheduling method according to claim 1, further comprising: obtaining a second intermediate data row required for padding in the second convolution calculation process by the second calculation unit from the first calculation result set in the first calculation unit according to the data duplication transmission mode corresponding to the second calculation unit.
11. The scheduling method according to claim 10 , wherein the first and second intermediate data rows acquired by the first and second computing units from each other have the same size.
12. 1. A scheduling apparatus for a multi-layer convolutional neural network, comprising: a calculation control module configured to cause a plurality of calculation units to perform a first convolution calculation on a plurality of corresponding data sets to obtain a plurality of corresponding first calculation result sets, the plurality of first calculation result sets being for constituting a first convolution layer obtained by the first convolution calculation, the plurality of calculation units including a first calculation unit and a second calculation unit; an allocation scheduling module configured to determine data replication transmission modes corresponding to the plurality of first calculation result sets in the plurality of calculation units according to allocation rules in the plurality of calculation units of a second convolution layer obtained when the plurality of calculation units perform second convolution calculations on the first convolution layer; and a data transmission module configured to cause a first calculation unit to obtain a first intermediate data row required for padding in the second convolution calculation process by the first calculation unit from a first calculation result set in the second calculation unit based on the corresponding data duplication transmission mode.
13. An electronic device comprising the scheduling device according to claim 12.
14. An electronic device, a processor; a memory containing at least one computer program module; The at least one computer program module is stored in the memory and configured to be executed by the processor, the at least one computer program module being for implementing the scheduling method according to any one of claims 1 to 11.
15. A storage medium, A storage medium having stored thereon non-transitory computer readable instructions that, when executed by a computer, implements the scheduling method of any one of claims 1 to 11.
Citation Information
Patent Citations
Method and device for adapting parameter of neural network
JP2019061664A
Artificial intelligence inference computing device
JP2020042774A