Instruction set generation method for neural network operation and computing device thereof

By adopting layer partitioning technology in NPUs, the layers of the neural network are grouped into multiple partial networks, which solves the problem of increasing data exchange between the NPU and external memory, and improves computing efficiency and resource utilization.

CN120019386APending Publication Date: 2025-05-16OPENEDGES TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071432.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-06
Filing Date
2023-10-05
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the neural network processing unit (NPU), as the number of neural network layers increases, the amount of data exchange between the NPU and external memory increases, resulting in a decrease in bandwidth consumption and computing efficiency of computing devices.

Method used

Through layer partitioning technology, the layers of the neural network are grouped into multiple partial networks, and by defining slice layers, partial networks and connection layers, the number of read and write operations that occur during execution of each layer is reduced, thereby reducing the communication bandwidth.

Benefits of technology

It effectively reduces the amount of data exchange between the NPU and the external memory, improves the computing efficiency and utilization of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019386A_ABST
    Figure CN120019386A_ABST
Patent Text Reader

Abstract

A method of generating an NPU composite, the method comprising the steps of: generating a pth partial network having the same structure as a first network structure defined by a first set of layers included in a predefined neural network; determining, in a first memory included in another computing device, a p-th read address, which is a location where an address is stored as a p-th portion input activation value of data to be input to the most upstream layer of the p-th portion network; determining, in the first memory, a p-th write address, which is a position at which an address of a p-th portion output activation value, which is data output as a most downstream layer of the p-th portion network, should be stored; and generating an NPU instruction [p] based on the pth read address and the pth write address.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology for generating instructions for improving neural network operation efficiency and computing resource utilization in a computing device including a neural processing unit (NPU). Background Art

[0002] The present invention relates to neural network operations performed in an NPU installed in a computing device. Figure 1 In the example of neural network operation, the Convolutional Neural Network (CNN) is taken as an example.

[0003] Figure 1 FIG. 1 shows a computational structure of a CNN according to an embodiment. Figure 1 For explanation. First, a convolution operation can be performed on the input image data (51) stored in the internal memory, and the operation uses multiple convolution kernels to generate a convolution layer (52). The step of generating the convolution layer (52) may include performing a nonlinear operation (for example: ReLU, Sigmoid or tanH) on multiple feature maps obtained by the convolution operation. Then, a pooling layer (53) can be generated by performing pooling on the convolution layer (52). Each convolution layer (52) may include data that can be expressed in the form of an M*N matrix. Then, flattening can be performed on the pooling layer (53) to generate an array input to the internal neural network (54). Then, the array can be input to the internal neural network (54) to generate an output from the internal neural network (54).

[0004] Figure 1 All the operations shown to be distinguished from each other can be regarded as different layers. In addition, the neural network according to the present invention can be regarded as comprising Figure 1 All the layers shown can also be considered to represent the internal neural network (54). Figure 1 These are examples for helping understanding, and thus the scope of the neural network according to the present invention is not limited to the above contents.

[0005] In a neural network, data can be transmitted in one direction and then be calculated and transformed whenever it encounters a layer. This transformation and flow of data can be represented by the word stream. The neural network may include a first layer and a second layer. At this time, if the output activation value output by the first layer is directly or further transformed and input to the second layer, the first layer can be called a layer located more upstream than the second layer, and the second layer can be called a layer located more downstream than the first layer. The terms "upstream" and "downstream" are introduced for the convenience of explaining the present invention.

[0006] A neural processing unit (NPU) may be installed in a computing device such as a desktop computer, a laptop computer, a smart phone or a tablet computer. The NPU may have a structure suitable for neural network operations. At this time, if the NPU is to perform neural network operations, the control unit inside the NPU must execute predetermined instructions for neural network operations to control the internal resources of the NPU. The instructions may be stored in the NPU during the manufacturing process of the user device, or may be provided to the NPU after the user device is manufactured.

[0007] When a predetermined neural network is run on an NPU, the size of input and output data of a specific layer defined in the predetermined neural network may be larger than the internal memory inside the NPU. In this case, it is necessary to divide the input and output data into a size that can be stored in the internal memory for processing.

[0008] In order to perform the operation corresponding to a specific layer, the NPU can obtain the input data required for the operation, such as the input activation value and other input data (for example, weights, etc.) to be input to the specific layer from the memory outside the NPU (for example, DRAM) through the bus. In addition, the output activation value (output data) output by the specific layer can be provided to the memory outside the NPU again through the bus. Whenever the operation of each layer is performed, write operations and read operations are performed on the external memory through the bus. Therefore, as the number of layers in the neural network increases, a large amount of computing resources will be consumed, and there is a problem of reduced overall computing efficiency. This problem will also occur even if the input and output data are divided into a size that can be stored in the internal memory for calculation.

[0009] The layers that make up the neural network can be set by the neural network maker to have a variety of input and output connections, so it is difficult to achieve effective computational division for all connection situations. Therefore, there is a problem of difficulty in achieving efficient hardware operation in terms of power consumption and bandwidth.

[0010] In one embodiment of the neural network operation method, in order to perform layer operations, input data such as tensors, layer parameters, weights and biases are required. It may happen that the size of this data is larger than the capacity of the NPU internal memory (SRAM). In addition, the layer operation can generate an output tensor such as the result output activation value, and the size of the output tensor may be larger than the capacity of the NPU internal memory.

[0011] The output activation value output from a specific layer may be recorded in an external memory of the NPU. In order to input the output activation value to the next layer of the specific layer, the NPU should read the output activation value recorded in the external memory and store it in the internal memory. Therefore, in order to transmit the activation value between layers, a write operation and a read operation through the bus may occur respectively.

[0012] In one embodiment, the layer that inputs the partial input activation values ​​generated by partitioning the input activation values ​​in the row dimension may be a convolution layer. At this time, the number of rows contained in the partial input activation values ​​should be greater than or equal to the size of the convolution kernel used for the convolution layer operation. In addition, the size of the partial input activation values ​​should be less than or equal to the capacity of the internal memory of the NPU. In addition, as the number of partitioned layers increases, the number of additional repeated operations increases, so there is a problem that the read bandwidth and the amount of operations may increase.

[0013] As described above, when layer partitioning is performed for the operation of the NPU, there is a problem that read operations and write operations to the external memory inevitably occur. Summary of the invention

[0014] Technical issues

[0015] The object of the present invention is to provide a technology for generating NPU instructions, which can reduce the amount of data exchanged between the NPU and other external memories to reduce the bandwidth of the computing device and improve the computing efficiency of the NPU.

[0016] Technical Solution

[0017] The instructions executed by the NPU may be generated and provided by a developer who wants to provide an application using a predetermined neural network operation. The present invention includes content related to a development tool that helps the above developer generate the above instructions.

[0018] The present invention can utilize the concept of layer partitioning. Layer partitioning can refer to a method of defining multiple layers based on one layer to generate a layer that can be run on the NPU when a computing device (data computing unit) of the NPU is used to perform computing according to the computing rules of the layers constituting the neural network.

[0019] In the present invention, the operation of combining the plurality of partial output activation values ​​to generate one output activation value is called a concat layer operation. When the concat layer operation is executed on the user computing device, the concat layer operation can be implemented by the NPU writing the plurality of partial output activation values ​​to an external memory (e.g., DRAM) outside the NPU. That is, when all the plurality of partial output activation values ​​are stored in the appropriate designated parts of the external memory, it can be considered that the one output activation has been generated.

[0020] According to a neural network operation method provided by one aspect of the present invention, in order to reduce the amount of data transmitted between the NPU and the DRAM through the bus, a group consisting of interconnected continuous layers can be defined in the layers constituting the neural network processed by the NPU. As a result, the communication bandwidth of the system including the NPU and the DRAM can be reduced. To this end, the entire neural network can be grouped into a predefined layer input-output structure that is conducive to operation segmentation.

[0021] The groups provided according to one aspect of the present invention have at least three types. The first type of group can be called an Inverse-Y group, the second type of group can be called a serial group, and the third type of group can be called a residual group. The groups provided according to one aspect of the present invention are not limited to the above three types.

[0022] At this time, the network defined by the defined group may be partitioned into a plurality of partial networks, and as a basis for the partitioning, the capacity of an internal memory included in the NPU may be used.

[0023] At this time, the starting layer (upstreammost layer) and the ending layer (downstreammost layer) of the layers constituting each group may be determined based on a benchmark that minimizes the consumption of hardware resources. Factors that need to be considered in order to optimize hardware resources include: overlap activation size, weight reloading size, and DRAM input and output size.

[0024] According to one aspect of the present invention, a plurality of layers may be grouped to generate a layer group, and the generated layer group may be partitioned, so that the number of read operations and write operations to an external memory occurring during execution of each layer within the defined layer group may be reduced. As a result, the bandwidth used for NPU operation may be reduced. In this specification, the layer group may be referred to as a group for short.

[0025] According to one aspect of the present invention, a grouping process for generating, by a developer computing device, a group consisting of a plurality of layers constituting a neural network may be provided.

[0026] According to one aspect of the present invention, a group partitioning process for partitioning, by a developer computing device, a group consisting of a plurality of layers constituting a neural network may be provided.

[0027] At this time, the grouping process may be executed prior to the group partitioning process.

[0028] In order to perform the grouping process, a layer grouping pattern may be predefined, the layer grouping pattern representing a specific pattern of consecutive layers that can be grouped. In the case where there are parts in the layers of the neural network that are identical to the predefined layer grouping pattern, grouping of the parts may be performed.

[0029] The neural network has had its structure designed before executing the method according to the invention and may not have been through an optimization process for a specific NPU.

[0030] Through the group partitioning process, a second network can be generated based on the first network defined by the group. The second network can be called a partitioned network.

[0031] The partitioned network may include: P partial networks having the same network structure information as the first network; P slicing layers generating P input activation values ​​to be input to the P partial networks; and a connection layer combining P output activation values ​​output from the P partial networks.

[0032] Here, the network structure information of the first network may be information including layers constituting the group (first network), operation rules of the layers, and links of activation value transmission paths between the layers.

[0033] The group partitioning process may include the following steps:

[0034] In step S310 , the developer computing device may define a group consisting of a plurality of layers constituting a neural network.

[0035] The rule defining the one group may be a rule using features of network structure information of the neural network.

[0036] In step S320, the developer computing device may define P slice layers, which divide the input activation values ​​to be input to the group and generate P partial input activation values.

[0037] At this time, the size of each part of the input activation value may be smaller than the capacity of a storage body storing the input activation value in the internal memory of the NPU included in the user computing device.

[0038] At this time, the activation values ​​input to the slice layers may be the same. Also, the activation values ​​output by each slice layer may have different values.

[0039] In step S330 , the developer computing device may define P partial networks, and the P partial networks respectively receive the P partial input activation values.

[0040] At this time, the network structure information of each of the partial networks may be the same as the network structure information of the first network defined by the group.

[0041] At this time, the partial input activation values ​​input to each partial network may include only partial data of the input activation values ​​that should be input to the most upstream layer among the layers belonging to the group.

[0042] In step S340, the developer computing device may define a connection layer that combines the P partial output activation values ​​respectively output by the P partial networks.

[0043] In step S350, the developer computing device may define a plurality of links, which represent activation value transmission paths between the P slice layers, the P partial networks, and the connection layer.

[0044] A partitioned network may be defined by defining the P slice layers, the P partial networks, the connection layer, and the plurality of links.

[0045] According to one aspect of the present invention, a method for generating an NPU instruction may be provided, the method comprising the following steps: a computing device generates a p-th partial network having the same structure as a first network structure defined by a first group of layers included in a predefined neural network; the computing device determines a p-th read address in a first memory included in another computing device, the p-th read address being a location where an address of a p-th partial input activation value as data to be input to the most upstream layer of the p-th partial network is stored; the computing device determines a p-th write address in the first memory, the p-th write address being a location where an address of a p-th partial output activation value as data output from the most downstream layer of the p-th partial network should be stored; and the computing device generates an NPU instruction [p], the NPU instruction comprising a first instruction set, a second instruction set, and a third instruction set. At this time, the first instruction set includes instructions for causing the NPU included in the other computing device to read the p-th partial input activation value from the first memory based on the p-th read address and store it in the internal memory of the NPU. The second instruction set includes instructions for causing the NPU to generate the p-th partial output activation value based on the p-th partial input activation value stored in the internal memory. Furthermore, the third instruction set includes instructions for causing the NPU to store the p-th partial output activation value in the first memory based on the p-th write address.

[0046] At this time, the p-th part of the input activation value may be a part of the input activation value input to the most upstream layer among the layers of the first group,

[0047] At this time, the first memory may be a memory provided outside the NPU, the p-th part input activation value may be transmitted from the first memory to the internal memory of the NPU via the bus of the other computing device, and the p-th part output activation value may be transmitted from the internal memory to the first memory via the bus.

[0048] At this time, the p-th partial output activation value may be generated by calculating the p-th partial input activation value stored in the internal memory based on the calculation rules of each layer included in the p-th partial network.

[0049] At this time, the step of generating the p-th partial network may include: the computing device defines the first group consisting of a plurality of consecutive layers included in a predefined neural network; the computing device generates structural information about the first network consisting of a plurality of layers and a plurality of links included in the defined first group; and the computing device generates the p-th partial network having the same structure as the first network. At this time, the structural information about the first network may be information about links representing the layers constituting the first group, the operation rules of the layers, and the activation value transmission paths between the layers.

[0050] At this time, the first group may include multiple layers, the most upstream layer may be a layer among the multiple layers that receives activation values ​​from the outside of the first group, and the most downstream layer may be a layer among the multiple layers that provides activation values ​​to the outside of the first group.

[0051] According to another aspect of the present invention, a method for generating an NPU instruction may be provided, the method comprising the following steps: a computing device generates a partitioned network including a p-th partial network (p is 1, 2, ..., or P) based on a first network composed of a first group of layers included in a predefined neural network; and the computing device generates an NPU instruction [p] (p is 1, 2, ..., or P) for the p-th partial network to be executed by an NPU included in another computing device. The step of generating a partitioned network may include: the computing device defines a p-th slicing layer (p is 1, 2, ..., or P), which receives input activation values ​​to be input into the first group and outputs partial input activation values ​​as part of the input activation values; the computing device defines a p-th partial network (p is 1, 2, ..., or P), which receives the p-th partial input activation value output by the p-th slicing layer; the computing device defines a connection layer, which combines P partial output activation values ​​output by P said partial networks; and the computing device completes the partitioned network by defining P said slicing layers, P said partial networks, and multiple links representing activation value transmission paths between the connection layers.

[0052] At this time, the first group of layers may be a plurality of consecutive layers included in the predefined neural network.

[0053] At this time, the pth part of the input activation value may be a part of the input activation value input to the most upstream layer among the layers of the first group, and the input activation value may be capable of being reconstructed from the first part of the input activation value to the Pth part of the input activation value.

[0054] At this time, the structure of the p-th partial network may be the same as the structure of the first network (p is 1, 2, ..., or P). In addition, the step of generating the NPU instruction [p] may include: the computing device determines the p-th read address in the first memory included in the other computing device, the p-th read address being the location of the address of the p-th partial input activation value as the data to be input to the most upstream layer of the p-th partial network; the computing device determines the p-th write address in the first memory, the p-th write address being the location of the address of the p-th partial output activation value as the data output from the most downstream layer of the p-th partial network; and the computing device generates the NPU instruction [p], the NPU instruction including the first instruction set, the second instruction set, and the third instruction set. In addition, the first instruction set may include instructions for causing the NPU to read the p-th partial input activation value from the first memory based on the p-th read address and store it in the internal memory of the NPU. The second instruction set may include instructions for causing the NPU to generate the p-th partial output activation value based on the p-th partial input activation value stored in the internal memory. Furthermore, the third instruction set may include instructions for causing the NPU to store the p-th partial output activation value in the first memory based on the p-th write address.

[0055] At this time, the first memory may be a memory provided outside the NPU. Furthermore, the p-th part of input activation values ​​may be transmitted from the first memory to the internal memory of the NPU via a bus of the other computing device, and the p-th part of output activation values ​​may be transmitted from the internal memory to the first memory via the bus.

[0056] At this time, the step of generating the p-th partial network may include: the computing device defines the first group consisting of a plurality of consecutive layers included in the predefined neural network; the computing device generates structural information about the first network consisting of a plurality of layers and a plurality of links included in the defined first group; and the computing device generates the p-th partial network having the same structure as the first network. At this time, the structural information about the first network may be information about links representing the layers constituting the first group, the operation rules of the layers, and the activation value transmission paths between the layers.

[0057] According to another aspect of the present invention, a computing device including a storage unit and a main processor can be provided. The storage unit stores a program including instructions, which causes the main processor to perform the following steps: generate a p-th partial network, which has the same structure as the first network structure defined by the first group of layers included in the predefined neural network; determine the p-th read address in the first memory included in another computing device, which is the location of the address of the p-th partial input activation value as data to be input to the most upstream layer of the p-th partial network; determine the p-th write address in the first memory, which is the location of the address of the p-th partial output activation value as data output by the most downstream layer of the p-th partial network; and generate an NPU instruction [p], which includes a first instruction set, a second instruction set, and a third instruction set. The first instruction set includes instructions for causing the NPU included in the other computing device to read the p-th partial input activation value from the first memory based on the p-th read address and store it in the internal memory of the NPU. The second instruction set includes instructions for causing the NPU to generate the p-th partial output activation value based on the p-th partial input activation value stored in the internal memory. And the third instruction set includes instructions for causing the NPU to store the p-th partial output activation value in the first memory based on the p-th write address.

[0058] According to another aspect of the present invention, a computing device including a storage unit and a main processor may be provided. The storage unit stores a program including instructions, which causes the main processor to perform the following steps: generating a partitioned network including a p-th partial network (p is 1, 2, ..., or P) based on a first network composed of a first group of layers included in a predefined neural network; generating an NPU instruction [p] (p is 1, 2, ..., or P) for the p-th partial network to be executed by an NPU included in another computing device. The step of generating a partitioned network includes: the computing device defines a p-th slice layer (p is 1, 2, ..., or P), the p-th slice layer receives input activation values ​​to be input into the first group, and outputs partial input activation values ​​as part of the input activation values; the computing device defines a p-th partial network (p is 1, 2, ..., or P), the p-th partial network receives the p-th partial input activation value output by the p-th slice layer; the computing device defines a connection layer, the connection layer combines the P partial output activation values ​​output by the P partial networks; and the computing device completes the partitioned network by defining multiple links, the multiple links representing activation value transmission paths between the P slice layers, the P partial networks and the connection layer.

[0059] According to one aspect of the present invention, a neural network operation method executed in an NPU including an internal memory may be provided. The neural network operation method comprises the following steps: repeatedly executing a predetermined first process [p] from p=1 to p=P in sequence (p=1, ..., P, P is a natural number greater than or equal to 2). At this time, the first process [p] includes the following steps: reading part of the input activation value [1][p] from the external memory connected via the bus and storing it in the first storage body of the internal memory; calculating the part of the input activation value [1][p] stored in the first storage body according to the operation rules of layer [1], and storing the generated part of the output activation value [1][p] in the first storage body; repeating the second process from s=1 to s=L-1 in sequence (L is a natural number greater than or equal to 2), the second process is to calculate the part of the output activation value [s][p] stored in the first storage body according to the operation rules of the layer [s+1] connected to the output end of the layer [s], and store the generated part of the output activation value [s+1][p] in the first storage body; and writing the part of the output activation value [L][p] stored in the first storage body into the external memory through the bus.

[0060] At this time, the partial input activation value [1][p] can be a part of the input activation value [1] of the layer [1] input to the neural network, or can be a value generated based on the part (p=1, ..., P, P is a natural number greater than or equal to 2).

[0061] At this time, the neural network operation method may further include the following steps: before sequentially repeating the first process [p], reading the weight [1] of the operation rule for the layer [1] and the weight [s+1] of the operation rule for the layer [s+1] from the external memory via the bus.

[0062] (s=1, ..., L-1), and store it in the second storage body of the internal memory. At this time, the output activation value [1][p] can be generated based on the input activation value [1][p] stored in the first storage body and the weight [1] stored in the second storage body, and the output activation value [s+1][p] can be generated based on the output activation value [s][p] stored in the first storage body and the weight [s+1] stored in the second storage body (s=1, ..., L-1).

[0063] At this time, the step of repeating the predetermined first process [p] can be performed based on a set of NPU instructions executed by the NPU, and the external memory stores the partial input activation value

[0064] The address of [1][p] may be included in the NPU instruction, and the address of the partial output activation value [L][p] that should be stored in the external memory may also be included in the NPU instruction.

[0065] At this time, the output activation value composed of the partial output activation value [L][p] (p=1, ..., P) can be the input activation value [L+1] input to the layer [L+1]. In addition, the neural network operation method can further include the following steps: after repeating the first process [p], the third process [q] is repeatedly executed from q=1 to q=Q in sequence (Q is a natural number greater than or equal to 2). At this time, the third process [q] may include the following steps: reading the input activation value [L+1][q] from the external memory connected via the bus and storing it in the first storage body of the internal memory; operating the input activation value [L+1][q] stored in the first storage body according to the operation rules of the layer [L+1], and storing the generated output activation value [L+1][q] in the first storage body; repeating the fourth process from s=L+1 to s=M-1 in sequence (M is a natural number greater than or equal to L+2), the fourth process is to operate the output activation value [s][q] stored in the first storage body according to the operation rules of the layer [s+1] connected to the output end of the layer [s], and store the generated output activation value [s+1][q] in the first storage body; and writing the output activation value [M][q] stored in the first storage body into the external memory through the bus.

[0066] At this time, the partial input activation value [L+1][q] can be a part of the input activation value [L+1] of the layer [L+1] input to the neural network, or it can be a value generated based on the part (q=1, ..., Q, Q is a natural number greater than or equal to 2).

[0067] At this time, the layer [1], the layer [s+1] (s=1, ..., L-1), the layer [L+1] and the layer [s+1] (s=L+1, ..., M-1) can be included in the neural network.

[0068] At this time, the partial output activation value [sc][p] may be based on the partial input activation value [sc][p] stored in the first memory bank and the weight stored in the second memory bank of the internal memory.

[0069] [sc], the operation rule of the layer [sc] may be a convolution operation rule (sc=1, ..., or L), the input activation value [1] may be a three-dimensional tensor consisting of a width dimension, a height dimension and an input channel dimension, the weight [sc] may be a four-dimensional tensor consisting of a width dimension, a height dimension, an input channel dimension and an output channel dimension, the size of the input channel dimension of the input activation value [1] may be the same as the size of the input channel dimension of the weight [sc], and the partial input activation value [1][p] may be a part of the input activation value [1] divided along the width dimension or the height dimension, or generated based on the part (p=1, ..., P, P is a natural number greater than or equal to 2).

[0070] According to another aspect of the present invention, there may be provided an NPU device, which includes: an internal memory, a control unit, and a data operation unit. The control unit is configured to use the data operation unit to execute a predetermined first process [p] repeatedly from p=1 to p=P in sequence (p=1, ..., P, P is a natural number greater than or equal to 2). The first process [p] includes the following steps: reading part of the input activation values ​​[1][p] from an external memory connected via a bus and storing them in a first storage body of the internal memory;

[0071] [1][p] performs operations according to the operation rules of layer [1] and generates partial output activation values

[0072] [1][p] is stored in the first storage body; the second process is repeatedly executed from s=1 to s=L-1 in sequence (L is a natural number greater than or equal to 2), and the second process is to operate the partial output activation value [s][p] stored in the first storage body according to the operation rule of the layer [s+1] connected to the output end of the layer [s], and store the generated partial output activation value [s+1][p] in the first storage body; and write the partial output activation value [L][p] stored in the first storage body into the external memory through the bus.

[0073] According to another aspect of the present invention, a computing device may be provided, which includes: the NPU device; the bus; and the external memory.

[0074] Effects of the Invention

[0075] According to the present invention, a technology for generating NPU instructions can be provided, which can reduce the amount of data exchanged between the NPU and other external memories to reduce the bandwidth of the computing device and improve the computing efficiency of the NPU. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 The calculation structure of a CNN according to an embodiment is shown.

[0077] Figure 2 The main structure of a computing device for executing a method related to a neural network operation according to an embodiment of the present invention is shown.

[0078] Figure 3 The concept of a user computing device acquiring an instruction file executed by an NPU according to an embodiment of the present invention is illustrated.

[0079] Figure 4 Show Figure 2 The computing device, internal memory and DMA of the NPU in the user computing device are shown.

[0080] Figure 5 FIG. 1 shows the structure of input activation values ​​input to a neural network layer according to an embodiment of the present invention.

[0081] Figure 6 A neural network operation method using row dimension partitioning according to an embodiment of the present invention is shown.

[0082] Figure 7 The concept of grouping process provided according to one aspect of the present invention is shown.

[0083] Figure 8a and Figure 8b The concepts of a group partitioning process for dividing a group consisting of layers into a plurality of partitions according to an aspect of the present invention are respectively shown.

[0084] Figure 8c is a flow chart of a group partitioning process provided according to an embodiment of the present invention.

[0085] Figure 9a , Figure 9b and Fig.9c The flowchart is a method for a developer computing device to generate a set of NPU instructions provided to a user computing device according to an embodiment of the present invention.

[0086] Fig.10a : is a conceptual diagram illustrating a portion of a simplified neural network structure to help understand the neural network used in an embodiment of the present invention.

[0087] Fig.10b 2 is a conceptual diagram for facilitating understanding of groups defined by partial layers included in a neural network according to an embodiment of the present invention.

[0088] Fig.10c It is a schematic diagram for explaining a group-defined network and its structure defined according to an embodiment of the present invention.

[0089] Fig.10dA method for defining multiple partial networks based on a network according to an embodiment of the present invention is shown.

[0090] Fig.10e Shows the correspondence between the network and the partial network.

[0091] Fig.11a , Fig.11b and Fig.11c According to the comparative example, Figure 2 A method for performing neural network operations in a user computing device.

[0092] Fig.12a , Figure 12b , Fig.12c , Fig.13a , Fig.13b and Fig.13c The present invention shows an embodiment of Figure 2 A method for performing neural network operations in a user computing device.

[0093] Fig.14 is a schematic diagram illustrating a neural network operation method provided according to an embodiment of the present invention.

[0094] Fig.15 and Fig.16 is a flowchart of a neural network operation method provided according to an embodiment of the present invention. DETAILED DESCRIPTION

[0095] The embodiments of the present invention are described below with reference to the accompanying drawings. However, the present invention is not limited to the embodiments described in this specification and can be embodied in a variety of other forms. The terms used in this specification are used to help understand the embodiments and are not intended to limit the scope of the present invention. In addition, the singular form used below also includes the plural form as long as the sentence does not clearly express the opposite meaning.

[0096] Figure 2 The main structure of a computing device for executing a method related to a neural network operation according to an embodiment of the present invention is shown.

[0097] Figure 2 The user computing device ( 1 ) shown may be a device such as a desktop computer, a laptop computer, a smartphone or a tablet computer.

[0098] The computing device (1) may include a dynamic random access memory (DRAM) (130), an NPU (110), a bus (700) connecting the DRAM (130) and the NPU (110), other hardware (99) connected to the bus (700), a main processor (160) and a storage unit (170).

[0099] The NPU (110) may also be referred to as a hardware accelerator.

[0100] In addition, the computing device (1) may further include a power supply unit, a communication unit, a user interface, and a peripheral device unit (not shown). The bus (700) may also be shared by the NPU (110), other hardware (99), and the main processor (160).

[0101] The storage unit (170) can be integrated with the computing device (1) or can be detachably integrated with the computing device (1).

[0102] The NPU (110) may include a direct memory access part (DMA) (20), a control part (40), an internal memory (30), an input buffer (650), a data operation part (610) and an output buffer (640).

[0103] Part or all of the data temporarily stored in the internal memory (30) may be provided from the DRAM (130) through the bus (700). At this time, in order to transfer the data stored in the DRAM (130) to the internal memory (30), the control unit (40) and the DMA unit (20) may control the internal memory (30) and the DRAM (130).

[0104] In this specification, the DRAM (130) may also be referred to as an external memory.

[0105] The data stored in the internal memory (30) can be provided to the data operation unit (610) through the input buffer (650).

[0106] The output value generated by operating the data operation unit (610) can be stored in the internal memory (30) via the output buffer (640). The output value stored in the internal memory (30) can be written into the DRAM (130) under the control of the control unit (40) and the DMA unit (20).

[0107] The control unit (40) can coordinate and control the operation of the internal resources of the NPU (110), such as the DMA unit (20), the internal memory (30) and the data operation unit (610).

[0108] In one embodiment, the data operation unit (610) may perform a first operation function in a first time period, and perform a second operation function in a second time period. For example, the data operation unit (610) may perform a first operation function corresponding to an operation rule of a first layer of a neural network in a first time period, and perform a second operation function corresponding to an operation rule of a second layer of a neural network in a second time period.

[0109] exist Figure 2 In the embodiment, the NPU (110) is provided with a data operation unit (610). However, in a variant embodiment not shown in the figure, Figure 2 The NPU (110) shown may be provided with a plurality of data operation units (610), and each of the data operation units (610) may execute operations requested by the control unit (40) in parallel.

[0110] In one embodiment, the data operation unit (610) may not output its output data all at once, but may output its output data sequentially in a given order according to time.

[0111] Figure 2 The developer computing device (2) shown may be a device such as a server, a desktop computer, or a notebook computer. The computing device (2) may include a DRAM (230), a bus (2700), other hardware (299), a main processor (260), and a storage unit (270).

[0112] Figure 3 The concept of a user computing device acquiring an instruction file executed by an NPU according to an embodiment of the present invention is illustrated.

[0113] In this specification, the user computing device (1) may be referred to as a first computing device, and the developer computing device (2) may be referred to as a second computing device.

[0114] In one example, the user computing device (1) may obtain an instruction file executed by the NPU from the developer computing device (2) via a predetermined communication channel.

[0115] In another example, the user computing device (1) can obtain the instruction file executed by the NPU from the developer computing device (2) via a predetermined communication channel through a relay device (3). The relay device (3) can also be a production device for the production process of the user computing device (1).

[0116] Figure 4 Show Figure 2 In the user computing device (1) shown, the NPU (110) includes a computing device (COMP, data computing unit) (610), an internal storage bank (SRAM) (Bank0-2) (30) and a DMA (20). The DMA (20) obtains data from an external memory (e.g., DRAM) via a bus and stores it in the internal memory (30). At this time, the stored data is data required for layer operations, such as input tensors (e.g., input activation values) and layer parameters (e.g., weights of each layer). At this time, the size of each of the data should be less than or equal to the capacity of each storage body.

[0117] exist Figure 4In the example, memory bank (0) may be a memory unit for storing input activation values, memory bank (1) may be a memory unit for storing weights, and memory bank (2) may be a memory unit for storing output activation values.

[0118] Figure 5 FIG. 1 shows the structure of input activation values ​​input to a neural network layer according to an embodiment of the present invention.

[0119] like Figure 5 As shown, the input activation value is a tensor with dimensions (C, H, X). Where H represents the height of the tensor, X represents the width of the tensor, C represents the depth of the tensor, and C represents the number of channels of the tensor.

[0120] The input activation value may be Figure 5 The channel-wise partition method is based on the line segment (AB) of Figure 5 The line segment (AC) is used as the basis for separation of row-wise partitioning, or Figure 5 The line segment (BC) is used as the reference for column-wise partitioning.

[0121] According to an embodiment of the present invention, in the provided neural network operation method, the input activation value can be dimensionally partitioned according to the row dimension partitioning method or the column dimension partitioning method.

[0122] Figure 6 A neural network operation method using row dimension partitioning according to an embodiment of the present invention is shown.

[0123] exist Figure 6 The upper part of , shows the concept of performing a convolution operation on input activation values ​​represented as a tensor with dimensions (C, H, X) to generate output activation values.

[0124] exist Figure 6 The lower part of shows the concept of performing row-dimensional partitioning on the input activation values ​​represented as a tensor with dimensions (C, H, X) to generate a first portion of input activation values ​​and a second portion of input activation values, then performing a convolution operation on the first portion of input activation values ​​to generate a first portion of output activation values, and performing a convolution operation on the second portion of input activation values ​​to generate a second portion of output activation values, and finally, combining the first portion of output activation values ​​with the second portion of output activation values ​​to generate a complete output activation value. At this time, a first weight corresponding to a first output channel can be convolved in the first portion of input activation values, and a second weight corresponding to a second output channel can be convolved in the second portion of input activation values.

[0125] According to the characteristics of the convolution operation, in order to reconstruct the output activation value by combining the first part of the output activation value with the second part of the output activation value, the first part of the input activation value should contain all channels of the input activation value, and the second part of the input activation value should also contain all channels of the input activation value. That is, when the total number of channels contained in the input activation value is Nc, the first part of the input activation value and the second part of the input activation value should both contain data about Nc channels. Therefore, in order Figure 6 The operation method shown in the upper part and the operation method shown in the lower part can provide the same result, and the input activation value should be partitioned by the row dimension partitioning method or the column dimension partitioning method instead of the channel dimension partitioning method.

[0126] Despite Figure 6 A row dimension partitioning method is shown, but differently from this, each of the multiple partial input activation values ​​generated by the column dimension partitioning method can also include all channels of the input activation value.

[0127] <Group partitioning process executed by the developer computing device>

[0128] Figure 7 The concept of grouping process provided according to one aspect of the present invention is shown.

[0129] The grouping process may be implemented in a developer computing device (2).

[0130] Figure 7 The left side of FIG. 1 shows some layers constituting a predetermined neural network. The neural network is only used as an example, and the structure of the neural network to which the present invention is applicable is not limited thereto.

[0131] exist Figure 7 In the example, layer L[4] and layer L

[12] are layers that copy the activation value input to themselves and output it twice. For example, layer L[4] provides the input activation value to layer L[8] and layer L[5] respectively.

[0132] exist Figure 7 In the example, layer L[8] and layer L

[16] are layers that add multiple input activation values ​​element-wise and output an output activation value.

[0133] For example, layer L[8] outputs the activation value received from layer L[4] and the activation value received from layer L[7] element-wise. Therefore, the magnitude of the output activation value output by layer L[4] and the magnitude of the output activation value output by layer L[7] should be the same. Furthermore, the magnitude of the output activation value output by layer L[8] is the same as the magnitude of the output activation value output by layer L[4] and the magnitude of the output activation value output by layer L[7].

[0134] Figure 7The right side is based on Figure 7 The neural network layers shown on the left side of exemplify the concept of generating a predetermined rule group according to an embodiment of the present invention.

[0135] In one embodiment of the present invention, a plurality of layers may form a group. Fig.14 In the example, layers L[1] to L[3] form a first group G1, layers L[4] to L

[11] form a second group G2, and layers L

[12] to L

[16] form a third group G3.

[0136] In the first group G1, the most upstream layer and the most downstream layer are layer L[1] and layer L[3], respectively; in the second group G2, the most upstream layer and the most downstream layer are layer L[4] and layer L

[11] , respectively; and in the third group G3, the most upstream layer and the most downstream layer are layer L

[12] and layer L

[16] , respectively.

[0137] Figure 8a and Figure 8b The concepts of the group partitioning process of dividing a group of layer components into multiple partitions according to one embodiment of the present invention are respectively shown.

[0138] In the following, Figure 8a and Figure 8b They can be collectively referred to as Figure 8.

[0139] The group partitioning process may be implemented in the developer computing device (2).

[0140] In the following, reference is made to Figure 8a Provide explanation.

[0141] Figure 8a According to an embodiment of the present invention, the partitioning rule Figure 7 The first group G1 in is reconstructed into P partitions. P is a natural number greater than or equal to 2. Figure 8a In the example of , P=3. Therefore, the first group G1 may be converted into a first partition group PG1.

[0142] The developer computing device (2) may define a group G1 consisting of a plurality of layers L[1] to L[3] constituting a neural network.

[0143] The network N[1] defined based on the group G1 may include multiple layers in the group G1 and links respectively connected to the multiple layers.

[0144] The developer computing device (2) can define three slice layers:

[0145] SL[1][1]~SL[1][3], this slice layer segmentation requires the input activation value IA[1] of the group G1 to be input and generates three (p=3) partial input activation values ​​IA[1][1]~IA[1][3] respectively.

[0146] In the symbol IA[s][p] representing the partial input activation value, s is a value identifying the layer to which the partial input activation value should be input, and p is a value identifying the partition formed by the group partitioning process (p=1, ..., P, P is the number of partitions).

[0147] In the notation SL[g][p] representing the slice layer, g is a value identifying a group and p is a value identifying a partition formed by the group partitioning process. For example, SL[1][2] represents a layer that generates partial input activation values ​​IA[1][1] provided to the first partial network PN[1][1] of the first partition of the first group.

[0148] The developer computing device (2) can define three partial networks:

[0149] PN[1][1]~PN[1][3], the three partial networks receive the three partial input activation values ​​IA[1][1]~IA[1][3] respectively.

[0150] In the notation PN[g][p] representing the partial network, g is a value identifying a group, and p is a value identifying a partition formed by the group partitioning process (p=1, ..., P, where P is the number of partitions).

[0151] At this time, the network structure information of each partial network PN[1][1] to PN[1][3] may be the same as the network structure information of the network N[1] defined by the group G1. That is, the number of layers contained in each network, the operation rules of each layer, and the connection relationship between layers may be the same.

[0152] The developer computing device (2) may define a connection layer Conc.[1]

[0153] (concatenation layer), the connection layer combines the N partial output activation values ​​output by the three partial networks PN[1][1]~PN[1][3] to generate an output activation value OA[3].

[0154] In the symbol OA[s][p] representing the partial output activation value, s is a value identifying a layer that outputs the partial output activation value, and p is a value identifying a partition formed by the group partitioning process (p=1, ..., P, where P is the number of partitions).

[0155] In the notation OA[s] representing the output activation value, s is a value identifying a layer that outputs a portion of the output activation value constituting the output activation value.

[0156] In the symbol Conc.[g] representing the connection layer, g is a value identifying a group.

[0157] The developer computing device (2) may define a plurality of links, which represent activation value transmission paths between the three slice layers, the three partial networks and the connection layer.

[0158] Similarly, the developer computing device (2) can define the partitioned network PN[1] based on the network N[1] by defining the three slice layers, the three partial networks, the connection layer and the multiple links.

[0159] In the following, reference will be made to Figure 8b Give a description.

[0160] Figure 8b According to an embodiment of the present invention, the partitioning rule Figure 7 An example of reconstructing the second group G2 in FIG. 1 into P partitions. In this example, P=2. Therefore, the second group G2 can be converted into a second partition group PG2.

[0161] The developer computing device (2) can define a group G2 consisting of multiple layers L[4]~L

[11] that constitute a neural network.

[0162] The network N[2] defined based on the group G2 may include multiple layers in the group G2 and links respectively connected to the multiple layers.

[0163] The developer computing device (2) may define two slice layers SL[2][1]~SL[2][3], which divide the input activation value IA[4] to be input to the group G2 and generate two (p=2) partial input activation values ​​IA[4][1]~IA[4][2] respectively.

[0164] Again, the input activation value IA[4] can be compared with Figure 8a The output activation value OA[3] is the same.

[0165] The developer computing device (2) may define two partial networks PN[2][1]-PN[2][2], which respectively receive the two partial input activation values ​​IA[4][1]-IA[4][2].

[0166] At this time, the network structure information of each of the partial networks PN[2][1] to PN[2][2] may be the same as the network structure information of the network N[2] defined by the group G2.

[0167] The developer computing device (2) may define a connection layer Conc.[2], which combines the two partial output activation values ​​OA

[11] [1]~OA

[11] [2] respectively output by the two partial networks PN[2][1]~PN[2][2] to generate an output activation value OA

[11] .

[0168] The developer computing device (2) may define a plurality of links, which represent activation value transmission paths between the two slice layers, the two partial networks, and the connection layer.

[0169] Similarly, the developer computing device (2) can define the partitioned network PN[2] based on the network N[2] by defining the two slice layers, the two partial networks, the connection layer and the multiple links.

[0170] like Figure 8a and Figure 8b As shown, the first topology representing the connection relationship between the multiple layers constituting the first group G1 and the second topology representing the connection relationship between the multiple layers constituting the second group G2 may be different from each other. However, no matter what topology a specific group has, a partition group (e.g., PG1) corresponding to the specific group (e.g., G1) can be generated by defining multiple partial networks (e.g., PN[1][1], PN[1][2], PN[1][3]), and the partial network has the same structural information as the structural information of the network (e.g., N[1]) defined by the specific group (e.g., G1). That is, a partition network (e.g., PN[1]) corresponding to the network (e.g., N[1]) can be generated.

[0171] Figure 8c is a flow chart of a group partitioning process provided according to an embodiment of the present invention.

[0172] The developer computing device (2) may perform a grouping process that generates a group consisting of a plurality of layers constituting a neural network. In order to perform the grouping process, a layer grouping pattern may be predefined, the layer grouping pattern representing a specific pattern of consecutive layers that may be grouped. In the case where there are portions in the layers of the neural network that have the same predefined layer grouping pattern, grouping of the portions may be performed.

[0173] Furthermore, the developer computing device (2) may provide a group partitioning process, which is a process for partitioning a group consisting of a plurality of layers constituting a neural network.

[0174] Through the group partitioning process, a second network can be generated based on the first network defined by the group. The second network can be called a partitioned network.

[0175] The partitioned network may include: P partial networks having the same network structure information as the first network; P slicing layers generating P input activation values ​​to be input to the P partial networks; and a connection layer combining P output activation values ​​output from the P partial networks.

[0176] Here, the network structure information of the first network may be information including layers constituting the group (first network), operation rules of the layers, and links of activation value transmission paths between the layers.

[0177] The group partitioning process may include the following steps:

[0178] In step S310 , the developer computing device may define a group consisting of a plurality of layers constituting a neural network.

[0179] In step S320, the developer computing device may define P slice layers, which segment the input activation values ​​to be input to the group and generate P partial input activation values.

[0180] In step S330 , the developer computing device may define P partial networks, and the P partial networks respectively receive the P partial input activation values.

[0181] In step S340, the developer computing device may define a connection layer that combines the P partial output activation values ​​respectively output by the P partial networks.

[0182] In step S350, the developer computing device may define a plurality of links, which represent activation value transmission paths between the P slice layers, the P partial networks, and the connection layer.

[0183] A partitioned network may be defined by defining the P slice layers, the P partial networks, the connection layer, and the plurality of links.

[0184] <NPU execution instruction generation method executed by developer computing device>

[0185] exist Figure 78 , a concept is proposed in which a developer computing device generates a partitioned network based on a group consisting of multiple layers.

[0186] In one embodiment of the present invention, generating the partition network may be generating a data structure including objects and functions that define the partition network shown in FIG. 8 .

[0187] The developer computing device (2) can use the generated partitioned network to generate an instruction set, which executes a neural network operation method for generating an output activation value that should be output by a group of the multiple layers from an input activation value input to the group. The instruction set can be transmitted to the user computing device (1), and the instruction set can be executed in the user computing device (1).

[0188] Figure 9a The flowchart of a method for a developer computing device (2) to generate a group of NPU instructions provided to a user computing device (1) according to an embodiment of the present invention is shown.

[0189] In step S10 , the developer computing device ( 2 ) may define a group consisting of a plurality of consecutive layers included in a neural network.

[0190] For example, the layer may be Figure 8a Group G1.

[0191] In this specification, the developer computing device (2) may be referred to as a second computing device.

[0192] In step S20, the second computing device (2) may generate structural information about a network consisting of the plurality of layers and the plurality of links included in the defined group.

[0193] Here, each layer can be regarded as a node constituting a network. Furthermore, the structure of the network can be specified based on the connection relationship between the nodes and links included in the network. Furthermore, each layer can be distinguished from each other based on the operation function performed by the layer and the position of the layer in the network. The structure of the network can be reproduced by the structural information.

[0194] In step S30, the second computing device (2) may generate a plurality of p-th partial networks (p=1, ..., P; P is a natural number greater than or equal to 2) having the same structural information as that of the network.

[0195] That is, a plurality of p-th partial networks having the same structure as that of the network can be generated.

[0196] In step S40, the second computing device (2) can determine the pth read address, which is the location of the address of the pth input activation value of the most upstream layer to be input to the pth partial network in the external memory (130) of the first computing device (user computing device) (1) (p=1,...,P, P is a natural number greater than or equal to 2).

[0197] Step S40 is, for example, Figure 8a The function of the slice layer [1] [1] is related to the slice layer

[0198] [1][1] is a functional module that generates and outputs a first partial input activation value IA[1][1] from an input activation value IA[1]. Here, the first partial input activation value IA[1][1] is a portion of the input activation value IA[1]. However, the function can be executed as a simulation in the second computing device. In contrast, when the function is executed in the first computing device, the function corresponds to an operation of reading data from a first read address, which is a location in an external memory (130) of the first computing device (1) where the address of the first partial input activation value IA[1][1] is stored.

[0199] In step S50, the second computing device (2) can determine the pth write address, which is the location of the address of the pth partial output activation value output by the most downstream layer of the pth partial network in the external memory (130) of the first computing device (1) (p=1, ..., P, P is a natural number greater than or equal to 2).

[0200] For example, in Figure 8a In the example, the most downstream layer of the first partial network PN[1][1] is layer L[3], and the first partial output activation value is OA[3][1].

[0201] In step S60, the second computing device (2) may generate NPU instructions (p) (p=1, ..., P, where P is a natural number greater than or equal to 2) including the first instruction set, the second instruction set and the third instruction set.

[0202] The first instruction set is an instruction set that enables the NPU (110) of the first computing device (1) to read the p-th part input activation value from the external memory (130) through the bus (99) of the first computing device (1) based on the p-th read address and store it in the internal memory (30) of the NPU (110).

[0203] The second instruction set is an instruction set for causing the NPU (110) to generate a p-th part output activation value as data for the output of the most downstream layer of the p-th part network by operating the p-th part input activation value stored in the internal memory (30) based on the operation rules of the layers included in the p-th part network.

[0204] The third instruction set is an instruction set for causing the NPU (110) to store the p-th partial output activation value in the external memory through the bus based on the p-th write address.

[0205] Figure 9b As Figure 9a A variant embodiment of the present invention shows a method for generating P groups of NPU instructions given network structure information of a group consisting of multiple consecutive layers.

[0206] exist Figure 9a After step S20, step S121 of setting the value of variable p to 1 may be performed.

[0207] In step S122 , it can be determined whether the value of the variable p is greater than a predetermined value P. If p>P is satisfied, the process can be moved to step S80 to end the process, and if p>P is not satisfied, the process can be moved to step S130 .

[0208] Figure 9b Steps S130 to S160 correspond to Figure 9a Steps S30 to S60 are performed based on the currently set p value.

[0209] In step S70 , the value of the variable p may be increased by 1, and then the process returns to step S122 .

[0210] Fig.9c As Figure 9a Another variant embodiment of the present invention shows a method for generating P groups of NPU instructions given network structure information of a group consisting of multiple consecutive layers.

[0211] exist Figure 9a After step S50, step S251 of setting the value of variable p to 1 may be performed.

[0212] In step S252 , it can be determined whether the value of the variable p is greater than a predetermined value P. If p>P is satisfied, the process can be moved to step S80 to end the process, and if p>P is not satisfied, the process can be moved to step S260 .

[0213] Figure 9b Step S260 corresponds to Figure 9a Step S60 is performed based on the currently set p value.

[0214] In step S70, the value of the variable p may be increased by 1, and then the process returns to step S252.

[0215] The generated set of NPU instructions may be sent to the user computing device (1). The user computing device (1) may use the set of NPU instructions to execute steps S800 and S900 described below.

[0216] In the following, a neural network operation method executed in a user computing device (1) using the set of NPU instructions is described in detail.

[0217] The following Fig.10a , Fig.10b , Fig.10c , Fig.10d and Fig.10e Shown from different perspectives Figure 7 , Figure 8a and Figure 8b Schematic diagram of a neural network and the concept of layers shown.

[0218] Fig.10a : is a conceptual diagram illustrating a portion of a simplified neural network structure to help understand the neural network used in an embodiment of the present invention.

[0219] Fig.10a The neural network (10) shown is composed of four layers L[1], L[2], L[3], and L[4] connected in series, and the operation rules OR of the layers are given by OR[1], OR[2], OR[3], and OR[4] respectively. The operation rules OR of each layer can represent the transfer function of the input and output data of the layer.

[0220] Fig.10b 2 is a conceptual diagram for facilitating understanding of groups defined by partial layers included in a neural network according to an embodiment of the present invention.

[0221] A group defined according to an embodiment of the present invention may include multiple layers that are directly connected to each other. Fig.10b An example of a first group G1 consisting of layer L[1], layer L[2], and layer L[3] is shown. At this time, in the first group G1, layer L[1] becomes the most upstream layer, and layer L[3] becomes the most downstream layer. The activation value input to layer L[1] is called input activation value [1]. The activation value output from layer L[s] is called output activation value OA[s] (s=1, 2, 3). The output activation value OA[s] is the input activation value s+1 (s=1, 2, 3) input to layer L[s+1].

[0222] Fig.10c It is a schematic diagram for explaining a group-defined network and its structure defined according to an embodiment of the present invention.

[0223] A network N[1] may be defined based on the first group G1.

[0224] The network N[1] may include multiple layers in the first group G1 and links respectively connected to the multiple layers. The link may represent a connection relationship between two layers mediated by an activation value transmitted between the two layers. That is, the link is a transmission path of the activation value between the multiple layers.

[0225] When the output activation value OA[s] of layer L[s] is provided to layer L[s+1], layer L[s] and layer L[s+1] can be regarded as being connected by a link identified by the output activation value OA[s]. At this time, the link can be called an outbound link of layer L[s] and can be called an inbound link of layer L[s+1].

[0226] according to Fig.10c In the example shown, the inbound link of the first layer L[1] is the first link LK[1], the inbound link of the second layer L[2] is the second link LK[2], the inbound link of the third layer L[3] is the third link LK[3], the outbound link of the first layer L[1] is the second link LK[2], the outbound link of the second layer L[2] is the third link LK[3], and the outbound link of the third layer L[3] is the fourth link LK[4].

[0227] At this time, structure information [1] which is structure information about the network N [1] may be defined.

[0228] The structural information [1] may be composed of a portion of the structural information of the neural network (10).

[0229] In this specification, the term "neural network structure information" refers to the structure information of the neural network (10), and "structure information [k]" refers to the structure information of the network N[k].

[0230] The structure information [1] may include information identifying an inbound link connected to any layer of the plurality of layers and an outbound link connected to any layer.

[0231] Furthermore, the structural information [1] may include information specifying the operation rules OR of the multiple layers. For example, the operation rule of a certain layer among the multiple layers may be defined as a convolution function, and the operation rule of another layer may be defined as a pooling function.

[0232] The structural information [1] of the network N[1] may include information that the network N[1] is composed of three layers L[1], L[2] and L[3] connected in series, and information that the operation rules of the layers are OR[1], OR[2] and OR[3] respectively. In addition, the structural information [1] may include the following information: the inbound link of the first layer L[1] is the first link LK[1], the inbound link of the second layer L[2] is the second link LK[2], the inbound link of the third layer L[3] is the third link LK[3], the outbound link of the first layer L[1] is the second link LK[2], the outbound link of the second layer L[2] is the third link LK[3], and the outbound link of the third layer L[3] is the fourth link LK[4].

[0233] Fig.10d A method for defining multiple partial networks based on a network N[k] according to an embodiment of the present invention is shown.

[0234] When P partial networks are defined based on the network N[k], each partial network may be labeled as a partial network PN[k][p] (p=1, . . . , P, where P is a natural number greater than or equal to 2).

[0235] Fig.10d The example shows two partial networks PN[1][1] and partial network PN[1][2] defined based on network N[1].

[0236] Fig.10e The corresponding relationship between the network N[k] and the partial network PN[k][p] is shown.

[0237] In the following, reference is made to Fig.10d and Fig.10e Provide explanation.

[0238] In this specification, the structural information of the neural network, the structural information of the network N[k] and the structural information of the partial network PN[k][p] may be respectively referred to as the neural network structural information, structural information [k] and structural information [k][p].

[0239] According to an embodiment of the present invention, the structure information [k][p] of the partial network PN[k][p] is the same as the structure information [k] of the network N[k]. However, the size of the partial input activation value input to the partial network PN[k][p] is smaller than the size of the input activation value input to the network N[k]d.

[0240] Therefore, the partial network PN[k][p] includes a layer L[s][p] corresponding to any layer L[s] included in the network N[k]. And the partial network PN[k][p] includes a link LK[s][p] corresponding to any link LK[s] included in the network N[k].

[0241] In a preferred embodiment, the operation rules of layer L[s][p] are the same as the operation rules of layer L[s] (p=1, . . . , P).

[0242] At this time, the activation value transmitted through the link LK[s][p] can be a part of the activation value transmitted through the link LK[s]. That is, the activation value [s][p] transmitted through the link LK[s][p] can be a part of the activation value [s] transmitted through the link LK[s]. For example, Fig.10d In the embodiment, the input activation value [1][1] transmitted via the link LK[1][1] may be a part of the input activation value [1] transmitted via the link LK[1], and the input activation value IA[1][2] transmitted via the link LK[1][2] may be the rest of the input activation value [1] transmitted via the link LK[2].

[0243] At this time, the operation rule OR[s][p] of layer L[s][p] can be the same as the operation rule OR[s] of layer L[s]. Therefore, the index p can be deleted from the operation rule OR[s][p] of layer L[s][p] and marked as the operation rule OR[s].

[0244] Therefore, layer L[s][p] can be considered to be the same as layer L[s] (p=1, . . . , P).

[0245] At this time, Fig.10e In the example, the size of the network N[1] may be larger than the size of the partial network PN[1][p]. Here, the size of the network may refer to the size of the memory capacity required to define the network and the size of the computing resources required to perform network functions.

[0246] At this time, the sizes of the two partial networks PN[k][p1] generated from the network N[k] and the size of the partial network PN[k][p2] may be the same as or different from each other.

[0247] <Method for performing neural network operation in user computing device>

[0248] Fig.11a , Fig.11b and Fig.11c According to the comparative example, Figure 2 A method for performing neural network operations in a user computing device.

[0249] In the following, Fig.11a , Fig.11b and Fig.11c They may be collectively referred to as FIG11.

[0250] Hereinafter, in this specification and the drawings, symbol IA represents an input activation value, and OA represents an output activation value.

[0251] The following will be described based on FIG. 11. Fig.10b The process shown is the process of generating the output activation value OA[3] from the input activation value IA[1].

[0252] FIG. 11 shows reference symbol s, and lists steps S101 to S106 .

[0253] When the reference number s is set to 1 and steps S101 to S106 are executed, Fig.10b The output activation value OA[1] is stored in the external memory (130).

[0254] Then, when the reference number s is set to 2 and steps S101 to S106 are executed, Fig.10b The output activation value OA[2] is stored in the external memory (130).

[0255] Then, when the reference number s is set to 3 and steps S101 to S106 are executed, Fig.10b The output activation value OA[3] is stored in the external memory (130).

[0256] Hereinafter, steps S101 to S106 are described in detail.

[0257] In step S101, the control unit (40) and the DMA unit (20) may read the weight [s] from the external memory (130) via the bus (700) and store the weight [s] in the second memory bank of the internal memory (30).

[0258] In step S102, the control unit (40) and the DMA unit (20) can read the input activation value IA[s] from the external memory (130) through the bus (700) and store it in the first storage bank of the internal memory (30).

[0259] In step S103, the control unit (40) may provide the input activation value IA[s] stored in the first memory bank to the data operation unit (610).

[0260] In step S104, the control unit (40) may provide the weight [s] stored in the second memory bank to the data calculation unit (610).

[0261] In step S105, the data operation unit (610) can generate an output activation value OA[s] based on the input activation value IA[s] and the weight [s] according to the operation rules of the layer [s], and the control unit (40) can store the output activation value OA[s] in the first storage body.

[0262] In step S106, the control unit (40) and the DMA unit (20) may store the output activation value OA[s] in the external memory (130) via the bus (700).

[0263] At this time, in the process of generating the output activation value OA[3] based on the input activation value IA[1], the input activation value IA[s] and the output activation value OA[s] are transmitted multiple times (s=1, 2, 3) through the bus (700).

[0264] Fig.12a , Figure 12b , Fig.12c , Fig.13a , Fig.13b and Fig.13c The present invention is shown in an embodiment of the present invention. Figure 2 A method for performing neural network operations in a user computing device.

[0265] In the following, Fig.12a , Figure 12b and Fig.12c can be collectively referred to as Figure 12, and Fig.13a , Fig.13b and Fig.13c They can be collectively referred to as Figure 13.

[0266] 12 and 13 will be used to illustrate the Fig.10b or Fig.10d The process of generating the output activation value OA[3] from the input activation value IA[1] shown in FIG.

[0267] Here, the input activation value IA[1] may be divided into an input activation value IA[1][1] and an input activation value IA[1][2] based on a row.

[0268] FIG. 12 shows reference symbol s, and lists steps S210 to S215 .

[0269] In step S210, the control unit (40) and the DMA unit (20) can read the data from the external memory (130) through the bus (700). Fig.10d The weights [s] of the operation rules of the layers included in the network N[1] are stored in the second storage bank (s=1, 2, 3) of the internal memory (30).

[0270] If at least some of the operation rules of the layers included in the network N[1] do not use weights, the corresponding weights may not be read from the external memory (130). In one embodiment, step S210 may be unnecessary.

[0271] In step S211, the control unit (40) and the DMA unit (20) can read the input activation value IA[1][1] which is a part of the input activation value IA[1] from the external memory (130) through the bus (700) and store it in the first storage body of the internal memory (30).

[0272] At this time, the capacity of the first storage body may be smaller than the total size of the input activation value IA[1], and may be larger than the size of the input activation value IA[1][1].

[0273] exist Figure 12b In the example, when the reference number s is set to 1 and steps S212 to S214 are executed, the generated Fig.10d The output activation value OA[1][1] is stored in the first storage body of the internal memory (30).

[0274] exist Figure 12b , when the reference number s is set to 2 and steps S212 to S214 are executed, Fig.10d The output activation value OA[2][1] is stored in the first storage body of the internal memory (30).

[0275] exist Figure 12b , when the reference number s is set to 3 and steps S212 to S214 are executed, Fig.10d The output activation value OA[3][1] is stored in the first storage body of the internal memory (30).

[0276] In the process of changing the reference number s to 1, 2, and 3 in sequence and repeatedly executing steps S212 to S214, the bus (700) is neither used nor the external memory (130) is accessed.

[0277] Hereinafter, steps S212 to S214 are described in detail.

[0278] In step S212, the control unit (40) may provide the input activation value IA[s][1] stored in the first memory bank to the data operation unit (610).

[0279] In step S213, the control unit (40) may provide the weight [s] stored in the second memory bank to the data operation unit (610).

[0280] In step S214, the data operation unit (610) may generate an output activation value OA[s][1] based on the input activation value IA[s][1] and the weight [s] according to the operation rule of layer [s][1], and the control unit (40) may store the output activation value OA[s][1] in the first storage body. At this time, the operation rule of layer [s][1] may be the same as the operation rule of layer [s].

[0281] exist Fig.12c In step S215, the control unit (40) and the DMA unit (20) can store the output activation value OA[3][1] in the external memory (130) via the bus (700).

[0282] FIG. 13 shows reference symbol s, and lists steps S221 to S225.

[0283] Figure 12 is from Fig.10d The process of generating the output activation value OA[3][1] from the input activation value IA[1][1] and storing it in the external memory (130) is shown in FIG. 13. Fig.10d The process of generating an output activation value OA[3][2] from an input activation value IA[1][2] and storing it in an external memory (130). When the output activation value OA[3][1] is combined with the output activation value OA[3][2], Fig.10d The output activation value OA[3].

[0284] In step S221, the control unit (40) and the DMA unit (20) can read the input activation value IA[1][2] which is the remaining part of the input activation value IA[1] from the external memory (130) through the bus (700) and store it in the first storage body of the internal memory (30).

[0285] At this time, the capacity of the first storage body may be smaller than the total size of the input activation value IA[1], and may be larger than the size of the input activation value IA[1][2].

[0286] exist Fig.13b , when the reference number s is set to 1 and steps S222 to S224 are executed, Fig.10d The output activation value OA[1][2] is stored in the first storage body of the internal memory (30).

[0287] exist Fig.13b , when the reference number s is set to 2 and steps S222 to S224 are executed, Fig.10d The output activation value OA[2][2] is stored in the first storage body of the internal memory (30).

[0288] exist Fig.13b , when the reference number s is set to 3 and steps S222 to S224 are executed, Fig.10d The output activation value OA[3][2] is stored in the first storage body of the internal memory (30).

[0289] In the process of changing the reference number s to 1, 2, and 3 in sequence and repeatedly executing steps S222 to S224, the bus (700) is not used and the external memory (130) is not accessed.

[0290] Steps S222 to S224 are described in detail.

[0291] In step S222, the control unit (40) may provide the input activation value IA[s][2] stored in the first memory bank to the data operation unit (610).

[0292] In step S223, the control unit (40) may provide the weight [s] stored in the second memory bank to the data operation unit (610).

[0293] In step S224, the data operation unit (610) may generate an output activation value OA[s][2] based on the input activation value IA[s][2] and the weight [s] according to the operation rule of layer [s][2], and the control unit (40) may store the output activation value OA[s][2] in the first storage body. At this time, the operation rule of layer [s][2] may be the same as the operation rule of layer [s].

[0294] exist Fig.13c In step S225, the control unit (40) and the DMA unit (20) can store the output activation value OA[3][2] in the external memory (130) via the bus (700).

[0295] Fig.14 is a schematic diagram illustrating a neural network operation method provided according to an embodiment of the present invention.

[0296] Reference Fig.14 For illustration, a neural network operation method provided according to an embodiment of the present invention may include steps S1 to S5.

[0297] In step S1, the control unit (40) and the DMA unit (20) can read the input activation value IA[1][p] (i.e., the input activation value IA[s=1][p]) from the external memory (130) through the bus and store it in the internal memory (30).

[0298] In step S2, the data operation unit (610) can calculate the output activation value OA[1][p] based on the input activation value IA[1][p] stored in the internal memory (30) according to the operation rule OR[1] of the layer L[1][p], and the control unit (40) can store the output activation value OA[1][p] in the internal memory (30).

[0299] In step S3, the data operation unit (610) can calculate the output activation value OA[2][p] based on the output activation value OA[1][p] stored in the internal memory (30) according to the operation rule OR[2] of the layer L[2][p], and the control unit (40) can store the output activation value OA[2][p] in the internal memory (30).

[0300] In step S4, the data operation unit (610) can calculate the output activation value OA[3][p] based on the output activation value OA[2][p] stored in the internal memory (30) according to the operation rule OR[3] of the layer L[3][p], and the control unit (40) can store the output activation value OA[3][p] in the internal memory (30).

[0301] In step S5, the control unit (40) and the DMA unit (20) can store the output activation value OA[3][p] (ie, the output activation value [s=3][p]) in the external memory (130) via the bus.

[0302] Here, the input activation value IA[1] may be divided into a total of P parts, and the input activation value IA[1][p] may be a part of the input activation value IA[1].

[0303] Here, the operation rule of layer [s] [p] may be the same as the operation rule of layer [s].

[0304] If steps S1 to S5 are repeated P times in the range of p=1, ..., P, the output activation value OA[3] of layer L[3] generated according to the operation rules of a series of layers L[1], L[2], and L[3] can be completed based on the input activation value IA[1].

[0305] exist Fig.14 In the figure, for the sake of convenience, it is illustrated that the partial network PN[1][p] includes three layers, but the number of layers included in the partial network PN[1][p] is not limited to this.

[0306] Fig.15 and Fig.16 is a flowchart of a neural network operation method provided according to an embodiment of the present invention.

[0307] The neural network operation method is a neural network operation method executed in an NPU (110), wherein the NPU includes a DMA unit (20), an internal memory (30), a data operation unit (operation device) and a control unit (control unit).

[0308] 610 and the control unit (40). The NPU (110) may be included in a user computing device (1), which may include a main processor (160), an external memory (DRAM) (130), a bus (700) and other hardware (99).

[0309] The neural network operation method may include the following steps S800: NPU uses the input activation value IA[1] composed of P input activation values ​​IA[1][p] distinguished by rows or columns, and repeats a predetermined first process [p] (p=1, ..., P, P is a natural number greater than or equal to 2) from p=1 to p=P.

[0310] At this time, the first process [p] may include the following steps:

[0311] In step S810, the DMA unit (20) and the control unit (40) may read the input activation value IA[1][p] from the external memory (130) via the bus (70) and store it in the first storage unit of the internal memory (30).

[0312] In step S820, the control unit (40) may store in the first storage body an output activation value OA[1][p] generated by operating the input activation value IA[1][p] stored in the first storage body according to the operation rule of layer [1].

[0313] In step S830, the control unit (40) may repeat, from s=1 to s=L-1, the second process (L is a natural number greater than or equal to 2) of storing the output activation value OA[s+1][p] generated by operating the output activation value OA[s][p] stored in the first storage body according to the operation rule of the layer [s+1] connected to the output end of the layer [s] in the first storage body.

[0314] In step S840, the DMA unit (20) and the control unit (40) may record the output activation value OA[L][p] stored in the first memory bank to the external memory (130) via the bus (700).

[0315] At this time, the layer [1] and the layer [s+1] (s=1, ..., L-1) may be included in the neural network.

[0316] Furthermore, the input activation value IA[1] may be an input activation value input to the neural network layer [1].

[0317] The input activation value IA[1] can be a tensor containing multiple rows.

[0318] At this time, before sequentially repeating step S800 of the first process [p], the neural network operation method may further include the following step S790: the DMA unit (20) and the control unit (40) read the weight [1] of the operation rule for the layer [1] and the weight [s] of the operation rule for the layer [s] from the external memory (130) via the bus 700.

[0319] (s=1, ..., L-1), and stored in the second storage bank of the internal memory (20).

[0320] In one embodiment, the output activation value OA[1][p] can be generated based on the input activation value IA[1][p] stored in the first memory bank and the weight [1] stored in the second memory bank, and the output activation value OA[s+1][p] can be generated based on the output activation value OA[s][p] stored in the first memory bank and the weight [s+1] stored in the second memory bank (s=1,...,L-1).

[0321] At this time, the NPU may include an instruction file having instruction codes for executing the steps S790 and S800.

[0322] After repeating the first process [p], information about the address of the external memory in which the output activation value OA[L][p] (p=1, . . . , P) is recorded may have been recorded in the instruction file used by the NPU.

[0323] In the step S800, the activation value composed of the output activation value OA[L][p] (p=1, ..., P) may be the input activation value IA[L+1] input to the layer [L+1]. At this time, the input activation value IA[L+1] may be distinguished by Q input activation values ​​IA[L+1][q] distinguished by row or column (q=1, ..., Q, Q is a natural number greater than or equal to 2).

[0324] According to an embodiment of the present invention, Fig.16 As shown, after repeating step S800 of the first process [p], the neural network operation method may further include the following step S900: using the input activation value IA[L+1], repeating the predetermined third process [q] from q=1 to q=Q in sequence.

[0325] At this time, the third process [q] may include the following steps:

[0326] In step S910, the DMA unit (20) and the control unit (40) can read the input activation value IA[L+1][q] from the external memory (130) through the bus (700) and store it in the first storage body of the internal memory (30).

[0327] In step S920, the control unit (40) may store in the first storage body the output activation value OA[L+1][q] generated by operating the input activation value IA[L+1][q] stored in the first storage body according to the operation rules of layer [L+1].

[0328] In step S930, the control unit (40) may repeat the fourth process (M is a natural number greater than or equal to L+2) from s=L+1 to s=M-1, storing the output activation value OA[s][q] stored in the first storage body according to the operation rule of the layer [s+1] connected to the output end of the layer [s], thereby generating the output activation value OA[s+1][q] stored in the first storage body.

[0329] In step S940, the DMA unit (20) and the control unit (40) may record the output activation value OA[M][q] stored in the first memory bank to the external memory (130) via the bus (700).

[0330] At this time, the layer [L+1] and the layer [s+1] (s=L+1, ..., M-1) can be included in the neural network.

[0331] By utilizing the embodiments of the present invention described above, those skilled in the art can easily implement various changes and modifications without departing from the essential characteristics of the present invention. The content of each claim of the claims can be combined with other claims that have no reference relationship within the scope of the present specification.

[0332] <Acknowledgements>

[0333] This invention was completed with the support of the following national research and development projects: [Project unique number] 2002, [Project number] 20-CM-BD-02, [Supervisory department] Ministry of Trade, Industry and Energy, [Project management (professional) organization] Civil-Military Cooperation Promotion Institute of the National Defense Science Research Institute, [Research project name] Civil-Military Dual-Use Technology Development, [Research project name] Development of AI accelerator (NPU) in edge SoC and middleware development for semantic information processing of image acquisition, [Contribution rate] 1 / 1,

[0334] [Project implementing agency] Openedges Technology, INC., [Research period] December 24, 2020 to December 23, 2023.

Claims

1. A method for generating an NPU instruction comprises the following steps: The computing device generates a p-th partial network, the p-th partial network having the same structure as a first network structure defined by a first group of layers included in the predefined neural network; The computing device determines, in a first memory included in another computing device, a p-th read address which is a location storing an address of a p-th partial input activation value as data to be input to the most upstream layer of the p-th partial network; The computing device determines, in the first memory, a p-th write address, which is a location where an address of a p-th partial output activation value, which is data output by the most downstream layer of the p-th partial network, should be stored; and The computing device generates an NPU instruction [p], the NPU instruction comprising a first instruction set, a second instruction set and a third instruction set, The first instruction set includes instructions for causing the NPU included in the other computing device to read the p-th part of input activation value from the first memory based on the p-th read address and store it in the internal memory of the NPU. The second instruction set includes instructions for causing the NPU to generate the p-th part output activation value based on the p-th part input activation value stored in the internal memory, The third instruction set includes instructions for causing the NPU to store the p-th partial output activation value in the first memory based on the p-th write address.

2. A method for generating NPU instructions according to claim 1, The first memory is a memory provided outside the NPU, The pth part of input activation values ​​is transmitted from the first memory to the internal memory of the NPU via the bus of the other computing device, The p-th part output activation value is transmitted from the internal memory to the first memory through the bus.

3. The NPU instruction generation method according to claim 1, characterized in that: The p-th partial output activation value is generated by calculating the p-th partial input activation value stored in the internal memory based on the calculation rules of each layer included in the p-th partial network.

4. The NPU instruction generation method according to claim 1, The step of generating the pth part network comprises: The computing device defines the first group consisting of a plurality of consecutive layers included in a predefined neural network; The computing device generates structural information about the first network composed of a plurality of layers and a plurality of links included in the defined first group; as well as The computing device generates the pth partial network having the same structure as the first network; The structural information about the first network is information about the layers constituting the first group, the operation rules of the layers, and the links of the activation value transmission paths between the layers.

5. The NPU instruction generation method according to claim 1, The first group comprises a plurality of layers, The most upstream layer is a layer among the plurality of layers that receives an activation value from outside the first group, The most downstream layer is a layer among the plurality of layers that provides an activation value to the outside of the first group.

6. The NPU instruction generation method according to claim 1, characterized in that: The p-th portion of input activation values ​​is a portion of input activation values ​​input to the most upstream layer among the layers of the first group.

7. A method for generating an NPU instruction comprises the following steps: The computing device generates a partitioned network including a p-th partial network (p is 1, 2, . . . , or P) based on a first network consisting of a first group of layers included in a predefined neural network; The computing device generates an NPU instruction [p] (p is 1, 2, . . . , or P) for the p-th part of the network to be executed by an NPU included in another computing device. The step of generating a partitioned network includes: The computing device defines a p-th slice layer (p is 1, 2, . . . , or P), the p-th slice layer receiving input activation values ​​to be input to the first group and outputting partial input activation values ​​as part of the input activation values; The computing device defines a p-th partial network (p is 1, 2, ..., or P), the p-th partial network receiving the p-th input activation value output by the p-th slice layer; The computing device defines a connection layer, which combines P partial output activation values ​​of the P partial network outputs; The computing device completes the partitioned network by defining a plurality of links, where the plurality of links represent activation value transmission paths between the P slice layers, the P partial networks, and the connection layers.

8. A method for generating NPU instructions according to claim 7, The p-th portion of input activation values ​​is a portion of input activation values ​​input to the most upstream layer of the first group of layers, And the input activation value can be reconstructed from the first part of input activation values ​​to the Pth part of input activation values.

9. A method for generating NPU instructions according to claim 7, The structure of the p-th partial network is the same as the structure of the first network (p is 1, 2, ..., or P), The step of generating an NPU instruction [p] comprises: The computing device determines, in a first memory included in the other computing device, a p-th read address which is a location at which an address of a p-th partial input activation value as data to be input to the most upstream layer of the p-th partial network is stored; The computing device determines, in the first memory, a p-th write address, which is a location where an address of a p-th partial output activation value, which is data output by the most downstream layer of the p-th partial network, should be stored; and The computing device generates an NPU instruction [p], the NPU instruction comprising a first instruction set, a second instruction set and a third instruction set, The first instruction set includes instructions for causing the NPU to read the p-th part of input activation value from the first memory based on the p-th read address and store it in the internal memory of the NPU. The second instruction set includes instructions for causing the NPU to generate the p-th part output activation value based on the p-th part input activation value stored in the internal memory, The third instruction set includes instructions for causing the NPU to store the p-th partial output activation value in the first memory based on the p-th write address.

10. A method for generating NPU instructions according to claim 9, The first memory is a memory provided outside the NPU, The pth part of input activation values ​​is transmitted from the first memory to the internal memory of the NPU via the bus of the other computing device, The p-th part output activation value is transmitted from the internal memory to the first memory through the bus.

11. A method for generating NPU instructions according to claim 7, The step of generating the pth part network comprises: The computing device defines the first group consisting of a plurality of consecutive layers included in the predefined neural network; The computing device generates structural information about the first network composed of a plurality of layers and a plurality of links included in the defined first group; as well as The computing device generates the pth partial network having the same structure as the first network, The structural information about the first network is information about the layers constituting the first group, the operation rules of the layers, and the links of the activation value transmission paths between the layers.

12. A computing device comprising: Storage Department; as well as Main processor; The storage unit stores a program including instructions, which enables the main processor to perform the following steps: Generate a p-th partial network, the p-th partial network having the same structure as a first network structure defined by a first set of layers included in a predefined neural network; determining, in a first memory included in another computing device, a p-th read address which is a location at which an address of a p-th partial input activation value as data to be input to the most upstream layer of the p-th partial network is stored; determining, in the first memory, a p-th write address, the p-th write address being a location where an address of a p-th partial output activation value as data outputted by the most downstream layer of the p-th partial network should be stored; and Generate NPU instructions [p], the NPU instructions include a first instruction set, a second instruction set and a third instruction set, The first instruction set includes instructions for causing an NPU included in the other computing device to read the p-th part of input activation values ​​from the first memory based on the p-th read address and store the p-th part of input activation values ​​in an internal memory of the NPU. The second instruction set includes instructions for causing the NPU to generate the p-th part output activation value based on the p-th part input activation value stored in the internal memory, The third instruction set includes instructions for causing the NPU to store the p-th partial output activation value in the first memory based on the p-th write address.

13. A computing device comprising: Storage unit; as well as Main processor; Wherein, a program containing instructions is stored, causing the main processor to execute the following steps: Generate a partitioned network including a p-th partial network based on a first network consisting of a first group of layers included in a predefined neural network (p is 1, 2, ..., or P); Generate an NPU instruction [p] (p is 1, 2, ..., or P) for execution by an NPU included in another computing device for the p-th part of the network; The step of generating a partitioned network comprises: The computing device defines a p-th slice layer (p is 1, 2, . . . , or P), the p-th slice layer receiving input activation values ​​to be input to the first group and outputting partial input activation values ​​as part of the input activation values; The computing device defines a p-th partial network (p is 1, 2, ..., or P), the p-th partial network receiving the p-th input activation value output by the p-th slice layer; The computing device defines a connection layer, which combines P partial output activation values ​​of the P partial network outputs; The computing device completes the partitioned network by defining P of the slice layers, P of the partial networks, and a plurality of links representing activation value transmission paths between the connection layers.