Data processing method and data processing device

By dynamically determining whether to store output data in memory based on calculation costs and data size during neural network forward processing, the method addresses the challenge of securing sufficient memory bandwidth as processor performance increases, enhancing processor efficiency and reducing system costs.

JP2025074362AActive Publication Date: 2025-05-13PREFERRED NETWORKS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025035695
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-05-13
Estimated Expiration
2040-08-26

AI Technical Summary

Technical Problem

As processor performance improves, securing sufficient memory bandwidth between the processor and external memory to transmit intermediate results of forward processing calculations becomes challenging, leading to increased system costs if new high-speed memory or interfaces are required.

Method used

A data processing method that determines whether to store output data in first memory during forward processing of a neural network based on the calculation cost of each layer and the size of the output data, thereby optimizing memory usage and reducing the need for high-speed memory or interfaces.

Benefits of technology

This approach improves the effective efficiency of the processor by maximizing the use of existing memory resources, reducing memory bandwidth requirements, and lowering system costs, while maintaining high training efficiency for neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025074362000001_ABST
    Figure 2025074362000001_ABST
Patent Text Reader

Abstract

To reduce the memory bandwidth of a memory while maximizing effective efficiency of computations.SOLUTION: A data processing method according to an embodiment of the present invention involves determining whether or not to store output data in one or more first memories in forward processing of a neural network on the basis of computation cost of each of at least one or more layers of the forward processing and the size of the output data.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a data processing method and a data processing device. [Background technology]

[0002] Deep learning training is generally performed using a processor with many built-in cores, such as a GPU (Graphics Processing Unit). When training is performed using this type of processor, the intermediate results of the forward processing calculation are usually stored in an external memory, such as a DRAM (Dynamic Random Access Memory), for backward processing. Then, during backward processing, the intermediate results (intermediate data) of the forward processing calculation required for the backward processing calculation are read from the external memory. The intermediate results of the forward processing calculation can be stored in the external memory every time because the memory bandwidth (communication band) between the processor and the external memory is sufficiently large. Summary of the Invention [Problem to be solved by the invention]

[0003] If the performance of processors improves in the future, it may become difficult to ensure a memory bandwidth between the processor and the external memory that is sufficient to transmit intermediate results of forward processing calculations to the external memory each time. If it is difficult to improve the memory bandwidth using the existing external memory or memory interface, it will be necessary to develop a new high-speed memory or a new high-speed memory interface, which will significantly increase system costs. [Means for solving the problem]

[0004] A data processing method according to an embodiment of the present invention determines, in forward processing of a neural network, whether to store the output data in one or more first memories based on the computational cost of each of at least one or more layers of the forward processing and the size of the output data. [Brief description of the drawings]

[0005] [Figure 1] 1 is a block diagram illustrating an example of a data processing device according to an embodiment of the present invention. [Diagram 2] FIG. 2 is an explanatory diagram showing an example of neural network training executed by the data processing device shown in FIG. [Diagram 3] FIG. 1 is a flow diagram showing an example of forward processing in training a neural network. [Figure 4] FIG. 1 is a flow diagram illustrating an example of backward processing and optimization processing in training a neural network. [Diagram 5] FIG. 2 is an explanatory diagram showing an example of training of a neural network by the data processing device of FIG. [Figure 6] FIG. 2 is an explanatory diagram showing another example of training a neural network by the data processing device of FIG. [Figure 7] FIG. 2 is an explanatory diagram showing yet another example of training a neural network by the data processing device of FIG. [Figure 8] FIG. 2 is an explanatory diagram showing another example of training a neural network using the data processing device of FIG. [Figure 9] FIG. 2 is a flow diagram illustrating an example of the operation of a data processing apparatus to perform training of a neural network. [Figure 10] 2 is a block diagram showing an example of a hardware configuration of the data processing device 100 in FIG. 1. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0006] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.

[0007] Fig. 1 is a block diagram showing an example of a data processing device according to an embodiment of the present invention. The data processing device 100 shown in Fig. 1 has at least one system board 10 including a processor 20 and a plurality of DRAMs (Dynamic Random Access Memories) 50 connected to the processor 20. For example, the data processing device 100 is a server. The processor 20 is an example of a computing device. The DRAM 50 is an example of a memory (external memory).

[0008] The processor 20 has a plurality of arithmetic units 30 and a plurality of SRAMs (Static Random Access Memories) 40 connected to the arithmetic units 30, respectively. The processor 20 is connected to a system bus. The processor 20 may be in the form of a chip or a package. Note that the memory connected to the processor 20 is not limited to the DRAM 50, and the memory connected to the arithmetic units 30 is not limited to the SRAM 40. The SRAM 40 is an example of an internal memory.

[0009] Thus, in this embodiment, the data processing device 100 includes multiple types of memories with different read / write speeds, namely, an SRAM 40 and a DRAM 50 that generally has a slower read / write speed than the SRAM 40. In this embodiment, for example, when training a neural network having multiple layers, it is possible to compensate for the insufficient read / write speed of the memory by not writing some of the calculation results of some layers to the DRAM 50. This makes it possible to improve the training speed.

[0010] Fig. 2 is an explanatory diagram showing an example of training of a neural network executed by the data processing device 100 shown in Fig. 1. In training of a neural network having multiple intermediate layers between an input layer and an output layer, forward processing, backward processing, and optimization processing are repeatedly executed multiple times while changing the training data. The forward processing, backward processing, and optimization processing will be described with reference to Figs. 3 and 4.

[0011] The forward processing is an example of a first processing, and the operations performed in multiple layers (multiple intermediate layers) in the forward processing are an example of multiple types of first operations. The first data is an example of data used in the first processing, and the second data is an example of data (operation result) obtained by executing the first processing. The backward processing is an example of a second processing, and the operations performed in multiple layers in the backward processing are an example of multiple types of second operations.

[0012] 3 is a flow diagram showing an example of forward processing in training a neural network. In forward processing, data and parameters such as weights are input to an input layer and each of a predetermined number of intermediate layers. In the input layer, the input data and parameter 1 are operated to generate intermediate data 1. In the intermediate layer next to the input layer, intermediate data 1 and parameter 2 are operated to generate intermediate data 2.

[0013] In the subsequent intermediate layers, the intermediate data generated by the previous intermediate layer and the parameters set for each intermediate layer are calculated, and the intermediate data generated by the calculation is output to the next intermediate layer. Note that there may be intermediate layers that do not use parameters. Examples of intermediate layers include a convolution layer, a pooling layer, and a fully connected layer.

[0014] In the output layer, output data is calculated using intermediate data N generated by the intermediate layer N (Nth layer) immediately before the output layer. In the output layer, which calculates the error in a classification problem, output data (solution) is calculated, for example, by using a softmax function as the activation function and cross entropy as the error function. In the output layer, as explained in Figure 4, the error from the correct answer (loss function) is calculated by comparing the output data with training data (correct answer data).

[0015] In this way, in forward processing, input data and parameters are calculated in each layer of the neural network to obtain data to be input to the next layer, and output data is output from the final layer (forward propagation). Note that forward processing is used not only for training neural networks, but also for inference using neural networks. Forward processing can be represented by a computation graph such as a Directed Acyclic Graph (DAG).

[0016] FIG. 4 is a flow diagram showing an example of backward processing and optimization processing in training a neural network. In backward processing, error backpropagation is performed in which errors are propagated in the reverse order of forward processing. In FIG. 4, the symbol Δ indicates data error or parameter error. The parameter update processing performed in the optimization processing is indicated by the dashed arrow.

[0017] First, in the backward processing, in the layer (output layer) for calculating the error, the output data generated in the forward processing is compared with the teacher data, and Δintermediate data N is generated, which is the error for the intermediate data N input to the output layer. Δintermediate data N is also the error of the output data output by the Nth intermediate layer.

[0018] Next, in each intermediate layer, starting from the intermediate layer closest to the output layer, an error (Δ intermediate data) for the output data and intermediate data that is input data are calculated, and a Δ parameter that is an error for the parameter of the intermediate layer is generated. The Δ parameter indicates the gradient of the parameter in the curve that indicates the change in error relative to the change in parameter. For example, in the intermediate layer adjacent to the input layer, Δ intermediate data 2 and intermediate data 1 are calculated to obtain Δ parameter 2.

[0019] In addition, in each intermediate layer, the error (Δ intermediate data) for the output data and the parameters of the intermediate layer are calculated, and Δ intermediate data, which is the error for the input data of the intermediate layer, is generated. The error (Δ intermediate data) for the input data of the intermediate layer is also the error of the output data of the previous intermediate layer (or input layer). For example, in the intermediate layer adjacent to the input layer, Δ intermediate data 2 and parameters 2 are calculated to obtain Δ intermediate data 1.

[0020] In the input layer, similarly to the intermediate layer, Δ intermediate data 1 and input data are calculated to obtain Δ parameter 1, and Δ intermediate data 1 and parameter 1 are calculated to obtain Δ input data, which is an error for the input data. In this way, in the backward processing, intermediate data, which is an intermediate result of the calculation by the forward processing, is required.

[0021] In the optimization process, the parameters are corrected in each intermediate layer and input layer using the Δ parameter (gradient of error) obtained in the backward process. In other words, the parameters are optimized. The parameter optimization is performed using a gradient descent method such as Momentum-SGD (Stochastic Gradient Descent) or ADAM.

[0022] In this way, in the backward processing, the error of the data input to the output layer (the output data of the intermediate layer immediately before the output layer) is calculated from the output data and the teacher data. Then, the process of calculating the error of the intermediate data using the calculated data error and the process of calculating the parameter error using the intermediate data error are performed in order starting from the output side layer (error backpropagation). In the parameter update process, the parameters are optimized based on the parameter error obtained in the backward processing.

[0023] 5 to 8 are explanatory diagrams showing an example of training of a neural network by the data processing device 100 of FIG. 1. For ease of explanation, it is assumed that the neural network to be trained has layers L1, L2, L3 and a layer Loss. In the computation graph, the layer L1 receives input data D0, and the output of the layer L1 is connected to the input of the layer L2. The output of the layer L2 is connected to the input of the layer L3, and the layer L3 outputs output data. The layer Loss calculates an error (loss function) using the output data D3 from the layer L3 and the teacher data. The explanation of FIGS. 5 to 8 is based on the following assumptions.

[0024] (1) In deep learning processing, multiple data points are generally processed in a batch. In order to consider the ratio of calculations to memory accesses in Fig. 5 to Fig. 8, it is sufficient to discuss the peak performance and memory bandwidth of the data processing device 100 (processor 20) per data point included in one batch. The data size, FLOPS (Floating-point Operations Per Second), and memory bandwidth are implicitly values ​​per data point.

[0025] (2) Layer L1 is a convolutional layer with a kernel size of 1x1, the number of input channels being "3", and the number of output channels being "128". Layer L2 is a convolutional layer with a kernel size of 3x3, the number of input channels being "128", and the number of output channels being "128". Layer L3 is a fully connected layer with the number of input channels being "128" and the number of output channels being "10". The input and output image sizes of layer L1 and the output image size of layer L2 are 32 pixels wide and 32 pixels high.

[0026] (3) As additional layers, an appropriate activation function such as ReLU (Rectified Linear Unit) is inserted after each layer L1 and L2 (convolutional layer). Average Pooling is inserted before layer L3 (fully connected layer). Each of layers L1 and L2 may be a layer that combines a convolutional layer and ReLU (Rectified Linear Unit). Layer L3 may be a layer that combines Average Pooling and a fully connected layer.

[0027] As a result, the neural network trained in Fig. 5 can execute a small-scale but practical image recognition task. Forward processing and backward processing can be executed by fusing with layers L1, L2, and L3, respectively. In other words, these layers do not cause additional access to the DRAM 50. Also, the amount of calculation of the additional layers is small enough to be ignored.

[0028] (4) The data used for training is represented in 32-bit floating-point format. The peak performance of the processor 20 is 0.5 TFLOPS (Tera Floating-point Operations Per Second). The bandwidth of the processor 20 to the DRAM 50 is 1 GB / s. The access to the DRAM 50 and the calculation by the calculator 30 can overlap to the maximum. In other words, the total elapsed time required for training is the longer of the access time to the DRAM 50 or the calculation time by the calculator 30.

[0029] 5 to 8, the letters in circles within the rectangular frames representing each layer L1, L2, L3, and Loss indicate the processing at that layer, with the first "F" indicating forward processing and the first "B" indicating backward processing. The numbers in parentheses within the rectangular frames representing each layer L1, L2, and L3 indicate examples of the cost (computational cost or operation cost) required for processing at each layer.

[0030] In forward processing, the numbers input or output for each layer L1, L2, L3, and Loss indicate the data size. In this example, for ease of explanation, the data size is assumed to be the same as the number of channels between two adjacent layers.

[0031] In the forward process, the transfer index shown under each layer L1, L2, and L3 is calculated by dividing the calculation cost of each layer by the data size output by each layer. The transfer index is one of the criteria for determining whether or not to transfer the data obtained by the forward process of each layer to the DRAM 50, and is an example of a storage value.

[0032] The larger the transfer value of a layer, the more preferable it is to transfer data obtained by calculation to the DRAM 50, and the smaller the transfer value of a layer, the more preferable it is not to transfer data obtained by calculation to the DRAM 50. Then, by determining whether or not to transfer data to the DRAM 50 for each layer based on the transfer value, it is possible to improve the effective efficiency indicated by the ratio of the calculation time by the processor 20 to the elapsed time in training the neural network. For example, the effective efficiency is calculated by dividing the minimum value (fastest value) of the calculation time by the processor 20 by the elapsed time in training the neural network. Note that in Figs. 5 to 8, the amount of data temporarily stored in the SRAM 40 to be used for the calculation of each layer is not limited. Also, data read from and written to the DRAM 50 is assumed to be via the SRAM 40.

[0033] In the training shown in FIG. 5, a transfer index threshold that determines whether or not data obtained by calculation in each layer is transferred to the DRAM 50 is not set. Therefore, in the forward processing, all data D1, D2, and D3 obtained by calculation in the layers L1, L2, and L3, respectively, are stored in the DRAM 50. Data D0 used in the layer L1 is transferred from the DRAM 50. In the backward processing, data D3, D2, D1, and D0 used in the layers Loss, L3, L2, and L1 are transferred from the DRAM 50.

[0034] If all data obtained by calculation in forward processing is stored in DRAM 50, and all data used in backward processing is read from DRAM 50, the total access time to DRAM 50 for the entire training will be 2.122 ms. Furthermore, the total calculation time by processor 20 for the entire training will be 1.817 ms. The total calculation time is the minimum value because it does not include data recalculation, which will be explained in Figure 6 and subsequent figures. Therefore, the elapsed time required for training will be the total access time to DRAM 50, which is the bottleneck (2.122 ms), and the effective efficiency will be 85.6% (1.817 / 2.122).

[0035] 6, a transfer index threshold (first threshold) that determines whether or not data obtained by calculation in each layer is transferred to DRAM 50 is set to "0.1". Therefore, in forward processing, data D2 and D3 obtained by calculation in layers L2 and L3 whose transfer index is equal to or greater than the threshold is stored in DRAM 50. Data D1 obtained by calculation in layer L1 whose transfer index is smaller than the threshold is not stored in DRAM 50. In other words, data processing device 100 thins out a portion of the data used for training and stores it in DRAM 50.

[0036] In the backward processing, data D1 used for the calculation in layer L2 is recalculated by executing forward processing of layer L1 by the processor 20. Data D3, D2, and D0 used for the calculation in layers Loss, L3, and L1 in the backward processing are transferred from the DRAM 50.

[0037] In FIG. 6, in the backward processing, data D1, which has a relatively large amount of data, is recalculated by forward processing F1 of layer L1. Since data D1 is no longer read from or written to DRAM 50, the access time of DRAM 50 during the entire training is significantly reduced compared to FIG. 5, and the total access time is 1.073 ms. Furthermore, the total calculation time by processor 20 during the entire training is slightly increased to 1.818 ms. Therefore, the elapsed time required for training is the total calculation time by processor 20, which is the bottleneck (1.818 ms), and the effective efficiency is improved from FIG. 5 to 99.9% (1.817 / 1.818).

[0038] In the training shown in FIG. 7, the threshold of the transfer index that determines whether or not data obtained by calculation in each layer is transferred to DRAM 50 is set to "1.0". However, in FIG. 7, in order to evaluate the effective efficiency, data D1 obtained by calculation in layer L1 whose transfer index is smaller than the threshold is transferred to DRAM 50. Therefore, in the forward processing, data D3 obtained by calculation in layer L3 whose transfer index is equal to or larger than the threshold, and data D1 obtained by calculation in layer L1 are stored in DRAM 50. Data D2 obtained by calculation in layer L2 whose transfer index is smaller than the threshold is not stored in DRAM 50.

[0039] In the backward processing, data D2 used for the calculation in layer L3 is recalculated by executing forward processing of layer L2 by the processor 20. Data D3, D1, and D0 used for the calculation in layers Loss, L2, and L1 in the backward processing are transferred from the DRAM 50.

[0040] In FIG. 7, in the backward processing, data D2, which has a relatively large amount of data, is recalculated by forward processing F2 of layer L2. Since data D2 is not read from or written to DRAM 50, the access time of DRAM 50 during the entire training is reduced as in FIG. 6, and the total access time is about 1 ms. On the other hand, the calculation cost of forward processing F2 of layer L2 is larger than the calculation cost of forward processing F1 of layer L1. Therefore, the total calculation time by processor 20 during the entire training is increased to 2.423 ms compared to FIG. 5 and FIG. 6. Therefore, the elapsed time is the total calculation time (2.423 ms) by processor 20, which is the bottleneck, and the effective efficiency is worse than FIG. 5, at 75.0% (1.817 / 2.423).

[0041] 8, the threshold of the transfer index that determines whether or not data obtained by calculation in each layer is transferred to the DRAM 50 is set to "1.0." Therefore, in the forward process, data D3 obtained by calculation in layer L3, whose transfer index is equal to or greater than the threshold, is stored in the DRAM 50. Data D1 and D2 obtained by calculation in layers L1 and L2, whose transfer indexes are smaller than the threshold, are not stored in the DRAM 50.

[0042] In the backward processing, data D2 used for the calculation in layer L3 is recalculated by sequentially executing forward processing of layers L1 and L2 by processor 20. In the backward processing, data D1 used for the calculation in layer L2 is the data stored in SRAM 40 in the forward processing of layer L1. Data D0 used for the calculation in layers Loss and L1 in the backward processing is transferred from DRAM 50.

[0043] The operation shown in Fig. 8 is a combination of the operations shown in Fig. 6 and Fig. 7, and the calculation cost of the processor 20 in the backward processing is higher than that in Fig. 7. Therefore, the effective efficiency is lower than 75.0% in Fig. 7.

[0044] As described above, it can be seen that the effective efficiency (99.9%) in Fig. 6 is the highest among the examples shown in Fig. 5 to Fig. 8. In this way, by setting the threshold value of the transfer index that determines whether or not to transfer data obtained by calculations in each layer to the DRAM 50 to an appropriate value, it is possible to reduce the memory bandwidth of the DRAM 50 while maximizing the effective efficiency of the processor 20.

[0045] Since the amount of data transferred to the DRAM 50 can be reduced, the capacity of the DRAM 50 mounted on the data processing device 100 can be reduced. For example, an inexpensive DRAM 50 with a small capacity can be adopted. As a result, the cost of the data processing device 100 can be reduced. In other words, the data processing device 100 with the reduced capacity of the DRAM 50 can also improve the effective efficiency of the processor 20.

[0046] Furthermore, since the capacity of the DRAM 50 can be reduced, it is possible to reduce the power consumption of the data processing device 100. Furthermore, if the memory bandwidth of the DRAM 50 is not reduced, a higher performance processor 20 can be used.

[0047] Fig. 9 is a flow diagram showing an example of the operation of the data processing device 100 that executes training of a neural network. For example, the flow shown in Fig. 9 may be realized by the data processing device 100 (processor 20) executing a data processing program. Fig. 9 shows an example of a data processing method and a data processing program.

[0048] First, in step S10, the data processing device 100 selects a layer for performing forward processing. Next, in step S12, the data processing device 100 performs forward processing of the selected layer.

[0049] Next, in step S14, the data processing device 100 determines whether or not it is worth storing the data obtained by the forward processing in the layer in the DRAM 50. For example, when the transfer index (FIGS. 5 to 8) in the layer targeted for the forward processing is equal to or greater than a preset threshold, the data processing device 100 executes step S16 to store the data in the DRAM 50. When the transfer index in the layer targeted for the forward processing is smaller than the threshold, the data processing device 100 executes step S18 without storing the data in the DRAM 50.

[0050] In step S16, the data processing device 100 stores the data obtained by the forward process in the DRAM 50, and executes step S18. In step S18, if there is a next layer to perform the forward process, the data processing device 100 executes step S10, and if there is no next layer to perform the forward process, the data processing device 100 executes step S20.

[0051] In step S20, the data processing device 100 selects a layer on which the backward processing is to be performed. Next, in step S22, the data processing device 100 determines whether or not data to be used for the backward processing is stored in the DRAM 50. If the data is stored in the DRAM 50, the data processing device 100 executes step S24, and if the data is non-stored data not stored in the DRAM 50, the data processing device 100 executes step S26.

[0052] In step S24, the data processing device 100 reads data to be used in the backward processing from the DRAM 50, and executes step S28. In step S26, the data processing device 100 executes forward processing to generate data to be used in the backward processing (non-stored data), and executes step S28.

[0053] In step S28, the data processing device 100 executes backward processing of the layer to be calculated using the data obtained in step S24 or step S26. Next, in step S30, the data processing device 100 executes step S20 if there is a next layer to execute backward processing on, and ends the operation shown in FIG. 9 if there is no next layer to execute backward processing on.

[0054] A part or the whole of the data processing device 100 in the above-mentioned embodiment may be configured with hardware, or may be configured with information processing of software (programs) executed by a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), etc. In the case of being configured with information processing of software, software for realizing at least a part of the functions of each device in the above-mentioned embodiment may be stored in a non-transient storage medium (non-transient computer-readable medium) such as a flexible disk, a CD-ROM (Compact Disc-Read Only Memory), or a USB (Universal Serial Bus) memory, and the software information processing may be executed by reading the software into a computer. The software may also be downloaded via a communication network. Furthermore, the software may be implemented in a circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array), so that the information processing may be executed by hardware.

[0055] The type of storage medium that stores software such as a data processing program is not limited. The storage medium is not limited to a removable medium such as a magnetic disk or an optical disk, but may be a fixed storage medium such as a hard disk or a memory. The storage medium may be provided inside the computer or outside the computer.

[0056] Fig. 10 is a block diagram showing an example of a hardware configuration of the data processing device 100 in Fig. 1. As an example, the data processing device 100 may be realized as a computer including a processor 20, a DRAM (main storage device) 50, an auxiliary storage device 60 (memory), a network interface 70, and a device interface 80, which are connected via a bus 90. For example, the processor 20 executes a data processing program to perform the training described in Figs. 5 to 8.

[0057] The data processing device 100 includes one of each component, but may include multiple of the same components. Although FIG. 10 shows one data processing device 100, the software may be installed in multiple devices including the data processing device 100, and each of the multiple data processing devices 100 may execute the same or different parts of the software. In this case, the data processing device 100 may be in the form of distributed computing in which each of the data processing devices 100 communicates with each other via a network interface 70 or the like to execute processing. In other words, the data processing device 100 may be configured as a computer system that realizes functions by executing instructions stored in one or more storage devices by one or more data processing devices 100. The data processing device 100 may also be configured to process information transmitted from a terminal by one or more data processing devices 100 provided on a cloud, and transmit the processing results to the terminal.

[0058] The operations described in Figures 5 to 8 and the operations described in the flow of Figure 9 may be executed in parallel using one or more processors 20, or using multiple computers via the communication network 200. Also, various calculations may be distributed to multiple calculation cores in the processor 20 and executed in parallel. Also, a part or all of the processes, means, etc. disclosed herein may be executed by at least one of a processor and a storage device provided on a cloud that can communicate with the data processing device 100 via a network. In this way, a computer system including the data processing device 100 may be in the form of parallel computing using one or more computers.

[0059] The processor 20 may be an electronic circuit (such as a processing circuit, processing circuitry, CPU, GPU, FPGA, or ASIC) including a computer control device and an arithmetic device. The processor 20 may also be a semiconductor device including a dedicated processing circuit. The processor 20 is not limited to an electronic circuit using electronic logic elements, and may be realized by an optical circuit using optical logic elements. The processor 20 may also include an arithmetic function based on quantum computing.

[0060] The processor 20 can perform arithmetic processing based on data and software (programs) input from each device, etc., in the internal configuration of the data processing device 100, and output arithmetic results and control signals to each device, etc. The processor 20 can control each component constituting the data processing device 100 by executing the OS (Operating System) of the data processing device 100, applications, etc.

[0061] The data processing device 100 may be realized by one or more processors 20. Here, the processor 20 may refer to one or more electronic circuits provided on one chip, or to one or more electronic circuits provided on two or more chips or two or more devices. When multiple electronic circuits are used, the electronic circuits may communicate with each other by wire or wirelessly.

[0062] The main memory device 50 is a storage device that stores instructions executed by the processor 20 and various data, and information stored in the main memory device 50 is read by the processor 20. The auxiliary memory device 60 is a storage device other than the main memory device 50. These storage devices refer to any electronic components capable of storing electronic information, and may be semiconductor memories. The semiconductor memories may be either volatile or non-volatile memories. The storage device for saving various data in the data processing device 100 may be realized by the main memory device 50 or the auxiliary memory device 60, or may be realized by an internal memory such as an SRAM 40 built into the processor 20.

[0063] The data processing device 100 is not limited to the configuration shown in FIG. 1. A plurality of processors 20 may be connected (coupled) to one storage device (memory), or a single processor 20 may be connected. A plurality of storage devices (memories) may be connected (coupled) to one processor 20. When the data processing device 100 is configured with at least one storage device (memory) and a plurality of processors 20 connected (coupled) to the at least one storage device (memory), the data processing device 100 may include a configuration in which at least one processor 20 among the plurality of processors 20 is connected (coupled) to at least one storage device (memory). This configuration may also be realized by the storage devices (memories) and processors 20 included in a plurality of data processing devices 100. Furthermore, the data processing device 100 may include a configuration in which the storage device (memory) is integrated with the processor 20 (for example, a cache memory including an L1 cache and an L2 cache).

[0064] The network interface 70 is an interface for connecting to the communication network 200 wirelessly or by wire. The network interface 70 may be an appropriate interface, such as one conforming to an existing communication standard. The network interface 70 may exchange information with an external device 210 connected via the communication network 200. The communication network 200 may be any one of a wide area network (WAN), a local area network (LAN), a personal area network (PAN), etc., or a combination thereof, as long as information is exchanged between the data processing device 100 and the external device 210. An example of a WAN is the Internet, an example of a LAN is IEEE802.11 or Ethernet (registered trademark), and an example of a PAN is Bluetooth (registered trademark) or Near Field Communication (NFC).

[0065] The device interface 80 is an interface such as a USB that directly connects to the external device 220 .

[0066] The external device 220 may be connected to the data processing device 100 via a network, or may be connected directly to the data processing device 100.

[0067] The external device 210 or the external device 220 may be, for example, an input device. The input device is, for example, a device such as a camera, a microphone, a motion capture device, various sensors, a keyboard, a mouse, or a touch panel, and provides acquired information to the data processing device 100. Alternatively, the input device may be a device equipped with an input unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.

[0068] Moreover, the external device 210 or the external device 220 may be, for example, an output device. The output device may be, for example, a display device such as an LCD (Liquid Crystal Display), a CRT (Cathode Ray Tube), a PDP (Plasma Display Panel), or an organic EL (Electro Luminescence) panel, or may be a speaker that outputs sound or the like. Alternatively, the output device may be a device including an output unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.

[0069] Furthermore, the external device 210 or the external device 220 may be a storage device (memory). For example, the external device 210 may be a network storage device or the like, and the external device 220 may be a storage device such as a HDD. The external device 220, which is a storage device (memory), is an example of a recording medium readable by a computer such as the processor 20.

[0070] Furthermore, the external device 210 or the external device 220 may be a device having some of the functions of the components of the data processing device 100. In other words, the data processing device 100 may transmit or receive some or all of the processing results of the external device 210 or the external device 220.

[0071] As described above, in this embodiment, by setting the threshold value of the transfer index that determines whether or not data obtained by calculations in each layer is transferred to the DRAM 50 to an appropriate value, it is possible to reduce the memory bandwidth of the DRAM 50 while maximizing the effective efficiency of the processor 20. This makes it possible to reduce the amount of DRAM 50 mounted on the data processing device 100, and thus the cost of the data processing device 100.

[0072] The transfer index is calculated by dividing the computation cost of each layer by the data size output by each layer. Therefore, the transfer index can be easily calculated regardless of the complexity of the neural network (computation graph). The method of determining whether to transfer the computation result by the processor 20 to the DRAM 50 based on the transfer index is not limited to neural network training and can be applied to other data processing.

[0073] In this specification (including the claims), when the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used, it includes any of a, b, c, ab, ac, bc, or abc. It may also include multiple instances of any element, such as aa, abb, aabbcc, etc. Furthermore, it also includes the addition of elements other than the enumerated elements (a, b, and c), such as having d, as in abcd.

[0074] In this specification (including claims), when expressions such as "data as input / based on / according to / in response to" (including similar expressions) are used, unless otherwise specified, it includes cases where various data itself is used as input, and cases where various data that have been processed in some way (e.g., noise-added, normalized, intermediate representation of various data, etc.) are used as input. In addition, when it is stated that a result is obtained "based on / according to / in response to data," it includes cases where the result is obtained based only on the data in question, and may also include cases where the result is obtained by being influenced by other data, factors, conditions, and / or states other than the data in question. In addition, when it is stated that "data is output," it includes cases where various data itself is used as output, and cases where various data that have been processed in some way (e.g., noise-added, normalized, intermediate representation of various data, etc.) are output, unless otherwise specified.

[0075] In this specification (including the claims), the terms "connected" and "coupled" are intended as open-ended terms including any of direct connection / coupling, indirect connection / coupling, electrically connection / coupling, communicatively connection / coupling, functionally connection / coupling, physically connection / coupling, etc. The terms should be interpreted appropriately depending on the context in which the terms are used, but any connection / coupling form that is not intentionally or naturally excluded should be interpreted as being included in the terms without any restriction.

[0076] In this specification (including the claims), when the expression "A configured to B" is used, it may include that the physical structure of element A has a configuration capable of performing operation B, and that the permanent or temporary setting / configuration of element A is configured / set to actually perform operation B. For example, when element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and is configured to actually perform operation B by setting a permanent or temporary program (instruction). Also, when element A is a dedicated processor or dedicated arithmetic circuit, it is sufficient that the circuit structure of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached.

[0077] In this specification (including the claims), terms implying containing or possessing (e.g., "comprising / including" and "having") are intended to be open-ended terms that include containing or possessing things other than the object designated by the object of the term. When the object of such terms implies no quantity or a singular number (such as an article "a" or "an"), the expression should be construed as not being limited to a specific number.

[0078] In this specification (including the claims), even if expressions such as "one or more" or "at least one" are used in some places and expressions that do not specify a quantity or suggest a singular number (expressions using the articles a or an) are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or suggest a singular number (expressions using the articles a or an) should be interpreted as not necessarily being limited to a specific number.

[0079] In this specification, when a particular advantage / result is described as being obtained from a particular configuration of an embodiment, it should be understood that the same effect can also be obtained from one or more other embodiments having the same configuration, unless otherwise stated. However, it should be understood that the presence or absence of the effect generally depends on various factors, conditions, and / or states, and that the effect is not necessarily obtained by the configuration. The effect is merely obtained by the configuration described in the embodiment when various factors, conditions, and / or states are satisfied, and the effect is not necessarily obtained in the invention according to the claim that specifies the configuration or a similar configuration.

[0080] In this specification (including the claims), when a term such as "maximize" is used, it includes finding a global maximum, finding an approximation of a global maximum, finding a local maximum, and finding an approximation of a local maximum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding an approximation of these maxima probabilistically or heuristically. Similarly, when a term such as "minimize" is used, it includes finding a global minimum, finding an approximation of a global minimum, finding a local minimum, and finding an approximation of a local minimum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding an approximation of these minima probabilistically or heuristically. Similarly, when a term such as "optimize" is used, it includes finding a global optimum, finding an approximation of a global optimum, finding a local optimum, and finding an approximation of a local optimum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding an approximation of these optimums probabilistically or heuristically.

[0081] In this specification (including claims), when a plurality of pieces of hardware perform a predetermined process, each piece of hardware may cooperate to perform the predetermined process, or a portion of the hardware may perform all of the predetermined process. Also, a portion of the hardware may perform a portion of the predetermined process, and another piece of hardware may perform the remainder of the predetermined process. In this specification (including claims), when an expression such as "one or more pieces of hardware perform a first process, and the one or more pieces of hardware perform a second process" is used, the hardware performing the first process and the hardware performing the second process may be the same or different. In other words, it is sufficient that the hardware performing the first process and the hardware performing the second process are included in the one or more pieces of hardware. The hardware may include an electronic circuit, or a device including an electronic circuit.

[0082] In this specification (including the claims), when multiple storage devices (memories) store data, each of the multiple storage devices (memories) may store only a portion of the data, or may store the entire data.

[0083] Although the embodiments of the present disclosure have been described in detail above, the present disclosure is not limited to the individual embodiments described above. Various additions, modifications, replacements, and partial deletions are possible within the scope of the conceptual idea and intent of the present invention derived from the contents defined in the claims and their equivalents. For example, in all the above-mentioned embodiments, when numerical values ​​or formulas are used in the explanation, they are shown as examples and are not limited to these. In addition, the order of each operation in the embodiment is shown as an example and is not limited to these. [Explanation of symbols]

[0084] 20 processors 30 Arithmetic unit 40 SRAM 50 DRAM 60 Auxiliary storage 70 Network Interface 80 Device Interfaces 90 Bus 100 Data processing device 200 Communication Network 210 External device 220 External device

Claims

1. In a forward process of the neural network, determining whether to store the output data in one or more first memories based on a computation cost of each of at least one or more layers of the forward process and a size of the output data; Data processing methods.

2. The output data is stored in the one or more first memories based on a value obtained by dividing the computation cost by a size of the output data.

2. The data processing method according to claim 1.

3. the one or more first memories are DRAMs; 3. The data processing method according to claim 1.

4. The output data stored in the one or more first memories is used in performing backward processing subsequent to the forward processing. The data processing method according to any one of claims 1 to 3.

5. At least a portion of the output data that is not stored in the one or more first memories is stored in one or more second memories. The data processing method according to any one of claims 1 to 4.

6. performing a recalculation using at least a portion of the output data stored in the one or more second memories; The data processing method according to claim 5.

7. the one or more second memories are SRAMs; 7. The data processing method according to claim 5 or 6.

8. The optimization process of the neural network is performed by ADAM. The data processing method according to any one of claims 1 to 7.

9. The forward processing is performed by a GPU. The data processing method according to any one of claims 1 to 8.

10. The forwarding process is performed by a plurality of processors; the one or more first memories are coupled to the plurality of processors; The data processing method according to any one of claims 1 to 9.

11. In a forward process of the neural network, determining whether to store the output data in one or more first memories based on a computation cost of each of at least one or more layers of the forward process and a size of the output data; Data processing device.

12. The output data is stored in the one or more first memories based on a value obtained by dividing the computation cost by a size of the output data. A data processing apparatus according to claim 11.

13. the one or more first memories are DRAMs; A data processing device according to claim 11 or 12.

14. The output data stored in the one or more first memories is used in performing backward processing subsequent to the forward processing. A data processing device according to any one of claims 11 to 13.

15. At least a portion of the output data that is not stored in the one or more first memories is stored in one or more second memories. A data processing apparatus according to any one of claims 11 to 14.

16. performing a recalculation using at least a portion of the output data stored in the one or more second memories; 16. A data processing apparatus according to claim 15.

17. the one or more second memories are SRAMs; 17. A data processing device according to claim 15 or 16.

18. The optimization process of the neural network is performed by ADAM. A data processing apparatus according to any one of claims 11 to 17.

19. The forward processing is performed by a GPU. A data processing apparatus according to any one of claims 11 to 18.

20. The forwarding process is performed by a plurality of processors; the one or more first memories are coupled to the plurality of processors; 20. A data processing apparatus according to any one of claims 11 to 19.

Citation Information

Patent Citations

  • Data storage determination program, data storage determination method, and data storage determination apparatus

    JP2017162342A

  • Data processing method, data processing device, data processing system and data processing program

    JP7648352B2

  • Model generation device, model generation method, and program

    WO2019182059A1

  • Modifying machine learning models to improve locality

    WO2020076392A1