Data processing method, data processing device, data processing program, and method for generating neural network model

By compressing intermediate data during machine learning processes and storing it in memory, the method addresses the challenge of securing sufficient memory bandwidth, reducing memory usage and system costs while maintaining processing accuracy.

JP7684127B2Active Publication Date: 2025-05-27PREFERRED NETWORKS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021121506
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-26
Publication Date
2025-05-27
Estimated Expiration
2041-07-26

AI Technical Summary

Technical Problem

It is challenging to secure sufficient memory bandwidth between the processor and external memory required to store intermediate data generated during machine learning processes in external memory each time.

Method used

A data processing method that compresses data during the first calculation process, stores the compressed data in memory, and executes a second calculation process using the compressed data, thereby reducing memory usage and ensuring sufficient memory bandwidth.

Benefits of technology

This approach reduces memory usage and ensures sufficient memory bandwidth for storing intermediate data, suppressing the increase in system cost without compromising the efficiency and accuracy of backward processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007684127000001
    Figure 0007684127000001
  • Figure 0007684127000002
    Figure 0007684127000002
  • Figure 0007684127000003
    Figure 0007684127000003
Patent Text Reader

Abstract

To reduce a use amount of a memory which holds data during calculation for first calculation processing to reduce a memory bandwidth, in a data processing method which executes second calculation processing by using the data during the calculation for the first calculation processing.SOLUTION: A data processing method about a machine learning model: compresses data during calculation for first calculation processing to generate compression data; stores the generated compression data in a memory region; and executes second calculation processing by using the compression data stored in the memory region.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a data processing method, a data processing device, a data processing program, and a method for generating a neural network model. [Background technology]

[0002] In general, in data processing such as model training in machine learning, intermediate data that is an intermediate result of forward processing calculations may be stored in an external memory such as a DRAM (Dynamic Random Access Memory) for backward processing. Among the intermediate data stored in the external memory, intermediate data required for the backward processing calculations may be read from the external memory, and the backward processing calculations may be executed. Summary of the Invention [Problem to be solved by the invention]

[0003] It may be difficult to secure sufficient memory bandwidth between the processor and the external memory required to store intermediate data generated for such machine learning in the external memory each time. [Means for solving the problem]

[0004] A data processing method according to an embodiment of the present invention is a data processing method relating to a machine learning model, which compresses data in the middle of a first calculation process to generate compressed data, stores the generated compressed data in a memory area, and executes a second calculation process using the compressed data stored in the memory area. [Brief description of the drawings]

[0005] [Figure 1] 1 is a block diagram showing an example of a data processing device according to a first embodiment of the present invention. [Diagram 2] FIG. 2 is an explanatory diagram showing an example of neural network training executed by the data processing device shown in FIG. [Diagram 3] FIG. 11 is an explanatory diagram showing an example of forward processing in training the neural network of this embodiment. [Figure 4] FIG. 1 is an explanatory diagram showing an example of backward processing and optimization processing in training the neural network of this embodiment. [Diagram 5] FIG. 4 is an explanatory diagram illustrating an example of a determination table used in forward processing in the first embodiment; [Figure 6] 2 is a flow diagram showing an example of forward processing of a neural network by the data processing device of FIG. 1. [Figure 7] FIG. 11 is an explanatory diagram illustrating an example of a decision table used in forward processing of a neural network by the data processing device of the second embodiment of the present invention. [Figure 8] FIG. 11 is a flow diagram illustrating an example of forward processing of a neural network by the data processing device of the second embodiment. [Figure 9] FIG. 9 is a flow diagram showing a continuation of the forward process of FIG. 8. [Figure 10] FIG. 2 is a block diagram showing an example of a hardware configuration of the data processing device of the above-described embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0006] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.

[0007] Fig. 1 is a block diagram showing an example of a data processing device according to a first embodiment of the present invention. The data processing device 100 shown in Fig. 1 has at least one system board 10 including a processor 20 and a plurality of DRAMs (Dynamic Random Access Memories) 50 connected to the processor 20. For example, the data processing device 100 may be a server.

[0008] The processor 20 has a plurality of arithmetic units 30 and a plurality of static random access memories (SRAMs) 40 connected to the plurality of arithmetic units 30, respectively. The processor 20 is connected to a system bus. The processor 20 may be in the form of a chip or a package. The arithmetic units 30 are an example of an arithmetic processing unit.

[0009] In this embodiment, the memory bandwidth of SRAM 40 is greater than that of DRAM 50. For this reason, if data used in calculations by processor 20 can be stored in SRAM 40, it is preferable to store it in SRAM 40. However, when SRAM 40 is built into processor 20, it may be difficult to store all data used by processor 20 in SRAM 40. In this case, data that cannot be stored in SRAM 40 may be stored in DRAM 50, which has a smaller memory bandwidth.

[0010] The internal memory connected to the arithmetic unit 30 is not limited to the SRAM 40 and may be, for example, a cache memory. The external memory connected to the processor 20 is not limited to the DRAM 50 and may be, for example, a magnetoresistive random access memory (MRAM), a hard disk drive (HDD), or a solid state drive (SSD). The SRAM 40 is an example of a first memory, and the memory area allocated to the SRAM 40 is an example of the first memory area. The DRAM 50 is an example of a second memory, and the memory area allocated to the DRAM 50 is an example of the second memory area.

[0011] In this manner, the data processing device 100 of this embodiment has a plurality of types of memories (in this embodiment, the SRAM 40 and the DRAM 50) with different memory bandwidths.

[0012] If the SRAM 40 has a sufficient memory capacity and can be mounted on the processor 20 or the system board 10, the first memory area and the second memory area may be allocated to the SRAM 40.

[0013] The data processing device 100 performs a plurality of calculation processes to perform training of a neural network having a plurality of layers. One of the calculation processes is, for example, forward processing of the neural network, and another of the calculation processes is backward processing of the neural network. In addition, the calculation process performed by the data processing device 100 is not limited to training of the neural network. For example, the data processing device 100 may perform calculation processes such as scientific and technical calculations.

[0014] FIG. 2 is an explanatory diagram showing an example of training of a neural network executed by the data processing device 100 shown in FIG. 1. For example, FIG. 2 shows an example of a method for generating a neural network model by machine learning. In the machine learning for training a neural network having a plurality of intermediate layers between an input layer and an output layer of the present embodiment, a forward process, a backward process, and an optimization process are repeatedly executed a plurality of times while changing the training data. Then, the data processing device 100 generates a neural network model based on the training. The forward process, the backward process, and the optimization process will be described with reference to FIG. 3 and FIG. 4. The forward process is an example of a first calculation process, and the backward process is an example of a second calculation process. The forward process, the backward process, and the optimization process are an example of a third calculation process. In this specification, the generation of a model or the generation of a neural network includes adjusting the model or neural network parameters.

[0015] FIG. 3 is an explanatory diagram showing an example of forward processing in training the neural network of this embodiment. For example, FIG. 3 shows an example of forward processing in a method for generating a neural network model by machine learning. In forward processing, data and parameters such as weights are input to an input layer and a predetermined number of intermediate layers. In the input layer, the input data and parameter 1 are calculated to generate intermediate data 1. In the intermediate layer next to the input layer, the intermediate data 1 and parameter 2 are calculated to generate intermediate data 2.

[0016] In the subsequent intermediate layers, the intermediate data generated by the previous intermediate layer and the parameters set for each intermediate layer are calculated, and the intermediate data generated by the calculation is output to the next intermediate layer. Note that there may be intermediate layers that do not use parameters. Examples of intermediate layers include a convolution layer, a pooling layer, and a fully connected layer.

[0017] In this embodiment, intermediate data generated by the calculation processing of the input layer and the middle layer is stored in the SRAM 40 without being compressed. Then, the middle layer and the output layer that execute the calculation processing read out the uncompressed intermediate data from the SRAM 40 and use it for the calculation processing. The data processing device 100 can execute the forward processing without degrading the calculation accuracy by using the uncompressed intermediate data in the forward processing. The intermediate data is an example of data generated in the middle of calculation by each layer through the calculation processing.

[0018] This intermediate data is also used in the backward processing described in FIG. 4. In this embodiment, the intermediate data used in the backward processing is compressed and then stored in the SRAM 40, the DRAM 50, or both the SRAM 40 and the DRAM 50. In the backward processing, the intermediate data is compressed and then stored in the memory, thereby reducing the memory usage. This makes it possible to ensure a sufficient memory bandwidth between the processor 20 and the DRAM 50, which is necessary for storing the intermediate data generated by the forward processing in the DRAM 50, without mounting a faster DRAM 50 and without widening the data bus width. In other words, the memory bandwidth required for storing the intermediate data in the DRAM 50 can be reduced compared to the memory bandwidth required for storing all the intermediate data in the DRAM 50. This makes it possible to suppress an increase in the system cost of the data processing device 100.

[0019] For example, the intermediate data used in the backward processing may be lossy compressed. Lossy compression has a lower compression cost than lossless compression and may be able to keep the compression ratio constant, so that the load on the processor 20 due to the compression processing can be reduced.

[0020] In addition, in the calculation of backward processing that uses intermediate data obtained in forward processing, errors in the intermediate data often only have a local effect and do not propagate and accumulate over a wide area. For example, in the backward processing of a convolutional layer, the intermediate data generated by forward processing only affects the weight gradient of that convolutional layer.

[0021] In addition, the gradient value calculated by the backward process may not require high accuracy compared to the forward process. For example, in updating the weights in the stochastic gradient descent method, the gradient value is expected to be smaller than the weight value, so even if the relative error of the gradient is large, the influence on the calculation of the backward process can be reduced. Therefore, even when performing the backward process using compressed intermediate data, appropriate weights can be calculated.

[0022] The data processing device 100 can perform the above-mentioned operations, for example, by using conversion of floating-point number data. Specifically, the calculation process in the forward process may be performed using double-precision floating-point number data, and the intermediate data generated may be compressed by converting the double-precision floating-point number data to single-precision floating-point number data. The data processing device 100 may also compress the intermediate data by converting the single-precision floating-point number data to 8-bit fixed-point number data. This allows the intermediate data to be lossily compressed easily using an existing conversion method. Furthermore, the data processing device 100 may compress the intermediate data by reducing the number of bits (number of digits) of the mantissa of the floating-point number data.

[0023] The compression rate of the intermediate data may be set higher as the training of the neural network progresses. That is, the training of the neural network shown in FIG. 2 may be repeatedly performed while gradually increasing the compression rate of the intermediate data. For example, the data processing device 100 may calculate the first predetermined number of iterations using single-precision floating-point number data, and the next predetermined number of iterations using half-precision floating-point number data. The data processing device 100 may further calculate the next 100 iterations using 8-bit fixed-point number data. The predetermined number of iterations is, for example, 100 iterations.

[0024] Furthermore, when floating-point number data is used in forward processing, the data processing device 100 may gradually increase the compression ratio of the intermediate data by sequentially decreasing the number of bits of the mantissa each time a predetermined number of iterations are executed. In this way, by gradually increasing the compression ratio of the intermediate data, the memory bandwidth for transferring the intermediate data can be further reduced, and an increase in the system cost of the data processing device 100 can be further suppressed.

[0025] The data processing device 100 may compress a plurality of intermediate data collectively, instead of compressing the intermediate data one by one. In this case, it may be possible to further increase the compression rate of the intermediate data, which can contribute to reducing memory bandwidth and system costs.

[0026] In the output layer, output data is calculated using intermediate data N generated by the intermediate layer N (Nth layer) immediately before the output layer. In the output layer, which calculates the error in a classification problem, the output data (solution) is calculated, for example, by using a softmax function as the activation function and cross entropy as the error function. In the output layer, as explained in Figure 4, the error from the correct answer (loss function) is calculated by comparing the output data with training data (correct answer data).

[0027] In this way, in forward processing, input data and parameters are operated in each layer of the neural network to calculate data (intermediate data) to be input to the next layer, and output data is output from the final layer (forward propagation). Note that forward processing may be used not only for training a neural network, but also for inference using a neural network. Forward processing can be expressed by a computation graph such as a Directed Acyclic Graph (DAG).

[0028] FIG. 4 is an explanatory diagram showing an example of backward processing and optimization processing in training the neural network of this embodiment. For example, FIG. 4 shows an example of backward processing in the method of generating a neural network model by machine learning. In the backward processing, error backpropagation is performed in which errors are propagated in the reverse order of the forward processing. In FIG. 4, the symbol Δ indicates data error or parameter error. The parameter update processing performed in the optimization processing is indicated by a dashed arrow.

[0029] First, in the backward processing, in the layer (output layer) where the error is calculated, the output data generated in the forward processing is compared with the teacher data, and Δintermediate data N is generated, which is the error for the intermediate data N input to the output layer. Δintermediate data N is also the error of the output data output by the Nth intermediate layer.

[0030] Next, in each intermediate layer, starting from the intermediate layer closest to the output layer, an error (Δ intermediate data) for the output data and intermediate data that is input data are calculated, and a Δ parameter that is an error for the parameter of the intermediate layer is generated. The Δ parameter indicates the gradient of the parameter in the curve that indicates the change in error relative to the change in parameter. For example, in the intermediate layer adjacent to the input layer, Δ intermediate data 2 and intermediate data 1 are calculated to calculate Δ parameter 2.

[0031] In addition, in each intermediate layer, an error (Δ intermediate data) for the output data and the parameters of the intermediate layer are calculated, and Δ intermediate data, which is an error for the input data of the intermediate layer, is generated. The error (Δ intermediate data) for the input data of the intermediate layer is also the error of the output data of the previous intermediate layer (or input layer). For example, in an intermediate layer adjacent to an input layer, Δ intermediate data 2 and parameters 2 are calculated to calculate Δ intermediate data 1. Here, the intermediate data is read from SRAM 40 or DRAM 50 for each layer, for example.

[0032] In the input layer, similarly to the intermediate layer, Δ intermediate data 1 and input data are calculated to calculate Δ parameter 1, and Δ intermediate data 1 and parameter 1 are calculated to calculate Δ input data, which is an error for the input data. In this way, in the backward processing, intermediate data, which is an intermediate result of the calculation by the forward processing, is required.

[0033] In the optimization process, the parameters are corrected in each intermediate layer and input layer using the Δ parameter (gradient of error) calculated in the backward process. In other words, the parameters are optimized. The parameter optimization is performed using a gradient descent method such as Momentum-SGD (Stochastic Gradient Descent) or ADAM.

[0034] In this way, in the backward processing, the error of the data input to the output layer (the output data of the intermediate layer immediately before the output layer) is calculated from the output data and the teacher data. Then, the process of calculating the error of the intermediate data using the calculated data error and the process of calculating the parameter error using the intermediate data error are performed in order starting from the output side layer (error backpropagation). In the parameter update process, the parameters are optimized based on the parameter error obtained in the backward processing.

[0035] 5 is an explanatory diagram showing an example of a judgment table (one example of judgment information) used in the forward process of the first embodiment. For example, the judgment tables TBL1(A), TBL1(B), TBL1(C), ... shown in FIG. 5 may be allocated to a storage area (such as SRAM 40 or a register in the arithmetic unit 30) in the processor 20. In the following, when the judgment tables TBL1(A), TBL1(B), TBL1(C), ... are not distinguished from one another, they will be simply referred to as judgment table TBL1. For example, a judgment table TBL1 is provided for each of the neural networks A, B, C, ....

[0036] Each judgment table TBL1 of this embodiment has, for each processing target layer, an area for storing an input deletion bit (1 bit) which is an example of information indicating whether or not data is deleted, and an area for storing a transfer judgment bit (2 bits) which is an example of information indicating a transfer destination. The input deletion bit holds information indicating whether or not the uncompressed target intermediate data input to the processing target layer is deleted from the SRAM 40 after the calculation process of the processing target layer is executed. For example, the input deletion bit "0" indicates that the uncompressed target intermediate data is not deleted from the SRAM 40, and the input deletion bit "1" indicates that the uncompressed target intermediate data is deleted from the SRAM 40. The input deletion bit is an example of deletion information indicating whether or not the uncompressed intermediate data is deleted from the SRAM 40.

[0037] In the data processing device 100 of this embodiment, when the input deletion bit is "0", after the execution of the calculation process of the layer to be processed, the data processing device 100 continues to hold the uncompressed intermediate data, which is the calculation process result in the other layer used in the calculation process, without deleting it from the SRAM 40. Also, when the input deletion bit is "1", the data processing device 100 deletes from the SRAM 40 the uncompressed intermediate data, which is the calculation process result in the other layer used in the calculation process, after the execution of the calculation process of the layer to be processed.

[0038] The data processing device 100 of this embodiment can reduce the memory capacity of the SRAM 40 built into the processor 20 by deleting uncompressed intermediate data from the SRAM 40 when the data is no longer needed for subsequent calculation processes. Note that the uncompressed intermediate data may be used for calculation processes of multiple layers. In this case, only the input deletion bit corresponding to the layer executed most recently is set to "1". This can prevent intermediate data from being erroneously deleted from the SRAM 40 even when common intermediate data is used in multiple layers.

[0039] The transfer determination bits of this embodiment hold information indicating the transfer destination (storage destination) of the intermediate data. The transfer determination bits "00" indicate that the compressed intermediate data is to be transferred to SRAM 40. The transfer determination bits "01" indicate that the compressed intermediate data is to be transferred to DRAM 50. The transfer determination bits "10" indicate that the compressed intermediate data is to be transferred to both SRAM 40 and DRAM 50. The information indicating the transfer destination (storage destination) of the intermediate data held in the transfer determination bits is an example of storage destination information.

[0040] By providing the transfer determination bit, the data processing device 100 can easily determine the transfer destination of the compressed intermediate data for each layer. When the compressed intermediate data is transferred to only one of the SRAM 40 or the DRAM 50, that is, when the compressed intermediate data is not transferred to both the SRAM 40 and the DRAM 50, the transfer determination bit may be 1 bit. In this case, a transfer determination bit of "0" indicates a transfer to the SRAM 40, and a transfer determination bit of "1" indicates a transfer to the DRAM 50.

[0041] Each decision table TBL1 may have an input deletion bit and a transfer determination bit common to all layers to be processed. That is, the input deletion bit and the transfer determination bit may be set for each neural network. Furthermore, at least one of the decision tables TBL1 may hold a plurality of input deletion bits and a plurality of transfer determination bits corresponding to at least one of the layers to be processed. In this case, the plurality of input deletion bits and the plurality of transfer determination bits are set corresponding to each of a plurality of data or a plurality of data groups used in the corresponding layer to be processed.

[0042] Also, a plurality of judgment tables TBL1 may be provided corresponding to a plurality of types of compression ratios. For example, in forward processing of neural network A, when the compression ratio is sequentially increased every time a predetermined number of iterations are executed, a judgment table TBL1(A) is provided for each compression ratio, and the judgment table TBL1(A) corresponding to the number of iterations is referenced. Alternatively, a compression ratio table (an example of compression ratio information) showing the correspondence between a plurality of compression ratios and the number of iterations may be provided corresponding to each judgment table TBL1.

[0043] Fig. 6 is a flow diagram showing an example of forward processing of a neural network by the data processing device 100 of Fig. 1. That is, Fig. 6 shows an example of a data processing method by the data processing device 100. For example, the processing shown in Fig. 6 is realized by the processor 20 of the data processing device 100 executing a data processing program. For example, Fig. 6 shows an example of forward processing in a method for generating a neural network model by machine learning.

[0044] First, in step S10, the processor 20 transfers input data such as parameters used in a processing target layer that executes forward processing to the SRAM 40. In the input layer shown in Fig. 3, the processor 20 transfers input data 1 and parameter 1 to the SRAM 40 as input data.

[0045] Next, in step S12, the processor 20 executes forward processing using the input data transferred in step S10 to generate intermediate data. Next, in step S14, the processor 20 stores the intermediate data (uncompressed) generated in step S12 in the SRAM 40.

[0046] Next, in step S16, the processor 20 refers to the input deletion bit of the decision table TBL1 and decides whether or not to delete the uncompressed intermediate data input to the processing target layer from the SRAM 40. If the input deletion bit is "1", the processor 20 decides to delete the uncompressed intermediate data from the SRAM 40 and moves the process to step S18. If the input deletion bit is "0", the processor 20 decides not to delete the uncompressed intermediate data from the SRAM 40 and moves the process to step S20.

[0047] In step S18, the processor 20 deletes the uncompressed intermediate data input to the processing target layer from the SRAM 40, and moves the process to step S20. In step S20, the processor 20 compresses the intermediate data calculated by the forward processing of the processing target layer, and generates compressed data.

[0048] Next, in step S22, the processor 20 refers to the transfer determination bits of the determination table TBL1. If the transfer determination bits are "00", the processor 20 decides to transfer the intermediate data to the SRAM 40, and shifts the process to step S24. If the transfer determination bits are "01", the processor 20 decides to transfer the intermediate data to the DRAM 50, and shifts the process to step S28. If the transfer determination bits are "10", the processor 20 decides to transfer the intermediate data to both the SRAM 40 and the DRAM 50, and shifts the process to step S26.

[0049] In step S24, the processor 20 transfers the intermediate data compressed in step S20 to the SRAM 40, and proceeds to step S30. In step S26, the processor 20 transfers the intermediate data compressed in step S20 to the SRAM 40, and proceeds to step S28.

[0050] In step S28, the processor 20 transfers the intermediate data compressed in step S20 to the DRAM 50, and proceeds to step S30. As a result, the compressed intermediate data can be transferred to at least one of the SRAM 40 and the DRAM 50 depending on the value of the transfer determination bit.

[0051] In step S30, if there is an unprocessed layer, the processor 20 returns the process to step S10 and executes forward processing of the next layer to be processed. If there is no unprocessed layer, that is, if the forward processing of the neural network is completed, the processor 20 ends the operation shown in FIG.

[0052] As described above, in the embodiment described with reference to FIG. 1 to FIG. 6, the amount of memory usage for holding intermediate data generated by forward processing can be reduced. This makes it possible to secure a sufficient memory bandwidth between the processor 20 and the DRAM 50 required for storing intermediate data generated by forward processing in the DRAM 50 without mounting a faster DRAM 50 and without widening the data bus width. In other words, the memory bandwidth required for storing intermediate data in the DRAM 50 can be reduced compared with the memory bandwidth required for storing all intermediate data in the DRAM 50. For example, in backward processing, by compressing the intermediate data and then holding it in the memory, the amount of memory usage can be further reduced, and the memory bandwidth of the memory for storing the intermediate data can be further reduced. This makes it possible to suppress an increase in the system cost of the data processing device 100 without reducing the efficiency and accuracy of backward processing.

[0053] It should be noted that the first process and the second process in this specification are not limited to the forward process and the backward process in training a machine learning model.

[0054] In the intermediate layer and the output layer that perform the calculation process, the data processing device 100 reads the intermediate data stored in the SRAM 40 in an uncompressed manner and uses the read uncompressed intermediate data for the calculation process. By using the uncompressed intermediate data for the forward process, the forward process can be performed without degrading the calculation accuracy.

[0055] The data processing device 100 can reduce the memory capacity of the SRAM 40 built into the processor 20 by deleting uncompressed intermediate data from the SRAM 40 when the data is no longer needed for subsequent calculation processing. By providing an input deletion bit for each layer to be processed, it is possible to prevent intermediate data from being erroneously deleted from the SRAM 40 even when common intermediate data is used in a plurality of layers.

[0056] By providing the transfer determination bit, the data processing device 100 can easily determine the transfer destination of the compressed intermediate data for each layer.

[0057] By lossy compressing the intermediate data used in the backward processing, the compression cost can be reduced compared to lossless compression, and the compression rate can be made constant, so that the load on the processor 20 due to the compression processing can be reduced.

[0058] The intermediate data can be easily lossily compressed by expressing the intermediate data in a floating-point data format and compressing the intermediate data by reducing the number of bits (number of digits) of the mantissa.

[0059] By setting the compression rate of the intermediate data higher as the calculation processing of the layer progresses, the memory bandwidth for transferring the intermediate data can be further reduced, and the increase in the system cost of the data processing device 100 can be further suppressed.

[0060] 7 is an explanatory diagram showing an example of a decision table used in forward processing of a neural network by a data processing device according to the second embodiment of the present invention. Detailed description of elements similar to those in FIG. 5 will be omitted.

[0061] The data processing device that refers to the decision tables TBL2(A), TBL2(B), TBL2(C), ... in Fig. 7 has a configuration similar to that of the data processing device 100 shown in Fig. 1. That is, the data processing device of this embodiment has a processor 20 including a plurality of arithmetic units 30 and a plurality of SRAMs 40, and at least one system board 10 including a DRAM 50.

[0062] In the following, when there is no need to distinguish between the judgment tables TBL2(A), TBL2(B), TBL2(C), ..., they will be simply referred to as judgment table TBL2. For example, like the judgment table TBL1, the judgment table TBL2 is provided for each of the neural networks A, B, C, ....

[0063] The judgment table TBL2 has an area for storing a compression judgment bit added to the judgment table TBL1 in Fig. 5. That is, each judgment table TBL2 has an area for storing an input deletion bit (1 bit), a compression judgment bit (1 bit), and a forwarding judgment bit (2 bits) for each processing target layer.

[0064] The compression determination bit "0" indicates that the intermediate data is compressed, and the compression determination bit "1" indicates that the intermediate data is not compressed. That is, in this embodiment, it is possible to switch between compression and non-compression of the intermediate data for each layer to be processed. For example, when the size of the intermediate data to be generated is large, the compression determination bit is set to "0", and when the size of the intermediate data to be generated is small, the compression determination bit is set to "1". As a result, when the size of the intermediate data is large, it is possible to suppress an increase in the memory bandwidth and to suppress an increase in the system cost of the data processing device. On the other hand, when the size of the intermediate data is small, it is possible to reduce the compression cost.

[0065] When the compression determination bit is "0", the meaning of each value of the transfer determination bit is the same as the meaning of each value of the transfer determination bit in the determination table TBL1 of Fig. 5. That is, the transfer determination bits "00" indicate that the compressed intermediate data is to be transferred to the SRAM 40. The transfer determination bits "01" indicate that the compressed intermediate data is to be transferred to the DRAM 50. The transfer determination bits "10" indicate that the compressed intermediate data is to be transferred to both the SRAM 40 and the DRAM 50.

[0066] On the other hand, when the compression determination bit is "0", the meaning of each value of the transfer determination bit is as follows: The transfer determination bits "00" indicate that uncompressed intermediate data is to be transferred to the DRAM 50. The transfer determination bits "01" indicate that uncompressed intermediate data is not to be transferred to the DRAM 50.

[0067] The uncompressed intermediate data is always transferred to the SRAM 40. Therefore, when the compression determination bit is "1" and the transfer determination bit is "00", the uncompressed intermediate data is transferred to both the SRAM 40 and the DRAM 50. When the compression determination bit is "1" and the transfer determination bit is "01", the uncompressed intermediate data is transferred only to the SRAM 40.

[0068] In this embodiment, the meaning of the transfer determination bit changes depending on whether the intermediate data is compressed by the compression determination bit. That is, the transfer determination bit can be used in common whether the intermediate data is compressed or not, and the increase in the size of the determination table TBL2 can be suppressed.

[0069] As in the above-described embodiment, each judgment table TBL2 may have an input deletion bit, a compression determination bit, and a forwarding determination bit that are common to all layers to be processed. That is, the input deletion bit, the compression determination bit, and the forwarding determination bit may be set for each neural network. Furthermore, at least one of the judgment tables TBL2 may hold a plurality of input deletion bits, a plurality of compression determination bits, and a plurality of forwarding determination bits corresponding to at least one of the layers to be processed. In this case, the plurality of input deletion bits, the plurality of compression determination bits, and the plurality of forwarding determination bits are set corresponding to each of a plurality of data or a plurality of data groups used in the corresponding layer to be processed.

[0070] Also, a plurality of judgment tables TBL2 may be provided corresponding to a plurality of types of compression ratios. For example, in the forward processing of the neural network A, when the compression ratio is sequentially increased for each predetermined number of iterations, a judgment table TBL2(A) is provided for each compression ratio, and the judgment table TBL2(A) corresponding to the number of iterations is referenced. Alternatively, a compression ratio table showing the correspondence between a plurality of compression ratios and the number of iterations may be provided corresponding to each judgment table TBL2.

[0071] Furthermore, when the compressed intermediate data is transferred to only one of SRAM 40 or DRAM 50, i.e., when the compressed intermediate data is not transferred to both SRAM 40 and DRAM 50, the transfer determination bit may be 1 bit. In this case, when the compression determination bit is "0", the transfer determination bit "0" indicates transfer to SRAM 40, and the transfer determination bit "1" indicates transfer to DRAM 50. When the compression determination bit is "1", the transfer determination bit "0" indicates transfer to DRAM 50, and the transfer determination bit "1" indicates no transfer to DRAM 50.

[0072] 8 and 9 are flow diagrams showing an example of forward processing of a neural network by the data processing device of the second embodiment. That is, FIG. 8 and FIG. 9 show an example of a data processing method by the data processing device. For example, the processing shown in FIG. 8 and FIG. 9 is realized by the processor 20 of the data processing device executing a data processing program. For example, FIG. 8 and FIG. 9 show an example of forward processing in a method for generating a neural network model by machine learning. The same step numbers are assigned to the same processing as FIG. 6, and detailed description is omitted.

[0073] The process from step S10 to step S18 is similar to the process from step S10 to step S18 in Fig. 6. However, after step S16 and step S18, the process proceeds to step S19 instead of step S20.

[0074] In step S19, the processor 20 refers to the compression decision bit in the decision table TBL2 and decides whether or not to compress the intermediate data generated by the forward process of the processing target layer. If the processor 20 decides to compress the intermediate data, the process proceeds to step S20, and if the processor 20 decides not to compress the intermediate data, the process proceeds to step S21.

[0075] Here, the decision as to whether or not to compress may be made depending on the hardware configuration of data processing device 100 and the configuration of the neural network. For example, the hardware configuration may be represented by the storage capacity and memory bandwidth of SRAM 40, the storage capacity and memory bandwidth of DRAM 50, and the processing performance of arithmetic unit 30. For example, the configuration of the neural network may be represented by a calculation procedure of the neural network, or may be represented by a calculation graph showing the neural network.

[0076] The process of step S20 is the same as the process of step S20 in Fig. 6. As in the above-described embodiment, the data processing device 100 may compress the generated intermediate data by converting the double-precision floating-point number data into single-precision floating-point number data. The data processing device 100 may compress the intermediate data by converting the single-precision floating-point number data into 8-bit fixed-point number data. Furthermore, the data processing device 100 may compress the intermediate data by reducing the number of bits (number of digits) of the mantissa part of the floating-point number data.

[0077] Furthermore, the compression rate of the intermediate data may be set higher as the training of the neural network progresses. The data processing device 100 may compress a plurality of intermediate data collectively, instead of compressing the intermediate data one by one.

[0078] After step S20, the process proceeds to step S22 in Fig. 9. In step S21, the processor 20 refers to the transfer determination bits in the determination table TBL2. If the transfer determination bits are "00", the processor 20 proceeds to step S28 in Fig. 9, and if the transfer determination bits are "01", the processor 20 proceeds to step S30 in Fig. 9.

[0079] The processing from step S22 to step S30 in Fig. 9 is the same as the processing from step S22 to step S30 in Fig. 6. However, the processing of step S28 is also executed when it is determined in step 21 in Fig. 8 that the intermediate data is to be transferred to the DRAM 50. The processing of step S30 is also executed when it is determined in step 21 in Fig. 8 that the intermediate data is not to be transferred to the DRAM 50.

[0080] As described above, the embodiment shown in Figures 7 to 9 can also provide the same effects as the above-mentioned embodiment. For example, by compressing the intermediate data and then storing it in memory, the memory usage can be reduced and the transfer time of the intermediate data can be shortened. This makes it possible to reduce the memory bandwidth required for transferring the intermediate data, and to suppress an increase in the system cost of the data processing device 100.

[0081] Furthermore, in this embodiment, by providing a compression determination bit in the determination table TBL2, the data processing device 100 can switch between compressing and not compressing the intermediate data for each layer to be processed. This makes it possible to suppress an increase in memory bandwidth when the size of the intermediate data is large, and to suppress an increase in the system cost of the data processing device. On the other hand, when the size of the intermediate data is small, the compression cost can be reduced.

[0082] By changing the meaning of the transfer determination bit depending on whether or not the intermediate data is compressed using the compression determination bit, the transfer determination bit can be used in common whether or not the intermediate data is compressed, thereby suppressing an increase in the size of the determination table TBL2.

[0083] A part or all of the data processing device in the above-mentioned embodiment may be configured with hardware, or may be configured with information processing of software (programs) executed by a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), etc. In the case of being configured with information processing of software, software for realizing at least a part of the functions of each device in the above-mentioned embodiment may be stored in a non-transient storage medium (non-transient computer-readable medium) such as a flexible disk, a CD-ROM (Compact Disc-Read Only Memory), or a USB (Universal Serial Bus) memory, and the software information processing may be executed by reading the software into a computer. The software may also be downloaded via a communication network. Furthermore, the software may be implemented in a circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array), so that the information processing is executed by hardware.

[0084] The type of storage medium that stores software such as a data processing program is not limited. The storage medium is not limited to a removable medium such as a magnetic disk or an optical disk, but may be a fixed storage medium such as a hard disk or a memory. The storage medium may be provided inside the computer or outside the computer.

[0085] Fig. 10 is a block diagram showing an example of a hardware configuration of the data processing device of the above-mentioned embodiment. The hardware configuration of the data processing device 100 of Fig. 1 will be described below. As an example, the data processing device 100 may be realized as a computer including a processor 20, a main storage device 50 (e.g., DRAM 50), an auxiliary storage device 60 (memory), a network interface 70, and a device interface 80, which are connected via a bus 90. For example, the processor 20 executes a data processing program to perform the operations described in Fig. 6 or Figs. 8 to 9.

[0086] The data processing device 100 includes one of each component, but may include multiple of the same components. Although FIG. 10 shows one data processing device 100, the software may be installed in multiple data processing devices 100, and each of the multiple data processing devices 100 may execute the same or different parts of the software. In this case, the data processing device 100 may be in the form of distributed computing in which each of the data processing devices 100 communicates via a network interface 70 or the like to execute the processing. That is, the data processing device 100 in the above-mentioned embodiment may be configured as a computer system that realizes a function by one or more data processing devices 100 executing instructions stored in one or more storage devices. Also, the data processing device 100 may be configured to process information transmitted from a terminal in one or more data processing devices 100 provided on a cloud, and transmit the processing results to the terminal.

[0087] The operations described in the flow of FIG. 6 and the operations described in the flow of FIG. 8 to FIG. 9 may be executed in parallel using one or more processors 20, or using multiple computers via a network. Also, various calculations may be distributed to multiple arithmetic cores in the processor 20 and executed in parallel. Also, a part or all of the processes, means, etc. disclosed herein may be executed by at least one of a processor and a storage device provided on a cloud that can communicate with the data processing device 100 via a network. In this way, the data processing device 100 in the above-mentioned embodiment may be in the form of parallel computing using one or more computers.

[0088] The processor 20 may be an electronic circuit (such as a processing circuit, processing circuitry, CPU, GPU, FPGA, or ASIC) including a computer control device and an arithmetic device. The processor 20 may also be a semiconductor device including a dedicated processing circuit. The processor 20 is not limited to an electronic circuit using electronic logic elements, and may be realized by an optical circuit using optical logic elements. The processor 20 may also include an arithmetic function based on quantum computing.

[0089] The processor 20 may perform arithmetic processing based on data and software (programs) input from each device, etc., in the internal configuration of the data processing device 100, and may output arithmetic results and control signals to each device, etc. The processor 20 may control each component constituting the data processing device 100 by executing the OS (Operating System) of the data processing device 100, applications, etc.

[0090] The data processing device 100 in the above-described embodiment may be realized by one or more processors 20. Here, the processor 20 may refer to one or more electronic circuits arranged on one chip, or to one or more electronic circuits arranged on two or more chips or two or more devices. When multiple electronic circuits are used, the electronic circuits may communicate with each other by wire or wirelessly.

[0091] The main memory device 50 (for example, the DRAM 50 in FIG. 1) may store instructions executed by the processor 20 and various data, and information stored in the main memory device 50 may be read by the processor 20. The auxiliary memory device 60 is a memory device other than the main memory device 50. These memory devices refer to any electronic components capable of storing electronic information, and may be semiconductor memories. The semiconductor memories may be either volatile or non-volatile memories. The memory device for saving various data in the data processing device 100 in the above-described embodiment may be realized by the main memory device 50 or the auxiliary memory device 60, or may be realized by an internal memory such as the SRAM 40 built into the processor 20.

[0092] When the data processing device 100 in the above-described embodiment is configured with at least one storage device (memory) and a plurality of processors 20 connected (coupled) to the at least one storage device (memory), a plurality of processors 20 may be connected (coupled) to one storage device (memory), or one processor 20 may be connected. Also, a plurality of storage devices (memories) may be connected (coupled) to one processor 20, or one storage device (memory) may be connected (coupled). Also, a configuration in which at least one processor 20 of the plurality of processors 20 is connected (coupled) to at least one storage device (memory) may be included. Also, this configuration may be realized by the storage devices (memories) and processors 20 included in a plurality of data processing devices 100. Furthermore, a configuration in which the storage device (memory) is integrated with the processor 20 (for example, a cache memory including an L1 cache and an L2 cache) may be included.

[0093] The network interface 70 is an interface for connecting to the communication network 200 wirelessly or by wire. The network interface 70 may be an appropriate interface, such as one conforming to an existing communication standard. The network interface 70 may exchange information with an external device 210 connected via the communication network 200. The communication network 200 may be any one of a wide area network (WAN), a local area network (LAN), a personal area network (PAN), etc., or a combination thereof, as long as information is exchanged between the data processing device 100 and the external device 210. An example of a WAN is the Internet, an example of a LAN is IEEE802.11 or Ethernet (registered trademark), and an example of a PAN is Bluetooth (registered trademark) or Near Field Communication (NFC).

[0094] The device interface 80 is an interface such as a USB that directly connects to the external device 220 .

[0095] The external device 220 may be connected to the data processing device 100 via a network, or may be connected directly to the data processing device 100.

[0096] The external device 210 or the external device 220 may be, for example, an input device. The input device is, for example, a device such as a camera, a microphone, a motion capture device, various sensors, a keyboard, a mouse, or a touch panel, and provides acquired information to the data processing device 100. Alternatively, the input device may be a device equipped with an input unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.

[0097] Moreover, the external device 210 or the external device 220 may be, for example, an output device. The output device may be, for example, a display device such as an LCD (Liquid Crystal Display), a CRT (Cathode Ray Tube), a PDP (Plasma Display Panel), or an organic EL (Electro Luminescence) panel, or may be a speaker that outputs sound or the like. Alternatively, the output device may be a device including an output unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.

[0098] Furthermore, the external device 210 or the external device 220 may be a storage device (memory). For example, the external device 210 may be a network storage device or the like, and the external device 220 may be a storage device such as a HDD. The external device 220, which is a storage device (memory), is an example of a recording medium readable by a computer such as the processor 20.

[0099] Furthermore, the external device 210 or the external device 220 may be a device having some of the functions of the components of the data processing device 100 in the above-described embodiment. In other words, the data processing device 100 may transmit or receive some or all of the processing results of the external device 210 or the external device 220.

[0100] In this specification (including the claims), when the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used, it includes any of a, b, c, ab, ac, bc, or abc. It may also include multiple instances of any element, such as aa, abb, aabbcc, etc. Furthermore, it also includes the addition of elements other than the enumerated elements (a, b, and c), such as having d, as in abcd.

[0101] In this specification (including claims), when expressions such as "using data as input / based on / according to / in response to data" (including similar expressions) are used, unless otherwise specified, it includes cases where various data itself is used as input, and cases where various data that have been processed in some way (e.g., noise-added, normalized, feature values ​​extracted from data, intermediate representations of various data, etc.) are used as input. In addition, when it is stated that a result is obtained "using data as input / based on / according to / in response to data" (including similar expressions), it includes cases where the result is obtained based only on the data, and may also include cases where the result is obtained by being influenced by other data, factors, conditions, and / or states other than the data. In addition, when it is stated that "data is output" (including similar expressions), unless otherwise specified, it includes cases where various data itself is used as output, and cases where various data that have been processed in some way (e.g., noise-added, normalized, feature values ​​extracted from data, intermediate representations of various data, etc.) are output.

[0102] In this specification (including the claims), the terms "connected" and "coupled" are intended as open-ended terms including any of direct connection / coupling, indirect connection / coupling, electrically connection / coupling, communicatively connection / coupling, functionally connection / coupling, physically connection / coupling, etc. The terms should be interpreted appropriately depending on the context in which the terms are used, but any connection / coupling form that is not intentionally or naturally excluded should be interpreted as being included in the terms without any restriction.

[0103] In this specification (including the claims), when the expression "A configured to B" is used, it may include that the physical structure of element A has a configuration capable of performing operation B, and that the permanent or temporary setting / configuration of element A is configured / set to actually perform operation B. For example, when element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and is configured to actually perform operation B by setting a permanent or temporary program (instruction). Also, when element A is a dedicated processor or dedicated arithmetic circuit, it is sufficient that the circuit structure of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached.

[0104] In this specification (including the claims), terms implying containing or possessing (e.g., "comprising / including" and "having") are intended to be open-ended terms that include containing or possessing things other than the object designated by the object of the term. When the object of such terms implies no quantity or a singular number (such as an article "a" or "an"), the expression should be construed as not being limited to a specific number.

[0105] In this specification (including the claims), even if expressions such as "one or more" or "at least one" are used in some places and expressions that do not specify a quantity or suggest a singular number (expressions using the articles a or an) are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or suggest a singular number (expressions using the articles a or an) should be interpreted as not necessarily being limited to a specific number.

[0106] In this specification, when a particular advantage / result is described as being obtained from a particular configuration of an embodiment, it should be understood that the same effect can also be obtained from one or more other embodiments having the same configuration, unless otherwise stated. However, it should be understood that the presence or absence of the effect generally depends on various factors, conditions, and / or states, and that the effect is not necessarily obtained by the configuration. The effect is merely obtained by the configuration described in the embodiment when various factors, conditions, and / or states are satisfied, and the effect is not necessarily obtained in the invention according to the claim that specifies the configuration or a similar configuration.

[0107] In this specification (including the claims), when a term such as "maximize" is used, it includes finding a global maximum, finding an approximation of a global maximum, finding a local maximum, and finding an approximation of a local maximum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding an approximation of these maxima probabilistically or heuristically. Similarly, when a term such as "minimize" is used, it includes finding a global minimum, finding an approximation of a global minimum, finding a local minimum, and finding an approximation of a local minimum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding an approximation of these minima probabilistically or heuristically. Similarly, when a term such as "optimize" is used, it includes finding a global optimum, finding an approximation of a global optimum, finding a local optimum, and finding an approximation of a local optimum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding an approximation of these optimums probabilistically or heuristically.

[0108] In this specification (including claims), when a plurality of hardware pieces perform a predetermined process, the hardware pieces may cooperate to perform the predetermined process, or some of the hardware pieces may perform all of the predetermined process. Also, some of the hardware pieces may perform part of the predetermined process, and other hardware pieces may perform the rest of the predetermined process. In this specification (including claims), when an expression such as "one or more hardware pieces perform a first calculation process, and the one or more hardware pieces perform a second calculation process" is used, the hardware pieces performing the first calculation process and the hardware pieces performing the second calculation process may be the same or different. In other words, it is sufficient that the hardware pieces performing the first calculation process and the hardware pieces performing the second calculation process are included in the one or more pieces of hardware. The hardware pieces may include an electronic circuit, or a device including an electronic circuit.

[0109] In this specification (including claims), when multiple storage devices (memories) store data, each of the multiple storage devices (memories) may store only a portion of the data, or may store the entire data. Also, a configuration in which some of the multiple storage devices (memories) store data may be included.

[0110] Although the embodiments of the present disclosure have been described in detail above, the present disclosure is not limited to the individual embodiments described above. Various additions, modifications, replacements, and partial deletions are possible within the scope of the conceptual idea and intent of the present invention derived from the contents defined in the claims and their equivalents. For example, in all the above-mentioned embodiments, when numerical values ​​or formulas are used in the explanation, they are shown as examples and are not limited to these. In addition, the order of each operation in the embodiment is shown as an example and is not limited to these. [Explanation of symbols]

[0111] 10 System Board 20 processors 30 Arithmetic unit 40 SRAM 50 DRAM (main memory) 60 Auxiliary storage 70 Network Interface 80 Device Interfaces 90 Bus 100 Data processing device 200 Communication Network 210 External device 220 External device

Claims

1. A data processing method for a machine learning model, comprising: compressing data during the calculation of a first calculation process to generate compressed data; storing the generated compressed data in a memory area; executing a second calculation process using the compressed data stored in the memory area The data processing method.

2. The first calculation process is a forward process of a neural network, The second calculation process is a backward process of the neural network The data processing method according to Claim 1.

3. For each of the plurality of layers constituting the neural network, generating data during the calculation of the first calculation process, For each of the plurality of layers, determining whether to compress the data during the calculation The data processing method according to Claim 2.

4. The memory area has a first memory area and a second memory area, The first memory area is allocated to a first memory, The second memory area is allocated to a second memory, The memory bandwidth of the first memory is larger than the memory bandwidth of the second memory The data processing method according to Claim 3.

5. Storing the data during the calculation generated by each of the plurality of layers in the first memory area without compression, Any one of the plurality of layers executes the first calculation process using the uncompressed data during the calculation by the first calculation process of other layers stored in the first memory area The data processing method according to Claim 4.

6. For each of the plurality of layers, holding deletion information indicating whether to delete the uncompressed data during the calculation of other layers used in the first calculation process from the first memory, and based on the deletion information, determining whether to delete the uncompressed data during the calculation of other layers from the first memory The data processing method according to Claim 5.

7. For each of the plurality of layers, holding storage destination information indicating the storage destination of the generated data during the calculation, When compressing the data during the calculation, based on the storage destination information, determining whether to store the compressed data during the calculation in the first memory area or the second memory area The data processing method according to any one of Claims 4 to 6.

8. When compressing the data during calculation, further determine whether to store the compressed data during calculation in both the first memory area and the second memory area based on the storage destination information. When not compressing the data during calculation, determine whether to store the compressed data during calculation in the second memory area based on the storage destination information. The data processing method according to claim 7.

9. Based on the calculation procedure of the neural network and the configuration of the hardware that executes the first calculation process and the second calculation process, determine whether to compress the data during calculation for each of the plurality of layers. The data processing method according to any one of claims 3 to 8.

10. Irreversibly compress the data during calculation to generate the compressed data. The data processing method according to any one of claims 1 to 9.

11. The data during calculation is floating-point data, Generate the compressed data by reducing the number of bits of the mantissa part of the data during calculation. The data processing method according to claim 10.

12. Update the parameters used in the first calculation process by the second calculation process, Repeatedly execute a plurality of third calculation processes each including the first calculation process and the second calculation process while gradually increasing the compression rate of the data. The data processing method according to claim 10 or claim 11.

13. A memory area, Data processing related to a machine learning model, which compresses the data during calculation of the first calculation process to generate compressed data, Stores the generated compressed data in the memory area, An arithmetic processing unit that executes a second calculation process using the compressed data stored in the memory area. A data processing device.

14. Compress the data during calculation of the first calculation process to generate compressed data, Store the generated compressed data in a memory area, Execute a second calculation process using the compressed data stored in the memory area. A data processing program for causing a computer to execute data processing for machine learning.

15. A method for generating a neural network model by machine learning, Compress the data during calculation of the forward process of the neural network to generate compressed data, Store the generated compressed data in a memory area, Execute the backward process of the neural network using the compressed data stored in the memory area. Execute an update process for parameters of layers included in the neural network based on the backward processing A method for generating a neural network model.

Citation Information

Patent Citations

  • Arithmetic unit for neural network

    JP1993265997A

  • Neural network processor using compression and decompression of activation data to reduce memory bandwidth utilization

    JP2020517014A

  • Neural Network Quantization Parameter Determination Method and Related Products

    US20200394522A1