Data processing device for convolution processing, adjusted address list acquisition method, and program

The data processing device for convolutional processing addresses the challenge of hardware-implementing unstructured pruning by optimizing data access through an adjusted address list, enabling efficient weight reduction and high-performance unstructured pruning.

JP2025096744APending Publication Date: 2025-06-30MEGACHIPS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023212634
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-30

AI Technical Summary

Technical Problem

Existing technologies face difficulties in hardware-implementing unstructured pruning for neural network models, which hinders the realization of lightweight models.

Method used

A data processing device for convolutional processing that includes temporary memories, an access control unit, and a storage unit for an adjusted address list. This device performs data processing to enable high-performance unstructured pruning by optimizing data access based on the adjusted address list.

Benefits of technology

The solution allows for efficient weight reduction of neural network models while enabling high-performance unstructured pruning, facilitating the implementation on hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025096744000001_ABST
    Figure 2025096744000001_ABST
Patent Text Reader

Abstract

To realize a data processing device for convolution processing which can perform data processing for realizing weight reduction of a neural network model, while adopting high-performance non-structured pruning.SOLUTION: In a data device for convolution processing, while considering the number of banks which can be accessed in parallel of a tentative memory in which feature amount data are stored, data are read out, from the tentative memory in which the feature amount data are stored, on the basis of an address list after adjustment processing acquired by performing address list adjustment processing, so that the data which can be outputted simultaneously is larger. Thereby the number of data which can be read out simultaneously can be larger.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to data processing technology for convolutional neural networks, and particularly to technology for processing feature data used in convolutional neural networks (data processing apparatus for convolutional processing).

Background Art

[0002] In recent years, technologies using neural network models that can realize a wide variety of applications with high precision have attracted attention. In technologies using neural network models, learning data is used to perform learning processing of the neural network model to obtain a learned model, and prediction processing (inference processing) is performed using the obtained learned model. As a result, with technologies using neural network models, it becomes possible to realize a wide variety of applications with high precision. As a technology using a neural network model that realizes high value in fields such as image recognition, a technology using a convolutional neural network model (CNN: Convolutional Neural Network) has attracted attention.

[0003] Furthermore, lightweight technologies have been developed to enable the use of convolutional neural network models even on mobile terminals and the like that do not have abundant computing resources. As such a technology, for example, a technology called Mobilenet (a lightweight technology for CNN models) has been developed (see, for example, Non-Patent Document 1).

[0004] In a technology called Mobilenet (a technique for lightweighting a CNN model), a method is adopted that divides ordinary convolution processing, called Depthwise separable convolution, into two types: (1) Depthwise convolution (spatial-direction convolution processing) and (2) Pointwise convolution (channel-direction convolution processing), thereby reducing the number of parameters in the CNN model. As a result, a lightweight and high-performance CNN model that can be mounted even on mobile terminals that do not have abundant computing resources can be realized.

[0005] Also, as a technique for lightweighting a neural network model, a method called pruning (pruning process) has been proposed, which reduces the number of parameters and the amount of computation by removing (setting the value to 0) some of the weights of the neural network. As pruning, there are unstructured pruning, which performs pruning irregularly in units of weights, and structured pruning, which performs pruning in units of layers or filters. Unstructured pruning is highly performant, but since it is necessary to perform pruning irregularly, it is difficult to be hardware-implemented (processing by hardware). On the other hand, structured pruning is easy to be hardware-implemented (processing by hardware), but since pruning is performed in units of layers or filters, its performance is inferior to that of unstructured pruning. For example, Non-Patent Document 2 discloses a technique for hardware-implementing structured pruning while realizing lightweighting of a neural network model.

Prior Art Documents

Non-Patent Documents

[0006]

Non-Patent Document 1

[0007] However, in the above conventional technology, it is difficult to hardware-implement (adopt) unstructured pruning while realizing the lightweighting of the neural network model.

[0008] Therefore, in view of the above problems, an object of the present invention is to realize a data processing device for convolutional processing that can perform data processing for realizing the lightweighting of a neural network model while adopting high-performance unstructured pruning. [Means for Solving the Problems]

[0009] To solve the above problems, a first invention is a data processing device for convolutional processing used in a neural network model, comprising: a plurality of temporary memories for storing feature data; an access control unit for controlling data writing and / or data reading of the plurality of temporary memories; and a storage unit for an address list.

[0010] The memory unit for the address list stores an adjusted address list obtained by performing an address list adjustment process so that more data can be output in parallel from a plurality of temporary memories, based on the position of the memory area where the feature amount data of the block to be subjected to the convolution process of the feature amount data remaining after pruning by the unstructured pruning method is stored, and the number of banks of the temporary memory that can be accessed in parallel.

[0011] Each of the plurality of temporary memories includes a plurality of banks that can be accessed in parallel.

[0012] The access control unit performs data write control so that the feature amount data of the block to be subjected to the convolution process is stored in the temporary memory.

[0013] Also, the access control unit inputs an adjusted address list obtained by performing an address list adjustment process so that more data can be output in parallel from a plurality of temporary memories, considering the position of the memory area where the feature amount data of the block to be subjected to the convolution process of the feature amount data remaining after pruning by the unstructured pruning method is stored, and the number of banks of the temporary memory that can be accessed in parallel. Based on the input adjusted address list, the access control unit performs data read control so that the feature amount data is read from the plurality of temporary memories.

[0014] In this data processing device for convolution processing, based on the adjusted address list (post-adjustment address list) obtained by performing an address list adjustment process on the address list of the data remaining after pruning by an unstructured pruning method, data is read from the temporary memory in which the feature amount data is stored, so that the number of data that can be read simultaneously (in parallel) can be increased. That is, in this data device for convolution processing, considering the number of banks of the temporary memory in which the feature amount data is stored and that can be accessed in parallel, based on the adjusted address list obtained by performing an address list adjustment process so that more data can be output simultaneously, data is read from the temporary memory in which the feature amount data is stored, so that the number of data that can be read simultaneously can be increased.

[0015] In this way, in this data device for convolution processing, by providing a temporary memory and performing data access based on the adjusted address list, while adopting high-performance unstructured pruning, the weight reduction of the neural network model can be easily realized.

[0016] Note that the "feature amount data" may be data after quantization processing is performed on the feature amount data (quantized feature amount data).

[0017] A second invention is the first invention, wherein the address list is two-dimensional array data.

[0018] The row data of the address list is an address sequence of the feature amount data remaining after pruning by an unstructured pruning method within a predetermined block for which convolution processing is to be performed.

[0019] The address list constitutes two-dimensional array data by including a plurality of row data.

[0020] The adjusted address list detects addresses where the positions of the banks of the temporary memory in which the feature amount data is stored overlap among the addresses in the same column of the address list, and when an address where the positions of the banks of the temporary memory in which the feature amount data is stored overlap is detected, it performs an address swapping process, which is a process of swapping one of the overlapping addresses with another address in the same row of the address list, and obtains the adjusted address list by executing the process of obtaining the adjusted address list.

[0021] As a result, this convolutional processing data device can efficiently obtain the adjusted address list.

[0022] The third invention is the first or second invention, and each of the plurality of temporary memories has an additional data storage memory area.

[0023] The access control unit also stores, in the additional data storage memory area, the feature amount data determined to be unable to be read out in parallel from the plurality of temporary memories in the adjusted address list.

[0024] As a result, in this convolutional processing data device, for data that could not be read out simultaneously, it is set as additional data (Extra block), and the data is copied to the additional data storage memory area of the temporary memory. For example, data reading processing is performed at a timing when the data can be read out. Therefore, regardless of the situation of the data remaining after pruning by the unstructured pruning method, data (feature amount data or quantized feature amount data) that can surely be the target of the convolutional processing can be read out. That is, in this convolutional processing data device, regardless of the situation of the data remaining after pruning by the unstructured pruning method, data (feature amount data or quantized feature amount data) that can surely be the target of the convolutional processing can be read out.

[0025] The fourth invention is the third invention, wherein the access control unit performs data read control so as to read out the feature amount data stored in the additional data storage memory area during a period other than when controlling to read out the feature amount data from the plurality of temporary memories based on the adjusted address list.

[0026] Accordingly, in this convolutional processing data device, for data that could not be read simultaneously, it is set as additional data (Extra block), and the data is copied to the additional data storage memory area of the temporary memory. For example, data read processing is performed at a timing when the data can be read out. Therefore, regardless of the situation of the data remaining after pruning by the unstructured pruning method, the data (feature amount data or quantized feature amount data) that is surely the target of the convolutional processing can be read out. That is, in this convolutional processing data device, regardless of the situation of the data remaining after pruning by the unstructured pruning method, the data (feature amount data or quantized feature amount data) that is surely the target of the convolutional processing can be read out.

[0027] The fifth invention is an adjusted address list acquisition method for acquiring an adjusted address list used in the convolutional processing data processing device according to the first or second invention, and is an adjusted address list acquisition method executed by a device including a memory and a processor, and includes a first step and a second step.

[0028] In the first step, the processor acquires an address list that is a list of addresses within a block that is the target of the convolutional processing of the feature amount data remaining after pruning by the unstructured pruning method.

[0029] In the second step, the processor takes into account the position of the memory area where the feature data of the block to be subjected to the convolution process in the temporary memory is stored and the number of banks of the temporary memory that can be accessed in parallel, and performs address list adjustment processing so that more data can be output in parallel from the plurality of temporary memories, thereby obtaining an adjusted address list.

[0030] Thereby, by using the adjusted address list obtained by this adjusted address list acquisition method, it becomes possible to perform data read control so that more data can be output in parallel from the plurality of temporary memories.

[0031] The sixth invention is the fifth invention, wherein the address list is two-dimensional array data.

[0032] The row data of the address list is an address list of the feature data remaining after pruning by an unstructured pruning method within a predetermined block to be subjected to the convolution process.

[0033] The address list constitutes two-dimensional array data by including a plurality of pieces of row data.

[0034] Then, in the second step, at the addresses in the same column of the address list, an address where the positions of the banks of the temporary memory in which the feature data is stored overlap is detected, and when an address where the positions of the banks of the temporary memory in which the feature data is stored overlap is detected, an address swapping process, which is a process of swapping one of the overlapping addresses with another address in the same row of the address list, is performed to obtain an adjusted address list.

[0035] Thereby, in this adjusted address list acquisition method, an adjusted address list can be efficiently obtained.

[0036] The seventh invention is the sixth invention, wherein in the second step, if the position of the bank of the temporary memory in which the feature amount data is stored does not reach a state where no overlapping addresses are detected even when the address swapping process is repeatedly executed, at least one of the overlapping addresses is set as additional data which is an address to be stored in the additional data storage memory area, and the address set as the additional data is rewritten with a value indicating that it has been set as the additional data.

[0037] Thereby, with this adjusted address list acquisition method, an adjusted address list in which additional data can be set can be acquired.

[0038] The eighth invention is a program for causing a computer to execute the adjusted address list acquisition method which is the fifth invention.

[0039] Thereby, a program for causing a computer to execute an adjusted address list acquisition method having the same effect as the fifth invention can be realized.

Advantages of the Invention

[0040] According to the present invention, it is possible to realize a data processing device for convolution processing that can perform data processing for realizing weight reduction of a neural network model while adopting high-performance unstructured pruning.

Brief Description of the Drawings

[0041]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Embodiments for Carrying Out the Invention

[0042] [First Embodiment] The first embodiment will be described below with reference to the drawings.

[0043] <1.1: Configuration of Data Processing Device for CNN> FIG. 1 is a schematic configuration diagram of a CNN data processing device 100 according to the first embodiment.

[0044] FIG. 2 is a schematic configuration diagram of the convolution processing data processing unit 3 of the CNN data processing device 100 according to the first embodiment.

[0045] FIG. 3 is a diagram schematically showing the memory area of the temporary memory Tmem_k of the convolution processing data processing unit 3 of the CNN data processing device 100 according to the first embodiment.

[0046] As shown in FIG. 1, the CNN data processing device 100 includes a quantization processing unit 1, an address list adjustment processing unit 2, a convolution processing data processing unit 3, a quantization data memory unit 4, and a convolution processing unit 5. The CNN data processing device 100 inputs feature data Din_f, data D_list_pruning including a list of pruning position information (address information), and weight coefficient data Din_w (weight filter (kernel)), executes convolution processing (convolution processing using feature data and weight coefficient data), and obtains (outputs) the processing result data Dout of the convolution processing.

[0047] The quantization processing unit 1 inputs the feature data Din_f, executes quantization processing on the feature data Din_f, and outputs the data after quantization processing as data D1 to the CNN data processing unit 2.

[0048] The address list adjustment processing unit 2 inputs data D_list_pruning including a pruning address list which is a list of pruning position information (address information), and executes address list adjustment processing on the pruning address list included in the data D_list_pruning. Then, the address list adjustment processing unit 2 outputs the data after the address list adjustment processing as data D_list_adj to the convolution processing data processing unit 3.

[0049] As shown in FIG. 2, the convolution processing data processing unit 3 includes an address list storage unit 31, a memory access control unit 32, a memory unit 33 including M (M: a natural number of 2 or more) temporary memories (Tmem_0 to Tmem_M-1), and a register unit 34.

[0050] The address list storage unit 31 has a memory capable of writing data to a predetermined area by specifying an address and reading data stored in the predetermined area by specifying an address. The address list storage unit 31 inputs the data D_list_adj output from the address list adjustment processing unit 2 and stores the data D_list_adj. Further, the address list storage unit 31 inputs a read control signal Ctl_r_ID from the memory access control unit 32, reads the data stored at the address indicated by the control signal Ctl_r_ID, and outputs the data as data D_ID to the memory access control unit 32.

[0051] The memory access control unit 32 is a control unit for performing access control (data write process control, data read process control) on the M temporary memories of the memory unit 33 and the address list storage unit 31. The memory access control unit 32 outputs a read control signal Ctl_r_ID to the address list storage unit 31, and reads the data stored at the address indicated by the read control signal Ctl_r_ID as data D_ID (input from the address list storage unit 31). Also, the memory access control unit 32 is a functional unit for independently (in parallel) performing data write process control and data read process control on the M temporary memories Tmem_0 to Tmem_M-1 of the memory unit 33. The memory access control unit 32 outputs a control signal Ctl_w for performing data write process control and / or a control signal Ctl_r for performing data read process control on the memory unit 33. Specifically, the memory access control unit 32 outputs a control signal Ctl_w (k) for performing data write process control and / or a control signal Ctl_r (k) for performing data read process control on the temporary memory Tmem_k (k: natural number, 0 ≦ k ≦ M-1) of the memory unit 33. Note that the control signal Ctl_w for performing data write process control on the temporary memory Tmem_k (k: natural number, 0 ≦ k ≦ M-1) of the memory unit 33 is denoted as control signal Ctl_w (k) , and the control signal Ctl_r for performing data read process control on the temporary memory Tmem_k of the memory unit 33 is denoted as control signal Ctl_r (k) .

[0052] Also, the memory access control unit 32 is a control unit for performing access control (data write process control, data read process control) on the register unit 34. The memory access control unit 32 outputs a control signal Ctl_reg for performing access control on the register unit 34 to the register unit 34.

[0053] Also, the memory access control unit 32 reads out the adjusted address list stored in the address list storage unit 31, and outputs the data including the address list as data D2_list_adj to the convolution processing unit 5.

[0054] As shown in FIG. 2, the memory unit 33 includes M (M: a natural number of 2 or more) provisional memories Tmem_0 to Tmem_M-1.

[0055] The provisional memory Tmem_k (k: natural number, 0 ≦ k ≦ M-1) is a memory that can write predetermined data to a predetermined address of the provisional memory Tmem_k and read out the data stored at the address from the predetermined address of the provisional memory Tmem_k. The provisional memory Tmem_k receives the control signal Ctl_w for data writing processing from the memory access control unit 32 (k) and writes the data D1 output from the quantization processing unit 1 to the address of the provisional memory Tmem_k specified by the control signal Ctl_w (k) Also, the provisional memory Tmem_k receives the control signal Ctl_r for data reading processing from the memory access control unit 32 (k) and reads out the data stored at the address from the address of the provisional memory Tmem_k specified by the control signal Ctl_w (k) and outputs the read data to the register unit 34. Note that the provisional memory Tmem_k has a plurality of banks (for example, 16 banks (16 in FIG. 3)), and access buses (16 access buses in FIG. 3) are connected to each of the plurality of banks. Through the access buses, a plurality of data can be read out simultaneously (in parallel) from the plurality of banks, or a plurality of data can be written to the plurality of banks simultaneously (in parallel). Also, as shown in FIG. 3, an additional data storage memory area Mem_extra_blk (the area shown by the hatched rectangle in FIG. 3) is secured in the provisional memory Tmem_k.

[0056] The register unit 34 has a memory (register) that can write data to a predetermined area by specifying an address and can read the data stored in the predetermined area by specifying an address. The register unit 34 inputs the data output from the memory unit 33 and the control signal Ctl_reg output from the memory access control unit 32. The register unit 34 writes the data output from the memory unit 33 to a predetermined address of the register unit 34 according to the control signal Ctl_reg. Also, the register unit 34 outputs the data at a predetermined address of the register unit 34 as data D2 to the quantization data memory unit 4 according to the control signal Ctl_reg.

[0057] The quantization data memory unit 4 has a memory that can store data. The memory can write data to a predetermined area by specifying an address and can read the data stored in the predetermined area by specifying an address. The quantization data memory unit 4 inputs the data D2 output from the convolution processing data processing unit 3 and stores the data D2. Also, the quantization data memory unit 4 outputs the stored data as data D3 to the convolution processing unit 5 (the quantization data memory unit 4 inputs a data read command from a control unit (not shown) or the convolution processing unit 5 and outputs the data at a predetermined address as data D3 to the convolution processing unit 5 according to the data read command).

[0058] The convolution processing unit 5 inputs the weight coefficient data Din_w (weight filter (kernel)), the data D3 output from the quantization data memory unit 4, and the data D2_list_adj output from the convolution processing data processing unit 3. The convolution processing unit 5 executes a convolution process using the data D3 and the weight coefficient data Din_w while referring to the adjusted address list included in the data D2_list_adj. Then, the convolution processing unit 5 outputs the data after the convolution process as data Dout.

[0059] <1.2: Operation of the Data Processing Device for CNN> The operation of the CNN data processing device 100 configured as described above will be described below.

[0060] FIG. 4 is a diagram for explaining CNN processing (convolution processing for the CNN model) by (1) Depthwise convolution (spatial direction convolution processing) and (2) Pointwise convolution (channel direction convolution processing).

[0061] FIGS. 5 and 6 are diagrams schematically showing the processing target block block h for each position in the height direction of the convolution processing.

[0062] FIG. 7 is a diagram schematically showing the address list.

[0063] FIG. 8 is a flowchart of the address list adjustment process executed by the CNN data processing device.

[0064] FIGS. 9 to 17 are diagrams for explaining the address list adjustment process.

[0065] FIGS. 18 to 25 are diagrams showing the data storage state of the temporary memory Tmem_k and the method for obtaining the address list after the adjustment process.

[0066] FIG. 26 is a diagram for explaining the first data read process.

[0067] FIG. 27 is a diagram for explaining the second data read process.

[0068] FIG. 28 is a diagram for explaining the third data read process.

[0069] FIG. 29 is a diagram for explaining the fourth data read process.

[0070] FIG. 30 is a diagram for explaining the fifth data read process.

[0071] FIG. 31 is a diagram for explaining the sixth data read process.

[0072] FIG. 32 is a diagram for explaining the seventh data read process.

[0073] FIG. 33 is a diagram for explaining the eighth data read process.

[0074] FIG. 34 is a diagram for explaining the ninth data read process.

[0075] FIG. 35 is a diagram for explaining the convolution process.

[0076] FIG. 36 is a diagram for explaining the convolution process.

[0077] Hereinafter, for the sake of convenience of explanation, as shown in FIG. 4, when the method of separating the normal convolution process into two, (1) Depthwise convolution (spatial direction convolution process) and (2) Pointwise convolution (channel direction convolution process) (Depthwise separable convolution) is adopted, the operation of the CNN data processing device 100 will be described for the case of performing the channel direction convolution process (Pointwise convolution) (an example).

[0078] Also, hereinafter, for the sake of convenience of explanation, the operation of the CNN data processing device 100 will be described for the case where the following settings are made (an example). Note that the settings for the CNN data processing device 100 are not limited to the following settings, and other settings may also be used. (1) The block of the feature amount data (the feature amount data input to the CNN data processing device 100) to be subjected to the convolution process has a channel number of "8" (c in = 8), a height of "8" (h = 8), and a width of "8" (w = 8) block (processing target block) (see the upper figure in FIG. 4). (2) The size of the kernel (filter) for performing pointwise convolution in the channel direction is 1×1×8 (the size in the height direction is "1", the size in the width direction is "1", and the size in the channel direction is "8") (a Conv1d kernel with a channel size of "8") (see the upper figure in Fig. 4). (3) The processing target block block for each position in the height direction of the convolution process h (h: an integer indicating the position in the height direction, 0 ≦ h ≦ 7) is an 8×8 block (a block with a channel size of "8" and a width size of "8"). The data pruned (excluded from the processing target of the convolution process) is the data indicated by the gray rectangles in Figs. 5 and 6, and the data not pruned (targeted for the convolution process) is the data indicated by the white rectangles in Figs. 5 and 6. Specifically, in the processing target block block h the data not pruned (targeted for the convolution process) shall be as follows (an example). Note that each data in the block block h is denoted as D h (X), h is the height position, and X is an 8-bit address expressed in hexadecimal (an address (8-bit data indicating the address) for specifying the position within the block block h ). Block block0: D0(8), D0(16), D0(19), D0(1a), D0(28), D0(2a), D0(33), D0(3c) Block block1: D1(7), D1(10), D1(12), D1(15), D1(1a), D1(1c), D1(32), D1(3b) Block block2: D2(13), D2(19), D2(1f), D2(23), D2(26), D2(2a), D2(32), D2(35) block3: D3(b), D3(10), D3(14), D3(1a), D3(1f), D3(32), D3(36), D3(3c) block4: D4(0), D4(2), D4(7), D4(19), D4(1a), D4(2d), D4(3a), D4(3b) block5: D5(8), D5(10), D5(16), D5(17), D5(30), D5(31), D5(33), D5(39) block6: D6(1), D6(d), D6(10), D6(16), D6(1f), D6(24), D6(29), D6(37) block7: D7(c), D7(f), D7(18), D7(1b), D7(1c), D7(23), D7(24), D7(26) Note that the positions of the data that were not pruned (and are the targets of the convolution process) are assumed to be known in advance (the positions of the data that were not pruned (and are the targets of the convolution process) are stored during or after the pruning process and are assumed to be accessible (retrievable) after the pruning process). (4) The memory unit 33 has eight (M = 8) temporary memories Tmem_0 to Tmem_7. The temporary memory Tmem_k (k: integer, 0 ≤ k ≤ 7) has 16 banks, and access buses are connected to each of the plurality of banks. Through the access buses, a plurality of data (a plurality of 16 or fewer data) can be read from the 16 banks simultaneously (in parallel), or a plurality of data (a plurality of 16 or fewer data) can be written to the 16 banks simultaneously (in parallel) (see FIG. 3). Also, it is assumed that an additional data storage memory area (extra block memory area) is set (secured) in the temporary memory Tmem_k (k: integer, 0 ≤ k ≤ 7) (see FIG. 3).

[0079] Hereinafter, the operation of the CNN data processing apparatus 100 will be described.

[0080] (1.2.1: Address list adjustment process) First, the address list adjustment process executed by the CNN data processing device 100 will be described.

[0081] Fig. 7 is a diagram showing an address list. As shown in Fig. 7, the address list is treated as two-dimensional array data. The address list is the block of data (the data remaining after pruning) to be subjected to the convolution process of the block h described in the above setting (3), and the address list within the block h is used as the column data of the address list, and this column data is data (two-dimensional array data) arranged vertically for each position h in the height direction.

[0082] block h When writing the data of the block into the temporary memory Tmem_k and reading out the data to be subjected to the convolution process (the data remaining after pruning), the data in the same column of the above address list is the data to be read out simultaneously via 16 access buses (shared in the temporary memory Tmem_k) from the temporary memory Tmem_k (a temporary memory having 16 banks). Therefore, by adjusting the address list so that data that does not use the same access bus is read out, data can be efficiently read out from the temporary memory Tmem_k. Specifically, when the address is X, the number of the access bus when reading data from the temporary memory Tmem_k is X mod 16 (mod 16 is the remainder with 16 (decimal) as the modulus) and can be calculated. Thus, the adjustment process of the address list is performed by swapping the data within the same row so that data having the same number of the access bus when reading data from the temporary memory Tmem_k does not exist in the same column.

[0083] Hereinafter, a case (an example) of executing the address list adjustment process for the address list in Fig. 7 will be described with reference to the flowchart in Fig. 8.

[0084] (Step S1): In step S1, the process of obtaining the address list is executed. Specifically, the address list adjustment processing unit 2 obtains the address list in FIG. 7. Note that the address list is stored during or after the pruning process and can be referred to (obtained) after the pruning process. Therefore, the address list adjustment processing unit 2 obtains the address list from a storage unit (not shown) in which the address list is stored.

[0085] (Step S2): In step S2, the initialization process is executed. Specifically, the address list adjustment processing unit 2 sets the number of access buses of the temporary memory in the memory unit 33 of the convolutional data processing unit 3 (in this embodiment, "16"), the number of temporary memories (in this embodiment, "8" (M = 8)), and the maximum number of columns j_max of the access list (in the case of the address list in FIG. 7, j_max = 8). Further, the variable j (a variable for loop 1 processing) is set to "0". Note that j_max is the maximum value of the number of target data for the convolutional process of block k and in the case of the address list in FIG. 7, the number of target data for the convolutional process of blocks block0 to block7 is all "8", and j_max = 8. Note that the number of target data for the convolutional process of blocks block0 to block7 may be different from each other.

[0086] ≪j = 0≫ (Step S3): In step S3, the loop process (loop 1 process) is started. The loop 1 process is repeatedly executed as long as the variable j satisfies the condition j < j_max.

[0087] (Step S4): In step S4, a determination process is executed to check whether there is data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 0 in the address list. Specifically, as shown in FIG. 9, the address list adjustment processing unit 2 determines whether there is data read from the temporary memory Tmem_k using the same access bus for the first column (corresponding to j = 0) of the address list. For the address X in the first column (corresponding to j = 0) of the address list, the address list adjustment processing unit 2 X mod 16 (mod 16 means the remainder with 16 (decimal) as the modulus) is calculated, and it is determined whether there are addresses with the same remainder (corresponding to addresses with the same lower 4-bit value). As can be seen from FIG. 9, the address in the first row (the address within block block0) is 0x08 (denoted as "8" in FIG. 9), and the address in the sixth row (the address within block block5) is 0x08 (denoted as "8" in FIG. 9). Since the remainders are the same (the lower 4-bit values are the same), the address list adjustment processing unit 2 determines that there is data read from the temporary memory Tmem_k using the same access bus (determines that there are overlapping addresses in terms of access bus usage) ("Yes" in step S4). Then, the process proceeds to step S5.

[0088] (Step S5): In step S5, a data swapping process (address swapping process) is executed. Specifically, as shown in FIG. 9, the address list adjustment processing unit 2 moves the address 0x08 (denoted as "8" in FIG. 9) in the sixth row, which is one of the overlapping addresses (the address within block block5), to the end (the last column) of the sixth row of the address list, and further shifts the data column (address column) from the second column to the eighth column in the sixth row (the rectangular part indicated by the dotted line in FIG. 9) one position to the left. By this process, the address list becomes the state shown in state 2 of FIG. 9.

[0089] Furthermore, the address list adjustment processing unit 2 executes a determination process as to whether there is data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 0 in the address list in state 2.

[0090] As can be seen from FIG. 10, the address of the fifth row (the address within block block4) is 0x00 (denoted as "0" in FIG. 10), and the address of the sixth row (the address within block block5) is 0x10 (denoted as "10" in FIG. 10). Since the remainder is the same (the lower 4-bit values are the same), the address list adjustment processing unit 2 determines that there is data read from the temporary memory Tmem_k using the same access bus (determines that there are overlapping addresses for access bus utilization). Then, as shown in FIG. 10, the address list adjustment processing unit 2 moves the address of the sixth row (the address within block block5), which is one of the overlapping addresses and is 0x10 (denoted as "10" in FIG. 10), to the end (the last column) of the sixth row of the address list, and further shifts the data column (address column) from the second column to the eighth column of the sixth row (the rectangular portion indicated by the dotted line in FIG. 10) one position to the left. By this process, the address list becomes the state of state 3 in FIG. 10. As can be seen from FIG. 10, since there is no data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 0 in the address list in state 3, the address list in state 3 is adjusted to a state where there is no data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 0.

[0091] (Step S6): In step S6, a determination process is executed to check whether, through the data swapping process in step S5, the data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 0 in the address list can be adjusted to a state where there is no such data. As can be seen from FIG. 10, the address list in state 3 is adjusted to a state where there is no data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 0. Therefore, the address list adjustment processing unit 2 determines that it has been possible to adjust to a state where there is no data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 0 in the address list, and advances the process to step S8.

[0092] (Step S8): In step S8, the address list adjustment processing unit 2 performs a process of incrementing the variable j by +1 (j = 1), and advances the process to step S9.

[0093] (Step S9): In step S9, an end determination of loop 1 processing is performed. If j < j_max is satisfied, the process returns to step S3. If j < j_max is not satisfied, the process advances to step S10. Here, j = 1 and j < j_max (= 8) is satisfied, so the process returns to step S3.

[0094] ≪j = 1≫ (Step S3): In step S3, the loop process (loop 1 process) continues (j = 1).

[0095] (Step S4): In step S4, a determination process is executed to check whether there is data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 1 in the address list. Specifically, as shown in FIG. 11, the address list adjustment processing unit 2 determines whether there is data read from the temporary memory Tmem_k using the same access bus for the second column (corresponding to j = 1) of the address list. The address list adjustment processing unit 2 calculates X mod 16 (mod 16 is the remainder with 16 (decimal) as the divisor) for the address X in the first column (corresponding to j = 1) of the address list, and determines whether there is an address with the same remainder (corresponding to an address with the same lower 4-bit value). As can be seen from FIG. 11, the address in the second row (the address within block block1) is 0x10 (represented as "10" in FIG. 11), and the address in the fourth row (the address within block block3) is 0x10 (represented as "10" in FIG. 11). Since the remainders are the same (the lower 4-bit values are the same), the address list adjustment processing unit 2 determines that there is data read from the temporary memory Tmem_k using the same access bus (determines that there is an overlapping address for access bus usage) ("Yes" in step S4). Then, the process proceeds to step S5.

[0096] (Step S5): In step S5, data swapping processing (address swapping processing) is executed. Specifically, as shown in FIG. 11, the address list adjustment processing unit 2 moves the address of the fourth row (the address within block 3), which is one of the duplicate addresses, to the end (the last column) of the fourth row of the address list, and further shifts the data column (address column) from the third column to the eighth column of the fourth row (the rectangular portion indicated by the dotted line in FIG. 11) one position to the left. By this processing, the address list becomes the state of state 4 in FIG. 11. As can be seen from FIG. 11, in the column corresponding to j = 1 in the address list of state 4, there is no data read from the temporary memory Tmem_k using the same access bus. Therefore, the address list of state 4 is adjusted to a state where there is no data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 1.

[0097] (Step S6): In step S6, a determination process is executed to determine whether the address list has been adjusted to a state where there is no data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 1 by the data swapping processing in step S5. As can be seen from FIG. 11, since the address list of state 4 is adjusted to a state where there is no data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 1, the address list adjustment processing unit 2 determines that the adjustment to a state where there is no data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 1 has been successful, and proceeds with the processing to step S8.

[0098] (Step S8): In step S8, the address list adjustment processing unit 2 performs a process of incrementing the variable j by +1 (j = 2), and proceeds with the processing to step S9.

[0099] (Step S9): In step S9, it is determined whether the loop 1 process has ended. If j < j_max is satisfied, the process returns to step S3. If j < j_max is not satisfied, the process proceeds to step S10. Here, j = 2, and since j < j_max (= 8) is satisfied, the process returns to step S3.

[0100] <<j = 2>> For the case of j = 2 as well, the processes of steps S3 to S9 are executed in the same manner as for the case of j = 1. By executing the process for the case of j = 2, as shown in FIG. 12, the address list in state 5 is obtained.

[0101] <<j = 3>> For the case of j = 3 as well, the processes of steps S3 to S9 are executed in the same manner as for the case of j = 1. By executing the process for the case of j = 3, as shown in FIG. 13, the address list in state 5 is obtained.

[0102] <<j = 4>> For the case of j = 4 as well, the processes of steps S3 to S9 are executed in the same manner as for the case of j = 1. By executing the process for the case of j = 4, as shown in FIG. 14, the address list in state 5 is obtained.

[0103] <<j = 5>> Even when j = 5, the processes of steps S3 to S9 are executed in the same manner as when j = 1. By executing the process for j = 5, as shown in FIG. 15, the address list in state 5 is obtained. Note that when j = 5, as can be seen from FIG. 15, the address of the first row (the address within block block0) 0x2A (denoted as "2A" in FIG. 15), the address of the third row (the address within block block2) 0x2A (denoted as "2A" in FIG. 15), and the address of the fifth row (the address within block block4) 0x3A (denoted as "3A" in FIG. 15) have the same remainder (the following 4-bit values are the same) and overlap (the data read from the temporary memory Tmem_k using the same access bus overlaps). Therefore, for the address column of the third row (the address column within block block2) and the address column of the fifth row (the address column within block block4), a data swapping process is executed (see FIG. 15).

[0104] ≪j = 6≫ When j = 6, as can be seen from FIG. 16, since there is no data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 6 in the address list of state 8 ( "No" in step S4), the processes of step S4 and step S8 are executed.

[0105] ≪j = 7≫ When j = 7, the following process is executed.

[0106] (Step S3): In step S3, the loop process (loop 1 process) continues (j = 2).

[0107] (Steps S4 to S6): In step S4, a determination process is executed to check whether there is data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 7 in the address list. Specifically, as shown in FIG. 17, the address list adjustment processing unit 2 determines whether there is data read from the temporary memory Tmem_k using the same access bus for the eighth column (corresponding to j = 7) of the address list.

[0108] As can be seen from FIG. 17, the address in the third row (the address within block block2) is 0x2A (denoted as "2A" in FIG. 17), and the address in the fifth row (the address within block block4) is 0x3A (denoted as "3A" in FIG. 17), and their remainders are the same (the following 4-bit values are the same), and they overlap (the data read from the temporary memory Tmem_k using the same access bus overlaps). Also, the address in the fourth row (the address within block block3) is 0x10 (denoted as "10" in FIG. 17), and the address in the sixth row (the address within block block5) is 0x10 (denoted as "10" in FIG. 17), and their remainders are the same (the following 4-bit values are the same), and they overlap (the data read from the temporary memory Tmem_k using the same access bus overlaps).

[0109] In the address list adjustment process, when j = 7, since the column to be processed is the last column, data swapping processing cannot be performed. Therefore, the address list adjustment processing unit 2 determines that it cannot be adjusted to a state where there is no data read from the temporary memory Tmem_k using the same access bus in the column corresponding to j = 7 in the address list ( "No" in step S6), and proceeds to step S7.

[0110] (Step S7): In step S7, additional data setting processing (Extra block setting processing) is executed. Specifically, the address list adjustment processing unit 2 performs processing to rewrite the address to an address (for example, 0xFF) indicating that one of the duplicate addresses is set as additional data (Extra block). In the case of FIG. 17, the address list adjustment processing unit 2 sets the address 0x3A (denoted as "3A" in FIG. 17) in the fifth row (the address within block 4) as additional data (Extra block), and rewrites the original address 0x3A to 0xFF (a value (address) indicating that it is set as additional data (Extra block)). Also, the address list adjustment processing unit 2 sets the address 0x10 (denoted as "10" in FIG. 17) in the sixth row (the address within block 5) as additional data (Extra block), and rewrites the original address 0x10 to 0xFF (a value (address) indicating that it is set as additional data (Extra block)). By this processing, the address list in state 9 is obtained.

[0111] (Step S8): In step S8, the address list adjustment processing unit 2 performs processing to increment the variable j by +1 (j = 8), and advances the processing to step S9.

[0112] (Step S9): In step S9, an end determination of loop 1 processing is performed. If j < j_max is satisfied, the processing returns to step S3. If j < j_max is not satisfied, the processing advances to step S10. Here, j = 8, and since j < j_max (= 8) is not satisfied, the processing advances to step S10.

[0113] (Step S10): In step S10, the address list adjustment processing unit 2 outputs data including an adjusted address list (the address list after adjustment processing (the address list in state 9 of FIG. 17)) to the convolutional processing data processing unit 3 as data D_list_adj. It is assumed that the adjusted address list (the address list after adjustment processing) includes information on the original addresses of the data (addresses) set in the additional data (Extra block).

[0114] As described above, the address list adjustment process is executed.

[0115] (1.2.2: Convolutional Processing Data) Next, the convolutional processing data processing executed by the CNN data processing device 100 will be described.

[0116] ≪Data Writing Process≫ The address list storage unit 31 inputs the data D_list_adj (data including the address list after address list adjustment processing) output from the address list adjustment processing unit 2 and stores the data D_list_adj. Here, it is assumed that the address list in state 9 of FIG. 17 is stored and held by the address list storage unit 31.

[0117] Also, the feature data Din_f is input to the quantization processing unit 1.

[0118] The quantization processing unit 1 executes quantization processing on the feature data Din_f and outputs the data after quantization processing to the convolutional processing data processing unit 3 as data D1.

[0119] The memory access control unit 32 generates a control signal Ctl_w for controlling the data writing process for the memory unit 33, outputs it to the memory unit 33, and writes the data D1 to the temporary memory Tmem_k of the memory unit 33. Specifically, the memory access control unit 32 executes the following process. (1) The memory access control unit 32 generates a control signal Ctl_w for writing the data (8×8 data) of the processing target block block where the height direction position h of the convolution process is "0" h (h = 0) to the temporary memory Tmem_0, and outputs the control signal Ctl_w (0) to the temporary memory Tmem_0. The temporary memory Tmem_0 writes the data (8×8 data) of the processing target block block where the height direction position h of the convolution process is "0" (0) (h = 0) to the temporary memory Tmem_0 according to the control signal Ctl_w. Note that the data of the block block (0) (h = 0) is assumed to be included in the data D1. The data (8×8 data) of the processing target block block h (h = 0) is set as data D0(0)~D0(3f) (see Fig. 5). Since the temporary memory Tmem_0 has 16 banks, as shown in Fig. 18, 64 data D0(0)~D0(3f) are sequentially stored in 16 banks (referred to as the first bank to the sixteenth bank). That is, as shown in Fig. 18, in the temporary memory Tmem_k, (A) 16 data D0(0)~D0(f) are respectively stored in the first bank to the sixteenth bank. (B) 16 data D0(10)~D0(1f) are respectively stored in the first bank to the sixteenth bank. (C) 16 data D0(20)~D0(2f) are respectively stored in the first bank to the sixteenth bank. (D) 16 data D0(30)~D0(3f) are respectively stored in the first bank to the sixteenth bank. (2) When h is 1 to 7, the same processing as in (1) above is executed. That is, the memory access control unit 32 generates a control signal Ctl_w for writing the data (8×8 data) of the processing target block block where the height direction position of the convolution process is h h to the temporary memory Tmem_h, and outputs the control signal Ctl_w h to the temporary memory Tmem_h. h The control signal Ctl_w (h) is generated, and the control signal Ctl_w (h)Output it to the temporary memory Tmem_h. The temporary memory Tmem_h is controlled by the control signal Ctl_w (h) According to, for the processing target block block at the height position h in the convolution process h Write the data (8×8 data) of to the temporary memory Tmem_h. Note that the data of block h Is assumed to be included in the data D1. The data of the processing target block block h The data (8×8 data) of is the data D h (0)~D h (3f) (see FIGS. 5 and 6). Note that since the temporary memory Tmem_h has 16 banks, as shown in FIGS. 19 (when h = 1) to 25 (when h = 7), 64 data D h (0)~D h (3f) are sequentially stored in 16 banks (referred to as the first bank to the sixteenth bank). That is, as shown in FIGS. 19 (when h = 1) to 25 (when h = 7), in the temporary memory Tmem_k (k = h (k coincides with h)), (A) 16 data D k (0)~D k (f) are respectively stored in the first bank to the sixteenth bank. (B) 16 data D k (10)~D k (1f) are respectively stored in the first bank to the sixteenth bank. (C) 16 data D k (20)~D k (2f) are respectively stored in the first bank to the sixteenth bank. (D) 16 data D k (30)~D k (3f) are respectively stored in the first bank to the sixteenth bank.

[0120] In addition, in FIGS. 18 (when h = k = 0) to 25 (when h = k = 7), the white rectangular parts are the data remaining after pruning (the data to be subjected to the convolution process), and the gray rectangular parts are the data cut off by pruning (the data not to be subjected to the convolution process).

[0121] The memory access control unit 32 also outputs the read control signal Ctl_r_ID to the address list storage unit 31, and reads (inputs from the address list storage unit 31) the data stored at the address specified by the read control signal Ctl_r_ID to obtain the address list after the adjustment process. The memory access control unit 32 then detects data set as additional data (Extra block) from the address list after the adjustment process, and performs processing to copy the detected additional data (Extra block) to the memory area Mem_extra_blk for storing additional data (processing (1) and (2) below). (1) The data in the 8th column of the 5th row (h=k=4) of the adjusted address list is “0xFF” and is set as extra data (Extra block). Therefore, the memory access control unit 32 outputs a control signal Ctl_w to copy the data D4 (3a) of the original address 0x3A (the original address is assumed to be obtainable from the address list storage unit 31) to the extra data storage memory area Mem_extra_blk. (4) and generates the control signal Ctl_w (4) The temporary memory Tmem_4 outputs the control signal Ctl_w (4) 22, the data D4 (3a) at the original address 0x3A is copied to the additional data storage memory area Mem_extra_blk (the additional data storage memory area Mem_extra_blk in the same bank as the bank in which the data D4 (3a) at the original address 0x3A is stored) (see FIG. 22). (2) Since the data in the eighth column of the sixth row (h=k=5) of the adjusted address list is “0xFF” and is set as extra data (Extra block), the memory access control unit 32 outputs a control signal Ctl_w to copy the data D5(10) of the original address 0x10 (the original address is assumed to be obtainable from the address list storage unit 31) to the extra data storage memory area Mem_extra_blk. (5) and generates the control signal Ctl_w (5)Output it to the temporary memory Tmem_5. The temporary memory Tmem_5 follows the control signal Ctl_w (5) to copy the data D5(10) at the original address 0x10 to the additional data storage memory area Mem_extra_blk (in the same bank as the bank where the data D5(10) at the original address 0x10 is stored) according to the control signal Ctl_w (see FIG. 23).

[0122] The memory access control unit 32 has information (data write information (e.g., a flag)) for grasping whether a predetermined amount of data (e.g., a predetermined amount of data in block units (e.g., data of a block corresponding to the adjusted address list)) has been written to the memory unit 33, and performs data write control based on the information (e.g., the flag). Thereby, for example, a predetermined amount of data in block units (e.g., data of a block corresponding to the adjusted address list) can be written to the temporary memory Tmem_k in a timely manner. For example, when the data of the block corresponding to the adjusted address list has completed the convolution process, new data can be written (overwritten) to the area of the temporary memory Tmem_k where the data of the block (data that no longer needs to be stored and retained) is stored.

[0123] <<Data read processing>> After determining that the data of the block corresponding to the adjusted address list has been written to the memory unit 33 by referring to the data write information, the memory access control unit 32 executes the data read processing from the memory unit 33.

[0124] <<First data read processing>> As shown in FIG. 26, the memory access control unit 32 refers to the adjusted address list and reads the data specified by the address in the first column of the adjusted address list (the address in the dotted rectangle in FIG. 26) from the temporary memories Tmem_0 to Tmem_7 via 16 access buses. Specifically, the following read processing is executed. (1) Since the data (address) in the first row of the first column of the address list after the adjustment process is "8", the memory access control unit 32 generates a control signal Ctl_r to read the data D0(8) of the temporary memory Tmem_0. (0) and outputs the control signal to the temporary memory Tmem_0. The temporary memory Tmem_0 reads the data D0(8) according to the control signal Ctl_r. (0) (Since the data D0(8) is stored in the ninth bank, the data is read through the ninth access bus). (2) Since the data (address) in the second row of the first column of the address list after the adjustment process is "7", the memory access control unit 32 generates a control signal Ctl_r to read the data D1(7) of the temporary memory Tmem_1. (1) and outputs the control signal to the temporary memory Tmem_1. The temporary memory Tmem_1 reads the data D1(7) according to the control signal Ctl_r. (1) (Since the data D1(7) is stored in the eighth bank, the data is read through the eighth access bus). (3) Since the data (address) in the third row of the first column of the address list after the adjustment process is "13", the memory access control unit 32 generates a control signal Ctl_r to read the data D2(13) of the temporary memory Tmem_2. (2) and outputs the control signal to the temporary memory Tmem_2. The temporary memory Tmem_2 reads the data D2(13) according to the control signal Ctl_r. (2) (Since the data D2(13) is stored in the fourth bank, the data is read through the fourth access bus). (4) Since the data (address) in the fourth row of the first column of the address list after the adjustment process is "B", the memory access control unit 32 generates a control signal Ctl_r to read the data D3(B) of the temporary memory Tmem_3. (3) and outputs the control signal to the temporary memory Tmem_3. The temporary memory Tmem_3 reads the data D3(B) according to the control signal Ctl_r. (3)In accordance with this, data D3(B) is read out (since data D3(B) is stored in the 12th bank, the data is read out via the 12th access bus). (5) Since the data (address) in the 5th row of the address in the first column of the address list after the adjustment process is "0", the memory access control unit 32 generates a control signal Ctl_r to read data D4(0) from the temporary memory Tmem_4 (4) and outputs the control signal to the temporary memory Tmem_4. The temporary memory Tmem_4 reads out data D4(0) in accordance with the control signal Ctl_r (4) (since data D4(0) is stored in the first bank, the data is read out via the first access bus). (6) Since the data (address) in the 6th row of the address in the first column of the address list after the adjustment process is "16", the memory access control unit 32 generates a control signal Ctl_r to read data D5(16) from the temporary memory Tmem_5 (5) and outputs the control signal to the temporary memory Tmem_5. The temporary memory Tmem_5 reads out data D5(16) in accordance with the control signal Ctl_r (5) (since data D5(16) is stored in the 7th bank, the data is read out via the 7th access bus). (7) Since the data (address) in the 7th row of the address in the first column of the address list after the adjustment process is "1", the memory access control unit 32 generates a control signal Ctl_r to read data D6(1) from the temporary memory Tmem_6 (6) and outputs the control signal to the temporary memory Tmem_6. The temporary memory Tmem_6 reads out data D6(1) in accordance with the control signal Ctl_r (6) (since data D6(1) is stored in the second bank, the data is read out via the second access bus). (8) Since the data (address) in the 8th row of the address in the first column of the address list after the adjustment process is "C", the memory access control unit 32 generates a control signal Ctl_r to read data D7(C) from the temporary memory Tmem_7 (7)Generate it and output the control signal to the temporary memory Tmem_7. The temporary memory Tmem_7 reads out the data D7(C) according to the control signal Ctl_r (7) (Since the data D7(C) is stored in the 13th bank, the data is read out via the 13th access bus).

[0125] In the data readout processes (1) to (8) above, as can be seen from FIG. 26, since the eight data items to be read out are read out via separate access buses without duplication, the eight data items (data D0(8), D1(7), D2(13), D3(B), D4(0), D5(16), D6(1), D7(C)) can be read out simultaneously (in parallel).

[0126] The data read out from the memory unit 33 as described above is output to the register unit 34.

[0127] ≪Second data readout process≫ For the second data readout process, the process is executed in the same manner as the first data readout process.

[0128] In the second data readout process, as shown in FIG. 27 (1) Data D1(10) is read out from the first bank of the temporary memory Tmem_1 via the first access bus. (2) Data D4(2) is read out from the third bank of the temporary memory Tmem_4 via the third access bus. (3) Data D3(14) is read out from the fifth bank of the temporary memory Tmem_3 via the fifth access bus. (4) Data D0(16) is read out from the seventh bank of the temporary memory Tmem_0 via the seventh access bus. (5) Data D3(17) is read out from the eighth bank of the temporary memory Tmem_3 via the eighth access bus. Data D2(19) is read from the 10th bank of the temporary memory Tmem_2 via the 10th access bus. (7) Data D6(D) is read from the 14th bank of the temporary memory Tmem_6 via the 14th access bus.

[0129] In the data readout processes (1) to (8) above, as can be seen from FIG. 27, the eight data items to be read are read via separate access buses without duplication. Therefore, the eight data items (data D0(16), D1(10), D2(19), D3(14), D4(2), D5(17), D6(D), D7(F)) can be read simultaneously (in parallel).

[0130] The data read from the memory unit 33 as described above is output to the register unit 34.

[0131] ≪Data Readout Processes for the 3rd to 8th Times≫ For the data readout processes for the 3rd to 8th times, the processes are executed in the same manner as the data readout process for the 1st time.

[0132] In the data readout process for the 3rd time, as can be seen from FIG. 28, the eight data items to be read are read via separate access buses without duplication. Therefore, the eight data items (data D0(19), D1(12), D2(1F), D3(1A), D4(7), D5(30), D6(16), D7(18)) can be read simultaneously (in parallel).

[0133] The data read from the memory unit 33 as described above is output to the register unit 34.

[0134] In the fourth data read process, as can be seen from FIG. 29, since eight data items to be read are read via separate access buses without duplication, the eight data items (data D0(1A), D1(15), D2(23), D3(1F), D4(19), D5(31), D6(24), D7(1B)) can be read simultaneously (in parallel).

[0135] The data read from the memory unit 33 as described above is output to the register unit 34.

[0136] In the fifth data read process, as can be seen from FIG. 30, since eight data items to be read are read via separate access buses without duplication, the eight data items (data D0(28), D1(1A), D2(26), D3(32), D4(2D), D5(33), D6(29), D7(1C)) can be read simultaneously (in parallel).

[0137] The data read from the memory unit 33 as described above is output to the register unit 34.

[0138] In the sixth data read process, as can be seen from FIG. 31, since eight data items to be read are read via separate access buses without duplication, the eight data items (data D0(2A), D1(1C), D2(32), D3(36), D4(3B), D5(39), D6(37), D7(23)) can be read simultaneously (in parallel).

[0139] The data read from the memory unit 33 as described above is output to the register unit 34.

[0140] In the seventh data read process, as can be seen from FIG. 32, since eight data items to be read are read via separate access buses without duplication, the eight data items (data D0(33), D1(32), D2(35), D3(3C), D4(1A), D5(8), D6(10), D7(24)) can be read simultaneously (in parallel).

[0141] The data read from the memory unit 33 as described above is output to the register unit 34.

[0142] In the eighth data read process, as can be seen from FIG. 33, six pieces of data to be read are read via separate access buses without duplication, so that the six pieces of data (data D0(3C), D1(3B), D2(2A), D3(10), D6(1F), D7(26)) can be read simultaneously (in parallel).

[0143] The data read from the memory unit 33 as described above is output to the register unit 34.

[0144] ≪Ninth data read process (Extra block read process)≫ In the ninth data read process, as shown in FIG. 34, additional data (Extra block) is read from the memory area Mem_extra_blk for storage. Specifically, the following read process is executed. (1) The data (address) in the fifth row of the eighth column of the address list after the adjustment process is "FF", which is set in the additional data (Extra block), and the original address is "3A", so the memory access control unit 32 generates a control signal Ctl_r (4) to read the data D4(3A) of the temporary memory Tmem_4 from the memory area Mem_extra_blk for storage of the temporary memory Tmem_4, and outputs the control signal to the temporary memory Tmem_4. The temporary memory Tmem_4 reads the data D4(3A) from the memory area Mem_extra_blk according to the control signal Ctl_r (4) (Since the data D4(3A) is stored in the eleventh bank, the data is read via the eleventh access bus). (2) The data (address) in the 6th row of the address in the 8th column of the address list after the adjustment process is "FF", and it is set in the additional data (Extra block). Since the original address is "10", the memory access control unit 32 reads the data D5(10) of the temporary memory Tmem_5 from the storage memory area Mem_extra_blk of the temporary memory Tmem_5, and generates a control signal Ctl_r (5) and outputs the control signal to the temporary memory Tmem_5. The temporary memory Tmem_5 reads the data D5(10) from the storage memory area Mem_extra_blk according to the control signal Ctl_r (5) (Since the data D5(10) is stored in the first bank, the data is read through the first access bus).

[0145] In the 9th data read process, as can be seen from FIG. 34, since the two data to be read are read through different access buses without duplication, the two data (data D4(3A), D5(10)) can be read simultaneously (in parallel).

[0146] The data read from the memory unit 33 as described above is output to the register unit 34.

[0147] As described above, all the data of the blocks corresponding to the address list after the adjustment process (the data of blocks block0 to block8 (the feature amount data after quantization processing)) have been read from the memory unit 33 and output to the register unit 34.

[0148] When a predetermined amount of data is output from the memory unit 33 and the register unit 34 stores the data, the register unit 34 outputs the predetermined amount of data to the quantization data memory unit 4. For example, when the register unit 34 stores data output simultaneously from the memory unit 33 via a plurality of access buses, the register unit 34 may output the data to the quantization data memory unit 4. Alternatively, after all the data of the block corresponding to the address list after the adjustment process is output from the memory unit 33 and stored, the register unit 34 may output the data to the quantization data memory unit 4.

[0149] When all the data of the block corresponding to the address list after the adjustment process is output from the memory unit 33 (when the output is completed), the memory access control unit 32 of the convolution processing data processing unit 3 outputs the data including the address list after the adjustment process (including the original address set in the additional data when there is additional data (Extra block) set data (address)) to the convolution processing unit 5 as data D2_list_adj.

[0150] When the convolution processing unit 5 inputs data D2_list_adj from the memory access control unit 32 of the convolution processing data processing unit 3, the convolution processing unit 5 reads out the feature amount data (quantized feature amount data) corresponding to the address list after the adjustment process from the quantization data memory unit 4. Further, the convolution processing unit 5 inputs the data (kernel weight coefficient data) of the kernel (filter) of the convolution process to be applied to the data of the block corresponding to the address list after the adjustment process, Din_w.

[0151] The kernel weight coefficient data K included in the data Din_w (L) is K (L) =[K0 (L) ,K1 (L) ,K2 (L) ,K3 (L) ,K4 (L) ,K5 (L) ,K6 (L) ,K7 (L) 0≦L≦c out -1​ c out : Number of output channels Assuming this, the convolution process corresponds to the following mathematical formula.

Equation

[0152] The convolution processing unit 5 obtains the data Do after convolution processing by processing the data D0(8), D0(16), D0(19), D0(1A), D0(28), D0(2A), D0(33), D0(3C) corresponding to the address of the first row of the address list after adjustment processing as follows (see Fig. 35). (1) The convolution processing unit 5 detects data in which the upper 5 bits of the address X of each data D0(X) have the same value (the value obtained by shifting X 3 bits to the right). (2) For the address X of the data detected in (1), the remainder when divided by 8 is calculated, and the weight coefficient data of the kernel with the calculated remainder value as the subscript is obtained, and the multiplication and addition operation of the weight coefficient data and the data at address X is performed.

[0153] Data D h A specific example of the convolution process in (i) will be described. (1) As shown in Fig. 35, in the data D0(8), D0(16), D0(19), D0(1A), D0(28), D0(2A), D0(33), D0(3C) corresponding to the address of the first row of the address list after adjustment processing, there is no data in which the upper 5 bits of the address i are "0", so the convolution processing result data Do 00 is "0". (2) As shown in FIG. 35, among the data D0(8), D0(16), D0(19), D0(1A), D0(28), D0(2A), D0(33), D0(3C) corresponding to the address of the first row of the address list after the adjustment process, the data whose upper 5 bits of the address i are "1" is only the data D0(8). Since the address of the data D0(8) is "8", the remainder with 8 as the modulus is "0", and the weight coefficient data of the kernel to be multiplied by the data D0(8) is K0 (L) is. Therefore, the convolution process result data Do 01 is Do 01 = D0(8) × K0 (L) is (3) As shown in FIG. 35, among the data D0(8), D0(16), D0(19), D0(1A), D0(28), D0(2A), D0(33), D0(3C) corresponding to the address of the first row of the address list after the adjustment process, the data whose upper 5 bits of the address i are "2" is only the data D0(16). Since the address of the data D0(16) is "16", the remainder with 8 as the modulus is "6", and the weight coefficient data of the kernel to be multiplied by the data D0(16) is K6 (L) is. Therefore, the convolution process result data Do 02 is Do 02 = D0(16) × K6 (L) is (4) As shown in FIG. 35, among the data D0(8), D0(16), D0(19), D0(1A), D0(28), D0(2A), D0(33), D0(3C) corresponding to the address of the first row of the address list after the adjustment process, the data whose upper 5 bits of the address i are "3" are the data D0(19) and the data D0(1A). Since the address of the data D0(19) is "19", the remainder with 8 as the modulus is "1", and the weight coefficient data of the kernel to be multiplied by the data D0(19) is K1 (L) is. Since the address of the data D0(1A) is "1A", the remainder with 8 as the modulus is "2", and the weight coefficient data of the kernel to be multiplied by the data D0(1A) is K2 (L) is

[0154] Therefore, the convolution process result data Do 03 is Do 03 = D0(19) × K1 (L) + D0(1A) × K2 (L) as follows. (5) As shown in FIG. 35, in the data D0(8), D0(16), D0(19), D0(1A), D0(28), D0(2A), D0(33), D0(3C) corresponding to the address of the first row of the address list after the adjustment process, there is no data where the upper 5 bits of the address i are "4". Therefore, the convolution process result data Do 04 is "0". (6) As shown in FIG. 35, in the data D0(8), D0(16), D0(19), D0(1A), D0(28), D0(2A), D0(33), D0(3C) corresponding to the address of the first row of the address list after the adjustment process, the data where the upper 5 bits of the address i are "5" are the data D0(28) and the data D0(2A). Since the address of the data D0(28) is "28", the remainder modulo 8 is "0", and the kernel weight coefficient data to be multiplied by the data D0(28) is K0 (L) as follows. Since the address of the data D0(2A) is "2A", the remainder modulo 8 is "2", and the kernel weight coefficient data to be multiplied by the data D0(2A) is K2 (L) as follows.

[0155] Therefore, the convolution process result data Do 05 is Do 05 = D0(28) × K1 (L) + D0(2A) × K2 (L) as follows. (7) As shown in FIG. 35, among the data D0(8), D0(16), D0(19), D0(1A), D0(28), D0(2A), D0(33), D0(3C) corresponding to the address in the first row of the address list after the adjustment process, the data whose upper 5 bits of the address i are "6" is only the data D0(33). Since the address of the data D0(33) is "33", the remainder modulo 8 is "3", and the weight coefficient data of the kernel to be multiplied by the data D0(33) is K3 (L) is. Therefore, the convolution process result data Do 06 is Do 06 = D0(33) × K3 (L) is (8) As shown in FIG. 35, among the data D0(8), D0(16), D0(19), D0(1A), D0(28), D0(2A), D0(33), D0(3C) corresponding to the address in the first row of the address list after the adjustment process, the data whose upper 5 bits of the address i are "7" is only the data D0(3C). Since the address of the data D0(3C) is "3C", the remainder modulo 8 is "4", and the weight coefficient data of the kernel to be multiplied by the data D0(3C) is K4 (L) is. Therefore, the convolution process result data Do 07 is Do 07 = D0(3C) × K4 (L) is

[0156] By processing as described above, the convolution processing unit 5 obtains the data (output data) Do in the first row of the convolution processed data (h = 0 data) Do 00 ~ Do 07 (see the upper right figure in FIG. 35).

[0157] The convolution processing unit 5 obtains the convolution processed data Do in the same manner as above for the data sequence of the data D h (i) corresponding to the addresses in the second and subsequent rows of the address list after the adjustment process (see FIG. 36).

[0158] Furthermore, the convolution processing unit 5 performs the same processing as above for the kernels K (1) ~K (cout-1) to obtain the data (output data) Do (1) (=D*K (1) ) after convolution processing, Do (2) (=D*K (2) ), ···, Do (cout-1) (=D*K (cout-1) )(*: convolution operator).

[0159] Thus, the convolution processing unit 5 obtains the data (output data) Do (1) ~Do (cout-1) .

[0160] <<Summary>> As described above, in the CNN data processing apparatus 100, for the address list of the data remaining after pruning by the unstructured pruning method, an address list adjustment process is performed, and based on the adjusted address list, from the temporary memory Tmem_k of the memory unit 33 in which the feature amount data after quantization processing is stored, data is read, so that the number of data that can be read simultaneously can be increased. That is, in the CNN data processing apparatus 100, considering the number of banks that can be accessed in parallel in the temporary memory Tmem_k of the memory unit 33 in which the feature amount data after quantization processing is stored, an address list adjustment process is performed so that more data can be output simultaneously. Therefore, based on the adjusted address list, by reading data from the temporary memory Tmem_k of the memory unit 33 in which the feature amount data after quantization processing is stored, the number of data that can be read simultaneously can be increased.

[0161] Furthermore, in the CNN data processing device 100, for data that could not be read simultaneously, it is set as additional data (Extra block), and the data is copied to the memory area Mem_extra_blk for storing additional data in the temporary memory Tmem_k. Then, data reading processing is performed at a timing when the data can be read. Therefore, regardless of the situation of the data remaining after pruning by the unstructured pruning method, the data (feature quantity data after quantization processing) that can surely be the target of the convolution processing can be read. That is, in the CNN data processing device 100, regardless of the situation of the data remaining after pruning by the unstructured pruning method, the data (feature quantity data after quantization processing) that can surely be the target of the convolution processing can be read, and the convolution processing can surely be executed.

[0162] Also, in the CNN data processing device 100, in the convolution processing unit 5, the convolution processing is performed by performing a sum-of-products operation using only the data remaining after pruning by the unstructured pruning method. Therefore, the amount of calculation can be reduced, and as a result, high-speed processing can be realized.

[0163] In this way, in the CNN data processing device 100, by providing the temporary memory Tmem_k in the memory unit 33 and performing data access based on the address list after the adjustment process, it is possible to easily realize the lightweight of the neural network model while adopting high-performance unstructured pruning.

[0164] ≪First Modification Example≫ Next, the first modification example of the first embodiment will be described.

[0165] In this modification example, the operation of the CNN data processing device 100 will be described for the case (an example) where the following settings are made. Note that the settings for the CNN data processing device 100 are not limited to the following settings, and other settings may also be used. (1) The block of feature amount data to be subjected to convolution processing (feature amount data input to the CNN data processing device 100) has a channel number of "64" (c in = 64), a height of "1" (h = 1), and a width of "1" (w = 1) (processing target block). (2) The size of the kernel (filter) for performing convolution processing in the channel direction (Pointwise convolution) is 1×1×64 (the size in the height direction is "1", the size in the width direction is "1", and the size in the channel direction is "64") (Conv1d kernel with a channel direction size of "64"), and the number of kernels (filters) is "8". (3) The processing target block block h for each kernel of the convolution processing (h: kernel number, 0 ≦ h ≦ 7) is, as in the first embodiment, an 8×8 block (assuming that 64 data processed by one kernel are stored in an 8×8 block), and the data pruned by pruning (excluded from the processing target of the convolution processing) is, as in the first embodiment, the data shown by the gray rectangles in FIGS. 5 and 6, and the data not pruned by pruning (subjected to the convolution processing) is the data shown by the white rectangles in FIGS. 5 and 6. Specifically, in the processing target block block h the data not pruned by pruning (subjected to the convolution processing) shall be as follows (an example). Note that each data in the block block h is denoted as D h (X), h is the height position, and X is an 8-bit address expressed in hexadecimal (an address (8-bit data indicating the address) for specifying the position within the block block h ). Block block0: D0(8), D0(16), D0(19), D0(1a), D0(28), D0(2a), D0(33), D0(3c) Block block1: D1(7), D1(10), D1(12), D1(15), D1(1a), D1(1c), D1(32), D1(3b) Block block2: D2(13), D2(19), D2(1f), D2(23), D2(26), D2(2a), D2(32), D2(35) block3: D3(b), D3(10), D3(14), D3(1a), D3(1f), D3(32), D3(36), D3(3c) block4: D4(0), D4(2), D4(7), D4(19), D4(1a), D4(2d), D4(3a), D4(3b) block5: D5(8), D5(10), D5(16), D5(17), D5(30), D5(31), D5(33), D5(39) block6: D6(1), D6(d), D6(10), D6(16), D6(1f), D6(24), D6(29), D6(37) block7: D7(c), D7(f), D7(18), D7(1b), D7(1c), D7(23), D7(24), D7(26) Note that the positions of the data that were not pruned (i.e., the data to be subjected to the convolution process) are assumed to be known in advance (the positions of the data that were not pruned (i.e., the data to be subjected to the convolution process) are stored during or after the pruning process and can be referenced (acquired) after the pruning process). (4) The memory unit 33 has eight (M = 8) temporary memories Tmem_0 to Tmem_7. The temporary memory Tmem_k (k: integer, 0 ≦ k ≦ 7) has 16 banks, and access buses are connected to each of the plurality of banks. Through the access buses, a plurality of data (a plurality of data of 16 or less) can be read simultaneously (in parallel) from the 16 banks, or a plurality of data (a plurality of data of 16 or less) can be written simultaneously (in parallel) to the 16 banks (see FIG. 3). Also, it is assumed that an additional data storage memory area (extra block memory area) is set (secured) in the temporary memory Tmem_k (k: integer, 0 ≦ k ≦ 7) (see FIG. 3). Also, for each kernel of the convolution process, a processing target block block h (h: kernel number, 0 ≦ h ≦ 7) is to be assigned.

[0166] When the above settings are made (in the case of this modification), in the CNN data processing device 100, in the quantization processing unit 1, the address list adjustment processing unit 2, the convolution processing data processing unit 3, and the quantized data memory unit 4, the same processing as in the first embodiment is executed.

[0167] Then, when the convolution processing unit 5 inputs the data D2_list_adj from the memory access control unit 32 of the convolution processing data processing unit 3, it reads out the feature amount data (quantized processing after feature amount data) corresponding to the adjusted address list from the quantized data memory unit 4. Also, the convolution processing unit 5 inputs the data (kernel weight coefficient data) Din_w of the kernel (filter) of the convolution process to be applied to the data of the block corresponding to the adjusted address list. Then, the convolution processing unit 5 executes the processing corresponding to the following mathematical formula to obtain the convolution processing result data y(n) by the nth kernel K (n)

Equation

[0168] In the CNN data processing device 100 of this modified example, as described above, by simply reading out the address position (the position of the address of block n (corresponding to the temporary memory Tmem_n)) and the weight coefficient of the kernel at the same position, and accumulating and adding the two, the convolution process result data can be obtained, so that the convolution process can be executed extremely efficiently.

[0169] [Other Embodiments] In the above embodiment, in the CNN data processing device 100, the case where the weight coefficient data Din_w (weight filter (kernel)) is input to the convolution processing unit 5 and the convolution process is executed using the weight coefficient data Din_w (weight filter (kernel)) has been described. However, the present invention is not limited to this. For example, the convolution processing unit 5 may perform vector decomposition processing on the weight filter (kernel), decompose it into a basis matrix and a real coefficient vector, and input the decomposed basis matrix and real coefficient vector to perform the convolution process. In this case, the convolution process is executed using the basis matrix (a matrix whose elements are only basis values (integer values)) and the data D3 output from the quantization data memory unit 4, and then the process using the real coefficient vector is performed, so that most of the convolution operations can be integer operations, and furthermore, the speed of the convolution process can be increased.

[0170] Also, in the above embodiment, in the CNN data processing device 100, the process for the case of performing the convolution process in the channel direction (Pointwise convolution) has been described. However, the present invention is not limited to this. For example, the present invention may be applied to the process for the convolution process in the spatial direction. Also in this case, considering the position within the block of the feature amount data remaining after pruning and the storage position of the temporary memory Tmem_k, an address list adjustment process may be performed so that there is no data (address) for which data reading processing is performed by the same access bus, and an adjusted address list may be obtained.

[0171] In the above-described embodiment, the CNN data processing apparatus 100 has been described assuming the case where it is applied to a CNN. However, the present invention is not limited to this, and the present invention may be applied to neural network models other than CNNs.

[0172] In the above-described embodiment, the CNN data processing apparatus 100 has been described assuming the case where it is applied to the processing of the convolutional layer of a CNN. However, the present invention is not limited to this, and the present invention may be applied to the processing of the fully-connected layer.

[0173] In the above-described embodiment, the case where the number of banks of the temporary memory Tmem_k is "16" in the CNN data processing apparatus 100 has been described. However, the present invention is not limited to this, and the number of banks of the temporary memory Tmem_k may be other numbers. Even when the number of banks of the temporary memory Tmem_k is set to other numbers, an address list adjustment process may be performed so that the number of data that can be accessed simultaneously increases.

[0174] In the above-described embodiment, the CNN data processing apparatus 100 has been described assuming the case where a method (depthwise separable convolution) that divides normal convolution processing into two types, (1) depthwise convolution (spatial-direction convolution processing) and (2) pointwise convolution (channel-direction convolution processing), is adopted. However, the present invention is not limited to this. In the CNN data processing apparatus 100, for example, the convolution processing data processing of the above-described embodiment may be applied to normal convolution processing.

[0175] In the above-described embodiment, the case where CNN data processing is performed on the data after quantization processing in the CNN data processing apparatus 100 has been described. However, the present invention is not limited to this. For example, data (feature data) on which quantization processing has not been performed may be input to the CNN data processing unit 2 of the CNN data processing apparatus 100, and CNN data processing by the CNN data processing unit 2 may be executed on the data.

[0176] Also, the configuration of the memory unit 33 of the CNN data processing device 100 is not limited to that described in the above embodiment, and the number of temporary memories and the number of data that can be accessed simultaneously in each temporary memory (the number of access buses) can be set to any number.

[0177] Also, each block (each functional unit) of the CNN data processing device 100 described in the above embodiment may be individually integrated into one chip by a semiconductor device such as an LSI, or may be integrated into one chip so as to include part or all of them. Further, each block (each functional unit) of the pose data generation system, the CG data system, and the pose data generation device described in the above embodiment may be realized by a plurality of semiconductor devices such as LSIs.

[0178] Here, it is described as an LSI, but depending on the degree of integration, it may also be referred to as an IC, a system LSI, a super LSI, or an ultra LSI.

[0179] Also, the method of integrating into a circuit is not limited to an LSI, and it may be realized by a dedicated circuit or a general-purpose processor. After manufacturing the LSI, an FPGA (Field Programmable Gate Array) that can be programmed, or a reconfigurable processor that can reconfigure the connection and setting of circuit cells inside the LSI may be used.

[0180] Also, part or all of the processing of each functional block in each of the above embodiments may be realized by a program. And part or all of the processing of each functional block in each of the above embodiments is performed by a central processing unit (CPU) in a computer. Also, the program for performing each process is stored in a storage device such as a hard disk or a ROM, and is read from the ROM or the RAM and executed.

[0181] In addition, each process of the above-described embodiment may be realized by hardware, or may be realized by software (including the case where it is realized together with an OS (Operating System), middleware, or a predetermined library). Furthermore, it may be realized by a mixed process of software and hardware.

[0182] For example, when each functional unit of the above-described embodiment is realized by software, a hardware configuration shown in FIG. 38 (for example, a hardware configuration in which a CPU, GPU, processor, ROM, RAM, memory, input unit, output unit, etc. are connected by a bus Bus) may be used to realize each functional unit by software processing.

[0183] Also, when each functional unit of the above-described embodiment is realized by software, the software may be realized using a single computer having the hardware configuration shown in FIG. 38, or may be realized by distributed processing using a plurality of computers.

[0184] Also, the execution order of the processing method in the above-described embodiment is not necessarily limited to the description of the above-described embodiment, and the execution order can be changed without departing from the gist of the invention. Also, in the processing method in the above-described embodiment, some steps may be executed in parallel with other steps without departing from the gist of the invention. Also, in the processing method in the above-described embodiment, the processing executed in parallel may be executed serially (sequentially).

[0185] A computer program for causing a computer to execute the above-described method and a computer-readable recording medium recording the program are included in the scope of the present invention. Here, examples of the computer-readable recording medium include a flexible disk, a hard disk, a CD-ROM, an MO, a DVD, a DVD-ROM, a DVD-RAM, a large-capacity DVD, a next-generation DVD, and a semiconductor memory.

[0186] The above computer program is not limited to being recorded on the above recording medium, and may be transmitted via a telecommunication line, a wireless or wired communication line, a network such as the Internet, or the like.

[0187] Also, the term "section" may be a concept including "circuitry". Circuitry may be realized in whole or in part by hardware, software, or a combination of hardware and software.

[0188] The functions of the elements disclosed herein may be implemented using a general-purpose processor, a dedicated processor, an integrated circuit, an ASIC ("application-specific integrated circuit"), a conventional circuit configuration, and / or a circuit configuration or processing circuit configuration including a combination thereof, which are configured to execute the disclosed elements or programmed to execute the disclosed functions. When a processor includes transistors and other circuit configurations therein, it is regarded as a processing circuit configuration or a circuit configuration. In the present disclosure, a circuit configuration, unit, or means is hardware that executes the recited functions or hardware programmed to execute the functions. The hardware may be any hardware disclosed herein or other known hardware programmed or configured to execute the recited functions. When the hardware is a processor that may be regarded as a certain type of circuit configuration, the circuit configuration, means, or unit is a combination of hardware and software, software used to configure the hardware, and / or a processor.

[0189] Note that the specific configuration of the present invention is not limited to the foregoing embodiments, and various changes and modifications are possible without departing from the gist of the invention.

Description of Reference Numerals

[0190] 100 Data processing device for CNN 3 CNN data processing unit (convolution processing data device) 31 Address list storage unit 32 Memory access control unit (access control unit) 33 Memory unit Tmem_0~Tmem_M-1 Temporary memory

Claims

1. A data processing device for convolution processing used in a neural network model, a plurality of temporary memories for storing feature data, an access control unit that controls data writing and / or data reading of the plurality of temporary memories, an address list storage unit that stores an adjusted address list obtained by performing an address list adjustment process so that more data can be output in parallel from the plurality of temporary memories, based on the position of the memory area where the feature data of the block to be subjected to convolution processing of the remaining feature data pruned by the unstructured pruning method is stored in the temporary memory, and the number of banks that can be accessed in parallel in the temporary memory, comprising: each of the plurality of temporary memories includes a plurality of banks that can be accessed in parallel, the access control unit performs data writing control so that the feature data of the block to be subjected to convolution processing is stored in the temporary memory, the access control unit performs data reading control so that feature data is read from the plurality of temporary memories based on the adjusted address list stored in the address list storage unit, A data processing device for convolution processing.

2. The address list is two-dimensional array data, the row data of the address list is an address list of the remaining feature data pruned by the unstructured pruning method within a predetermined block to be subjected to convolution processing, the address list constitutes two-dimensional array data by including a plurality of the row data, the adjusted address list is a list obtained by executing a process of obtaining an adjusted address list by performing an address swapping process, which is a process of detecting an address where the positions of the banks of the temporary memory in which the feature data is stored overlap among the addresses in the same column of the address list, and swapping one of the overlapping addresses with another address in the same row of the address list when an address where the positions of the banks of the temporary memory in which the feature data is stored overlap is detected, The data processing device for convolution processing according to claim 1.

3. Each of the plurality of temporary memories has an additional data storage memory area. The access control unit stores, in the additional data storage memory area, the feature amount data determined to be unable to be read out in parallel from the plurality of temporary memories in the adjusted address list. The data processing device for convolution processing according to claim 1 or 2.

4. The access control unit performs data read control so as to read out the feature amount data stored in the additional data storage memory area during a period other than when it controls to read out the feature amount data from the plurality of temporary memories based on the adjusted address list. The data processing device for convolution processing according to claim 3.

5. An adjusted address list acquisition method for acquiring an adjusted address list used in the data processing device for convolution processing according to claim 1 or 2, the adjusted address list acquisition method being executed by a device including a memory and a processor, a first step in which the processor acquires an address list that is a list of addresses within a block for which convolution processing of the feature amount data remaining after pruning by an unstructured pruning method is to be performed; a second step in which the processor acquires an adjusted address list by performing an address list adjustment process so as to increase the data that can be output in parallel from the plurality of temporary memories based on the position of the memory area in which the feature amount data of the block for which convolution processing of the temporary memory is to be performed is stored and the number of banks that can be accessed in parallel in the temporary memory; An adjusted address list acquisition method comprising the above.

6. The address list is two-dimensional array data. The row data of the address list is an address column of the feature amount data remaining after pruning by an unstructured pruning method within a predetermined block for which convolution processing is to be performed. The address list constitutes two-dimensional array data by including a plurality of the row data. The second step is In the addresses in the same column of the address list, detect an address where the positions of the banks of the temporary memory in which the feature amount data is stored overlap. When an address where the positions of the banks of the temporary memory in which the feature amount data is stored overlap is detected, perform an address swapping process, which is a process of swapping one of the overlapping addresses with another address in the same row of the address list, to obtain an adjusted address list. The adjusted address list acquisition method according to claim 5.

7. The second step is If, even when the address swapping process is repeatedly executed, a state where no address where the positions of the banks of the temporary memory in which the feature amount data is stored overlap is not detected, set at least one of the overlapping addresses as additional data, which is an address stored in the memory area for additional data storage, and rewrite the address set as additional data with a value indicating that it has been set as additional data. The adjusted address list acquisition method according to claim 6.

8. A program for executing the adjusted address list acquisition method according to claim 5 by a computer.