Data processing device for convolution processing

The data processing device for CNNs addresses the processing time bottleneck by utilizing multiple bank memories and an access control unit for parallel data access, resulting in a high-performance and high-speed CNN model with reduced read process executions and overall processing time.

JP2025088892APending Publication Date: 2025-06-12MEGACHIPS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023203692
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Conventional CNN models, such as Mobilenet, face a bottleneck in processing time due to the longer time required for reading out feature data compared to convolution processing, making it difficult to shorten the total processing time.

Method used

A data processing device for convolutional neural networks is designed with multiple bank memories and an access control unit that allows parallel data access and optimized data storage and retrieval strategies, reducing the number of times feature data is read and improving overall processing efficiency.

Benefits of technology

The proposed solution enables a high-performance and high-speed CNN model by reducing the number of executions of the feature data read process and shortening the entire convolution processing time, including the read process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025088892000001_ABST
    Figure 2025088892000001_ABST
Patent Text Reader

Abstract

To provide a data processing device for convolution processing capable of reducing the number of times that processing for reading feature quantity data is executed, shortening the time required for the entire convolution processing including the processing for reading the feature quantity data, and performing data processing for implementing a high-performance and high-speed CNN model.SOLUTION: In a data processing device for convolution processing, (1) a plurality of access buses is provided in each of a plurality of bank memories Tmem_k of a memory unit 22, and data for a plurality of channels can be accessed simultaneously (in parallel); and (2) since different (independent) bank memories Tmem_k are allocated for each height direction of a convolution processing target region (region to be convolved with a kernel), a plurality of data items having different height directions can be accessed simultaneously (in parallel).SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to data processing technology for convolutional neural networks, and particularly to technology for processing feature data used in convolutional neural networks (data processing apparatus for convolutional processing).

Background Art

[0002] In recent years, technologies using neural network models that can accurately realize a variety of applications have attracted attention. In technologies using neural network models, learning data is used to perform learning processing of the neural network model to obtain a learned model, and prediction processing (inference processing) is performed using the obtained learned model. As a result, technologies using neural network models can accurately realize a variety of applications. As a technology using a neural network model that realizes high value in fields such as image recognition, a technology using a convolutional neural network model (CNN: Convolutional Neural Network) has attracted attention.

[0003] Furthermore, lightweight technologies have been developed to enable the use of convolutional neural network models even on mobile terminals and the like that do not have abundant computing resources. As such a technology, for example, a technology called Mobilenet (a lightweight technology for CNN models) has been developed (see, for example, Non-Patent Document 1).

[0004] In a technology called Mobilenet (a technique for lightweighting CNN models), a method is adopted that divides ordinary convolution processing, called Depthwise separable convolution, into two types: (1) Depthwise convolution (spatial-direction convolution processing) and (2) Pointwise convolution (channel-direction convolution processing). By doing so, the number of parameters in the CNN model is reduced. As a result, it is possible to realize a lightweight and high-performance CNN model that can also be installed in mobile terminals and the like that do not have abundant computing resources.

Prior Art Documents

Non-Patent Documents

[0005]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] However, in the above conventional technology (Mobilenet), it is necessary to perform processing for frequently reading out feature data, and the time required for reading out feature data is longer than the time required for the sum-of-products calculation for convolution processing. As a result, the time required to execute CNN processing becomes longer. That is, in the above conventional technology (Mobilenet), even if the processing of the CNN operation itself (convolution processing) is speeded up, the time required for reading out feature data becomes the bottleneck (critical path), and it is difficult to shorten the total processing time (CNN processing) including the processing for reading out feature data.

[0007] In view of the above problems, the present invention aims to realize a high-performance and high-speed CNN model that can reduce the number of executions of the process of reading feature amount data and shorten the time required for the entire convolution process including the read process of the feature amount data, and perform data processing for realizing a convolution process data processing device capable of performing such data processing.

Means for Solving the Problems

[0008] In order to solve the above problems, a first invention is a data processing device for convolution processing used in a convolutional neural network model, including a plurality of bank memories for storing feature amount data, and an access control unit for performing data writing and / or data read control of the plurality of bank memories.

[0009] The feature amount data is three-dimensional data specified by positions in the width direction, height direction, and channel direction.

[0010] Each of the plurality of bank memories has a plurality of access buses so that data can be accessed in parallel.

[0011] The access control unit performs data write control so that the feature amount data whose height direction position is the first value is stored in the bank memory assigned to the first value among the plurality of bank memories, and further stores a plurality of feature amount data having the same width direction position and continuous channel direction positions in a memory area of an address that can be accessed in parallel by a plurality of buses.

[0012] In this data processing device for convolution processing, (1) each of the plurality of bank memories Tmem_k of the memory unit 22 is provided with a plurality of access buses, and data for a plurality of channels can be accessed simultaneously (in parallel), and (2) different (independent) bank memories Tmem_k are assigned for each height direction of the convolution processing target area (the area to be convolved with the kernel), so that a plurality of data with different height directions can be accessed simultaneously (in parallel).

[0013] Note that the "feature amount data" may be the feature amount data after quantization processing.

[0014] The second invention is the first invention, wherein the access control unit performs data read control on a plurality of bank memories so as to read out feature amount data having the same position in the width direction and continuous positions in the height direction for a plurality of channels in a read unit period.

[0015] Accordingly, in this convolution processing data device, in a period of one data read process (read unit period), data of h×1 (h rows and 1 column, h: position in the height direction) in the convolution processing target area can be read out for a plurality of channels.

[0016] The third invention is the first or second invention, wherein when a data group obtained by reading out feature amount data having the same position in the width direction and continuous positions in the height direction from a plurality of bank memories for a plurality of channels is set as a plurality of channel h×1 data groups, the access control unit acquires, as the number of output systems, the number of overlapping plurality of channel h×1 data groups in the same read unit period according to the position of the convolution processing target area, and controls the plurality of bank memories so that the plurality of channel h×1 data groups corresponding to the acquired number of output systems are output from the plurality of bank memories.

[0017] In this convolution processing data device, the number of output systems Num_sys, which is the number of sets of overlapping data (sets of h×1 data in the convolution processing target area), is acquired according to the position (slid position) of the convolution processing target area, and the sets of overlapping data (sets of h×1 data in the convolution processing target area) corresponding to the acquired number of output systems Num_sys can be output in different systems (in parallel).

[0018] Accordingly, in this convolution processing data device, the number of times of reading overlapping data can be reduced by sliding the position of the convolution processing target area.

[0019] The fourth invention is the third invention, further comprising a register unit capable of storing data by address specification.

[0020] The register unit inputs a plurality of channel h×1 data groups output from a plurality of bank memories, uses the size in the width direction of the kernel of the convolution process applied to the plurality of channel h×1 data groups as an offset value, and sequentially writes the feature amount data with consecutive positions in the height direction included in the plurality of channel h×1 data groups to the memory area of the address of the register unit offset by the offset value.

[0021] In this data device for convolution processing, a register unit is provided. In the register unit, according to the size (shape) of the convolution processing target area (the size (shape) of the kernel), the data read from the memory unit 22 is written to discontinuous addresses (addresses added with a predetermined offset value (corresponding to the size in the width direction of the kernel, in the case of a 3×3 kernel, the offset value is "3")), and after all the data (feature amount data) in the convolution processing target area is aligned (after all the data in the convolution processing target area is written at consecutive addresses of the register unit), all the data in the convolution processing target area can be output.

[0022] Thereby, in this data device for convolution processing, all the data (data to be subjected to convolution processing) in the convolution processing target area is output as data arranged in the order of performing the convolution operation, and can be written, for example, to the memory unit for quantization data. Then, the data arranged in the order of performing the convolution operation is read from, for example, the memory unit for quantization data, and the convolution processing can be executed at high speed by performing the convolution processing using the kernel weight coefficient data applied to the data.

[0023] In this way, in this data device for convolution processing, by simply providing a functional unit for performing the above processing, it is possible to obtain data arranged in the order of performing convolution operations while reducing the number of times of reading duplicate data. Therefore, in this data device for convolution processing, it is possible to reduce the number of times of executing the process of reading feature amount data, and shorten the time required for the entire convolution processing including the reading process of feature amount data, and perform data processing for realizing a high-performance and high-speed CNN model.

[0024] A fifth invention is the fourth invention, wherein after the feature amount data of the area to be subjected to the convolution process by the kernel is stored in the memory area of the continuous addresses of the register unit, the feature amount data stored in the memory area of the continuous addresses is output. In this data device for convolution processing, after the feature amount data of the area to be subjected to the convolution process by the kernel is stored in the memory area of the continuous addresses of the register unit, that is, after the feature amount data stored by being offset by the offset value is not in a discontinuous state but in a continuous storage state (stored in the memory area of the continuous addresses of the register unit), the feature amount data stored in the memory area of the continuous addresses is output. Therefore, in this data device for convolution processing, it is possible to guarantee that the feature amount data is output from the register unit after the feature amount data of the area to be subjected to the convolution process by the kernel is complete.

[0025] A sixth invention is the fifth invention, wherein the register unit outputs the feature amount data stored in the memory area of the continuous addresses all at once or in the order of the continuous addresses of the register unit.

[0026] As a result, in this convolutional processing data device, the feature amount data of the area to be subjected to the convolutional processing by the kernel (a plurality of feature amount data to be subjected to the convolutional processing by the kernel (for example, when the kernel is a 3×3 kernel, nine feature amount data included in the area to be convolved by the kernel) are output in a state where they are arranged together or in the order of performing the convolutional processing by the kernel (in the order of performing the product-sum calculation of the weight coefficient data of the kernel).

Effect of the Invention

[0027] According to the present invention, it is possible to realize a data processing device for convolutional processing that can perform data processing for realizing a high-performance and high-speed CNN model capable of reducing the number of execution times of the process of reading out feature amount data and shortening the time required for the entire convolutional processing including the reading out process of the feature amount data.

Brief Description of the Drawings

[0028]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Embodiments for Carrying Out the Invention

[0029] [First Embodiment] The first embodiment will be described below with reference to the drawings.

[0030] <1.1: Configuration of the CNN Data Processing Apparatus> FIG. 1 is a schematic configuration diagram of a CNN data processing apparatus 100 according to the first embodiment.

[0031] FIG. 2 is a schematic configuration diagram of the CNN data processing unit 2 of the CNN data processing apparatus 100 according to the first embodiment.

[0032] As shown in FIG. 1, the CNN data processing device 100 includes a quantization processing unit 1, a CNN data processing unit 2 (convolution processing data device), a quantized data memory unit 3, and a convolution processing unit 4. The CNN data processing device 100 inputs feature data Din_f and weight coefficient data Din_w (weight filter (kernel)), executes convolution processing (convolution processing using the feature data and the weight coefficient data), and obtains (outputs) the processing result data Dout of the convolution processing.

[0033] The quantization processing unit 1 inputs the feature data Din_f, executes quantization processing on the feature data Din_f, and outputs the data after the quantization processing as data D1 to the CNN data processing unit 2.

[0034] As shown in FIG. 2, the CNN data processing unit 2 includes a memory access control unit 21, a memory unit 22 including M (M: natural number) bank memories (Tmem_0 to Tmem_M-1), and a register unit 23.

[0035] The memory access control unit 21 is a control unit for performing access control (data write processing control, data read processing control) on the M bank memories of the memory unit 22. The memory access control unit 21 is a functional unit for independently (in parallel) performing data write processing control and data read processing control on the M bank memories Tmem_0 to Tmem_M-1 of the memory unit 22. The memory access control unit 21 outputs a control signal Ctl_w for performing data write processing control and / or a control signal Ctl_r for performing data read processing control on the memory unit 22. Specifically, the memory access control unit 21 outputs a control signal Ctl_w for performing data write processing control on the bank memory Tmem_k (k: natural number, 0 ≦ k ≦ M-1) of the memory unit 22 (k) and / or a control signal Ctl_r for performing data read processing control (k)Output it. Note that for the bank memory Tmem_k (k is a natural number, 0 ≦ k ≦ M - 1) of the memory unit 22, a control signal Ctl_w for performing data write processing control is the control signal Ctl_w (k) is denoted as, and a control signal Ctl_r for performing data read processing control for the bank memory Tmem_k of the memory unit 22 is the control signal Ctl_r (k) is denoted as.

[0036] Also, the memory access control unit 21 is a control unit for performing access control (data write processing control, data read processing control) on the register unit 23. The memory access control unit 21 outputs a control signal Ctl_reg for performing access control of the register unit 23 to the register unit 23.

[0037] As shown in FIG. 2, the memory unit 22 includes M (M is a natural number) bank memories Tmem_0 to Tmem_M - 1.

[0038] The bank memory Tmem_k (k is a natural number, 0 ≦ k ≦ M - 1) is a memory that can write predetermined data to a predetermined address of the bank memory Tmem_k and read the data stored at the address from the predetermined address of the bank memory Tmem_k. The bank memory Tmem_k writes the data D1 output from the quantization processing unit 1 to the address of the bank memory Tmem_k specified by the control signal Ctl_w for data write processing from the memory access control unit 21 according to the control signal Ctl_w (k) and outputs the read data to the register unit 23. (k) Also, the bank memory Tmem_k reads the data stored at the address from the address of the bank memory Tmem_k specified by the control signal Ctl_r for data read processing from the memory access control unit 21 according to the control signal Ctl_r (k) and outputs the read data to the register unit 23. (k) according to the control signal Ctl_w

[0039] The register unit 23 has a memory (register) that can write data to a predetermined area by specifying an address and can read data stored in the predetermined area by specifying an address. The register unit 23 inputs the data output from the memory unit 22 and the control signal Ctl_reg output from the memory access control unit 21. The register unit 23 writes the data output from the memory unit 22 to a predetermined address of the register unit 23 according to the control signal Ctl_reg. Further, the register unit 23 outputs the data at a predetermined address of the register unit 23 as data D2 to the quantization data memory unit 3 according to the control signal Ctl_reg.

[0040] The quantization data memory unit 3 has a memory that can store data. The memory can write data to a predetermined area by specifying an address and can read data stored in the predetermined area by specifying an address. The quantization data memory unit 3 inputs the data D2 output from the CNN data processing unit 2 and stores the data D2. Further, the quantization data memory unit 3 outputs the stored data as data D3 to the convolution processing unit 4 (the quantization data memory unit 3 inputs a data read command from a control unit (not shown) or the convolution processing unit 4 and outputs the data at a predetermined address as data D3 to the convolution processing unit 4 according to the data read command).

[0041] The convolution processing unit 4 inputs the weight coefficient data Din_w (weight filter (kernel)) and the data D3 output from the quantization data memory unit 3. The convolution processing unit 4 executes a convolution process using the data D3 and the weight coefficient data Din_w and outputs the data after the convolution process as data Dout.

[0042] <1.2: Operation of the CNN Data Processing Device> The operation of the CNN data processing device 100 configured as described above will be described below.

[0043] FIG. 3 is a diagram for explaining CNN processing (convolution processing for a CNN model) by (1) Depthwise convolution (spatial direction convolution processing) and (2) Pointwise convolution (channel direction convolution processing).

[0044] FIG. 4 is a diagram for explaining data stored in each bank memory (in the case of four bank memories (an example)) of the CNN data processing unit 2 of the CNN data processing apparatus 100 for CNN.

[0045] FIG. 5 is a diagram for explaining data access to the bank memory of the CNN data processing unit 2 of the CNN data processing apparatus 100 for CNN.

[0046] FIGS. 6 and 7 are diagrams for explaining the relationship between read data and blocks when performing data read processing of the CNN data processing unit 2 of the CNN data processing apparatus 100 for CNN.

[0047] FIGS. 8 to 10 are flowcharts of CNN data processing executed by the CNN data processing apparatus 100 for CNN.

[0048] FIG. 11 is a timing chart of data write processing and data read processing of the CNN data processing unit 2 of the CNN data processing apparatus 100 for CNN.

[0049] FIG. 12 is a diagram for explaining data access to the bank memory of the CNN data processing unit 2 of the CNN data processing apparatus 100 for CNN.

[0050] FIGS. 13 to 15 are diagrams for explaining data read processing from the bank memory of the CNN data processing unit 2 of the CNN data processing apparatus 100 for CNN.

[0051] FIGS. 16 and 17 are timing charts of data read processing of the CNN data processing unit 2 of the CNN data processing apparatus 100 for CNN.

[0052] Figures 18 to 20 are diagrams for explaining the relationship between the read data, blocks, and the data write addresses of the register unit 23 when the data read process of the CNN data processing unit 2 of the CNN data processing apparatus 100 is performed.

[0053] Figure 21 is a diagram for explaining the relationship (including the relationship for channels 0 to 4) between the read data, blocks, and the data write addresses of the register unit 23 when the data read process of the CNN data processing unit 2 of the CNN data processing apparatus 100 is performed.

[0054] As shown in FIG. 3, when adopting a method of dividing the normal convolution process into two types: (1) Depthwise convolution (spatial direction convolution process), and (2) Pointwise convolution (channel direction convolution process), in the Depthwise convolution (spatial direction convolution process), the time taken for the read process of the feature data becomes longer than the time taken for the sum-of-products calculation for the convolution process. As a result, the time for executing the CNN process becomes longer. To address this, in the CNN data processing apparatus 100, the CNN data processing unit 2 performs (A) data write processes for N (N: a natural number of 2 or more) channels (Ch 0 ~Ch N-1 ) in parallel, and (B) data read processes for N (N: a natural number of 2 or more) channels (Ch 0 ~Ch N-1 ) in parallel. Further, in the CNN data processing apparatus 100, the CNN data processing unit 2 simultaneously reads N-channel data from each of the M bank memories Tmem_0 to Tmem_M - 1, and performs, by parallel processing, a process of simultaneously outputting data of a number of systems corresponding to the kernel of the CNN process.

[0055] Hereinafter, for the sake of convenience of explanation, the operation of the CNN data processing apparatus 100 will be described for the following case (an example). Note that the settings for the CNN data processing apparatus 100 are not limited to the following settings, and other settings may also be used. (1) The number of channels N of the feature amount data (the feature amount data input to the CNN data processing apparatus 100) to be subjected to the convolution process is "4" (N = 4, channels Ch 0 ~Ch 3 ). (2) The number of bank memories of the memory unit 22 of the CNN data processing unit 2 is "4" (M = 4, bank memories Tmem_0 to Tmem_3 (bank 0 ~bank 3 ). (3) As shown in FIG. 4, the bank memory Tmem_k (bank k ) of the memory unit 22 of the CNN data processing unit 2 can store 8×4 (8 rows and 4 columns) data (32 data), and the data stored at the address of the i-th row and j-th column of the bank memory Tmem_k (bank k ) is denoted as D bnk (i, j) (i, j: natural numbers, 0≦i≦7, 0≦j≦N - 1 (N = 4)) (i corresponds to the position in the width direction of the feature amount data (feature map), j corresponds to the channel (position in the channel direction) of the feature amount data (feature map), and k corresponds to the position in the height direction of the feature amount data (feature map)). (4) As shown in FIG. 5, the bank memory Tmem_k (bank k ) of the memory unit 22 of the CNN data processing unit 2 has an access bus that can access the data in the same row simultaneously, and for the data in different rows of the bank memory Tmem_k (bank k ), data read processing and data write processing can be performed independently (in parallel) (in the case of FIG. 5, for the data D 0 ) of the 0th row of the bank memory Tmem_0 (bank bn0 (0, 0)~D bn0 (0, 3), data read processing is performed, and at the same time (in parallel), for the data D bn0(7,0) to D bn0 (It shows the state of performing data writing processing for (7,3)). (5) The size of the kernel (weight coefficient filter) for depthwise convolution (spatial convolution processing) is 3×3 (the weight coefficients of the kernel are represented by a 3×3 matrix). (6) As shown in FIGS. 5 and 6, the area (feature map) to be filtered by the kernel (the object of convolution processing) is an area of size 4×8 (4×8 feature data), the stride of the convolution processing is "1", and there is no padding. That is, the blocks in FIGS. 5 and 6 0 , block 1 , block 2 , block 3 , block 4 , block 5 , block 6 , ··· By shifting the object to be filtered by the 3×3 kernel, the object of the convolution processing is specified.

[0056] The feature data Din_f is input to the quantization processing unit 1.

[0057] The quantization processing unit 1 executes quantization processing on the feature data Din_f, and outputs the data after quantization processing as data D1 to the CNN data processing unit 2.

[0058] As shown in FIG. 8, the CNN data processing unit 2 performs (1) data writing processing (for data of N channels (Ch 0 ~Ch N-1 (In this embodiment, N = 4)) to the bank memory Tmem_k (bank k ) of the memory unit 22 of the CNN data processing unit 2), and (2) data reading processing (for data of N channels (Ch 0 ~Ch N-1 (In this embodiment, N = 4)) to the bank memory Tmem_k (bank kExecute the data reading process from

[0059] Here, the processing of the CNN data processing unit 2 will be described with reference to the flowcharts of FIGS. 8 to 10.

[0060] (Step S1w): In step S1w, the CNN data processing unit 2 executes a data writing process. For example, as shown in FIGS. 11 and 12, during the period T 0 (at time t 0 ~t 1 ), from the quantization processing unit 1 to the CNN data processing unit 2, (1) Data D bn0 (4,0), D bn0 (4,1), D bn0 (4,2), D bn0 (4,3) (2) Data D bn1 (4,0), D bn1 (4,1), D bn1 (4,2), D bn1 (4,3) (3) Data D bn2 (4,0), D bn2 (4,1), D bn2 (4,2), D bn2 (4,3) (4) Data D bn3 (4,0), D bn3 (4,1), D bn3 (4,2), D bn3 (4,3) (The data D bnh (i,j) indicates data where the position in the width direction of the channel Ch j is i and the position in the height direction is h) When the data D1 including it is output, the CNN data processing unit 2 performs the following processing. In the above data, since the position h in the height direction is 0 to 3, the bank memories to be subjected to the writing process are set to bank0 to bank3. For this purpose, in the CNN data processing unit 2, variables hws (a variable for designating the start bank memory to be subjected to the writing process) and hwe (a variable for designating the end bank memory to be subjected to the writing process) for designating the bank memory to be subjected to the writing process are set to hws = 0 and hwe = 3 (hws, hwe: natural numbers, 0 ≦ hws ≦ M - 1, 0 ≦ hwe ≦ M - 1, hws < hwe) (step S11w).

[0061] (Step S12w_bnk_hws): In step S12w_bnk_hws, the bank memory bank hws (bank memory Tmem_hws) (hws = 0) is subjected to a writing process for the number of channels (Ch 0 ~Ch N-1 ). Specifically, the following processing is executed. (1) The memory access control unit 21 of the CNN data processing unit 2, during the period T 0 (time t 0 ~t 1 ), for the data D bn0 (4, 0), D bn0 (4, 1), D bn0 (4, 2), D bn0 (4, 3), four pieces of data are written in parallel (using four access buses (see FIG. 5)) to consecutive addresses of the bank memory Tmem_0 of the memory unit 22 (bank memory bank hws , hws = 0), and a control signal Ctl_w (0) is output to the bank memory Tmem_0. Then, according to the control signal Ctl_w (0) , the bank memory Tmem_0 of the memory unit 22 stores the data D bn0 (4, 0), D bn0 (4, 1), D bn0 (4, 2), D bn0The four pieces of data (4,3) are written in parallel (using four access buses) to consecutive address areas of the bank memory Tmem_0 of the memory unit 22 (see FIG. 12).

[0062] (Steps S12w_bnk_1, S12w_bnk_2): In steps S12w_bnk_1 and S12w_bnk_2, the bank memory bank 1 , bank 2 (Bank memory Tmem_1, Tmem_2) each have the number of channels (Ch 0 ~Ch N-1 ) data write processing is executed. Specifically, the following processing is executed. (2) The memory access control unit 21 of the CNN data processing unit 2 0 (Time t 0 ~t 1 During the period, data D bn1 (4,0), D bn1 (4,1), D bn1 (4,2), D bn1 The four pieces of data (4, 3) are transferred in parallel (using four access buses (see FIG. 5)) to the bank memory Tmem_1 (bank 1 ) to write to consecutive addresses. (1) The control signal Ctl_w is output to the bank memory Tmem_1. (1) Accordingly, the bank memory Tmem_1 of the memory unit 22 stores the data D bn1 (4,0), D bn1 (4,1), D bn1 (4,2), D bn1 The four pieces of data (4,3) are written in parallel (using four access buses) to areas with consecutive addresses in the bank memory Tmem_1 of the memory unit 22 (see FIG. 12). (3) The memory access control unit 21 of the CNN data processing unit 2 0 (Time t 0 ~t 1 During the period, data D bn2 (4,0), Dbn2 (4,1), D bn2 (4,2), D bn2 Four data of (4,3) are written in parallel (using four access buses (see Fig. 5)) to consecutive addresses of the bank memory Tmem_2 of the memory unit 22 (bank memory bank 2 ) by the control signal Ctl_w (2) output to the bank memory Tmem_2. Then, according to the control signal Ctl_w (2) , the bank memory Tmem_2 of the memory unit 22 writes the data D bn2 (4,0), D bn2 (4,1), D bn2 (4,2), D bn2 Four data of (4,3) are written in parallel (using four access buses) to the area where the addresses of the bank memory Tmem_2 of the memory unit 22 are consecutive (see Fig. 12).

[0063] (Step S12w_bnk_hwe): In step S12w_bnk_hwe, the bank memory bank hwe (bank memory Tmem_hwe) (hwe = 3) performs the writing process for the number of channels (Ch 0 ~Ch N-1 ). Specifically, the following process is performed. (4) The memory access control unit 21 of the CNN data processing unit 2 controls the period T 0 (time t 0 ~t 1 ), and the data D bn3 (4,0), D bn3 (4,1), D bn3 (4,2), D bn3 (4,3) are written in parallel (using four access buses (see Fig. 5)) to consecutive addresses of the bank memory Tmem_3 of the memory unit 22 (bank hwe , hwe = 3) by the control signal Ctl_w (3) output to the bank memory Tmem_3. Then, the control signal Ctl_w (3)Accordingly, for the bank memory Tmem_3 of the memory unit 22, data D bn3 (4,0), D bn3 (4,1), D bn3 (4,2), D bn3 (4,3) of the four pieces of data are written in parallel (using four access buses) to a region where the addresses of the bank memories Tmem_3 of the memory unit 22 are consecutive (see FIG. 12).

[0064] In the above description, the case where hws = 0 and hwe = 3 are set and four consecutive pieces of data are written with respect to the position in the height direction of the feature data (when data is written in parallel to the bank memories Tmem_0 to Tmem_3) has been described. However, it is not limited thereto, and hws and hwe may be set to other values, and data may be written in parallel to a plurality of bank memories specified according to hws and hwe. Further, since the bank memories Tmem_k (k: natural number, 0 ≦ k ≦ M - 1) are each independently accessible, for example, the data writing process to the bank memory Tmem_k in the above process may be executed in parallel.

[0065] Also, in the above, the period T 0 (time t 0 ~t 1 of the period) has been described for the data writing process. However, for the period T 1 (time t 1 ~t 2 ) of (1) in FIG. 11, the period T 2 (time t 1 ~t 2 ) of (2), and the period T 3 (time t 2 ~t 3 ) of (3), the same processing as above is also executed. In the cases of the above (1) to (4), the following data is written to the bank memories Tmem_0 (bank 0 ) to Tmem_3 (bank 0 ) (see FIG. 12). (1) Period T 1 (time t 1 ~t 2 of the period): · Memory for bank Tmem_0 (bank 0 ) is written with data D bn0 (5,0), data D bn0 (5,1), data D bn0 (5,2), data D bn0 (5,3). · Memory for bank Tmem_1 (bank 1 ) is written with data D bn1 (5,0), data D bn1 (5,1), data D bn1 (5,2), data D bn1 (5,3). · Memory for bank Tmem_2 (bank 2 ) is written with data D bn2 (5,0), data D bn2 (5,1), data D bn2 (5,2), data D bn2 (5,3). · Memory for bank Tmem_3 (bank 3 ) is written with data D bn3 (5,0), data D bn3 (5,1), data D bn3 (5,2), data D bn3 (5,3). (2) Period T 2 (at time t 2 ~t 3 ): · Memory for bank Tmem_0 (bank 0 ) is written with data D bn0 (6,0), data D bn0 (6,1), data D bn0 (6,2), data D bn0 (6,3). · Memory for bank Tmem_1 (bank 1 ) is written with data D bn1 (6,0), data D bn1 (6,1), data D bn1 (6,2), data D bn1 (6,3). · Memory for bank Tmem_2 (bank 2 ) is written with data Dbn2 (6,0), Data D bn2 (6,1), Data D bn2 (6,2), Data D bn2 (6,3) is written. · Bank memory Tmem_3 (bank 3 ) is written with Data D bn3 (6,0), Data D bn3 (6,1), Data D bn3 (6,2), Data D bn3 (6,3) is written. (3) Period T 3 (Time t 3 ~t 4 period): · Bank memory Tmem_0 (bank 0 ) is written with Data D bn0 (7,0), Data D bn0 (7,1), Data D bn0 (7,2), Data D bn0 (7,3) is written. · Bank memory Tmem_1 (bank 1 ) is written with Data D bn1 (7,0), Data D bn1 (7,1), Data D bn1 (7,2), Data D bn1 (7,3) is written. · Bank memory Tmem_2 (bank 2 ) is written with Data D bn2 (7,0), Data D bn2 (7,1), Data D bn2 (7,2), Data D bn2 (7,3) is written. · Bank memory Tmem_3 (bank 3 ) is written with Data D bn3 (7,0), Data D bn3 (7,1), Data D bn3 (7,2), Data D bn3 (7,3) is written.

[0066] (Step S2w): In step S2w, it is determined whether there is any data to be the target of the data writing process in the CNN data processing unit 2. If there is still data to be processed, the process returns to S1w and the same process as above is executed. On the other hand, if there is no data to be processed, the data writing process in the CNN data processing unit 2 is terminated.

[0067] (Step S1r): In step S1r, the CNN data processing unit 2 executes a data reading process. For example, as shown in FIGS. 11 and 12, during the period T 0 (at time t 0 ~t 1 ), from the memory unit 22 of the CNN data processing unit 2, (1) data D bn0 (0,0), D bn0 (0,1), D bn0 (0,2), D bn0 (0,3) (2) data D bn1 (0,0), D bn1 (0,1), D bn1 (0,2), D bn1 (0,3) (3) data D bn2 (0,0), D bn2 (0,1), D bn2 (0,2), D bn2 (0,3) (data D bnh (i,j) indicates data where the position in the width direction of the channel Ch j is i and the position in the height direction is h) When reading, the CNN data processing unit 2 performs the following process. In the above data, since the position h in the height direction is 0 to 2, the bank memories to be the targets of the reading process are bank 0 ~bank 2It is set to (the size of the kernel (weight coefficient filter) of the depthwise convolution (spatial convolution process) is 3×3, and the size in the height direction is "3", so it is set like this). For this reason, in the CNN data processing unit 2, variables hrs (a variable that designates the bank memory to be the target of the read process) and hre (a variable that designates the end bank memory of the read process target) for designating the bank memory to be the target of the read process are set to hrs = 0 and hre = 2 (hrs, hre: natural numbers, 0 ≦ hrs ≦ M - 1, 0 ≦ hre ≦ M - 1, hrs < hre) (step S110r).

[0068] (Step S111r_bnk_hrs (hrs = 0)): In step S111r_bnk_hrs, the bank memory bank hrs (bank memory Tmem_hrs) (hrs = 0), data for the number of channels (Ch 0 ~Ch N-1 )(N = 4) read process is executed. Specifically, the following process is executed. (1) The memory access control unit 21 of the CNN data processing unit 2, during period T 0 (time t 0 ~t 1 period), data D bn0 (0,0), D bn0 (0,1), D bn0 (0,2), D bn0 (0,3) four pieces of data, in parallel (using four access buses (see Fig. 5)), from the area where the addresses of the bank memory Tmem_0 of the memory unit 22 (bank memory bank hrs , hrs = 0) are consecutive, instructs the control signal Ctl_r (0) to be output to the bank memory Tmem_0. Then, according to the control signal Ctl_r (0) , the bank memory Tmem_0 of the memory unit 22, data D bn0 (0,0), D bn0 (0,1), D bn0 (0,2), D bn0Four pieces of data at (0, 3) (data for four channels) are read out in parallel (using four access buses) from a region where the addresses of the bank memories Tmem_0 in the memory unit 22 are consecutive (see FIG. 12) (for the relationship between the data for four channels read out in parallel and the spatial position of the feature data (feature map), see FIG. 13).

[0069] (Step S111r_bnk_1): In step S111r_bnk_1, the bank memory bank 1 reads out data for the number of channels (Ch 0 to Ch N-1 )(N = 4). Specifically, the following processing is executed. (2) The memory access control unit 21 of the CNN data processing unit 2 outputs a control signal Ctl_r 0 during period T 0 (time t 1 to t bn1 ) to instruct to read out four pieces of data D bn1 (0, 0), D bn1 (0, 1), D bn1 (0, 2), D 1 )(bank memory bank (1) ) in parallel (using four access buses (see FIG. 5)) from a region where the addresses of the bank memory Tmem_1 in the memory unit 22 are consecutive. Then, according to the control signal Ctl_r (1) , the bank memory Tmem_1 in the memory unit 22 reads out four pieces of data D bn1 (0, 0), D bn1 (0, 1), D bn1 (0, 2), D bn1 (0, 3) (data for four channels) in parallel (using four access buses) from a region where the addresses of the bank memory Tmem_1 in the memory unit 22 are consecutive (see FIG. 12) (for the relationship between the data for four channels read out in parallel and the spatial position of the feature data (feature map), see FIG. 14).

[0070] (Step S111r_bnk_hre (hre = 2)): In step S111r_bnk_hre, from the bank memory bank hre (bank memory Tmem_hre) (hre = 2), data reading processing for the number of channels (Ch 0 ~Ch N-1 )(N = 4) is executed. Specifically, the following processing is executed. (3) The memory access control unit 21 of the CNN data processing unit 2 is during period T 0 (at time t 0 ~t 1 ), for data D bn2 (0,0), D bn2 (0,1), D bn2 (0,2), D bn2 (0,3), the four pieces of data are read in parallel (using four access buses (see Fig. 5)) from a region where the addresses of the bank memory Tmem_2 in the memory unit 22 (bank memory bank 2 ) are consecutive, and a control signal Ctl_r (2) is output to the bank memory Tmem_2. Then, according to the control signal Ctl_r (2) , the bank memory Tmem_2 in the memory unit 22 reads the four pieces of data D bn2 (0,0), D bn2 (0,1), D bn2 (0,2), D bn2 (0,3) (data for four channels) in parallel (using four access buses) from a region where the addresses of the bank memory Tmem_2 in the memory unit 22 are consecutive (see Fig. 12) (For the relationship between the four channels of data read in parallel and the spatial position of the feature data (feature map), see Fig. 15).

[0071] (Steps S112r_bnk_hrs (hrs = 0) to S112r_bnk_hre (hre = 2)): In steps S112r_bnk_hrs (hrs = 0) to S112r_bnk_hre (hre = 2), a process for determining the number of output systems Num_sys is executed. As shown in FIGS. 6, 16, and 17, during period T 0 the data read in 0 is the data in the first column of block 1 , so there is no data (duplicate data) common to other blocks (block 5 ). Therefore, in each step of steps S112r_bnk_hrs (hrs = 0) to S112r_bnk_hre (hre = 2), the number of output systems Num_sys = 1 is determined.

[0072] (Steps S113r_bnk_hrs (hrs = 0) to S113r_bnk_hre (hre = 2)): In steps S113r_bnk_hrs (hrs = 0) to S113r_bnk_hre (hre = 2), a set of data of the number of output systems Num_sys (= 1) is output simultaneously (register write process is executed). Specifically, the following process is executed.

[0073] The memory access control unit 21 of the CNN data processing unit 2 generates a control signal Ctl_r 0 for reading the data D bnk (0, 0) of the bank memory Tmem_k during period T (k) and outputs the control signal Ctl_r (k) to the bank memory Tmem_k. Also, a control signal Ctl_reg is generated to instruct that the data D bnk (0, 0) of the bank memory Tmem_k be written to a region of a predetermined address in the register unit 23, and the control signal Ctl_reg is output to the register unit 23. The bank memory Tmem_k reads the data D (k) according to the control signal Ctl_r bnk (0, 0), and the register unit 23 writes the data D output from the bank memory Tmem_k according to the control signal Ctl_reg bnkWrite (0, 0) to a predetermined address. Assume that the predetermined address is indicated by the control signal Ctl_reg.

[0074] Period T 0 During this period, for the data of channel 0 (Ch 0 ) and channel k (Ch k ) (1 ≤ k ≤ 3), the data output from the bank memory Tmem_k and the write destination address of the register section 23 of the data are as follows. ≪Period T 0 ≫ (output to one system (Num_sys = 1)) Channel 0 (Ch 0 ): (1) Data D 0 of the bank memory Tmem_0 (bank bn0 )(0, 0) → Address adr00 of the register section 23 (Ch0) (2) Data D 1 of the bank memory Tmem_1 (bank bn1 )(0, 0) → Address adr03 of the register section 23 (Ch0) (3) Data D 2 of the bank memory Tmem_2 (bank bn2 )(0, 0) → Address adr06 of the register section 23 (Ch0) Channel k (Ch k )(k: natural number, 1 ≤ k ≤ 3): (1) Data D 0 of the bank memory Tmem_0 (bank bn0 )(0, k) → Address adr00 of the register section 23 (Chk) (2) Data D 1 of the bank memory Tmem_1 (bank bn1 )(0, k) → Address adr03 of the register section 23 (Chk) (3) Data D 2Data D of ( ) bn2 (0, k) → Address adr06 of register unit 23 (Chk) Figure 18 shows the relationship between the above data (data of channel 0) and the write destination address of register unit 23. As shown in Figure 18, during period T 0 in, data D bn0 (0, 0), data D bn1 (0, 0), data D bn2 (0, 0) are written in the areas of address adr00 (Ch0) , address adr03 (Ch0) (= address adr00 (Ch0) + 3 addresses), address adr06 (Ch0) (= address adr03 (Ch0) + 3 addresses), respectively. That is, during period T 0 in, data D bn0 (0, 0), data D bn1 (0, 0), data D bn2 (0, 0) are written every 3 addresses (skipping 3 addresses) of register unit 23 (the thick rectangles in Figure 18 indicate the data to be written). This is because the area to be subjected to the convolution process (kernel size) is 3 × 3. Therefore, according to the area to be subjected to the convolution process (kernel size), the 3 × 3 data is reshaped into 1 × 9 data so that it can be output to the quantization data memory unit 3. Note that the addresses adr00 (Chk) ~ adr08 (Chk) (k: natural number, 0 ≤ k ≤ N - 1) of register unit 23 are assumed to be consecutive addresses.

[0075] (Step S12r): In step S12r, a determination process is executed to determine whether a predetermined amount of data has been output from the bank memory Tmem_k of memory unit 22 to register unit 23. When period T 0 ends, since all the data in the area to be subjected to the convolution process (kernel size) has not been output from the bank memory Tmem_k of memory unit 22 to register unit 23 (period T in Figure 18 0Refer to the part of , and return the process to step S11r.

[0076] (Step S1r) (Period T 1 ): In step S1r, the CNN data processing unit 2 executes data reading processing. For example, as shown in FIGS. 11 and 17, during period T 1 (time t 1 ~t 2 ), from the memory unit 22 of the CNN data processing unit 2, (1) Data D bn0 (1,0), D bn0 (1,1), D bn0 (1,2), D bn0 (1,3) (2) Data D bn1 (1,0), D bn1 (1,1), D bn1 (1,2), D bn1 (1,3) (3) Data D bn2 (1,0), D bn2 (1,1), D bn2 (1,2), D bn2 (1,3) (Data D bnh (i,j) indicates data where the position in the width direction of channel Ch j is i and the position in the height direction is h) is read. In this case, the CNN data processing unit 2 performs the following processing. In the above data, since the position h in the height direction is 0 to 2, the bank memories to be read are bank 0 ~bank 2It is set to (since the size of the kernel (weight coefficient filter) of the depthwise convolution (spatial direction convolution process) is 3×3 and the size in the height direction is "3", it is set in this way). For this reason, in the CNN data processing unit 2, variables hrs (a variable that designates the bank memory to be the target of the read process) and hre (a variable that designates the end bank memory to be the target of the read process) for designating the bank memory to be the target of the read process are set to hrs = 0 and hre = 2 (hrs, hre: natural numbers, 0 ≤ hrs ≤ M - 1, 0 ≤ hre ≤ M - 1, hrs < hre) (step S110r).

[0077] (Step S111r_bnk_hrs (hrs = 0)): In step S111r_bnk_hrs, the bank memory bank hrs (The bank memory Tmem_hrs) (hrs = 0), data for the number of channels (Ch 0 ~Ch N-1 )(N = 4) is read. Specifically, the following processes are executed. (1) The memory access control unit 21 of the CNN data processing unit 2, during the period T 1 (at time t 1 ~t 2 ), data D bn0 (1,0), D bn0 (1,1), D bn0 (1,2), D bn0 (1,3) of four data items are read in parallel (using four access buses (see Fig. 5)) from the region where the addresses of the bank memory Tmem_0 of the memory unit 22 (the bank memory bank hrs , hrs = 0) are consecutive, and a control signal Ctl_r (0) is output to the bank memory Tmem_0. Then, according to the control signal Ctl_r (0) , the bank memory Tmem_0 of the memory unit 22 outputs data D bn0 (1,0), D bn0 (1,1), D bn0 (1,2), D bn0Four data of (1, 3) (data for four channels) are read out in parallel (using four access buses) from an area where the addresses of the bank memories Tmem_0 in the memory unit 22 are consecutive.

[0078] (Step S111r_bnk_1): In step S111r_bnk_1, from the bank memory bank 1 the read process for the number of channels (Ch 0 ~Ch N-1 )(N = 4) of data is executed. Specifically, the following processes are executed. (2) The memory access control unit 21 of the CNN data processing unit 2 controls the data D 1 during the period T 1 (time t 2 ~t bn1 (1, 0), D bn1 (1, 1), D bn1 (1, 2), D bn1 (1, 3) of four data to be read out in parallel (using four access buses (see Fig. 5)) from an area where the addresses of the bank memory Tmem_1 (bank memory bank 1 ) in the memory unit 22 are consecutive by outputting a control signal Ctl_r (1) to the bank memory Tmem_1. Then, according to the control signal Ctl_r (1) , the bank memory Tmem_1 in the memory unit 22 reads out four data of D bn1 (1, 0), D bn1 (1, 1), D bn1 (1, 2), D bn1 (1, 3) (data for four channels) in parallel (using four access buses) from an area where the addresses of the bank memory Tmem_1 in the memory unit 22 are consecutive.

[0079] (Step S111r_bnk_hre (hre = 2)): In step S111r_bnk_hre, from the bank memory bank hre (bank memory Tmem_hre) (hre = 2), the read process for the number of channels (Ch 0~Ch N-1 The read process of the data of (N = 4) is executed. Specifically, the following processes are executed. (3) The memory access control unit 21 of the CNN data processing unit 2 is during the period T 1 (at time t 1 ~t 2 period), for the data D bn2 (1,0), D bn2 (1,1), D bn2 (1,2), D bn2 (1,3), the four data are read in parallel (using four access buses (see Fig. 5)) from the area where the addresses of the bank memories Tmem_2 of the memory unit 22 are consecutive. The control signal Ctl_r 2 ) is output to the bank memory Tmem_2. Then, according to the control signal Ctl_r (2) , the bank memory Tmem_2 of the memory unit 22 reads the four data of D (2) (1,0), D bn2 (1,0), D bn2 (1,1), D bn2 (1,2), D bn2 (1,3) (data for four channels) in parallel (using four access buses) from the area where the addresses of the bank memory Tmem_2 of the memory unit 22 are consecutive.

[0080] (Steps S112r_bnk_hrs (hrs = 0) to S112r_bnk_hre (hre = 2)): In steps S112r_bnk_hrs (hrs = 0) to S112r_bnk_hre (hre = 2), the process of determining the number of output systems Num_sys is executed. As shown in Figs. 6, 16, and 17, the data read in the period T 1 is the data in the second column of block 0 and is also the data in the first column of block 1 . Therefore, in each step of steps S112r_bnk_hrs (hrs = 0) to S112r_bnk_hre (hre = 2), the number of output systems Num_sys = 2 is determined.

[0081] (Steps S113r_bnk_hrs (hrs = 0) to S113r_bnk_hre (hre = 2)): In steps S113r_bnk_hrs (hrs=0) to S113r_bnk_hre (hre=2), data for the number of output systems Num_sys (=2) is output simultaneously (register write processing is executed). Specifically, the following processing is executed.

[0082] The memory access control unit 21 of the CNN data processing unit 2 1 In this case, the data D of the bank memory Tmem_k is bnk Control signal Ctl_r for reading (1,0) (k) The control signal Ctl_r is generated. (k) to the bank memory Tmem_k, and the data D bnk The bank memory Tmem_k generates a control signal Ctl_reg that instructs that (1,0) is written to a region at a predetermined address of the register unit 23, and outputs the control signal Ctl_reg to the register unit 23. (k) According to Data D bnk The register unit 23 reads out (1,0) and outputs the data D output from the bank memory Tmem_k in accordance with the control signal Ctl_reg. bnk (1,0) is written to a specific address. Note that the specific address is assumed to be specified by the control signal Ctl_reg.

[0083] Period T 1 In the channel 0 (Ch 0 ) data and channel k(Ch k ) (1≦k≦3), the data output from the bank memory Tmem_k and the address of the write destination of the register unit 23 for that data are as follows: ≪Period T 1 ≫(2 output systems (Num_sys=2)) Channel 0 (Ch 0 ): (1) Bank memory Tmem_0(bank0 )'s data D bn0 (1,0) → Address adr01 of register section 23 (Ch0) → Address adr10 of register section 23 (Ch0) (2) Bank memory Tmem_1 (bank 1 )'s data D bn1 (1,0) → Address adr04 of register section 23 (Ch0) → Address adr13 of register section 23 (Ch0) (3) Bank memory Tmem_2 (bank 2 )'s data D bn2 (1,0) → Address adr07 of register section 23 (Ch0) → Address adr16 of register section 23 (Ch0) Channel k (Ch k )(k: natural number, 1 ≦ k ≦ 3): (1) Bank memory Tmem_0 (bank 0 )'s data D bn0 (1,k) → Address adr01 of register section 23 (Chk) → Address adr10 of register section 23 (Chk) (2) Bank memory Tmem_1 (bank 1 )'s data D bn1 (1,k) → Address adr04 of register section 23 (Chk) → Address adr13 of register section 23 (Chk) (3) Bank memory Tmem_2 (bank 2 )'s data D bn2 (1,k) → Address adr07 of register section 23 (Chk) → Address adr16 of register section 23 (Chk) Figure 18, Figure 19 (period T1 shows the relationship between the above data (data of channel 0) and the write destination address of the register section 23 of the register section 23.

[0084] As shown in FIG. 18, during period T 1 the data D bn0 (1, 0), the data D bn1 (1, 0), the data D bn2 (1, 0) are respectively written in the areas of address adr01 (Ch0) , address adr04 (Ch0) (= address adr01 (Ch0) + 3 addresses), address adr07 (Ch0) (= address adr04 (Ch0) + 3 addresses). That is, during period T 1 the data D bn0 (1, 0), the data D bn1 (1, 0), the data D bn2 (1, 0) are written every 3 addresses (skipping 3 addresses) in the register section 23 (the thick rectangles in FIG. 18 indicate the data to be written). This is because the area to be subjected to the convolution process (the size of the kernel) is 3×3. Therefore, according to the area to be subjected to the convolution process (the size of the kernel), the 3×3 data is reshaped into 1×9 data so that it can be output to the quantization data memory section 3.

[0085] Also, as shown in FIG. 19, during period T 1 the data D bn0 (1, 0), the data D bn1 (1, 0), the data D bn2 (1, 0) are respectively written in the areas of address adr10 (Ch0) , address adr13 (Ch0) (= address adr10 (Ch0) + 3 addresses), address adr16 (Ch0) (= address adr13 (Ch0) + 3 addresses). That is, during period T 1 the data D bn0 (1, 0), the data D bn1 (1, 0), the data Dbn2 (1,0) is written for every three addresses of the register unit 23 (skipping three addresses each time) (the thick-rectangle in FIG. 19 indicates the data to be written). Since the area to be subjected to the convolution process (the size of the kernel) is 3×3, in order to reshape the 3×3 data into 1×9 data according to the area to be subjected to the convolution process (the size of the kernel) and output it to the quantization data memory unit 3. Note that the address adr10 of the register unit 23 (Chk) ~adr18 (Chk) (k: natural number, 0≦k≦N-1) shall be consecutive addresses.

[0086] (Step S12r): In step S12r, a determination process is executed to check whether a predetermined amount of data has been output from the bank memory Tmem_k of the memory unit 22 to the register unit 23. At the end of the period T 1 when it ends, since not all the data in the area to be subjected to the convolution process (the size of the kernel) has been output from the bank memory Tmem_k of the memory unit 22 to the register unit 23 (refer to the part of the period T in FIG. 18 1 and the part of the period T in FIG. 19 1 ), the process returns to step S11r.

[0087] (Step S1r)(period T 2 ): In step S1r, the CNN data processing unit 2 executes a data reading process. For example, as shown in FIGS. 11 and 17, during the period T 2 (time t 2 ~t 3 ), from the memory unit 22 of the CNN data processing unit 2, (1) Data D bn0 (2,0), D bn0 (2,1), D bn0 (2,2), D bn0 (2,3) (2) Data D bn1 (2,0), D bn1 (2,1), D bn1 (2,2), D bn1 (2,3) (3) Data D bn2 (2,0), D bn2 (2,1), D bn2 (2,2), D bn2 (2,3) (Data D bnh (i, j) indicates the data where the position in the width direction of channel Ch j is i and the position in the height direction is h) is read. In this case, the CNN data processing unit 2 performs the following processing. Note that in the above data, since the position h in the height direction is 0 to 2, the bank memory to be the target of the read process is bank 0 ~ bank 2 is set (Since the size of the kernel (weight coefficient filter) of Depthwise convolution (spatial convolution process) is 3 × 3 and the size in the height direction is "3", it is set in this way). For this purpose, in the CNN data processing unit 2, variables hrs (a variable specifying the start bank memory of the read process) and hre (a variable specifying the end bank memory of the read process) for specifying the bank memory to be the target of the read process are set to hrs = 0 and hre = 2 (hrs, hre: natural numbers, 0 ≦ hrs ≦ M - 1, 0 ≦ hre ≦ M - 1, hrs < hre) (step S110r).

[0088] (Step S111r_bnk_hrs (hrs = 0)): In step S111r_bnk_hrs, from the bank memory bank hrs (bank memory Tmem_hrs) (hrs = 0), read processes for the number of channels (Ch 0 ~ Ch N-1 )(N = 4) of data are executed. Specifically, the following processing is executed. (1) The memory access control unit 21 of the CNN data processing unit 2, during the period T 2 (time t 2 ~ t 3 period), for data D bn0 (2,0), D bn0 (2,1), D bn0 (2,2), Dbn0 Four data of (2, 3) are read in parallel (using four access buses (see FIG. 5)) from a region where the addresses of the bank memories Tmem_0 (bank memories bank hrs , hrs = 0) are consecutive. A control signal Ctl_r (0) is output to the bank memory Tmem_0. Then, according to the control signal Ctl_r (0) , the bank memory Tmem_0 of the memory unit 22 reads four data of D bn0 (2, 0), D bn0 (2, 1), D bn0 (2, 2), D bn0 (2, 3) in parallel (using four access buses) from a region where the addresses of the bank memory Tmem_0 of the memory unit 22 are consecutive.

[0089] (Step S111r_bnk_1): In step S111r_bnk_1, read processing of data for the number of channels (Ch 1 ~Ch 0 ~Ch N-1 )(N = 4) is executed from the bank memory bank (2) The memory access control unit 21 of the CNN data processing unit 2 controls the data D 2 during the period T 2 (time t 3 ~t bn1 (2, 0), D bn1 (2, 1), D bn1 (2, 2), D bn1 (2, 3) are read in parallel (using four access buses (see FIG. 5)) from a region where the addresses of the bank memory Tmem_1 (bank memory bank 1 ) of the memory unit 22 are consecutive. A control signal Ctl_r (1) is output to the bank memory Tmem_1. Then, according to the control signal Ctl_r (1) , the bank memory Tmem_1 of the memory unit 22 reads the data D bn1 (2, 0), Dbn1 (2,1), D bn1 (2,2), D bn1 Four data items at (2,3) (data for 4 channels) are read in parallel (using four access buses) from a region where the addresses of the bank memories Tmem_1 in the memory unit 22 are consecutive.

[0090] (Step S111r_bnk_hre (hre = 2)): In step S111r_bnk_hre, the bank memory bank hre (bank memory Tmem_hre) (hre = 2) reads out data for the number of channels (Ch 0 ~Ch N-1 )(N = 4). Specifically, the following processing is executed. (3) The memory access control unit 21 of the CNN data processing unit 2 controls the period T 2 (time t 2 ~t 3 ) to read four data items D bn2 (2,0), D bn2 (2,1), D bn2 (2,2), D bn2 (2,3) in parallel (using four access buses (see Fig. 5)) from a region where the addresses of the bank memory Tmem_2 in the memory unit 22 are consecutive. A control signal Ctl_r 2 ) is output to the bank memory Tmem_2. Then, in accordance with the control signal Ctl_r (2) , the bank memory Tmem_2 in the memory unit 22 reads four data items D (2) (2,0), D bn2 (2,1), D bn2 (2,2), D bn2 (2,3) in parallel (using four access buses) from a region where the addresses of the bank memory Tmem_2 in the memory unit 22 are consecutive. bn2 Four data items at (2,3) (data for 4 channels) are read in parallel (using four access buses) from a region where the addresses of the bank memory Tmem_2 in the memory unit 22 are consecutive.

[0091] (Step S112r_bnk_hrs(hrs = 0) to S112r_bnk_hre(hre = 2)): In step S112r_bnk_hrs(hrs = 0) to S112r_bnk_hre(hre = 2), the process of determining the number of output systems Num_sys is executed. As shown in FIGS. 6, 16, and 17, during period T 2 the data read out is the data in the third column of block 0 and is the data in the second column of block 1 and is the data in the first column of block 2 Therefore, in each step of step S112r_bnk_hrs(hrs = 0) to S112r_bnk_hre(hre = 2), the number of output systems Num_sys = 3 is determined.

[0092] (Step S113r_bnk_hrs(hrs = 0) to S113r_bnk_hre(hre = 2)): In step S113r_bnk_hrs(hrs = 0) to S113r_bnk_hre(hre = 2), data for Num_sys(=3) sets of output systems are output simultaneously (register write processing is executed). Specifically, the following process is executed.

[0093] The memory access control unit 21 of the CNN data processing unit 2 generates a control signal Ctl_r 2 during period T bnk to read the data D (k) (2,0) of the bank memory Tmem_k, outputs the control signal Ctl_r (k) to the bank memory Tmem_k, and also generates a control signal Ctl_reg to instruct that the data D bnk (2,0) of the bank memory Tmem_k be written into a region of a predetermined address in the register unit 23, and outputs the control signal Ctl_reg to the register unit 23. The bank memory Tmem_k, according to the control signal Ctl_r (k) reads the data D bnkRead (2, 0), and according to the control signal Ctl_reg, the register unit 23 receives the data D output from the bank memory Tmem_k bnk Write (2, 0) to a predetermined address. Assume that the control signal Ctl_reg has instructed the predetermined address

[0094] Period T 2 During this period, for the data of channel 0 (Ch 0 ) and channel k (Ch k )(1 ≤ k ≤ 3), the data output from the bank memory Tmem_k and the write destination address of the register unit 23 for this data are as follows ≪Period T 2 ≫ (Output to 3 systems (Num_sys = 3)) Channel 0 (Ch 0 ): (1) The data D of the bank memory Tmem_0 (bank 0 ) bn0 (2, 0) → Address adr02 of the register unit 23 (Ch0) → Address adr11 of the register unit 23 (Ch0) → Address adr20 of the register unit 23 (Ch0) (2) The data D of the bank memory Tmem_1 (bank 1 ) bn1 (2, 0) → Address adr05 of the register unit 23 (Ch0) → Address adr14 of the register unit 23 (Ch0) → Address adr23 of the register unit 23 (Ch0) (3) The data D of the bank memory Tmem_2 (bank 2 ) bn2 (2, 0) → Address adr08 of the register unit 23 (Ch0) → Address adr17 of the register unit 23 (Ch0) → Address adr26 of the register unit 23(Ch0) Channel k (Ch k )(k is a natural number, 1 ≤ k ≤ 3): (1) The data D 0 of the bank memory Tmem_0 (bank bn0 (2,k) → The address adr02 of the register section 23 (Chk) → The address adr11 of the register section 23 (Chk) → The address adr20 of the register section 23 (Chk) (2) The data D 1 of the bank memory Tmem_1 (bank bn1 (2,k) → The address adr05 of the register section 23 (Chk) → The address adr14 of the register section 23 (Chk) → The address adr23 of the register section 23 (Chk) (3) The data D 2 of the bank memory Tmem_2 (bank bn2 (2,k) → The address adr08 of the register section 23 (Chk) → The address adr17 of the register section 23 (Chk) → The address adr26 of the register section 23 (Chk) Figures 18 to 20 (the part during period T 2 ) show the relationship between the above data (the data of channel 0) and the write destination addresses in the register section 23.

[0095] As shown in Figure 18, during period T 2 , the data D bn0 (2,0), the data D bn1 (2,0), the data D bn2 (2,0) are respectively the address adr02 (Ch0) , the address adr05 (Ch0) (= the address adr02 (Ch0) + 3 addresses), the address adr08(Ch0) (= Address adr05 (Ch0) + 3 addresses) is written in the area. That is, during period T 2 in, data D bn0 (2,0), data D bn1 (2,0), data D bn2 (2,0) is written every 3 addresses (skipping 3 addresses) of the register unit 23 (the thick rectangle in Fig. 18 indicates the data to be written). This is because the area to be subjected to the convolution process (kernel size) is 3×3, so that the 3×3 data can be reshaped into 1×9 data according to the area to be subjected to the convolution process (kernel size) and output to the quantization data memory unit 3.

[0096] Also, as shown in Fig. 19, during period T 2 in, data D bn0 (2,0), data D bn1 (2,0), data D bn2 (2,0) are respectively written in the areas of address adr11 (Ch0) , address adr14 (Ch0) (= Address adr11 (Ch0) + 3 addresses), address adr17 (Ch0) (= Address adr14 (Ch0) + 3 addresses). That is, during period T 2 in, data D bn0 (2,0), data D bn1 (2,0), data D bn2 (2,0) are written every 3 addresses (skipping 3 addresses) of the register unit 23 (the thick rectangle in Fig. 19 indicates the data to be written). This is because the area to be subjected to the convolution process (kernel size) is 3×3, so that the 3×3 data can be reshaped into 1×9 data according to the area to be subjected to the convolution process (kernel size) and output to the quantization data memory unit 3.

[0097] Also, as shown in Fig. 20, during period T 2 in, data D bn0 (2,0), data Dbn1 (2,0), data D bn2 (2,0) is respectively address adr20 (Ch0) , address adr23 (Ch0) (= address adr20 (Ch0) + 3 addresses), address adr26 (Ch0) (= address adr23 (Ch0) + 3 addresses) is written in the area. That is, during period T 2 in, data D bn0 (2,0), data D bn1 (2,0), data D bn2 (2,0) is written every 3 addresses (skipping 3 addresses) of the register unit 23 (the thick rectangle in Fig. 20 indicates the data to be written). This is because the area to be subjected to the convolution process (the size of the kernel) is 3×3. According to the area to be subjected to the convolution process (the size of the kernel), the 3×3 data is reshaped into 1×9 data so that it can be output to the quantization data memory unit 3. Note that the addresses adr20 (Chk) ~adr28 (Chk) (k: natural number, 0≦k≦N - 1) are assumed to be consecutive addresses.

[0098] (Step S12r): In step S12r, a determination process is executed to determine whether a predetermined amount of data has been output from the bank memory Tmem_k of the memory unit 22 to the register unit 23. At the time when period T 2 ends, as shown in Fig. 18, for block 0 , since all the data in the area to be subjected to the convolution process (the size of the kernel) has been output from the bank memory Tmem_k of the memory unit 22 to the register unit 23 (refer to the part of period T 2 in Fig. 18), the process proceeds to step S13r.

[0099] (Step S13r): In step S13r, a register output process is executed. Specifically, for block 0Regarding this, since all the data in the area to be subjected to the convolution process (the size of the kernel) has been output from the bank memory Tmem_k of the memory unit 22 to the register unit 23, the register unit 23 outputs the data including the following data as data D2 to the quantization data memory unit 3. ≪Feature amount data (quantized data) (block 0 )(Ch 0 )≫(Period T 2 ) Data D bn0 (0,0), Data D bn0 (1,0), Data D bn0 (2,0) Data D bn1 (0,0), Data D bn1 (1,0), Data D bn1 (2,0) Data D bn2 (0,0), Data D bn2 (1,0), Data D bn2 (2,0) (The above data is stored in the area where the addresses of the register unit 23 are consecutive (adr00 (Ch0) ~adr08 (Ch0) ). (See Figure 21)) ≪Feature amount data (quantized data) (block 0 )(Ch 1 )≫(Period T 2 ) Data D bn0 (0,1), Data D bn0 (1,1), Data D bn0 (2,1) Data D bn1 (0,1), Data D bn1 (1,1), Data D bn1 (2,1) Data D bn2 (0,1), Data D bn2 (1,1), Data D bn2 (2,1) (The above data is stored in the area where the addresses of the register unit 23 are consecutive (adr00 (Ch1) ~adr08 (Ch1) ).) ≪Feature amount data (quantized data) (block0 )(Ch 2 )>>(Period T 2 ) Data D bn0 (0,2), Data D bn0 (1,2), Data D bn0 (2,2) Data D bn1 (0,2), Data D bn1 (1,2), Data D bn1 (2,2) Data D bn2 (0,2), Data D bn2 (1,2), Data D bn2 (2,2) (The above data is stored in a continuous area (adr00 (Ch2) ~adr08 (Ch2) ) of the register section 23.) <<Feature data (quantized data) (block 0 )(Ch 3 )>>(Period T 2 ) Data D bn0 (0,3), Data D bn0 (1,3), Data D bn0 (2,3) Data D bn1 (0,3), Data D bn1 (1,3), Data D bn1 (2,3) Data D bn2 (0,3), Data D bn2 (1,3), Data D bn2 (2,3) (The above data is stored in a continuous area (adr00 (Ch3) ~adr08 (Ch3) ) of the register section 23.)

[0100] (Step S2r): In step S2r, it is determined whether there is any data to be processed in the data reading process by the CNN data processing unit 2. If there is still data to be processed, the process returns to S11r and the same process as above is executed. On the other hand, if there is no data to be processed, the data reading process by the CNN data processing unit 2 is terminated.

[0101] When there is still data to be processed and during period T 3 Regarding the process of, in the CNN data processing unit 2, the same process as the above period T 2 is executed.

[0102] Period T 3 In the process of, in step S13r, the register output process is executed, and for block 1 since all the data in the area (kernel size) to be subjected to the convolution process has been output from the bank memory Tmem_k of the memory unit 22 to the register unit 23, the register unit 23 outputs the data including the following data as data D2 to the quantization data memory unit 3. ≪Feature data (quantization data) (block 1 )(Ch 0 )≫ (during period T 3 ) Data D bn0 (1,0), data D bn0 (2,0), data D bn0 (3,0) Data D bn1 (1,0), data D bn1 (2,0), data D bn1 (3,0) Data D bn2 (1,0), data D bn2 (2,0), data D bn2 (3,0) (The above data is stored in the area where the addresses of the register unit 23 are continuous (adr10 (Ch0) ~adr18 (Ch0) ). (See Fig. 21)) ≪Feature data (quantization data) (block 1 )(Ch 1 )≫ (during period T3 ) Data D bn0 (1,1), Data D bn0 (2,1), Data D bn0 (3,1) Data D bn1 (1,1), Data D bn1 (2,1), Data D bn1 (3,1) Data D bn2 (1,1), Data D bn2 (2,1), Data D bn2 (3,1) (The above data is stored in a region where the addresses of the register unit 23 are consecutive (adr10 (Ch1) ~adr18 (Ch1) ).) ≪Feature amount data (quantized data) (block 1 )(Ch 2 )≫ (Period T 3 ) Data D bn0 (1,2), Data D bn0 (2,2), Data D bn0 (3,2) Data D bn1 (1,2), Data D bn1 (2,2), Data D bn1 (3,2) Data D bn2 (1,2), Data D bn2 (2,2), Data D bn2 (3,2) (The above data is stored in a region where the addresses of the register unit 23 are consecutive (adr10 (Ch2) ~adr18 (Ch2) ).) ≪Feature amount data (quantized data) (block 1 )(Ch 3 )≫ (Period T 3 ) Data D bn0 (1,3), Data D bn0 (2,3), Data D bn0 (3,3) Data D bn1 (1,3), Data D bn1 (2,3), Data D bn1(3,3) Data D bn2 (1,3), Data D bn2 (2,3), Data D bn2 (3,3) (The above data is stored in a region where the addresses of the register unit 23 are consecutive (adr10 (Ch3) ~adr18 (Ch3) ).) In the data processing unit 2 for CNN, for the processing after the period T 4 the same process is executed. When there is no data to be processed left, the process is terminated.

[0103] The memory unit 3 for quantized data inputs the data D2 output from the register unit 23 of the data processing unit 2 for CNN and stores the data D2. Since the data D2 output from the register unit 23 of the data processing unit 2 for CNN is data obtained by reshaping 3×3 data into 1×9 data according to the region to be subjected to the convolution process (the size of the kernel (in this embodiment, 3×3)), the memory unit 3 for quantized data stores the data D2 in, for example, a region of consecutive addresses.

[0104] The convolution processing unit 4 reads out the data in the region to be subjected to the convolution process by the input weight coefficient data Din_w (weight filter (kernel)) from the memory unit 3 for quantized data. Then, the convolution processing unit 4 performs a convolution process (convolution operation) on the data read out from the memory unit 3 for quantized data using the weight coefficient data Din_w (in this embodiment, a 3×3 kernel), obtains the data after the convolution process, and outputs the obtained data as the data Dout.

[0105] <<Summary>> As described above, in the CNN data processing apparatus 100, the CNN data processing unit 2 can execute in parallel the data writing process to the memory unit 22 of the data (data after quantization processing of feature amount data) output from the quantization processing unit 1 and the data reading process from the memory unit 22. Further, the memory unit 22 has a plurality of bank memories Tmem_k and can write and / or read a plurality of data simultaneously (in parallel). Therefore, in the CNN data processing apparatus 100, high-speed data writing and data reading processes can be realized. And in the CNN data processing apparatus 100, (1) a plurality of access buses are provided for each of the plurality of bank memories Tmem_k of the memory unit 22, and data for a plurality of channels can be accessed simultaneously (in parallel), and (2) different (independent) bank memories Tmem_k are assigned for each height direction of the convolution processing target region (region to be convolved with the kernel), so that data in a plurality of different height directions can be accessed simultaneously (in parallel). For this reason, in the CNN data processing apparatus 100, during the period of one data reading process (period T i ), data of h×1 (h rows and 1 column, h: position in the height direction) of the convolution processing target region can be read for a plurality of channels.

[0106] Furthermore, in the CNN data processing apparatus 100, according to the position (shifted position) of the convolution processing target region, the number of output systems Num_sys, which is the number of sets of overlapping data (sets of h×1 data of the convolution processing target region), is obtained, and the sets of overlapping data (sets of h×1 data of the convolution processing target region) corresponding to the obtained number of output systems Num_sys are output to the register unit 23 in different systems (in parallel).

[0107] Thereby, in the CNN data processing apparatus 100, the number of times of reading overlapping data can be reduced by sliding the position of the convolution processing target region.

[0108] Also, the CNN data processing device 100 includes a register unit 23. In the register unit 23, according to the size (shape) of the convolution processing target area (the size (shape) of the kernel), the data read from the memory unit 22 is written to discontinuous addresses (addresses obtained by adding a predetermined offset value (corresponding to the size in the width direction of the kernel; for a 3×3 kernel, the offset value is "3")). After all the data (quantized data of feature amount data) in the convolution processing target area is aligned (after all the data in the convolution processing target area is written at consecutive addresses in the register unit 23), all the data in the convolution processing target area is output to the quantized data memory unit 3.

[0109] As a result, in the CNN data processing device 100, all the data (data to be subjected to convolution processing) in the convolution processing target area is used as data arranged in the order of performing the convolution operation. It can be output and written to the quantized data memory unit 3. Then, the data arranged in the order of performing the convolution operation is read from the quantized data memory unit 3, and the convolution processing unit 4 performs convolution processing using the kernel weight coefficient data to be applied to the data, so that the convolution processing can be executed at high speed.

[0110] In this way, in the CNN data processing device 100, by simply providing the CNN data processing unit 2, it is possible to reduce the number of times of reading duplicate data and obtain data arranged in the order of performing the convolution operation. Therefore, the CNN data processing device 100 can reduce the number of times of executing the process of reading the feature amount data and shorten the time required for the entire convolution processing including the reading process of the feature amount data, and perform data processing for realizing a high-performance and high-speed CNN model.

[0111] [Other Embodiments] In the above embodiment, in the CNN data processing device 100, the case where the weight coefficient data Din_w (weight filter (kernel)) is input to the convolution processing unit 4 and the convolution processing is executed using the weight coefficient data Din_w (weight filter (kernel)) has been described. However, the present invention is not limited to this. For example, the convolution processing unit 4 may execute vector decomposition processing on the weight filter (kernel), decompose it into a base matrix and a real coefficient vector, and input the decomposed base matrix and real coefficient vector to perform convolution processing. In this case, the convolution processing is executed using the base matrix (a matrix whose elements are only base values (integer values)) and the data D3 output from the quantization data memory unit 3, and then the processing by the real coefficient vector is performed, so that most of the convolution operations can be integer operations, and furthermore, the acceleration of the convolution processing can be realized.

[0112] Also, in the above embodiment, in the CNN data processing device 100, the case where the convolution processing target region is a region of size 4×8 using a kernel of a predetermined size (3×3) has been described. However, the present invention is not limited to these, and the size of the kernel and the size of the convolution processing target region may be other sizes.

[0113] Also, in the above embodiment, in the CNN data processing device 100, the case where Depthwise convolution (spatial direction convolution processing) is described assuming the execution of CNN data processing has been described. However, the present invention is not limited to this. In the CNN data processing device 100, for example, the CNN data processing of the above embodiment may be applied to normal convolution processing.

[0114] In the above-described embodiment, the case where CNN data processing is performed on the data after quantization processing in the CNN data processing apparatus 100 has been described. However, the present invention is not limited to this. For example, data (feature amount data) on which quantization processing has not been performed may be input to the CNN data processing unit 2 of the CNN data processing apparatus 100, and CNN data processing by the CNN data processing unit 2 may be executed on the data.

[0115] Further, the configuration of the memory unit 22 of the CNN data processing apparatus 100 is not limited to that described in the above embodiment, and the number of bank memories and the number of data that can be accessed simultaneously in each bank memory (the number of access buses) can be set to any number.

[0116] Each block (each functional unit) of the CNN data processing apparatus 100 described in the above embodiment may be individually integrated into one chip by a semiconductor device such as an LSI, or may be integrated into one chip so as to include part or all of them. Further, each block (each functional unit) of the pose data generation system, CG data system, and pose data generation apparatus described in the above embodiment may be realized by a plurality of semiconductor devices such as LSIs.

[0117] Here, although an LSI has been used, depending on the degree of integration, it may also be referred to as an IC, system LSI, super LSI, or ultra LSI.

[0118] Further, the method of integrating the circuit is not limited to an LSI, and it may be realized by a dedicated circuit or a general-purpose processor. After manufacturing the LSI, an FPGA (Field Programmable Gate Array) that can be programmed or a reconfigurable processor that can reconfigure the connection and setting of circuit cells inside the LSI may be used.

[0119] Also, part or all of the processing of each functional block in each of the above embodiments may be realized by a program. And part or all of the processing of each functional block in each of the above embodiments is performed by a central processing unit (CPU) in a computer. Further, the program for performing each processing is stored in a storage device such as a hard disk or a ROM, and is read from the ROM or the RAM and executed.

[0120] Also, each processing in the above embodiments may be realized by hardware, or may be realized by software (including the case where it is realized together with an OS (operating system), middleware, or a predetermined library). Further, it may be realized by a mixed processing of software and hardware.

[0121] For example, when each functional unit in the above embodiments is realized by software, a hardware configuration shown in FIG. 22 (for example, a hardware configuration in which a CPU, a GPU, a ROM, a RAM, an input unit, an output unit, etc. are connected by a bus Bus) may be used to realize each functional unit by software processing.

[0122] Also, when each functional unit in the above embodiments is realized by software, the software may be realized using a single computer having the hardware configuration shown in FIG. 22, or may be realized by distributed processing using a plurality of computers.

[0123] Also, the execution order of the processing method in the above embodiments is not necessarily limited to the description of the above embodiments, and the execution order can be changed without departing from the gist of the invention. Also, in the processing method in the above embodiments, some steps may be executed in parallel with other steps without departing from the gist of the invention. Also, in the processing method in the above embodiments, the processing executed in parallel may be executed serially (sequentially).

[0124] A computer program that causes a computer to execute the above-described method and a computer-readable recording medium storing the program are included in the scope of the present invention. Here, examples of the computer-readable recording medium include a flexible disk, a hard disk, a CD-ROM, an MO, a DVD, a DVD-ROM, a DVD-RAM, a large-capacity DVD, a next-generation DVD, and a semiconductor memory.

[0125] The above computer program is not limited to that recorded on the above recording medium, and may be transmitted via an electric communication line, a wireless or wired communication line, a network typified by the Internet, or the like.

[0126] Also, the term "section" may be a concept including "circuitry". Circuitry may be realized in whole or in part by hardware, software, or a combination of hardware and software.

[0127] The functions of the elements disclosed herein may be implemented using a general-purpose processor, a dedicated processor, an integrated circuit, an ASIC (Application Specific Integrated Circuit), a conventional circuit configuration, and / or a combination thereof that is configured to execute the disclosed elements or programmed to execute the disclosed functions. A processor is considered a processing circuit configuration or a circuit configuration when it includes transistors and other circuit configurations therein. In the present disclosure, a circuit configuration, unit, or means is hardware that executes the recited functions or hardware programmed to execute the functions. The hardware may be any hardware disclosed herein or other known hardware that is programmed to execute the recited functions or configured to execute the functions. When the hardware is a processor that may be considered a type of circuit configuration, the circuit configuration, means, or unit is a combination of hardware and software, software used to configure the hardware, and / or a processor.

[0128] Note that the specific configuration of the present invention is not limited to the foregoing embodiments, and various changes and modifications are possible without departing from the gist of the invention.

Explanation of Reference Numerals

[0129] 100 Data processing device for CNN 2 Data processing unit for CNN (data device for convolution processing) 21 Memory access control unit (access control unit) 22 Memory unit Tmem_0~Tmem_M-1 Memory Tmem_k for banks 23 Register unit

Claims

1. A data processing device for convolution processing used in a convolutional neural network model, a plurality of bank memories for storing feature data, an access control unit for performing data writing and / or data reading control of the plurality of bank memories, comprising: the feature data is three-dimensional data specified by a position in the width direction, a position in the height direction, and a position in the channel direction, each of the plurality of bank memories has a plurality of access buses so as to be able to access data in parallel, the access control unit performs data writing control so that the feature data whose position in the height direction is a first value is stored in the bank memory assigned to the first value among the plurality of bank memories, and further, a plurality of the feature data having the same position in the width direction and consecutive positions in the channel direction are stored in a memory area of an address that can be accessed in parallel by the plurality of buses, A data processing device for convolution processing.

2. The access control unit, in a read unit period, performs data reading control on the plurality of bank memories so as to read the feature data having the same position in the width direction and consecutive positions in the height direction for a plurality of channels, The data processing device for convolution processing according to claim 1.

3. When a data group obtained by reading the feature data having the same position in the width direction and consecutive positions in the height direction from the plurality of bank memories for a plurality of channels is a plurality of channel h×1 data groups, The access control unit, according to the position of the convolution processing target area, in the same read unit period, acquires the number of overlapping plurality of channel h×1 data groups as the number of output systems, and controls the plurality of bank memories so that the plurality of channel h×1 data groups for the acquired number of output systems are output from the plurality of bank memories, The data processing device for convolution processing according to claim 1 or 2.

4. further comprising a register unit capable of storing data by address designation, the register unit, Input the plurality of channel h×1 data groups output from the plurality of bank memories, use the size in the width direction of the kernel of the convolution process applied to the plurality of channel h×1 data groups as an offset value, and sequentially write the feature amount data with continuous positions in the height direction included in the plurality of channel h×1 data groups into the memory area of the address of the register unit offset by the offset value. The data processing device for convolution processing according to claim 3.

5. The register unit After the feature amount data of the area to be subjected to the convolution process by the kernel is stored in the memory area of the continuous addresses of the register unit, output the feature amount data stored in the memory area of the continuous addresses. The data processing device for convolution processing according to claim 4.

6. The register unit Output the feature amount data stored in the memory area of the continuous addresses all at once or in the order of the continuous addresses. The data processing device for convolution processing according to claim 5.