Data processing device for convolution processing

The data processing apparatus for convolution processing addresses the bottleneck of feature data readout in conventional CNN technologies by utilizing multiple bank memories and an access control unit for parallel data access, resulting in a high-performance and high-speed CNN model with reduced data read times.

WO2025115418A1PCT designated stage expired Publication Date: 2025-06-05MEGACHIPS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/036254
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-01
Filing Date
2024-10-10
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Conventional technologies, such as MobileNet, face a bottleneck in the time required for reading out feature data during convolutional neural network (CNN) processing, which hinders the overall speed of the CNN process despite accelerated convolution processing.

Method used

A data processing apparatus for convolution processing is designed with a plurality of bank memories and an access control unit. This apparatus allows for parallel data access through multiple access buses and assigns independent bank memories to each height direction, enabling simultaneous access to data with different height directions. The access control unit optimizes data read and write operations to reduce the number of data readouts and improve processing efficiency.

Benefits of technology

The proposed solution significantly reduces the number of times feature data needs to be read, thereby shortening the total time required for the convolution process. This leads to a high-performance and high-speed CNN model capable of efficiently processing data, including the data read process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024036254_05062025_PF_FP_ABST
    Figure JP2024036254_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention realizes a data processing device for convolution processing capable of reducing the number of executions of processing for reading feature quantity data, shortening the time required for the entire convolution processing including the processing for reading the feature quantity data, and performing data processing for realizing a high-performance and high-speed CNN model. [Solution] In the data processing device for convolution processing, (1) a plurality of access buses are provided in each of a plurality of bank memories (Tmem_k) of a memory unit (22), and data for a plurality of channels can be accessed simultaneously (in parallel); and (2) since different (independent) bank memories Tmem_k are allocated for each height direction of a convolution processing target region (region to be convolved with a kernel), a plurality of data items having different height directions can be accessed simultaneously (in parallel).
Need to check novelty before this filing date? Find Prior Art

Description

Convolution processing data processing device

[0001] The present invention relates to a data processing technology for a convolutional neural network, and more particularly to a technology for processing feature data used in a convolutional neural network (a data processing device for convolutional processing).

[0002] In recent years, technologies using neural network models that realize a wide variety of applications with high accuracy have been attracting attention. In technologies using neural network models, training data is used to perform a training process for the neural network model, a trained model is obtained, and the acquired trained model is used to perform a prediction process (inference process). As a result, technologies using neural network models can realize a wide variety of applications with high accuracy. Technologies using convolutional neural networks (CNNs) have been attracting attention as technologies using neural network models that have achieved high value in fields such as image recognition.

[0003] Furthermore, lightweight technologies have been developed to enable the use of convolutional neural network models on mobile devices that do not have abundant computational resources, such as a technology called Mobilenet (a lightweight technology for CNN models) (see, for example, Non-Patent Document 1).

[0004] A technology called Mobilenet (a technology for reducing the weight of CNN models) reduces the number of parameters in CNN models by adopting a technique called depthwise separable convolution, which divides normal convolution processing into two processes: (1) depthwise convolution (convolution processing in the spatial direction) and (2) pointwise convolution (convolution processing in the channel direction). This makes it possible to realize a lightweight, high-performance CNN model that can be installed on mobile devices that do not have abundant computing resources.

[0005] Howard, Andrew G., et al. "MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications." arXiv preprint arXiv:1704.04861 (2017).

[0006] However, in the above-described conventional technology (Mobilenet), it is necessary to frequently read out feature data, and the time required for the feature data readout process is longer than the time required for the product-sum calculation for the convolution process, resulting in a longer time required for the CNN process. In other words, in the above-described conventional technology (Mobilenet), even if the processing speed of the CNN calculation itself (convolution process) is increased, the time required for the feature data readout process becomes a bottleneck (critical path), making it difficult to shorten the time required for the total processing (CNN process) including the feature data readout process.

[0007] In view of the above-mentioned problems, an object of the present invention is to provide a data processing device for convolution processing that can reduce the number of times feature data reading processing is executed, shorten the time required for the entire convolution processing including the feature data reading processing, and perform data processing to realize a high-performance, high-speed CNN model.

[0008] In order to solve the above problems, a first invention is a data processing device for convolution processing used in a convolutional neural network model, comprising a plurality of bank memories for storing feature data, and an access control unit for controlling data writing and / or data reading from the plurality of bank memories.

[0009] The feature amount data is three-dimensional data specified by the position in the width direction, the position in the height direction, and the position in the channel direction.

[0010] Each of the plurality of bank memories has a plurality of access buses so that data can be accessed in parallel.

[0011] The access control unit controls data writing so that feature data having a first value for height position is stored in a bank memory assigned to the first value among a plurality of bank memories, and further stores a plurality of feature data having the same width position and consecutive channel position in a memory area with addresses that can be accessed in parallel via a plurality of buses.

[0012] In this data processing device for convolution processing, (1) each of the multiple bank memories Tmem_k of the memory unit 22 is provided with multiple access buses, allowing simultaneous (parallel) access to data for multiple channels, and (2) different (independent) bank memories Tmem_k are assigned to each height direction of the convolution processing target area (the area to be convolved with the kernel), allowing simultaneous (parallel) access to multiple data in different height directions.

[0013] The "feature amount data" may be feature amount data after quantization processing.

[0014] The second invention is the first invention, wherein the access control unit controls data reading from multiple bank memories so that feature data having the same widthwise position and continuous heightwise positions is read for multiple channels during a read unit period.

[0015] As a result, this convolution processing data device can read out data for multiple channels of h × 1 (h rows, 1 column, h: height position) from the convolution processing target area during one data read processing period (read unit period).

[0016] A third invention is the first or second invention, in which a data group obtained by reading out feature data for multiple channels that have the same widthwise position and continuous heightwise positions from multiple bank memories is defined as a multiple-channel h×1 data group, and the access control unit obtains the number of overlapping multiple-channel h×1 data groups in the same read unit period as the number of output systems depending on the position of the convolution processing target area, and controls the multiple bank memories so that multiple-channel h×1 data groups equal to the number of output systems obtained are output from the multiple bank memories.

[0017] This convolution processing data device acquires the number of output systems Num_sys, which is the number of overlapping data sets (h × 1 data sets in the convolution processing target area), depending on the position of the convolution processing target area (slid position), and can output the overlapping data sets (h × 1 data sets in the convolution processing target area) equal to the acquired number of output systems Num_sys, each in separate systems (in parallel).

[0018] As a result, in this convolution processing data device, the number of times overlapping data is read can be reduced by sliding the position of the convolution processing target region.

[0019] A fourth aspect of the present invention is the third aspect of the present invention, further comprising a register section capable of storing data by address designation.

[0020] The register unit inputs a multi-channel h x 1 data group output from multiple bank memories, uses the widthwise size of the kernel of the convolution processing to be performed on the multi-channel h x 1 data group as an offset value, and sequentially writes feature data with consecutive heightwise positions contained in the multi-channel h x 1 data group into the memory area of ​​the register unit at an address offset by the offset value.

[0021] This convolution processing data device includes a register unit, in which data read from the memory unit 22 is written to discrete addresses (addresses to which a predetermined offset value (corresponding to the size of the kernel in the width direction (for a 3 × 3 kernel, the offset value is "3")) is added) according to the size (shape) of the region to be convolution processing (size (shape) of the kernel), and once all of the data (feature amount data) for the region to be convolution processing has been collected (after all of the data for the region to be convolution processing has been written at consecutive addresses in the register unit), all of the data for the region to be convolution processing can be output.

[0022] As a result, this convolution processing data device can output all data in the convolution processing target area (data to be convolved) as data arranged in the order in which the convolution operation is performed, and write the data to, for example, a quantized data memory unit.Then, by reading the data arranged in the order in which the convolution operation is performed from, for example, the quantized data memory unit, and performing convolution processing using kernel weighting coefficient data to be multiplied by the data, it is possible to perform convolution processing at high speed.

[0023] In this way, in this convolution processing data device, simply by providing a functional unit that performs the above processing, it is possible to acquire data arranged in the order in which the convolution operation is to be performed while reducing the number of times that duplicate data is read. Therefore, this convolution processing data device can perform data processing to realize a high-performance, high-speed CNN model, which can reduce the number of times that the feature data read process is executed and shorten the time required for the entire convolution process, including the feature data read process.

[0024] A fifth aspect of the invention is the fourth aspect of the invention, wherein the register unit outputs the feature data stored in the memory area with consecutive addresses after the feature data of the region to be convolved by the kernel has been stored in the memory area with consecutive addresses in the register unit. In this convolution data device, after the feature data of the region to be convolved by the kernel has been stored in the memory area with consecutive addresses in the register unit, that is, after the feature data stored with the offset value has been stored in a continuous state rather than a discrete state (a state in which the feature data is stored in the memory area with consecutive addresses in the register unit), the feature data stored in the memory area with consecutive addresses is output. Therefore, in this convolution data device, it is possible to ensure that the feature data of the region to be convolved by the kernel is output from the register unit after the feature data of the region to be convolved by the kernel has been collected.

[0025] A sixth aspect of the present invention is the fifth aspect of the present invention, wherein the register unit outputs the feature amount data stored in the memory area of ​​consecutive addresses all at once or in the order of consecutive addresses in the register unit.

[0026] As a result, in this convolution processing data device, it is guaranteed that feature data of an area that is the target of convolution processing by a kernel (plurality of feature data that are the target of convolution processing by a kernel (for example, if the kernel is a 3 × 3 kernel, nine feature data included in the area that is the target of convolution processing by the kernel) are output all at once or in a state arranged in the order in which convolution processing by the kernel is performed (the order in which product-sum calculations of kernel weight coefficient data are performed).

[0027] According to the present invention, it is possible to realize a data processing device for convolution processing that can reduce the number of times the process of reading feature data is executed and shorten the time required for the entire convolution process including the process of reading feature data, and that can perform data processing to realize a high-performance, high-speed CNN model.

[0028] FIG. 1 is a schematic configuration diagram of a CNN data processing device 100 according to a first embodiment. FIG. 2 is a schematic configuration diagram of a CNN data processing unit 2 of the CNN data processing device 100 according to the first embodiment. FIG. 3 is a diagram for explaining CNN processing (convolution processing for a CNN model) using (1) depthwise convolution (convolution processing in the spatial direction) and (2) pointwise convolution (convolution processing in the channel direction). FIG. 4 is a diagram for explaining data stored in each bank memory (when there are four bank memories (as an example)) of the CNN data processing unit 2 of the CNN data processing device 100. FIG. 5 is a diagram for explaining data access to the bank memory of the CNN data processing unit 2 of the CNN data processing device 100. FIG. 6 is a diagram for explaining the relationship between read data and blocks when performing data read processing of the CNN data processing unit 2 of the CNN data processing device 100. FIG. 7 is a diagram for explaining the relationship between read data and blocks when performing data read processing of the CNN data processing unit 2 of the CNN data processing device 100. 1 is a flowchart of CNN data processing executed by the CNN data processing device 100. FIG. 2 is a flowchart of CNN data processing executed by the CNN data processing device 100. FIG. 3 is a timing chart of data write processing and data read processing of the CNN data processing unit 2 of the CNN data processing device 100. FIG. 4 is a diagram for explaining data access to the bank memory of the CNN data processing unit 2 of the CNN data processing device 100. FIG. 5 is a diagram for explaining data read processing from the bank memory of the CNN data processing unit 2 of the CNN data processing device 100. FIG. 6 is a diagram for explaining data read processing from the bank memory of the CNN data processing unit 2 of the CNN data processing device 100. FIG. 7 is a diagram for explaining data read processing from the bank memory of the CNN data processing unit 2 of the CNN data processing device 100. 10 is a timing chart of a data read process of the CNN data processing unit 2 of the CNN data processing device 100.A diagram for explaining the relationship between read data and data write addresses of the block and register unit 23 when performing data read processing of the CNN data processing unit 2 of the CNN data processing device 100. A diagram for explaining the relationship between read data and data write addresses of the block and register unit 23 when performing data read processing of the CNN data processing unit 2 of the CNN data processing device 100. A diagram for explaining the relationship between read data and data write addresses of the block and register unit 23 when performing data read processing of the CNN data processing unit 2 of the CNN data processing device 100. A diagram for explaining the relationship between read data and data write addresses of the block and register unit 23 when performing data read processing of the CNN data processing unit 2 of the CNN data processing device 100 (including the relationship for channels 0 to 4). A diagram showing a CPU bus configuration.

[0029] First Embodiment A first embodiment will be described below with reference to the drawings.

[0030] 1.1: Configuration of CNN Data Processing Device FIG. 1 is a schematic diagram of a CNN data processing device 100 according to the first embodiment.

[0031] FIG. 2 is a schematic configuration diagram of the CNN data processing unit 2 of the CNN data processing device 100 according to the first embodiment.

[0032] 1, the CNN data processing device 100 includes a quantization processing unit 1, a CNN data processing unit 2 (convolution processing data device), a quantized data memory unit 3, and a convolution processing unit 4. The CNN data processing device 100 receives input of feature amount data Din_f and weighting coefficient data Din_w (weight filter (kernel)), executes convolution processing (convolution processing using the feature amount data and weighting coefficient data), and acquires (outputs) processing result data Dout of the convolution processing.

[0033] The quantization processing unit 1 receives the feature data Din_f, performs quantization processing on the feature data Din_f, and outputs the quantized data to the CNN data processing unit 2 as data D1.

[0034] As shown in FIG. 2, the CNN data processing unit 2 includes a memory access control unit 21, a memory unit 22 including M (M: natural number) bank memories (Tmem_0 to Tmem_M-1), and a register unit 23.

[0035] The memory access control unit 21 is a control unit for performing access control (data write process control, data read process control) for the M bank memories of the memory unit 22. The memory access control unit 21 is a functional unit for independently (in parallel) performing data write process control and data read process control for the M bank memories Tmem_0 to Tmem_M-1 of the memory unit 22. The memory access control unit 21 outputs a control signal Ctl_w for performing data write process control and / or a control signal Ctl_r for performing data read process control to the memory unit 22. Specifically, the memory access control unit 21 outputs a control signal Ctl_w for performing data write process control to the bank memories Tmem_k (k: natural number, 0≦k≦M−1) of the memory unit 22. (k) and / or a control signal Ctl_r for controlling the data read process. (k) The control signal Ctl_w for controlling the data write process to the bank memory Tmem_k (k: natural number, 0≦k≦M−1) of the memory unit 22 is output as the control signal Ctl_w (k) and the control signal Ctl_r for controlling the data read process for the bank memory Tmem_k of the memory unit 22 is represented as the control signal Ctl_r (k) It is written as follows.

[0036] The memory access control unit 21 is a control unit for performing access control (data write process control, data read process control) for the register unit 23. The memory access control unit 21 outputs a control signal Ctl_reg to the register unit 23 for performing access control of the register unit 23.

[0037] As shown in FIG. 2, the memory unit 22 includes M (M: natural number) bank memories Tmem_0 to Tmem_M-1.

[0038] The bank memory Tmem_k (k: natural number, 0≦k≦M−1) is a memory that can write predetermined data to a predetermined address of the bank memory Tmem_k and read the data stored in the predetermined address from the address of the bank memory Tmem_k. The bank memory Tmem_k receives a control signal Ctl_w for data write processing from the memory access control unit 21. (k) The data D1 output from the quantization processing unit 1 is converted into the control signal Ctl_w (k) The bank memory Tmem_k is written to the address specified by the control signal Ctl_r for data read processing from the memory access control unit 21. (k) According to the control signal Ctl_w (k) The data stored in the address of the bank memory Tmem_k specified by is read from the address, and the read data is output to the register unit 23 .

[0039] The register unit 23 has a memory (register) that can write data to a predetermined area by specifying an address, and can read data stored in the predetermined area by specifying an address. The register unit 23 receives data output from the memory unit 22 and a control signal Ctl_reg output from the memory access control unit 21. In accordance with the control signal Ctl_reg, the register unit 23 writes the data output from the memory unit 22 to a predetermined address in the register unit 23. In addition, in accordance with the control signal Ctl_reg, the register unit 23 outputs the data at the predetermined address in the register unit 23 to the quantized data memory unit 3 as data D2.

[0040] The quantized data memory unit 3 has a memory capable of storing data, and by specifying an address, data can be written to a predetermined area of ​​the memory, and data stored in a predetermined area can be read out by specifying an address. The quantized data memory unit 3 inputs data D2 output from the CNN data processing unit 2 and stores the data D2. The quantized data memory unit 3 also outputs the stored data to the convolution processing unit 4 as data D3 (the quantized data memory unit 3 inputs a data read command from a control unit (not shown) or the convolution processing unit 4, and outputs data at a predetermined address to the convolution processing unit 4 as data D3 in accordance with the data read command).

[0041] The convolution processing unit 4 receives the weighting coefficient data Din_w (weighting filter (kernel)) and the data D3 output from the quantized data memory unit 3. The convolution processing unit 4 executes convolution processing using the data D3 and the weighting coefficient data Din_w, and outputs the data after the convolution processing as data Dout.

[0042] <1.2: Operation of CNN Data Processing Device> The operation of the CNN data processing device 100 configured as above will be described below.

[0043] FIG. 3 is a diagram for explaining CNN processing (convolution processing for a CNN model) using (1) depthwise convolution (convolution processing in the spatial direction) and (2) pointwise convolution (convolution processing in the channel direction).

[0044] FIG. 4 is a diagram for explaining data stored in each bank memory (for example, when there are four bank memories) of the CNN data processing unit 2 of the CNN data processing device 100.

[0045] FIG. 5 is a diagram for explaining data access to the bank memory of the CNN data processing unit 2 of the CNN data processing device 100.

[0046] 6 and 7 are diagrams for explaining the relationship between read data and blocks when the CNN data processing unit 2 of the CNN data processing device 100 performs data read processing.

[0047] 8 to 10 are flowcharts of the CNN data processing executed by the CNN data processing device 100.

[0048] FIG. 11 is a timing chart of the data write process and data read process of the CNN data processing unit 2 of the CNN data processing device 100.

[0049] FIG. 12 is a diagram for explaining data access to the bank memory of the CNN data processing unit 2 of the CNN data processing device 100.

[0050] 13 to 15 are diagrams for explaining the data read process from the bank memory of the CNN data processing unit 2 of the CNN data processing device 100.

[0051] 16 and 17 are timing charts of the data read process of the CNN data processing unit 2 of the CNN data processing device 100.

[0052] 18 to 20 are diagrams for explaining the relationship between read data, blocks, and data write addresses of the register unit 23 when performing data read processing in the CNN data processing unit 2 of the CNN data processing device 100.

[0053] FIG. 21 is a diagram for explaining the relationship between the read data, blocks, and data write addresses of the register unit 23 when performing data read processing in the CNN data processing unit 2 of the CNN data processing device 100 (including the relationship for channels 0 to 4).

[0054] As shown in FIG. 3, when a method of dividing normal convolution processing into two, (1) Depthwise convolution (convolution processing in the spatial direction) and (2) Pointwise convolution (convolution processing in the channel direction), is adopted, in Depthwise convolution (convolution processing in the spatial direction), the time required for the feature data read processing is longer than the time required for the product-sum calculation for the convolution processing, and as a result, the time required for executing the CNN processing is longer. To address this, in the CNN data processing device 100, the CNN data processing unit 2 performs the following steps: (A) N (N: natural number equal to or greater than 2) channels (Ch 0 ~Ch N-1 ) are written in parallel, and (B) N (N: a natural number equal to or greater than 2) channels (Ch 0 ~Ch N-1 Furthermore, in the CNN data processing device 100, the CNN data processing unit 2 simultaneously reads data for N channels from each of the M bank memories Tmem_0 to Tmem_M-1, and simultaneously outputs data for the number of systems corresponding to the kernel of the CNN processing, by parallel processing.

[0055] For ease of explanation, the operation of the CNN data processing device 100 will be described below in the following case (one example). Note that the settings for the CNN data processing device 100 are not limited to the following settings, and other settings may be used. (1) The number of channels N of the feature data to be subjected to convolution processing (feature data input to the CNN data processing device 100) is "4" (N=4, channel Ch 0 ~Ch 3 (2) The number of bank memories in the memory unit 22 of the CNN data processing unit 2 is "4" (M=4, bank memories Tmem_0 to Tmem_3 (bank 0 ~bank 3 (3) The bank memory Tmem_k(bank k) (k: natural number, 0≦k≦M−1 (M=4)) can store 8×4 (8 rows and 4 columns) data (32 data), as shown in FIG. k ) is stored in the address of the i-th row and j-th column of D bnk (i, j) (i, j: natural numbers, 0≦i≦7, 0≦j≦N−1 (N=4)) (i corresponds to the position in the width direction of the feature data (feature map), j corresponds to the channel (position in the channel direction) of the feature data (feature map), and k corresponds to the position in the height direction of the feature data (feature map). (4) Bank memory Tmem_k(bank k ) has an access bus that can simultaneously access data in the same row, and the bank memory Tmem_k (bank k ) can be independently (in parallel) read and write data from different rows of the bank memory Tmem_0 (bank 0 ) Data D on the 0th row bn0 (0,0) to D bn0 Data read processing is performed for (0, 3), and at the same time (in parallel), data D bn0 (7,0) to D bn0 (7, 3)). (5) The size of the kernel (weighting coefficient filter) of depthwise convolution (convolution processing in the spatial direction) is 3x3 (the weighting coefficient of the kernel is expressed as a 3x3 matrix). (6) The area (feature map) to be filtered by the kernel (the object of convolution processing) is an area of ​​4x8 size (4x8 feature data) as shown in Figures 5 and 6, the stride of the convolution processing is "1", and there is no padding. In other words, 0 , block 1 , block 2 , block 3 , block 4 , block 5 , block 6By shifting the target to which the 3x3 kernel filter is applied, the target of the convolution process is identified.

[0056] The feature data Din_f is input to the quantization processing unit 1 .

[0057] The quantization processing unit 1 performs quantization processing on the feature amount data Din_f, and outputs the quantized data to the CNN data processing unit 2 as data D1.

[0058] As shown in FIG. 8, the CNN data processing unit 2 performs the following steps: (1) data writing process (N channels (Ch 0 ~Ch N-1 (In this embodiment, N=4)) data of the bank memory Tmem_k(bank k (1) data writing process to N channels (Ch 0 ~Ch N-1 (In this embodiment, N=4)) data of the bank memory Tmem_k(bank k ) are executed in parallel.

[0059] Here, the processing of the CNN data processing unit 2 will be described with reference to the flowcharts of FIGS.

[0060] (Step S1w): In step S1w, the CNN data processing unit 2 executes a data writing process. For example, as shown in FIGS. 11 and 12, 0 (Time t 0 ~t 1 (1) Data D bn0 (4,0), D bn0 (4, 1), D bn0 (4, 2), D bn0 (4, 3) (2) Data D bn1 (4,0), D bn1 (4, 1), D bn1 (4, 2), D bn1 (4, 3) (3) Data D bn2(4,0), D bn2 (4, 1), D bn2 (4, 2), D bn2 (4, 3) (4) Data D bn3 (4,0), D bn3 (4, 1), D bn3 (4, 2), D bn3 (4, 3) (Data D bnh (i, j) is the channel Ch j When data D1 including "i" (which indicates data in which the width position is i and the height position is h) is output, the CNN data processing unit 2 performs the following process. In the above data, the height position h is 0 to 3, so the bank memories to be written to are set to bank0 to bank3. For this purpose, the CNN data processing unit 2 sets a variable hws (a variable specifying the start bank memory to be written to) and a variable hwe (a variable specifying the end bank memory to be written to) to specify the bank memories to be written to hws = 0 and hwe = 3 (hws, hwe: natural numbers, 0 ≦ hws ≦ M-1, 0 ≦ hwe ≦ M-1, hws < hwe) (step S11w).

[0061] (Step S12w_bnk_hws): In step S12w_bnk_hws, the bank memory bank hws (Bank memory Tmem_hws) (hws=0) stores the number of channels (Ch 0 ~Ch N-1 Specifically, the following process is performed: (1) The memory access control unit 21 of the CNN data processing unit 2 writes the data during the period T 0 (Time t 0 ~t 1 During the period, data D bn0 (4,0), D bn0 (4, 1), D bn0 (4, 2), D bn0 The four pieces of data (4, 3) are transferred in parallel (using four access buses (see FIG. 5)) to the bank memory Tmem_0 (bank memory bank hws , hws=0)(0) is output to the bank memory Tmem_0. (0) According to this, the bank memory Tmem_0 of the memory unit 22 stores the data D bn0 (4,0), D bn0 (4, 1), D bn0 (4, 2), D bn0 The four pieces of data (4, 3) are written in parallel (using four access buses) to areas with consecutive addresses in the bank memory Tmem_0 of the memory unit 22 (see FIG. 12).

[0062] (Steps S12w_bnk_1, S12w_bnk_2): In steps S12w_bnk_1 and S12w_bnk_2, the bank memory bank 1 , bank 2 (Bank memories Tmem_1, Tmem_2) each have the number of channels (Ch 0 ~Ch N-1 Specifically, the following process is performed: (2) The memory access control unit 21 of the CNN data processing unit 2 writes the data during the period T 0 (Time t 0 ~t 1 During the period, data D bn1 (4,0), D bn1 (4, 1), D bn1 (4, 2), D bn1 The four pieces of data (4, 3) are transferred in parallel (using four access buses (see FIG. 5)) to the bank memory Tmem_1 (bank memory bank 1 ) to write to consecutive addresses. (1) to the bank memory Tmem_1. (1) According to this, the bank memory Tmem_1 of the memory unit 22 stores the data D bn1 (4,0), D bn1 (4, 1), D bn1 (4, 2), D bn1(4, 3) are written in parallel (using four access buses) to areas with consecutive addresses in the bank memory Tmem_1 of the memory unit 22 (see FIG. 12). (3) The memory access control unit 21 of the CNN data processing unit 2 writes the four data (4, 3) in parallel (using four access buses) to areas with consecutive addresses in the bank memory Tmem_1 of the memory unit 22 (see FIG. 12). 0 (Time t 0 ~t 1 During the period, data D bn2 (4,0), D bn2 (4, 1), D bn2 (4, 2), D bn2 The four pieces of data (4, 3) are stored in parallel (using four access buses (see FIG. 5)) in the bank memory Tmem_2 (bank memory bank 2 ) to write to consecutive addresses. (2) to the bank memory Tmem_2. (2) Accordingly, the bank memory Tmem_2 of the memory unit 22 stores the data D bn2 (4,0), D bn2 (4, 1), D bn2 (4, 2), D bn2 The four pieces of data (4, 3) are written in parallel (using four access buses) to areas with consecutive addresses in the bank memory Tmem_2 of the memory unit 22 (see FIG. 12).

[0063] (Step S12w_bnk_hwe): In step S12w_bnk_hwe, the bank memory bank hwe (Bank memory Tmem_hwe) (hwe=3) has the number of channels (Ch 0 ~Ch N-1 Specifically, the following process is performed: (4) The memory access control unit 21 of the CNN data processing unit 2 writes the data during the period T 0 (Time t 0 ~t 1 During the period, data D bn3 (4,0), D bn3 (4, 1), D bn3 (4, 2), D bn3The four pieces of data (4, 3) are stored in parallel (using four access buses (see FIG. 5)) in the bank memory Tmem_3 (bank hwe , hwe=3) (3) to the bank memory Tmem_3. (3) According to this, the bank memory Tmem_3 of the memory unit 22 stores the data D bn3 (4,0), D bn3 (4, 1), D bn3 (4, 2), D bn3 The four pieces of data (4, 3) are written in parallel (using four access buses) to areas with consecutive addresses in the bank memory Tmem_3 of the memory unit 22 (see FIG. 12).

[0064] In the above, a case has been described in which hws = 0 and hwe = 3 are set and four consecutive data are written in relation to the height direction of the feature amount data (a case in which data is written in parallel to bank memories Tmem_0 to Tmem_3), but this is not limiting, and hws and hwe may be set to different values ​​and data may be written in parallel to multiple bank memories specified according to hws and hwe. Also, since each of bank memories Tmem_k (k: natural number, 0≦k≦M−1) can be accessed independently, for example, the data writing process to bank memory Tmem_k in the above process may be executed in parallel.

[0065] In the above, the period T 0 (Time t 0 ~t 1 The data write process during the period T 1 (Time t 1 ~t 2 (2) Period T 2 (Time t 1 ~t 2 (3) Period T 3 (Time t 2 ~t 3In the above cases (1) to (4), the bank memory Tmem_0 (bank 0 )~Tmem_3(bank 0 The following data is written in the period T (see FIG. 12): 1 (Time t 1 ~t 2 During this period: Bank memory Tmem_0 (bank 0 ) and data D bn0 (5,0), Data D bn0 (5, 1), data D bn0 (5,2), Data D bn0 (5, 3) is written to the bank memory Tmem_1 (bank 1 ) and data D bn1 (5,0), Data D bn1 (5, 1), data D bn1 (5,2), Data D bn1 (5, 3) is written to the bank memory Tmem_2 (bank 2 ) and data D bn2 (5,0), Data D bn2 (5, 1), data D bn2 (5,2), Data D bn2 (5, 3) is written to the bank memory Tmem_3 (bank 3 ) and data D bn3 (5,0), Data D bn3 (5, 1), data D bn3 (5,2), Data D bn3 (5, 3) is written. (2) Period T 2 (Time t 2 ~t 3 During this period: Bank memory Tmem_0 (bank 0 ) and data D bn0 (6,0), Data D bn0 (6, 1), data D bn0 (6,2), Data D bn0 (6, 3) is written to the bank memory Tmem_1 (bank 1 ) and data D bn1 (6,0), Data D bn1 (6, 1), data Dbn1 (6,2), Data D bn1 (6, 3) is written to the bank memory Tmem_2 (bank 2 ) and data D bn2 (6,0), Data D bn2 (6, 1), data D bn2 (6,2), Data D bn2 (6, 3) is written to the bank memory Tmem_3 (bank 3 ) and data D bn3 (6,0), Data D bn3 (6, 1), data D bn3 (6,2), Data D bn3 (6, 3) is written. (3) Period T 3 (Time t 3 ~t 4 During this period: Bank memory Tmem_0 (bank 0 ) and data D bn0 (7,0), Data D bn0 (7, 1), data D bn0 (7,2), Data D bn0 (7, 3) is written to the bank memory Tmem_1 (bank 1 ) and data D bn1 (7,0), Data D bn1 (7, 1), data D bn1 (7, 2), data D bn1 (7, 3) is written to the bank memory Tmem_2 (bank 2 ) and data D bn2 (7,0), Data D bn2 (7, 1), data D bn2 (7, 2), data D bn2 (7, 3) is written to the bank memory Tmem_3 (bank 3 ) and data D bn3 (7,0), Data D bn3 (7, 1), data D bn3 (7,2), Data D bn3 (7,3) is written.

[0066] (Step S2w): In step S2w, it is determined whether or not there is any data remaining to be written in the CNN data processing unit 2. If there is any data remaining to be written, the process returns to S1w and the same process as above is executed. On the other hand, if there is no data remaining to be written, the data writing process in the CNN data processing unit 2 is terminated.

[0067] (Step S1r): In step S1r, the CNN data processing unit 2 executes a data read process. For example, as shown in FIGS. 11 and 12, 0 (Time t 0 ~t 1 (1) Data D bn0 (0,0), D bn0 (0,1), D bn0 (0,2), D bn0 (0, 3) (2) Data D bn1 (0,0), D bn1 (0,1), D bn1 (0,2), D bn1 (0, 3) (3) Data D bn2 (0,0), D bn2 (0,1), D bn2 (0,2), D bn2 (0, 3) (data D bnh (i, j) is the channel Ch j In the above data, the position h in the height direction is 0 to 2, so the bank memory to be read is bank 0 ~bank 2(This is set because the size of the kernel (weighting coefficient filter) of depthwise convolution (convolution processing in the spatial direction) is 3 × 3 and the size in the height direction is "3".) For this reason, in the CNN data processing unit 2, the variable hrs (variable specifying the starting bank memory of the read processing target) and the variable hre (variable specifying the ending bank memory of the read processing target) specifying the bank memory to be read are set to hrs = 0 and hre = 2 (hrs, hre: natural numbers, 0 ≦ hrs ≦ M-1, 0 ≦ hre ≦ M-1, hrs < hre) (step S110r).

[0068] (Step S111r_bnk_hrs (hrs=0)): In step S111r_bnk_hrs, the bank memory bank hrs (Bank memory Tmem_hrs) (hrs=0) for the number of channels (Ch 0 ~Ch N-1 ) (N=4) data read processing is executed. Specifically, the following processing is executed. (1) The memory access control unit 21 of the CNN data processing unit 2 executes the read processing of the data of the period T 0 (Time t 0 ~t 1 During the period, data D bn0 (0,0), D bn0 (0,1), D bn0 (0,2), D bn0 The four data (0, 3) are transferred in parallel (using four access buses (see FIG. 5)) to the bank memory Tmem_0 (bank memory bank hrs , hrs=0) (0) to the bank memory Tmem_0. (0) According to this, the bank memory Tmem_0 of the memory unit 22 stores the data D bn0 (0,0), D bn0 (0,1), D bn0 (0,2), D bn0Four pieces of data (0, 3) (four channels of data) are read in parallel (using four access buses) from consecutive address areas of the bank memory Tmem_0 of the memory unit 22 (see FIG. 12) (see FIG. 13 for the relationship between the four channels of data read in parallel and the spatial position of the feature data (feature map)).

[0069] (Step S111r_bnk_1): In step S111r_bnk_1, the bank memory bank 1 From the number of channels (Ch 0 ~Ch N-1 ) (N=4) data read processing is executed. Specifically, the following processing is executed. (2) The memory access control unit 21 of the CNN data processing unit 2 executes the read processing of the data of the period T 0 (Time t 0 ~t 1 During the period, data D bn1 (0,0), D bn1 (0,1), D bn1 (0,2), D bn1 The four data (0, 3) are transferred in parallel (using four access buses (see FIG. 5)) to the bank memory Tmem_1 (bank memory bank 1 ) to read from the area where the addresses are consecutive. (1) to the bank memory Tmem_1. (1) According to this, the bank memory Tmem_1 of the memory unit 22 stores the data D bn1 (0,0), D bn1 (0,1), D bn1 (0,2), D bn1 Four pieces of data (0, 3) (four channels of data) are read in parallel (using four access buses) from consecutive address areas of the bank memory Tmem_1 of the memory unit 22 (see FIG. 12) (see FIG. 14 for the relationship between the four channels of data read in parallel and the spatial position of the feature data (feature map)).

[0070] (Step S111r_bnk_hre (hre=2)): In step S111r_bnk_hre, the bank memory bank hre (Bank memory Tmem_hre) (hre=2) for the number of channels (Ch 0 ~Ch N-1 ) (N=4) data read processing is executed. Specifically, the following processing is executed. (3) The memory access control unit 21 of the CNN data processing unit 2 executes the read processing of the data of the period T 0 (Time t 0 ~t 1 During the period, data D bn2 (0,0), D bn2 (0,1), D bn2 (0,2), D bn2 The four data (0, 3) are transferred in parallel (using four access buses (see FIG. 5)) to the bank memory Tmem_2 (bank memory bank 2 ) to read from the area where the addresses are consecutive. (2) to the bank memory Tmem_2. (2) Accordingly, the bank memory Tmem_2 of the memory unit 22 stores the data D bn2 (0,0), D bn2 (0,1), D bn2 (0,2), D bn2 Four pieces of data (0, 3) (four channels of data) are read in parallel (using four access buses) from consecutive address areas of the bank memory Tmem_2 of the memory unit 22 (see FIG. 12) (see FIG. 15 for the relationship between the four channels of data read in parallel and the spatial position of the feature data (feature map)).

[0071] (Steps S112r_bnk_hrs (hrs=0) to S112r_bnk_hre (hre=2)): In steps S112r_bnk_hrs (hrs=0) to S112r_bnk_hre (hre=2), a process for determining the number of output systems Num_sys is executed. 0 The data read out by 0Since it is the data in the first column of the other blocks (block 1 ~block 5 ) and there is no data in common (overlapping data). Therefore, in each of steps S112r_bnk_hrs (hrs=0) to S112r_bnk_hre (hre=2), the number of output systems Num_sys=1 is determined.

[0072] (Steps S113r_bnk_hrs (hrs=0) to S113r_bnk_hre (hre=2)): In steps S113r_bnk_hrs (hrs=0) to S113r_bnk_hre (hre=2), data sets for the number of output systems Num_sys (=1) are output simultaneously (register write processing is executed). Specifically, the following processing is executed.

[0073] The memory access control unit 21 of the CNN data processing unit 2 0 In this case, the data D of the bank memory Tmem_k is bnk Control signal Ctl_r for reading (0,0) (k) and generates the control signal Ctl_r (k) to the bank memory Tmem_k, and the data D bnk The bank memory Tmem_k generates a control signal Ctl_reg that instructs that (0, 0) be written to an area at a predetermined address in the register unit 23, and outputs the control signal Ctl_reg to the register unit 23. (k) According to the data D bnk (0,0) is read out, and the register unit 23 reads out the data D output from the bank memory Tmem_k in accordance with the control signal Ctl_reg. bnk (0,0) is written to a predetermined address, which is designated by the control signal Ctl_reg.

[0074] Period T 0 In this case, channel 0 (Ch 0 ) data and channel k (Ch k) (1≦k≦3), the data output from the bank memory Tmem_k and the address of the register unit 23 to which the data is written are as follows: 0 >> (Output to 1 system (Num_sys = 1)) Channel 0 (Ch 0 (1) Bank memory Tmem_0 (bank 0 ) Data D bn0 (0,0) → Address adr00 of register unit 23 (Ch0) (2) Bank memory Tmem_1 (bank 1 ) Data D bn1 (0,0) → Address adr03 of register unit 23 (Ch0) (3) Bank memory Tmem_2 (bank 2 ) Data D bn2 (0,0) → Address adr06 of register unit 23 (Ch0) Channel k (Ch k ) (k: natural number, 1≦k≦3): (1) Bank memory Tmem_0 (bank 0 ) Data D bn0 (0, k) → address adr00 of register unit 23 (Chk) (2) Bank memory Tmem_1 (bank 1 ) Data D bn1 (0, k) → address adr03 of register unit 23 (Chk) (3) Bank memory Tmem_2 (bank 2 ) Data D bn2 (0, k) → address adr06 of register unit 23 (Chk) 18 shows the relationship between the above data (data of channel 0) and the address of the write destination in the register unit 23. As shown in FIG. 0 In this case, data D bn0 (0,0), data D bn1 (0,0), data D bn2 (0,0) is address adr00 (Ch0) , address adr03 (Ch0) (= address adr00 (Ch0) +3 address), address adr06 (Ch0)(= address adr03 (Ch0) +3 address). 0 In this case, data D bn0 (0,0), data D bn1 (0,0), data D bn2 (0,0) is written to every third address of the register unit 23 (the bold rectangles in FIG. 18 indicate the data to be written). This is because the area (kernel size) to be subjected to the convolution process is 3x3, so that the 3x3 data can be reshaped into 1x9 data according to the area (kernel size) to be subjected to the convolution process and output to the quantization data memory unit 3. Note that the address adr00 of the register unit 23 (Chk) ~adr08 (Chk) (k: natural number, 0≦k≦N−1) are assumed to be consecutive addresses.

[0075] (Step S12r): In step S12r, a determination process is performed to determine whether a predetermined amount of data has been output from the bank memory Tmem_k of the memory unit 22 to the register unit 23. 0 At the time when the period T 0 (See the section below), and the process returns to step S11r.

[0076] (Step S1r) (Period T 1 ): In step S1r, the CNN data processing unit 2 executes a data read process. For example, as shown in FIGS. 11 and 17, 1 (Time t 1 ~t 2 (1) Data D bn0 (1,0), D bn0 (1,1), D bn0 (1, 2), D bn0 (1, 3) (2) Data D bn1 (1,0), D bn1 (1,1), D bn1(1, 2), D bn1 (1, 3) (3) Data D bn2 (1,0), D bn2 (1,1), D bn2 (1, 2), D bn2 (1, 3) (Data D bnh (i, j) is the channel Ch j In this case, the CNN data processing unit 2 performs the following process. In the above data, the height position h is 0 to 2, so the bank memory to be read is bank 0 ~bank 2 (This is set because the size of the kernel (weighting coefficient filter) of depthwise convolution (convolution processing in the spatial direction) is 3 × 3 and the size in the height direction is "3".) For this reason, in the CNN data processing unit 2, the variable hrs (variable specifying the starting bank memory of the read processing target) and the variable hre (variable specifying the ending bank memory of the read processing target) specifying the bank memory to be read are set to hrs = 0 and hre = 2 (hrs, hre: natural numbers, 0 ≦ hrs ≦ M-1, 0 ≦ hre ≦ M-1, hrs < hre) (step S110r).

[0077] (Step S111r_bnk_hrs (hrs=0)): In step S111r_bnk_hrs, the bank memory bank hrs (Bank memory Tmem_hrs) (hrs=0) for the number of channels (Ch 0 ~Ch N-1 ) (N=4) data read processing is executed. Specifically, the following processing is executed. (1) The memory access control unit 21 of the CNN data processing unit 2 executes the read processing of the data of the period T 1 (Time t 1 ~t 2 During the period, data D bn0 (1,0), D bn0 (1,1), D bn0 (1, 2), D bn0The four data (1, 3) are transferred in parallel (using four access buses (see FIG. 5)) to the bank memory Tmem_0 (bank memory bank hrs , hrs=0) (0) to the bank memory Tmem_0. (0) According to this, the bank memory Tmem_0 of the memory unit 22 stores the data D bn0 (1,0), D bn0 (1,1), D bn0 (1, 2), D bn0 The four pieces of data (1, 3) (data for four channels) are read in parallel (using four access buses) from areas with consecutive addresses in the bank memory Tmem_0 of the memory unit 22.

[0078] (Step S111r_bnk_1): In step S111r_bnk_1, the bank memory bank 1 From the number of channels (Ch 0 ~Ch N-1 ) (N=4) data read processing is executed. Specifically, the following processing is executed. (2) The memory access control unit 21 of the CNN data processing unit 2 executes the read processing of the data of the period T 1 (Time t 1 ~t 2 During the period, data D bn1 (1,0), D bn1 (1,1), D bn1 (1, 2), D bn1 The four data (1, 3) are transferred in parallel (using four access buses (see FIG. 5)) to the bank memory Tmem_1 (bank memory bank 1 ) to read from the area where the addresses are consecutive. (1) to the bank memory Tmem_1. (1) According to this, the bank memory Tmem_1 of the memory unit 22 stores the data D bn1 (1,0), D bn1 (1,1), D bn1 (1, 2), D bn1The four pieces of data (1, 3) (data for four channels) are read in parallel (using four access buses) from areas with consecutive addresses in the bank memory Tmem_1 of the memory unit 22.

[0079] (Step S111r_bnk_hre (hre=2)): In step S111r_bnk_hre, the bank memory bank hre (Bank memory Tmem_hre) (hre=2) for the number of channels (Ch 0 ~Ch N-1 ) (N=4) data read processing is executed. Specifically, the following processing is executed. (3) The memory access control unit 21 of the CNN data processing unit 2 executes the read processing of the data of the period T 1 (Time t 1 ~t 2 During the period, data D bn2 (1,0), D bn2 (1,1), D bn2 (1, 2), D bn2 The four data (1, 3) are transferred in parallel (using four access buses (see FIG. 5)) to the bank memory Tmem_2 (bank memory bank 2 ) to read from the area where the addresses are consecutive. (2) to the bank memory Tmem_2. (2) Accordingly, the bank memory Tmem_2 of the memory unit 22 stores the data D bn2 (1,0), D bn2 (1,1), D bn2 (1, 2), D bn2 The four pieces of data (1, 3) (data for four channels) are read in parallel (using four access buses) from areas with consecutive addresses in the bank memory Tmem_2 of the memory unit 22.

[0080] (Steps S112r_bnk_hrs (hrs=0) to S112r_bnk_hre (hre=2)): In steps S112r_bnk_hrs (hrs=0) to S112r_bnk_hre (hre=2), a process for determining the number of output systems Num_sys is executed. 1 The data read out by 0 The second column of data and block 1 Therefore, in each of steps S112r_bnk_hrs (hrs=0) to S112r_bnk_hre (hre=2), the number of output systems Num_sys=2 is determined.

[0081] (Steps S113r_bnk_hrs (hrs=0) to S113r_bnk_hre (hre=2)): In steps S113r_bnk_hrs (hrs=0) to S113r_bnk_hre (hre=2), data sets of the number of output systems Num_sys (=2) are output simultaneously (register write processing is executed). Specifically, the following processing is executed.

[0082] The memory access control unit 21 of the CNN data processing unit 2 1 In this case, the data D of the bank memory Tmem_k is bnk Control signal Ctl_r for reading (1,0) (k) and generates the control signal Ctl_r (k) to the bank memory Tmem_k, and the data D bnk The bank memory Tmem_k generates a control signal Ctl_reg that instructs that (1, 0) be written to an area at a predetermined address in the register unit 23, and outputs the control signal Ctl_reg to the register unit 23. (k) According to the data D bnk The register unit 23 reads out (1, 0) and the data D output from the bank memory Tmem_k in accordance with the control signal Ctl_reg. bnk(1, 0) is written to a predetermined address, which is assumed to be specified by the control signal Ctl_reg.

[0083] Period T 1 In this case, channel 0 (Ch 0 ) data and channel k (Ch k ) (1≦k≦3), the data output from the bank memory Tmem_k and the address of the register unit 23 to which the data is written are as follows: 1 >> (Output to two systems (Num_sys = 2)) Channel 0 (Ch 0 (1) Bank memory Tmem_0 (bank 0 ) Data D bn0 (1,0) → Address adr01 of register unit 23 (Ch0) →Address adr10 of register unit 23 (Ch0) (2) Bank memory Tmem_1 (bank 1 ) Data D bn1 (1,0) → Address adr04 of register unit 23 (Ch0) →Address adr13 of register unit 23 (Ch0) (3) Bank memory Tmem_2 (bank 2 ) Data D bn2 (1,0) → Address adr07 of register unit 23 (Ch0) →Address adr16 of register unit 23 (Ch0) Channel k (Ch k ) (k: natural number, 1≦k≦3): (1) Bank memory Tmem_0 (bank 0 ) Data D bn0 (1, k) → address adr01 of register unit 23 (Chk) →Address adr10 of register unit 23 (Chk) (2) Bank memory Tmem_1 (bank 1 ) Data D bn1 (1, k) → address adr04 of register unit 23 (Chk) →Address adr13 of register unit 23 (Chk) (3) Bank memory Tmem_2 (bank2 ) Data D bn2 (1, k) → address adr07 of register unit 23 (Chk) →Address adr16 of register unit 23 (Chk) 18 and 19 (period T 1 The part ( ) shows the relationship between the above data (data of channel 0) and the address of the register section 23 to which the data is written.

[0084] As shown in FIG. 1 In this case, data D bn0 (1,0), Data D bn1 (1,0), Data D bn2 (1, 0) are address adr01 (Ch0) , address adr04 (Ch0) (= address adr01 (Ch0) +3 address), address adr07 (Ch0) (= address adr04 (Ch0) +3 address). 1 In this case, data D bn0 (1,0), Data D bn1 (1,0), Data D bn2 (1, 0) is written to every third address of the register unit 23 (the bold rectangles in FIG. 18 indicate the data to be written). This is because the area to be subjected to the convolution process (kernel size) is 3×3, and therefore the 3×3 data is shaped into 1×9 data in accordance with the area to be subjected to the convolution process (kernel size) so that it can be output to the quantization data memory unit 3.

[0085] Also, as shown in FIG. 1 In this case, data D bn0 (1,0), Data D bn1 (1,0), Data D bn2 (1, 0) are addresses adr10 (Ch0) , address adr13 (Ch0) (= address adr10 (Ch0) +3 address), address adr16 (Ch0) (= address adr13 (Ch0)+3 address). 1 In this case, data D bn0 (1,0), Data D bn1 (1,0), Data D bn2 (1, 0) is written to every third address of the register unit 23 (the bold rectangles in FIG. 19 indicate the data to be written). This is because the area (kernel size) to be subjected to the convolution process is 3x3, so that the 3x3 data can be reshaped into 1x9 data according to the area (kernel size) to be subjected to the convolution process and output to the quantization data memory unit 3. Note that the address adr10 of the register unit 23 (Chk) ~adr18 (Chk) (k: natural number, 0≦k≦N−1) are assumed to be consecutive addresses.

[0086] (Step S12r): In step S12r, a determination process is performed to determine whether a predetermined amount of data has been output from the bank memory Tmem_k of the memory unit 22 to the register unit 23. 1 At the time when the period T 1 , the period T 1 (See the section below), and the process returns to step S11r.

[0087] (Step S1r) (Period T 2 ): In step S1r, the CNN data processing unit 2 executes a data read process. For example, as shown in FIGS. 11 and 17, 2 (Time t 2 ~t 3 (1) Data D bn0 (2,0), D bn0 (2,1), D bn0 (2,2), D bn0 (2, 3) (2) Data D bn1 (2,0), D bn1 (2,1), D bn1(2,2), D bn1 (2, 3) (3) Data D bn2 (2,0), D bn2 (2,1), D bn2 (2,2), D bn2 (2, 3) (Data D bnh (i, j) is the channel Ch j In this case, the CNN data processing unit 2 performs the following process. In the above data, the height position h is 0 to 2, so the bank memory to be read is bank 0 ~bank 2 (This is set because the size of the kernel (weighting coefficient filter) of depthwise convolution (convolution processing in the spatial direction) is 3 × 3 and the size in the height direction is "3".) For this reason, in the CNN data processing unit 2, the variable hrs (variable specifying the starting bank memory of the read processing target) and the variable hre (variable specifying the ending bank memory of the read processing target) specifying the bank memory to be read are set to hrs = 0 and hre = 2 (hrs, hre: natural numbers, 0 ≦ hrs ≦ M-1, 0 ≦ hre ≦ M-1, hrs < hre) (step S110r).

[0088] (Step S111r_bnk_hrs (hrs=0)): In step S111r_bnk_hrs, the bank memory bank hrs (Bank memory Tmem_hrs) (hrs=0) for the number of channels (Ch 0 ~Ch N-1 ) (N=4) data read processing is executed. Specifically, the following processing is executed. (1) The memory access control unit 21 of the CNN data processing unit 2 executes the read processing of the data of the period T 2 (Time t 2 ~t 3 During the period, data D bn0 (2,0), D bn0 (2,1), D bn0 (2,2), D bn0The four data (2, 3) are transferred in parallel (using four access buses (see FIG. 5)) to the bank memory Tmem_0 (bank memory bank hrs , hrs=0) (0) to the bank memory Tmem_0. (0) According to this, the bank memory Tmem_0 of the memory unit 22 stores the data D bn0 (2,0), D bn0 (2,1), D bn0 (2,2), D bn0 The four pieces of data (2, 3) (data for four channels) are read in parallel (using four access buses) from areas with consecutive addresses in the bank memory Tmem_0 of the memory unit 22.

[0089] (Step S111r_bnk_1): In step S111r_bnk_1, the bank memory bank 1 From the number of channels (Ch 0 ~Ch N-1 ) (N=4) data read processing is executed. Specifically, the following processing is executed. (2) The memory access control unit 21 of the CNN data processing unit 2 executes the read processing of the data of the period T 2 (Time t 2 ~t 3 During the period, data D bn1 (2,0), D bn1 (2,1), D bn1 (2,2), D bn1 The four data (2, 3) are transferred in parallel (using four access buses (see FIG. 5)) to the bank memory Tmem_1 (bank memory bank 1 ) to read from the area where the addresses are consecutive. (1) to the bank memory Tmem_1. (1) According to this, the bank memory Tmem_1 of the memory unit 22 stores the data D bn1 (2,0), D bn1 (2,1), D bn1 (2,2), D bn1The four pieces of data (2, 3) (data for four channels) are read in parallel (using four access buses) from areas with consecutive addresses in the bank memory Tmem_1 of the memory unit 22.

[0090] (Step S111r_bnk_hre (hre=2)): In step S111r_bnk_hre, the bank memory bank hre (Bank memory Tmem_hre) (hre=2) for the number of channels (Ch 0 ~Ch N-1 ) (N=4) data read processing is executed. Specifically, the following processing is executed. (3) The memory access control unit 21 of the CNN data processing unit 2 executes the read processing of the data of the period T 2 (Time t 2 ~t 3 During the period, data D bn2 (2,0), D bn2 (2,1), D bn2 (2,2), D bn2 The four data (2, 3) are transferred in parallel (using four access buses (see FIG. 5)) to the bank memory Tmem_2 (bank memory bank 2 ) to read from the area where the addresses are consecutive. (2) to the bank memory Tmem_2. (2) Accordingly, the bank memory Tmem_2 of the memory unit 22 stores the data D bn2 (2,0), D bn2 (2,1), D bn2 (2,2), D bn2 The four pieces of data (2, 3) (data for four channels) are read in parallel (using four access buses) from areas with consecutive addresses in the bank memory Tmem_2 of the memory unit 22.

[0091] (Steps S112r_bnk_hrs (hrs=0) to S112r_bnk_hre (hre=2)): In steps S112r_bnk_hrs (hrs=0) to S112r_bnk_hre (hre=2), a process for determining the number of output systems Num_sys is executed. 2 The data read out by 0 The third column of data and block 1 The second column of data and block 2 Therefore, in each of steps S112r_bnk_hrs (hrs=0) to S112r_bnk_hre (hre=2), the number of output systems Num_sys=3 is determined.

[0092] (Steps S113r_bnk_hrs (hrs=0) to S113r_bnk_hre (hre=2)): In steps S113r_bnk_hrs (hrs=0) to S113r_bnk_hre (hre=2), data for the number of output systems Num_sys (=3) is output simultaneously (register write processing is executed). Specifically, the following processing is executed.

[0093] The memory access control unit 21 of the CNN data processing unit 2 2 In this case, the data D of the bank memory Tmem_k is bnk Control signal Ctl_r for reading (2,0) (k) and generates the control signal Ctl_r (k) to the bank memory Tmem_k, and the data D bnk The bank memory Tmem_k generates a control signal Ctl_reg that instructs that (2, 0) be written to an area at a predetermined address in the register unit 23, and outputs the control signal Ctl_reg to the register unit 23. (k) According to the data D bnk (2, 0), and the register unit 23 reads the data D output from the bank memory Tmem_k in accordance with the control signal Ctl_reg. bnk(2, 0) is written to a predetermined address, which is assumed to be specified by the control signal Ctl_reg.

[0094] Period T 2 In this case, channel 0 (Ch 0 ) data and channel k (Ch k ) (1≦k≦3), the data output from the bank memory Tmem_k and the address of the register unit 23 to which the data is written are as follows: 2 >> (Output to 3 systems (Num_sys = 3)) Channel 0 (Ch 0 (1) Bank memory Tmem_0 (bank 0 ) Data D bn0 (2,0) → Address adr02 of register unit 23 (Ch0) →Address adr11 of register unit 23 (Ch0) →Address adr20 of register unit 23 (Ch0) (2) Bank memory Tmem_1 (bank 1 ) Data D bn1 (2,0) → Address adr05 of register unit 23 (Ch0) →Address adr14 of register unit 23 (Ch0) →Address adr23 of register section 23 (Ch0) (3) Bank memory Tmem_2 (bank 2 ) Data D bn2 (2,0) → Address adr08 of register unit 23 (Ch0) →Address adr17 of register unit 23 (Ch0) →Address adr26 of register unit 23 (Ch0) Channel k (Ch k ) (k: natural number, 1≦k≦3): (1) Bank memory Tmem_0 (bank 0 ) Data D bn0 (2, k) → address adr02 of register unit 23 (Chk) →Address adr11 of register unit 23 (Chk) →Address adr20 of register unit 23 (Chk)(2) Bank memory Tmem_1 (bank 1 ) Data D bn1 (2, k) → address adr05 of register unit 23 (Chk) →Address adr14 of register unit 23 (Chk) →Address adr23 of register section 23 (Chk) (3) Bank memory Tmem_2 (bank 2 ) Data D bn2 (2, k) → address adr08 of register unit 23 (Chk) →Address adr17 of register unit 23 (Chk) →Address adr26 of register unit 23 (Chk) 18 to 20 (period T 2 The part ( ) shows the relationship between the above data (data of channel 0) and the address of the register section 23 to which the data is written.

[0095] As shown in FIG. 2 In this case, data D bn0 (2,0), Data D bn1 (2,0), Data D bn2 (2, 0) are addresses adr02 (Ch0) , address adr05 (Ch0) (= address adr02 (Ch0) +3 address), address adr08 (Ch0) (= address adr05 (Ch0) +3 address). 2 In this case, data D bn0 (2,0), Data D bn1 (2,0), Data D bn2 (2, 0) is written to every third address of the register unit 23 (the bold rectangles in FIG. 18 indicate the data to be written). This is because the area to be subjected to the convolution process (kernel size) is 3×3, and therefore the 3×3 data is shaped into 1×9 data in accordance with the area to be subjected to the convolution process (kernel size) so that it can be output to the quantization data memory unit 3.

[0096] Also, as shown in FIG. 2 In this case, data D bn0 (2,0), Data D bn1 (2,0), Data D bn2 (2, 0) are addresses adr11 (Ch0) , address adr14 (Ch0) (= address adr11 (Ch0) +3 address), address adr17 (Ch0) (= address adr14 (Ch0) +3 address). 2 In this case, data D bn0 (2,0), Data D bn1 (2,0), Data D bn2 (2, 0) is written to every third address of the register unit 23 (the bold rectangles in FIG. 19 indicate the data to be written). This is because the area to be subjected to the convolution process (kernel size) is 3×3, and therefore the 3×3 data is shaped into 1×9 data in accordance with the area to be subjected to the convolution process (kernel size) so that it can be output to the quantization data memory unit 3.

[0097] Also, as shown in FIG. 2 In this case, data D bn0 (2,0), Data D bn1 (2,0), Data D bn2 (2, 0) are addresses adr20 (Ch0) , address adr23 (Ch0) (= address adr20 (Ch0) +3 address), address adr26 (Ch0) (= address adr23 (Ch0) +3 address). 2 In this case, data D bn0 (2,0), Data D bn1 (2,0), Data D bn2(2, 0) is written to every third address of the register unit 23 (the bold rectangles in FIG. 20 indicate the data to be written). This is because the area (kernel size) to be subjected to the convolution process is 3x3, so that the 3x3 data can be reshaped into 1x9 data according to the area (kernel size) to be subjected to the convolution process and output to the quantization data memory unit 3. Note that the address adr20 of the register unit 23 (Chk) ~adr28 (Chk) (k: natural number, 0≦k≦N−1) are assumed to be consecutive addresses.

[0098] (Step S12r): In step S12r, a determination process is performed to determine whether a predetermined amount of data has been output from the bank memory Tmem_k of the memory unit 22 to the register unit 23. 2 When the process is completed, as shown in FIG. 0 Since all data in the region (kernel size) to be subjected to the convolution process has been output from the bank memory Tmem_k of the memory unit 22 to the register unit 23 (the period T 2 (See the section below), and the process proceeds to step S13r.

[0099] (Step S13r): In step S13r, a register output process is executed. 0 Since all data in the region (kernel size) to be subjected to the convolution process has been output from the bank memory Tmem_k of the memory unit 22 to the register unit 23, the register unit 23 outputs data including the following data as data D2 to the quantized data memory unit 3. 0 ) (Ch 0 )≫(Period T 2 ) Data D bn0 (0,0), data D bn0 (1,0), Data D bn0 (2,0) Data D bn1 (0,0), data D bn1 (1,0), Data D bn1 (2,0) Data Dbn2 (0,0), data D bn2 (1,0), Data D bn2 (2, 0) (The above data is stored in the area of ​​consecutive addresses in the register section 23 (adr00 (Ch0) ~adr08 (Ch0) ) (see FIG. 21) ≪Feature data (quantized data) (block 0 ) (Ch 1 )≫(Period T 2 ) Data D bn0 (0, 1), data D bn0 (1,1), data D bn0 (2, 1) Data D bn1 (0, 1), data D bn1 (1,1), data D bn1 (2, 1) Data D bn2 (0, 1), data D bn2 (1,1), data D bn2 (2, 1) (The above data is stored in the area of ​​consecutive addresses in the register section 23 (adr00 (Ch1) ~adr08 (Ch1) ) is stored in the block. 0 ) (Ch 2 )≫(Period T 2 ) Data D bn0 (0,2), data D bn0 (1, 2), data D bn0 (2, 2) Data D bn1 (0,2), data D bn1 (1, 2), data D bn1 (2, 2) Data D bn2 (0,2), data D bn2 (1, 2), data D bn2 (2, 2) (The above data is stored in the area of ​​consecutive addresses in the register section 23 (adr00 (Ch2) ~adr08 (Ch2) ) is stored in the block. 0 ) (Ch 3 )≫(Period T 2 ) Data D bn0 (0,3), Data D bn0 (1, 3), Data Dbn0 (2, 3) Data D bn1 (0,3), Data D bn1 (1, 3), Data D bn1 (2, 3) Data D bn2 (0,3), Data D bn2 (1, 3), Data D bn2 (2, 3) (The above data is stored in the area of ​​consecutive addresses in the register section 23 (adr00 (Ch3) ~adr08 (Ch3) ) is stored in

[0100] (Step S2r): In step S2r, it is determined whether or not there is any data remaining to be processed by the CNN data processing unit 2. If there is any data remaining to be processed, the process returns to S11r and the same process as above is executed. On the other hand, if there is no data remaining to be processed, the data reading process by the CNN data processing unit 2 is terminated.

[0101] If there is data remaining to be processed and the period T 3 Regarding the processing of the CNN data processing unit 2, 2 The same process is performed.

[0102] Period T 3 In the process of step S13r, register output process is executed, and block 1 Since all data in the region (kernel size) to be subjected to the convolution process has been output from the bank memory Tmem_k of the memory unit 22 to the register unit 23, the register unit 23 outputs data including the following data as data D2 to the quantized data memory unit 3. 1 ) (Ch 0 )≫(Period T 3 ) Data D bn0 (1,0), Data D bn0 (2,0), Data D bn0 (3,0) Data D bn1 (1,0), Data D bn1 (2,0), Data D bn1 (3,0) Data D bn2(1,0), Data D bn2 (2,0), Data D bn2 (3, 0) (The above data is stored in the area (adr10) where the addresses of the register section 23 are consecutive. (Ch0) ~adr18 (Ch0) ) (see FIG. 21) ≪Feature data (quantized data) (block 1 ) (Ch 1 )≫(Period T 3 ) Data D bn0 (1,1), data D bn0 (2, 1), data D bn0 (3, 1) Data D bn1 (1,1), data D bn1 (2, 1), data D bn1 (3, 1) Data D bn2 (1,1), data D bn2 (2, 1), data D bn2 (3, 1) (The above data is stored in the area (adr10) where the addresses of the register section 23 are consecutive. (Ch1) ~adr18 (Ch1) ) is stored in the block. 1 ) (Ch 2 )≫(Period T 3 ) Data D bn0 (1, 2), data D bn0 (2, 2), data D bn0 (3, 2) Data D bn1 (1, 2), data D bn1 (2, 2), data D bn1 (3, 2) Data D bn2 (1, 2), data D bn2 (2, 2), data D bn2 (3, 2) (The above data is stored in the area (adr10) of the register unit 23 where the addresses are consecutive. (Ch2) ~adr18 (Ch2) ) is stored in the block. 1 ) (Ch 3 )≫(Period T 3 ) Data D bn0 (1, 3), Data D bn0 (2, 3), data D bn0(3, 3) Data D bn1 (1, 3), Data D bn1 (2, 3), data D bn1 (3, 3) Data D bn2 (1, 3), Data D bn2 (2, 3), data D bn2 (3, 3) (The above data is stored in the area (adr10) where the addresses of the register section 23 are consecutive. (Ch3) ~adr18 (Ch3) ) In the CNN data processing unit 2, the period T 4 Subsequent processing is similarly executed, and when there is no data remaining to be processed, the processing is terminated.

[0103] The quantized data memory unit 3 receives the data D2 output from the register unit 23 of the CNN data processing unit 2 and stores the data D2. The data D2 output from the register unit 23 of the CNN data processing unit 2 is data obtained by shaping 3×3 data into 1×9 data according to the region to be subjected to convolution processing (kernel size (3×3 in this embodiment)), so the quantized data memory unit 3 stores the data D2 in a region of consecutive addresses, for example.

[0104] The convolution processing unit 4 reads data of an area to be subjected to convolution processing using the input weighting coefficient data Din_w (weighting filter (kernel)) from the quantized data memory unit 3. Then, the convolution processing unit 4 performs convolution processing (convolution operation) on the data read from the quantized data memory unit 3 using the weighting coefficient data Din_w (3×3 kernel in this embodiment), obtains data after the convolution processing, and outputs the obtained data as data Dout.

[0105] Summary As described above, in the CNN data processing device 100, the CNN data processing unit 2 can execute, in parallel, the data writing process of the data output from the quantization processing unit 1 (data after quantization of feature data) to the memory unit 22 and the data reading process from the memory unit 22. Furthermore, the memory unit 22 has multiple bank memories Tmem_k, and can write and / or read multiple pieces of data simultaneously (in parallel). Therefore, the CNN data processing device 100 can realize high-speed data writing and data reading processes. In the CNN data processing device 100, (1) a plurality of access buses are provided for each of the plurality of bank memories Tmem_k of the memory unit 22, and data for a plurality of channels can be accessed simultaneously (in parallel), and (2) different (independent) bank memories Tmem_k are assigned for each height direction of the convolution processing target area (the area to be convolved with the kernel), so that a plurality of data in different height directions can be accessed simultaneously (in parallel). For this reason, in the CNN data processing device 100, the period of one data read processing (period T i ), it is possible to read out data for multiple channels of h×1 (h rows and 1 column, h: height position) in the region to be convolution processed.

[0106] Furthermore, the CNN data processing device 100 acquires the number of output systems Num_sys, which is the number of overlapping data sets (h × 1 data sets in the convolution processing target area), according to the position of the convolution processing target area (slid position), and outputs the overlapping data sets (h × 1 data sets in the convolution processing target area) equal to the acquired number of output systems Num_sys to the register unit 23, each in a separate system (in parallel).

[0107] As a result, in the CNN data processing device 100, the number of times overlapping data is read can be reduced by sliding the position of the convolution processing target region.

[0108] The CNN data processing device 100 also includes a register unit 23, in which data read from the memory unit 22 is written to discrete addresses (addresses to which a predetermined offset value (corresponding to the size of the kernel in the width direction (in the case of a 3 × 3 kernel, the offset value is "3")) is added) according to the size (shape) of the region to be convolution processed (size (shape) of the kernel), and after all data of the region to be convolution processed (quantized data of feature amount data) has been collected (after all data of the region to be convolution processed has been written at consecutive addresses in the register unit 23), all data of the region to be convolution processed is output to the quantized data memory unit 3.

[0109] As a result, the CNN data processing device 100 can output all data in the convolution processing target region (data to be subjected to convolution processing) as data arranged in the order of convolution calculation, and write it to the quantized data memory unit 3. Then, the data arranged in the order of convolution calculation is read from the quantized data memory unit 3, and the convolution processing unit 4 performs convolution processing using kernel weighting coefficient data to be applied to the data, thereby enabling high-speed convolution processing.

[0110] In this way, the CNN data processing device 100 can acquire data arranged in the order in which convolution operations are performed while reducing the number of times duplicate data is read, simply by providing the CNN data processing unit 2. Therefore, the CNN data processing device 100 can perform data processing to realize a high-performance, high-speed CNN model that can reduce the number of times feature data is read and shorten the time required for the entire convolution process, including the feature data read process.

[0111] [Other Embodiments] In the above embodiment, a case has been described in which weight coefficient data Din_w (weight filter (kernel)) is input to the convolution processing unit 4 in the CNN data processing device 100, and convolution processing is performed using the weight coefficient data Din_w (weight filter (kernel)). However, this is not limited to this. For example, the convolution processing unit 4 may perform vector decomposition processing on the weight filter (kernel) to decompose it into a basis matrix and a real coefficient vector, and the decomposed basis matrix and real coefficient vector may be input to perform convolution processing. In this case, convolution processing is performed using a basis matrix (a matrix whose elements are only basis values ​​(integer values)) and data D3 output from the quantized data memory unit 3, and then processing is performed using the real coefficient vector. This makes it possible to perform most of the convolution calculations as integer calculations, and further to achieve faster convolution processing.

[0112] In the above embodiment, the CNN data processing device 100 is described as using a kernel of a predetermined size (3 × 3) and a region to be subjected to convolution processing having a size of 4 × 8. However, the size of the kernel and the size of the region to be subjected to convolution processing may be different sizes.

[0113] In the above embodiment, the CNN data processing device 100 has been described assuming that the CNN data processing is performed for depthwise convolution (convolution processing in the spatial direction), but this is not limitative. For example, the CNN data processing device 100 may apply the CNN data processing of the above embodiment to normal convolution processing.

[0114] In the above embodiment, the CNN data processing device 100 performs CNN data processing on data after quantization processing, but the present invention is not limited to this. For example, data (feature data) that has not been subjected to quantization processing may be input to the CNN data processing unit 2 of the CNN data processing device 100, and the CNN data processing may be performed on the data by the CNN data processing unit 2.

[0115] Furthermore, the configuration of the memory unit 22 of the CNN data processing device 100 is not limited to that described in the above embodiment, and the number of bank memories and the number of data that can be simultaneously accessed from each bank memory (number of access buses) can be set to any number.

[0116] Furthermore, each block (each functional unit) of the CNN data processing device 100 described in the above embodiment may be individually integrated into a single chip using a semiconductor device such as an LSI, or may be integrated into a single chip to include some or all of the blocks. Furthermore, each block (each functional unit) of the pose data generation system, CG data system, and pose data generation device described in the above embodiment may be realized by multiple semiconductor devices such as LSIs.

[0117] Although the term "LSI" is used here, it may also be called an IC, system LSI, super LSI, or ultra LSI depending on the degree of integration.

[0118] Furthermore, the method of integration is not limited to LSI, but may be realized by a dedicated circuit or a general-purpose processor. It is also possible to use a field programmable gate array (FPGA) that can be programmed after LSI manufacturing, or a reconfigurable processor that can reconfigure the connections and settings of circuit cells inside the LSI.

[0119] Furthermore, some or all of the processing of each functional block in each of the above embodiments may be realized by a program. And, some or all of the processing of each functional block in each of the above embodiments is performed by a central processing unit (CPU) in a computer. Furthermore, the programs for performing each processing are stored in a storage device such as a hard disk or ROM, and are read out and executed in the ROM or RAM.

[0120] Furthermore, each process in the above-described embodiment may be realized by hardware, or may be realized by software (including cases where it is realized together with an OS (operating system), middleware, or a predetermined library). Furthermore, it may be realized by a combination of software and hardware.

[0121] For example, when each functional unit of the above embodiment is realized by software, each functional unit may be realized by software processing using the hardware configuration shown in FIG. 22 (for example, a hardware configuration in which a CPU, GPU, ROM, RAM, input unit, output unit, etc. are connected via a bus).

[0122] Furthermore, when each functional unit of the above embodiment is realized by software, the software may be realized using a single computer having the hardware configuration shown in Figure 22, or may be realized by distributed processing using multiple computers.

[0123] Furthermore, the execution order of the processing method in the above embodiment is not necessarily limited to the description of the above embodiment, and the execution order can be changed within the scope of the gist of the invention. Furthermore, in the processing method in the above embodiment, some steps may be executed in parallel with other steps within the scope of the gist of the invention. Furthermore, in the processing method in the above embodiment, processes executed in parallel may be executed serially (sequentially).

[0124] The scope of the present invention includes a computer program for causing a computer to execute the above-described method and a computer-readable recording medium on which the program is recorded. Examples of computer-readable recording media include flexible disks, hard disks, CD-ROMs, MOs, DVDs, DVD-ROMs, DVD-RAMs, large-capacity DVDs, next-generation DVDs, and semiconductor memories.

[0125] The computer program is not limited to one recorded on the recording medium, but may be one transmitted via a telecommunications line, a wireless or wired communication line, a network such as the Internet, or the like.

[0126] The term "part" may also include the concept of "circuitry." A circuitry may be realized in whole or in part by hardware, software, or a combination of hardware and software.

[0127] The functions of the elements disclosed herein may be implemented using circuitry or processing circuitry, including general-purpose processors, special-purpose processors, integrated circuits, ASICs ("application-specific integrated circuits"), conventional circuitry, and / or combinations thereof, configured to perform the disclosed elements or programmed to perform the disclosed functions. A processor is considered to be processing circuitry or circuitry when it includes transistors and other circuitry therein. In this disclosure, a circuitry, unit, or means is hardware that performs the recited function or hardware programmed to perform the function. The hardware may be any hardware disclosed herein or other known hardware that is programmed to perform or configured to perform the recited function. When the hardware is a processor, which may be considered as a type of circuitry, the circuitry, means, or unit is a combination of hardware and software, software used to configure the hardware, and / or processor.

[0128] The specific configuration of the present invention is not limited to the above-described embodiment, and various changes and modifications are possible without departing from the gist of the invention.

[0129] 100 CNN data processing device 2 CNN data processing unit (convolution processing data device) 21 Memory access control unit (access control unit) 22 Memory unit Tmem_0 to Tmem_M-1 Bank memory Tmem_k 23 Register unit

Claims

1. A data processing device for convolution processing used in a convolutional neural network model, comprising: a plurality of bank memories for storing feature data; and an access control unit that controls data writing and / or data reading from the plurality of bank memories, wherein the feature data is three-dimensional data specified by a position in a width direction, a position in a height direction, and a position in a channel direction, and each of the plurality of bank memories has a plurality of access buses so as to enable parallel access to data, and the access control unit controls data writing so that the feature data having a first value in a height direction is stored in the bank memory assigned to the first value among the plurality of bank memories, and further stores a plurality of pieces of the feature data having the same position in the width direction and consecutive positions in the channel direction in a memory area at an address that can be accessed in parallel by the plurality of buses.

2. The data processing device for convolution processing according to claim 1, wherein the access control unit controls data reading from the memory for the multiple banks so as to read out multiple channels of the feature data that are at the same position in the width direction and consecutive positions in the height direction during a read unit period.

3. A convolution processing data processing device as described in claim 1 or 2, wherein a data group obtained by reading out from the multiple bank memories multiple channels of feature data having the same widthwise position and continuous heightwise positions is defined as a multiple channel h x 1 data group, the access control unit obtains the number of overlapping multiple channel h x 1 data groups in the same read unit period according to the position of the convolution processing target area as the number of output systems, and controls the multiple bank memories so that the multiple channel h x 1 data groups equal to the number of output systems obtained are output from the multiple bank memories.

4. A data processing device for convolution processing as described in claim 3, further comprising a register unit capable of storing data by address specification, said register unit inputting the multi-channel h x 1 data group output from the multiple bank memories, setting the width-wise size of the kernel of the convolution processing to be applied to the multi-channel h x 1 data group as an offset value, and sequentially writing feature data having consecutive positions in the height direction included in the multi-channel h x 1 data group into a memory area of ​​an address of said register unit offset by said offset value.

5. The data processing device for convolution processing according to claim 4, wherein after feature data of an area to be subjected to convolution processing by the kernel is stored in a memory area of ​​consecutive addresses in the register unit, the register unit outputs the feature data stored in the memory area of ​​consecutive addresses.

6. The convolution processing data processing device according to claim 5, wherein the register unit outputs the feature data stored in the memory area of ​​the consecutive addresses all at once or in the order of the consecutive addresses.

Citation Information

Patent Citations

  • Semiconductor devices and methods of manufacturing semiconductor devices

    KR1020230055576A

  • Dynamic processing element array expansion

    US20200410337A1