Memory-Embedded Device, Processing Method, Parameter Setting Method, and Image Sensor Device
The memory-embedded device addresses the inefficiencies in memory access for AI technologies by using a memory access controller to optimize data transfer according to parameter specifications, resulting in reduced CPU overhead and improved performance.
Patent Information
- Application Number
- JP2022527005
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-29
- Filing Date
- 2021-05-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-05-21
AI Technical Summary
Existing AI technologies, such as neural networks, face challenges in efficiently accessing memory due to the need for dedicated instructions for address calculation, leading to increased CPU overhead and power consumption.
A memory-embedded device comprising a processor, a memory access controller, and a memory, where the memory access controller reads and writes data used in convolutional operations to and from memory according to parameter specifications, optimizing memory access.
This solution enables efficient and appropriate memory access, reducing CPU overhead and power consumption by offloading memory access operations to the memory access controller, thereby improving performance in AI processing tasks.
Smart Images

Figure 0007697462000003 
Figure 0007697462000004 
Figure 0007697462000005
Abstract
Description
Technical Field
[0001] The present disclosure relates to a memory - embedded device, a processing method, a parameter setting method, and an image sensor device.
Background Art
[0002] In AI technologies such as neural networks, since a huge amount of calculations are performed, access to memory increases. For example, a technique for accessing an N - dimensional tensor has been provided (Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] According to the prior art, by preparing an instruction corresponding to the calculation (generation) of an address and dedicated hardware that performs only address calculation, a part of the processing is offloaded to the hardware.
[0005] However, in the above - mentioned prior art, the CPU needs to issue a dedicated instruction every time for address calculation, and there is room for improvement. Therefore, it is desired to enable appropriate access to memory.
[0006] Therefore, the present disclosure proposes a memory - embedded device, a processing method, a parameter setting method, and an image sensor device that can enable appropriate access to memory.
Means for Solving the Problems
[0007] To solve the above problems, a memory - embedded device according to one aspect of the present disclosure is a memory - embedded device including a processor, a memory access controller, and a memory accessed according to the processing of the memory access controller, wherein the memory access controller is configured to read and write data used in the operation of the convolutional operation circuit to and from the memory according to the specification of parameters.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17A
Figure 17B
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Modes for Carrying Out the Invention
[0009] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Note that the memory built-in device, processing method, parameter setting method, and image sensor device according to the present application are not limited by this embodiment. Also, in the following embodiments, the same reference numerals are given to the same parts to omit redundant explanations.
[0010] The present disclosure will be described in accordance with the item order shown below. 1. Embodiments 1-1. Outline of the processing system according to the embodiment of the present disclosure 1-2. Overall outline and problems 1-3. First embodiment 1-3-1. Variation 1-4. Second embodiment 1-4-1. Premises, etc. 2. Other embodiments 2-1. Other Configuration Examples (Image Sensor, etc.) 2-2. Others 3. Effects of the Present Disclosure
[0011] [1. Embodiment] [1-1. Overview of the Processing System According to the Embodiment of the Present Disclosure] FIG. 1 is a diagram showing an example of a processing system according to an embodiment of the present disclosure. As shown in FIG. 1, the processing system 10 includes a memory built-in device 20, a plurality of sensors 600, and a cloud system 700. Note that the processing system 10 shown in FIG. 1 may include a plurality of memory built-in devices 20 and a plurality of cloud systems 700.
[0012] The plurality of sensors 600 include various sensors such as an image sensor 600a, a microphone 600b, an acceleration sensor 600c, and other sensors 600d. When the image sensor 600a, the microphone 600b, the acceleration sensor 600c, the other sensors 600d, etc. are described without particular distinction, they are referred to as "sensors 600". The sensors 600 may have various sensors not limited to the above, such as a position sensor, a temperature sensor, a humidity sensor, an illuminance sensor, a pressure sensor, a proximity sensor, and a sensor for detecting biological information such as odor, sweat, heartbeat, pulse, and brain wave. For example, each sensor 600 transmits the detected data to the memory built-in device 20.
[0013] The cloud system 700 includes a server device (computer) used to provide cloud services. The cloud system 700 communicates with the memory built-in device 20 and transmits and receives information to and from the remote memory built-in device 20.
[0014] The memory - embedded device 20 is communicably connected, either wired or wirelessly, via a communication network (e.g., the Internet) to the sensor 600 and the cloud system 700. The memory - embedded device 20 has a communication processor (network processor), and communicates with external devices such as the sensor 600 and the cloud system 700 via the communication network by means of the communication processor. The memory - embedded device 20 transmits and receives information to and from the sensor 600, the cloud system 700, etc. via the communication network. Also, the memory - embedded device 20 and the sensor 600 may communicate by means of a wireless communication function such as Wi - Fi (registered trademark) (Wireless Fidelity), Bluetooth (registered trademark), LTE (Long Term Evolution), 5G (5th - generation mobile communication system), LPWA (Low Power Wide Area), etc.
[0015] The memory - embedded device 20 includes an arithmetic unit 100 and a memory 500.
[0016] The arithmetic unit 100 is a computer (information - processing device) that executes arithmetic processing related to machine learning. For example, the arithmetic unit 100 is used for the calculation of the functions of artificial intelligence (AI: Artificial Intelligence). The functions of artificial intelligence are, for example, functions such as learning based on learning data, and inference, recognition, classification, data generation, etc. based on input data, but are not limited thereto. Also, the functions of artificial intelligence use a deep neural network. That is, in the example of FIG. 1, the processing system 10 is an artificial - intelligence system (AI system) that performs processing related to artificial intelligence. The memory - embedded device 20 performs DNN (Deep Neural Network) processing on inputs from a plurality of sensors 600.
[0017] The arithmetic unit 100 includes a plurality of processors 101, a plurality of first - cache memories 200, a plurality of second - cache memories 300, and a third - cache memory 400.
[0018] The plurality of processors 101 includes processors 101a, 101b, 101c, etc. When explaining processors 101a to 101c, etc. without particular distinction, they are described as "processor 101". In the example of FIG. 1, three processors 101 are illustrated, but the number of processors 101 may be four or more, or less than three.
[0019] The processor 101 may be various processors such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit). Note that the processor 101 is not limited to a CPU or a GPU, and may have any configuration as long as it is applicable to arithmetic processing. In the example of FIG. 1, the processor 101 includes a convolution operation circuit 102 and a memory access controller 103. The convolution operation circuit 102 performs a Convolution operation. The memory access controller 103 is used for accessing the first cache memory 200, the second cache memory 300, the third cache memory 400, and the memory 500, the details of which will be described later. Also, the processor including the convolution operation circuit 102 may be a neural network accelerator. The neural network accelerator is suitable for efficiently processing the above-described artificial intelligence functions.
[0020] The plurality of first cache memories 200 include a first cache memory 200a, a first cache memory 200b, a first cache memory 200c, and the like. The first cache memory 200a corresponds to the processor 101a, the first cache memory 200b corresponds to the processor 101b, and the first cache memory 200c corresponds to the processor 101c. For example, the first cache memory 200a transmits corresponding data to the processor 101a in response to a request from the processor 101a. When the first cache memories 200a to 200c and the like are described without particular distinction, they are referred to as "the first cache memory 200". In the example of FIG. 1, three first cache memories 200 are illustrated, but the number of the first cache memories 200 may be four or more, or may be less than three. For example, the first cache memory 200 has SRAM (Static Random Access Memory), but the first cache memory 200 is not limited to SRAM and may have a memory other than SRAM.
[0021] The plurality of second cache memories 300 include a second cache memory 300a, a second cache memory 300b, a second cache memory 300c, and the like. The second cache memory 300a corresponds to the processor 101a, the second cache memory 300b corresponds to the processor 101b, and the second cache memory 300c corresponds to the processor 101c. For example, when the data requested from the processor 101a is not in the first cache memory 200a, the second cache memory 300a transmits corresponding data to the first cache memory 200a. When the second cache memories 300a to 300c and the like are described without particular distinction, they are referred to as "the second cache memory 300". In the example of FIG. 1, three second cache memories 300 are illustrated, but the number of the second cache memories 300 may be four or more, or may be less than three. For example, the second cache memory 300 has SRAM, but the second cache memory 300 is not limited to SRAM and may have a memory other than SRAM.
[0022] The third cache memory 400 is the cache memory that is farthest from the processor 101, that is, the LLC (Last Level Cache). The third cache memory 400 is commonly used by the processors 101a to 101c and the like. For example, when the data requested from the processor 101a is not in the first cache memory 200a and the second cache memory 300a, the third cache memory 400 transmits the corresponding data to the second cache memory 300a. For example, the third cache memory 400 has SRAM, but the third cache memory 400 is not limited to SRAM and may have a memory other than SRAM.
[0023] The memory 500 is a storage device provided outside the arithmetic unit 100. For example, the memory 500 is connected to the arithmetic unit 100 by a bus or the like and transmits and receives information to and from the arithmetic unit 100. In the example of FIG. 1, the memory 500 has a DRAM (Dynamic Random Access Memory) or a flash memory (Flash Memory). Note that the memory 500 is not limited to DRAM or flash memory and may have a memory other than DRAM or flash memory. For example, when the data requested from the processor 101a is not in the first cache memory 200a, the second cache memory 300a, and the third cache memory 400, the memory 500 transmits the corresponding data to the third cache memory 400.
[0024] Here, the memory hierarchy of the processing system 10 shown in FIG. 1 will be described with reference to FIG. 2. FIG. 2 is a diagram showing an example of the memory hierarchy. Specifically, FIG. 2 is a diagram showing an example of the hierarchy of off-chip memory and on-chip memory. In FIG. 2, an example is shown in which the processor 101 is a CPU and the memory 500 is a DRAM.
[0025] As shown in FIG. 2, the first cache memory 200, the second cache memory 300, and the third cache memory 400 are on-chip memories. Also, the memory 500 is an off-chip memory.
[0026] As shown in FIG. 2, a cache memory is often used as a memory close to an arithmetic unit such as the processor 101. The cache memory has a hierarchical structure as shown in FIG. 2. In the example of FIG. 2, the first cache memory 200 is the first-level cache memory (L1 Cache) closest to the processor 101. The second cache memory 300 is the second-level cache memory (L2 Cache) next closest to the processor 101 after the first cache memory 200. The third cache memory 400 is the third-level cache memory (L3 Cache) next closest to the processor 101 after the second cache memory 300.
[0027] For example, the higher the cache memory, the faster it is but the smaller the memory capacity. Therefore, by swapping unnecessary data and necessary data, access to large-sized data is realized. The overall outline and the like will be described below.
[0028] [1-2. Overall Outline and Problems] Next, the overall outline and problems will be described with reference to FIGS. 3 to 8. First, the convolutional operation will be described with reference to FIG. 3. FIG. 3 is a diagram showing an example of the dimensions used in the convolutional operation. As shown in FIG. 3, for example, data handled by a CNN (Convolutional Neural Network) has up to four dimensions. An explanation of the dimensions and examples of their uses are shown in Table 1. FIG. 3 conceptually shows Table 1. Table 1 shows the four dimensions used in the convolutional operation. Note that although Table 1 shows five parameters, when focusing on individual data (e.g., input-feature-map, etc.), the maximum number of dimensions is four.
[0029]
Table 1
[0030] As shown in Table 1, the parameter "W" corresponds to the width of the Input-feature-map (input feature map). For example, the parameter "W" corresponds to one-dimensional data such as a microphone or an action / environment / acceleration sensor (such as acceleration sensor 600c, etc.). Hereinafter, the parameter "W" is also referred to as the "first parameter".
[0031] The feature map after the convolution operation using the Input-feature-map is shown as the Output-feature-map (output feature map). The parameter "X" corresponds to the width of the feature map (Output-feature-map) after the convolution operation. The parameter "X" corresponds to the parameter "W" of the next layer. When distinguishing the parameter "X" from the parameter "W", the parameter "X" may be referred to as the "first parameter after the operation". Also, the parameter "W" may be referred to as the "first parameter before the operation".
[0032] Also, the parameter "H" corresponds to the height of the Input-feature-map. For example, the parameter "H" corresponds to the data of the second dimension of an image sensor (such as image sensor 600a, etc.). Hereinafter, the parameter "H" is also referred to as the "second parameter".
[0033] The parameter "Y" corresponds to the height of the feature map (Output-feature-map) after the convolution operation. The parameter "Y" corresponds to the parameter "H" of the next layer. When distinguishing the parameter "Y" from the parameter "H", the parameter "Y" may be referred to as the "second parameter after the operation". Also, the parameter "H" may be referred to as the "second parameter before the operation".
[0034] Also, the parameter "C" corresponds to the number of channels of the Input-feature-map, the number of channels of the Weight, and the number of channels of the Bias. For example, when the parameter "C" is used to perform convolution on the R, G, and B directions of an image, or when performing convolution processing on one-dimensional data from multiple sensors, etc., the dimension of the sum of convolutions (convolution) is increased by one and defined as a channel. Hereinafter, the parameter "C" is also referred to as the "third parameter".
[0035] Also, the parameter "M" corresponds to the number of channels of the Output-feature-map, the number of batches of the Weight, and the number of batches of the Bias. For example, the parameter "M" is used in this dimension to adapt the above-mentioned channel concept between the layers of the CNN. The parameter "M" corresponds to the parameter "C" of the next layer. Hereinafter, the parameter "M" is also referred to as the "fourth parameter".
[0036] Also, the parameter "N" corresponds to the number of batches of the Input-feature-map and the number of batches of the Output-feature-map. For example, when processing multiple sets of input data in parallel using the same coefficient, the parameter "N" defines this set direction as another dimension. Hereinafter, the parameter "N" is also referred to as the "fifth parameter".
[0037] Here, the convolution process that performs the convolution operation will be described with reference to FIG. 4. FIG. 4 is a conceptual diagram showing the convolution process. For example, the main elements that make up a neural network are the convolution layer and the fully connected layer, and in these layers, the sum of products (operation) of elements of a high-dimensional tensor such as a four-dimensional tensor is performed. For example, as shown in "Sum of products operation: o = i * w + p" in FIG. 4, the sum of products operation includes operations such as the product of the input data i and the weight w, and the sum of the result of the product and the intermediate result p of the operation to calculate the output data o.
[0038] The sum-of-products operation for one degree generates a total of four memory accesses, which are three loads (reads) of data and one store (write) of data. For example, in the convolution process example shown in FIG. 4, 4HWK 2 performs CM times of sum-of-products operations. Therefore, 4HWK 2 generates CM times of memory accesses. For example, even in a relatively small network for mobile terminals, since H, W are from 10 to 200, K is from 1 to 7, C is from 3 to 1000, M is from 32 to 1000, etc., the number of memory accesses reaches from tens of thousands to hundreds of billions of times.
[0039] Also, generally, memory access consumes more power than the operation itself. For example, when accessing off-chip memory such as DRAM, it requires hundreds of times the power of the operation. Therefore, by reducing off-chip memory access and accessing memory closer to the arithmetic unit, the power consumption can be reduced. Thus, reducing this off-chip memory access becomes a major issue.
[0040] The sum-of-products of the above-mentioned tensor elements has high data reusability because access to the same data occurs frequently. Especially when performing convolution operations, this tendency becomes prominent. When using a cache memory configured in a general set-associative manner, depending on the shape of the tensor used in the operation, the memory utilization efficiency may be impaired. For example, when only a part of the memory is used during the operation as shown in FIG. 5, the memory utilization efficiency may be significantly impaired. FIG. 5 is a diagram showing an example of storing tensor data in the cache memory. Also, since it is only known at runtime where on the memory the data is placed, it is difficult to perform program optimization.
[0041] Therefore, as a technique for reducing accesses to off-chip memory without using cache memory, a method with an internal buffer can also be considered. Since the data loaded from DRAM is directly transferred to the internal buffer, by optimizing the utilization of the internal buffer, it becomes possible to reduce the access frequency to DRAM. However, the interface between the internal buffer and DRAM needs to communicate with each other using the data address. An example of this is shown in FIG. 6. FIG. 6 is a diagram showing an example of a convolution operation program and its abstraction.
[0042] Also, FIG. 7 shows the address calculation when accessing 4D tensor data. FIG. 7 is a diagram showing an example of the address calculation when accessing the elements of a tensor. As described above, in order to convert from index information such as i, j, k, l to an address, it is necessary to perform 6 multiplications and 3 additions only using the part of the index information. Therefore, in the case of accessing 4D data, a large number of instructions are required to access one element.
[0043] As described above, by preparing an instruction corresponding to the address calculation and dedicated hardware that only performs the address calculation, and offloading the address calculation to the hardware, it is possible to improve performance and suppress power consumption. However, the multiplications and additions for the address calculation must be performed every time there is an access. Therefore, for example, when performing a task that requires a high-dimensional tensor product, the optimization and efficient use of cache memory, and the memory configuration that suppresses the increase in the address calculation itself will be described in the following first embodiment.
[0044] [1-3. First Embodiment] Next, the first embodiment will be described with reference to FIGS. 8 to 16. First, the outline of the first embodiment will be described with reference to FIG. 8. FIG. 8 is a conceptual diagram according to the first embodiment. In FIG. 8, the first cache memory 200 is taken as an example for explanation. However, the memory is not limited to the first cache memory 200, and may be applied to various memories such as the second cache memory 300, the third cache memory 400, and the memory 500. In the following examples, access to 4D data is taken as an example, but access to lower-dimensional data and access to higher-dimensional data according to the resources of the hardware are allowed.
[0045] The first cache memory 200 shown in FIG. 8 is a type of cache memory. Instead of accessing data by address like a conventional cache memory, access is performed using the index information of the tensor to be accessed. The first cache memory 200 shown in FIG. 8 has a plurality of partial cache memory areas 201, and an example of accessing using index information such as idx1, idx2, idx3, idx4, etc. is shown.
[0046] In FIG. 8, an example is shown in which, in access by index information, when the corresponding data is not on the cache memory (the first cache memory 200), access is performed to a lower-level memory (for example, the memory 500) using an address. When multiple cache memories are hierarchically used as shown in FIG. 1, the index information is passed to a lower-level memory to search for the corresponding data.
[0047] In this case, when the corresponding data is not in the first cache memory 200 during access by index information, the index information is passed to the cache memory immediately below the first cache memory (second cache memory 300), and the corresponding data is searched for in the second cache memory 300. Also, when the corresponding data is not in the second cache memory 300 during access by index information, the index information is passed to the cache memory immediately below the second cache memory (third cache memory 400), and the corresponding data is searched for in the third cache memory 400. Further, when the corresponding data is not in the third cache memory 400 during access by index information, access is performed on the memory 500 using the address.
[0048] Hereinafter, a specific example will be described with reference to FIGS. 9 and 10. FIGS. 9 and 10 are diagrams showing an example of the processing according to the first embodiment. In this embodiment, the first cache memory 200 is taken as a representative example of the cache memory according to the present invention and is referred to as the cache memory 200. Also, in this embodiment, the partial cache memory area 201 is referred to as a tile.
[0049] First, in FIG. 9, register 111 is a register that holds the configuration information of the cache memory. For example, memory built-in device 20 has register 111. Register 111 holds information indicating that one tile is composed of cache line 202set*way (pieces), and the entire cache is composed of M*N (pieces) of tiles. In this embodiment, the values way, set, N, and M correspond to dimension1, dimension2, dimension3, and dimension4 in FIG. 8, respectively. For example, these values may be fixed values when the cache memory is configured. For example, the value M of register 111 is used in the example of FIG. 9 for memory built-in device 20 to select only one tile from the tiles in one direction (e.g., the height direction) by taking the remainder of dividing index information idx4 by value M. Similarly, regarding the values set and N, they are used for the selection of the set and the selection of the tile, respectively. Note that since way is not used during memory access, it does not have to be held in register 111. Note that a set refers to a plurality (two or more) of cache lines arranged continuously in the width direction in one tile, and a way refers to a plurality (two or more) of cache lines arranged continuously in the height direction in one tile.
[0050] The cache line 202 shown in FIG. 9 represents the minimum unit of data. For example, similar to a normal cache memory, cache line 202 is composed of a header information part for determining whether the data is the desired one and a data information part for storing the actual data. The header information of cache line 202 includes information corresponding to a tag such as index information for specifying the data and information for selecting the replacement target. Note that any configuration is allowed for the information used in the header and the assignment method of the information.
[0051] In FIG. 9, the cache memory 200 represents the entire cache memory, includes a plurality of partial cache memory areas 201, and as described above, this partial cache memory area 201 is referred to as a tile. Also, a tile is one that has a plurality (two or more) of cache lines 202, and the cache memory 200 is one that includes a plurality (two or more) of tiles. That is, in the cache memory 200 of FIG. 9, each of the rectangular areas indicated by the height set and the width way corresponds to a partial cache memory area 201 called a tile. That is, in the example of FIG. 9, a total of 16 tiles in 4 in the height direction × 4 in the width direction are shown.
[0052] In FIG. 9, the selector 112 is used to select which one of the M tiles (for example, tiles in the height direction) arranged in the first direction of the cache memory 200 is to be used. For example, the selector 112 uses the remainder obtained by dividing the index information idx4 shown in FIG. 8 by the value M to select which one of the M tiles is to be used. For example, the memory built-in device 20 has the selector 112.
[0053] In FIG. 9, the selector 113 selects which one of the N tiles (for example, tiles in the width direction) arranged in a second direction different from the first direction of the cache memory 200 is to be used. For example, the selector 113 uses the remainder obtained by dividing the index information idx3 shown in FIG. 8 by the value N to select which one of the N tiles is to be used. For example, the memory built-in device 20 has the selector 113. One tile out of a plurality of tiles in the cache memory 200 is selected by the selector 112 and the selector 113.
[0054] In FIG. 9, in the tile selected by the combination of selector 112 and selector 113, selector 114 selects which set to use. For example, selector 114 selects which set of tiles to use by using the remainder obtained by dividing the index information idx2 shown in FIG. 8 by the value set. For example, the memory built-in device 20 has selector 114.
[0055] In FIG. 9, comparator 115 is used to compare the header information of all way cache lines 202 in the set selected by selector 112, selector 113, and selector 114 with index information idx1 to idx4 and the like. That is, it is a circuit that determines a so-called cache hit (whether data exists in cache memory 200). Comparator 115 compares the header information of all way cache lines 202 in the set with index information idx1 to idx4 and the like. Then, as a result of the comparison, comparator 115 outputs "hit (data exists)" when there is a match, and "miss (data does not exist)" when there is no match. That is, comparator 115 determines whether the desired data exists in the lines within the set and generates a hit or miss signal. For example, the memory built-in device 20 has comparator 115.
[0056] In FIG. 10, register 116 is a register that holds the start address (base addr) of the tensor to be accessed, the size of dimension 1 (size1), the size of dimension 2 (size2), the size of dimension 3 (size3), the size of dimension 4 (size4), and the data size (datasize) of the tensor. For example, the memory built-in device 20 has register 116.
[0057] When information indicating a cache miss (value miss) is output from the comparator 115 in FIG. 9, the address generation logic 117 generates an address using the information in the register 116 and the index information idx1 to idx4. For example, the memory built-in device 20 has the address generation logic 117. The memory access controller 103 may have the function of the address generation logic 117. The address calculation formula is represented by the following formula (1).
[0058]
Number
[0059] In formula (1), datasize is the data size (e.g., number of bytes) indicated in the register 116, and becomes a numerical value such as "4" for float (e.g., 4-byte single-precision floating-point real number) and 2 for short (e.g., 2-byte signed integer). Regarding the address calculation by the address generation logic 117, any configuration is allowed as long as an address can be generated from the index information.
[0060] Next, with reference to FIG. 11, the processing procedure according to the first embodiment will be described. FIG. 11 is a flowchart showing the processing procedure according to the first embodiment. In the example of FIG. 11, the arithmetic unit 100 is described as the main body of the processing, but the main body of the processing may be read as the first cache memory 200, the memory built-in device 20, etc. according to the content of the processing.
[0061] As shown in FIG. 11, the arithmetic unit 100 sets base addr (step S101). The arithmetic unit 100 sets the base addr shown in the register 116 of FIG. 10.
[0062] The arithmetic unit 100 sets size1 (step S102). The arithmetic unit 100 sets the size1 shown in the register 116 of FIG. 10.
[0063] The arithmetic unit 100 sets sizeN (step S103). The arithmetic unit 100 sets sizeN shown in the register 116 of FIG. 10. Note that "N" in sizeN is an arbitrary value. Although only steps S102 and S103 are illustrated in FIG. 11, the size setting is performed for the number of sizes (number of dimensions). For example, in the example of FIG. 10, "N" in sizeN is "4", and the arithmetic unit 100 sets each of size1, size2, size3, and size4.
[0064] The arithmetic unit 100 sets datasize (step S104). The arithmetic unit 100 sets datasize shown in the register 116 of FIG. 10.
[0065] The arithmetic unit 100 waits for cache access (step S105). Then, the arithmetic unit 100 specifies the set using set, N, and M (step S106).
[0066] When the cache hits (step S107: Yes) and the process is a read (step S108: Yes), the arithmetic unit 100 transfers the data (step S109). For example, when the cache hits (when the corresponding data is in the first cache memory 200) and the process is a read, the first cache memory 200 transfers the data to the processor 101.
[0067] Also, when the cache hits (step S107: Yes) and the process is not a read (step S108: No), the arithmetic unit 100 writes the data (step S110). For example, when the cache hits (when the corresponding data is in the first cache memory 200) and the process is not a read but a write, the first cache memory 200 writes the data.
[0068] Then, the arithmetic unit 100 updates the header information (step S111), returns to step S105, and repeats the process.
[0069] When the cache misses (step S107: No), the arithmetic unit 100 calculates an address (step S112). Then, the arithmetic unit 100 requests access to the lower-level memory (step S113). For example, when the cache misses (when the corresponding data is not in the first cache memory 200), the arithmetic unit 100 generates an address and requests access to the memory 500.
[0070] When the initial reference is not a miss (step S114: No), the arithmetic unit 100 selects a replacement target (step S115) and determines an insertion position (step S116). When the initial reference is a miss (step S114: Yes), the arithmetic unit 100 determines an insertion position (step S116).
[0071] Then, after waiting for the data (step S117), the arithmetic unit 100 writes the data (step S118). Then, the processing after step S108 is performed.
[0072] With the configurations and processes of FIGS. 9 to 11 above, since it appears as the memory in FIG. 8 to software developers, the memory built-in device 20 can easily optimize tasks that require access to tensor data. Also, by increasing the cache hit rate through such optimization, the memory built-in device 20 can reduce the number of processes corresponding to address calculation.
[0073] When modifying the process, after performing "set datasize" in step S104, desired information is written to a register, and the part of "specify the set using set, N, M" in step S106 is changed to a process using additional information.
[0074] Here, a specific example of tensor access will be described with reference to FIG. 12. FIG. 12 is a diagram showing an example of memory access according to the first embodiment. In FIG. 12, the index information idx1 to idx4 connected to the comparator 122 and the address generation logic 123 are omitted, and the description will be made from the state after the initialization of each register is completed.
[0075] The access example in FIG. 12 is an access to the 4D tensor v of the program PG1 in the upper left of FIG. 12, and it is assumed that the access to v[0][1][1][1] in FIG. 12 is at the timing of a miss.
[0076] First, as shown in FIG. 12, the index information 0, 1, 1, and 1 of V[0][1][1][1] are respectively set to idx1 to idx4, and the memory is accessed using the index information idx1 to idx4. In this case, the access using the index information is performed by the following unique instruction or a dedicated accelerator. (Instruction) ld idx4,idx3,idx2,idx1 st idx4,idx3,idx2,idx1
[0077] Next, as shown in FIG. 12, for each value of the index information idx2 to idx4, the corresponding set is selected using the remainders (residues) obtained by dividing the values by the values set, N, and M. In the example of FIG. 12, the selector selects the corresponding set using the index information of idx2 = 1, idx3 = 1, idx4 = 1, and the information of the register 121 with set = 4, N = 1, and M = 1. For example, the memory built-in device 20 has the register 121.
[0078] Next, as shown in FIG. 12, the header information of all cache lines in the set and the index information idx1 to idx4 are input to the comparator 122, and a cache miss is determined. The comparator 122 is a circuit having the same function as the comparator 115 in FIG. 9.
[0079] Next, as shown in FIG. 12, address generation logic 123 calculates an address using index information idx1 to idx4, base addr, various sizes (size1 to size4), and datasize information. The address generation logic 123 is the same as the address generation logic 117 in FIG. 10.
[0080] Next, as shown in FIG. 12, the memory built-in device 20 accesses the DRAM (for example, memory 500) at the calculated address. Note that the symbols i, j, k, and l in the DRAM correspond to the symbols used in the program PG1 in FIG. 12 and are described for the purpose of explanation corresponding to the program PG1. Actually, in order to access the DRAM, it is calculated using index information idx1 to idx4, base addr, various sizes (size1 to size4), and datasize information.
[0081] Finally, as shown in FIG. 12, data is inserted from the DRAM into the cache memory (such as the first cache memory 200).
[0082] [1-3-1. Modification Example] Here, a modification example according to the first embodiment will be described with reference to FIG. 13. FIG. 13 is a diagram showing a modification example according to the first embodiment. FIG. 13 shows an example in which the cache memory is composed of only sets and ways without using tiles. Note that in FIG. 13, only the differences from FIGS. 9 and 10 are shown, and the same points are omitted from the description as appropriate.
[0083] In FIG. 13, register 131 is a register that holds the allocation information of the cache memory to be used. For example, the memory built-in device 20 has register 131. The value msize1 indicates how many cache lines in the way direction are grouped together, and the value msize2 indicates how many groups (also called blocks) of msize1 cache lines there are in the way direction. Also, the value msize3 represents how many sets in the set direction are grouped together, and the value msize4 indicates how many groups of msize3 cache lines there are in the set direction. In this case, msize2 = way / msize1 and msize4 = set / msize3. Also, since msize1 is information not used during memory access, only msize2 is retained and msize1 does not need to be retained in register 131.
[0084] In FIG. 13, cache memory 200 is a memory composed of a set of set * way cache lines, similar to a normal cache memory.
[0085] In FIG. 13, selector 132 selects a group of msize3 cache lines using the remainder value obtained by dividing the index information corresponding to index information idx4 in FIG. 8 by the value msize4. That is, selector 132 selects which group to use in one direction (for example, the height direction). For example, the memory built-in device 20 has selector 132.
[0086] In FIG. 13, selector 133 selects a group of msize1 cache lines using the remainder value obtained by dividing the index information corresponding to index information idx2 in FIG. 8 by the value msize2. That is, selector 133 selects which group to use in another direction (for example, the width direction). For example, the memory built-in device 20 has selector 133.
[0087] In FIG. 13, from the group selected by selector 132, which set to use is selected by the remainder value obtained by dividing the index information corresponding to index information idx3 in FIG. 8 by value msize3. That is, selector 134 uses the remainder obtained by dividing index information idx3 by value msize3 to select which set among the groups to use. For example, the memory built-in device 20 has selector 134.
[0088] Here, the cache line will be described with reference to FIG. 14. FIG. 14 is a diagram showing an example of the configuration of a cache line. FIG. 14 shows an example of the configuration when the cache line 202 contains data of a plurality of words. In the example of FIG. 14, the case where 4-word data is stored in one line is shown. In the case of being used for cache hit determination of hit or miss, idx1, which is the index information with the lowest dimension, is stored after discarding the lower 2 bits.
[0089] Cache hit determination in the case of configuring the cache line 202 as shown in FIG. 14 is performed by the hardware configuration as shown in FIG. 15. FIG. 15 is a diagram showing an example of hit determination regarding the cache line. Specifically, FIG. 15 is a diagram showing an example of cache hit determination when there are a plurality of words in the cache line. For example, among v[i][j][k][l], i is compared with idx4, j is compared with idx3, k is compared with idx2, and l is compared with idx1 after being shifted 2 bits to the right (discarding the lower 2 bits).
[0090] Next, the initial settings in the case of performing CNN processing will be described with reference to FIG. 16. FIG. 16 is a diagram showing an example of the initial settings in the case of performing CNN processing. FIG. 16 shows four initial settings for input, weight, bias, and output.
[0091] For example, one cache memory is used for each tensor, and information about each dimension is written to the setting register for each of them. For example, in the case of the input-feature-map, in FIG. 16, the size in the 1D direction is W, the size in the 2D direction is H, the size in the 3D direction is C, and the size in the 4D direction is N. Therefore, the memory built-in device 20 writes W to size1, H to size2, C to size3, and N to size4 respectively. In this way, the memory built-in device 20 designates the first parameter regarding the first dimension of the data, the second parameter regarding the second dimension of the data, the third parameter regarding the third dimension of the data, and the fifth parameter regarding the number of data. Also, appropriate values are designated for base addr and datasize.
[0092] As described above, in the first embodiment, the memory built-in device 20 is a type of cache memory, and constitutes a memory such as the first cache memory 200 as a cache memory specialized for accessing tensors. In this case, unlike a normal cache memory, the memory built-in device 20 can control access using the index information of the tensor to be accessed instead of the address. Also, the cache configuration is made to match the shape of the tensor. Further, the memory built-in device 20 includes an address generator (such as the address generation logic 117) in order to be compatible with a general memory that requires access by address. Thereby, the memory built-in device 20 can enable appropriate access to the memory. The memory built-in device 20 can change the correspondence with the address of the cache memory according to the designation of parameters. The memory built-in device 20 can change the address space of the cache memory according to the designation of parameters. That is, the memory built-in device 20 can set parameters to change the address space of the cache memory. The memory built-in device 20 can deform the address space of the cache memory according to the designation of parameters.
[0093] In the first embodiment, since the memory - built - in device 20 has the above - described configuration, for software developers, the access to tensors and their arrangement in the memory match, which facilitates the generation of more optimal code and enables the use of the memory without waste. Also, since the memory - built - in device 20 performs address generation only when data does not exist in the cache memory, the cost associated with address generation can be reduced.
[0094] [1 - 4. Second Embodiment] Next, a second embodiment will be described. Hereinafter, the memory - built - in device 20A will be described as an example, but the memory - built - in device 20A may have the same configuration as the memory - built - in device 20.
[0095] [1 - 4 - 1. Premises, etc.] First, prior to the description of the second embodiment, the premises and the like related to the second embodiment will be described.
[0096] The configuration of the convolution operation circuit as described above is fixed. For example, once the hardware (such as a semiconductor chip) of the data path including a data buffer and a multiplier - accumulator (MAC) is completed, it cannot be changed. On the other hand, software determines the data arrangement according to the pre - and post - processing offloaded to the CNN operation circuit. This is because it can optimize the efficiency of software development and the scale of software. Also, instead of software, hardware such as a sensor may directly store the data of the CNN operation in the memory. At this time, the sensor stores the data in a fixed arrangement based on its own hardware specifications in the memory. Thus, the operation circuit needs to efficiently access the data placed by software or a sensor without considering the configuration of the operation circuit.
[0097] However, if the data access order of the arithmetic circuit is fixed, there is a problem that efficient access cannot be achieved. For example, in circuit configuration X that can perform multiply-accumulate operations (MAC operations) on three 8-bit pixels simultaneously (in one cycle), when performing convolution processing on an RGB image, it is most efficient in terms of the number of cycles to perform convolution on the R channel first, then the G channel, and finally the B channel. Therefore, layout A (see, for example, FIGS. 21 and 23) that can read consecutive pixels of each channel in order is optimal. On the other hand, in the case of circuit configuration Y with three circuits that perform multiply-accumulate operations on one pixel per cycle, the arrangement of layout B that can read one pixel each of R, G, and B is preferable. However, due to the software or sensor specifications described above, when the combination of circuit configuration X and layout B is used, if the data access order of the arithmetic circuit is fixed, it takes extra cycles to read data from the memory, or the array of arithmetic units cannot be fully utilized, resulting in an overall increase in the number of cycles.
[0098] As methods for solving this problem, there are methods such as a first method in which software rearranges the arrangement on the memory before the CNN task, a second method in which a part of the loop process is offloaded to hardware, and a third method in which software calculates the address. However, the first method has problems such as high computational cost and poor memory usage efficiency because it has two types of data copies. Also, the second method has a problem of high computational cost because the loop process is performed by the processor's instructions. Additionally, the third method has a problem that the address calculation cost increases. Therefore, a configuration that enables appropriate access to the memory will be described in the following second embodiment.
[0099] Hereinafter, the configuration and processing of the second embodiment will be specifically described with reference to FIGS. 17A to 23. First, the outline of the second embodiment will be described with reference to FIGS. 17A and 17B. FIGS. 17A and 17B are diagrams showing an example of address generation according to the second embodiment. Hereinafter, when FIGS. 17A and 17B are not distinguished and described, they may be referred to as FIG. 17.
[0100] FIG. 17 shows a case where an address is generated using a dimension #0 counter 150, a dimension #1 counter 151, a dimension #2 counter 152, a dimension #3 counter 153, and an address calculation unit 160. For example, the memory built-in device 20A makes a memory access request using the address generated by the address calculation unit 160 using the count values of each of the dimension #0 counter 150, dimension #1 counter 151, dimension #2 counter 152, and dimension #3 counter 153. For example, the address calculation unit 160 may be an arithmetic circuit that takes as input the count (value) of each of the dimension #0 counter 150, dimension #1 counter 151, dimension #2 counter 152, and dimension #3 counter 153, calculates the address corresponding to the input, and outputs the calculated address. Hereinafter, the dimension #0 counter 150, dimension #1 counter 151, dimension #2 counter 152, dimension #3 counter 153, and the address calculation unit 160 may be collectively referred to as an "address generator".
[0101] FIG. 17A shows a case where a clock pulse is input to the dimension #0 counter 150 and the dimension #0 counter 150, dimension #1 counter 151, dimension #2 counter 152, and dimension #3 counter 153 are connected in this order. Specifically, it is connected such that the carry-over pulse signal of the dimension #0 counter 150 is input to the dimension #1 counter 151, the carry-over pulse signal of the dimension #1 counter 151 is input to the dimension #2 counter 152, and the carry-over pulse signal of the dimension #2 counter 152 is input to the dimension #3 counter 153.
[0102] Further, FIG. 17B shows a case where a clock pulse is input to the dimension #3 counter 153 and the dimension #3 counter 153, dimension #0 counter 150, dimension #1 counter 151, and dimension #2 counter 152 are connected in this order. Specifically, it is connected such that the carry-over pulse signal of the dimension #3 counter 153 is input to the dimension #0 counter 150, the carry-over pulse signal of the dimension #0 counter 150 is input to the dimension #1 counter 151, and the carry-over pulse signal of the dimension #1 counter 151 is input to the dimension #2 counter 152.
[0103] As shown in FIG. 17, indices of multiple dimensions are calculated by a counter, and the connection of the carry-over pulse signals of multiple counters can be freely changed as desired. The memory built-in device 20A calculates an address from multiple indices (counter values) and a multiplier of a preset dimension (distance between dimensions).
[0104] An example of the memory access controller 103 is shown in FIG. 18. FIG. 18 is a diagram showing an example of the memory access controller. The memory built-in device 20A shown in FIG. 18 includes a processor 101 and an arithmetic circuit 180. Thus, in FIG. 18, the memory access controller 103 is included in the arithmetic circuit 180. In the example of FIG. 18, although the memory access controller 103 is shown outside the processor 101, the memory access controller 103 may be included in the processor 101. The arithmetic circuit 180 may be integrated with the processor 101.
[0105] The arithmetic circuit 180 shown in FIG. 18 includes, in addition to the memory access controller 103, a control register 181, a temporary buffer 182, a MAC array 183, and the like. The control register 181 is a register included in the arithmetic circuit 180. For example, the control register 181 is a register (control device) used for receiving instructions read from a storage device (memory system) such as the memory 500 via the memory access controller 103 or temporarily storing instructions for execution. The temporary buffer 182 is a buffer included in the arithmetic circuit 180. For example, the temporary buffer 182 is a storage device or storage area for temporarily storing data. The MAC array 183 is a MAC (multiply-accumulate unit) array included in the arithmetic circuit 180.
[0106] The memory access controller 103 includes a dimension #0 counter 150, a dimension #1 counter 151, a dimension #2 counter 152, a dimension #3 counter 153, an address calculation unit 160, a connection switching unit 170, and the like. Information indicating the sizes of dimensions #0 to #3 and the increment width of the dimension of access order #0 are input to the dimension #0 counter 150, the dimension #1 counter 151, the dimension #2 counter 152, and the dimension #3 counter 153. Information indicating the size of dimension #0 is input to the dimension #0 counter 150. For example, a first parameter related to the first dimension of data is set in the dimension #0 counter 150. Information indicating the size of dimension #1 is input to the dimension #1 counter 151. For example, a second parameter related to the second dimension of data is set in the dimension #1 counter 151. Information indicating the size of dimension #2 is input to the dimension #2 counter 152. For example, a third parameter related to the third dimension of data is set in the dimension #2 counter 152. In the example of FIG. 18, the memory access controller 103 mounted on the arithmetic circuit 180 mounts an address generator. In the example of FIG. 18, in the connection switching unit 170 of the memory access controller 103 that switches the connection of the carry-over signals of the four counters, software can set the connection order in advance so that memory access can be performed in an arbitrary order. Further, information indicating the access order of dimensions #0 to #3, information indicating the head address, and the like are input to the address calculation unit 160. Further, information indicating the access order of dimensions #0 to #3 is input to the connection switching unit 170. The connection switching unit 170 switches the connection order of the dimension #0 counter 150, the dimension #1 counter 151, the dimension #2 counter 152, and the dimension #3 counter 153 based on the information indicating the access order of dimensions #0 to #3.
[0107] An example of the control flow of software in the case of the configuration of FIG. 18 described above is shown in FIG. 19. FIG. 19 is a flowchart showing the procedure of the process according to the second embodiment.
[0108] As shown in FIG. 19, when the data is in an amount that can fit into the temporary buffer 182 inside the hardware (step S201: Yes), the processor 101 sets the variable i to "0" (step S202). That is, when the data is in an amount that can fit into the temporary buffer 182 inside the hardware, the processor 101 performs the following processing without dividing the data.
[0109] On the other hand, when the data is not in an amount that can fit into the temporary buffer inside the hardware (step S201: No), the processor 101 performs division of the convolution process (step S203). When the data is not in an amount that can fit into the temporary buffer inside the hardware, the processor 101 divides the data into a plurality of pieces (step S203). For example, the processor 101 divides the data into i + 1 pieces (where i is 1 or more in this case). Then, the processor 101 sets the variable i to "0".
[0110] Then, the processor 101 performs parameter setting for division i (step S204). The processor 101 performs parameter setting for use in processing the data of division i corresponding to the variable i. For example, the processor 101 performs parameter setting for use in processing the data of division 0 corresponding to the variable 0. For example, the processor 101 sets at least one of the dimension size, the dimension access order, the increment or decrement width of the counter, and the dimension multiplier. For example, the processor 101 sets at least one of the parameters related to the first dimension of the data of division i, the parameters related to the second dimension of the data of division i, and the parameters related to the third dimension of the data of division i.
[0111] Then, the processor 101 kicks the arithmetic circuit 180 (step S205). The processor 101 issues a trigger to the arithmetic circuit 180.
[0112] Then, in response to a request from the processor 101, the arithmetic circuit 180 executes loop processing (step S301).
[0113] And if the operation for division i has not been completed (step S206: No), the processor 101 repeats step S206 until the process ends. Note that the processor 101 and the arithmetic circuit 180 may communicate until the operation for division i is completed. The processor 101 may perform polling or interrupt-based confirmation with the arithmetic circuit 180.
[0114] And if the operation for division i has been completed (step S206: Yes), the processor 101 determines whether i is the last division (step S207).
[0115] If i is not the last division (step S207: No), the processor 101 increments the variable i by 1 (step S208). Then, the processor 101 returns to step S204 to repeat the process.
[0116] If i is the last division (step S207: Yes), the processor 101 ends the process. For example, if the processor 101 has not performed data division, since the data with i = 0 is the last data, the process ends.
[0117] In the "parameter setting for division i" in step S204 in FIG. 19, by setting the "access order of dimensions" in advance in the register in the arithmetic circuit 180 before the operation, the memory access controller 103 can access data flexibly. As an example, in a certain recognition task, the order of reading 3D data of an RGB image can be set as the width direction first, then the height direction, and then the RGB channel direction (in the representation of Table 1, in the order of W, H, C). In another recognition task, it may be set to read the RGB channel direction first, then the width direction, and finally the height direction (in the representation of Table 1, in the order of C, W, H).
[0118] Here, an example of the control change process by the connection switching unit 170 is shown in FIG. 20. FIG. 20 is a diagram showing an example of the process according to the second embodiment. The arrows in FIG. 20 indicate the direction from the source to the destination of the physical signal line. Also, the dotted arrows in layout A in FIG. 21 indicate the order of reading data. FIG. 21 is a diagram showing an example of memory access according to the second embodiment.
[0119] In the example of FIG. 20, since the 3D data of the RGB image is the target, address generation is performed using three counters, namely, the dimension #0 counter 150, the dimension #1 counter 151, and the dimension #2 counter 152, without using the dimension #3 counter 153. In FIG. 20, the connection switching unit 170 shows the case where the clock pulse CP is input to the dimension #0 counter 150 and the dimension #0 counter 150, the dimension #1 counter 151, the dimension #2 counter 152, and the dimension #3 counter 153 are connected in this order.
[0120] When each of the dimension #0 counter 150, the dimension #1 counter 151, and the dimension #2 counter 152 in FIG. 20 corresponds to the dimensions of the width (W), height (H), and RGB channel (C) of the 3D RGB image data, the image can be read in the order of W, H, and C. That is, in the case of the connection of the counters of the memory access controller 103 in FIG. 20, as shown in FIG. 21, the entire data DT11 corresponding to red (R), the entire data DT12 corresponding to green (G), and the entire data DT13 corresponding to blue (B) are accessed in this order.
[0121] Next, another example of the control change process by the connection switching unit 170 is shown in FIG. 22. FIG. 22 is a diagram showing another example of the process according to the second embodiment. The arrows in FIG. 22 indicate the direction from the source to the destination of the physical signal line. Also, the dotted arrows in layout A in FIG. 23 indicate the order of reading data. FIG. 23 is a diagram showing another example of memory access according to the second embodiment.
[0122] In the example of FIG. 22, since the three-dimensional data of the RGB image is the target, address generation is performed using three counters, namely, the dimension #0 counter 150, the dimension #1 counter 151, and the dimension #2 counter 152, without using the dimension #3 counter 153. In FIG. 22, the connection switching unit 170 shows the case where the clock pulse CP is input to the dimension #2 counter 152 and the dimension #2 counter 152, the dimension #0 counter 150, the dimension #1 counter 151, and the dimension #3 counter 153 are connected in this order.
[0123] When each of the dimension #0 counter 150, the dimension #1 counter 151, and the dimension #2 counter 152 in FIG. 22 corresponds to the dimensions of the width (W), height (H), and RGB channel (C) of the three-dimensional RGB image data, the image can be read in the order of C, W, H. That is, in the case of the connection of the counters of the memory access controller 103 in FIG. 22, as shown in FIG. 23, the first data of the data DT21 corresponding to red (R), the first data of the data DT22 corresponding to green (G), the first data of the data DT23 corresponding to blue (B), the second data of the data DT21 corresponding to red (R), and so on are accessed in this order.
[0124] As shown in the two examples of FIGS. 20 to 23, the memory built-in device 20A can perform memory access in different orders by changing the connection even in the case of the same layout A.
[0125] As described above, in the second embodiment, the memory built-in device 20A can read and write tensor data from the memory in an arbitrary order, and can perform data access optimal for the arithmetic unit without being restricted by software or sensor specifications. As a result, the memory built-in device 20A can complete the processing of the same tensor with a small number of cycles by making the most of the parallelization of the arithmetic unit. Therefore, the memory built-in device 20A can also contribute to power reduction of the entire system. In addition, since the tensor address calculation can be performed without the intervention of the processor after the parameters are set once, power-saving data access is possible.
[0126] [2. Other Embodiments] The processes according to the above-described embodiments may be implemented in various different forms (modifications) other than the above-described embodiments.
[0127] [2-1. Other Configuration Examples (Image Sensor, etc.)] For example, the above-described memory-integrated devices 20 and 20A may be integrated with the sensor 600. An example of this case is shown in FIG. 24. FIG. 24 is a diagram showing an example of application to a memory stacked type image sensor device. FIG. 24 shows an intelligent image sensor device (memory stacked type image sensor device) 30 in which an image sensor 600a including an image area and a memory-integrated device 20 serving as a logic area are stacked by a stacking technique. The memory-integrated device 20 has a function of communicating with an external device and can acquire data from sensors 600 other than the image sensor 600a.
[0128] For example, it is assumed to be mounted on an IoT (Internet of Things) sensor node that executes an AI recognition algorithm in an edge device using time-series sensor data and image sensor data for identification recognition and the like. Therefore, by integrating the memory-integrated devices 20 and 20A including the mounted circuit (semiconductor logic circuit) and the like as shown in FIG. 24 with the sensor 600 such as the image sensor 600a in a stacked structure or the like, an intelligent sensor with low power consumption and high flexibility can be realized. The intelligent image sensor device 30 as shown in FIG. 24 can be applied to environmental sensing and in-vehicle sensing solutions.
[0129] [2-2. Others] In addition, among the processes described in each of the above embodiments, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. Additionally, regarding the processing procedures, specific names, and information including various data and parameters shown in the above documents and drawings, they can be arbitrarily changed unless otherwise specified. For example, the various information shown in each figure is not limited to the illustrated information.
[0130] Also, each component of each device shown in the drawings is a functional concept and does not necessarily need to be physically configured as shown in the drawings. That is, the specific form of the distribution and integration of each device is not limited to that shown, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads and usage situations.
[0131] In addition, the above-described embodiments and modification examples can be appropriately combined within a range that does not conflict with the processing content.
[0132] Also, the effects described in this specification are merely examples and are not limiting, and there may be other effects.
[0133] [3. Effects of the Present Disclosure] As described above, the memory - embedded device according to the present disclosure (memory - embedded devices 20 and 20A in the embodiments) includes a processor (processor 101 in the embodiments), a memory access controller (memory access controller 103 in the embodiments), and a memory (first - level cache memory 200, second - level cache memory 300, third - level cache memory 400, memory 500 in the embodiments) that is accessed according to the processing of the memory access controller. The memory access controller is configured to read and write data used in the operation of the convolutional operation circuit to the memory.
[0134] As a result, the memory - embedded device according to the present disclosure accesses a memory such as a cache memory according to the processing of the memory access controller, and reads and writes data used in the operation of the convolution operation circuit to / from a memory such as a cache memory according to the processing of the memory access controller, thereby enabling appropriate access to the memory.
[0135] Also, the processor includes a convolution operation circuit (in the embodiment, the convolution operation circuit 102). As a result, the memory - embedded device reads and writes data used in the operation of the convolution operation circuit in its own device to / from a memory such as a cache memory according to the processing of the memory access controller, thereby enabling appropriate access to the memory.
[0136] Also, the parameter is at least one of a first parameter related to the first dimension of the data before or after the operation, a second parameter related to the second dimension of the data before or after the operation, a third parameter related to the third dimension of the data before the operation, a fourth parameter related to the third dimension of the data after the operation, and a fifth parameter related to the number of data before or after the operation. As a result, the memory - embedded device can enable appropriate access to the memory by specifying the data to be read and written to / from a memory such as a cache memory according to the specified parameter.
[0137] Also, the memory includes cache memories (in the embodiment, the first cache memory 200, the second cache memory 300, and the third cache memory 400). As a result, the memory - embedded device can enable appropriate access to the memory by accessing the cache memory according to the processing of the memory access controller.
[0138] Also, the cache memory is configured to read and write data specified using parameters. Thereby, the memory built-in device can enable appropriate access to the memory by reading and writing data specified using parameters to the cache memory.
[0139] Also, the cache memory constitutes a physical memory address space set using parameters. Thereby, the memory built-in device can enable appropriate access to the memory by accessing the cache memory that constitutes the physical memory address space set using parameters.
[0140] Also, the memory built-in device performs initial setting for a register corresponding to parameters. Thereby, the memory built-in device can enable appropriate access to the memory by performing initial setting for a register corresponding to parameters.
[0141] Also, the convolution operation circuit is used for calculation of artificial intelligence functions. Thereby, the memory built-in device can enable appropriate access to the memory for data used in the calculation of artificial intelligence functions in the convolution operation circuit.
[0142] Also, the artificial intelligence function is learning or inference. Thereby, the memory built-in device can enable appropriate access to the memory for data used in the calculation of artificial intelligence learning or inference in the convolution operation circuit.
[0143] Also, the artificial intelligence function uses a deep neural network. Thereby, the memory built-in device can enable appropriate access to the memory for data used in the calculation using a deep neural network in the convolution operation circuit.
[0144] In addition, the memory built-in device includes an image sensor (image sensor 601a in the embodiment) for inputting an external image. Thereby, the memory built-in device can enable appropriate access to the memory for processing using the image sensor. The image sensor is, for example, a CMOS (Complementary Metal Oxide Semiconductor) image sensor and has a function of acquiring an image in pixel units by a large number of photodiodes.
[0145] In addition, the memory built-in device includes a communication processor for communicating with an external device via a communication network. Thereby, the memory built-in device can enable appropriate access to the memory by communicating with the outside to acquire information.
[0146] An image sensor device (intelligent image sensor device 30 in the embodiment) is an image sensor device including a processor that provides an artificial intelligence function, a memory access controller, a memory accessed according to the processing of the memory access controller, and an image sensor, wherein the memory access controller is configured to read and write data used in the operation of the convolutional operation circuit to and from the memory according to the specification of parameters. Thereby, the image sensor device can enable appropriate access to the memory by reading and writing data used in the operation of the convolutional operation circuit, such as an image captured by the device itself, to and from a memory such as a cache memory according to the processing of the memory access controller.
[0147] Note that the present technology can also adopt the following configuration. (1) A processor, A memory access controller, A memory accessed according to the processing of the memory access controller, A memory built-in device including the same, The memory access controller is configured to read and write data used in the operation of the convolutional operation circuit to and from the memory according to the specification of the parameters. Memory built-in device. (2) The processor includes the convolutional operation circuit. The memory built-in device according to (1). (3) The parameter is At least one of a first parameter related to the first dimension of the data before the operation or the data after the operation, a second parameter related to the second dimension of the data before the operation or the data after the operation, a third parameter related to the third dimension of the data before the operation, a fourth parameter related to the third dimension of the data after the operation, and a fifth parameter related to the number of the data before the operation or the data after the operation. The memory built-in device according to (2). (4) The memory includes a cache memory. The memory built-in device according to (3). (5) The cache memory is configured to read and write the data specified by the parameter. The memory built-in device according to (4). (6) The cache memory constitutes a physical memory address space set using the parameter. The memory built-in device according to (5). (7) Perform initial setting for the register corresponding to the parameter. The memory built-in device according to any one of (3) to (6). (8) The convolutional operation circuit is used for the calculation of the artificial intelligence function. The memory built-in device according to any one of (2) to (7). (9) The artificial intelligence function is learning or inference. The memory built-in device according to (8). (10) The function of the artificial intelligence uses a deep neural network. The memory - embedded device according to (8) or (9). (11) Including an image sensor The memory - embedded device according to any one of (1) to (10). (12) Including a communication processor that communicates with an external device via a communication network The memory - embedded device according to any one of (1) to (11). (13) Set the register corresponding to the parameter, Execute a program including a convolution operation having an array according to the parameter. Processing method. (14) Among the parameters for specifying the data read and written by the processor that reads and writes data used in the operation of the convolution operation circuit to the memory, Set at least one of the first parameter related to the first dimension of the data before the operation or the data after the operation, the second parameter related to the second dimension of the data before the operation or the data after the operation, the third parameter related to the third dimension of the data before the operation, the fourth parameter related to the third dimension of the data after the operation, and the fifth parameter related to the number of the data before the operation or the data after the operation. Parameter setting method for executing control. (15) A processor that provides an artificial intelligence function, A memory access controller, A memory accessed according to the processing of the memory access controller, An image sensor, An image sensor device including The memory access controller is configured to read and write data used in the operation of the convolution operation circuit to the memory according to the specification of the parameter. Image sensor device.
Explanation of symbols
[0148] 10 Processing system 20, 20A Memory built-in device 100 Arithmetic unit 101 Processor 102 Convolution arithmetic circuit 103 Memory access controller 200 First cache memory 300 Second cache memory 400 Third cache memory 500 Memory 600 Sensor 600a Image sensor 700 Cloud system
Claims
1. A processor, a memory access controller, and a memory accessed according to the processing of the memory access controller, wherein the memory built-in device includes: the memory access controller is configured to read and write data used in the operation of the convolutional operation circuit to and from the memory according to the specification of parameters; the processor includes the convolutional operation circuit; the parameters are at least one of a first parameter related to a first dimension of the data before the operation or the data after the operation, a second parameter related to a second dimension of the data before the operation or the data after the operation, a third parameter related to a third dimension of the data before the operation, a fourth parameter related to a third dimension of the data after the operation, and a fifth parameter related to the number of the data before the operation or the data after the operation. A memory built-in device.
2. The memory includes a cache memory, The memory built-in device according to Claim 1.
3. The cache memory is configured to read and write data specified using the parameters, The memory built-in device according to Claim 2.
4. The cache memory constitutes a physical memory address space set using the parameters, The memory built-in device according to Claim 3.
5. Performing initial setting for a register corresponding to the parameter, The memory built-in device according to Claim 1.
6. The convolutional operation circuit is used for the calculation of artificial intelligence functions, The memory built-in device according to Claim 1.
7. The artificial intelligence function is learning or inference, The memory built-in device according to Claim 6.
8. The artificial intelligence function uses a deep neural network, The memory built-in device according to Claim 6.
9. Including an image sensor, The memory built-in device according to Claim 1.
10. Including a communication processor that communicates with an external device via a communication network, The memory built-in device according to Claim 1.
11. Performing setting of a register corresponding to a parameter, Executing a program including a convolutional operation having an array according to the parameter, The parameters are At least one of the first parameter related to the first dimension of the data before the operation or the data after the operation, the second parameter related to the second dimension of the data before the operation or the data after the operation, the third parameter related to the third dimension of the data before the operation, the fourth parameter related to the third dimension of the data after the operation, and the fifth parameter related to the number of the data before the operation or the data after the operation. Processing method.
12. Among the parameters for specifying the data read and written by the processor that reads and writes the data used in the operation of the convolutional operation circuit to the memory, Set at least one of the first parameter related to the first dimension of the data before the operation or the data after the operation, the second parameter related to the second dimension of the data before the operation or the data after the operation, the third parameter related to the third dimension of the data before the operation, the fourth parameter related to the third dimension of the data after the operation, and the fifth parameter related to the number of the data before the operation or the data after the operation. Parameter setting method for executing control.
13. A processor that provides an artificial intelligence function, A memory access controller, A memory accessed according to the processing of the memory access controller, An image sensor, An image sensor device including: The memory access controller is configured to read and write the data used in the operation of the convolutional operation circuit to the memory according to the specification of the parameters. The processor includes the convolutional operation circuit. The parameters are: At least one of the first parameter related to the first dimension of the data before the operation or the data after the operation, the second parameter related to the second dimension of the data before the operation or the data after the operation, the third parameter related to the third dimension of the data before the operation, the fourth parameter related to the third dimension of the data after the operation, and the fifth parameter related to the number of the data before the operation or the data after the operation. Image sensor device.
Citation Information
Patent Citations
Address generator
JP2001184260A
Apparatus, systems and computer-implemented methods for processing instruction for accessing n-dimensional tensor
JP2017138964A
Arithmetic processing circuit and recognition system
JP2018067154A