Arithmetic processing unit
Patent Information
- Application Number
- JP2022063290
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-06
- Publication Date
- 2026-09-30
- Estimated Expiration
- 2042-04-06
AI Technical Summary
【0013】 本発明によれば、空間認識や物体認識などの知的処理をハードウェア実装する際に、演算処理において必要となるデータの複数の組み合わせに関して並列に読み出し処理および書き込み処理を実行することができるため、ロボットの制御などに必要な知的処理を高速に実行することが可能となる。
Smart Images

Figure 0007926756000003 
Figure 0007926756000004 
Figure 0007926756000005
Abstract
Description
[[Technical Field]]
[0001] The present invention relates to an arithmetic processing device having a memory circuit and an arithmetic circuit that are suitable for executing predetermined arithmetic processing. [[Background Art]]
[0002] Active research has been conducted on modeling intellectual processing in living organisms and applying the model to control of robots and the like. For example, Non-Patent Document 1 reports that, as the RatSLAM algorithm obtained by modeling the memory function of the hippocampus of rodents and applying the model to a spatial recognition system, visual information from a camera and a self-position are learned according to the Hebb's rule, and the correlation between the visual information and the self-position is stored, thereby enabling advanced spatial recognition. Further, Patent Document 1 discloses an implementation method of SLAM, which is a type of self-position estimation algorithm.
[0003] Further, as an intellectual processing algorithm for recognizing a specific pattern from an image or the like, for example, Non-Patent Document 2 proposes a convolutional neural network combining depth-wise convolution arithmetic processing and point-wise convolution arithmetic processing.
[0004] Further, Non-Patent Document 3 proposes TMEM (transpose-memory) as a technology enabling parallel access to a memory, specifically as an implementation method that locally enables parallel access to a memory. In this memory, when two-dimensional data is stored, elements are stored while being shifted in parallel-arranged RAMs, thereby enabling parallel access to column components and row components of the data. [[Prior Art Documents]] [[Patent Documents]]
[0005] [[Patent Document 1]] Japanese Unexamined Patent Application Publication No. 2021-77353 [[Non-Patent Documents]]
[0006] [Non-Patent Document 1] Milford Michael J., et al., “RatSLAM: A hippocampal model for simultaneous localization and mapping,” IEEE International Conference on Robotics and Automation, 2004. [Non-Patent Document 2] Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, Hartwig Adam, “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications.” arXiv:1704.04861, 2017. [Non-Patent Document 3] Li Yu, et al., Journal of Signal Processing Systems, 2008. [Non-Patent Document 4] Joseph Redmon, Ali Farhadi, “YOLOv3: An Incremental Improvement,” arXiv:1804.02767v1, 2018. [Overview of the Initiative] [Problems that the invention aims to solve]
[0007] However, performing advanced spatial recognition by memorizing the correlation between visual information and self-position as described above requires a vast amount of memory capacity for learning, making hardware implementation difficult. In other words, to implement the above process in hardware, a memory device is needed to store visual information about objects placed in space in association with self-position information. Furthermore, in order to introduce ambiguity into the spatial position, it is desirable to store visual information about objects placed in space within a predetermined range centered on a predetermined position.
[0008] In this case, there are two possible units for accessing the information held in the memory: a combination of data distributed in a predetermined area that corresponds to a single event, and a combination of data at the same location that corresponds to multiple events. In conventional computing devices, when performing arithmetic processing on these multiple data combinations, parallel reading was not possible for at least one of the data combinations, which led to a decrease in processing speed.
[0009] Furthermore, when performing processing on a convolutional neural network that combines depth-wise and point-wise convolution operations, there are two possible units for accessing feature data: combinations of data distributed in a predetermined region and associated with a single feature, and combinations of data at the same location and associated with multiple features. When performing calculations on these multiple data combinations, parallel reading from memory was not possible for at least one of the data combinations, which led to a decrease in processing speed.
[0010] Furthermore, the aforementioned TMEM was used for rearranging data during video compression, and while it enabled localized parallel access to memory, it did not read all the data at predetermined locations related to specific events or features in parallel. Moreover, it did not perform predetermined arithmetic operations on the read data to realize specific intelligent processing.
[0011] The present invention has been made in view of the above circumstances, and aims to provide an arithmetic processing device characterized in that it performs parallel read and write operations on multiple combinations of data corresponding to predetermined arithmetic operations on different memory blocks of a memory circuit, and the arithmetic circuit performs predetermined arithmetic operations on the data read in parallel as described above with respect to the data stored in the memory circuit. [Means for solving the problem]
[0012] To achieve the above objective, the arithmetic processing unit of the present invention comprises a memory circuit that is composed of multiple memory blocks and can access data in parallel by independently accessing each memory block, An arithmetic processing device comprising: an arithmetic circuit that performs predetermined arithmetic processing on the aforementioned data; a control circuit for reading and writing data held in the memory circuit; and a control circuit for controlling the arithmetic processing in the arithmetic circuit, wherein the memory circuit stores data relating to different events or features correlated with spatial information; the arithmetic circuit performs predetermined arithmetic processing on multiple combinations of the data read in parallel with respect to the data stored in the memory circuit; and the control circuit is capable of performing parallel read and write processing on multiple combinations of data corresponding to the predetermined arithmetic processing on different memory blocks of the memory circuit. It is characterized by the following: [Effects of the Invention]
[0013] According to the present invention, when implementing intelligent processing such as spatial recognition and object recognition in hardware, it is possible to perform parallel read and write operations for multiple combinations of data required in the computation process, thereby enabling high-speed execution of intelligent processing necessary for robot control and the like. [Brief explanation of the drawing]
[0014] [Figure 1] Block diagram of the arithmetic processing unit in the first embodiment. [Figure 2] A schematic diagram of the hippocampus-entorhinal cortex model in the first embodiment. [Figure 3] A diagram illustrating the storage of weight information for each event in the first embodiment. [Figure 4] A figure showing an example of the weight distribution in the first embodiment. [Figure 5] A figure showing an example of the weight distribution values in the first embodiment. [Figure 6] A diagram illustrating the addition of weight information in the first embodiment. [Figure 7] A configuration diagram of a memory circuit according to the first embodiment. [Figure 8] A configuration diagram of a memory circuit according to the first embodiment. [Figure 9] A configuration diagram of a memory circuit according to the first embodiment. [Figure 10] A configuration diagram of an arithmetic circuit according to the first embodiment. [Figure 11] A configuration diagram of a memory circuit according to the first embodiment. [Figure 12] Diagrams of an event type detection device and a spatial position detection device according to the first embodiment. [Figure 13] A flowchart of processing according to the first embodiment. [Figure 14] A configuration diagram of a memory circuit according to the first embodiment. [Figure 15] A configuration diagram of a memory circuit according to the first embodiment. [Figure 16] A configuration diagram of a memory circuit according to the first embodiment. [Figure 17] A block diagram of an arithmetic processing device according to the second embodiment. [Figure 18] A configuration diagram of an arithmetic circuit according to the second embodiment. [Figure 19] A block diagram of an arithmetic processing device according to the third embodiment. [Figure 20] A configuration diagram of a convolutional neural network according to the third embodiment. [Figure 21] A configuration diagram of a memory circuit according to the third embodiment. [Figure 22] A configuration diagram of a multiply-accumulate arithmetic circuit according to the third embodiment. [Figure 23] A configuration diagram of a memory circuit according to the third embodiment. Mode for Carrying Out the Invention
[0015] First Embodiment FIG. 1 is a block diagram of the arithmetic processing device described in the present embodiment. In Figure 1, the arithmetic processing unit 100 in this embodiment includes a memory circuit 2 composed of a plurality of memory blocks 1 that store data relating to events occurring in a space having a predetermined extent, an arithmetic circuit 3 that performs predetermined calculations on the data, and a control circuit 4 for reading and writing data held in the memory circuit 1 and for controlling the arithmetic processing in the arithmetic circuit 3. Here, the arithmetic processing unit 100 performs processing based on the hippocampus-entorhinal cortex model described later.
[0016] Figure 1 also shows some of the memory elements that make up each memory block, along with labels corresponding to their coordinates, expressed in the format (x, y). To avoid making Figure 1 too complex, only a portion of the memory block configuration composed of memory elements is shown here; detailed diagrams and explanations will be described later in this embodiment.
[0017] Figure 2 shows a schematic diagram of the hippocampus-entorhinal cortex model, which is the basic concept of the computational processing model that learns the correlation between information about events occurring in space and spatial information, as implemented in the computational processing unit 100 described in this embodiment.
[0018] First, the processing model realized in this embodiment will be explained using Figure 2. As shown in Figure 2, the hippocampal-entorhinal-cortex model used as the computational processing model in this embodiment is a simplified hippocampal-entorhinal-cortex model consisting of place cells 8, cue cells 9, and connecting cells 10, which is a simplified version of the hippocampal-entorhinal-cortex model observed in the hippocampus and entorhinal cortex of organisms such as rats as a physiological finding, with the aim of implementing it as hardware.
[0019] In the hippocampus-entorhinal cortex model shown in Figure 2, the internal state of place cells 8 transitions based on velocity and direction information to represent self-positional information. Additionally, cue cells 9, which respond to specific events, represent information about the type of event.
[0020] Furthermore, linked cells 10 have the function of learning the correlation between information about events occurring in space and spatial information by combining spatial information input from place cells 8 with non-spatial information about specific events input from cue cells 9.
[0021] For example, let's call the entity that detects events while moving through space an agent. In the aforementioned hippocampus-entorhinal cortex model, as shown in Figure 3, agent 11 detects events a, b, and c while moving between positions 12-15 in space. Each time an event is detected, it stores weight information with a predetermined spatial extent for the location where the event was detected, as information about the event that correlates with the spatial information.
[0022] In other words, as agent 11 moves from position 12 to position 15, it detects events a, b, and c, and the weight distribution 16 for each event is stored as shown in Figure 3 (A) to (C). In Figure 3, the magnitude of the stored weight distribution is shown as the magnitude in the direction perpendicular to the two-dimensional space.
[0023] Here, the weight information having a predetermined spatial extent has a maximum value at the center, as shown in Figure 4, and has a distribution in which the value decreases as it moves away from the center within the predetermined spatial extent. Here, the weight distribution is often a two-dimensional Gaunsian distribution as shown in Figure 4, but it is not necessarily limited to this, and the weight information may have other distributions as long as it represents the correlation between information about events occurring in space and spatial information and has a predetermined spatial extent.
[0024] Furthermore, although the spatial distribution of weight information in this embodiment was described as a two-dimensional Gaunsian distribution as mentioned above, the space in which the agent moves is represented as discrete two-dimensional coordinates in this embodiment, as shown in Figure 3. For example, as shown in (A) to (C), the space in which the agent moves is represented as discrete two-dimensional coordinates (x, y). Therefore, the spatial distribution of weight information also takes discrete values corresponding to the discrete two-dimensional coordinates.
[0025] Figure 5 shows a weight distribution having discrete values corresponding to the discrete two-dimensional coordinates in this embodiment. As shown in Figure 5, the weight distribution in this embodiment can take on three different values, 1, 2, and 4, within a region defined by a 3x3 discrete two-dimensional coordinate system.
[0026] As shown in Figure 3(A), if the same event is detected in a different location, weight information with a predetermined spatial extent stored at multiple locations is stored as information about the event that correlates with the spatial information. In this case, the weight information of the same event detected at the present time is added to the weight information of past events, so as the agent moves, the weight information of the events is stored as information that correlates with the spatial information, as shown in Figure 3(A).
[0027] In Figure 3(A), since there is no spatial overlap between the weight distributions for the same event detected at different locations, each weight distribution is stored independently. However, if the two distributions overlap spatially, the weight information for that event is stored as information correlated with spatial information, as shown in Figure 6(A), by adding the value of the newly detected weight distribution to the previously stored weight distribution.
[0028] Figure 6(A) shows the weight information for an event, where two weight distributions are added together, creating a further peak between the peaks of the two weight distributions. Figure 6(B) shows Figure 6(A) as values corresponding to two-dimensional spatial positions (only the distribution of positions with non-zero values is shown). Note that the x and y coordinates are omitted in Figure 6(B).
[0029] Furthermore, if multiple identical events are detected, the weight distribution at those locations is added together as described above, so that the weight information related to those events is stored as information correlated with spatial information, as described above.
[0030] Furthermore, in the hippocampus-entorhinal cortex model, when different events are detected, the weight information is stored as memories related to each different event, as shown in Figures 3(A) to (C), and is correlated with spatial information. In other words, in the hippocampus-entorhinal cortex model, spatially correlated information related to an event is stored individually for each event.
[0031] Note that Figure 6(A) illustrates the summation of a weight distribution with a two-dimensional distribution, showing the values as the height direction in a two-dimensional space. Since the figure represents a three-dimensional diagram of the weight distribution on a two-dimensional plane, the distribution shape does not necessarily match that of the actual weight distribution in Figure 6(B).
[0032] Next, we will describe the arithmetic processing unit that implements the aforementioned hippocampus-entorhinal cortex model as hardware. As mentioned above, the arithmetic processing unit 100 in this embodiment has a memory circuit 2 composed of a plurality of memory blocks 1, an arithmetic circuit 3, and a control circuit 4.
[0033] Here, an example of the configuration of the memory circuit 2 is shown in Figure 7. As shown in Figure 7, the memory circuit 2 in this embodiment has a configuration in which multiple memory blocks 1 corresponding to two-dimensional spatial coordinates are located. Furthermore, each memory block 1 is composed of memory elements that store values. In Figure 7, each memory block 1 constitutes each row of the memory circuit 2, and the memory circuit 2 has a total of 25 rows of memory blocks 1. In Figure 7, the boundaries of each memory block 1 are shown by solid lines.
[0034] Furthermore, in the memory circuit 2 shown in Figure 7, one cell enclosed by a solid line and a dashed line corresponds to a memory element that stores the value obtained by adding weight information values corresponding to a predetermined event type at a predetermined coordinate position.
[0035] In this embodiment, storing data correlated with spatial information about an event in the memory circuit 2 is equivalent to learning the correlation between information about an event occurring in space and spatial information in the aforementioned hippocampus-entorhinal cortex model. In the following explanation, the terms "memory" and "learning" will be used, but unless otherwise specified, they mean the same thing.
[0036] Next, as shown in Figure 7, each memory block 1 has a data input port 17 and a data output port 18 between it and the arithmetic circuit 3.
[0037] Furthermore, each memory element in memory block 1 corresponds to a position coordinate in two-dimensional space. In Figure 7, the coordinates expressed in the form (x, y) are used as labels for each memory element. In this embodiment, it is assumed that data corresponding to a 5x5 (=25) position coordinate is stored as spatial information, and the position in the two-dimensional space where the agent moves is detected as a 5x5 two-dimensional spatial coordinate.
[0038] As shown in Figure 7, the position coordinates of each memory element in the corresponding two-dimensional space are shifted by one coordinate position for each memory block. Furthermore, in Figure 7, it is assumed that all memory elements initially store a value of 0. In Figure 7, the values held are indicated numerically below the coordinate notation mentioned above.
[0039] Furthermore, each column of the memory circuit 2 corresponds to the type of event detected in space. In Figure 7, each column is labeled with one of six types of events, numbered 0-5. In Figure 7, the boundaries between the labels for each event type are shown with dashed lines. In other words, this embodiment assumes a case where there are six types of events detected by the agent in space. Note that in Figure 7, the input port for control signals from the control circuit 4 and the output port for outputting data to the outside of the arithmetic processing unit 100 are omitted.
[0040] Figure 8 also shows that, for event type (4), memory elements corresponding to the two-dimensional spatial coordinates (1, 0), (2, 0), (3, 0), (1, 1), (2, 1), (3, 1), (1, 2), (2, 2), and (3, 2) store non-zero values. In this case, it means that the agent has memory related to event type (4) at the aforementioned coordinate positions. In Figure 8, memory elements that store non-zero values are shown in black. The coordinates and values are shown in white text. (The same applies to subsequent figures.)
[0041] Furthermore, regarding event type (4), when the two weight distributions shown in Figure 6 are added together, the memory circuit stores the weight information as information correlated with the spatial information, as shown in Figure 9. (Figure 9 shows the case where an additional weight distribution is added to the storage in Figure 8.)
[0042] Here, in Figures 8 and 9, we will explain how to access the data in the memory circuit that is already stored in the memory circuit and is subject to the addition of weight distributions for the same event data, in order to add the weight distributions corresponding to predetermined event data (in this case, event type (4)). In this embodiment, each memory block 1 has an independent input port 17 and an output port 18, and as will be described later, it is possible to read data in a predetermined coordinate region relating to a specified event in parallel based on a control signal from the control circuit 4.
[0043] In other words, as shown in Figures 8 and 9, the data related to the predetermined events to be added to the weight distribution is distributed and held across multiple memory blocks, and can therefore be accessed in parallel through the aforementioned input and output ports. In Figures 8 and 9, the input and output ports for inputting and outputting data when adding the weight distribution are shown in black. As shown in Figures 8 and 9, these input and output ports are configured independently for each memory block, and it can be seen that the data can be accessed in parallel through the black input and output ports.
[0044] Next, an example of the configuration of the arithmetic circuit 3 in this embodiment will be described with reference to Figure 10. The arithmetic circuit 3 in this embodiment has the function of adding a predetermined value to each of the multiple input data, and is composed of multiple adders 19, a register 20 that holds the input values, a register 21 that holds the addition result, and a register 22 that holds multiple predetermined values that are the values to be added. The input value register 20 has an input port 23 for inputting data read from the memory circuit 2, and the addition result register 21 has an output port 24 that outputs the addition result to the memory circuit 2.
[0045] The arithmetic circuit 3 reads from the memory circuit 2 and stores multiple input values input from the input port 23 in the input value register 20. After adding multiple predetermined values held in the value to be added register 22, it stores the addition result in the addition result register 21 and outputs it to the memory circuit 2 from the output port 24. In this embodiment, the adder 19 and registers 20, 21, and 22 are configured to perform operations in parallel on 3x3=9 values corresponding to the spatial distribution of the weight distribution.
[0046] Next, the function of the control circuit 4 in this embodiment will be described. As shown in Figure 1, the control circuit 4 in this embodiment has an input terminal 5 for event type data and an input terminal 6 for event detection position data, and has the function of outputting a control signal to execute control to read the corresponding data from the memory circuit 2 and input it to the arithmetic circuit 3 based on the input data. Furthermore, it has the function of outputting a control signal to execute control to write the calculation result by the arithmetic circuit 3 back to the original position (read position) in the memory circuit 2.
[0047] Furthermore, the control circuit 4 also has the function of outputting control signals for outputting in parallel stored data relating to all events at predetermined positions in two-dimensional space. In Figure 1, the aforementioned control signals are represented by a single signal line. The control circuit 4 also has the function of outputting control signals for controlling the execution of calculations in the arithmetic circuit 3.
[0048] In this memory circuit, as mentioned above, the position coordinates of each memory element in the corresponding two-dimensional space are shifted by one coordinate position for each memory block, making it possible to output stored data for all events at a single predetermined position in the two-dimensional space in parallel. Figure 11 shows that stored data for all events at a predetermined position (2,2) can be output in parallel from a memory circuit that holds arbitrary data. In Figure 11, the memory element that holds stored data for all events at the predetermined position (2,2) is shown in black. The coordinates and values are shown in white text. The output port for outputting stored data for all events at the predetermined position (2,2) in parallel is also shown in black.
[0049] As described above, the memory circuit in this embodiment makes it possible to read in parallel multiple data points in a predetermined coordinate region relating to a predetermined event and data relating to all events at a predetermined position in two-dimensional space. In other words, input processing and output processing can be performed in parallel for multiple combinations of data, and predetermined arithmetic processing can be performed in parallel on that data.
[0050] In the hippocampus-entorhinal cortex model shown in Figure 2, the event detection location data, which corresponds to spatial information input to the connected cells, and the event type data, which corresponds to non-spatial information, are input to the arithmetic processing unit 100 from a device configured outside the arithmetic processing unit 100. In other words, with respect to the hippocampus-entorhinal cortex model, the functions realized by the arithmetic processing unit 100 described in this embodiment correspond to the connected cells mentioned above. In the arithmetic processing unit 100 of this embodiment, the functions for calculating spatial information and non-spatial information in the hippocampus-entorhinal cortex model are not essential, but the arithmetic processing unit may include these functions. Examples of embodiments of these functions will be described below.
[0051] In this embodiment, data relating to the type of event and data relating to the location where the event was detected are input from an external source, and specific examples of each type of data are shown below. First, in this embodiment, the data relating to the type of event is obtained by performing object detection processing on an image acquired by a camera, and the object recognized as a predetermined type is used as the event type data.
[0052] For example, the object detection algorithm described in Non-Patent Document 4 can be used to determine the type of object in the image. The event type detection device 25 in Figure 12 is assumed to have a processor and implement the object detection algorithm as software. Furthermore, six types of object types are set in advance for the event type detection device 25, and when a corresponding object type is detected, the aforementioned object type data is output to the arithmetic processing unit 100.
[0053] Next, the arithmetic processing unit 100 in this embodiment is assumed to be mounted on an agent capable of moving through space, such as a robot. Here, the agent is assumed to have a spatial position detection device 26, as shown in Figure 12, which detects its current location within a pre-set space. The method of realizing the spatial position detection device 26 is not the main focus of the present invention, and any means of realization may be used, but in this embodiment, as an easily implementable means, the spatial position detection device 26 has the function of constantly calculating the distance and direction of movement from the point where the agent first departed, and is capable of calculating the current position (coordinates) within space.
[0054] The aforementioned event type detection device 25 detects the object type and outputs the object type information to the arithmetic processing unit 100. In synchronization with this timing, the spatial position detection device 26 outputs the position (coordinates) in space to the arithmetic processing unit 100. This allows the arithmetic processing unit 100 to learn the event type data and the event detection position data as spatially correlated data.
[0055] Next, the operation of the arithmetic processing unit 100 in this embodiment will be explained using the flowcharts shown in Figures 1 and 13.
[0056] In the arithmetic processing unit 100 of this embodiment, data relating to the type of event that occurred in space (hereinafter referred to as event type data) and data relating to the position of the agent when the event was detected (hereinafter referred to as event detection position data) are input from outside the arithmetic processing unit 100 (processing 27 in the flowchart of Figure 13). Here, the label of the detected event type is (4), and the position coordinates where the event was detected are (2, 1).
[0057] Inside the arithmetic processing unit 100, event type data and event detection location data are input to the control circuit 4 (processing 28 in the flowchart of Figure 13). Based on the event type data and event detection location data, the control circuit 4 outputs a control signal to the memory circuit 2 to read the appropriate data set.
[0058] Here, the appropriate data set refers to multiple data points to which the weight information having a predetermined spatial extent, as described above, should be added, corresponding to the detected event type. In other words, in this embodiment, the weight information having a predetermined spatial extent has values distributed in a 3x3 area corresponding to a two-dimensional Gaussian distribution as shown in Figure 5. Therefore, the memory circuit 2 reads out data from a 3x3 area surrounding the coordinates based on the coordinates of the event type data and the event detection position data (processing 29 in the flowchart of Figure 13).
[0059] Here, it is possible to read data from a predetermined coordinate region related to a specified event in parallel. That is, as shown in Figure 7, the data related to a predetermined event is distributed and held in multiple memory blocks, and can be accessed in parallel through the aforementioned input and output ports.
[0060] The data from the read 3x3 area (a total of 9 data points) is then input to the arithmetic circuit 3 based on the control signal from the control circuit 4 mentioned above (process 30 in the flowchart of Figure 13).
[0061] In this embodiment, the arithmetic circuit 3 includes an adder 19, which adds weight data distributed within a 3x3 area to the input 3x3 area data (processing 31 in the flowchart of Figure 13).
[0062] Once the addition process is completed in the arithmetic circuit 3, the control circuit 4 then inputs control signals to the arithmetic circuit 3 and the memory circuit 2 to read the addition result from the arithmetic circuit 3 and write it to the memory circuit 2.
[0063] Based on the aforementioned control signal, the calculation result read from the arithmetic circuit 3 is written to the memory circuit 2 in a memory element corresponding to the same event and at the same coordinate position as when it was read. In this case, the calculation result can be written to the memory circuit in parallel. That is, as shown in Figure 8, the data related to a predetermined event is distributed and held in multiple memory blocks, and can be accessed in parallel through the aforementioned input port.
[0064] This completes the updating of data correlated with spatial information regarding the event in question. This corresponds to the learning process in the hippocampus-entorhinal cortex model described above (process 32 in the flowchart of Figure 13). As a result, the distribution of values in memory circuit 2 will be as shown in Figure 8, similar to the example described above.
[0065] Furthermore, if the agent moves and recognizes the same event in a different location, the event type data and data regarding the agent's position at the time the event was detected are input to the processing unit 100, as described above.
[0066] Inside the arithmetic processing unit 100, event type data and event detection location data are input to the control circuit 4, as described above, and a control signal is input from the control circuit 4 to the memory circuit 2 to read the appropriate data set. Here, the event type data is the same as described above, while the event detection location data indicates different location coordinates. Here, the label of the detected event type is (4), and the location coordinates where the event was detected are (2, 3). Therefore, the memory circuit 2 reads out 3x3 data surrounding coordinates different from those described above in parallel, based on the coordinates of the event detection location data.
[0067] The data in the 3x3 area (a total of 9 data points) that has been read is then input to the arithmetic circuit 3 based on the control signal from the aforementioned control circuit, and the adder 19 adds the weight data distributed in the 3x3 area to the input 3x3 area data, in the same way as described above.
[0068] Once the addition process is complete, similar to the process described above, the same event is written in parallel to the same memory location based on the event type data and event detection location data. This completes the updating of the data related to the event that is correlated with spatial information. As a result, the distribution of values in the memory circuit will be as shown in Figure 9, similar to the example described above.
[0069] Furthermore, if the agent moves and recognizes a different event (for example, event type (1)) in a different location, the event type data and data regarding the location where the event was detected are input to the processing unit 100, as described above.
[0070] Inside the arithmetic processing unit 100, event type data and event detection location data are input to the control circuit 4, as described above, and a control signal is input from the control circuit 4 to the memory circuit 2 to read the appropriate data set. Here, the label of the detected event type is (1), and the location coordinates where the event was detected are (1, 1). Therefore, the memory circuit 2 reads out data in parallel from a 3x3 area around the coordinates of the event detection location, relating to an event different from the one described above, based on the coordinates of the event detection location data.
[0071] The data in the 3x3 area (a total of 9 data points) that has been read is then input to the arithmetic circuit 3 based on the control signal from the aforementioned control circuit, and the adder 19 adds weight data distributed within a 3x3 range to the input 3x3 area data.
[0072] Once the addition process is complete, the calculation results are written in parallel to the memory circuit 2 based on the event type data and event detection location data, similar to the process described above. This completes the updating of data related to the event that is correlated with spatial information. As a result, the distribution of values in the memory circuit becomes as shown in Figure 14.
[0073] The above process is repeated until a predetermined termination timing is reached, and the memory circuit stores information correlated with spatial information regarding events detected in conjunction with the agent's movement (process 33 in the flowchart of Figure 13). In this embodiment, the termination timing is determined by a termination signal input to the control circuit 4 from an external source. The method for determining the termination timing and the input method / means are not limited to the present invention, so Figure 1 omits the illustration of the termination signal input port, etc.
[0074] Furthermore, after processing is complete, in this embodiment, the memory circuit 2 outputs in parallel the stored data relating to all events at predetermined locations in the two-dimensional space. In this case, the control circuit 4 inputs a read signal to the memory circuit 2 for outputting in parallel the stored data relating to all events at predetermined locations in the two-dimensional space, and the data is output in parallel from the output port. In the memory circuit, as mentioned above, the position coordinates of each memory element in the two-dimensional space are shifted by one coordinate position for each memory block, as shown in Figure 11. Therefore, it is possible to output in parallel the stored data relating to all events at predetermined locations in the two-dimensional space.
[0075] In this embodiment, the output data is assumed to be output externally through the output terminal 7 of the arithmetic processing unit 100 (processing 34 in the flowchart of Figure 13). In this embodiment, the output terminal 7 has a total of 6 ports in order to output the output data externally in parallel, but the method of outputting data externally is not limited to the above configuration and method.
[0076] As mentioned above, the information stored in memory circuit 2 is information that shows the correlation between events detected as the agent moves and spatial information. For example, by detecting the coordinates that store the maximum value for a particular event, it becomes possible to obtain the spatial coordinates in which that event is most likely to occur. The application of the data output from memory circuit 2 is not the main focus of this invention, so further detailed explanation is omitted.
[0077] As mentioned above, in the first embodiment, the data relating to the type of event was obtained by performing object detection processing on an image acquired by a camera and recognizing it as a predetermined object type, but the present invention is not limited to this. For example, the type of command may be recognized by speech recognition with respect to the voice spoken by the user to the agent, and that may be used as the type of event.
[0078] In this case, similar to the event type detection device that performs object detection processing, it is possible to determine the command type and output it to the arithmetic processing unit 100 in this embodiment. Furthermore, the data related to the event type may be any other event, such as olfactory information or tactile information.
[0079] As described above, the arithmetic processing unit in this embodiment is capable of reading in parallel multiple combinations of data in a predetermined coordinate region relating to a predetermined event, and combinations of data relating to all events at a predetermined position in two-dimensional space. In other words, input processing and output processing can be performed in parallel for multiple combinations of data, and predetermined arithmetic processing can be performed in parallel on these data. Therefore, the arithmetic processing of the hippocampus-entorhinal cortex model described in this embodiment can be realized without causing a decrease in processing speed due to data reading and writing from the memory circuit.
[0080] Furthermore, in the above explanation, the position coordinates in two-dimensional space corresponding to each memory element in the memory circuit were shifted by one coordinate position for each memory block, but this is not the only way to shift the coordinate positions. For example, as shown in Figure 16, it is also possible to shift the coordinate positions by two for each memory block. In other words, as mentioned above, any method of storing data (method of determining the storage position) that allows parallel reading of multiple data in a predetermined coordinate region relating to a predetermined event and data relating to all events at a predetermined position in two-dimensional space is acceptable.
[0081] Regarding other methods of shifting coordinate positions, it is possible to store the data in an appropriate memory location based on the following formula. For example, let (m, n) be the spatial coordinates to which the agent moves, and let m and n take values within the range defined by the following formula 1.
[0082]
number
[0083] Here, M and N have values corresponding to the size of the space in which the agent moves (in this embodiment, M=5, N=5). At this time, as shown in Figure 15, the row address of the memory location that stores the spatial coordinates (m, n) and the event label (let's call it c) is calculated by the following equation 2.
[0084]
number
[0085] Here, s represents the amount of shift when shifting the coordinate position. That is, if the shift amount is 1, as in this embodiment, s=1, and if the shift amount is 2, s=2. Figure 15 shows an example of data storage when s=1, and Figure 16 shows an example of data storage when s=2.
[0086] Figure 16 shows that, when s=2, a memory circuit holding arbitrary memories can output in parallel all memory data related to events at a predetermined location (2.2). In Figures 15 and 16, the memory element that holds the stored data for all events at the predetermined position (2,2) is shown in black. The coordinates and values are shown in white text. Additionally, the output port for outputting the stored data for all events at the predetermined position (2,2) in parallel is shown in black.
[0087] In this embodiment, the condition c ≤ M x N is assumed to be satisfied. However, if this condition is not satisfied, parallel reading and writing similar to this embodiment can be achieved by dividing and storing data related to different events at the same location into multiple memory blocks. The determination of the address of the memory element that holds the data based on equation 2 is achieved by incorporating a circuit that executes equation 2 into the control circuit.
[0088] [Second Embodiment] Figure 17 shows a block diagram of the arithmetic processing unit 200 described in this embodiment. In Figure 17, the arithmetic processing unit 200 in this embodiment is modified in which the arithmetic circuit 3 of the arithmetic processing unit 100 in the first embodiment is replaced with an arithmetic circuit 35 that has an addition / subtraction circuit that performs addition and subtraction. The arithmetic processing unit 200 in this embodiment is the same as the first embodiment except that the arithmetic circuit 3 in the first embodiment performs subtraction as well as addition. Therefore, only the differences from the first embodiment will be described in this embodiment, and other points will be omitted as they are the same as the first embodiment.
[0089] First, the arithmetic circuit 35 in the arithmetic processing unit 200 will be explained using Figure 18. As shown in Figure 18, the arithmetic circuit 35 in the arithmetic processing unit 200 has an adder / subtractor 36, and is capable of performing subtraction in addition to addition. That is, for example, in a hippocampus-entorhinal cortex model, if the effect of memory decay over time with respect to memory correlated with space regarding detected events can be realized by subtracting a predetermined value from the data held in the memory circuit 3 after a predetermined time has elapsed. Furthermore, it is possible to add a weight distribution with negative values for a specific event. Moreover, under certain conditions, a negative weight distribution can be set for a given event, and memories related to both positive and negative weight distributions can be added (or subtracted).
[0090] As described above, this embodiment, like the first embodiment, allows for the retention and updating of memories correlated with space regarding detected events in the hippocampus-entorhinal cortex model, and further enables subtraction processing of stored data. This makes it possible to realize the effect of memory decay over time and processing of memories with negative weight distributions.
[0091] [Third Embodiment] Figure 19 shows a block diagram of the arithmetic processing unit described in this embodiment. In Figure 19, the arithmetic processing unit 300 in this embodiment is modified in which the arithmetic circuit of the arithmetic processing unit in the first embodiment is changed to a multiply-accumulate circuit 37 that performs multiply-accumulate operations. Also, the control circuit 39 in Figure 19 has some different functions from the control circuit 4 in the first embodiment.
[0092] Furthermore, the arithmetic processing unit 300 in this embodiment implements depth-wise convolution processing and point-wise convolution processing in the convolutional neural network described in Non-Patent Document 2.
[0093] Furthermore, outside the arithmetic processing unit 300 in this embodiment, a post-processing circuit 38 is configured to perform post-processing on the calculation results of the sum-of-accumulate circuit 37. The calculation results from the post-processing circuit 38 are input to the memory circuit 2.
[0094] In this embodiment, only the differences from the first embodiment will be described, and other points will be the same as in the first embodiment and will not be described.
[0095] First, using Figure 20, the depth-wise convolution and point-wise convolution operations in the convolutional neural network implemented in this embodiment will be explained. Figure 20 shows an extracted portion of the convolution operations that are repeatedly performed hierarchically in the convolutional neural network. As shown in Figure 20, in the convolutional neural network, depth-wise convolution operations and point-wise convolution operations are repeatedly performed alternately.
[0096] As can be seen from Figure 20, the depth-wise convolution operation is implemented as a convolution operation on neurons of hierarchical n by performing a sum-of-products operation on the neuron output values and weight values contained in a predetermined spatial region (receptive field: a 3x3 region in Figure 20) of a feature surface composed of the output values of neurons from the previous hierarchical n-1.
[0097] Furthermore, the pointwise convolution operation is implemented as a convolution operation on neurons of the n+1th hierarchical level by performing a sum-of-products operation on the neuron output values and weight values at predetermined positions on all feature surfaces composed of the output values of neurons from the previous nth hierarchical level. A detailed explanation of the above convolution operation is provided in Non-Patent Document 2, so further explanation is omitted.
[0098] Next, we will explain the case where the depth-wise convolution operation and point-wise convolution operation described above are implemented using the arithmetic processing unit 300 shown in Figure 19, in comparison with Embodiment 1.
[0099] First, in the convolution operation in the arithmetic processing unit 300 of this embodiment, the output values of the neurons in the previous layer are assumed to be held in the memory circuit 2 in Figure 19. Here, each feature surface of the neuron output value corresponds to the event type in Embodiment 1. Furthermore, the spatial distribution of the neuron output values within each feature surface corresponds to the spatial distribution of each event type in Embodiment 1.
[0100] In this embodiment, the memory circuit 2 shown in Figure 19 only shows the portion that holds data for one layer of the convolutional neural network shown in Figure 20, which is the subject of this explanation. However, in reality, it will hold data for an appropriate number of layers according to the computational processing flow. Furthermore, the position coordinates of each memory element in the corresponding two-dimensional space are shifted by one coordinate position for each memory block.
[0101] Next, the functions of the control circuit 39 in this embodiment will be described. The control circuit 39 in this embodiment has the function of controlling the memory circuit 2 and the multiply-accumulate circuit 37 based on a pre-set processing flow. The control circuit 39 has the function of outputting a control signal to execute the control of reading data from a predetermined area relating to a predetermined feature surface in parallel from the memory circuit 2 and inputting it into the multiply-accumulate circuit 37. Furthermore, the control circuit 39 also has the function of outputting a control signal to execute the control of reading stored data relating to all feature surfaces at a predetermined position in two-dimensional space in parallel and inputting it into the multiply-accumulate circuit 37. The control circuit 39 also has the function of outputting a control signal to control the execution of calculations in the multiply-accumulate circuit 37. Furthermore, the control circuit 39 has the function of outputting a control signal to control the execution of processing in the post-processing circuit 38, and further has the function of outputting a control signal to execute the process of writing the processing results of the post-processing circuit 38 to the memory circuit 2.
[0102] Next, the method for realizing depth-wise convolution processing by the arithmetic processing unit 300 in this embodiment will be described. As mentioned above, depth-wise convolution processing is realized by sum-of-products calculations of the neuron output values and weight values included in a predetermined spatial region (receptive field) of a feature surface, which is composed of the output values of neurons from the previous layer. Therefore, the arithmetic processing unit 300 reads out the neuron output values from the previous layer included in the two-dimensional space (receptive field) to be subjected to the sum-of-products calculation with respect to a predetermined feature surface held in the memory circuit 2.
[0103] As shown in Figure 21, each memory block 1 in this embodiment has independent input ports 17 and output ports 18, similar to Embodiment 1, and as will be described later, it is possible to read data for a predetermined coordinate region related to a specified event in parallel based on a control signal from the control circuit 4. That is, as shown in Figure 21, since the data for a predetermined feature surface is distributed and held in multiple memory blocks, it can be accessed in parallel through the aforementioned input ports and output ports. Note that in Figure 21, the output ports and input ports for outputting and inputting data for sum-of-products calculations with load values are shown in black.
[0104] Next, an example of the configuration of the sum-of-accumulate circuit 37 in this embodiment will be described using Figure 22. The sum-of-accumulate circuit 37 in this embodiment has the function of multiplying each of the multiple input data by a predetermined value (weight value) and further accumulating the multiplication results, and is composed of multiple multipliers 40, a register 41 that holds input values, an accumulator 44 that accumulates the multiplication results, and a register 42 that holds multiple predetermined values (weight values) that are the multiplicands. The input value register 41 has an input port 43 that receives data read from the memory circuit 2, and the accumulator 44 has an output port 45 that outputs the accumulation result to the memory circuit 2.
[0105] The multiply-accumulate circuit 37 reads from the memory circuit 2, stores multiple input values input from the input port 43 in the input value register 41, multiplies them by multiple predetermined values (load values) held in the load value register 42, accumulates the multiplication result in the accumulator 44, and outputs it to the memory circuit 2 from the output port 45. In this embodiment, the multiplier 40 and registers 41 and 42 are configured to perform operations on 3x3=9 values corresponding to the load distribution (receptive field) in parallel.
[0106] The cumulative results are input to the memory circuit 2 after undergoing predetermined post-processing, such as activation functions, in the post-processing circuit 38 in Figure 19. However, since post-processing is not the main focus of this invention, a detailed description and explanation are omitted.
[0107] Furthermore, when storing the cumulative results and the calculation results obtained by applying predetermined calculation processing to the cumulative results in the memory circuit 2, as explained in Figure 19, each feature surface of the neuron output value corresponds to the event type in Embodiment 1, and the spatial distribution of the neuron output values within each feature surface corresponds to the spatial distribution of each event type in Embodiment 1.
[0108] As shown in Figures 21 and 23, each feature surface corresponds to each row of the memory circuit, and in this embodiment, it corresponds to feature surface labels 0-5. Furthermore, the position coordinates in the two-dimensional space corresponding to each memory element are shifted by one coordinate position for each memory block.
[0109] Next, the method for realizing point-wise convolution processing by the arithmetic processing unit in this embodiment will be described. As mentioned above, point-wise convolution processing is realized by sum-of-products calculation of the neuron output values and weight values at predetermined positions on all feature surfaces composed of the output values of neurons from the previous layer. Therefore, the arithmetic processing unit 300 reads out the neuron output values from the previous layer at predetermined positions that are to be subjected to sum-of-products calculation for all feature surfaces held in the memory circuit 2.
[0110] In memory circuit 2, as mentioned above, the position coordinates of each memory element in the corresponding two-dimensional space are shifted by one coordinate position for each memory block, as shown in Figure 23. Therefore, it is possible to output the stored data for all feature surfaces at a predetermined position in parallel. That is, as shown in Figure 23, the data for the same position coordinates of all feature surfaces is distributed and held across multiple memory blocks, and can therefore be output in parallel through the output port mentioned above.
[0111] The calculation process in the sum-of-accumulate circuit 37 is the same as the depth-width convolution calculation process described above.
[0112] As described above, the arithmetic processing unit in this embodiment is capable of reading in parallel, both a combination of data for a predetermined coordinate region relating to a predetermined feature surface and a combination of data for all feature surfaces at a predetermined position in two-dimensional space.
[0113] In other words, input and output processing can be performed in parallel for multiple combinations of data, and predetermined arithmetic operations can be performed in parallel on that data. Therefore, both the depth-wise convolution operation and the point-wise convolution operation described in this embodiment can be implemented without causing a decrease in processing speed due to data reading and writing from the memory circuit.
[0114] Furthermore, in the above explanation, the position coordinates in two-dimensional space corresponding to each memory element of the memory circuit were shifted by one coordinate position for each memory block, but as with Embodiment 1, the method of shifting the coordinate positions is not limited to this.
[0115] Furthermore, although this embodiment describes an example in which depth-wise convolution and point-wise convolution operations are performed alternately in a hierarchical manner, it is also possible to apply both depth-wise and point-wise convolution operations to the same hierarchical level of the neural network. In this case, it becomes possible to read in parallel the combinations of multiple data points in a predetermined coordinate region relating to a predetermined feature surface and the combinations of data relating to all feature surfaces at a predetermined position in two-dimensional space, for each of the neuron output values constituting the feature surface at the same hierarchical level held in the memory circuit. [Explanation of symbols]
[0116] 100, 200, 300 Arithmetic Processing Units 1 memory block 2 Memory Circuit 3, 35 Arithmetic circuit 4.39 Control circuits Input ports 5, 6, 17, 23, 43 Output ports 7, 18, 24, 45 8 place cells 9. Queue cells 10 connected cells 11 Agents 12-15 Position in space 16. Weight Distribution 19 Adder Registers 20-22, 41, and 42 25 Event Type Detection Device 26. Spatial position detection device 27-34 Processing Items 36 Addition and Subtraction Machine 37. Multiply-accumulate circuit 38 Post-processing circuit 40 Multiplier 44 Accumulator
Claims
1. A memory circuit composed of multiple memory blocks, which allows parallel access to data by independently accessing each memory block, A calculation circuit that performs predetermined calculation processing on the aforementioned data, A control circuit for reading and writing data held in the memory circuit and for controlling the arithmetic processing in the arithmetic circuit, A processing unit equipped with, The memory circuit stores data relating to different events or features that are correlated with spatial information. The aforementioned arithmetic circuit performs predetermined arithmetic processing on multiple combinations of data read in parallel with respect to the data stored in the memory circuit. The control circuit is capable of performing parallel read and write operations on different memory blocks of the memory circuit with respect to multiple combinations of data corresponding to the predetermined arithmetic operations. A processing unit characterized by the following features.
2. The arithmetic processing apparatus according to claim 1, characterized in that the plurality of data corresponding to the predetermined arithmetic processing includes a combination of data distributed in a predetermined region that is associated with a single event or feature, and a combination of data at the same location that is associated with multiple events or features.
3. The arithmetic processing device according to claim 1 or 2, characterized in that the arithmetic processing is an addition or subtraction process of weight data having a distribution in a predetermined region.
4. The arithmetic processing device according to claim 3, characterized in that the weight data having a distribution in the predetermined region has a distribution that decreases monotonically from the center position.
Citation Information
Patent Citations
Memory control method, memory control device, image processor and program
JP2007279902A
Image memory system
JP2007323260A
Drone vision slam method based on GPU acceleration
JP2021077353A
Energy-efficient compute-near-memory binary neural network circuits
JP2021086611A
Address Generation for High-Performance Vector Processing
US20200356367A1