In-Memory Processing Based on Multiple Weight Sets

The crossbar array structure with local weight switching and prefetching in the crossbar array structure addresses the inefficiencies of conventional architectures, enhancing the speed and energy efficiency of matrix-vector multiplication.

JP7711327B2Active Publication Date: 2025-07-22アクセレラ エーアイ ビーヴィ
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024538653
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2025-07-22
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

Conventional computer architectures face challenges in accelerating and energy-efficiently performing matrix-vector multiplication due to the separation of processing and data storage, leading to high power consumption and data transfer congestion.

Method used

A crossbar array structure with N×M memory systems that store K sets of weights, enabling local weight switching and prefetching to reduce data exchange frequency, allowing for faster calculations and lower power consumption by performing MAC operations.

Benefits of technology

The proposed method significantly accelerates matrix-vector calculations by reducing data transfer frequency and power consumption, enabling efficient processing by locally enabling and accumulating weights without intermediate transfers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007711327000001
    Figure 0007711327000001
  • Figure 0007711327000002
    Figure 0007711327000002
  • Figure 0007711327000003
    Figure 0007711327000003
Patent Text Reader

Abstract

The invention is particularly directed to an in-memory processing method, the purpose of which is to perform matrix-vector calculations. The method relies on a device having a crossbar array structure (15). The latter comprises N input lines (152) and M output lines (153) interconnected at cross points defining N×M cells (155), where N≧2 and M≧2. The cells comprise respective memory systems, each of which stores K weights W i、j、k where K≧2. Thus, the crossbar array structure includes an N×M memory system capable of storing K N×M weight sets. To perform a multiply-accumulate (MAC) operation, the method first enables the N×M active weights for the N×M cells by selecting a weight from the K weights for each of the memory systems and setting the selected weight as the active weight. Then, a signal encoding a vector of N components is applied to the N input lines of the crossbar array structure. The latter then performs a MAC operation based on the vector and the N×M active weights. Finally, the resulting output signals at the outputs of the M output lines are read to obtain corresponding values. This allows different weight sets to be enabled locally in the crossbar array, which allows the rotation of weights to be performed locally and accordingly reduces the frequency of data exchange by the storage device. As a result, the idle time of the crossbar array structure is reduced. Thus, the proposed approach can substantially reduce the frequency of data transfer, which leads to faster computation. Furthermore, weights may possibly be read ahead while performing a MAC operation according to currently active weights. Among particular advantages, the look ahead step may be at least partially hidden by the pipelining. The present invention is further directed to related devices, systems, and computer program products.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The present invention generally relates to in-memory processing techniques (i.e., methods, devices, and systems) and related acceleration techniques. In particular, the present invention relates to an in-memory processing device including a crossbar array structure for performing matrix-vector multiplication by continuous partial product accumulation and coefficient prefetching in memory.

[0002] Matrix-vector multiplication is often required in many applications such as technical computing tasks and cognitive tasks. Examples of such cognitive tasks include training of cognitive models such as neural networks for computer vision and natural language processing and the inference performed thereby, as well as other machine learning models used for weather forecasting and financial forecasting.

[0003] Such operations pose multiple challenges due to their recurrence, generality, as well as size and memory requirements. On the one hand, it is necessary to accelerate these operations, especially in high-performance computing applications. On the other hand, it is necessary to realize an energy-efficient way of executing them.

[0004] Conventional computer architectures are based on the von Neumann computing concept, in which the processing function and data storage are separated into different physical devices. Such architectures have the drawbacks of congestion and high power consumption because data must be continuously transferred from the memory device to the control device and the arithmetic device through a physically constrained and costly interface.

[0005] As one possibility to accelerate matrix-vector multiplication, a dedicated hardware acceleration device such as a dedicated circuit having a crossbar array configuration may be used. This circuit includes input lines and output lines, which are interconnected at crosspoints defining cells. Each cell includes a respective memory device designed to store each matrix coefficient. The vector is encoded as a signal applied to the input lines of the crossbar array to cause a multiply-accumulate (MAC) operation on the latter's input lines. There are several possible implementations. For example, the coefficients of the matrix ("weights") can be stored in columns of cells. Adjacent to each column of cells is a column of arithmetic units that can multiply the weights by the input vector values (producing partial products) and finally accumulate all the partial products to produce the result of the full inner product. Such an architecture can simplify and efficiently map matrix-vector multiplication. The weights can be updated by reprogramming the memory elements to perform matrix-vector multiplication as needed. Such a solution breaks the "memory wall" and enables more efficient processing in or near the memory by integrating the arithmetic and memory devices into a single in-memory computing (IMC) device.

[0006] Devices having such a crossbar array structure are in daily use. Here, the inventor has set the task of improving such devices to enhance their energy efficiency and speed up calculations. SUMMARY OF THE INVENTION

[0007] According to a first aspect, the present invention is embodied as a method of in-memory processing aimed at performing matrix-vector calculations. The method depends on a device having a crossbar array structure. The latter includes N input lines and M output lines, which are interconnected at crosspoints defining N×M cells, where N≧2 and M≧2. Each cell includes a respective memory system designed to store K weights each, where K≧2. Thus, the crossbar array structure includes an N×M memory system capable of storing K sets of N×M weights. To perform a multiply-accumulate (MAC) operation, the method first enables N×M active weights for the N×M cells by, for each of the memory systems, selecting one weight from these K weights and setting the selected weight as the active weight. Next, a signal encoding an N-component vector is applied to the N input lines of the crossbar array structure. Thereby, the latter performs a MAC operation based on this vector and the N×M active weights. Finally, the output signals obtained at the outputs of the M output lines are read out to obtain the corresponding values.

[0008] The above scheme enables locally enabling different sets of weights in the crossbar array, thereby locally switching between active weights and accordingly reducing the frequency of data exchange by the storage device. As a result, the idle time of the core computer device, i.e., the crossbar array structure, is shortened. That is, some intermediate weight updates can be avoided because up to K consecutive calculation cycles can be executed without the need to transfer a new set of weights. Rather, the corresponding set of weights can be locally enabled as the active weights in each calculation cycle. Furthermore, partial results can be locally accumulated to avoid the transfer of intermediate results. Thus, the proposed approach can substantially reduce the frequency of data transfer, thereby bringing about an acceleration of calculations and a reduction in power consumption required to perform matrix-vector calculations.

[0009] In a particularly advantageous embodiment, the method further includes prefetching weights while performing a MAC operation according to the N×M weights that are currently active as weights. That is, when 1≦q≦K−1, instead of q of the N×M weight sets that were previously active, q of the N×M weight sets (i.e., the weights to be used next) are prefetched and stored in the N×M memory system. In other words, the weights can be preloaded (i.e., prefetched during a certain number of calculation cycles) with foresight in order to further shorten the idle time of the crossbar structure if necessary. Among the specific advantages, the prefetching step may be at least partially hidden in a pipelined manner.

[0010] Preferably, for 1≦k≦K, the N×M active weights are enabled by selecting, for each of at least a subset of the N×M memory systems, the k-th weight of the K weights of each memory system, and setting each of the thus selected weights as the currently active weight. As a result, the array structure can change context almost instantaneously.

[0011] In a typical embodiment, several matrix-vector calculation cycles are executed continuously. Each cycle includes operations as described above. That is, first, for each of the memory systems, a new set of N×M active weights is enabled for the N×M cells by selecting a weight from among its K weights and setting the selected weight as the active weight. Next, a signal encoding an N-component vector is applied to the N input lines of the crossbar array structure to cause the latter to perform a MAC operation based on the current vector and the new set of N×M active weights. Finally, the output signal obtained at the output of the M output lines is read out to obtain the corresponding value. Up to K such cycles can be executed without the need to transfer new weights to the memory system of the array.

[0012] Preferably, each of the cycles further includes continuously performing the accumulation by accumulating the results of the partial products corresponding to the read output signals. After completing several matrix-vector calculation cycles, the method may return the result obtained based on the continuous accumulation, for example, to an external (i.e., external to the array) storage device. Thus, there is no need to transfer intermediate results.

[0013] Optionally, new weights may be pre-read before completing K of the several matrix-vector calculation cycles. That is, the method may pre-read q sets of N×M weights for 1≦q≦K−1 and store the latter in an N×M memory system instead of q sets of N×M that were previously enabled as active weights. Interestingly, since the new weight values may optionally be pre-read in the middle, additional matrix-vector calculation cycles can be executed without interruption (beyond the K cycles) while continuously accumulating the partial results. Also, the pre-reading step is hidden in a pipelined manner. The final result can be returned at the very end of the entire matrix-vector calculation without being bothered by idle time due to intermediate data transfers.

[0014] The large operand can be addressed, for example, by decomposing the required operation into a K×T matrix-vector computation. That is, in the method, a K×T matrix-vector computation cycle may be executed, where T corresponds to the number of input vectors. Each input vector is decomposed into K sub-vectors of N components and associated with each of the K block matrices, the latter corresponding to K sets of N×M weights. Note that the sub-vectors actually correspond to the vectors incorporated above, and each sub-vector is a part of the input vector. In that case, the K×T matrix-vector computation cycle is executed as follows. First, K sets of N×M weights are loaded. The K sets of N×M weights correspond to the K block matrices. Accordingly, the memory system is programmed to store the K sets of N×M weights. Next, several operations are executed, which are for each of the K sub-vectors of each of the T input vectors. First, one of each of the K block matrices, i.e., the N×M weights corresponding to the block matrix associated with the current sub-vector, is enabled (as the currently active weights). Second, signals encoding the vectors corresponding to each sub-vector are applied to the N input lines, whereby, in a crossbar array structure, a MAC operation is executed based on each sub-vector and the currently active weights. Third, in the method, the output signal obtained at the output of the M output lines is read to obtain the corresponding partial value.

[0015] The reading preferably includes, if any, accumulating the partial value obtained for each sub-vector with the previously obtained partial value for the previous one of the K sub-vectors to obtain an updated result. Finally, in the method, the result obtained based on the last obtained updated result is returned.

[0016] In an embodiment, prior to encoding for the purpose of programming an N×M memory system according to K sets of N×M weights and applying a sub-vector to an input signal and then executing several matrix-vector calculation cycles by applying such input signals to N input lines, an external (i.e., external to the array) processing device is used to map a given problem to a fixed number of sub-vectors and a series of K sets of N×M weights. Note that the external processing device may optionally be integrated with a crossbar array structure. Alternatively, the device forms part of a separate device or machine.

[0017] The N×M memory system can be either a digital memory system or an analog memory system. In either case, MAC operations can be performed in parallel or as bit-serial operations.

[0018] In an embodiment where the N×M memory system is a digital memory system, each of the N×M cells further includes an arithmetic device connected to the corresponding one of the N×M memory system.

[0019] For example, the MAC operation can be performed bit-serial in P cycles where P≥2, where P corresponds to the bit-width of each of the N components of each vector (or sub-vector) used as input. In that case, partial product values are obtained, which are locally accumulated (in the crossbar array) upon completion of each of the P cycles. However, this accumulation should be distinguished from the accumulation that is performed upon completion of a vector-level operation, i.e., an operation related to the vector (or sub-vector), when processing several vectors (or sub-vectors) in succession.

[0020] According to another aspect, the present invention is embodied as a computer program for in-memory processing. The computer program product includes a computer-readable storage medium in which program instructions are embodied. The program instructions are executable by the processing means of an in-memory processing hardware device to cause the latter to execute any of the steps of the methods described above.

[0021] According to a further aspect, the present invention is embodied as an in-memory processing hardware device. To conform to the method, the device comprises a crossbar array structure including N input lines and M output lines interconnected at cross-points defining N×M cells, where N≧2 and M≧2. Each cell includes a respective memory system designed to store K weights, where K≧2. That is, the crossbar array structure includes N×M memory systems which, as a whole, are adapted to store K sets of N×M weights for performing MAC operations. The device further includes a selection circuit connected to the N×M memory systems. The selection circuit is configured to select one weight from each of the K weights of the memory systems and set the selected weight as an active weight to enable N×M active weights for the N×M cells. Additionally, the device includes an input device configured to apply a signal encoding an N-component vector to the N input lines of the crossbar array structure to cause the latter to execute a MAC operation based on this vector and the N×M active weights as enabled by the selection circuit during the operation. The device further includes a reading device configured to read the output signal obtained at the M output lines.

[0022] In an embodiment, each of the N×M memory systems is designed such that these K weights can be programmed independently. The device may further include a program circuit connected to each memory system. The program circuit is configured to program the K weights of the N×M memory system. Advantageously, when 1≦q≦K−1, the program circuit prefetches q weights of the N×M weight set that are not currently set as active weights, and accordingly, may be configured to program the N×M memory system such that the latter stores the prefetched weights instead of the q weights of the N×M weight set.

[0023] In a preferred embodiment, each of the N×M memory systems includes K memory elements, each of which is adapted to store the corresponding weight of the K weights, and the selection circuit includes an N×M multiplexer connected to each of the K memory elements of the corresponding one of the N×M memory systems, and selection control lines connected to each of the multiplexers such that any one of the K weights of each memory system can be selected and set as the active weight during operation.

[0024] Preferably, the selection circuit is further configured to select a subset of the n×m weights from one of the K sets of N×M weights by selecting, along with the k-th weight of the K weights of each memory system of a subset of the n×m memory systems of the N×M memory system, where 2≦n≦N, 2≦m≦M, and 1≦k≦K.

[0025] In an embodiment, the in-memory processing hardware device further includes a sequencer circuit and an accumulator circuit. The sequencer circuit is connected to the input device and the selection circuit and schedules the operations of the input device and the selection circuit so as to continuously execute several cycles of matrix-vector calculation based on one or more vector sets. During the operation, each of the cycles of matrix-vector calculation includes one or more cycles of MAC operations. Different N×M weight sets are selected from K N×M weight sets and set as the active weights of N×M in each cycle of matrix-vector calculation. The accumulator circuit is configured to accumulate the partial product values obtained when each MAC operation cycle is completed. Preferably, the accumulator circuit is arranged at the output of the output line.

[0026] In a preferred embodiment, each of the N×M memory systems includes K memory elements adapted to store the corresponding weights of K weights respectively. Each of the K memory elements of each of the N×M memory systems can be, for example, a digital memory element. In that case, each of the N×M cells further includes an arithmetic unit, which is connected to each of the K memory elements of the corresponding one of the N×M memory systems via the corresponding part of the selection circuit.

[0027] In an embodiment, each of the K memory elements of an N×M memory system is designed to store a weight of P bits. The input device is configured to supply an N-component vector bit-serially to the input lines in P cycles by applying the above signals, where each of the N components corresponds to a P-bit input word and P≧2. The N×M cells are configured to perform a MAC operation bit-serially in P cycles. In addition, the hardware device further includes an accumulator circuit configured to accumulate values corresponding to the partial bit-serial product values obtained in each of the P cycles. Also, the selection circuit is configured to maintain the same N×M set of weights as the active weights between each of the P cycles.

[0028] Preferably, the in-memory processing hardware device further includes a configuration and control logic unit connected to each of the input device and the selection circuit, a pre-data processing device connected to the configuration and control logic unit, and a post-data processing device connected to the output of the output line.

[0029] According to another aspect, the present invention is embodied as a computing system including one or more in-memory processing hardware devices as described above. Preferably, the computing system further includes a storage device and a general-purpose processing device connected to the storage device for reading and writing data to and from the storage device. Each of the in-memory processing hardware devices is configured to read and write data to and from the storage device. The general-purpose processing device is configured to map a given computing task to vectors and weights for the memory systems of one or more in-memory processing hardware devices.

[0030] These and other objects, features, and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the present invention, which should be read in conjunction with the accompanying drawings. The examples are provided to clarify the present invention so that those skilled in the art can easily understand the present invention in conjunction with the detailed description.

Brief Description of the Drawings

[0031]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 3C

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 6

Figure 7

[0032] The accompanying drawings show simplified representations of devices or portions thereof as included in the embodiments. Similar or functionally similar elements in the figures are assigned the same reference numerals unless otherwise indicated.

[0033] A computerized device, system, method, and computer program embodying the present invention will now be described by way of non-limiting examples.

[0034] The following description is structured as follows. General embodiments and high-level alternatives are addressed in Section 1, and in Section 2, particularly preferred embodiments are dealt with. Section 3 summarizes the final views. Note that this method and this alternative are collectively referred to as "this method". All reference signs Sn refer to the method steps of the flowchart of FIG. 7, and the reference numerals are related to the devices, components, and concepts included in the embodiments of the present invention.

[0035] 1. General embodiments and high-level alternatives A first aspect of the present invention will now be described with reference to FIGS. 2A - 4B and FIG. 7. This aspect relates to a method of in-memory processing, the objective of which is to accelerate multiply-accumulate operations or MAC operations.

[0036] The method will rely on devices 10, 10a having crossbar array structures 15, 15a. The crossbar array is shown explicitly in FIG. 2A. This structure 15, 15a includes N input lines 152 and M output lines 153, where N≥2 and M≥2. The input lines 152 and the output lines 153 are interconnected at cross points (junctions). The cross points thereby define an N×M cell 155. Each cell 155 includes a respective memory system 157. It should be noted that each memory system 157 is designed to store K weights, where K≥2. In practice, K may typically be equal to 4, 8, 16, or 32 (as assumed in FIGS. 3A, 3B, and FIGS. 5A, 5B). Overall, the crossbar array structure 15 includes an N×M memory system capable of storing K sets of N×M weights, i.e., a total of K×N×M weights.

[0037] In other words, the crossbar array structure 15 includes N×M cells 155 in a crossbar configuration, where each cross point of the crossbar configuration corresponds to a cell, and each cell includes a memory system 157 capable of storing K weights. Such weights are W in FIG. 2A i、j、kIt is described as follows, where i ranges from 1 to N, j ranges from 1 to M, and k ranges from 1 to K. In practice, the number of input lines 152 and output lines 153 typically ranges from approximately several hundred to several thousand lines. For example, arrays of 256×256, 512×512 (such as in FIGS. 4A and 4B), or 1024×1024 can be considered, but N does not necessarily have to be equal to M. The concepts of input lines and output lines will be further described below.

[0038] The proposed method basically focuses on enabling certain weights before performing MAC operations based on a given vector and matrix coefficients corresponding to the enabled weights. That is, the N×M weights are enabled for the N×M cells 155 in step S70 (see the flowchart of FIG. 7). This is achieved by selecting a weight from these K possible weights for each memory system and then setting the selected weight as the active weight. Note that the selection and setting of weights can be performed in a single operation, especially when using a selection circuit 159 that depends on multiplexers connected to each memory system, as in the embodiments discussed below with reference to FIG. 3C.

[0039] Once the weight set is enabled, the vector components are injected into the crossbar array structure 15 (step S82). More precisely, the signals encoding an N-component vector (hereinafter referred to as the N-vector) are applied to the N input lines 152 of the crossbar array structure 15 in S82. Thereby, the crossbar array structure 15 performs a MAC operation based on the N-vector and the currently enabled N×M active weights in S84. As a result of the MAC operation, each value encoded by the signals supplied to the N input lines is multiplied by the currently active weight values enabled from the K weight sets stored in the memory system 157.

[0040] Regarding the crossbar configuration, M MAC operations are executed in parallel during each calculation cycle. Note that the operations executed for each cell correspond to two scalar operations, namely, one multiplication and one addition. Thus, M MAC operations imply N×M multiplications and N×M additions, which means a total of 2×N×M scalar operations.

[0041] The output signals obtained on the M output lines 153 are then read out in step S90 to obtain the corresponding values. In practice, it is often necessary to execute several calculation cycles continuously, whereby the weights are locally made valid (i.e., selected and set as active weights) cycle by cycle before supplying the components of the N vectors, executing the MAC operations, and reading out the output values. Such output values may advantageously correspond to partial values that can be locally accumulated in the devices 10, 10a. In that regard, the read operation to be performed should be understood in a broad sense. The read operation may aim not only to extract the output values, but also to accumulate them (if necessary) with previous output values and / or to store such values.

[0042] It should be noted that different weight sets can be locally enabled in the crossbar array 15 by the proposed scheme, thereby locally performing weight rotation and accordingly reducing the frequency of data exchange by a memory device externally attached to or incorporated in the devices 10, 10a. This also shortens the idle time of the devices 10, 10a. That is, some intermediate weight updates are avoided because up to K consecutive calculation cycles can be executed without the need to transfer a new weight set. Rather, the corresponding weight sets are locally enabled as active weights in each calculation cycle. Further, the weights can optionally be preloaded (i.e., pre-read during the calculation cycle) to further shorten the idle time of the crossbar structure 15 if necessary. Thus, the proposed approach enables a substantial reduction in the frequency of weight data transfer, thereby resulting in faster calculations. Also, since partial results can be locally accumulated, such results do not need to be transferred either, thereby reducing the power consumption of the devices 10, 10a.

[0043] To state the views in order. Each memory system 157 preferably includes K different memory elements for the sake of simplicity. Such elements can be connected so as to be independently programmable. This allows the weights (to be used next) to be pre-read as in the preferred embodiment discussed below. The memory elements can be programmed to store binary data or multi-bit data, for example, similar to the synaptic weights of a synaptic crossbar array structure.

[0044] The weights are related to numerical values and represent matrix coefficients. Such weights need to overcome (a part of) the problem to be solved and accordingly be programmed into the memory system 157. In that regard, the hardware device 10 may advantageously include a programming circuit 158 (FIG. 3B) that programs an N×M memory system 157 to be configured to store each of the K sets of weights. The programming circuit may be controlled, for example, by the logic device 12 in the devices 10, 10a. Alternatively, the programming circuit may be external, in which case the device may include pads and traces dedicated to programming the memory elements.

[0045] Programming the memory elements of the IMC device is known per se. However, in this context, what is required is to appropriately and timely program several memory elements (or several memory values of the memory system) for each cell. Further required is to appropriately select the active weights. For that purpose, a selection circuit 159 (FIGS. 3B, 3C) can be used to perform the selection of the required weights and enable the selected weights as active weights.

[0046] The vectors used in the input (also referred to above as N-vectors) each have N components according to the number N of input lines. Such vectors may in fact correspond to parts of larger input vectors. That is, the problem to be solved (e.g., matrix-matrix multiplication) may typically involve large operands. Thus, in a first problem, it may be necessary to decompose into smaller matrix-vector operations that include input vectors and parts of matrices partitioned according to the size of the array 15. For example, the input matrix may be decomposed into input vectors, which themselves may be decomposed into sub-vectors (i.e., N-vectors), to which each block matrix is assigned and for the purpose of performing a plurality of operations, the outputs of which can finally be reconfigured to form the final result.

[0047] Thus, in practice, the basic operation principle is to supply N vectors to the array 15 and perform matrix-vector operations based on the supplied vector components on the one hand and the currently active weights on the other hand, the latter being properly enabled according to the current N vectors.

[0048] To perform the MAC operation, the input signal is applied to N input lines, and the components of the N vector are encoded by this signal. That is, each input signal encodes a different vector component and is applied to the corresponding input line. The input signals correspond to the so-called data channels in the synaptic crossbar structure. Each vector component and each matrix coefficient can be encoded, for example, as a P-bit value. That is, the MAC operation can be performed bit-serial (as assumed in FIG. 4A or FIG. 7) or in parallel (as in FIG. 4B). In the bit-serial implementation form, as will be discussed in detail later with reference to the preferred embodiment, each multiplication operation is performed in a P-bit serial cycle (for example, P = 8 or 16). Alternatively, each P-bit word (vector component) is injected in parallel into M cells of the corresponding input line.

[0049] This approach supports analog memory elements and analog operations. In an analog electrical implementation, digital inputs are converted to an analog representation by a digital-to-analog converter (DAC) or a pulse-width modulator (PWM) and then applied to the input lines. Each cell operation typically corresponds to a single analog operation, whereby the input signal is multiplied by the weight values transmitted by the memory component, resulting in an electrical interaction with that component and branching to the output of a column, thus efficiently providing an analog addition operation. The same principle is available for optical input signals. The digital implementation relies on a digital memory system. That is, the N×M memory system 157 is a digital memory system (e.g., each containing K digital memory elements). In that case, each of the N×M cells 155 includes an arithmetic device 156 (including a multiplier and an adder tree) connected to the corresponding memory system 157 as assumed in FIGS. 2A, 3A, and 3B.

[0050] To be complete, note that the input line 152 refers to the channel through which data is communicated to the M cells by the signal. However, such an input line does not necessarily correspond to the number of physical conductors or logical channels required to actually reach the cells. In a bit-serial implementation, each input line may include a single physical line, which is sufficient to supply an input signal that transmits N vector component data. However, in an approach for parallel data collection, each input line may include up to P parallel conductors, each of which is connected to the M cells of the corresponding input line. In such a case, P bits are injected in parallel to each of the M corresponding cells via the parallel conductors. Still, various intermediate configurations can be considered that include both parallel and bit-serial supply of data.

[0051] The hardware devices 10, 10a are preferably fabricated as an integrated structure, e.g., a microchip, incorporating all the components necessary to perform the core computational steps. Such components notably include input devices 151, 151a (FIGS. 4A, 4B) for applying an input signal to N input lines 152, a program circuit 158 (FIG. 3B), a selection circuit 159 (FIGS. 3B, 3C), and readout devices 154, 154a (FIGS. 4A, 4B) which may include an accumulator. The selection circuit 159 and the input device 151 may, for example, as in an embodiment, form part of the same configuration and control logic circuitry and be controlled by the same logic device 12 (FIG. 2B). The devices 10, 10a themselves relate to another aspect of the present invention, which will be described in detail later.

[0052] For a particular embodiment of the method, it is discussed here. First, the weights can be pre-loaded (i.e., pre-read during the computational cycle) to further shorten the idle time of the crossbar structure, if necessary. The pre-read mechanism is shown in FIG. 6. As a very significant advantage, the step of pre-reading may be (at least partially) hidden by a pipelining approach.

[0053] Specifically, during the current computational cycle, i.e., while performing the MAC operation according to the currently active N×M weights that are currently valid, new weights can be pre-read. If 1≦q≦K−1, instead of the q sets of the previously active N×M weights, up to q sets of N×M weights (i.e., the next weights to be used) can be pre-read S115 and stored in the N×M memory system 157.

[0054] Before starting the matrix-vector calculation cycle, K sets of N×M weights can first be loaded in the array. Thus, the subsequent pre-reading steps are typically executed iteratively. By pre-reading the weight sets, a look-ahead approach is enabled, which allows for a further acceleration of the calculation, as shown in FIG. 6.

[0055] For example, assume that one or more matrix-matrix multiplications must be performed. In the example of FIG. 6, four consecutive different matrix-matrix multiplications are performed, each requiring T computational cycles and using only a single set of weights. In this case, for example, it is possible to repeatedly load the set of weights from external memory to prefetch the weights for the next matrix-matrix. As another example, the first matrix multiplication uses two sets of weights, while the third and fourth matrix multiplications require only one set of weights. In that case, the set of weights for the third and fourth matrix multiplications can be prefetched while the two sets of weights for the first matrix multiplication are being used for the computation. Such a prefetch strategy, as opposed to previous crossbar array structures that utilize only a single set of weights, (as shown in FIG. 6) can potentially lead to significant speedup because the time to load the weights can be partially hidden in a pipelined fashion.

[0056] In that regard, note that the cycles shown at the top of FIG. 6 correspond to the cycles executed in a normal crossbar array, where each cell stores a single weight value in this case. The load step (for loading the weights into this array) and the processing step must be interleaved, thereby making it impossible to hide the load time in a pipelined fashion. In contrast, in an array that stores multiple sets of weights, the matrix coefficients can optionally be preloaded into unused memory elements while other sets of weights are currently active. This can result in a further significant speedup in processing time compared to a system that relies on a single set of weights.

[0057] Referring to FIGS. 3B, 3C, and 7, it is preferable that the required weights can be advantageously enabled all at once, preferably almost instantaneously. For example, for an N×M weight, when 1≦k≦K, the k-th weight of the K weights of each memory system 157 can be selected along with it, and each weight thus selected can be set as the currently active weight, thereby enabling it to be effectively used. Depending on the situation, the entire N×M memory system 157, or only a subset thereof, for example, the selection of weights for a subarray, may be performed. For example, when multiplying a vector smaller than the L-component norm (L<N) by an L×L matrix, the corresponding weight array can be selected and set without the need to enable the remaining weights, because the remaining components of the vector smaller than the norm can be set to zero (zero-padding). In all cases, some weights can optionally be selected and set along with it. As a result, the change in context is almost instantaneous. That is, switching from one set of weights to another does not incur any substantial downtime. This can be achieved, for example, by a selection circuit including a multiplexer 159 as shown in FIGS. 3B and 3C.

[0058] In practice, as assumed in FIGS. 5A, 5B, and 7, some calculation cycles may have to be executed continuously. Up to K cycles can be executed continuously by locally switching the weights. That is, the method may involve executing S60 to S100, some matrix-vector calculation cycles, where each cycle includes: (i) enabling S70 a new set of N×M active weights; (ii) performing an MAC operation S84 based on the enabled weights and the associated N vectors; and (iii) reading S90 the output signals obtained at the outputs of the M output lines 153 to obtain the corresponding values. Each time, the new set of N×M weights is enabled for the N×M cells 155 by selecting a weight from its K weights for each memory system 157 and setting the selected weight as the active weight S70. To perform the MAC operation S84, a signal encoding the N vectors is applied to the N input lines 152 of the crossbar array structure 15 S82, whereby the latter performs the MAC operation S84 based on this N vector and the new set of N×M active weights.

[0059] That is, the input signal is repeatedly applied to the N input lines 152 to continuously supply the N vectors and cause the MAC operation to be performed accordingly. As described above, each N vector may actually correspond to a portion of a larger input vector (e.g., from a given input matrix), to which a corresponding block matrix is assigned as shown in FIG. 5A. In that case, the N vectors supplied to the crossbar array in each matrix-vector calculation cycle are different. In other cases, the same N vector may be applied several times in succession, depending on the operation decomposition scheme determined upstream of step S30 (FIG. 7).

[0060] In all cases, the new active weight set can be locally enabled for each of these in a matrix-vector calculation cycle that is executed without undergoing an intermediate program step for changing the weights. As mentioned above, this can of course involve possible look-ahead operations, which are nevertheless hidden in a pipelined fashion. That is, q sets of N×M weights can optionally be pre-read S115 and stored (instead of the previous q sets of weights) before completing K matrix-vector calculation cycles. For example, a single N×M weight set can be pre-read (q = 1) in each iteration. Alternatively, two N×M weight sets can be pre-read after completing every second iteration. Various other pre-read schemes are possible. Note that such pre-read schemes can optionally be applied dynamically depending on the workload.

[0061] Also, for operations involving large operands, the partial results obtained at the end of each intermediate cycle can advantageously be locally accumulated. That is, after a part of the calculation cycle, the result of the partial product can be accumulated S90 in devices 10, 10a (at the output of the crossbar array structure 15). That is, the accumulation is performed continuously for the purpose of later reconstructing the final result. The final result is obtained based on the continuous accumulation. The final result can be returned to the external storage device 2, for example, after completing a certain number of calculation cycles. Interestingly, since new weight values can optionally be pre-read S115 in the middle, further matrix-vector calculation cycles can be executed without interruption while continuing to accumulate the partial results.

[0062] Thanks to partial accumulation, only the final result of the matrix-vector calculation needs to be transferred and written to the external memory without incurring idle time due to the transfer of updated weights and intermediate partial results. When matrix-matrix multiplication is performed, the final result of each matrix-vector product may optionally also be stored locally at device 10. In matrix-matrix multiplication, only the result then has to be returned to the external memory device 2. In both cases, some results are obtained locally based on successive accumulation before transferring the final result to an external entity. The partial results do not need to be transferred to an external entity and are actually deleted without even having to be stored, due to successive accumulation.

[0063] Considering the flow of FIG. 7, the example of FIG. 5A is examined. Here, the aim is to perform matrix-matrix multiplication. This, as previously described, is decomposable into K×T matrix-vector calculation cycles, where T corresponds to the number of columns of one of the input matrices and accordingly is decomposed into T input vectors. And also, each input vector is decomposable into K sub-vectors, i.e., N-vectors of N components each, in which case each N-vector is associated with each block matrix. That is, each input vector of K×N components is associated with K block matrices, the latter corresponding to K sets of weights of N×M. Then, the K×T matrix-vector calculation cycle can be executed as S50 to S110 as follows. First, K sets of weights of N×M (corresponding to K block matrices) are loaded S55, and accordingly programmed in the memory system 157, which has to store the K sets of weights of N×M. Next, for each of the T input vectors, for each of their N-vectors (i.e., each of the K sub-vectors, see steps S60 to S100), the calculation cycle is executed (see steps S58 to S110). That is, the loop for the N-vectors can be nested within the loop for the input vectors, and these input vectors themselves can be nested within the loop for the input matrix as required (steps S50 to S120).

[0064] (Regarding the N vector), the innermost loop follows the same principle as described above. That is, the N×M active weights are enabled as the currently active weights S70, and this weight corresponds to the block matrix associated with the current N vector as previously assigned. Then, the signal encoding the current N vector is applied to the N input lines 152 S82, and the crossbar array structure 15 executes a MAC operation based on this N vector and the currently active weights S84. The output signal obtained at the output of the M output lines 153 is also further read out to obtain the corresponding partial value, which is advantageously accumulable in the device 10 S90. That is, the partial values obtained for each N vector (but the first one) are locally accumulable with the partial values previously obtained for the previous N vectors S90. In this way, an updated result is obtained in each cycle. The finally obtained updated result is ultimately returned to the external memory S120.

[0065] Such an operation is visually shown in Figure 5A, where K is assumed to be equal to 4 in this example. That is, the matrix-matrix multiplication will be calculated in four calculation cycles of T, and each of the T input vectors has 4×N components as seen in Figure 5A. Further, the sequence of operations assumes T = 4 input vectors in this example. Each block matrix is stored as a single set of N×M weights. In the arithmetic unit in the IMC array, a partial inner product is calculated for each pair of the N vector and the associated set of weights. Figure 5A shows how the context (i.e., the set of weights) is switched in each calculation cycle in the array and how the complete result is locally calculated in the accumulator. The final result is not written back to the external memory until after the last iteration. That is, the accumulator accumulates four partial products continuously and finally writes back the result to the external memory. This process is repeated T times, i.e., once for each input vector.

[0066] Figure 5B shows the timing. First, all weight sets (WS0 to WS3) are continuously loaded into the IMC, which corresponds to step S55 in FIG. 7. Next, the calculation cycles are started in an interleaved manner. Only four input vectors are included (T = 4), and each is decomposed into K = 4 sub-vectors of N components. The weight sets WS0 to WS3 are successively enabled (i.e., rotated) according to each of the K parts of each of the T input vectors.

[0067] In this example, since four available weight sets are used, all the required weight sets can be pre-loaded in array 15 in advance without any look-ahead being required, i.e., S55. That is, the decomposition assumed in the flow of FIG. 7 presumably does not require weight look-ahead because the matrix-vector operations already utilize the rotation of all K weight sets in this example, leaving no room (and in fact no need) to look ahead to the next weight set.

[0068] However, look-ahead can be advantageous if the input vectors have to be decomposed into more than K parts. In addition, in the context of FIGS. 5A and 7, new weights can optionally be looked ahead during the processing of the last K sub-vectors corresponding to the last input vector. That is, upon completion of the operation cycle for any of these K sub-vectors, an instruction can be given to look ahead to a new weight set and write it in place of the previously active N×M weights. This shortens the idle period (corresponding to step S55) before starting the calculations associated with another matrix-matrix multiplication, i.e., before S50.

[0069] Of course, FIGS. 5A and 7 reflect one possible mapping of the operations. Various other computational strategies leading to various sequences of operations may be considered at step S30. Further, it should be noted that regardless of the decomposition scheme employed at runtime, the core computer device 10 can generally be designed to enable look-ahead of weights as needed.

[0070] The optimal operation mapping is determined at S30 by an external processing device 2, 13, i.e., a device separate from the core computer array 15. Still, this external processing device 13 may optionally be integrated with the core IMC array 15 in devices 10, 10a as assumed in FIG. 2B. In all cases, the processing devices 2, 13 are used to determine the computational strategy (i.e., identify and associate sub-vectors and block matrices), and this operation may also be referred to as a conditional operation. In practice, this operation will map a given problem to a fixed number of sub-vectors and a set of N×M weights K at S30. This step is then executed to program the N×M memory system 157 for subsequent execution of MAC operations at S55 and to encode the calculated vectors into the input signals before the calculated vectors are input. The processing devices 2, 13 may optionally perform other tasks as discussed later.

[0071] In an embodiment, the MAC operation is performed bit-serial, i.e., in P serial cycles at S84, where P≧2. In practice, P is typically 2 requal to, where 3 ≦ r ≦ 6. Assume P is equal to 8 in the example of FIG. 4A. The value P corresponds to the bit width of each of the N vectors used at input. Note that for each of the P cycles completed, a partial product value needs to be locally accumulated S86 in that case. The calculation cycles S82~S88 are sub-cycles that need to be distinguished from the matrix-vector calculation cycles (S50~S100), and these themselves may benefit from a partial accumulation S90. That is, each internal calculation cycle S80 includes P cycles, while the matrix-vector calculation cycles include K cycles (themselves nested in T cycles).

[0072] Alternatively, the method depends on a parallel implementation form (see FIG. 4B), which does not require any parallel-to-serial conversion as in FIG. 4A. In the parallel implementation form, each N vector is processed by weighted multiplication in a single cycle. In yet another alternative, a hybrid approach involving a parallel supply of bit-serial values can be considered.

[0073] Another aspect of the present invention relates to a computer program for in-memory processing. The computer program product includes a computer-readable storage medium in which program instructions are embodied, where the program instructions are executable by the processing means 12, 13, 14 of the in-memory processing hardware devices 10, 10a to cause the latter to execute the steps described above, as necessary, including MAC operations S84, as well as accumulations S86, S90 and look-ahead S115 operations. More generally, such processing means 12, 13, 14 process some (or optionally all) of the pre-processing and post-processing operations, as suggested in FIG. 2B. These operations are executable, for example, on an instruction-based processor or a dedicated accelerator having various instruction- or command-based control mechanisms.

[0074] Apart from the operation S30 aimed at determining the calculation strategy, the device 13 may optionally perform other tasks related to, for example, element-by-element operations or non-linear operations. For example, in a machine learning (ML) application, the device 13 may perform feature extraction to convert some input data (e.g., an image, an audio file, or text) into a vector, which is then used to train a cognitive model or for inference purposes using the crossbar array structures 15, 15a. In that regard, one or more neuron layers may optionally be mapped onto the arrays 15, 15a according to the partitions of the arrays. Still, the devices 12 - 14 may optionally collect the output from the arrays, (if necessary) process such output, and reinject them into the arrays as new inputs to map, for example, multiple layers of a deep neural network that need to be executed. ML operations (such as feature extraction) may notably need to perform depthwise convolution, pooling / unpooling, etc. Similarly, the post-processing device 14 may be utilized to perform affine scaling of the output vector, apply a non-linear activation function, etc.

[0075] More generally, the devices 12 - 14 may perform various operations, which depend on the actual application. Also, such operations may be partially executed in the client device 3 and the intermediate device 2. Various calculation strategies that may depend on the application can be devised.

[0076] Referring again to FIGS. 1 - 4B, a further aspect of the present invention related to the in-memory processing hardware devices 10, 10a will now be described in detail. The functional and structural features of this device have already been described with respect to the method. Such features will only be briefly described below.

[0077] To conform to this method, devices 10, 10a comprise a crossbar array structure 15 as shown in Figure 2A, etc. That is, the array 15 includes N input lines 152 and M output lines 153 that are interconnected at cross points defining N×M cells 155. Each cell 155 includes a corresponding memory system 157, each of which is designed to store K weights. The array 15 is designed to perform MAC operations.

[0078] Devices 10, 10a further include a selection circuit 159 (partially) shown in Figure 3B, etc. The selection circuit 159 is connected to the N×M memory system 157. This circuit 159 is generally configured to select a certain weight from the K weights of each memory system and set the selected weight as the active weight. Thereby, N×M active weights can be enabled for the N×M cells 155.

[0079] Devices 10, 10a also include input devices 151, 151a configured to apply signals that encode an N-vector to the N input lines 152 of the array 15. Thereby, during the operation, the array 15 performs a MAC operation based on the N-vector and the corresponding N×M active weights so as to be enabled by the selection circuit 159.

[0080] Furthermore, the reading device 154 is configured to read the output signals obtained at the outputs of the M output lines 153 and, if necessary, accumulate partial output values, as discussed above. Also, the reading device should be understood in a broad sense. For example, the device may include accumulators 154, 154a, and / or memory elements that store such output values. In an analog implementation form, the reading device may further include an analog-to-digital converter.

[0081] Each of the N×M memory systems 157 is preferably designed such that these K weights are independently programmable. As seen in FIG. 3B, devices 10, 10a may each include a program circuit 158 connected to the respective memory system 157. As described above, only a portion of the program circuit 158, which is a portion connected to the corresponding memory system, is shown in FIG. 3B. Overall, the program circuit 158 is configured to program the K weights of each of the N×M memory systems. Since the K weights of each memory system are independently programmable, any of the K weights that are not currently set as active weights can be optionally (re)programmed even if another of the K weights is currently set as the active weight, thereby enabling the weights to be preloaded (read ahead) with foresight during the operation.

[0082] That is, the program circuit 158 may preferably pre-read q of the N×M weight sets that are not currently set as active weights when 1≦q≦K−1, and accordingly, program the N×M memory system 157 such that the latter stores the pre-read weights instead of the q of the N×M weight sets. Thus, during the operation, while the crossbar array structure 15 is already performing a MAC operation based on the currently active weights, the program circuit 158 may program each of the N×M memory systems 157 to change the weights that are not currently set as active weights.

[0083] Note that the program circuit 158 must be sufficiently independent of the calculation circuit so that it can reprogram the weights that are not currently active with foresight while the calculation circuit 15 is performing a MAC operation based on the currently active weights. This independence enables the weights required for the next operation cycle to be preloaded with foresight. The pre-read operation can be performed, for example, on several weight sets at a time. As described above, various pre-read schemes are possible.

[0084] The program circuit 158 can be connected to the configuration and control logic circuit 12 of the local memory device 11, for example, as assumed in FIG. 2B. In an analog implementation form, in principle, the input line 152 can be reused to program the memory system 157, but it is preferable to provide a separate program circuit so that the memory system 157 can be reprogrammed during the calculation cycle. Similarly, a digital memory cell (i.e., a cell including a digital memory element) can be connected to dedicated lines that embody word lines and bit lines for writing operations in, for example, an SRAM memory device. It should be noted that, in contrast to the illustration of FIG. 3B, the selection circuit 159 may, in some cases, reuse word lines and bit lines for read operations. Therefore, the selection circuit and the program circuit may actually partially overlap.

[0085] In the examples of FIGS. 3A and 3B, each of the N×M memory systems 157 includes K different memory elements for the sake of simplicity. Each memory element is adapted to store a corresponding weight. In that case, the selection circuit 159 may include an N×M multiplexer. Each multiplexer is connected to all the memory elements of the corresponding memory system 157 as shown in FIG. 3B. Further, by connecting the selection control lines to each multiplexer, any one of the K weights of each memory system 157 can be selected and set as the active weight during the operation. Selection bits can be transmitted through the control lines to select the active weight as shown in FIG. 3C.

[0086] In the example of FIG. 3C, the multiplexer is a channel multiplexer that uses an inverter and a logic "NAND" gate to reach the common output X. That is, the combinational logic circuit switches one of several input lines A, B, C, D to a single common output line X. The data lines A, B, C, D are W in FIG. 3B 1、1、0 、W 1、1、1 、W 1、1、2 、W 1、1、3It corresponds to (transmitting a binary input address). The data selection lines are determined by Add0 and Add1 corresponding to the least significant bit (LSB) and the most significant bit (MSB), respectively. An N×M multiplexer is used to switch the weights of the N×M memory system. In principle, in order to enable individual control of each multiplexer, there are at most 2×N×M control lines, that is, two control lines, for each multiplexer 159. However, in practice, the control lines can be shared across all multiplexers if it is desired to select the weight sets simultaneously, especially as discussed below. Thus, all control lines are preferably shared, which enables the same index K to be selected simultaneously for each element in the M×N memory system. In such a case, the number of control lines can be reduced to Log2(K) lines.

[0087] Similarly, the program circuit 158 may include an N×M demultiplexer, in which case the same control bit lines are used throughout the array 15. Also, for the sake of simplicity, in FIG. 3B, a single demultiplexer 158 connected to the corresponding memory system 157 is shown. However, in practice, there are N×M demultiplexers 158 and N×M multiplexers 159 connected to each memory system 157. Alternatively, the program circuit 158 and the selection circuit 159 may include other types of electronic components, in which case such components are arranged in or at least connected to each cell as necessary to program and select the memory values. In yet another alternative, each memory system is configured to store K different values at corresponding local addresses instead of being composed of K different memory elements.

[0088] As described above, the required weights are preferably the S70 enabled all at once. For that purpose, the selection circuit 159 can advantageously be configured to select (at least) a subset of the n×m weights from one of the K sets of N×M weights. This is most effectively achieved by selecting, for each memory system of the subset of the n×m memory systems, the k-th weight of the K weights, where 2≦n≦N, 2≦m≦M, and 1≦k≦K. As shown above, enabling the weights of the n×m subarray can be advantageous for those matrix-vector calculations where not all of the N×M weights have to be switched, which depends on how the problem is first mapped onto the N×M cells 155. Note that the switching operation may rarely have to be performed for a single cell (i.e., n = 1 and m = 1). However, in practice, the selection of the weights will generally be performed simultaneously for a large subset of the N×M memory system (i.e., n>1 and m>1), or even for all of the N×M memory systems containing particularly large operand matrices, as in the examples of the applications discussed above with respect to FIGS. 5A, 5B, and 7. However, alternatively, the selection circuit 159 can systematically select a set of N×M weights from one of the K sets of N×M weights by selecting, for each of the K weights of each N×M memory system, the k-th weight to systematically switch all of the memory systems 157 at the same time. Thus, generally, the selection circuit 159 is configured to select a set of n×m weights and set the latter as the active weights for the n×m memory systems of the array 15, where 1≦n≦N and 1≦m≦M.

[0089] Devices 10, 10a typically include a sequencer circuit connected to input devices 151, 151a and a selection circuit. The sequencer circuit schedules the operations of input devices 151, 151a and the selection circuit to continuously execute several cycles of matrix-vector calculation as described above. That is, such operations are based on N vectors. Each cycle of matrix-vector calculation includes one or more cycles of MAC operations (depending on whether the MAC operations to be executed are supplied bit-serial) and different N×M weight sets, where the latter are selected from K N×M weight sets and set as the active weights of N×M in each cycle. The sequencer circuit, program circuit 158, and input circuit 151 preferably form part of the same configuration and control logic circuit, typically including on-chip logic device 12 as assumed in FIG. 2B. That is, the sequence function is preferably executed by logic device 12, similar to other configuration and control functions.

[0090] Furthermore, devices 10, 10a may include accumulator circuits 154, 154a configured to accumulate the partial product values obtained when each matrix-vector calculation is completed. Also, in bit-serial applications, each cycle of matrix-vector calculation includes several MAC cycles S80 by bit-serial operations, as in the embodiment previously discussed with respect to FIG. 7. Here, further accumulation S86 must be performed. In another form depending on the parallel implementation (which does not require parallel-serial conversion 151a), weight multiplication is performed in a single cycle for a complete input. In all cases, the active weights remain in the same state during each matrix-vector calculation cycle.

[0091] As can be seen in FIGS. 4A and 4B, the accumulator circuits 154, 154a can be arranged at the output of the output line 153. The accumulator is known per se. The accumulator circuits 154, 154a may, notably, form part of a reading device (not shown). Alternatively, the accumulator may, notably, be arranged in each cell, as the case may be, in order to accumulate S86 the values obtained during bit-serial arithmetic at the level of each cell.

[0092] As described above, each of the N×M memory systems 157 preferably includes K different memory elements, for the sake of simplicity. Each memory element is adapted to store a corresponding weight. Such a memory element may, notably, be a digital memory element, such as a static random access memory (SRAM) device. Alternatively, the memory element is an analog memory element. In that case, each multiply-accumulate operation, i.e., Σ i W i、j、k x i is performed by analogy and the output signal is transferred into the digital domain using an analog-to-digital converter (ADC) circuit, if necessary. The memory element may optionally be a non-volatile memory element. More generally, the present invention is applicable to various types of electronic memory devices (e.g., SRAM devices, flash cells, memristor devices, etc.). Any type of memristor device, such as a phase change memory cell, a resistive random access memory (RRAM), and an electrochemically random access memory (ECRAM) device, may be contemplated.

[0093] In a preferred embodiment, each of the K memory elements is a digital memory element such as an SRAM device. In that case, each cell 155 further includes an arithmetic unit 156 (including a multiplier and adder tree) connected to each of the K memory elements of the corresponding memory system 157 via a corresponding selection circuit portion 159 (e.g., via a multiplexer). Each cell is physically connected to each memory element via a selection circuit component (such as a multiplexer or other selection circuit component), but note that it is logically connected to only one such element at a time by the selection made by the selection circuit.

[0094] In the bit-serial implementation form (Figure 4A), each memory element is designed to store a P-bit weight. The input device 151 is configured to supply the components of the N-vector bit-serially to the input line 152 in P cycles (P≥2) by applying an input signal, and each vector component corresponds to a P-bit input word. The N×M cells 155 must also be designed to perform the MAC operation bit-serially (i.e., in P cycles). Thus, the hardware device 10 must include an accumulator circuit 154 for accumulating values corresponding to the partial bit-serial product values obtained in each of the P cycles, which corresponds to step S86 in Figure 7. On the other hand, the selection circuit 159 must maintain the same set of N×M weights as the active weights during each of the P cycles.

[0095] As an example of an implementation form, (i) the crossbar array is an array of N×M = 512×512, (ii) K = 4, whereby, in total, four switchable N×M weight sets are available, and (iii) as shown in FIG. 4A, assuming that the bit width (IBW) of the input sample is equal to P = 8 bits, the bit width (WBW) of the weight is also equal to 8 bits. The IMC device 10 takes in an N-vector of 512 components (each component corresponding to an 8-bit input word). Each vector component is supplied serially in bits over P = 8 cycles. The IMC device 10 maps a total of 512×512×4 weights (each 8 bits). Only one of the weight sets is active for processing during each of the P cycles S80. As described above, the serial bit cycle S80 must not be confused with the matrix-vector calculation cycles (S60 to S100). In each cycle S80, M weights (one per column) are multiplied by all N input bits (one per row). The resulting partial products can be stored as 17-bit values (with WBW + Log2(N) = 17), which are accumulated in the accumulator 154 below the IMC15 over IBW = 8 cycles and finally generate the complete vector product, i.e., the final result for this calculation cycle S80 (S88: yes).

[0096] Also, in the arithmetic unit in the IMC array, a partial inner product is calculated for each pair of the N vectors and the associated block matrices. The accumulator 154 may be further used to accumulate S90 the K partial products before writing back the final result to the external memory. The IMC switches contexts (weight sets) in each matrix-vector calculation cycle. This process can be repeated for each input vector. For example, a programmable accumulator can be programmed to accumulate some intermediate result values obtained, for example, after shift and inversion operations. Thus, when K weight sets are used in an 8-bit bit-serial IMC implementation, the accumulator 154 can internally accumulate a K×8 (17-bit) value that is appropriately shifted in response to the repetition of the bit-serial sequence. When P = 8, K = 4, and N = 512, the final bit width of the output accumulator is 27 bits. The 27 bits are calculated as follows. In each iteration of the bit-serial process, 512 8-bit multiplication values are accumulated, which requires 17 bits to represent. The 17-bit value is shifted and accumulated in 8 cycles, in which case the combined value requires 25 bits to be fully represented. The accumulator can repeat this cycle for K = 4 different weight sets and finally requires 27 bits to represent the final result.

[0097] In general, the parameters N, M, P, and K can take various possible values. The values shown above are only examples. In another form that depends on the parallel implementation shown in FIG. 4B, for example, no parallel-to-serial conversion is necessary. Rather, the vector components are supplied in parallel to each of the M cells of the corresponding input line 152 via the input device 151a. The complete input is executed in a single cycle (by weight multiplication). By using the accumulator 154a to accumulate S90 the partial products, the current N vector and the associated matrix block are provided. In that case, no intermediate accumulation for the MAC operation is necessary.

[0098] Regardless of whether it is based on a bit-serial implementation form or a parallel implementation form, the hardware devices 10, 10a can incorporate a configuration and control logic unit 12 that is connected to each of the input devices 151, 151a and the selection circuit 159, as shown in FIG. 2B. Further, the pre-data processing device 13 may be connected to the configuration and control logic unit 12 (so as to appropriately partition the problem into N vectors and block matrices and issue operations). For the sake of completeness, the post-data processing device 14 is connected to the output of the output line 153, for example, to the output of the accumulators 154, 154a, so as to appropriately rearrange the output data as necessary, and, as seen in FIG. 1, store those data in a local memory or a memory in the immediate vicinity (e.g., memory 11), or may be instructed to return those data to the external entity 2.

[0099] In that regard, another aspect of the present invention relates to the computing system 1. The system 1 may particularly include one or more in-memory processing hardware devices 10, 10a as described above. The computing system may have a client / server configuration, for example, as assumed in FIG. 1. That is, the user 4 can interact with the server 2 (via the personal device 3) for the purpose of executing calculations. In the latter case, in particular, substantial matrix-matrix multiplication or matrix-vector multiplication may need to be performed, in which case the server 2 may decide to offload such calculations to the hardware devices 10, 10a that function as accelerators.

[0100] Server 2 can be regarded as integrating an external memory device with an external general-purpose processing device 2, where the latter is connected to the former so as to read and write data to the storage device 2 during operation. Further, each of the in-memory processing hardware devices 10, 10a is configured to communicate with Server 2, so that, as needed, it can read data from and write data back to the storage device 2 to handle the calculation tasks sent by Server 2. It should be noted that the general-purpose processing device may, in some cases, be configured to map the first calculation task (the problem to be solved) to an N-vector and the corresponding block matrix.

[0101] It should be noted that the external memory device and the general-purpose processing device form part of the same general-purpose computer 2 (i.e., the server) in the example of FIG. 1. However, in principle, the external processing device and the storage device may, in some cases, be provided on physically different machines. In yet another form, System 1 may also be configured as a cloud computing system and may, in some cases, use containerization technology. That is, the present invention can, in particular, be embodied as a cloud computing system and can be utilized in some form as part of a cloud-based service. System 1 may further include a composable disaggregated infrastructure that can, in particular, include the hardware devices 10, 10a together with other hardware acceleration devices, such as ASICs and FPGAs. In all cases, data exchange can be optimized to speed up calculations as described above.

[0102] The above embodiments have been briefly described with reference to the accompanying drawings and can correspond to several alternative forms. Several combinations of the above features are possible. Examples will be given in the following section.

[0103] 2. Specific Embodiments Particularly preferred embodiments rely on architectures such as those shown in FIGS. 2A, 3A, and 4A having SRAM memory elements. Due to the high proportion of tightly coupled arithmetic units rather than memory elements typically in the region, the area of the IMC chip between the arithmetic units (multipliers and adder trees) and the memory elements is balanced, unlike the previous crossbar arrays which can be arranged more densely. Apart from the better balance between the memory and the arithmetic region, the proposed solution (by multiple weight sets) improves the flexibility of the IMC device 15, enables native mapping of larger matrix-vector multiplications, and allows prefetching of weight sets. Each weight set represents a separate block matrix. Nevertheless, the weight sets are connected to the same arithmetic unit as can be seen in FIGS. 3A and 3B.

[0104] Moreover, the proposed architecture and functionality also result in (see FIGS. 5A and 5B, due to less interaction with external memory) improved efficiency and (see FIG. 6, due to the possibility of prefetching weights for inactive weight sets) faster execution time. This architecture particularly relies on an accumulator circuit which can accumulate partial products resulting from both bit-serial cycles and multiple weight sets. Thus, there is no need to write intermediate results to external memory. This significantly reduces the read / write volume, i.e., from 2K - 1 (where only one weight set can be stored locally in this case) to 1 (where K weight sets are stored locally in this case).

[0105] As shown in FIG. 7, a typical operation flow is as follows, assuming a bit-serial implementation form (see also FIGS. 2A to 4A). A device 10 including a crossbar array structure 15 is provided in step S10. For example, the device 10 is set to communicate with the server 2 for data. In step S20, the server 2 receives a request (from a user 4 who can be a computerized client) to perform a matrix-matrix multiplication. In step S30, a calculation strategy is determined by either the processing means of the server 2 or the embedded processing device 13. As a result, N vectors are associated with each block matrix. In step S40, a matrix-matrix multiplication cycle is started. In step S50, the next iteration is started, whereby a given input matrix of T columns is selected. As a result of the calculation strategy, K sets of N×M weights are loaded in step S55. The K sets of weights correspond to the K block matrices that are continuously used during the calculation cycle. The memory element is programmed accordingly. The next input vector of the current input matrix (of the K×N components) is selected in step S58 and padded if necessary. The next sub-vector of this input vector (i.e., an N-vector of N components) is selected in step S60 together with the corresponding block matrix. In step S70, the corresponding N×M weights are locally activated as active weights.

[0106] In step S80, a block matrix calculation is started. In step S82, a loop is started to serially supply the next bit (of the vector components of the current N-vector) to the N input lines of the array 15. In step S84, a bit-serial MAC operation is performed. In step S86, partial results are accumulated. The process is repeated (S88: no) until all P bit-serial cycles are completed (S88: yes). When all P bit-serial cycles are completed, the processing of the current N-vector is completed.

[0107] The intermediate matrix-vector product obtained with this N vector is accumulated with the previous matrix-vector product as needed, S90. That is, all intermediate matrix-vector products are accumulated except for the very first matrix-vector product. The intermediate matrix-vector product calculation cycle (S60~S100) is repeated until all sub-vectors are processed (S100: yes) for multiplication by the relevant block matrix. The loop for the input vector (S50~S110) is repeated for all input vectors. When all vectors are processed (S110: yes), the final result for the current input matrix is returned to the calling entities 2, 13, S120. Alternatively, this result may be locally stored until all input matrices (S50~S120) are processed. Only then are the results for all input matrices returned.

[0108] 3. Final view The computerized devices 10, 10a and the system 1 can be suitably designed to implement the embodiments of the present invention as described herein. In that regard, it is understandable that the methods described herein are not essentially interactive, i.e., are automated. The automated part of such methods can be implemented only in hardware or as a combination of hardware and software. In an exemplary embodiment, the automated part of the methods described herein is implemented in software as a service or executable program (e.g., an application), the latter being executed by a suitable digital processing device. However, all embodiments described herein optionally include calculations performed by a crossbar array structure adapted to store multiple sets of weights using the look-ahead and accumulation capabilities of the devices 10, 10a.

[0109] Still, the methods described herein may typically include executable programs, scripts, or more generally, any form of executable instructions that instruct the core calculations in devices 10, 10a to be performed. The required computer-readable program instructions may be downloaded to the processing elements from, for example, a computer-readable storage medium via a network, such as the Internet and / or a wireless network.

[0110] Aspects of the invention are described herein with reference to, in particular, flowcharts and block diagrams. It will be understood that each block or combination of blocks in the flowcharts and block diagrams can be implemented by computer-readable program instructions. The flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of devices 10, 10a, as well as systems 1 including such devices, methods of operating them, and computer program products, according to various embodiments of the invention.

[0111] The present invention has been described with reference to a limited number of embodiments, alternative forms, and the accompanying drawings. However, as will be understood by those skilled in the art, various changes can be made and equivalents can be substituted without departing from the scope of the present invention. In particular, the features (in device or method form) recited in a given embodiment may be combined with or substituted for other features in another embodiment, alternative form, or the drawings without departing from the scope of the present invention. Therefore, various combinations of the features described with respect to any of the above embodiments or alternative forms are conceivable, which remain within the scope of the appended claims. Furthermore, many minor modifications can be made to adapt a particular situation or material to the teachings of the present invention without departing from the scope of the present invention. Accordingly, the present invention is not limited to the specific embodiments disclosed, and it is intended that the present invention include all embodiments within the scope of the appended claims. Furthermore, many other alternative forms can be conceived other than those explicitly touched upon above. For example, other types of memory elements, selection circuits, and program circuits can be conceived.

Claims

1. A method for in-memory processing, providing a crossbar array structure (15) including N input lines (152) and M output lines (153) interconnected at cross-points defining N×M cells (155), where N≧2 and M≧2, and each cell (155) includes a respective memory system (157) designed to store K weights where K≧2, whereby the crossbar array structure (15) includes an N×M memory system storing K sets of N×M weights (S10); for each of the memory systems (157), selecting one weight from the K weights and setting the selected weight as an active weight, thereby enabling N×M active weights for the N×M cells (155) (S70); applying a signal encoding an N-component vector to the N input lines (152) of the crossbar array structure (15) (S82), and causing the crossbar array structure (15) to perform a multiply-accumulate operation or MAC operation based on the vector and the N×M active weights (S84); reading out the output signals obtained at the outputs of the M output lines (153) (S86, S90) to obtain corresponding values comprising a method for in-memory processing.

2. while performing a MAC operation according to the N×M weights currently active as active weights, prefetching q sets of N×M weights to be used next instead of q sets of N×M weights that were previously active, where 1≦q≦K−1, and further storing the prefetched weights in the N×M memory system (157), the method according to claim 1.

3. the N×M active weights are enabled (S70) by selecting, for each memory system (157) of at least a subset of the N×M memory systems (157), the k-th weight of the K weights of the respective memory system (157), where 1≦k≦K, and setting each weight so selected as the currently active weight, the method according to claim 1 or 2.

4. The method includes performing several matrix-vector calculation cycles (S58 - S110), each of the matrix-vector calculation cycles comprising: For each of the memory systems (157), selecting one weight from its K weights and setting the selected weight as an active weight, thereby enabling a new N×M set of active weights for the N×M cells (155) (S70); Applying a signal encoding an N-component vector to the N input lines (152) of the crossbar array structure (15) (S82), and causing the crossbar array structure (15) to perform a MAC operation based on the vector and the new N×M set of active weights (S84); Reading out the output signals obtained at the outputs of the M output lines (153) (S86, S90) to obtain corresponding values. The method according to any one of claims 1 to 3. **Claim 5** Each of the matrix-vector calculation cycles further includes continuously performing accumulation by accumulating the results of the partial products corresponding to the read output signals (S90). The method according to claim 4. **Claim 6** Before completing K of the several matrix-vector calculation cycles, prefetching q sets of N×M weights (S115), where 1 ≤ q ≤ K - 1, and storing the prefetched weights in the N×M memory system (157) instead of q previously enabled sets of N×M weights as active weights. The method according to claim 5. **Claim 7** After completing the several matrix-vector calculation cycles (S110: Yes), further returning the result obtained based on the successive accumulations to an external storage device (2) (S120). The method according to claim 5 or 6. **Claim 8** The method includes performing K×T matrix-vector calculation cycles (S50 - S110), where T corresponds to the number of input vectors, each of the input vectors being decomposed into K sub-vectors of N components and associated with K respective block matrices, the K block matrices each corresponding to a set of K N×M weights, whereby the K×T matrix-vector calculation cycles are: Load K sets of N×M weights corresponding to each of the K block matrices (S55), and accordingly program the memory system (157) to store the K sets of N×M weights, and for each of the K sub-vectors of each of the T input vectors (S60), enable the N×M active weights corresponding to the associated ones of each of the K block matrices as the currently active weights (S70), apply a signal encoding the vector corresponding to each of the sub-vectors to the N input lines (152) (S82) to cause the crossbar array structure (15) to perform a MAC operation based on each of the sub-vectors and the currently active weights (S84), and The method according to claim 7, which is executed by reading out the output signals obtained at the outputs of the M output lines (153) to obtain corresponding partial values.

9. Reading out the output signal, if any, includes accumulating the partial values obtained for each of the sub-vectors together with the partial values previously obtained for the previous ones of the K sub-vectors (S90) to obtain an updated result, The method according to claim 8, further including returning a result obtained based on the last obtained updated result (S120).

10. In an external processing device, program the N×M memory system (157) according to K sets of N×M weights (S55), and before encoding a certain number of sub-vectors into an input signal and applying such an input signal to the N input lines (152) for the purpose of executing the several matrix-vector calculation cycles, The method according to any one of claims 4 to 9, further including mapping a given problem to the sub-vectors and the K sets of N×M weights (S30).

11. The N×M memory system (157) is a digital memory system, each of the N×M cells (155) further includes an arithmetic unit (156) connected to the corresponding one of the N×M memory systems (157). The MAC operation is preferably performed bit - serially in P cycles where P≧2 (S84), P corresponding to the bit - width of each of the N components of the vectors used in the input, whereby partial - product values are accumulated upon completion of each of the P cycles (S86), a method according to any one of claims 1 to 10.

12. A computer - readable storage medium in which program instructions are embodied, the program instructions being executable by processing means of an in - memory processing hardware device (10) to cause the in - memory processing hardware device (10) to execute the steps of any one of claims 1 to 11.

13. An in - memory processing hardware device (10), A cross - bar array structure (15) including N input lines (152) and M output lines (153) interconnected at cross - points defining N×M cells (155), where N≧2 and M≧2, each cell (155) including a respective memory system (157) designed to store K weights where K≧2, whereby the cross - bar array structure includes an N×M memory system (157) adapted to store K sets of N×M weights for performing a multiply - accumulate operation or a MAC operation. A selection circuit (159) connected to the N×M memory system (157), the selection circuit (159) being configured to select a weight from each of the K weights of the memory system and set the selected weight as an active weight to enable N×M active weights for the N×M cells (155). An input device (151) configured to apply a signal encoding an N - component vector to the N input lines (152) of the cross - bar array structure (15) to cause the cross - bar array structure (15) to perform a MAC operation based on the vector and the N×M active weights enabled by the selection circuit (159). A reading device (154) configured to read an output signal obtained at the output of the M output lines (153). An in - memory processing hardware device (10) comprising the above.

14. Each of the N×M memory systems (157) is designed such that these K weights can be programmed independently. The in-memory processing hardware device (10) further includes a program circuit (158) connected to each of the memory systems (157), and the program circuit (158) is configured to program the K weights of the N×M memory system. The in-memory processing hardware device (10) according to claim 13. **Claim 15** The program circuit (158) prefetches q weights of the N×M weight set that are not currently set as active weights, and accordingly programs the N×M memory system (157) such that the N×M memory system (157) is configured to store the prefetched weights instead of the q weights of the N×M weight set, where 1≤q≤K−1. The in-memory processing hardware device (10) according to claim 14. **Claim 16** Each of the N×M memory systems (157) includes K memory elements adapted to store the corresponding weights of the K weights respectively, The selection circuit (159) includes an N×M multiplexer each connected to each of the K memory elements corresponding to the corresponding one of the N×M memory systems (157), and a selection control line connected to each of the multiplexers such that any one of the K weights of each of the memory systems (157) can be selected and set as the active weight during operation. The in-memory processing hardware device (10) according to any one of claims 13 to 15. **Claim 17** The selection circuit (159) is further configured to select a subset of the n×m weights from among the K weights of the N×M weight set by selecting the k-th weight of the K weights of each of the memory systems (157) of a subset of the n×m memory systems (157) of the N×M memory system, where 2≤n≤N, 2≤m≤M, and 1≤k≤K. The in-memory processing hardware device (10) according to any one of claims 13 to 16. **Claim 18** The in-memory processing hardware device (10) further includes a sequencer circuit and an accumulator circuit (154). The sequencer circuit is connected to the input device (151) and the selection circuit (159) so as to continuously execute several cycles of matrix-vector calculation based on one or more vector sets, and schedules the operations of the input device (151) and the selection circuit (159). During the operation, each of the cycles of the matrix-vector calculation includes one or more cycles of MAC operations. Different sets of N×M weights are selected from the K sets of N×M weights and are set as the N×M active weights in each of the cycles of the matrix-vector calculation. The accumulator circuit (154) is configured to accumulate the partial product values obtained when each MAC operation cycle is completed. The in-memory processing hardware device (10) according to any one of claims 13 to 17.

19. The accumulator circuit (154) is arranged at the output of the output line (153). The in-memory processing hardware device (10) according to claim 18.

20. Each of the N×M memory systems (157) includes K memory elements adapted to store the corresponding weights of the K weights respectively. The in-memory processing hardware device (10) according to any one of claims 13 to 18.

21. Each of the K memory elements of each of the N×M memory systems (157) is a digital memory element. Each of the N×M cells (155) further includes an arithmetic device (156) connected to each of the K memory elements of the corresponding one of the N×M memory systems (157) via a corresponding part of the selection circuit (159). The in-memory processing hardware device (10) according to claim 20.

22. Each of the K memory elements of each of the N×M memory systems (157) is designed to store P-bit weights. The input device (151) is configured to supply an N-component vector bit-serially to the input line (152) in P cycles by applying the signal. Each of the N components corresponds to a P-bit input word, where P≧2. The N×M cell (155) is configured to execute MAC operations bit-serially in the P cycles. The in-memory processing hardware device (10) further includes an accumulator circuit (154) configured to accumulate values corresponding to the partial bit-serial product values obtained in each of the P cycles. The memory-internal processing hardware device (10) according to claim 21, wherein the selection circuit (159) is configured to maintain the same N×M set of weights as the active weights between each of the P cycles.

23. The in-memory processing hardware device (10) further includes a configuration and control logic unit (12) connected to each of the input device (151) and the selection circuit (159), a pre-data processing device (13) connected to the configuration and control logic unit (12), and a post-data processing device (14) connected to the output of the output line (153). The in-memory processing hardware device (10) according to any one of claims 13 to 22.

24. A computing system (1) comprising one or more in-memory processing hardware devices (10) respectively described in any one of claims 13 to 23.

25. The computing system further includes a storage device (2) and a general-purpose processing device (2) connected to the storage device to read and write data to and from the storage device (2). Each of the in-memory processing hardware devices (10) is configured to read and write data to and from the storage device (2). The general-purpose processing device (2) is configured to map a given computing task to vectors and weights with respect to the memory system (157) of the one or more in-memory processing hardware devices (10). The computing system according to claim 24.

Citation Information

Patent Citations

  • Neuromorphic device using 3D crossbar memory

    KR1020200024419A