Elements for in-memory computing

By introducing cluster cycle management circuits and column multiplexers into the memory array system, the data storage and computing processes are optimized, the problem of poor memory utilization is solved, and the efficiency and accuracy of neural network computing are improved.

CN112070219BActive Publication Date: 2025-09-26STMICROELECTRONICS SRL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010518406.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-02
Filing Date
2020-06-09
Publication Date
2025-09-26
Estimated Expiration
2040-06-09

AI Technical Summary

Technical Problem

Existing memories have poor memory utilization when processing neural network calculations, resulting in poor aspect ratio utilization and possible loss of accuracy, especially when the dataset changes and cannot be fully optimized.

Method used

A memory array system is used to optimize data storage and calculation processes through cluster cycle management circuit devices and column multiplexers. Configurable multiplexers and sensing circuits are used to achieve efficient data reading and writing and calculation value combinations, reducing the number of multiplexer cycles to improve memory utilization.

Benefits of technology

It improves the aspect ratio utilization of memory, reduces precision loss, and optimizes the efficiency and performance of neural network computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112070219B_ABST
    Figure CN112070219B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to elements for in-memory computation. A memory array is arranged in a plurality of columns and a plurality of rows. Computational circuits each deduce a computational value based on a cell value in a corresponding column. A column multiplexer cycles through a plurality of data lines, each corresponding to a computational circuit. A cluster cycle management circuit device determines the number of multiplexer cycles based on the number of columns storing data for the computation cluster. As the column multiplexer cycles through the data lines, a sensing circuit obtains a computational value from the computational circuit via the column multiplexer. The sensing circuit combines the computational values ​​obtained within a determined number of multiplexer cycles. A first clock can activate the multiplexer to cycle through its data lines for a determined number of multiplexer cycles, and a second clock can activate each individual cycle. Multiplexers or additional circuit devices can be used to modify the order in which data is written to the columns.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to memory arrays, such as memory arrays used in learning / inference engines, such as artificial neural networks (ANNs). Background Art

[0002] It is well known that various computer vision, speech recognition, and signal processing applications benefit from the use of learning / inference engines. As discussed in this disclosure, learning / inference engines can be classified into the technical categories of machine learning, artificial intelligence, neural networks, probabilistic inference engines, accelerators, etc. Such machines are arranged to quickly execute hundreds, thousands, or even millions of concurrent operations. Conventional learning / inference engines can provide hundreds of trillions of floating-point operations (i.e., one trillion (10 12 ) floating-point operations).

[0003] Known computer vision, speech recognition, and signal processing applications benefit from the use of learning / inference engines, such as deep convolutional neural networks (DCNNs). DCNNs are computer-based tools that process large amounts of data and adaptively "learn" by combining proximal related features within the data, making broad predictions about the data, and refining the predictions based on reliable conclusions and new combinations. DCNNs are arranged in multiple "layers," and make different types of predictions at each layer.

[0004] For example, if multiple two-dimensional facial images are provided as input to a DCNN, the DCNN will learn various features of the face (e.g., edges, curves, angles, points, color contrast, bright spots, dark spots, etc.). These one or more features are learned at one or more first layers of the DCNN. Then, in one or more second layers, the DCNN will learn various identifiable features of the face (e.g., eyes, eyebrows, forehead, hair, nose, mouth, cheeks, etc.); each feature can be distinguished from all other features. That is, the DCNN learns to recognize and distinguish eyes from eyebrows or any other facial features. In one or more third or subsequent layers, the DCNN learns features of the entire face and higher-order features (e.g., race, gender, age, emotional state, etc.). In some cases, the DCNN is even taught to recognize a specific identity of a person. For example, a random image can be identified as a face, and a face can be identified as Person_A, Person_B, or some other identity.

[0005] In other examples, a DCNN can be fed multiple pictures of animals and taught to identify lions, tigers, and bears; a DCNN can be fed multiple pictures of cars and taught to identify and distinguish between different types of vehicles; and many other DCNNs can be formed. DCNNs can be used to learn word patterns in sentences, identify music, analyze personal shopping patterns, play video games, create traffic routes, and many other learning-based tasks. Summary of the Invention

[0006] The system can be summarized as including: a memory array having a first plurality of cells arranged as a plurality of cell rows intersecting a plurality of cell columns; a plurality of first computation circuits, wherein each first computation circuit is operable to deduce a computation value based on cell values ​​in a corresponding cell column of the first plurality of cells; a first column multiplexer, wherein the first column multiplexer is operable to cycle through a plurality of data lines, each of the plurality of data lines corresponding to a first computation circuit of the plurality of first computation circuits; a first sensing circuit, wherein the first sensing circuit is operable to obtain computation values ​​from the plurality of first computation circuits via the first column multiplexer as the first column multiplexer cycles through the plurality of data lines and to combine the obtained computation values ​​within a determined number of multiplexer cycles; and a cluster cycle management circuit device, wherein the cluster cycle management circuit device is operable to determine a determined number of multiplexer cycles based on the number of columns storing data of the computation cluster. The determined number of multiplexer cycles can be less than the number of physical data lines of the first column multiplexer.

[0007] The calculated value may be a partial sum of cell values ​​in the corresponding column, and in operation, the first sensing circuit calculates a first sum based on the obtained partial sum via a first set of the plurality of data lines during a first set of cycles of the first column multiplexer, and calculates a second sum based on the obtained partial sum via a second set of the plurality of data lines during a second set of cycles of the first column multiplexer, wherein the first cycle set and the second cycle set have a certain number of multiplexer cycles. In operation, the first column multiplexer may cycle through a second plurality of data lines, each data line corresponding to a respective second calculation circuit from a plurality of second calculation circuits, wherein each second calculation circuit calculates a partial sum based on cell values ​​in a corresponding column of cells in the second plurality of cells. The second plurality of data lines for the first column multiplexer may be provided by the second column multiplexer.

[0008] The system may include: a plurality of second calculation circuits, wherein each second calculation circuit is operable to deduce a calculation value based on a cell value in a corresponding column of cells in a second plurality of cells of a memory array; a second column multiplexer, wherein the second column multiplexer is operable to cycle through a plurality of data lines, each data line corresponding to a second calculation circuit in the plurality of second calculation circuits; and a second sensing circuit, wherein the second sensing circuit is operable to obtain calculation values ​​from the plurality of second calculation circuits via the second column multiplexer as the second column multiplexer cycles through the plurality of data lines, and to combine the calculation values ​​obtained within a determined number of multiplexer cycles.

[0009] In operation, the cluster cycle management circuitry can generate a plurality of control signals in response to a clock signal and can provide the plurality of control signals to the first sensing circuit and the first column multiplexer to cycle through the plurality of data lines for a determined number of multiplexer cycles, thereby causing the first sensing circuit to obtain a calculated value from the corresponding first calculation circuit. In operation, the first column multiplexer can modify the address of each of the plurality of data lines of the first column multiplexer to write a plurality of consecutive data items to the first plurality of cells. The system can also include data line selection circuitry that, in operation, selects a different order of cycling through the plurality of data lines of the first column multiplexer to write the plurality of consecutive data items to the first plurality of cells.

[0010] The method can be summarized as including: storing data in a plurality of memory cells, the memory cells being arranged as a plurality of cell rows intersecting a plurality of cell columns; inferring a plurality of calculation values ​​based on cell values ​​from the plurality of cell columns, wherein each corresponding calculation value is calculated based on the cell values ​​from the corresponding cell column; determining the number of columns in the plurality of cell columns that store data for a data calculation cluster; selecting a number of multiplexer cycles based on the determined number of columns; generating a result for the data calculation cluster by using a column multiplexer cycle through a first subset of the plurality of calculation values ​​for a selected number of multiplexer cycles, and by using a sensing engine to combine corresponding calculation values ​​from the first subset of calculation values; and outputting the result of the data calculation cluster.

[0011] The method may include: generating a second result for the data computation cluster by cycling through a second subset of the plurality of computation values ​​for a selected number of multiplexer cycles using a second column multiplexer and combining corresponding computation values ​​from the second subset of computation values ​​using a second sensing engine; and outputting the second result for the data computation cluster. The method may include: generating a second result for the data computation cluster by cycling through a second subset of the plurality of computation values ​​for a selected number of multiplexer cycles using a column multiplexer and combining corresponding computation values ​​from the second subset of computation values ​​using a sensing engine; and outputting the second result for the data computation cluster.

[0012] The method may include modifying the number of data lines utilized by a column multiplexer based on a selected number of multiplexer cycles for a data computation cluster. The method may include enabling the column multiplexer in response to a non-memory clock signal to cycle through a first subset of computed values ​​for a selected number of multiplexer cycles; and enabling each cycle of each data line of the column multiplexer for the selected number of multiplexer cycles in response to a memory clock signal to obtain the first subset of computed values. The method may include modifying an address of each of a plurality of data lines of the column multiplexer to write a consecutive plurality of data to a first plurality of cells. The method may include selecting a different order in which the column multiplexer cycles through a plurality of columns of cells to write the consecutive plurality of data to the plurality of cells.

[0013] The computing device can be summarized as including: a device for storing data in multiple cells, the multiple cells being arranged as multiple cell rows intersecting multiple cell columns; a device for calculating corresponding calculation values ​​based on cell values ​​from each corresponding column in the multiple cell columns; a device for determining a calculation cluster loop size based on the number of columns in the multiple cell columns that store data for the data calculation cluster; a device for looping through the corresponding calculation values ​​for the determined calculation cluster loop size; and a device for combining the corresponding calculation values ​​to generate a result of the data calculation cluster for the determined calculation cluster loop size.

[0014] The computing device may include: means for initiating a loop through corresponding computation values ​​for a determined computation cluster loop size; and means for initiating each loop for each corresponding computation value of the determined computation cluster loop size. The computing device may include: means for modifying an order in which data is stored in a plurality of cell columns to write a plurality of data consecutively to the plurality of cells.

[0015] A non-transitory computer-readable medium having content that causes a cluster cycle management circuit device to perform actions that can be summarized as including: storing data in a plurality of memory cells arranged as a plurality of cell rows intersecting a plurality of cell columns; determining a number of columns in the plurality of cell columns that store data for a computational operation; selecting a number of multiplexer cycles based on the determined number of columns; generating a result of the computational operation by employing a column multiplexer cycle through a first subset of the plurality of cell columns for a selected number of multiplexer cycles and combining values ​​from the first subset of the cell columns by employing a sense engine; and outputting the result of the computational operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Non-limiting and non-exhaustive embodiments are described with reference to the following figures, in which like reference numerals refer to like parts throughout the various views unless otherwise indicated. The sizes and relative positions of the elements in the drawings are not necessarily drawn to scale. For example, the shapes of the various elements are selected, enlarged, and positioned to improve the readability of the drawings. For ease of identification in the drawings, the specific shapes of the drawn elements have been selected. In addition, for ease of illustration, some elements known to those skilled in the art are not shown in the drawings. hereinafter, one or more embodiments are described with reference to the accompanying drawings, in which:

[0017] Figure 1 is a functional block diagram of one embodiment of an electronic device or system having a processing core and memory according to one embodiment;

[0018] Figure 2A-2C illustrates a use case context diagram of a memory array with configurable multiplexers and sensing circuits;

[0019] Figures 3A-3D A use case context diagram illustrating the use of a self-shifting multiplexer to continuously write multiple data into a memory array;

[0020] Figures 4A-4D A use case context diagram illustrating a method of continuously writing multiple data into a memory array using pre-decode shifting is illustrated;

[0021] Figure 5 The diagram generally illustrates the use of Figure 2A-2C A logic flow diagram of one embodiment of a process for reading data from a memory array using a configurable multiplexer and sensing circuit is shown;

[0022] Figure 6 A logic flow diagram generally illustrates one embodiment of a process employing a multiplexer that modifies a data line address to sequentially write multiple data lines such as a Figures 3A-3D The memory array shown;

[0023] Figure 7 The figure generally shows the use of pre-decoding shift to continuously write multiple data Figures 4A-4D A logic flow diagram of one embodiment of a memory array is shown; and

[0024] Figure 8 Illustrated is a logic flow diagram generally showing one embodiment of a process for initiating multiplexer reads and writes to a memory array using a first clock and initiating each cycle of the multiplexer reads and writes using a second clock. DETAILED DESCRIPTION

[0025] The following description and the accompanying drawings set forth certain specific details to provide a thorough understanding of the various disclosed embodiments. However, those skilled in the relevant art will recognize that the disclosed embodiments can be practiced in various combinations without one or more of these specific details, or with other methods, components, devices, materials, etc. In other cases, in order to avoid unnecessary confusion in the description of the embodiments, well-known structures or components associated with the environment of the present disclosure (including but not limited to interfaces, power supplies, physical component layouts, etc.) are not shown or described. Additionally, the various embodiments may be methods, systems, or devices.

[0026] Throughout the specification, claims, and drawings, unless the context clearly indicates otherwise, the following terms have the meanings clearly associated herein. The term "herein" refers to the specification, claims, and drawings associated with this application. Unless the context clearly indicates otherwise, the phrases "in one embodiment," "in another embodiment," "in various embodiments," "in some embodiments," "in other embodiments," and their other variations refer to one or more features, structures, functions, limitations, or characteristics of the present disclosure, and are not limited to the same or different embodiments. As used herein, the term "or" is an inclusive "or" operator and is equivalent to the phrase "A or B, or both" or "A or B or C, or any combination thereof," and lists with additional elements are treated similarly. Unless the context clearly indicates otherwise, the term "based on" is not exclusive and allows for reference to additional features, functions, aspects, or limitations that are not described. Additionally, throughout the specification, the meanings of "a," "an," and "the" include the singular and the plural.

[0027] The computations performed by DCNNs or other neural networks often involve repeated calculations on large amounts of data. For example, many learning / inference engines compare known information, or kernels, to unknown data, or feature vectors (e.g., comparing a known grouping of pixels to a portion of an image). A common type of comparison is the dot product between the kernel and the feature vector. However, kernel size, feature size, and depth often vary between different layers of the neural network. In some cases, dedicated compute units can be used to perform these operations on varying datasets. However, due to the variability in the dataset, memory utilization may not be fully optimized. For example, small-swing computations can be enabled by selecting all elements connected to a single compute line, while large compute clusters may utilize a large number of elements across multiple compute lines, which can result in poor memory aspect ratio utilization. Similarly, due to limited voltage margins, such variations can lead to a loss of accuracy.

[0028] Figure 11 is a functional block diagram of one embodiment of an electronic device or system 100 of the type to which the embodiments to be described may be applied. System 100 includes one or more processing cores or circuits 102. Processing core 102 may include, for example, one or more processors, state machines, microprocessors, programmable logic circuits, discrete circuit devices, logic gates, registers, and various combinations thereof. The processing cores may control the overall operation of system 100, the execution of applications by system 100, and the like.

[0029] The system 100 includes one or more memories (e.g., one or more volatile and / or non-volatile memories) that may store, for example, all or part of instructions and data related to the control of the system 100, applications and operations executed by the system 100, etc. As shown, the system 100 includes one or more cache memories 104, one or more primary memories 106, and one or more secondary memories 108. One or more of the memories 104, 106, 108 include memory arrays that are shared by one or more processes executed by the system 100 during operation (see, e.g., Figure 2A-2C The memory array 202, Figures 3A-3D Memory array 302 or Figures 4A-4D 402 in the memory array).

[0030] The system 100 may include one or more sensors 120 (e.g., image sensors, audio sensors, accelerometers, pressure sensors, temperature sensors, etc.), one or more interfaces 130 (e.g., wireless communication interfaces, wired communication interfaces, etc.), one or more BIST circuits 140 and other circuits 150 (which may include antennas, power supplies, etc.), and a primary bus system 170. The primary bus system 170 may include one or more data, address, power, and / or control buses that couple to various components of the system 100. The system 100 may also include additional bus systems, such as a bus system 162 that communicatively couples the cache memory 104 and the processing core 102, a bus system 164 that communicatively couples the cache memory 104 and the primary memory 106, a bus system 166 that communicatively couples the primary memory 106 and the processing core 102, and a bus system 168 that communicatively couples the primary memory 106 and the secondary memory 108.

[0031] System 100 also includes cluster cycle management circuitry 160 that, in operation, employs one or more memory management routines to utilize configurable multiplexers and sense circuits to read data from memory 104, 106, or 108 (see, for example, Figure 2A-2C ), or dynamically write data to memory in a shared memory array (see, for example, Figures 3A-3D and Figures 4A-4D). Memory management circuitry 160 may perform the routines and functions described herein (including the respective Figure 5-Figure 8 Cluster cycle management circuitry 160 may be one or more processors, state machines, microprocessors, programmable logic circuits, discrete circuitry, logic gates, registers, and various combinations thereof.

[0032] Primary memory 106 is typically working memory for system 100 (e.g., one or more memories on which processing core 102 operates) and may typically be a volatile memory of limited size that stores code and data related to processes executed by system 100. For convenience, references herein to data stored in memory may also refer to code stored in memory. Secondary memory 108 may typically be non-volatile memory that stores instructions and data that may be retrieved and stored in primary memory 106 when needed by system 100. Cache memory 104 may be a relatively faster memory than secondary memory 108 and may typically have a limited size that may be larger than the size of primary memory 106.

[0033] Cache memory 104 temporarily stores code and data for later use by system 100. Instead of retrieving required code or data from secondary memory 108 for storage in primary memory 106, system 100 may first check cache memory 104 to see if the data or code is already stored in cache memory 104. Cache memory 104 can significantly improve system (e.g., system 100) performance by reducing the time and other resources required to retrieve data and code for use by system 100. When code and data for use by system 100 are retrieved (e.g., from secondary memory 108), or when data or code is written (e.g., to primary memory 106 or secondary memory 108), a copy of the data or code may be stored in cache memory 104 for later use by system 100. Various cache management routines may be used to control the data stored in one or more cache memories 104.

[0034] Figure 2A-2C A use case context diagram illustrating a memory array with a configurable multiplexer and a sensing circuit for reading data from the memory array is illustrated. Figure 2A The system 200A in FIG. 2 includes a memory array 202 , a plurality of computational circuits (shown as part of circuit 206 ), one or more column multiplexers 208 a - 208 b , and one or more sensing circuits 210 a - 210 b .

[0035] As shown, the memory 202 includes a plurality of cells 230 arranged in a column-row arrangement, in which a plurality of cell rows intersect a plurality of cell columns. Each cell can be addressed via a specific column and a specific row. The number of cells 230 illustrated in the memory array 202 is for illustrative purposes only, and a system employing the embodiments described herein may include more or fewer cells in more or fewer columns and more or fewer rows. The details of the functions and components used to access a particular memory cell are known to those skilled in the art and are not described herein for the sake of brevity.

[0036] Each column 204a-204h of memory array 202 is in electrical communication with a corresponding partial sum circuit 206. Each partial sum circuit 206 includes circuitry configured to, during operation, deduce or determine the sum of the cell values ​​from each cell 230 in the corresponding column. For example, partial sum PS1 is the sum of the cells 230 in column 204a, partial sum PS2 is the sum of the cells 230 in column 204b, and so on. In one embodiment, partial sum circuit 206 may be a sample and hold circuit.

[0037] Because the embodiments described herein can be used for neural network computations, data can be processed in a computation cluster. A computation cluster is a plurality of computations performed on a data set. For example, a computation cluster can include, for example, a comparison between kernel data and feature vector data over a preselected number of pixels. In the illustrated embodiment, data is added to memory 202 such that all data for a single computation cluster is associated with the same sensing circuit 210a-210b. For example, data labeled k1[a1], k1[a2], ..., k1[n1], k1[n2], ..., k1[p1], k1[p2], ..., k1[x1], and k1[x2] is for a first computation cluster and is stored in cells 230 in columns 204a-204d. In contrast, the data labeled k2[a1], k2[a2], ..., k2[n1], k2[n2], ..., k2[p1], k2[p2], ..., k2[x1], and k2[x2] are for a separate second computational cluster and are stored in cells 230 in columns 204e-204h. In this way, the results obtained by sense circuit 210a are used for one computational cluster, while the results obtained by sense circuit 210b are used for another computational cluster.

[0038] Each multiplexer 208a-208b includes a plurality of data lines 216 and outputs 218a-218b. In the embodiment shown, each multiplexer 208a-208b includes four physical data lines 216. Therefore, the multiplexers 208a-208b are considered to have four physical multiplexer cycles. The data lines 216 are illustrated with reference to C1, C2, C3, and C4. These reference numerals are not an indication of the values ​​transmitted along the data lines 216. Instead, they refer to a specific clock cycle in which the multiplexers 208a-208b select a specific data line to transmit the value from the corresponding portion and 206 to the corresponding output 218a-218b. In other embodiments and configurations, the multiplexers can have two physical data lines, eight physical data lines, or other numbers of physical data lines. In addition, the multiplexers 208a-208b are illustrated as read-only multiplexers and are configured to read data from the memory 202. However, in some embodiments, the multiplexers 208 a - 208 b may be read / write multiplexers and configured for writing data to the memory 202 and for reading data from the memory 202 .

[0039] Multiplexers 208a-208b are configured to cycle through data lines 216 based on a selected number of cycles in a complete compute cluster cycle. A compute cluster cycle corresponds to a column associated with a particular compute cluster. For a given compute cluster cycle, the number of data lines 216 that multiplexers 208a-208b cycle through is determined based on the size of the compute cluster cycle, which is the number of columns 204 storing data in memory 202 for a given compute cluster. The cluster cycle size is independent of the number of physical data line cycles available to each multiplexer. As described above and as Figure 2A As shown, data for the first computing cluster is stored in four columns (columns 204a-204d), while data for the second computing cluster is stored in four separate columns (columns 204e-204h). In this example, the computing cluster size is four cycles.

[0040] In various embodiments, each multiplexer 208a-208b receives a control signal that indicates when to initiate a cycle through the data lines 216, when to cycle through each individual data line 216, or a combination thereof. In the illustrated embodiment, the control signals received by the multiplexers 208a-208b include, for example, various clock signals. For example, each multiplexer 208a-208b includes a first clock input 228. The first clock input 228 receives one or more clock signals from the first clock 212 to cycle through the data lines 216 during a read operation for a given compute cluster cycle. The first clock 212 can be external or internal to the memory. In some embodiments, the first clock 212 is a system clock.

[0041] In some embodiments, the signal received from the first clock 212 via the first clock input 228 initializes the multiplexers 208a-208b to cycle through the data line 216, but each individual data line cycle is triggered by a signal received from the second clock 214 via the second clock input 226. In this manner, each individual compute cluster cycle is triggered by a clock signal from the first clock 212. In this example, the clock signal received from the first clock 212 is an example of a cluster control signal that initiates a cycle of the multiplexers, and the clock signal received from the second clock 214 is an example of a multiplexer cycle control signal that initiates each individual cycle of the multiplexers.

[0042] For example, when multiplexer 208a receives clock signal CLK from first clock 212, multiplexer 208a initiates a computation cluster loop, which triggers multiplexer 208a to cycle to first data line 216 with the next clock signal from second clock 214. Thus, when receiving clock signal C1 from second clock 214, multiplexer 208a cycles to the data line corresponding to partial sum PS1, causing its partial sum value to be obtained by sensing circuit 210a. When receiving clock signal C2 from second clock 214, multiplexer 208a cycles to the next data line corresponding to partial sum PS2, causing its partial sum value to be obtained by sensing circuit 210a and combined with its previously held value. Multiplexer 208a continues to operate in a similar manner for clock signals C3 and C4 received from second clock 214, causing sensing circuit 210a to obtain values ​​from PS3 and PS4 and combine them with the values ​​from PS1 and PS2.

[0043] Similarly, when multiplexer 208b receives clock signal CLK from first clock 212, multiplexer 208b starts the computation cluster loop, which triggers multiplexer 208b to loop to first data line 216 using the next clock signal from second clock 214. When clock signal C1 is received from second clock 214, multiplexer 208b loops to the data line corresponding to partial sum PS5 so that its partial sum value is obtained by sensing circuit 210b. In this example, the clock signal CLK that starts multiplexer 208b is the same as the clock signal that starts multiplexer 208a. In this way, a single clock signal starts both multiplexers 208a-208b to begin looping through the corresponding data lines 216. Multiplexer 208b cycles through data lines corresponding to partial sums PS6, PS7, and PS8 in response to receiving clock signals C2, C3, and C4 from second clock 214 so that similar to multiplexer 208a and sense circuit 210b, partial sum values ​​are obtained and combined by sense circuit 210b.

[0044] Once the maximum number of multiplexer cycles is reached for a given compute cluster cycle (e.g., Figure 2A In the example shown, four (four), the sense circuits 210a-210b output their obtained values, and the multiplexers 208a-208b wait for the next first clock signal from the first clock 212 to start the multiplexers 208a-208b to cycle through the data lines 216 again for the next set of computing cluster cycles. As a result, the system 200A obtains two results for the two computing clusters, one result obtained by the sense circuit 210a for the columns 204a-204d, and the other result obtained by the sense circuit 210b for the columns 204e-204h. And because the multiplexers 208a-208b use the same clock signal (the first clock signal, or a combination of the first and second clock signals) to cycle through their respective data lines 216, the two results for the two computing clusters are obtained during a single computing cluster cycle.

[0045] As shown, multiplexers 208a-208b access their corresponding four data lines 216 using one clock signal from first clock 212 and four clock signals from second clock 214. However, embodiments are not limited thereto. Instead, in other embodiments, receiving a clock signal from first clock 212 may instruct multiplexers 208a-208b to cycle to the first data line 216 and then cycle to the second data line 216 upon receiving the next clock signal from second clock 214. In this manner, second clock 214 is used to cycle through each corresponding data line for a given compute cluster after the first cycle. In this illustrated example, multiplexers 208a-208b would access their corresponding four data lines 216 using one clock signal from first clock 212 and three clock signals from second clock 214.

[0046] In at least one embodiment, the second clock 214 may be a memory clock used to refresh data in the memory array 202. In other embodiments, the second clock 214 may be a clock dedicated to the multiplexers 208a-208b or some other clock. The second clock 214 may be external or internal to the memory. Typically, the second clock 214 is physically closer to the multiplexers 208a-208b and uses less power to cycle through the multiplexer data lines than the first clock 212. In various embodiments, the second clock signal may be referred to as a cycling control signal.

[0047] In other embodiments, each signal received from the first clock 212 (or some other clock) via the first clock input 228 triggers individual data line cycling. For example, when multiplexers 208a-208b receive a first clock signal from the first clock 212, multiplexer 208a cycles to the data line corresponding to partial sum PS1 so that its partial sum value is obtained by the sensing circuit 210a, and multiplexer 208b cycles to the data line corresponding to partial sum PS5 so that its partial sum value is obtained by the sensing circuit 210b. However, when multiplexers 208a-208b receive a second clock signal from the first clock 212, multiplexer 208a cycles to the data line corresponding to partial sum PS2 so that its partial sum value is obtained by the sensing circuit 210a, and multiplexer 208b cycles to the data line corresponding to partial sum PS6 so that its partial sum value is obtained by the sensing circuit 210b. Multiplexers 208a-208b continue to operate in a similar manner for additional clock signals received from first clock 212. Once the maximum number of multiplexer cycles is reached for the compute cluster cycle (e.g., at Figure 1 If the multiplexer 208a receives four clock signals (four in the example shown), the multiplexer 208a cycles back to the first data line on the next clock signal. In such an embodiment, the second clock 214 is optional and not used.

[0048] In some embodiments, the multiplexers 208a-208b may have other input / output interfaces (not shown), such as for selecting between read and write operations, selecting a particular data line 216, selecting a compute cluster loop size, and the like.

[0049] For example, as shown, multiplexers 208a-208b utilize clock signals directly to cycle through data lines 216. However, in other embodiments, each multiplexer 208a-208b may further include one or more data select lines (not shown). The data select lines are configured to receive addresses from selection circuitry to cycle through specific data lines 216. In at least one such embodiment, the selection circuitry utilizes a first clock signal from first clock 212 or a second clock signal from second clock 214, or some combination thereof, to cycle through different data lines 216, as described herein. Figures 4A-4D An example of how such a selection circuit arrangement may be utilized is shown. Figures 4A-4D The data line selection circuit arrangement shown in FIG is used to instruct a multiplexer to cycle through multiple data lines to write data to a memory, but a similar circuit arrangement can also be used to instruct a multiplexer to cycle through multiple data lines to read data from a memory.

[0050] As another example of other input / output interfaces, the multiplexers 208a-208b may include one or more compute cluster cycle size input lines (not shown). The compute cluster cycle size input lines may be used to identify how many columns 204 or data lines 216 the multiplexers 208a-208b and the sense circuits 210a-210b will utilize to obtain results from the memory 202 for a single compute cluster. For example, Figure 2A In the example shown, multiplexers 208a-208b utilize four data lines to obtain results from the partial sum for four columns of a single compute cluster. In this example, each multiplexer 208 may have one or more compute cluster cycle size input lines that instruct the multiplexers 208a-208b to operate as a compute cluster cycle (four multiplexer cycles or four data lines).

[0051] Similar to the multiplexers 208a-208b, the sensing circuits 210a-210b may also receive clock signals from the first clock 212 and the second clock 214. These clock signals are used to trigger the sensing circuits 210a-210b to obtain the partial sum values ​​selected by the multiplexers 208a-208b and output via the outputs 218a-218b for a given computation cluster cycle. When the computation cluster cycle is completed, in response to the final clock signal from the second clock 214 or in response to the next clock signal from the first clock, the sensing circuits 210a-210b move their combined stored values ​​to the output buffers 232a-232b, respectively. In some embodiments, similar to that described herein for the multiplexers 208a-208b, the sensing circuits 210a-210b may also include one or more computation cluster cycle size input lines (not shown) to indicate how many different partial sum values ​​(the number of columns associated with a given computation) will be obtained for a given computation cluster cycle.

[0052] System 200A also includes cluster loop management circuitry 160. In some embodiments, cluster loop management circuitry 160 determines the number of multiplexer loops to be used for a compute cluster loop (e.g., a compute cluster loop size) based on the number of columns of stored data for a given compute cluster. In at least one embodiment, the number of columns of stored data for a given compute cluster can be based on one or more thresholds compared to a batch size of kernel data used for neural network processing. In some embodiments, the kernel set is divided into batches that are processed sequentially based on the current neural network layer being used.

[0053] In various embodiments, the cluster cycle management circuitry 160 provides information or instructions regarding the size of the computer cluster or the clock signal utilized (e.g., whether the first clock triggers all multiplexer cycles or whether the first clock initiates a multiplexer cycle in response to the second clock signal) to the multiplexers 208a-208b and the sensing circuits 210a-210b. In some embodiments, the cluster cycle management circuitry 160 manages or generates control signals (e.g., the first clock signal and the second clock signal) and provides the control signals to the multiplexers 208a-208b and the sensing circuits 210a-210b. In some embodiments, the cluster cycle management circuitry can coordinate the output of data from the output buffers 232a-232b for further processing (e.g., input to the next layer in a neural network process).

[0054] although Figure 2A An embodiment is illustrated in which multiplexers 208a-208b cycle through all four of their physical data lines 216 such that each sense circuit 210a-210b combines four partial sums during a single compute cluster cycle, but some implementations may desire that sense circuits 210a-210b combine fewer partial sums (e.g., two partial sums). For example, in some cases and implementations, a small compute cluster may not utilize memory optimally when four columns are used due to a poor aspect ratio. To improve the aspect ratio, data for the small compute cluster is stored in two columns, rather than four as described above, and as Figure 2B As shown, the multiplexers 208a-208b and sensing circuits are also reconfigured accordingly.

[0055] Figure 2B The system 200B in FIG. 1 illustrates an embodiment in which the configurable multiplexers 208a-208b and the sensing circuits 210a-210b utilize two columns in the memory 202 instead of four columns. The system 200B is Figure 2A One embodiment of the system 200A described in .

[0056] As described above, each column 204a-204h of the memory array 202 is in electrical communication with a corresponding partial sum circuit 206 that is configured to deduce or determine the sum of the cell values ​​from each cell in the corresponding column. Figure 2AIn the illustrated embodiment, data labeled k1[a1], k1[a2], ... and k1[n1], k1[n2], ... are used for the first computing cluster and are stored in the cells in columns 204a-204b; data labeled k2[p1], k2[p2], k2[p3], k2[x1], and k2[x2] are used for the second computing cluster and are stored in the cells in columns 204c-204d; data labeled k3[a1], k3[a2], ... and k3[n1], k3[n2], ... are used for the third computing cluster and are stored in the cells in columns 204e-204f; and data labeled k4[p1], k4[p2], k4[p3], k4[x1], and k4[x2] are used for the fourth computing cluster and are stored in the cells in columns 204g-204h.

[0057] Each multiplexer 208a-208b includes four physical data lines 216. Because each compute cluster is stored in two columns 204 in the memory 202, each multiplexer 208a-208b is reconfigured to cycle through two data lines 216 for a given compute cluster. As described above, in some embodiments, a signal received from the first clock 212 via the first clock input 228 can initialize the multiplexers 208a-208b to cycle through the data lines 216, but each individual data line cycle is triggered by a signal received from the second clock 214 via the second clock input 226. The signal from the second clock 214 is utilized so that the number of cycles of the multiplexer data lines is equal to the number of columns used by a single compute cluster.

[0058] For example, when multiplexer 208a receives clock signal CLK1 from first clock 212 to start the computation cluster loop, multiplexer 208a cycles to the data line corresponding to partial sum PS1 upon receiving the next clock signal C1 from second clock 214, so that the partial sum value from PS1 is obtained by sensing circuit 210a. When clock signal C2 is received from second clock 214, multiplexer 208a cycles to the data line corresponding to partial sum PS2, so that its partial sum value is obtained by sensing circuit 210a and combined with its previously held value. After sensing circuit 210a has combined the values ​​from PS1 and PS2, the resulting value is output (as the corresponding computation cluster output) to, for example, output buffer 232a.

[0059] Then, multiplexer 208a receives second signal CLK2 from first clock 212 to start another computing cluster cycle, and another computing cluster cycle starts the second computing cluster cycle. When receiving clock signal C3 from second clock 214, multiplexer 208a cycles to the data line corresponding to partial sum PS3, so that its partial sum value is obtained by sensing circuit 210a. When receiving clock signal C4 from second clock 214, multiplexer 208a cycles to the data line corresponding to partial sum PS4, so that its partial sum value is obtained by sensing circuit 210a and combined with its previously maintained value. After sensing circuit 210a has combined the values ​​from PS3 and PS4, the resulting value is output to output buffer 232a as the corresponding computing cluster output.

[0060] Similar to multiplexer 208a, when multiplexer 208b receives clock signal CLK1 from first clock 212 to initiate the computation cluster loop, multiplexer 208b loops over the data lines corresponding to partial sums PS5 and PS6 upon receiving the next clock signals C1 and C2 from second clock 214. Upon receiving clock signal C2 from second clock 214, multiplexer 208b loops over the data line corresponding to partial sum PS6, allowing its partial sum value to be obtained by sense circuit 210b and combined with its previously held value. After sense circuit 210b has combined the values ​​from PS5 and PS6, the resulting value is output as the corresponding computation cluster output, for example, to output buffer 232b.

[0061] Multiplexer 208b then receives a second signal, CLK2, from first clock 212 to initiate another computational cluster cycle, which in turn initiates a second computational cluster cycle. Upon receiving clock signal C3 from second clock 214, multiplexer 208b cycles to the data line corresponding to partial sum PS7, allowing its partial sum value to be obtained by sensing circuit 210b. Upon receiving clock signal C4 from second clock 214, multiplexer 208b cycles to the data line corresponding to partial sum PS8, allowing its partial sum value to be obtained by sensing circuit 210b and combined with its previously held value. After sensing circuit 210b has combined the values ​​from PS7 and PS8, the resulting value is output to output buffer 232b as the corresponding computational cluster output.

[0062] Once the maximum number of physical multiplexer cycles is reached (e.g., in Figure 2B ), the multiplexers 208a-208b wait for the next first clock signal from the first clock 212 to start the multiplexers 208a-208b to cycle through the data lines 216 again based on the compute cluster cycle size.

[0063] As a result, system 200B obtains four results for the four computing clusters: one result is obtained by sense circuit 210a for columns 204a-204b, a second result is obtained by sense circuit 210a for columns 204c-204d, a third result is obtained by sense circuit 210b for columns 204e-204f, and a fourth result is obtained by sense circuit 210b for columns 204g-204h. Furthermore, because multiplexers 208a-208b utilize the same clock signal (the first clock signal, or a combination of the first and second clock signals) to cycle through their respective data lines 216, the first two results for the first two computing clusters (columns 204a-204b and 204e-204f) are obtained during the first computing cluster cycle time, while the second two results for the second two computing clusters (columns 204c-204d and 204g-204h) are obtained during the second computing cluster cycle time.

[0064] Similar to the above, the multiplexers 208a-208b can utilize the clock signal from the first clock 212 to instruct the multiplexers 208a-208b to cycle to the first data line 216 upon receiving the next clock signal from the second clock 214, allowing each multiplexer cycle to be controlled by the clock signal received from the second clock 214, or they can utilize one clock signal from the first clock 212 and one clock signal from the second clock 214 to access the corresponding two data lines 216 of a given computing cluster, or they can utilize each separate clock signal received from the first clock 212 to cycle to the next data line. Similarly, the clock signals received from the first clock 212, the second clock 214, or a combination thereof can also be used to trigger the sensing circuits 210a-210b, respectively, to obtain the current output values ​​of the multiplexers 208a-208b.

[0065] In some embodiments, the number of data lines for a given compute cluster can be controlled by the first clock signal 212. For example, the first clock signal can maintain a value of "1" for two second clock signals and then drop to a value of "0." The change to "0" indicates the start of a new compute cluster cycle. In another example, other circuit devices (not shown) can store the number of columns associated with a given compute cluster cycle and monitor the number of second clock signals that enable the multiplexer to cycle to the corresponding number of data lines. Once the number of second clock signals equals the number of columns associated with the given compute cluster cycle, the other circuit devices release the next first clock signal to start a new compute cluster cycle.

[0066] In other embodiments, as described above, multiplexers 208a-208b can have a compute cluster cycle size input line (not shown) that instructs the multiplexer regarding the number of columns associated with a given compute cluster cycle. For example, a single input line can be used such that "0" indicates two columns and "1" indicates four columns. Other numbers of compute cluster cycle size input lines can also be used to represent different compute cluster cycle sizes. Similarly, sense circuits 210a-210b can utilize the compute cluster cycle size input line or a clock signal to instruct the sense circuit output to hold a value.

[0067] although Figure 2A and Figure 2B An embodiment is shown in which multiplexers 208a-208b cycle through four data lines 216 so that each sensing circuit 210a-210b cycles through a combination of partial sums for one or more compute clusters, but some implementations may desire the sensing circuits to combine more partial sums (e.g., eight partial sums). For example, in some cases and implementations, a very large compute cluster may not be able to optimally utilize memory when using two or four columns due to a poor aspect ratio. To improve the aspect ratio, such as Figure 2C As shown, data of a large computing cluster can be stored in eight columns instead of two or four columns as described above.

[0068] As mentioned above, Figure 2C The system 200C in FIG. 1 illustrates an embodiment in which the configurable multiplexers 208a-208b and the sensing circuits 210a-210b are reconfigured to utilize eight columns in the memory 202 instead of two or four columns. The system 200C is Figure 2A The system 200A described in Figure 2B An embodiment of the system 200B described in.

[0069] As described above, each column 204a-204h of the memory array 202 is in electrical communication with a corresponding partial sum circuit 206 that is configured to deduce or determine the sum of the cell values ​​from each cell in the corresponding column. Figure 2A and Figure 2B In the illustrated embodiment, data labeled k1[a1], ..., k1[a4], k1[n1], ..., k1[n3], k1[p1], ..., k1[p3], k1[x1], k1[x2], k1[e1], ..., k1[e4], k1[h1], ..., k1[h3], k1[q1], ..., k1[q3], k1[y1], and k1[y2] are used to compute clusters and are stored in cells of columns 204a-204h.

[0070] However, in the example shown, multiplexer 208a is reconfigured to have eight data lines 216. In some embodiments, multiplexer 208a is reconfigured so that the additional data lines are provided by a circuit device that was previously used as another multiplexer. In other embodiments, a second multiplexer (e.g., Figure 2A or Figure 2B Multiplexer 208b) in a is controlled or "slaved" by multiplexer 208a to utilize its data lines. For example, after multiplexer 208a has cycled through its physical data lines, multiplexer 208a can provide a clock signal to a second multiplexer.

[0071] As described above, in some embodiments, a signal received from the first clock 212 via the first clock input 228 may initialize the multiplexer 208a to cycle through the data lines 216, but each individual data line cycle is triggered by a signal received from the second clock 214 via the second clock input 226. The signal from the second clock 214 is utilized so that the number of multiplexer data lines cycled equals the number of columns utilized by a single compute cluster cycle.

[0072] For example, when the multiplexer 208a receives the clock signal CLK from the first clock 212 to start the computation cluster cycle, the multiplexer 208a waits for the clock cycle C1 from the second clock 214 to cycle to the data line 216 corresponding to the partial sum PS1. When the clock signal C2 is received from the second clock 214, the multiplexer 208a cycles to the data line corresponding to the partial sum PS2, so that its partial sum value is obtained by the sensing circuit 210a and combined with its previously held value. The multiplexer 208a continues to cycle to the other data lines 216 corresponding to the partial sums PS3, PS4, PS5, PS6, PS7, and PS8, so that their partial sum values ​​are obtained by the sensing circuit 210a in response to the clock signals C3, C4, C5, C6, C7, and C8 from the second clock 214 and their partial sum values ​​are combined.

[0073] After the sense circuit 210a has combined the values ​​from PS1 to PS8, the resulting values ​​are output as corresponding compute cluster outputs, for example, to output buffer 232a. Once the maximum multiplexer cycle number of compute cluster cycles is reached (e.g., at Figure 2C If eight data lines are received (eight in the example shown), the multiplexer 208a waits for the next first clock signal from the first clock 212 to start the multiplexer 208a cycling through the data lines 216 again.

[0074] Similar to the above, the multiplexer 208a can use the clock signal from the first clock 212 to instruct the multiplexer 208a to cycle to the first data line 216 at the next clock signal received from the second clock 214, allowing each multiplexer cycle to be controlled by the clock signal received from the second clock 214, or the multiplexer 208a can use one clock signal from the first clock 212 and seven clock signals from the second clock 214 to access the data line 216 for the computing cluster, or the multiplexer 208a can use each individual clock signal received from the first clock 212 to cycle to the next data line. As described above, the multiplexer 208a can use one or more computing cluster cycle size input lines (not shown) to reconfigure the multiplexer 208a to the correct number of data lines for a given computing cluster cycle size. Similarly, the clock signals received from the first clock 212, the second clock 214, or a combination thereof can also be used to trigger the sense circuit 210a to obtain the current output value of the multiplexer 208a.

[0075] It should be appreciated that other numbers of multiplexers 208a-208b, sense circuits 210a-210b, and partial sum circuit 206 can be used for other sizes of memory 202. Similarly, multiplexers 208a-208b can utilize other numbers of physical data lines, as well as different computation cluster cycle sizes. In addition, partial sum circuit 206 can obtain a mathematical sum of each cell value in a corresponding column 204 in memory 202. However, in some embodiments, partial sum circuit 206 can obtain other mathematical combinations of cell values ​​(e.g., multiplications).

[0076] Although Figure 2A-2C The embodiment described in the embodiment is described as using the clock signals from the first clock 212 and the second clock 214 to read data from the memory 202, but the embodiment is not limited thereto. In some embodiments, such as in combination with Figures 3A-3D and Figures 4A-4D As shown and described, similar clock signals can also be used to write data to memory.

[0077] also, Figure 2A-2CAn embodiment is illustrated in which the number of multiplexer cycles to be employed during a read operation is determined and utilized based on the number of columns storing data for a given compute cluster, where the number of multiplexer cycles can be less than, equal to, or greater than the number of physical data lines of the multiplexer. In some other embodiments, during a write operation, the multiplexer can cycle through a determined number of multiplexer cycles (less than, equal to, or greater than the number of physical data lines of the multiplexer). For example, in some embodiments, the cluster cycle management circuitry 160 can compare the kernel batch size of a given compute cluster to one or more thresholds to determine the number of columns storing kernel data for the given compute cluster. In one such embodiment, each threshold can correspond to a particular number of multiplexer cycles, e.g., if the batch size is less than a first threshold, the number of multiplexer cycles is two, if the batch size is greater than the first threshold but less than a second threshold, the number of multiplexer cycles is four, and if the batch size is greater than the second threshold, the number of multiplexer cycles is eight. Other numbers of thresholds and corresponding numbers of multiplexer cycles (which can vary based on the neural network layer or processing being performed) can also be utilized.

[0078] Combined with the above Figure 2A-2C The described embodiments illustrate various embodiments for utilizing configurable multiplexers to read data from memory while also performing computations on the values ​​read from memory (e.g., obtaining and combining partial sums for a given compute cluster) in a high data processing neural network environment.

[0079] In order to read such a large amount of data from memory, the input data needs to be written to memory. As mentioned above, neural networks typically involve many similar calculations on large amounts of data. Moreover, neural networks are designed to learn specific characteristics from large amounts of training data. Some of this training data is useful, while other is not. As a result, as the system learns, erroneous data should eventually be removed from the system. However, certain situations may result in erroneous training or input data. For example, if consecutive images are read into memory in the same manner, the same pixel location is stored in the same memory location. However, if a pixel becomes "dead" (for example, due to a failure of the photosensor or related circuitry, or a failure of the memory cell), that memory location will always have the same value. As a result, the system may learn "dead" pixels during training or fail to correctly process consecutive input images.

[0080] Figures 3A to 3DA context diagram illustrates a use case in which a shift or line shift multiplexer 308 is used to sequentially write multiple data to a memory array 302, according to one embodiment. A shift multiplexer or line shift multiplexer refers to a multiplexer that includes circuitry that, when operated, causes the multiplexer to change the address of each physical data line of the multiplexer or to modify the selection of physical data lines for sequential data writing. In various embodiments, multiplexer 308 includes a barrel shifter that modifies the internal address of each data line 316 for sequential data writing of multiple data. Multiplexer 408 can be a write-only or read-write multiplexer.

[0081] Figure 3A An example 300A for writing to a first computing cluster is shown. Input data is written to columns 304a, 304b, 304c, and 304d via data lines addressed as 00, 01, 10, and 11, respectively, in that order. The first data of the first computing cluster is stored in column 304a, the second data of the first computing cluster is stored in column 304b, the third data of the first computing cluster is stored in column 304c, and the fourth data of the first computing cluster is stored in column 304d. Then, for example, the above combination Figure 2A-2C As described, the data in the memory is processed.

[0082] Figure 3B The diagram shows Figure 3A Example 300B of a second computing cluster write after the first computing cluster write. When receiving data for the second computing cluster write, the barrel shifter modifies the address of the data line so that the data lines addressed as 00, 01, 10, and 11 correspond to columns 304b, 304c, 304d, and 304a, respectively. The input data is written to columns 304b, 304c, 304d, and 304a in that order via the data lines addressed as 00, 01, 10, and 11, respectively. The first data of the second computing cluster is stored in column 304b, the second data of the second computing cluster is stored in column 304c, the third data of the second computing cluster is stored in column 304d, and the fourth data of the second computing cluster is stored in column 304a. Then, for example, in combination with the above Figure 2A-2C As described, the data in the memory is processed.

[0083] Figure 3C The diagram shows Figure 3BExample 300C of a third computing cluster write after the second computing cluster write. When receiving data for the third computing cluster write, the barrel shifter again modifies the address of the data line so that the data lines addressed as 00, 01, 10 and 11 correspond to columns 304c, 304d, 304a and 304b, respectively. The input data is written to columns 304c, 304d, 304a and 304b in this order via the data lines addressed as 00, 01, 10 and 11, respectively. The first data of the third computing cluster is stored in column 304c, the second data of the third computing cluster is stored in column 304d, the third data of the third computing cluster is stored in column 304a, and the fourth data of the third computing cluster is stored in column 304b. Then, for example, the above is combined Figure 2A-2C As described, the data in the memory is processed.

[0084] Figure 3D The diagram shows Figure 3C Example 300D of a fourth computing cluster write after the third computing cluster write. When receiving data for the fourth computing cluster write, the barrel shifter again modifies the address of the data line so that the data lines addressed as 00, 01, 10, and 11 correspond to columns 304d, 304a, 304b, and 304c, respectively. The input data is written to columns 304d, 304a, 304b, and 304c in this order via the data lines addressed as 00, 01, 10, and 11, respectively. The first data of the fourth computing cluster is stored in column 304d, the second data of the fourth computing cluster is stored in column 304a, the third data of the fourth computing cluster is stored in column 304b, and the fourth data of the fourth computing cluster is stored in column 304c. Then, for example, the above is combined Figure 2A-2C As described, the data in the memory is processed.

[0085] When receiving the data written by the fifth computing cluster, the barrel shifter modifies the address of the data line again so that Figure 3A As shown, data lines addressed as 00, 01, 10, and 11 correspond to columns 304a, 304b, 304c, and 304d, respectively. Figures 3A-3D As shown, the shifting of the data line address can continue in this manner for consecutive computing cluster writes. This shifting helps to semi-randomly store consecutive errors, which helps to eliminate such errors through learning / inference processes or helps to ignore such errors when analyzing the target image.

[0086] As described above, the first clock signal may be used to enable the multiplexer to cycle through its data lines for a read operation of a computing cluster cycle of a configurable number of data lines. Figures 3A-3D The use of such a clock signal for data write operations is also illustrated.

[0087] In some embodiments, the signal received from the first clock 212 via the first clock input 228 initializes the multiplexer 308 to cycle through the data line 216, but each individual data line cycle is triggered by a signal received from the second clock 214 via the second clock input 226. In this way, each individual compute cluster cycle is triggered by the clock signal from the first clock 212.

[0088] For example, in Figure 3A , when the multiplexer 308 receives the clock signal CLK from the first clock 212 to initiate a compute cluster write cycle, the multiplexer 308 cycles to address 00 of the data line 316 in response to C1 from the second clock 214, allowing data to be written to column 304a. Clock signal C2 causes the multiplexer 308 to cycle to address 01 of the data line 316, allowing data to be written to column 304b. Multiplexer 308 continues to operate in a similar manner, causing clock signals C3 and C4 to write data to columns 304c and 304d, respectively.

[0089] Figure 3B Similarly, however, when the multiplexer 308 receives the clock signal CLK from the first clock 212 to initiate the compute cluster write cycle and subsequently receives C1 from the second clock 214, the multiplexer 308 cycles to address 00 of the data line 316 so that data can be written to column 304b (not Figure 3A Clock signals C2, C3, and C4 received from the second clock 214 cause the multiplexer 308 to cycle addresses 01, 10, and 11, respectively, to the data line 316, thereby writing data to columns 304c, 304d, and 304a, respectively.

[0090] Likewise, in Figure 3C In the example, when the multiplexer 308 receives the clock signal CLK from the first clock 212, the computing cluster write cycle is started. When receiving C1 from the second clock 214, the multiplexer 308 cycles to the data line 316 address 00 so that data can be written to the column 304c (not Figure 3B The clock signals C2, C3, and C4 received from the second clock 214 cause the multiplexer 308 to cycle the addresses 01, 10, and 11 to the data line 316, respectively, so that data can be written to columns 304d, 304a, and 304b, respectively.

[0091] In addition, Figure 3D In the embodiment, when the multiplexer 308 receives the clock signal CLK from the first clock 212, the multiplexer 208 starts the computing cluster write cycle in response to receiving C1 from the second clock 214 and cycles to the data line 316 address 00 so that data can be written to the column 304d (not Figure 3CThe clock signals C2, C3, and C4 received from the second clock 214 cause the multiplexer 308 to cycle addresses 01, 10, and 11 to the data lines 316, respectively, so that data can be written to columns 304a, 304b, and 304c, respectively.

[0092] As shown, the multiplexer 308 writes to the four data lines 216 using one clock signal from the first clock 212 and four clock signals from the second clock 214. However, embodiments are not limited thereto. Instead, in other embodiments, receiving a clock signal from the first clock 212 may instruct the multiplexer 308 to cycle the address 00 to the first data line 316. In this manner, for a given compute cluster cycle write operation, the first clock 212 is used to cycle to the first data line and the second clock 214 is used to cycle through each corresponding data line. In this example, the multiplexer 308 will use three clock signals from the second clock 214 to write data to the columns 304a-304d.

[0093] In other embodiments, each signal received from the first clock 212 via the first clock input 228 triggers an individual data line cycle. For example, when the multiplexer 308 receives a first clock signal from the first clock 212, the multiplexer 308 cycles to data line address 00. However, when the multiplexer 308 receives a second clock signal from the first clock 212, the multiplexer 308 cycles to data line address 01. The multiplexer 308 continues to operate in a similar manner for additional clock signals received from the first clock 212. Once the maximum number of multiplexer cycles is reached in the compute cluster cycle (e.g., at Figures 3A-3D If four are present in the example shown), the multiplexer 308 cycles back to data line address 00 on the next clock signal. In such an embodiment, the second clock 214 is optional and not used.

[0094] although Figures 3A-3D The diagram illustrates the use of a multiplexer with a built-in barrel shifter to change the address of the data line, but the embodiments are not limited thereto. Instead, in some embodiments, other shifting circuitry (e.g., circuitry external to the multiplexer) can be used as a pre-decoder for the multiplexer to shift and select the data line for the multiplexer without changing the internal data line address of the multiplexer.

[0095] Figures 4A-4D The use case context diagram of pre-decoding shift for writing multiple data consecutively into a memory array is illustrated. The system includes a multiplexer 408, a memory array 402, and a data line selection circuit device 406. The multiplexer 408 can be a write-only or read-write multiplexer. Figures 3A-3DUnlike the example shown (in which the addresses of the data lines change), each data line 416 corresponds to an unchanging address. Thus, data line address 00 corresponds to column 404a in memory 402, data line address 01 corresponds to column 404b in memory 402, data line address 10 corresponds to column 404c in memory 402, and data line address 11 corresponds to column 404d in memory 402.

[0096] Figure 4A An example 400A for writing to a first computing cluster is illustrated. Data line selection circuitry 406 selects the order in which data lines 416 of multiplexer 408 are cycled to write data. In this example, data line selection circuitry 406 selects the data line addresses as 00, 01, 10, and 11 in that order and outputs the data line addresses to multiplexer 408 via address selection line 418. The first data of the first computing cluster is stored in column 404a, the second data of the first computing cluster is stored in column 404b, the third data of the first computing cluster is stored in column 404c, and the fourth data of the first computing cluster is stored in column 404d. Then, for example, the above is combined with Figure 2A-2C As described, the data in the memory is processed.

[0097] Figure 4B The diagram shows Figure 4A 400B. When receiving data for the second computing cluster to write, the data line selection circuit device 406 modifies the order in which the data lines 416 of the multiplexer 408 are cycled to write the data. In this example, the data line selection circuit device 406 selects the data line addresses as 01, 10, 11, and 00 in this order, and outputs the data line addresses to the multiplexer 408 via the address selection line 418. The first data of the second computing cluster is stored in column 404b, the second data of the second computing cluster is stored in column 404c, the third data of the second computing cluster is stored in column 404d, and the fourth data of the second computing cluster is stored in column 404a. Then, for example, the above is combined with Figure 2A-2C As described, the data in the memory is processed.

[0098] Figure 4C The diagram shows Figure 4B400C of an example of a third computing cluster being written after the second computing cluster is written. When data for the second computing cluster is received, the data line selection circuit device 406 modifies the order in which the data lines 416 of the multiplexer 408 are cyclically written with data. In this example, the data line selection circuit device 406 selects the data line addresses as 10, 11, 00, and 01 in this order, and outputs the data line addresses to the multiplexer 408 via the address selection line 418. The first data of the third computing cluster is stored in column 404c, the second data of the third computing cluster is stored in column 404d, the third data of the third computing cluster is stored in column 404a, and the fourth data of the third computing cluster is stored in column 404b. Then, for example, in combination with the above Figure 2A-2C As described, the data in the memory is processed.

[0099] Figure 4D The diagram shows Figure 4C 400D of an example of a fourth computing cluster write after a third computing cluster write. When data for a second computing cluster write is received, the data line selection circuit device 406 modifies the order in which the data lines 416 of the multiplexer 408 are cycled to write the data. In this example, the data line selection circuit device 406 selects the data line addresses as 11, 00, 01, and 10 in that order and outputs the data line addresses to the multiplexer 408 via the address selection line 418. The first data of the fourth computing cluster is stored in column 404d, the second data of the fourth computing cluster is stored in column 404a, the third data of the fourth computing cluster is stored in column 404b, and the fourth data of the fourth computing cluster is stored in column 404c. Then, for example, in combination with the above Figure 2A-2C As described, the data in the memory is processed.

[0100] When data for writing to the fifth computing cluster is received, as shown in FIG. Figure 4A As shown in D, the data line selection circuit device 406 modifies the order of the data lines 416 to data line addresses 00, 01, 10 and 11 in this order. Figures 4A-4D As shown, for consecutive computing cluster writes, the shifting of the data line address can continue in this way. Figures 3A-3D As discussed in , the shifting allows for a semi-random storage of consecutive errors, which allows the errors to be removed via a learning / inference process or ignored in the analysis of the target image.

[0101] As above Figures 3A-3D As described in , a first clock signal can be used to enable a multiplexer to cycle through its data lines for a write operation for a computing cluster cycle and a second clock for each individual cycle. Figures 4A-4D The described embodiments may also utilize the first clock and the second clock in a similar manner. Figures 4A-4D In the example shown, the data line selection circuitry 406 receives clock signals from the first clock 212 and the second clock 214 rather than from the multiplexer 408 itself. Instead, changes on the address select lines 418 are triggered by the clock signals and trigger the multiplexer to select the appropriate data line 416. In other embodiments, similar to Figures 3A-3D As shown, the multiplexer 408 may include input terminals (not shown) to receive the first clock signal and the second clock signal.

[0102] The data line selection circuitry 406 receives a signal from the first clock 212 to initialize the selection of the data line address, but each individual data line address cycle is triggered by a signal received from the second clock 214. In this way, each individual compute cluster cycle is triggered by the clock signal from the first clock 212.

[0103] Similar to the description above, in some embodiments, the data line selection circuitry 406 utilizes one clock signal from the first clock 212 and four clock signals from the second clock 214 to select four data line addresses for the multiplexer 408 to write to four different columns 404a-404d. In other embodiments, for a given compute cluster cycle write operation, receiving a clock signal from the first clock 212 can instruct the data line selection circuitry 406 to cycle to the first data line address and cycle through each of the remaining corresponding data line addresses using the second clock 214. In the illustrated example, the data line selection circuitry 406 will utilize one clock signal from the first clock 212 and three clock signals from the second clock 214 to select four data line addresses for the multiplexer to write data to the columns 404a-404d. In other embodiments, each signal received from the first clock 212 via the first clock input 228 triggers individual data line address selections.

[0104] In various embodiments, respectively, Figures 3A-3D and Figures 4A-4D The input sequence of the barrel shifter and data line selection circuitry can also change the address / data line selection sequence to a different sequence. Moreover, in some embodiments, not all address inputs / data line addresses can be shifted. For example, for large multiplexer ratios, only a subset of address inputs / data line addresses can be shifted to the dithered data stored in the memory under consideration.

[0105] Now refer to Figure 5-Figure 8 To describe the operation of certain aspects of the present disclosure. In at least one of the various embodiments, as described herein, respectively, in combination with Figure 5-Figure 8 The described processes 500, 600, 700, and 800 may be implemented by one or more components or circuits associated with an in-memory computing element.

[0106] Figure 5 A logic flow diagram generally illustrates one embodiment of a process 500 for using a Figure 2A-2C The configurable multiplexer and sensing circuit shown are used to read data from the memory array. After the start block, the process 500 starts at block 502, where cluster data is stored in the column-row memory array. Figure 6 and Figure 7 Various embodiments that facilitate storage of data in memory are described in more detail.

[0107] Process 500 proceeds to block 504 where a partial sum is calculated for each column in memory.

[0108] Process 500 continues at block 506 where the number of columns is selected for the compute cluster. In various embodiments, this is the number of columns that store data for a single compute cluster (which may be referred to as the compute cluster loop size).

[0109] Process 500 continues to decision block 508, where a determination is made as to whether the number of columns selected for the computational cluster is greater than the number of column multiplexer cycles. In at least one embodiment, this determination is based on a comparison of the number of columns selected and the number of physical data lines on the multiplexer. If the number of columns selected for the computational cluster is greater than the number of multiplexer cycles, process 500 proceeds to block 510; otherwise, process 500 proceeds to block 512.

[0110] At block 510, the multiplexer is enlarged. In some embodiments, the number of data lines supported by the multiplexer is reconfigured to accommodate the number of columns selected for the compute cluster. In at least one embodiment, the enlargement of the multiplexer can include utilizing a second multiplexer as the enlarged portion of the multiplexer. For example, this can be accomplished by reconfiguring the first multiplexer to utilize the circuitry of the second multiplexer to provide additional data lines or by configuring the first multiplexer to control the number of columns selected for the compute cluster. Figure 2C The process 500 proceeds to block 512 .

[0111] If, at decision block 508, the number of columns selected is equal to or less than the number of multiplexer cycles, or after block 510, process 500 proceeds to block 512. At block 512, the multiplexer and sensing circuits are enabled to cycle through the partial sums of each column in the memory associated with the compute cluster. As described above, the first clock signal can enable the multiplexer to cycle through the partial sums of the compute cluster cycles.

[0112] Process 500 proceeds to block 514 where the current partial sum for the current multiplexer cycle is obtained and combined with the computed cluster value.

[0113] Process 500 continues to decision block 516, where a determination is made as to whether the number of partial sums obtained is equal to the number of columns selected for the compute cluster. In various embodiments, this determination is made based on a comparison of the number of multiplexer cycles selected (the number of partial sums obtained) and the number of columns selected (the compute cluster cycle size). If the number of partial sums obtained is equal to the number of columns selected, process 500 proceeds to block 520; otherwise, process 500 proceeds to block 518.

[0114] At block 518, the multiplexer loop is incremented to obtain the next partial sum for the compute cluster.After block 518, process 500 loops to block 514 to obtain the next partial sum for the incremented multiplexer loop.

[0115] If, at decision block 516, the number of partial sums obtained is equal to the number of columns selected for the computation cluster, process 500 proceeds from decision block 516 to block 520. At block 520, the combined computation cluster value is output to a buffer that can be accessed by another component of the system for further processing (e.g., input to the next layer of neural network processing).

[0116] The process 500 continues at decision block 522 where it is determined whether the maximum number of multiplexer cycles has been reached. In various embodiments, the maximum number of multiplexer cycles is the physical number of data lines associated with the multiplexer. If the maximum number of multiplexer cycles has been reached, the process 500 proceeds to decision block 524; otherwise, the process 500 loops back to block 512 to initiate the multiplexer and sense circuit loop through the next portion and set of the next compute cluster for the selected number of columns.

[0117] At decision block 524, a determination is made as to whether additional data is to be stored in memory, for example, for the next set of computing clusters. If additional data is to be stored in memory, process 500 loops back to block 502; otherwise, process 500 terminates or otherwise returns to the calling process to perform other actions.

[0118] Figure 6 A logic flow diagram generally illustrates one embodiment of a process 600 for employing a multiplexer that modifies a data line address for sequentially writing multiple data lines such as a Figures 3A-3D The memory array is shown.

[0119] After a start block, process 600 begins at block 602 where a command is received to store data in a column-row memory array for cluster computing operations. The command may be a write command or data on a specific input data line.

[0120] Process 600 proceeds to block 604, where each multiplexer data line is associated with a corresponding column in the memory. In various embodiments, the association between a column in the memory and a particular data line of the multiplexer includes selecting and assigning a multiplexer address to the particular data line. In one example embodiment, the first data line of the multiplexer is assigned to address 00, the second data line of the multiplexer is assigned to address 01, the third data line of the multiplexer is assigned to address 10, and the fourth data line of the multiplexer is assigned to address 11.

[0121] Process 600 continues at block 606 where next data for the cluster computing operation is received. In various embodiments, the next data is input data to be stored in memory for a particular computing cluster.

[0122] Process 600 continues to block 608, where the next data for cluster calculation is stored in the column associated with the multiplexer data line of the current multiplexer cycle. In various embodiments, the current multiplexer cycle is the current address of the multiplexer data line. For example, when receiving first data, the current multiplexer cycle is the first cycle and address 00 can be used; when receiving second data, the current multiplexer cycle is the second cycle and address 01 can be used, etc.

[0123] Next, process 600 continues at decision block 610, where a determination is made as to whether the maximum data line cycle for the multiplexer has been reached. In some embodiments, the maximum data line cycle is the number of physical data lines of the multiplexer. In other embodiments, the maximum data line cycle may be a selected number of data lines or columns for a particular compute cluster. If the final or maximum data line cycle for the multiplexer has been reached, process 600 proceeds to block 612; otherwise, process 600 proceeds to block 616.

[0124] At block 616, the multiplexer cycles to the next data line. In at least one embodiment, the multiplexer increments its current multiplexer cycle data line address to the next address. For example, if data was stored via address 00 at block 608, the next multiplexer data line address may be 01. The process 600 then cycles to block 606 to receive the next data to be stored in the memory for cluster computing operations.

[0125] If the maximum multiplexer data line cycle has been reached at decision block 610, process 600 proceeds from decision block 610 to block 612. At block 612, cluster computing operations are performed. In some embodiments, a read command may be received that is communicated via a multiplexer and sensing circuit (e.g., as described herein, including Figure 2A-2C and Figure 5 500 in the process to start reading the storage column.

[0126] Process 600 proceeds to decision block 614 where it is determined whether a command to store data for a continuous cluster computing operation has been received. If a continuous cluster computing operation is to be performed, process 600 proceeds to block 618; otherwise, process 600 terminates or returns to the calling process to perform other actions.

[0127] At block 618, the association between each multiplexer data line and a column in the memory is modified. In at least one non-limiting embodiment, the multiplexer includes a barrel shifter that modifies the data line / column association between the stored data for successive cluster computing operations. For example, Figures 3A-3D As shown, the barrel shifter modifies the addresses of the data lines. If the previous association indicated that the first data line of the multiplexer was assigned address 00, the second data line of the multiplexer was assigned address 01, the third data line of the multiplexer was assigned address 10, and the fourth data line of the multiplexer was assigned address 11, then the modified association may be that the first data line of the multiplexer is assigned address 01, the second data line of the multiplexer is assigned address 10, the third data line of the multiplexer is assigned address 11, and the fourth data line of the multiplexer is assigned address 00.

[0128] After block 618, process 600 loops to block 606 to receive data for the next cluster computer operation.

[0129] Figure 7 A logic flow diagram generally illustrates one embodiment of a process 700 for sequentially writing multiple data using pre-decode shifting. Figures 4A-4D The memory array is shown.

[0130] After the start block, process 700 begins at block 702 where a command is received to store data in a column-row memory array for cluster computing operations. In various embodiments, block 702 may be implemented in conjunction with the above. Figure 6 The embodiment described in block 602 in FIG.

[0131] Process 700 proceeds to block 704 where each multiplexer data line is associated with a corresponding column in the memory. In various embodiments, block 704 may employ a combination of the above Figure 6The embodiment described in block 604 in FIG.

[0132] Process 700 continues at block 706 where an order of multiplexer data lines for writing data for cluster computing operations is determined. In some embodiments, the initial order of the multiplexer data lines may be preselected or predetermined. In other embodiments, the initial data line order may be random.

[0133] In at least one embodiment, the multiplexer data line sequence is the order in which data line addresses are provided to the multiplexer for data writing. For example, the selection circuitry may determine that the initial sequence of data line addresses is 00, 01, 10, and 11. However, other sequences may be used. Similarly, other numbers of addresses may be used depending on the size of the multiplexer.

[0134] Process 700 continues to block 708 where an initial multiplexer data line is selected based on the determined data line sequence. In at least one embodiment, the initial multiplexer data line selection is the first data line address in the determined data line sequence and is output to the multiplexer.

[0135] Process 700 continues at block 710 where the next data for the cluster computing operation is received. In various embodiments, block 710 may employ a combination of the above Figure 6 The embodiment described in block 606 in FIG.

[0136] Process 700 continues to block 712 where the next data for cluster computation is stored in the column associated with the selected multiplexer data line. In at least one embodiment, the data line associated with the data line address for the selected data line is used to store the received data in the corresponding column in the memory.

[0137] Next, the process 700 continues at decision block 714 where it is determined whether the maximum data line cycle of the multiplexer has been reached. In various embodiments, decision block 714 may be implemented using Figure 6 If the last or maximum data line cycle of the multiplexer has been reached, the process 700 proceeds to block 716; otherwise, the process 700 proceeds to block 720.

[0138] At block 720, the multiplexer is instructed to cycle to the next data line based on the determined data line sequence. In at least one embodiment, the selection circuit device increments the data line address in the determined data line sequence. For example, if the previous data line address is 00 and the data line sequence is addresses 00, 01, 10, 11, then the next data line address is 01. The process 700 then cycles to block 710 to receive the next data to be stored in the memory for cluster computing operations.

[0139] If at decision block 714, the maximum multiplexer data line cycle has been reached, then process 700 proceeds from decision block 714 to block 716. At block 716, cluster computing operations are performed. In various embodiments, block 716 may employ a combination of the above Figure 6 The embodiment described in block 612 of FIG.

[0140] Process 700 continues to decision block 718 where it is determined whether a command to store data for a continuous cluster computing operation has been received. In various embodiments, decision block 718 may employ a combination of the above Figure 6 If a continuous cluster computing operation is to be performed, the process 700 proceeds to block 722; otherwise, the process 700 terminates or otherwise returns to the calling process to perform other actions.

[0141] At block 722, a new multiplexer data line sequence is determined. The new multiplexer data line sequence may be random, or the new multiplexer data line sequence may increment the previous address. For example, if the previous sequence was addresses 00, 01, 10, 11, the new sequence may be addresses 01, 10, 11, 00. In at least one non-limiting embodiment, the data line selection circuitry includes circuitry (e.g., a barrel shifter) that selects a data line sequence to store data for successive cluster computing operations. For example, the data line selection circuitry selects the data line sequence that instructs the multiplexer to cycle through its data lines to perform operations such as Figures 4A-4D Different orders of consecutive cluster computing operations are shown.

[0142] After block 722, process 700 loops to block 708 to receive data for the next cluster computer operation.

[0143] Figure 8 Illustrated is a logic flow diagram generally showing one embodiment of a process for initiating multiplexer reads and writes to a memory array using a first clock (e.g., an external system clock) and initiating each cycle of the multiplexer reads and writes using a second clock (e.g., an internal memory clock).

[0144] After the start block, process 800 begins at block 802 where a command for a cluster computing operation on data in a column-row memory is received. In some embodiments, the command may be, for example, to read data from the memory (e.g., as described above in conjunction with Figure 2A-2C In other embodiments, the command may be, for example, to write data to the memory (e.g., as described above in conjunction with Figures 3A-3D and Figures 4A-4D described).

[0145] Process 800 proceeds to block 804 where a first clock signal is used to enable a column multiplexer to facilitate a read or write operation to the memory. In various embodiments, the first clock is separate from the memory itself. In at least one embodiment, the first clock can be separate from the memory cell or chip and can be the overall system clock.

[0146] Process 800 continues at block 806 where a second clock is triggered to step through the multiplexer loop. In at least one embodiment, the first clock signal may trigger the first multiplexer loop. In other embodiments, the first multiplexer loop is triggered using the second clock signal after the first clock signal.

[0147] Process 800 continues to block 808 where data is read from or written to the memory at the column corresponding to the current multiplexer cycle. In various embodiments, block 808 may employ a Figure 5 Box 514, Figure 6 Box 608 or Figure 7 An embodiment of block 712 in .

[0148] Process 800 continues to decision block 810 where it is determined whether the maximum number of multiplexer cycles has been reached. In various embodiments, decision block 810 may employ Figure 5 522 in the decision box, Figure 6 Decision box 610 in or Figure 7 If the maximum multiplexer cycle is reached, the process 800 proceeds to block 812; otherwise, the process 800 proceeds to block 716.

[0149] At block 716, the multiplexer is incremented to the next multiplexer cycle at the next second clock signal. In various embodiments, the current memory cycle address is incremented in response to receiving the next second clock signal. After block 716, process 800 loops to block 808 to read or write data to the next column in the memory associated with the next multiplexer cycle.

[0150] If the maximum multiplexer cycle is reached at decision block 810, process 800 proceeds from decision block 810 to block 812. At block 812, the memory is placed in a low power mode or a memory retention mode. In some embodiments, the memory can transition between an operating mode and a memory retention mode or between a low power mode and a retention mode using other states. Such other states may include, but are not limited to, an idle state, a wait state, a light sleep state, etc. In other embodiments, block 812 may be optional, and the memory may not enter a retention or low power mode after the maximum multiplexer cycle is reached at decision block 810.

[0151] Embodiments of the aforementioned processes and methods may include Figure 5-Figure 8 Additional actions not shown in the Figure 5-Figure 8 All actions shown in the Figure 5-Figure 8 The actions shown in , can be combined and modified in various ways. For example, Figure 8 Process 800 in may omit act 812, combine acts 804 and 806, etc. when writing successive computation clusters to and reading from memory for successive neural network layer processing.

[0152] Some embodiments may take the form of or include a computer program product. For example, according to one embodiment, a computer-readable medium is provided that includes a computer program adapted to perform one or more of the above methods or functions. The medium may be a physical storage medium (e.g., such as a read-only memory (ROM) chip), or a disk (e.g., a digital versatile disk (DVD-ROM)), a compact disk (CD-ROM), a hard disk, a memory, a network, or a portable medium article to be read by an appropriate drive or via an appropriate connection (including one or more bar codes or other related codes encoded as stored on one or more such computer-readable media and readable by an appropriate reader device).

[0153] In addition, in some embodiments, some or all of the methods and / or functions may be implemented or provided in other manners (e.g., at least in part in firmware and / or hardware, including but not limited to one or more application-specific integrated circuits (ASICs), digital signal processors, discrete circuit devices, logic gates, standard integrated circuits, controllers (e.g., by executing appropriate instructions and including microcontrollers and / or embedded controllers), field programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), etc., as well as devices employing RFID technology and various combinations thereof).

[0154] As described below, the embodiments may be further or alternatively generalized.

[0155] The system can be summarized as including: a memory array having a plurality of cells, the plurality of cells being arranged as a plurality of cell rows intersecting a plurality of cell columns; a column multiplexer having a plurality of data lines and a plurality of select lines, wherein each respective data line corresponds to a respective column of the plurality of columns, and the plurality of select lines indicate which of the plurality of data lines is selected within a given cycle, wherein the column multiplexer, in operation, cycles through the plurality of data lines to write data to the plurality of cells based on a data line address received via the plurality of select lines; and a data line selection circuit device, in operation, selecting a first cycle order for a first write operation to write a first plurality of data to the plurality of cells and selecting a second cycle order for a second write operation to write a second plurality of data to the plurality of cells, wherein the first and second cycle orders indicate an order in which the plurality of data lines are selected on each cycle and wherein the first cycle order is different from the second cycle order.

[0156] The column multiplexer may include a data line selection circuit device to select the plurality of data lines in a first cyclic order during a first write operation and to select the plurality of data lines in a second cyclic order during a second write operation. The data line selection circuit device may include a plurality of output lines corresponding to the plurality of select lines of the column multiplexer, and the data line selection circuit device, in operation, outputs data line addresses of the plurality of data lines via the plurality of output lines according to the first selected cyclic order to write a first plurality of data into the plurality of cells during the first write operation, and outputs data line addresses of the plurality of data lines via the plurality of output lines according to the second selected cyclic order to write a second plurality of data into the plurality of cells during the second write operation.

[0157] The system may include a cluster cycle management circuit device that, in operation, generates a plurality of control signals in response to a clock signal and provides the plurality of control signals to a column multiplexer to cycle through the plurality of data lines to write data to corresponding columns of the plurality of cell columns.

[0158] The system may include: a plurality of computation circuits, each computation circuit deriving a computation value based on cell values ​​in a corresponding column of cells in a plurality of cells and corresponding to a corresponding data line in a plurality of data lines of a column multiplexer; and a sensing circuit, wherein the sensing circuit, in operation, obtains computation values ​​from the plurality of computation circuits via the column multiplexer as the column multiplexer cycles through the plurality of data lines, and combines the computation values ​​obtained within a determined number of multiplexer cycles during a read operation. The determined number of multiplexer cycles may be less than the plurality of data lines of the column multiplexer. In operation, the sensing circuit may derive a first value based on the computation values ​​obtained via a first set of the plurality of data lines during a first portion of the read operation, and derive a second value based on the computation values ​​obtained via a second set of the plurality of data lines during a second portion of the read operation, wherein the first set of cycles and the second set of cycles comprise the determined number of multiplexer cycles. The column multiplexer can be operable to cycle through a second plurality of data lines, each of the second plurality of data lines corresponding to a respective second computation circuit in a plurality of second computation circuits, wherein each second computation circuit derives a computation value based on a cell value in a corresponding column of cells in the second plurality of cells. The second plurality of data lines for the column multiplexer can be provided by a second column multiplexer. The system can include a first clock and a second clock, the first clock being operable to initiate a read operation by the column multiplexer cycling through the plurality of data lines for a determined number of multiplexer cycles, and the second clock being operable to initiate each cycle of each data line of the column multiplexer for a determined number of multiplexer cycles during a read operation, such that the sensing circuit obtains the computation value from the corresponding computation circuit.

[0159] The method can be summarized as including: receiving a first command to store first data for a first cluster computing operation in a plurality of cells, the plurality of cells being arranged as a plurality of cell rows intersecting a plurality of cell columns; associating each of a plurality of multiplexer data lines with a corresponding column in a plurality of columns; determining a first data line order for storing the first data in the plurality of cells; looping through the plurality of data lines according to the first data line order to store the first data in corresponding columns in the plurality of columns to perform a first cluster computing operation; performing the first cluster computing operation; in response to completion of the first cluster computing operation, receiving a second command to store second data for a second cluster computing operation in the plurality of cells; determining a second data line order for storing the second data in the plurality of cells, the second data line order being different from the first data line order; and looping through the plurality of data lines according to the second data line order to store the second data in corresponding columns in the plurality of columns to perform a second cluster computing operation.

[0160] Determining a second data line order for storing the second data in the plurality of cells may include employing a barrel shifter to modify an address of each of the plurality of data lines based on the first data line order. Determining the first data line order for storing the first data in the plurality of cells may include selecting the first data line order and instructing a column multiplexer to cycle through the plurality of data lines according to the first data line order, and wherein determining a second data line order for storing the second data in the plurality of cells may include selecting the second data line order and instructing the column multiplexer to cycle through the plurality of data lines according to the second data line order.

[0161] Cycling through the plurality of data lines according to a first data line sequence may include initiating cycling through the plurality of data lines according to the first data line sequence in response to a first clock signal from a first clock, and initiating each cycle for each data line in the first data line sequence in response to a first set of clock signals from a second clock to write data to each corresponding column in the plurality of columns; wherein cycling through the plurality of data lines according to a second data line sequence may include initiating cycling through the plurality of data lines according to the second data line sequence in response to a second clock signal from the first clock, and initiating each cycle for each data line in the first data line sequence in response to a second set of clock signals from the second clock to write data to each corresponding column in the plurality of columns.

[0162] Performing a first cluster compute operation may include: computing a plurality of computed values ​​based on cell values ​​from a plurality of cell columns, wherein each respective computed value is computed based on cell values ​​from a respective cell column; selecting a number of multiplexer cycles for the first cluster compute operation based on a number of columns storing data for the first cluster compute operation; generating a result of the first cluster compute operation by cycling through a first subset of the plurality of computed values ​​using a column multiplexer for the selected number of multiplexer cycles, and combining respective computed values ​​from the first subset of computed values ​​using a sensing engine; and outputting the result of the first cluster compute operation. The method may further include cycling through a second subset of the plurality of computed values ​​using a second column multiplexer for the selected number of multiplexer cycles, and combining respective computed values ​​from the second subset of computed values ​​using a second sensing engine; and outputting the second result of the first cluster compute operation. The number of multiplexer cycles defined may be less than the number of data lines.

[0163] A computing device can be summarized as including: a device for receiving a plurality of data to be stored in a plurality of cells, the plurality of cells being arranged as a plurality of cell rows intersecting a plurality of cell columns; a device for determining a data line sequence, the data line sequence being used to store the received plurality of data in the plurality of cells; a device for cycling through the plurality of data lines according to the data line sequence to store the received plurality of data in corresponding columns of the plurality of columns; and a device for modifying the data line sequence to store subsequently received plurality of data in the plurality of cells.

[0164] The computing device may further include: means for initiating a cycle through a plurality of data lines based on a signal from a first clock; and means for initiating each cycle of each of the plurality of data lines according to a data line sequence based on a signal from a second clock. The computing device may further include: means for calculating a corresponding calculation value based on a cell value from each corresponding column of a plurality of cell columns; means for cycling through each corresponding calculation value for a determined number of calculation values; and means for combining corresponding calculation values ​​from a subset of calculation values ​​to generate a result of a data calculation cluster for the determined number of calculation values.

[0165] The system can be summarized as including: a memory array having a plurality of cells, the plurality of cells being arranged as a plurality of cell rows intersecting a plurality of cell columns; a column multiplexer having a plurality of data lines, wherein each respective data line corresponds to a respective column of the plurality of columns, wherein the column multiplexer is operative to initiate a cycle through the plurality of data lines for a read or write operation in response to a clock signal from a first clock to read data from the plurality of cells or write data to the plurality of cells, and wherein the column multiplexer is operative to initiate each cycle through each data line of the column multiplexer to read data from a respective column of the plurality of cell columns or write data to a respective column of the plurality of cell columns in response to a clock signal from a second clock separate from the first clock. The first clock may be a system clock external to the plurality of cells, and the second clock may be a memory refresh clock associated with the plurality of cells.

[0166] The system may include a data line selection circuit device that is operable to select a first cycle order for a first write operation to write a first plurality of data to a plurality of cells, and to select a second cycle order for a second write operation to write a second plurality of data to the plurality of cells, wherein the first and second cycle orders indicate an order in which the plurality of data lines are selected on each cycle, and wherein the first cycle order is different from the second cycle order. The data line selection circuit device may include a plurality of output lines corresponding to the plurality of select lines of the column multiplexer, and the data line selection circuit device, in operation, outputs data line addresses of the plurality of data lines via the plurality of output lines according to the first selected cycle order to write the first plurality of data to the plurality of cells during the first write operation, and outputs data line addresses of the plurality of data lines via the plurality of output lines according to the second selected cycle order to write the second plurality of data to the plurality of cells during the second write operation.

[0167] The column multiplexer can modify an address of each of a plurality of data lines of the column multiplexer during operation to write a plurality of consecutive data into a plurality of cells. The column multiplexer can select the plurality of data lines in a first cyclic sequence during a first write operation and in a second cyclic sequence during a second write operation.

[0168] The system may include: cluster cycle management circuitry, which, in operation, determines a number of multiplexer cycles based on the number of columns storing data associated with a read operation; a plurality of computation circuits, each computation circuit deriving a computation value based on cell values ​​in a corresponding column of cells in a plurality of cells and corresponding to a corresponding data line in a plurality of data lines of a column multiplexer; and sensing circuitry, which, in operation, obtains computation values ​​from the plurality of computation circuits via the column multiplexer as the column multiplexer cycles through the plurality of data lines and combines the computation values ​​obtained within a determined number of multiplexer cycles during a read operation. The determined number of multiplexer cycles may be less than the plurality of data lines of the column multiplexer. In operation, the sensing circuitry may derive a first value based on the computation values ​​obtained via a first set of the plurality of data lines during a first portion of the read operation, and derive a second value based on the computation values ​​obtained via a second set of the plurality of data lines during a second portion of the read operation, wherein the first and second cycle sets comprise the determined number of multiplexer cycles. The column multiplexer is operable to cycle through a second plurality of data lines, each of the second plurality of data lines corresponding to a respective second computation circuit of the plurality of second computation circuits, wherein each second computation circuit computes a computation value based on cell values ​​in a corresponding column of cells in the second plurality of cells.

[0169] The method can be summarized as including: receiving a command associated with a cluster computing operation to read data from or write data to a plurality of cells, the plurality of cells being arranged as a plurality of cell rows intersecting a plurality of cell columns; activating a multiplexer to cycle through a plurality of data lines in response to a clock signal from a first clock, each of the plurality of data lines corresponding to a corresponding column in a plurality of columns; and activating each cycle of the multiplexer for each of the plurality of data lines in response to a clock signal from a second clock to read data from or write data to each corresponding column in the plurality of columns. The first clock can be a system clock external to the plurality of cells, and the second clock is a memory refresh clock associated with the plurality of cells.

[0170] The method may further include: determining a data line order for storing the first data in the plurality of cells before enabling the multiplexer to cycle through the plurality of data lines; and wherein each cycle of enabling the multiplexer for each of the plurality of data lines comprises cycling through the plurality of data lines according to the data line order to store the data in a corresponding column of the plurality of columns.

[0171] The method may further include receiving a second command to write second data to the plurality of cells; determining a second data line sequence for storing the first data in the plurality of cells; enabling a multiplexer to cycle through the plurality of data lines in response to a second clock signal from the first clock; and enabling each cycle of the multiplexer for each of the plurality of data lines to store the second data in each corresponding column of the plurality of columns according to the second data line sequence and in response to an additional clock signal from the second clock. Determining the second data line sequence for storing the second data in the plurality of cells may include employing a barrel shifter to modify an address of each of the plurality of data lines based on the first data line sequence.

[0172] The method may further include: calculating multiple calculation values ​​based on cell values ​​from multiple cell columns, wherein each corresponding calculation value is calculated based on the cell values ​​from the corresponding cell column; selecting a number of multiplexer cycles for the cluster computing read operation based on the number of columns storing data for the cluster computing read operation; and looping through a first subset of the multiple calculation values ​​for the selected number of multiplexer cycles by employing a multiplexer, and combining corresponding calculation values ​​from the first subset of calculation values ​​by employing a sensing engine to generate a result of the cluster computing read operation.

[0173] The method may also include: in response to a clock signal from the first clock, starting a second multiplexer to cycle through a second plurality of data lines, each of the second plurality of data lines corresponding to a corresponding column in a second plurality of columns in the plurality of cells; in response to a clock signal from the second clock, starting each cycle of the second multiplexer for each data line of the second plurality of data lines to read data from each corresponding column in the second plurality of columns; and generating a second result of the cluster computing read operation by cycling through a second subset of the plurality of calculated values ​​for a selected number of multiplexer cycles and combining corresponding calculated values ​​from the second subset of calculated values ​​by employing a second sensing engine.

[0174] A computing device can be summarized as including: a device for receiving a command to read data from or write data to a plurality of cells, the plurality of cells being arranged as a plurality of cell rows intersecting a plurality of cell columns; a device for enabling a multiplexer to cycle through a plurality of data lines according to a first clock source, each of the plurality of data lines corresponding to a corresponding column in a plurality of columns; and a device for enabling each cycle of the multiplexer for each of the plurality of data lines to read data from or write data to each corresponding column in the plurality of columns according to a second clock source.

[0175] The computing device may further include: means for determining a data line order for storing a plurality of received data in a plurality of cells; means for looping through the plurality of data lines according to the data line order to store the received plurality of data in corresponding columns of a plurality of columns; and means for modifying the data line order to store subsequently received plurality of data in the plurality of cells. The computing device may further include: means for calculating corresponding calculation values ​​based on cell values ​​from each corresponding column of a plurality of cell columns; means for determining a number of calculation values ​​based on the number of columns storing data associated with the data calculation cluster; means for looping through each corresponding calculation value for a determined number of calculation values; and means for combining corresponding calculation values ​​from a subset of calculation values ​​to generate a result for the data calculation cluster for the determined number of calculation values.

[0176] The various embodiments described above can be combined to provide further embodiments. These and other changes can be made to the embodiments in light of the above detailed description. Generally, in the following claims, the terms used should not be construed to limit the claims to the specific embodiments disclosed in the specification and claims, but should be construed to include all possible embodiments and the full range of equivalents claimed. Therefore, the claims are not limited by the disclosure.

Claims

1. An electronic system comprising: a memory array having a first plurality of cells arranged as a plurality of cell rows intersecting a plurality of cell columns; a plurality of first calculation circuits, wherein each first calculation circuit is operable to derive a calculation value based on a cell value in a corresponding cell column in the first plurality of cells; a first column multiplexer that, in operation, cycles through a plurality of data lines, each data line of the plurality of data lines corresponding to a first computation circuit of the plurality of first computation circuits; a first sensing circuit that, in operation, obtains the calculated values ​​from the plurality of first calculation circuits via the first column multiplexer as the first column multiplexer cycles through the plurality of data lines and combines the calculated values ​​obtained within a determined number of multiplexer cycles; as well as A cluster cycle management circuit is configured to determine the determined number of multiplexer cycles based on the number of columns storing computational cluster data, wherein the first column multiplexer is configured to modify an address of each of the plurality of data lines of the first column multiplexer to write a continuous plurality of data to the first plurality of cells. 2 . The electronic system of claim 1 , wherein the determined number of multiplexer cycles is less than a plurality of physical data lines of the first column multiplexer.

3. The electronic system of claim 1 , wherein the calculated value is a partial sum of cell values ​​in the corresponding column, and the first sensing circuit is operable to derive a first sum from the obtained partial sum via a first set of the plurality of data lines during a first set of cycles of the first column multiplexer, and to derive a second sum from the obtained partial sum via a second set of the plurality of data lines during a second set of cycles of the first column multiplexer, wherein the first and second cycle sets have the determined number of multiplexer cycles.

4. The electronic system of claim 3 , wherein the first column multiplexer is operative to cycle through a second plurality of data lines, each data line in the second plurality of data lines corresponding to a respective second computation circuit in a plurality of second computation circuits, wherein each second computation circuit computes a partial sum based on cell values ​​in a corresponding column of cells in the second plurality of cells.

5. The electronic system of claim 4, wherein the second plurality of data lines for the first column multiplexer is provided by a second column multiplexer.

6. The electronic system according to claim 1, comprising: a plurality of second calculation circuits, wherein each second calculation circuit is operable to calculate a calculation value based on a cell value in a corresponding cell column in a second plurality of cells of the memory array; a second column multiplexer that, in operation, cycles through a plurality of data lines, each data line of the plurality of data lines corresponding to a second computation circuit of the plurality of second computation circuits; as well as The second sensing circuit is configured to obtain the calculated values ​​from the plurality of second calculation circuits via the second column multiplexer as the second column multiplexer cycles through the plurality of data lines, and to combine the calculated values ​​obtained within the determined number of multiplexer cycles.

7. The electronic system according to claim 1, wherein: The cluster cycle management circuit generates a plurality of control signals in response to a clock signal during operation, and provides the plurality of control signals to the first sensing circuit and the first column multiplexer to cycle through the plurality of data lines for the determined number of multiplexer cycles, thereby causing the first sensing circuit to obtain the calculated value from the corresponding first calculation circuit.

8. The electronic system according to claim 1, comprising: A data line selection circuit is operable to select different cyclic sequences of the plurality of data lines passing through the first column multiplexer to write a continuous plurality of data to the first plurality of cells.

9. A method for operating an electronic system, comprising: storing data in a plurality of memory cells arranged as a plurality of cell rows intersecting a plurality of cell columns; calculating a plurality of calculated values ​​based on the cell values ​​from the plurality of cell columns, wherein each respective calculated value is calculated based on the cell values ​​from a respective cell column; determining a number of columns in the plurality of cell columns that store data for a data computing cluster; selecting a number of multiplexer loops based on the determined number of columns; generating a result for the data computation cluster by cycling through a first subset of the plurality of computation values ​​for a selected number of multiplexer cycles using a column multiplexer, and by combining corresponding computation values ​​from the first subset of computation values ​​using a sense engine; as well as The result of the data calculation cluster is output, wherein the method includes modifying an address of each of a plurality of data lines of the column multiplexer to write a continuous plurality of data to the plurality of memory cells.

10. The method according to claim 9, comprising: generating a second result for the data computation cluster by cycling through a second subset of the plurality of computation values ​​for a selected number of multiplexer cycles using a second column multiplexer, and by combining corresponding computation values ​​from the second subset of computation values ​​using a second sensing engine; as well as The second result of the data computing cluster is output.

11. The method according to claim 9, comprising: generating a second result for the data computation cluster by cycling through a second subset of the plurality of computation values ​​for a selected number of multiplexer cycles using the column multiplexer and combining respective computation values ​​from the second subset of computation values ​​using the sense engine; as well as The second result of the data computing cluster is output.

12. The method according to claim 9, comprising: The number of data lines utilized by the column multiplexer is modified based on the selected number of multiplexer cycles for the data computation cluster.

13. The method according to claim 9, comprising: responsive to a non-memory clock signal, enabling said column multiplexer for cycling through said first subset of calculated values ​​for a selected number of multiplexer cycles; as well as In response to a memory clock signal, each cycle of each data line of the column multiplexer is enabled for a selected number of multiplexer cycles to obtain the first subset of calculated values.

14. The method according to claim 9, comprising: Different cycling orders are selected for the column multiplexer to cycle through the plurality of columns of cells, thereby writing consecutive pluralities of data to the plurality of cells.

15. A computing device comprising: means for storing data in a plurality of cells arranged as a plurality of cell rows intersecting a plurality of cell columns; means for calculating a respective calculated value based on the cell values ​​from each respective column of the plurality of cell columns; means for determining a compute cluster loop size based on a number of columns in the plurality of cell columns storing data for the data compute cluster; means for looping through corresponding computation values ​​for the determined computation cluster loop size; means for combining the respective compute values ​​to generate a result for the data compute cluster for the determined compute cluster cycle size; as well as means for modifying an order in which data is stored in the plurality of cell columns to write a continuous plurality of data to the plurality of cells, wherein the means for modifying modifies an address of each of a plurality of data lines of the means for looping through corresponding calculation values ​​for the determined calculation cluster loop size to write a continuous plurality of data to the plurality of cells.

16. The computing device of claim 15, comprising: means for initiating a loop through the corresponding compute values ​​for the determined compute cluster loop size; as well as Means for initiating each loop of each respective compute value for the determined compute cluster loop size.

17. A non-transitory computer-readable medium having content that causes a cluster cycle management circuit to perform actions comprising: storing data in a plurality of memory cells arranged as a plurality of cell rows intersecting a plurality of cell columns; determining a number of columns in the plurality of cell columns that store data for a computing operation; selecting a number of multiplexer loops based on the determined number of columns; generating a result of the computation operation by cycling through a first subset of the plurality of cell columns using a column multiplexer for a selected number of multiplexer cycles and by combining values ​​from the first subset of cell columns using a sense engine; as well as The result of the calculation operation is output, wherein the action includes modifying an address of each of a plurality of data lines of the column multiplexer to write a continuous plurality of data to the plurality of memory cells.

Citation Information

Patent Citations

  • Apparatuses and methods for simultaneous in data path compute operations

    US20180364908A1