Redundant computation across planes
By performing redundant computational operations in parallel on multiple planes and utilizing content-addressable memory cells (CAMs) for parallel computation, the problems of long waiting time and bandwidth limitations in arithmetic operations in the prior art are solved, and more efficient processing performance is achieved.
Patent Information
- Application Number
- CN202211732917.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-02-23
- Filing Date
- 2022-12-30
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-12-30
AI Technical Summary
Existing technologies suffer from long waiting times and bandwidth limitations when performing arithmetic operations, especially when performing vector calculations, where serial processing leads to low computational efficiency.
Associative processing technology is employed to perform redundant computational operations in parallel on multiple planes. This involves using content-addressable memory cells (CAMs) for parallel computation, storing the same data on redundant planes to reduce waiting time, and performing computational operations in parallel in memory through associative processing technology.
It reduces the waiting time for arithmetic operations, increases processing bandwidth, avoids interface bottlenecks between the host device and the accelerator, and reduces power consumption.
Smart Images

Figure CN116382888B_ABST
Abstract
Description
[0001] Cross-reference
[0002] This patent application claims priority to U.S. Patent Application No. 17 / 652,229, entitled “Redundant Computing Across Planes,” filed February 23, 2022, by EILERT et al., and U.S. Provisional Patent Application No. 63 / 266,216, entitled “Redundant Computing Across Planes,” filed December 30, 2021, by EILERT et al., each of which is assigned to the assignee of this patent application, and each of which is expressly incorporated herein by reference in its entirety. Technical Field
[0003] The following generally pertains to one or more systems for memory, and more specifically, to redundant computation across planes. Background Art
[0004] Memory devices are widely used to store information in various electronic devices such as computers, user devices, wireless communication devices, cameras, and digital displays. Information is stored by programming memory cells within the memory device into various states. For example, a binary memory cell can be programmed into one of two supported states, typically represented by logic 1 or logic 0. In some instances, a single memory cell can support more than two states and can store any of them. To access the stored information, a component can read or sense at least one stored state in the memory device. To store information, a component can write or program states into the memory device.
[0005] Various types of memory devices and memory cells exist, including magnetic hard disks, random access memory (RAM), read-only memory (ROM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), static RAM (SRAM), ferroelectric RAM (FeRAM), magnetic RAM (MRAM), resistive RAM (RRAM), flash memory, phase-change memory (PCM), auto-select memory, chalcogenide memory technology, etc. Memory cells can be volatile or non-volatile. Even without an external power supply, non-volatile memory, such as FeRAM, can maintain its stored logic state for an extended period of time. Volatile memory devices (e.g., DRAM) can lose their stored state when disconnected from an external power source. Summary of the Invention
[0006] Describe an apparatus. The apparatus may include: a memory die including a plurality of planes arranged as a plurality of data blocks, the plurality of planes including content-addressable memory cells; and logic coupled to the memory die and configured to: perform a computational operation on first data stored in a first plane of the plurality of planes, wherein the computational operation is at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a set of consecutive bits of a vector; concurrently perform a computational operation on second data stored in a second plane of the plurality of planes, wherein the second data represents the set of consecutive bits of a vector; and read third data from the first plane and write it to the second plane, the third data representing the result of the computational operation on the first data.
[0007] Describe an apparatus. The apparatus may include: a memory die including a plurality of planes arranged as a plurality of data blocks, the plurality of planes including content-addressable memory cells; and logic coupled to the memory die and configured to: perform a computation operation on first data stored in a first plane, wherein the computation operation is at least partially based on the capability of the content-addressable memory cells, and wherein the first data represents a first set of consecutive bits of a vector; perform a computation operation on second data stored in a second plane, at least partially based on a first value from the output bits of the computation operation on the first data, wherein the second data represents a second set of consecutive bits of a vector; and perform a computation operation on third data stored in a third plane, at least partially based on a second value from the output bits of the computation operation on the first data, wherein the third data represents a second set of consecutive bits of a vector.
[0008] This invention describes a method. The method may include performing a computational operation on first data stored in a first plane of a plurality of planes including content-addressable memory cells, wherein the computational operation is at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a set of consecutive bits of a vector; concurrently performing a computational operation on second data stored in a second plane, wherein the second data represents the set of consecutive bits of a vector; and reading third data from the first plane and writing it into the second plane, the third data representing the result of the computational operation on the first data.
[0009] This invention describes a method. The method may include: performing a computation operation on first data stored in a first plane of a plurality of planes including content-addressable memory cells, wherein the computation operation is at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a first set of consecutive bits of a vector; performing a computation operation on second data stored in a second plane, at least partially based on a first value from the output bits of the computation operation on the first data, wherein the second data represents a second set of consecutive bits of a vector; and performing a computation operation on third data stored in a third plane, at least partially based on a second value from the output bits of the computation operation on the first data, wherein the third data represents a second set of consecutive bits of a vector.
[0010] This invention describes a method. The method may include: performing a computational operation on first data stored in a first plane comprising a plurality of planes including content-addressable memory cells, wherein the computational operation is at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a first set of consecutive bits of a vector; performing a computational operation concurrently with the computational operation on the first data on second data stored in a second plane and representing a second set of consecutive bits more significant than the first set of consecutive bits; wherein the computational operation on the second data is at least partially based on a first value from the output bits of the computational operation on the first data; performing a computational operation concurrently with the computational operation on the first data on third data stored in a third plane and representing a second set of consecutive bits of a vector, wherein the computational operation on the third data is at least partially based on a second value from the output bits of the computational operation on the first data; and reading fourth data from the second plane and writing it back to the first plane, the fourth data representing the result of the computational operation on the second data, wherein the fourth data is copied at least partially based on the output bits of the computational operation on the first data having the first value. Attached Figure Description
[0011] Figure 1 This section describes examples of systems that support redundant computation across planes, based on examples disclosed herein.
[0012] Figure 2 This describes an example of vector computation that supports redundant computation across planes, based on examples disclosed herein.
[0013] Figure 3 This describes an instance of a plane that supports redundant computation, as illustrated in the examples disclosed herein.
[0014] Figure 4 This describes an instance of a plane that supports redundant computation, as illustrated in the examples disclosed herein.
[0015] Figure 5This describes an instance of a plane that supports redundant computation, as illustrated in the examples disclosed herein.
[0016] Figure 6 This document describes an example of a processing flow that supports redundant computation across planes, based on examples disclosed herein.
[0017] Figure 7 A block diagram illustrating an apparatus that supports cross-plane redundant computation based on examples disclosed herein.
[0018] Figures 8 to 10 A flowchart illustrating one or more methods that support cross-plane redundancy computation, as illustrated in the examples disclosed herein. Detailed Implementation
[0019] In some systems, the host device can offload various processing tasks to electronic devices, such as accelerators. For example, the host device can offload computations (e.g., vector or scalar computations) to electronic devices that can perform the computations using a computing engine and processing techniques. Such offloading of computations may involve communicating operands or operand information from the host device to the electronic device, and further communicating results from the electronic device to the host device. Therefore, the bandwidth of the electronic device can be constrained by the communication interface between the electronic device and the host device, as well as the size of the computing engine and the serial processing. According to the techniques described herein, the host device can substantially increase processing bandwidth by offloading processing tasks to an associative processor memory (APM) system, which, among other things, uses associative processing in memory to perform data-parallel computations.
[0020] For example, some systems can use associative processing to perform arithmetic operations on operands of arithmetic operations (e.g., the system can produce results from one or more vector or scalar operands, whether present or absent in the system). Such systems can perform arithmetic operations on a bit-by-bit serial basis, such that arithmetic output bits based on the less significant bits (e.g., carry, borrow) can be used to perform arithmetic operations on the more significant bits. However, performing arithmetic operations on a serial basis increases the latency of the arithmetic operation, as well as other disadvantages. In other words, a subset of arithmetic operations, for example, can essentially be bit-by-bit serial in associative processing because it is based on a search-update sequence, which is based on carry / borrow generated by the search-update operation from the less significant bits. Therefore, the longer the vector element length, the higher the latency of the arithmetic operation.
[0021] According to the techniques described herein, an APM system can reduce the latency of computational operations, such as arithmetic operations, by performing redundant computational operations on vector operands in parallel. For example, the APM system can use a first set of planes to perform computational operations based on (e.g., assumed) a first value (e.g., 0) for each arithmetic output bit (e.g., carry, borrow). In parallel, the APM system can use a second set of planes to perform computational operations based on (e.g., assumed) a second value (e.g., 1) for each arithmetic output bit (e.g., carry, borrow). The APM system can then replace erroneous results from the first set of planes with correct results from the second set of planes, such that all results from the first set of planes are correct. Alternatively, the APM system can reconstruct correct results by marking correct bits in each plane based on carry / borrow calculated from less significant bits. Therefore, reconstruction may or may not involve data movement (e.g., reconstruction can be done by tracking the position of correct results across planes). By performing redundant computations as described herein, the APM system can reduce the latency of arithmetic (e.g., bit-sequencing) operations.
[0022] First, as referenced Figure 1 and 2 The features of this disclosure are described in the context of the system and vector computation described herein. In the plane and as referenced... Figures 3 to 6 The features of this disclosure are described in the context of the described processing flow. References to [references to other references are also provided]. Figures 7 to 10 The device diagrams and flowcharts related to the described cross-plane redundancy calculations further illustrate and describe these and other features of this disclosure.
[0023] Figure 1 This describes an example of a system 100 that supports redundant computing across planes, as disclosed herein. System 100 may include a host device 105 and an Associative Processing Memory (APM) system 110. Host device 105 may interact with APM system 110 (e.g., communicate, control) and other components of the device containing APM system 110. In some instances, host device 105 and APM system 110 may interact via interface 115, which may be an instance of a Compute Quick Link (CXL) interface or other types of interfaces.
[0024] In some instances, system 100 may be contained in or coupled to a computing device, electronic device, mobile computing device, or wireless device. The device may be a portable electronic device. For example, the device may be a computer, laptop computer, tablet computer, smartphone, cellular phone, wearable device, Internet-connected device, etc. Host device 105 may be or contain a system-on-a-chip (SoC), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or combinations thereof. In some instances, host device 105 may be referred to as a host, host system, or other suitable terms.
[0025] The APM system 110 can operate as an accelerator (e.g., a high-speed processor) for the host device 105, allowing the host device 105 to offload various processing tasks to the APM system 110, which can be configured to execute processing tasks faster than the host device 105. For example, the device 105 can send a program (e.g., a set of instructions, such as Reduced Instruction Set V (RISC-V) vector instructions) to the APM system 110 for execution by the APM system 110. As part of the program, or as instructed by the program, the APM system 110 can perform various computational operations on vectors (e.g., the APM system 110 can perform vector computations). Computational operations can refer to logical operations, arithmetic operations, or other types of operations involving vector manipulation. A vector can contain one or more elements, also referred to as vector elements, each element having a corresponding number of bits. The length or size of a vector can refer to the number of elements in the vector, and the length or size of an element can refer to the number of bits in the element.
[0026] APM controller 120 may be configured to interface with host device 105 on behalf of APM device 125. Upon receiving a program from host device 105, APM controller 120 may parse the program and instruct or otherwise prompt APM device 125 to perform various computational operations associated with or instructed by the program. In some instances, APM controller 120 may retrieve (e.g., from memory 130) vectors of computational operations and may pass these vectors to APM device 125 for associated processing. In some instances, APM controller 120 may indicate vectors of computational operations to APM device 125 so that APM device 125 can retrieve vectors from memory 130. In some instances, host device 105 may provide vectors to APM system 110. Therefore, memory 130 may be configured to store vectors accessible by APM controller 120, APM device 125, host device 105, or a combination thereof.
[0027] The vectors of computational operations at APM device 125 may be indicated (or accompanied) by a program received from host device 105 or other control signaling associated with the program (e.g., other separate control signaling). For example, a program indicating computational operations on a pair of vectors may include one or more addresses (or one or more pointers to one or more addresses) of memory 130 in which the vectors are stored. Although shown as included in APM system 110, memory 130 may be external to APM system 110, but is still coupled thereto. Although shown as a single component, the functionality of memory 130 may be provided by multiple memories 130.
[0028] APM device 125 may include memory cells, such as addressable memory cells (CAMs) configured to store the contents of vectors (e.g., vector operands, vector results) associated with computational operations. Vector operands may be vectors of operands of computational operations (e.g., vector operands may be vectors to which computational operations are performed). Vector results may be vectors generated by vector computation.
[0029] The APM system 110 can be configured to store information, such as truth tables, for various computational operations, wherein information about a given computational operation (e.g., a truth table) can indicate the result of the computational operation for various combinations of logical values. For example, among other types of operations, the APM system 110 can also store information (e.g., one or more truth tables) about logical operations (e.g., AND, OR, XOR, NOT, NAND, NOR, XNOR) and arithmetic operations (e.g., addition, subtraction). Memory cells storing information about computational operations (e.g., one or more truth tables) can store various combinations of logical values of the operands of the computational operation, as well as the corresponding result and carry (if applicable) for each combination of logical values. The APM system 110 can store the truth table for associative processing in one or more memories (e.g., in one or more die-on mask ROMs) that can be coupled to or contained therein. For example, the truth table may be stored in memory 130, in local memory of APM device 125, or both. In either instance, APM device 125 may cache shared instructions on-device (e.g., instead of fetching or receiving the instructions).
[0030] At least some of the APM devices 125 (if not every APM device 125) can use associative processing to perform computational operations on vectors stored in the APM device 125. Unlike serial processing (where vectors move back and forth between the processor and memory), associative processing may involve searching and writing vectors in memory (also referred to as "in-situ"), which allows for increased parallelism in processing bandwidth. Among other advantages, the execution of in-situ computational operations also allows system 100 to avoid bottlenecks at the interface between host device 105 and APM system 110, which reduces latency and power consumption compared to other processing techniques (such as serial processing). Associative processing may also be referred to as associative computation or other suitable terms.
[0031] In some instances, the APM device 125, which uses associative processing to perform computational operations, can utilize information (e.g., a truth table) to perform computational operations bit-by-bit using a "search and write" technique. For example, if the APM device 125 includes a CAM unit that stores vector operands for computational operations, the APM device 125 can search in the CAM unit for bits of the vector operands that match entries in the truth table corresponding to the computational operation, determine the result of the bit computational operation based on the matching entries in the truth table, and write the result back to content-addressable memory. The APM device 125 can then continue with the next significant bit of the vector and use associative processing to perform computational operations on those bits. In some instances, the bit computational operation may involve arithmetic output bits (e.g., carry, borrow) bits that have been determined to be part of a computational operation on the less significant bit.
[0032] Each APM device 125 may include one or more dies 135, which may also be referred to as a memory die, semiconductor die, or other suitable terms. Die 135 may include multiple data blocks 140, each of which may further include multiple planes 145. In some instances, data blocks 140 may be configured such that a single plane 145 is operable or active per data block at a time (e.g., one plane per data block can perform associative computations at a time). However, any number of data blocks 140 may be active at any given time (e.g., any number of data blocks may be performing associative computations at any given time). Therefore, data blocks 140 can operate in parallel, which can increase the number of computational operations performed during time intervals, thereby increasing the bandwidth of APM device 125 relative to other technologies. Using multiple APM devices 125, as relative to a single APM device 125, can further increase the bandwidth of APM system 110 relative to other systems. Each APM device 125 may include a local controller or logic for controlling the operation of the APM device 125.
[0033] Each plane 145 may contain a memory array of memory cells containing memory units (e.g., CAM cells). The memory cells in the memory array may be arranged in columns and rows, and may be non-volatile or volatile memory cells. The memory array containing CAM cells may be configured to search for CAM cells by content, as opposed to by address. For example, a memory array containing CAM cells storing vectors of computational operations may compare the logical values of the operands of the vector with entries in a truth table associated with the computational operations to determine which results correspond to those logical values.
[0034] As described, the APM device 125 can be configured to store vectors associated with computational operations in memory cells of the APM device 125. To facilitate associative processing, vectors can be stored column-wise across multiple planes. For example, given multiple n-bit (e.g., n = 32) elements (denoted as E0 to E... N For a vector v0, the APM device 125 can divide each element into multiple groups of consecutive bits (e.g., four groups of eight consecutive bits). The APM device 125 can store the first group of consecutive bits (e.g., the least significant group of consecutive bits) of each element of vector v0 in a first plane 145, where each row of plane 145 stores the first group of consecutive bits of the corresponding element of vector v0. Thus, in some instances, column 150 can store the first 8 bits of each element of vector v0 (e.g., column 150 can span 8 columns). In a similar manner, the APM device 125 can store the next significant group of consecutive bits from each element of vector v0 in a second plane 145. The remaining groups of consecutive bits of vector v0 are stored in the same way. Thus, vector v0 can be stored column-wise across multiple planes. Bits of other vectors v1 to vn can be stored in a similar column-wise manner across plane 145.
[0035] Using column-oriented storage to distribute vectors across multiple planes allows the APM device 125 to store multiple vectors per plane 145 compared to other techniques, which in turn allows the APM device 125 to operate on combinations of multiple vectors. For example, consider a plane of 256 rows × 256 columns. The APM device 125 can store 32 vectors with 32-bit elements across four planes, which allows the APM device 125 to operate on those 32-bit vectors (e.g., one plane at a time) without performing time-consuming vector moves, instead of storing eight vectors with 32-bit elements across a single plane, which limits the APM device 125 to operating on those eight vectors (where there is no time-consuming vector movement).
[0036] In some instances, the APM device 125 may store vectors according to a vector mapping scheme, which may be one of several vector mapping schemes supported by the APM device 125. A vector mapping scheme may refer to a scheme in which vectors are mapped (and written) to plane 145 of the APM device 125. For example, the APM device 125 may support a first vector mapping scheme (referred to as vector mapping scheme 1) and a second vector mapping scheme (referred to as vector mapping scheme 2). In vector mapping scheme 1, vectors may be distributed across the plane of the same data block 140. In vector mapping scheme 2, vectors may be distributed across the plane of different data blocks 140. A vector mapping scheme may also be referred to as a storage scheme, a layout scheme, or other suitable terms.
[0037] APM system 110 may select between vector mapping schemes before writing vectors to APM device 125 according to a selected vector mapping scheme. For example, APM system 110 may select a vector mapping scheme for a group of computational operations based on the size of the vector associated with a set of computational operations, the type of computational operations in the group (e.g., arithmetic and logical), the number of computational operations in the group, or a combination thereof, and other aspects. In some instances, APM system 110 may select a vector mapping scheme in response to an instruction provided by host device 105. For example, host device 105 may indicate a vector mapping scheme associated with the instruction set of the group of computational operations. After the vectors have been written to APM device 125 according to the selected vector mapping scheme, APM device 125 may use associative processing to perform computational operations on the vectors according to the selected vector mapping scheme. Alternatively, a compiler or preprocessor may determine the vector mapping scheme.
[0038] The associative processing techniques described herein can be implemented via logic at APM system 110, logic at APM device 125, or logic distributed between APM system 110 and APM device 125. The logic may include one or more controllers, access circuitry systems, communication circuitry systems, or combinations thereof, as well as other components and circuitry. The logic may be configured to perform aspects of the techniques described herein, causing components of APM system 110 and / or APM device 125 to perform aspects of the techniques described herein, or both.
[0039] In some instances (e.g., if the length of the vector elements is greater than the number of columns 150), the vector can be distributed across multiple planes 145 of the APM device 125. In such instances, the APM device 125 can perform computational operations (e.g., arithmetic operations) on the vector on a plane-by-plane basis, such that the arithmetic output bits propagate across the planes. However, performing computational operations on a plane-by-plane basis can increase system latency. According to the techniques described herein, the APM device 125 can reduce system latency by using redundant planes (e.g., planes storing duplicate data representing the same vector) and performing computational operations in parallel across the redundant planes based on different values of the arithmetic output bits (e.g., carry, borrow).
[0040] Figure 2 This describes an example of a vector computation 200 supporting redundant computation across planes, as disclosed herein. The vector computation 200 may be an example of vector addition and may be performed on operand vectors vA and vB, which may be stored in a memory unit (e.g., a CAM unit) of the plane of the APM device. The result of the vector addition may be a vector vD. Each operand vector may contain four bits (e.g., an operand vector may contain a single 4-bit element), and the position of each bit may be represented as i. The operand vectors may be stored in the plane of the APM device, as referenced... Figure 1 The above discussion can be associated with a set of vector instructions (e.g., RISC-V vector instructions). Vector computation 200 can be performed using truth table 205, which can be a truth table for adding two bits and possible carry. Truth table 205 can be stored in memory coupled to or contained in the APM device, and CAM technology can be used to compare entries (e.g., rows) of truth table 205 with the operands of vectors vA and vB.
[0041] The examples of associative processing using computational operations on vectors provided are for illustrative purposes only and are not intended to limit the scope in any way.
[0042] To perform the addition of vectors vA and vB using associative processing, the APM device may retrieve entries from memory (e.g., using a sequencer) from truth table 205 and compare the entries with the operands of vectors vA and vB (e.g., using in-situ CAM techniques). Upon finding a match, the APM device may return the corresponding results of the matching entry (e.g., vDi and carry c) before moving to the next valid operand of the vectors. i+1 Write the vector to be stored to the plane (or a different plane).
[0043] For example, for i = 0, the APM device can compare the entry in truth table 205 with the corresponding operand bits from vectors vA and vB (e.g., c0 = 0, vA0 = 1, and vB0 = 0). Upon detecting a match between an operand bit and an entry in truth table 205, the APM device can write the result corresponding to the matching entry (e.g., vD0 = 0 and carry c1 = 1) to the plane storing the operand vectors. Alternatively, the device can serially compare the entry from truth table 205 with the operand bit i = 0 (e.g., starting from the top entry and moving down one entry along truth table 205). In some instances, the APM device can compare the entry from truth table 205 with multiple operand bits in parallel (e.g., concurrently).
[0044] After determining the result of the i-th operand, the APM device can proceed to the next valid operand (which may contain a carry i+1 determined from the i-th operand). For example, after determining the result of the i=0 operand, the APM device can proceed to the i=1 operand (which may contain a carry c1 determined from the i=0 operand). However, in some scenarios (e.g., when the computation operation is a logical operation), the APM device can perform the computation operation in parallel on some or all of the operands.
[0045] For i=1, the APM device can compare the entry in truth table 205 with the corresponding operand bits from vectors vA and vB (e.g., c1=1, vA1=0, and vB1=0). Upon detecting a match between an operand bit and an entry in truth table 205, the APM device can write the result corresponding to the matching entry (e.g., vD1=1 and carry c2=0) to the plane (or a different plane) storing the operand vectors. The APM device can compare the entries from truth table 205 with the operand bit i=1 serially (e.g., starting from the top entry and moving down one entry along truth table 205). After determining the result for operand bit i=1, the APM device can continue to operand bit i=2 (which may contain the carry c2 determined based on operand bit i=1).
[0046] For i=2, the APM device can compare the entry in truth table 205 with the corresponding operand bits from vectors vA and vB (e.g., c2=0, vA2=0, and vB2=0). Upon detecting a match between an operand bit and an entry in truth table 205, the APM device can write the result corresponding to the matching entry (e.g., vD2=0 and carry c3=0) to the plane (or a different plane) storing the operand vectors. The APM device can compare the entry from truth table 205 with the operand bit i=2 serially (e.g., starting from the top entry and moving down one entry along truth table 205). After determining the result for operand bit i=2, the APM device can continue to operand bit i=3 (which may contain the carry c3 determined based on operand bit i=2).
[0047] For i=3, the APM device can compare the entry in truth table 205 with the corresponding operand bits from vectors vA and vB (e.g., c3=0, vA3=0, and vB3=1). When a match is detected between an operand bit and an entry in truth table 205, the APM device can write the result corresponding to the matching entry (e.g., vD3=1 and carry c4=0) to the plane (or a different plane) storing the operand vectors. The APM device can compare the entry from truth table 205 with the operand bit i=3 serially (e.g., starting from the top entry and moving down one entry along truth table 205).
[0048] Therefore, the APM device can use correlation processing to determine that adding vA (e.g., 0b0001) and vB (e.g., 0b1001) produces vD = 0b1010. After completing the additional operations, the APM device can communicate the vector vD to the host device and use the resulting vector vD to perform other computational operations, or combinations thereof.
[0049] Although APM devices can perform computations on a serial, bit-by-bit basis, the latency can be reduced if the APM device performs computations on different groups of bits in parallel. For example, if vector vA has a vector element length of 16 bits, the APM device can divide each vector into four groups of consecutive bits and perform computations on each group of consecutive bits in parallel (but within a group, computations can be performed on a serial, bit-by-bit basis, as shown in the reference). Figure 2(as described above). For example, an APM device can perform computational operations in parallel on groups A (bits 0 to 3), group B (bits 4 to 7), group C (bits 8 to 11), and group D (bits 12 to 15), but within each group, the APM device can perform computational operations on a bit-by-bit basis. To account for arithmetic output bits, the APM device can redundantly perform each computational operation using different values of arithmetic bits (e.g., carry, borrow). The APM device can then select the correct result from the redundant computational operations based on the actual values of the arithmetic bits (e.g., carry, borrow) calculated by the plane storing the lower significant bits.
[0050] Figure 3 This describes an example of a plane 300 that supports redundant computation, as disclosed in the examples herein. Plane 300 may be as shown in the references... Figure 1 An example of plane 145 is described. Therefore, plane 300 can be configured to store vectors for computational operations performed using associative processing. In some instances, plane 300 may be in the same data block, as discussed in Reference Vector Mapping Scheme 1. In other instances, plane 300 may be located in different data blocks, as discussed in Reference Vector Mapping Scheme 2.
[0051] In a given instance, n vectors with multiple (e.g., 256) multi-bit elements (e.g., 32-bit elements) are mapped to four planes. However, other numbers of these factors are taken into account and are within the scope of this disclosure.
[0052] The APM device can map n vectors (represented as v0 to v). n-1 This is then written to four planes. The number of planes to which a vector is mapped can be a function of the element length and the number of bits mapped to each plane. For example, the number of planes to which a vector is mapped might be equal to the element length divided by the number of bits mapped to each plane. In a given instance, the number of planes to which a vector is mapped is four, which is equal to the element length (e.g., 32) divided by the number of bits mapped to each plane (e.g., eight).
[0053] At least some (if not every) planes may store a set of consecutive bits from at least some (if not every) elements of at least some (if not every) vectors (e.g., each plane may store a corresponding set of consecutive bits from each element of each vector). For example, plane 0 may store consecutive bits 0 to 7 for each element of each vector; plane 1 may store consecutive bits 8 to 15 for each element of each vector; plane 2 may store consecutive bits 16 to 23 for each element of each vector; and plane 3 may store consecutive bits 24 to 31 for each element of each vector. Bits of different vectors may be stored across different columns of the plane, while bits of different elements may be stored across different rows of the plane. For example, bits from vector 0 may be stored in a first set of 8 columns of each plane; bits from vector 1 may be stored in a second set of 8 columns of each plane; bits from vector 2 may be stored in a third set of 8 columns of each plane; and so on. For each vector, bits from element 0 may be stored in the first row of a given plane; bits from element 1 may be stored in the second row of the plane; bits from element 2 may be stored in the third row of the plane, and so on.
[0054] Therefore, a plane with x rows (e.g., 256 rows) can store vectors with x elements or fewer (vectors of length 256 or less). If a vector has more than x elements, the elements of the vector can be split across multiple planes (e.g., the elements of a vector of length 512 can be stored in two planes, where the first plane stores bits from the first 256 elements and the second plane stores bits from the last 256 elements). Thus, systems using the vector mapping scheme described herein can support vectors larger in size than other systems (e.g., serial processing systems), which may be limited by the size of the processing circuitry system (e.g., a computing engine).
[0055] Vectors can be stored according to either vector mapping scheme 1 or vector mapping scheme 2. In vector mapping scheme 1, the planes to which vectors are mapped can be located in the same data block. For example, planes 0 to 3 can be in data block A. In vector mapping scheme 2, the planes to which vectors are mapped can be located in different data blocks. For example, plane 0 can be in data block A, plane 1 can be in data block B, plane 2 can be in data block C, and plane 3 can be in data block D. Commonly, data blocks A to D (e.g., data blocks spanning their distributed vectors) can be referred to as hyperplanes. The two vector mapping schemes allow the APM device to perform computational operations on multiple vectors in parallel (e.g., during partial or full overlap time). For example, given h data blocks, the APM device can immediately perform h different computational operations.
[0056] Therefore, in vector mapping scheme 1, the APM device can use a single data block to perform calculations on the vector. For example, the APM device can use data block A to perform calculations on bits 0 to 7 of the elements in the vector, use data block A to perform calculations on bits 8 to 15 of the elements in the vector, use data block A to perform calculations on bits 16 to 23 of the elements in the vector, and use data block A to perform calculations on bits 24 to 31 of the elements in the vector. If a carry is caused by a calculation operation, the APM device can transfer the carry (denoted as "C") between planes of data block A. For example, if the carry is caused by a calculation operation on bits 0 to 7, the APM device can transfer the carry from plane 0 to plane 1 in data block A.
[0057] In vector mapping scheme 2, the APM device can use multiple data blocks to perform calculations on the vector. For example, the APM device can use data block A to perform calculations on bits 0 to 7 of the elements in the vector, data block B to perform calculations on bits 8 to 15, data block C to perform calculations on bits 16 to 23, and data block D to perform calculations on bits 24 to 31. If a carry is caused by a calculation operation, the APM device can transfer the carry between data blocks. For example, if the carry originates from a calculation operation on bits 0 to 7, the APM device can transfer the carry from data block A to data block B.
[0058] The associative processing techniques described herein can be implemented through logic at the APM system, through logic at the APM device, or through logic distributed between the APM system and the APM device. The logic may include one or more controllers, access circuitry systems, communication circuitry systems, or combinations thereof, as well as other components and circuitry. The logic may be configured to perform aspects of the techniques described herein, causing components of the APM system and / or the APM device to perform aspects of the techniques described herein, or both.
[0059] An APM device can perform computational operations serially or in parallel. If the APM device performs computational operations serially, it can perform the computational operations on one plane at a time, sequentially (e.g., starting with the least significant plane (e.g., plane 0) and ending with the most significant plane (e.g., plane 3)). The APM device can perform computational operations on one plane at a time because the computational operation on plane n may depend on the arithmetic output bits produced by the computational operation on plane n-1. However, in some instances, performing computational operations on one plane at a time can increase latency and other disadvantages.
[0060] According to the techniques described herein, an APM device can reduce latency by performing computational operations in parallel across planes. To this end, in some instances, the APM device can use corresponding redundant planes for plane 1, plane 2, and plane 3. The redundant plane can store bits for the same computational operations as planes 1, 2, and 3. The APM device can use a first possible value (e.g., 0) for the arithmetic output bits for planes 1, 2, and 3, and a second possible value (e.g., 1) for the redundant plane. By using different values for the arithmetic output bits, the APM device can perform computational operations on all planes (e.g., planes 0 through 3, and the redundant plane) without waiting for computational operations on one or more other planes (e.g., the preceding plane) to complete. After performing the computational operation, the APM device can determine the actual (e.g., calculated) value of the arithmetic output bits and then select the result of the computational operation from the plane using the correct possible value of the arithmetic output bits.
[0061] Figure 4 This describes an instance of plane 400 supporting redundant computation, as disclosed in the examples herein. Plane 400 may contain planes P0 through P6. Plane P4 may be redundant with plane P0, meaning that plane P4 stores operand vector elements of the same computational operation as plane P0. Similarly, P5 may be redundant with plane P1 (such that plane P5 stores operand vector elements of the same computational operation as plane P2), and plane P6 may be redundant with plane P3 (such that plane P6 stores operand vector elements of the same computational operation as plane P3). Planes storing the same operand instance elements of computational operations may be referred to as sister planes or redundant planes, and in... Figure 4 The image is displayed with matching shadows. Planes P1, P2, and P3 can form a first channel (channel 405-a) of the plane, and planes P4, P5, and P6 can form a second channel (channel 405-b) of the plane. Planes P0 to P6 can be located in the same data block (e.g., according to layout 1) or in different data blocks (e.g., according to layout 2). For example, each plane can be located in data block A, or the planes can be distributed across data blocks so that each plane is located in a corresponding data block.
[0062] Each plane can store multiple sets of consecutive bits for a vector element. For example, plane P0 can store consecutive bits 0 to 7 for each element of vectors v0 to v31. Planes P1 and P4 can each store consecutive bits 8 to 15 for each element of vectors v0 to v31. Planes P2 and P5 can each store consecutive bits 16 to 23 for each element of vectors v0 to v31. And planes P3 and P6 can each store consecutive bits 24 to 31 for each element of vectors v0 to v31. Although 32 vectors are shown, each with 256 elements and 8 bits per element, other numbers of vectors, elements, and bits are considered and are within the scope of this disclosure.
[0063] An APM device comprising planes P0 through P6 can use redundant computation to reduce the latency of computational operations. For example, an APM device can use redundant computation to reduce the latency of computational operations (e.g., addition) on operand vectors v0 and v1. For ease of illustration, the computational operation is described with reference to a single element of vector v0. However, the techniques described herein can be extended to multiple elements of vectors v0 and v1, including all elements of vectors v0 and v1. Although described with reference to two operand vectors (v0 and v1), the techniques described herein can be implemented for any number of operand vectors.
[0064] When performing redundant calculations, the APM device can use a first value (e.g., 0) of the speculative carry that serves as the input bits for planes P0, P1, and P2. The APM device can use a second value (e.g., 1) of the speculative carry that serves as the input bits for planes P4, P5, and P6. The speculative carry of a plane can represent the actual carry from a lower valid plane in channel 405 of the plane, and a possible value can be assigned to the actual carry. For example, speculative carry c8 Spec It can represent actual carry (from bit 0 to 7), and speculative carry c16 Spec c24 can represent both actual carry (from digit 8 to 15) and speculative carry. Spec This can represent an actual carry (from bits 16 to 23). An actual carry for a plane can refer to a carry determined based on bits in a previous (e.g., less efficient) plane, whereas a speculative carry is set to one of two possible values without considering bits in the aforementioned plane. Actual carry c8 Act c16 Act and C24 Act These may be referred to as output bits or arithmetic output bits. Although the carry description is referenced, the APM device may use redundant calculations as described herein for other types of arithmetic output bits.
[0065] By using speculative carry and redundant planes, the APM device can perform metering operations on each plane in parallel (e.g., concurrently, with full or partial overlap in time). Specifically, the APM device can use the actual carry c0 (denoted as c0). ACT The APM device performs computational operations on bits 0 to 7 of the element n (denoted as [En]) of the vector v0 in plane 0. Concurrently, the APM device can: 1) use speculative carry c8 Spec =0 (e.g., c8) Act 1) performs computation on bits 8 to 15 of element n of vector v0 in plane 1, and 2) uses speculative carry c8. Spec=1 (e.g., a second possible value of c8Act) to perform computation operations on bits 8 to 15 of element n of vector v0 in plane 4. Additionally, concurrently, the APM device can: 1) use speculative carry c16 Spec =0 (e.g., c16) Act The first possible value) is used to perform computation operations on bits 16 to 23 of the element n of vector v0 in plane 2, and 2) uses speculative carry c16. Spec =1 (e.g., c16) Act The second possible value) is used to perform computation operations on bits 15 to 23 of element n of vector v0 in plane 5. Furthermore, concurrently, the APM device can: 1) use speculative carry c24 Spec =0 (e.g., c24) Act The first possible value) is used to perform computation operations on bits 24 to 31 of the element n of vector v0 in plane 3, and 2) uses speculative carry c24. Spec =1 (e.g., c24) Act The second possible value) is used to perform computation operations on bits 24 to 31 of the element n of vector v0 in plane 6.
[0066] An APM device can use associative processing to perform computational operations. For example, an APM device can search for a bit value in a vector operand that matches an entry in a truth table for a computational operation, and then determine the result of the computational operation based on the correspondence from the truth table. Thus, an APM device can perform computational operations based on the capabilities of content-addressable memory cells used to store vector operands (e.g., search and replace capabilities).
[0067] After performing a computational operation on a plane, the APM device can store the result of the computational operation in, for example, the plane itself. For instance, the APM device can store the result of a computational operation on bits 0 through 7 in the content-addressable memory cell of vector v31 in plane P0. The same applies to other planes. In some instances, the APM device can also store the actual carry of the plane in the plane (or a local register or other storage device) for later use (e.g., for use during reconstruction).
[0068] Therefore, unlike serial computation, the APM device can perform computation on a pair of sister planes before completing computation on the less efficient pair, which reduces latency. However, each sister pair may produce incorrect results because only one of the sister planes in each pair will use a speculative carry with the correct value (e.g., only one plane will use a carry that matches the actual carry c). Act The actual carry c of the value Act (Possible values). For example, if the actual carry is c8 ActIf the value is equal to 1, then plane P4 will have the correct result for the computational operations from 8 to 15 (because plane P4 uses c8). Spec =1, this is the same as c8 Act (Matching) and plane P1 will have incorrect results (because plane P1 uses c8) Spec =0, which is the same as c8 Act (Mismatch).
[0069] Therefore, only one sister plane can store the correct results of computation operations on the elements of a vector. For illustration, consider c8. Act =1, c16 Act =0 and c24 Act =0 (for example, for element n). In this instance, the plane with the correct result for element n of vector v0 (as shown by the dashed line) is plane P4 (which uses c8). Spec =1), plane P2 (which uses c16) Spec =0), and plane P3 (which uses c) Spec 24 = 0). In other words, a sister plane with the correct result of redundant computational operations can be a plane whose possible values match (e.g., are equal to) the actual carry value.
[0070] Although described with reference to a single element n, plane 400 can perform redundant computations for each element in the operand vector. Therefore, a given sister plane may produce correct results for some vector elements but incorrect results for others (e.g., c8). Act For element j, it can be equal to 0, but for element k, it can be equal to 1; it produces a correct result for element j in plane P0, but an incorrect result for element k in plane P0. To ensure that each sister pair has a correct result for at least one plane for each element, the APM device can copy the correct result from one sister plane to another, as shown in the reference. Figure 5 To describe in more detail: The process of replicating correct results between planes (or marking correct results across planes) can be referred to as reconstruction. Replicating correct results between planes can involve reading correct results from one sister plane and writing correct results to another sister plane.
[0071] Therefore, APM devices can use redundant computing to perform computational operations in parallel across multiple planes.
[0072] Figure 5 This describes an instance of plane 500 that supports redundant computation across planes, as disclosed in the examples herein. Plane 500 may contain planes P0 to P6, which can be instances of planes P0 to P6 after computational operations are performed on operand vectors v0 and v1, as referenced. Figure 5 As described. For example, planes P0 to P6 may contain a portion of vector 31, which may represent a reference. Figure 5 The result of the computational operation described.
[0073] Therefore, bits 0 to 7 of the elements of vector v31 in plane P0 can represent the results of computation operations on bits 0 to 7 of the elements of operand vectors v0 and v1. In sister planes P1 and P4, bits 8 to 15 of the elements of vector v31 can represent the corresponding results of computation operations on bits 8 to 15 of the elements of operand vectors v0 and v1 (e.g., plane P1 can store c8-based...). Spec =0(c8) Act The result of the first possible value), and plane P4 can store c8-based results. Spec =1(c8) Act The result of the second possible value). In sister planes P2 and P5, bits 16 to 23 of the elements of vector v31 can represent the corresponding results of computation operations on bits 16 to 23 of the elements of operand vectors v0 and v1 (e.g., plane P2 can store the result of computation operations based on c16). Spec =0(c16) Act The result of the first possible value), and plane P5 can store c16-based results. Spec =1(c16) Act The result of the second possible value). And in the sister planes P3 and P4, bits 24 to 31 of the elements of vector v31 can represent the corresponding results of the computation operation on bits 24 to 31 of the elements of operand vectors v0 and v1 (for example, plane P3 can store the result of the computation operation on bits 24 to 31 of the elements of operand vectors v0 and v1). Spec =0(c24) Act The result of the first possible value), and plane P6 can store the result based on c24. Spec =1(c24) Act The result of the second possible value).
[0074] The result of each computational operation can be stored in planes P0 through P6. However, as mentioned earlier, at least some results in each plane may be incorrect. To ensure that at least one sister plane has the correct result for each element, the APM device can read the correct result from one sister plane and write it to another sister plane. For example, if plane P4 stores the correct result for element 17, then the APM device can read the correct result from element 17 in plane P4 and write the correct result to element 17 in P1 (thus overwriting the incorrect result written to element 17 in P1).
[0075] An APM device can determine which results are correct by comparing the actual carry value of an element with the speculative carry value used for that element. For example, if a result is calculated using a speculative carry value that matches (e.g., equals) the actual carry value of the element, then the APM device can determine that the result for that element is correct. To illustrate, if the actual carry of the element is c8...Act If the value is equal to 1, then the APM device can determine that the correct result for the element is in plane P4 (because plane P4 uses c8). Spec =1).
[0076] In some instances, an APM device can replicate correct results in a single direction (e.g., from one sister plane to another, but not vice versa) of a pair of sister planes. For example, an APM device can replicate correct results from plane P4 to plane P1, but not from plane P1 to plane P4. Replicating correct results in a single direction of a pair of sister planes reduces reconstruction latency (e.g., the amount of time spent filling one of the sister planes with correct results) compared to other techniques, but may leave some elements with incorrect results for each sister plane. In other instances, an APM device can replicate correct results in two directions (e.g., correct results from each plane can be replicated to the other plane). For example, an APM device can replicate correct results from plane P4 to plane P1 and from plane P1 to plane P4. Replicating correct results in both directions for a pair of sister planes ensures that each sister plane has correct results for each element, but may increase reconstruction latency compared to other techniques.
[0077] If an APM device replicates correct results in a single direction across a pair of sister planes, the APM device can select the direction based on the ratio of elements with correct results to elements with incorrect results. For example, the APM can determine the sister plane with the lowest ratio of correct to incorrect results as the give plane, where the give plane is the plane from which correct results are replicated. By selecting the sister plane with the lowest ratio of correct to incorrect results as the give plane, the APM device can reduce reconstruction latency compared to using another sister plane as the give plane (because fewer elements need to be replicated). For example, if plane P4 has 56 correct results and plane P1 has 200 correct elements, the APM device can reduce reconstruction time by replicating the 56 correct results from plane P4 to plane P1 (compared to replicating the 200 correct results from plane P1 to plane P4).
[0078] In some instances, the APM device can copy the correct results from each pair of sister planes to a new plane, rather than copying the correct results between sister planes. For example, the APM device can copy the correct results from planes P1 and P4 to a new plane P7 (not shown). Similarly, the APM device can copy the correct results from planes P2 and P5 to a new plane P8 (not shown). And the APM device can copy the correct results from planes P3 and P6 to a new plane P9 (not shown). Alternatively, the APM device can copy the correct results from each pair of sister planes to different pairs of sister planes. For example, the APM device can copy the correct results from planes P1 and P4 to planes P2 and / or plane 5. Similarly, the APM device can copy the correct results from planes P2 and P5 to planes P3 and / or plane 6. And the APM device can copy the correct results from planes P3 and P6 to planes P1 and P4.
[0079] In some instances, an APM device can copy results between planes on a bit-serial, row-serial basis. For example, an APM device can (in parallel) copy the least significant bit from each correct element in each plane to another sister plane, then (in parallel) copy the next significant bit from each correct element, and so on. Alternatively, an APM device can copy results between planes on a bit-parallel, row-serial basis. For example, an APM device can (in parallel) copy bits from the least significant correct element in one sister plane to another sister plane, then (in parallel) copy bits from the next significant correct element, and so on.
[0080] Therefore, an APM device can collect the correct results of computational operations in one or more planes by copying the correct bits between planes. Alternatively, the APM device can reconstruct the correct results by marking the correct bits in each plane (instead of copying the correct bits between planes). In this way, the APM device can refer to the markings to determine the correct bits in each plane for operation during subsequent operations.
[0081] Figure 6 This describes an example of a processing flow 600 that supports redundant computation across planes, as disclosed herein. Processing flow 600 can be implemented by an APM system or APM device, as described herein.
[0082] In 605, the APM device can perform computational operations (e.g., using associative processing) on a set of operand vectors (e.g., v0 and v1). The APM device can perform computational operations using a set of planes (e.g., planes P0 to P6), as referenced. Figure 4 As described. For example, an APM device can use the first possible value of the actual carry (e.g., c). Spec =0) performs calculations on some planes (e.g., planes P1, P2, and P3) and can use a second possible value with actual carry (e.g., c).Spec =1) Perform calculations in other planes (e.g., planes P4, P5, and P6).
[0083] An APM device can perform computational operations across the group of planes on an element-by-element basis. For example, an APM device can concurrently perform a computational operation on element 0 (denoted as E[0]) in each plane of the group of planes. The APM device can then concurrently perform a computational operation on element 1 (denoted as [E1]) in each plane of the group of planes. And so on. Thus, in some instances, an APM device can perform computational operations on elements serially, but can perform computational operations on planes in parallel. Performing computational operations on vectors allows the APM device to determine the result of the computational operation on each element and the actual carry value of the computational operation.
[0084] At 610, the APM device can write the result of the computation operation to the group plane. For example, the APM device can write the result from plane P0 to plane P0, from plane P1 to plane P1, from plane P2 to plane P2, and so on. In some instances, the APM device can write the result of the computation operation on an element (e.g., [Ex]) before performing the computation operation on the next element (e.g., [Ex]). In other words, the operation at 610 can overlap with the operation at 605.
[0085] In 615, the APM device can determine the correct result for each element across planes. For example, the APM device can determine which of planes P0 and P4 has the correct result for element 0, which of planes P0 and P4 has the correct result for element 1, which of planes P0 and P4 has the correct result for element 2, and so on for each element and each pair of sister planes. The sister planes use potentially correct values for the actual carry of the elements (e.g., c). Spec =c Act A sister plane of an element may be a plane that has the correct result for that element.
[0086] At 620, the APM device can determine the ratio of correct to incorrect results for one or more pairs of sister planes, or for each pair of sister planes. For example, the APM device can determine the ratio of correct to incorrect results for plane P0 as 56 / 200, and the ratio of correct to incorrect results for plane P1 as 200 / 56. Alternatively, the APM device can determine the number of correct results (or the number of incorrect results) for each plane in a pair of sister planes.
[0087] At 625, the APM device can copy correct results between sister planes (e.g., the APM device can perform reconstruction). For example, the APM device can copy correct results from plane P1 to plane P4. The APM device can copy correct results from the sister plane with the lowest ratio of correct to incorrect results (e.g., the sister plane with the fewest correct results). Alternatively, the APM device can copy correct results from each sister plane to another sister plane. Alternatively, the APM device can copy correct results from each sister plane to a new plane. Alternatively, the APM device can copy correct results from each sister plane to one or more of the sister planes in the next valid pair of sister planes.
[0088] Therefore, APM devices can use redundant planes and associated processing to perform computational operations in parallel across multiple planes, which can reduce waiting time.
[0089] Figure 7 A block diagram 700 illustrates an apparatus 720 supporting cross-plane redundancy computation according to examples disclosed herein. Apparatus 720 may be an example of an aspect of an apparatus, as referenced... Figures 1 to 6 As described herein, device 720 or its various components may be instances of various aspects of a device for performing cross-plane redundancy computation, as described herein. For example, device 720 may include an associative processing circuitry system 725, an access circuitry system 730, a controller 735, or any combination thereof. Each of these components may communicate directly or indirectly with each other (e.g., via one or more buses).
[0090] The associative processing circuitry 725 may be configured or otherwise supported to perform computational operations on first data stored in a first plane (e.g., plane P1) of a plurality of planes containing content-addressable memory cells (e.g., using associative computation), wherein the computational operations are at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a set of consecutive bits of a vector (e.g., bits 8 to 15 of vector v0). In some instances, the associative processing circuitry 725 may be configured or otherwise supported to perform computational operations on second data stored in a second plane (e.g., plane P4) concurrently with the computational operations performed on the first data, wherein the second data represents the set of consecutive bits of a vector (e.g., bits 8 to 15 of vector v0). The access circuitry 730 may be configured or otherwise supported to perform means for reading third data from the first plane and writing it to the second plane, the third data representing the result of a computational operation on the first data.
[0091] In some instances, the controller 735 may be configured or otherwise supported for determining the output bit (e.g., c8) based at least in part on a second set of consecutive bits (e.g., bits 0 to 7 of vector v0) of a vector that is less significant than the first set of consecutive bits. Act A device for calculating the value of a third data bit, wherein the third data is copied from the first plane to the second plane based at least in part on the value of the output bit.
[0092] In some instances, the computational operation on the first data is based at least in part on the first value of the output bit, and the controller 735 may be configured or otherwise supported to allow for means of determining that the value of the output bit is equal to the first value, wherein the third data is copied from the first plane to the second plane based at least in part on the value being equal to the first value.
[0093] In some instances, the access circuit system 730 may be configured or otherwise supported to perform computational operations on fourth data representing a second set of consecutive bits (e.g., bits 0 to 7 of vector v0), wherein the value of the output bit is based at least in part on the computational operations performed on the fourth data.
[0094] In some instances, the fourth data is stored in the third plane. In some instances, computational operations on the fourth data are performed concurrently with computational operations on the first and second data.
[0095] In some instances, the access circuitry 730 may be configured or otherwise supported to allow for writing third data to a first plane, at least in part, based on a computational operation performed on first data. In some instances, the access circuitry 730 may be configured or otherwise supported to allow for writing fourth data to a second plane, at least in part, based on a computational operation performed on second data, wherein writing the third data from the first plane to the second plane replaces the fourth data with the third data.
[0096] In some instances, the associative processing circuitry 725 may be configured or otherwise supported for (e.g., using associative computation) concurrently performing computational operations on fourth data stored in a third plane (e.g., plane P2) with the computational operations performed on the first and second data, wherein the fourth data represents a second set of consecutive bits of a vector (e.g., bits 16 to 23 of vector v0). In some instances, the associative processing circuitry 725 may be configured or otherwise supported for (e.g., using associative computation) concurrently performing computational operations on fifth data stored in a fourth plane (e.g., plane P5) with the computational operations performed on the fourth data, wherein the fifth data represents a second set of consecutive bits of a vector (e.g., bits 16 to 23 of vector v0).
[0097] In some instances, the associative processing circuitry 725 may be configured or otherwise supported for (e.g., using associative computation) performing computational operations on first data stored in a first plane (e.g., plane P0) of a plurality of planes containing content-addressable memory cells, wherein the computational operations are at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a first set of consecutive bits of a vector (e.g., bits 0 to 7 of vector v0). In some instances, the associative processing circuitry 725 may be configured or otherwise supported for (e.g., using associative computation) at least partially based on output bits (e.g., c8) from computational operations on the first data. Act The means for performing a computation operation on second data stored in a second plane (e.g., plane P1) based on a first value of the vector, wherein the second data represents a second set of consecutive bits of a vector (e.g., bits 8 to 15 of vector v0). In some instances, the associative processing circuitry 725 may be configured or otherwise supported for (e.g., using associative computation) at least partially based on the output bits (e.g., c8) from the computation operation on the first data. Act A means for performing a calculation operation on third data stored in a third plane (e.g., plane P4) for the second value of a vector, wherein the third data represents a second set of consecutive bits of a vector (e.g., bits 8 to 15 of vector v0).
[0098] In some instances, the controller 735 may be configured or otherwise supported for determining output bits from computation operations on the first data (e.g., c8). Act The access circuit system 730 may be configured or otherwise supported for reading fourth data from the second plane and writing it to the third plane, at least in part based on the output bit having the first value, wherein the fourth data represents the result of a computational operation on the third data.
[0099] In some instances, the controller 735 may be configured or otherwise supported for determining output bits from computation operations on the first data (e.g., c8). Act The access circuit system 730 may be configured or otherwise supported for reading fourth data from the third plane and writing it back to the second plane, at least in part based on the output bit having a second value, the fourth data representing the result of a computational operation on the third data.
[0100] In some instances, the controller 735 may be configured or otherwise supported for determining output bits from computation operations on the first data (e.g., c8). ActA device having a first value. In some instances, the access circuit system 730 may be configured or otherwise supported for reading fourth data from a second plane (e.g., plane P1) and writing it to a fourth plane (e.g., plane P2 or plane P5) based at least in part on a determination, the fourth data representing the result of a computational operation on the second data.
[0101] In some instances, the associative processing circuitry 725 may be configured or otherwise supported to perform computational operations on fourth data stored in a fourth plane (e.g., plane P2) based at least in part on a first value (e.g., using associative computation), wherein the fourth data represents a third set of consecutive bits of a vector (e.g., bits 16 to 23 of vector v0). In some instances, the associative processing circuitry 725 may be configured or otherwise supported to perform computational operations on fifth data stored in a fifth plane (e.g., plane P5) based at least in part on a second value, wherein the fifth data represents a third set of consecutive bits of a vector (e.g., bits 16 to 23 of vector v0).
[0102] In some instances, the associative processing circuitry 725 may be configured or otherwise supported for (e.g., using associative computation) performing computational operations on first data stored in a first plane (e.g., plane P0) of a plurality of planes containing content-addressable memory cells, wherein the computational operations are at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a first set of consecutive bits of a vector (e.g., bits 0 to 7 of vector v0). In some instances, the associative processing circuitry 725 may be configured or otherwise supported for (e.g., using associative computation) performing computational operations on second data concurrently with the computational operations on the first data, the second data being stored in a second plane (e.g., plane P1) and representing a second set of consecutive bits more significant than the first set of consecutive bits (e.g., bits 8 to 15 of vector v0), wherein the computational operations on the second data are at least partially based on a first value (e.g., 0) from the output bits (e.g., c8Act) of the computational operations on the first data. In some instances, the associative processing circuitry 725 may be configured or otherwise supported for performing computational operations on third data concurrently with computational operations on first data, the third data being stored in a third plane (e.g., plane P4) and representing a second set of consecutive bits of a vector (e.g., bits 8 to 15 of vector v0), wherein the computational operation on the third data is at least partially based on the output bits (e.g., c8) from the computational operation on the first data. ActThe second value of the second data (e.g., 1). In some instances, the access circuitry 730 may be configured or otherwise supported to support means for reading fourth data from the second plane and writing it to the first plane, the fourth data representing the result of a computational operation on the second data, wherein at least in part is based on the output bits (e.g., c8) from the computational operation on the first data having a first value (e.g., 0). Act And copy the fourth data.
[0103] Figure 8 A flowchart illustrating a method 800 for supporting cross-plane redundancy computation based on examples disclosed herein. Operation of method 800 may be implemented by means or components thereof as described herein. For example, operation of method 800 may be performed by means of means such as those described herein. Figures 1 to 7 As described. In some instances, the device may execute a set of instructions to control the functional elements of the device to perform the described function. Alternatively, the device may use dedicated hardware to perform aspects of the described function.
[0104] At 805, the method may include performing a computational operation on first data stored in a first plane of a plurality of planes comprising content-addressable memory cells, wherein the computational operation is at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a set of consecutive bits of a vector. The operation of 805 may be performed according to examples disclosed herein. In some examples, aspects of the operation of 805 may be performed by an associative processing circuitry system 725, as referenced... Figure 7 As described.
[0105] At 810, the method may include performing computation operations concurrently with performing computation operations on first data and on second data stored in a second plane, wherein the second data represents the group of consecutive bits of a vector. The operation of 810 may be performed according to examples disclosed herein. In some instances, aspects of the operation of 810 may be performed by an associative processing circuitry system 725, as referenced... Figure 7 As described.
[0106] At 815, the method may include reading third data from a first plane and writing it to a second plane, the third data representing the result of a computational operation on the first data. The operation of 815 may be performed according to examples disclosed herein. In some instances, aspects of the operation of 815 may be performed by the access circuitry system 730, as referenced... Figure 7 As described.
[0107] In some instances, the device as described herein may perform one or more methods, such as method 800. The device may include features, circuitry, logic, means, or instructions (e.g., a non-transitory computer-readable medium storing instructions executable by a processor) for performing aspects of this disclosure, or any combination thereof:
[0108] Aspect 1: A method or apparatus comprising operations, features, circuitry, logic, means, or instructions, or any combination thereof, for performing: performing a computational operation on first data stored in a first plane of a plurality of planes comprising content-addressable memory cells, wherein the computational operation is at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a set of consecutive bits of a vector; performing a computational operation concurrently with performing the computational operation on the first data on second data stored in a second plane, wherein the second data represents the set of consecutive bits of a vector; and reading third data from the first plane and writing it into the second plane, the third data representing the result of the computational operation on the first data.
[0109] Aspect 2: The method or apparatus according to aspect 1 further includes operations, features, circuitry, logic, means or instructions, or any combination thereof, for determining the value of an output bit based at least in part on a second set of consecutive bits of a vector that is less efficient than the set of consecutive bits, wherein third data is copied from the first plane to the second plane based at least in part on the value of the output bit.
[0110] Aspect 3: The method or apparatus according to aspect 2, wherein the computation operation on the first data is based at least in part on the first value of the output bit, and the method, apparatus and non-transitory computer-readable medium further include operations, features, circuit systems, logic, means or instructions, or any combination thereof, for performing the following: determining that the value of the output bit is equal to the first value, wherein the third data is copied from the first plane to the second plane based at least in part on the value being equal to the first value.
[0111] Aspect 4: The method or apparatus according to any one of Aspects 2 to 3 further comprises an operation, feature, circuit system, logic, means or instruction, or any combination thereof, for performing a computation operation on fourth data representing a second set of consecutive bits, wherein the value of the output bit is based at least in part on the computation operation performed on the fourth data.
[0112] Aspect 5: The method or apparatus according to aspect 4 further includes operations, features, circuit systems, logic, means or instructions, or any combination thereof, for performing the following operations: storing a fourth data in a third plane and performing the computation operation on the fourth data concurrently with the computation operations on the first data and the second data.
[0113] Aspect 6: The method or apparatus according to any one of aspects 1 to 5 further comprises operations, features, circuit systems, logic, means or instructions, or any combination thereof, for performing: writing third data to a first plane based at least in part on performing a computational operation on first data and writing fourth data to a second plane based at least in part on performing a computational operation on second data, wherein writing the third data from the first plane to the second plane replaces the fourth data with the third data.
[0114] Aspect 7: The method or apparatus according to any one of aspects 1 to 6 further comprises an operation, feature, circuit system, logic, means or instruction, or any combination thereof, for performing the following operations: performing a calculation operation on fourth data stored in a third plane concurrently with performing calculation operations on first data and second data, wherein the fourth data represents a second set of consecutive bits of a vector; and performing a calculation operation on fifth data stored in a fourth plane concurrently with performing calculation operations on the fourth data, wherein the fifth data represents a second set of consecutive bits of a vector.
[0115] Figure 9 A flowchart illustrating a method 900 supporting cross-plane redundancy computation according to an example disclosed herein. Operation of method 900 may be implemented by means or components thereof as described herein. For example, operation of method 900 may be performed by means of means such as those described herein. Figures 1 to 7 As described. In some instances, the device may execute a set of instructions to control the functional elements of the device to perform the described function. Alternatively, the device may use dedicated hardware to perform aspects of the described function.
[0116] At 905, the method may include performing a computational operation on first data stored in a first plane of a plurality of planes comprising content-addressable memory cells, wherein the computational operation is at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a first set of consecutive bits of a vector. The operation of 905 may be performed according to examples disclosed herein. In some examples, aspects of the operation of 905 may be performed by an associative processing circuitry system 725, as referenced... Figure 7 As described.
[0117] In 910, the method may include performing a computation operation on second data stored in a second plane, at least in part based on a first value from the output bits of a computation operation on the first data, wherein the second data represents a second set of consecutive bits of a vector. The operation of 910 may be performed according to examples disclosed herein. In some instances, aspects of the operation of 910 may be performed by an associative processing circuitry system 725, as referenced... Figure 7 As described.
[0118] In 915, the method may include performing a computation operation on third data stored in a third plane, at least in part based on a second value from the output bits of the computation operation on the first data, wherein the third data represents a second set of consecutive bits of a vector. The operation of 915 may be performed according to examples disclosed herein. In some instances, aspects of the operation of 915 may be performed by an associative processing circuitry system 725, as referenced... Figure 7 As described.
[0119] In some instances, the device as described herein may perform one or more methods, such as method 900. The device may include features, circuitry, logic, means, or instructions (e.g., a non-transitory computer-readable medium storing instructions executable by a processor) for performing aspects of this disclosure, or any combination thereof:
[0120] Aspect 8: A method or apparatus comprising operations, features, circuitry, logic, means, or instructions, or any combination thereof, for performing: performing a computation operation on first data stored in a first plane comprising a plurality of planes including content-addressable memory cells, wherein the computation operation is at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a first set of consecutive bits of a vector; performing a computation operation on second data stored in a second plane, at least partially based on a first value from the output bits of the computation operation on the first data, wherein the second data represents a second set of consecutive bits of a vector; and performing a computation operation on third data stored in a third plane, at least partially based on a second value from the output bits of the computation operation on the first data, wherein the third data represents a second set of consecutive bits of a vector.
[0121] Aspect 9: The method or apparatus according to aspect 8 further comprises an operation, feature, circuit system, logic, means or instruction, or any combination thereof, for: determining that an output bit from a computational operation on first data has a first value and reading fourth data from a second plane and writing it to a third plane based at least in part on the output bit having the first value, the fourth data representing the result of a computational operation on the third data.
[0122] Aspect 10: The method or apparatus according to any one of Aspects 8 to 9 further comprises an operation, feature, circuit system, logic, means or instruction, or any combination thereof, for determining that an output bit from a computational operation on first data has a second value and reading fourth data from a second plane and writing it to a third plane based at least in part on the second value of the output bit, the fourth data representing the result of a computational operation on third data.
[0123] Aspect 11: The method or apparatus according to any one of aspects 8 to 10 further comprises an operation, feature, circuit system, logic, means or instruction, or any combination thereof, for determining that an output bit from a computational operation on first data has a first value and reading fourth data from a second plane and writing it to a fourth plane, at least in part based on the determination, the fourth data representing the result of the computational operation on the second data.
[0124] Aspect 12: The method or apparatus according to any one of aspects 8 to 11 further comprises an operation, feature, circuit system, logic, means or instruction, or any combination thereof, for performing: performing a calculation operation on fourth data stored in a fourth plane based at least in part on a first value, wherein the fourth data represents a third set of consecutive bits of a vector; and performing a calculation operation on fifth data stored in a fifth plane based at least in part on a second value, wherein the fifth data represents a third set of consecutive bits of a vector.
[0125] Figure 10 A flowchart illustrating a method 1000 for supporting cross-plane redundancy computation based on examples disclosed herein is provided. Operation of method 1000 may be implemented by means of apparatus or components thereof as described herein. For example, operation of method 1000 may be performed by means of apparatus, as referenced herein. Figures 1 to 7 As described. In some instances, the device may execute a set of instructions to control the functional elements of the device to perform the described function. Alternatively, the device may use dedicated hardware to perform aspects of the described function.
[0126] At 1005, the method may include performing a computational operation on first data stored in a first plane of a plurality of planes comprising content-addressable memory cells, wherein the computational operation is at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a first set of consecutive bits of a vector. The operation of 1005 may be performed according to examples disclosed herein. In some examples, aspects of the operation of 1005 may be performed by an associative processing circuitry system 725, as referenced... Figure 7 As described.
[0127] In 1010, the method may include: performing a computation operation on second data concurrently with a computation operation on first data, the second data being stored in a second plane and representing a second set of consecutive bits that are more significant than the first set of consecutive bits, wherein the computation operation on the second data is based at least in part on a first value from the output bits of the computation operation on the first data. The operation of 1010 may be performed according to examples disclosed herein. In some examples, aspects of the operation of 1010 may be performed by an associative processing circuitry system 725, as referenced... Figure 7 As described.
[0128] In 1015, the method may include: performing a computation operation on third data concurrently with a computation operation on first data, the third data being stored in a third plane and representing a second set of consecutive bits of a vector, wherein the computation operation on the third data is at least partially based on a second value from the output bits of the computation operation on the first data. The operation of 1015 may be performed according to examples disclosed herein. In some examples, aspects of the operation of 1015 may be performed by an associative processing circuitry system 725, as referenced... Figure 7 As described.
[0129] In 1020, the method may include reading fourth data from a second plane and writing it to a first plane, the fourth data representing the result of a computation operation on the second data, wherein the fourth data is copied at least in part based on the output bits from the computation operation on the first data having a first value. The operation of 1020 may be performed according to examples disclosed herein. In some instances, aspects of the operation of 1020 may be performed by the access circuitry system 730, as referenced... Figure 7 As described.
[0130] In some instances, the device as described herein may perform one or more methods, such as method 1000. The device may include features, circuitry, logic, means, or instructions (e.g., a non-transitory computer-readable medium storing instructions executable by a processor) for performing aspects of this disclosure, or any combination thereof:
[0131] Aspect 13: A method or apparatus comprising operations, features, circuitry, logic, means, or instructions, or any combination thereof, for performing: performing a computational operation on first data stored in a first plane comprising content-addressable memory cells, wherein the computational operation is at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a first set of consecutive bits of a vector; performing a computational operation on second data concurrently with the computational operation on the first data, the second data being stored in a second plane and representing a second set of consecutive bits more significant than the first set of consecutive bits; wherein the computational operation on the second data is at least partially based on a first value from the output bits of the computational operation on the first data; performing a computational operation on third data concurrently with the computational operation on the first data, the third data being stored in a third plane and representing a second set of consecutive bits of a vector, wherein the computational operation on the third data is at least partially based on a second value from the output bits of the computational operation on the first data; and reading fourth data from the second plane and writing it back to the first plane, the fourth data representing the result of the computational operation on the second data, wherein the fourth data is copied at least partially based on the output bits of the computational operation on the first data having a first value.
[0132] It should be noted that the methods described herein describe possible implementations, and the operations and steps can be rearranged or otherwise modified, and other implementations are possible. Furthermore, parts from two or more methods can be combined.
[0133] Describe a device. The following provides an overview of various aspects of the device as described herein:
[0134] Aspect 14: The device comprising: a memory die including a plurality of planes arranged as a plurality of data blocks, the plurality of planes including content-addressable memory cells; and logic coupled to the memory die and configured to: perform a computational operation on first data stored in a first plane of the plurality of planes, wherein the computational operation is at least partially based on the capability of the content-addressable memory cells, and wherein the first data represents a set of consecutive bits of a vector; concurrently perform a computational operation on second data stored in a second plane of the plurality of planes, wherein the second data represents the set of consecutive bits of a vector; and read third data from the first plane and write it to the second plane, the third data representing the result of the computational operation on the first data.
[0135] Aspect 15: The device according to aspect 14, wherein the logic is further configured to: determine the value of the output bit based at least in part on a second set of consecutive bits of a vector that is less effective than the set of consecutive bits, wherein the third data is copied from the first plane to the second plane based at least in part on the value of the output bit.
[0136] Aspect 16: The apparatus according to aspect 15, wherein the computation operation on the first data is based at least in part on the first value of the output bit, and wherein the computation operation on the second data is based at least in part on the second value of the output bit, and wherein the logic is further configured to: determine that the value of the output bit is equal to the first value, wherein the third data is copied from the first plane to the second plane based at least in part on the value being equal to the first value.
[0137] Aspect 17: The device according to any one of Aspects 15 to 16, wherein the logic is further configured to: perform a computation operation on fourth data representing a second group of consecutive bits, wherein the value of the output bit is based at least in part on the computation operation performed on the fourth data.
[0138] Aspect 18: The apparatus according to aspect 17, wherein the fourth data is stored in a third plane among a plurality of planes, and the computation operation on the fourth data is performed concurrently with the computation operations on the first data and the second data.
[0139] Aspect 19: The device according to any one of Aspects 14 to 18, wherein the logic is further configured to: write third data to a first plane based at least in part on performing a computation operation on first data; and write fourth data to a second plane based at least in part on performing a computation operation on second data, wherein writing the third data from the first plane to the second plane replaces the fourth data with the third data.
[0140] Aspect 20: The device according to any one of aspects 14 to 19, wherein the logic is further configured to: concurrently perform computation operations on fourth data stored in a third plane, wherein the fourth data represents a second set of consecutive bits of a vector, and concurrently perform computation operations on fifth data stored in a fourth plane among a plurality of planes, wherein the fifth data represents a second set of consecutive bits of a vector, in conjunction with the computation operations on the fourth data.
[0141] Aspect 21: The device according to aspect 20, wherein the logic is further configured to: read sixth data from the third plane and write it to the fourth plane, the sixth data representing the result of a computation operation on the fourth data.
[0142] Aspect 22: The device according to any one of aspects 14 to 21, wherein the first plane and the second plane are located in different data blocks of a plurality of data blocks.
[0143] Aspect 23: The device according to any one of aspects 14 to 22, wherein the first plane and the second plane are located in the same data block of a plurality of data blocks.
[0144] Describe a device. The following provides an overview of various aspects of the device as described herein:
[0145] Aspect 24: A device comprising: a memory die including a plurality of planes arranged as a plurality of data blocks, the plurality of planes including content-addressable memory cells; and logic coupled to the memory die and configured to: perform a computation operation on first data stored in a first plane, wherein the computation operation is based at least in part on the capability of the content-addressable memory cells, and wherein the first data represents a first set of consecutive bits of a vector; perform a computation operation on second data stored in a second plane, wherein the second data represents a second set of consecutive bits of a vector, based at least in part on a first value from the output bits of the computation operation on the first data; and perform a computation operation on third data stored in a third plane, wherein the third data represents a second set of consecutive bits of a vector, based at least in part on a second value from the output bits of the computation operation on the first data.
[0146] Aspect 25: The device according to aspect 24, wherein computational operations on first data, second data and third data are performed concurrently.
[0147] Aspect 26: The device according to aspects 24 to 25, wherein the second group of consecutive bits is more effective than the first group of consecutive bits.
[0148] Aspect 27: The device according to any one of Aspects 24 to 26, wherein the logic is further configured to: determine that the output bit from the computation operation on the first data has a first value; and read fourth data from the second plane and write it to the third plane based at least in part on the output bit having the first value, the fourth data representing the result of the computation operation on the third data.
[0149] Aspect 28: The device according to any one of Aspects 24 to 27, wherein the logic is further configured to: determine that the output bit from the computation operation on the first data has a second value; and read fourth data from the third plane and write it to the second plane based at least in part on the output bit having the second value, the fourth data representing the result of the computation operation on the third data.
[0150] Aspect 29: The device according to any one of Aspects 24 to 28, wherein the logic is further configured to: determine that the output bit from the computation operation on the first data has a first value; and read fourth data from the second plane and write it to the fourth plane, at least in part based on the determination, the fourth data representing the result of the computation operation on the second data.
[0151] Aspect 30: The device according to any one of aspects 24 to 29, wherein the logic is further configured to: perform a computation operation on fourth data stored in a fourth plane based at least in part on a first value, wherein the fourth data represents a third set of consecutive bits of a vector, and perform a computation operation on fifth data stored in a fifth plane based at least in part on a second value, wherein the fifth data represents a third set of consecutive bits of a vector.
[0152] Aspect 31: The apparatus according to aspect 30, wherein the calculation operations on the fourth and fifth data are concurrent with the calculation operations on the first, second and third data.
[0153] Aspect 32: The device according to any one of Aspects 30 to 31, wherein the logic is further configured to: determine that a second output bit from a computation operation on the second data has a first value; and read sixth data from the fourth plane and write it to the fifth plane, at least in part based on the second output bit having the first value, the sixth data representing the result of the computation operation on the second data.
[0154] It should be noted that the methods described herein describe possible implementations, and the operations and steps can be rearranged or otherwise modified, and other implementations are possible. Furthermore, parts from two or more methods can be combined.
[0155] The information and signals described herein can be represented using any of a variety of technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips mentioned above can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or optical particles, or any combination thereof. Some diagrams may illustrate a signal as a single signal; however, a signal can represent a bus of signals, where the bus can have various bit widths.
[0156] The terms "electronic communication," "conductive contact," "connection," and "coupling" refer to the relationship between components that support signal flow between them. Components are considered to be in electronic communication (or electrically connected, connected, or coupled to each other) if there is any conductive path between them that can readily support signal flow. At any given time, the conductive path between components that are in electronic communication (or electrically connected, connected, or coupled to each other) can be open or closed, depending on the operation of the device containing the connected components. The conductive path between connected components can be a direct conductive path or an indirect conductive path, which may include intermediate components such as switches, transistors, or other components. In some instances, for example, by using one or more intermediate components such as switches or transistors, the signal flow between connected components can be interrupted for a period of time.
[0157] The term "coupling" refers to a shift from an open-circuit relationship between components (where signals cannot currently communicate between components via conductive paths) to a closed-circuit relationship between components (where signals can communicate between components via conductive paths). When a component (e.g., a controller) couples other components together, that component initiates a change that allows signals to flow between the other components via conductive paths that were previously not permitted.
[0158] The actions can occur "concurrently" if two or more actions occur simultaneously, substantially simultaneously, in partially overlapping time, or in fully overlapping time.
[0159] The descriptions set forth herein, in conjunction with the accompanying drawings, illustrate exemplary configurations and do not represent all instances that may be practiced or that fall within the scope of the claims. The term "exemplary" as used herein means "serving as an example, illustration, or description," and not "preferred" or "superior to other instances." Detailed descriptions include specific details to provide an understanding of the described techniques. However, these techniques may be practiced without these specific details. In some cases, well-known structures and apparatuses are shown in block diagram form to avoid obscuring the concepts of the described instances.
[0160] In the accompanying drawings, similar components or features may have the same reference label. Furthermore, various components of the same type can be distinguished by a dash following the reference label and a second label for differentiation among similar components. If only the first reference label is used in the specification, then the description applies to any of the similar components having the same first reference label, regardless of the second reference label.
[0161] The functions described herein can be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented as software executed by a processor, the functions can be stored on or transmitted via a computer-readable medium as one or more instructions or code. Other examples and embodiments are within the scope of this disclosure and the appended claims. For example, due to the nature of software, the functions described herein can be implemented using software executed by a processor, hardware, firmware, hardwired, or any combination thereof. Features implementing the functions can also be located in various locations, including portions distributed such that the functions are implemented at different physical locations.
[0162] For example, the various illustrative blocks and modules described in connection with this disclosure may be implemented or performed using any of the following devices designed to perform the functions described herein: a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors incorporating a DSP core, or any other such configuration).
[0163] As used herein (included in the claims), "or" as used in a list of items (e.g., a list of items followed by phrases such as "at least one of" or "one or more of") indicates an inclusive list, such that a list of at least one of A, B, or C means A or B or C or AB or AC or BC or ABC (i.e., A and B and C). Furthermore, as used herein, the phrase "based on" should not be considered a reference to a closed set of conditions. For example, an exemplary step described as "based on condition A" may be based on both condition A and condition B without departing from the scope of this disclosure. In other words, as used herein, the phrase "based on" should be considered in the same manner as the phrase "at least partially based on".
[0164] Computer-readable media includes both non-transitory computer storage media and communication media, encompassing any media that facilitates the transfer of a computer program from one place to another. Non-transitory storage media can be any available media accessible by a general-purpose or special-purpose computer. By way of example, and not limitation, non-transitory computer-readable media may include RAM, ROM, electrically erasable programmable read-only memory (EEPROM), compact optical disc (CD) ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to carry or store desired program code in the form of instructions or data structures and is accessible by a general-purpose or special-purpose computer or a general-purpose or special-purpose processor. Furthermore, any connection may be appropriately referred to as computer-readable media. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then those coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave are all included in the definition of media. As used herein, disks and optical discs include CDs, laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of these also fall under the category of computer-readable media.
[0165] The description herein is provided to enable those skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art without departing from its scope, and the general principles defined herein may be applied to other variations. Therefore, this disclosure is not limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An apparatus comprising: A memory die comprising multiple planes arranged as multiple data blocks, the multiple planes including content-addressable memory cells; and Logic, which is coupled to the memory die and configured to: A computational operation is performed on first data stored in a first plane among the plurality of planes, wherein the computational operation is based at least in part on the capabilities of the content-addressable memory cells, and wherein the first data represents a set of consecutive bits of a vector; The computation operation is performed concurrently on the first data and on second data stored in a second plane among the plurality of planes, wherein the second data represents the group of consecutive bits of the vector; and The third data is read from the first plane and written to the second plane, the third data representing the result of the calculation operation on the first data.
2. The device of claim 1, wherein the logic is further configured to: The value of the output bit is determined at least in part based on a second set of consecutive bits of the vector that are less significant than the first set of consecutive bits, wherein the third data is copied from the first plane to the second plane at least in part based on the value of the output bit.
3. The apparatus of claim 2, wherein the calculation operation on the first data is at least partially based on a first value of the output bit, and wherein the calculation operation on the second data is at least partially based on a second value of the output bit, and wherein the logic is further configured to: The value of the output bit is determined to be equal to the first value, wherein the third data is copied from the first plane to the second plane based at least in part on the fact that the value is equal to the first value.
4. The device of claim 2, wherein the logic is further configured to: The calculation operation is performed on the fourth data representing the second group of consecutive bits, wherein the value of the output bit is based at least in part on the calculation operation performed on the fourth data.
5. The device of claim 4, wherein the fourth data is stored in a third plane of the plurality of planes, and wherein the calculation operation on the fourth data is performed concurrently with the calculation operations on the first data and the second data.
6. The device of claim 1, wherein the logic is further configured to: The third data is written to the first plane based at least in part on performing the computational operation on the first data; and The fourth data is written to the second plane at least in part based on performing the computation operation on the second data, wherein the third data is written from the first plane to the second plane to replace the fourth data.
7. The device of claim 1, wherein the logic is further configured to: Concurrently with performing the calculation operation on the first data and the second data, the calculation operation is performed on fourth data stored in the third plane, wherein the fourth data represents the second set of consecutive bits of the vector; and The computation operation is performed concurrently on the fourth data in the fourth plane stored in the plurality of planes, wherein the fifth data represents the second set of consecutive bits of the vector.
8. The device of claim 7, wherein the logic is further configured to: The sixth data is read from the third plane and written to the fourth plane, the sixth data representing the result of the calculation operation on the fourth data.
9. The device of claim 1, wherein the first plane and the second plane are located in different data blocks of the plurality of data blocks.
10. The device of claim 1, wherein the first plane and the second plane are located in the same data block of the plurality of data blocks.
11. An apparatus comprising: A memory die comprising multiple planes arranged as multiple data blocks, the multiple planes including content-addressable memory cells; and Logic, which is coupled to the memory die and configured to: A computational operation is performed on first data stored in a first plane, wherein the computational operation is based at least in part on the capabilities of the content-addressable memory cells, and wherein the first data represents a first set of consecutive bits of a vector; The computation operation is performed on second data stored in a second plane, at least in part, based on a first value from the output bits of the computation operation on the first data, wherein the second data represents a second set of consecutive bits of the vector; and The computation operation is performed on third data stored in a third plane, at least in part, based on the second value of the output bit from the computation operation on the first data, wherein the third data represents the second set of consecutive bits of the vector.
12. The device of claim 11, wherein the computational operations on the first data, the second data, and the third data are performed concurrently.
13. The device of claim 11, wherein the second set of consecutive bits is more valid than the first set of consecutive bits.
14. The device of claim 11, wherein the logic is further configured to: Determine that the output bit from the computation operation on the first data has the first value; and At least in part, fourth data is read from the second plane and written to the third plane based on the first value of the output bit, the fourth data representing the result of the computation operation on the third data.
15. The device of claim 11, wherein the logic is further configured to: Determine that the output bit from the computation operation on the first data has the second value; and Fourth data is read from the third plane and written to the second plane, at least in part based on the output bit having the second value, the fourth data representing the result of the computation operation on the third data.
16. The device of claim 11, wherein the logic is further configured to: Determine that the output bit from the computation operation on the first data has the first value; and Fourth data is read from the second plane and written to the fourth plane, at least in part based on the determination, the fourth data representing the result of the computation operation on the second data.
17. The device of claim 11, wherein the logic is further configured to: The computation operation is performed on fourth data stored in the fourth plane, at least in part, based on the first value, wherein the fourth data represents the third set of consecutive bits of the vector; and The computation operation is performed on fifth data stored in the fifth plane, at least in part, based on the second value, wherein the fifth data represents the third set of consecutive bits of the vector.
18. The device of claim 17, wherein the calculation operations on the fourth data and the fifth data are performed concurrently with the calculation operations on the first data, the second data and the third data.
19. The device of claim 17, wherein the logic is further configured to: Determine that the second output bit from the computation operation on the second data has the first value; and The sixth data is read from the fourth plane and written to the fifth plane, at least in part based on the second output bit having the first value, the sixth data representing the result of the computation operation on the second data.
20. A method comprising: A computational operation is performed on first data stored in a first plane of a plurality of planes including content-addressable memory cells, wherein the computational operation is at least partially based on the capabilities of the content-addressable memory cells, and wherein the first data represents a set of consecutive bits of a vector; The computation operation is performed concurrently on the first data and on the second data stored in the second plane, wherein the second data represents the group of consecutive bits of the vector; and The third data is read from the first plane and written to the second plane, the third data representing the result of the calculation operation on the first data.
21. The method of claim 20, further comprising: The value of the output bit is determined at least in part based on a second set of consecutive bits of the vector that are less significant than the first set of consecutive bits, wherein the third data is copied from the first plane to the second plane at least in part based on the value of the output bit.
22. The method of claim 21, wherein the calculation operation on the first data is at least partially based on a first value of the output bit, and wherein the calculation operation on the second data is at least partially based on a second value of the output bit, the method further comprising: The value of the output bit is determined to be equal to the first value, wherein the third data is copied from the first plane to the second plane based at least in part on the fact that the value is equal to the first value.
23. The method of claim 21, further comprising: The calculation operation is performed on the fourth data representing the second group of consecutive bits, wherein the value of the output bit is based at least in part on the calculation operation performed on the fourth data.
24. The method of claim 23, wherein the fourth data is stored in a third plane, and wherein the computation operation on the fourth data is performed concurrently with the computation operations on the first data and the second data.
25. The method of claim 20, further comprising: The third data is written to the first plane based at least in part on performing the computational operation on the first data; and The fourth data is written to the second plane at least in part based on performing the computation operation on the second data, wherein the third data is written from the first plane to the second plane to replace the fourth data.
26. The method of claim 20, further comprising: Concurrently with performing the calculation operation on the first data and the second data, the calculation operation is performed on fourth data stored in the third plane, wherein the fourth data represents the second set of consecutive bits of the vector; and The computation operation is performed concurrently on the fourth data and on the fifth data stored in the fourth plane, wherein the fifth data represents the second set of consecutive bits of the vector.
27. A method comprising: A computational operation is performed on first data stored in a first plane of a plurality of planes including content-addressable memory cells, wherein the computational operation is based at least in part on the capabilities of the content-addressable memory cells, and wherein the first data represents a first set of consecutive bits of a vector; The computation operation is performed on second data stored in a second plane, at least in part, based on a first value from the output bits of the computation operation on the first data, wherein the second data represents a second set of consecutive bits of the vector; and The computation operation is performed on third data stored in a third plane, at least in part, based on the second value of the output bit from the computation operation on the first data, wherein the third data represents the second set of consecutive bits of the vector.
28. The method of claim 27, further comprising: Determine that the output bit from the computation operation on the first data has the first value; and At least in part, fourth data is read from the second plane and written to the third plane based on the first value of the output bit, the fourth data representing the result of the computation operation on the third data.
29. The method of claim 27, further comprising: The output bit from the computation operation on the first data is determined to have the second value; and Fourth data is read from the third plane and written to the second plane, at least in part based on the output bit having the second value, the fourth data representing the result of the computation operation on the third data.
30. The method of claim 27, further comprising: Determine that the output bit from the computation operation on the first data has the first value; and Fourth data is read from the second plane and written to the fourth plane, at least in part based on the determination, the fourth data representing the result of the computation operation on the second data.
31. The method of claim 27, further comprising: The computation operation is performed on fourth data stored in a fourth plane, at least in part, based on the first value, wherein the fourth data represents a third set of consecutive bits of the vector; and The computation operation is performed on fifth data stored in the fifth plane, at least in part, based on the second value, wherein the fifth data represents the third set of consecutive bits of the vector.
32. A method comprising: A computational operation is performed on first data stored in a first plane of a plurality of planes including content-addressable memory cells, wherein the computational operation is based at least in part on the capabilities of the content-addressable memory cells, and wherein the first data represents a first set of consecutive bits of a vector; The computation operation on the second data is performed concurrently with the computation operation on the first data, the second data being stored in a second plane and representing a second set of consecutive bits that are more significant than the first set of consecutive bits, wherein the computation operation on the second data is based at least in part on a first value from the output bits of the computation operation on the first data; The computation operation on the third data is performed concurrently with the computation operation on the first data, the third data being stored in a third plane and representing the second set of consecutive bits of the vector, wherein the computation operation on the third data is based at least in part on the second value of the output bit from the computation operation on the first data; and A fourth data is read from the second plane and written to the first plane, the fourth data representing the result of the computation operation on the second data, wherein the fourth data is copied at least in part based on the output bits from the computation operation on the first data having the first value.
Citation Information
Patent Citations
Flash memory compression
CN106537327A
Data transfer with a bit vector operation device
US20170242902A1