Method and apparatus for performing diversity matrix operations in a memory array
By converting the memory array into matrix structure and performing spatial diversity matrix calculations, the problem of matrix computing performance bottleneck in the prior art is solved, and efficient matrix computing performance is achieved.
Patent Information
- Application Number
- CN202011399071.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-05
- Filing Date
- 2020-12-02
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2040-12-02
AI Technical Summary
The prior art has performance bottlenecks in performing matrix operations, especially in multi-input/multi-output (MIMO) and large-scale MIMO applications, resulting in inefficient computing.
Matrix and vector calculations are performed using non-transitory computer-readable media by converting the memory array into a matrix structure for matrix transformation, and performing spatial diversity matrix calculations therein.
It realizes the fast and efficient execution of matrix operations, reduces the bottleneck of processor-memory interface, and improves the computing performance in MIMO and large-scale MIMO applications.
Smart Images

Figure CN112926022B_ABST
Abstract
Description
[0001] Related Applications
[0002] This application's subject matter relates to co-owned and co-pending U.S. Patent Application No. 16 / 002,644, filed on June 7, 2018, and titled "AN IMAGE PROCESSOR FORMED IN AN ARRAY OF MEMORY CELLS"; U.S. Patent Application No. 16 / 211,029, filed on December 5, 2018, and titled "METHODS AND APPARATUS FOR INCENTIVIZING PARTICIPATION IN FOG NETWORKS"; U.S. Patent Application No. 16 / 242,960, filed on January 8, 2019, and titled "METHODS AND APPARATUS FOR ROUTINE BASED FOG NETWORKING"; U.S. Patent Application No. 16 / 276,461, filed on February 14, 2019, and titled "METHODS AND APPARATUS FOR CHARACTERIZING MEMORY DEVICES"; U.S. Patent Application No. 16 / 276,471, filed on February 14, 2019, and titled "METHODS AND APPARATUS FOR CHECKING THE RESULTS OF CHARACTERIZED MEMORY SEARCHES"; U.S. Patent Application No. 16 / 276,489, filed on February 14, 2019, and titled "METHODS AND APPARATUS FOR MAINTAINING CHARACTERIZED MEMORY DEVICES"; U.S. Patent Application No. 16 / 403, filed on May 3, 2019, and titled "METHODS AND APPARATUS FOR PERFORMING MATRIX TRANSFORMATIONS WITHIN A MEMORY ARRAY",No. 245, and U.S. Patent Application No. 16 / 689,981, filed on November 20, 2019, and entitled "METHODS AND APPARATUS FOR PERFORMING VIDEO PROCESSING MATRIX OPERATIONS WITHIN A MEMORY ARRAY", each of the foregoing applications is incorporated herein by reference in its entirety.
[0003] Copyright
[0004] A portion of the disclosure of this patent document contains copyrighted material. The copyright owner does not object to the facsimile reproduction by anyone of the patent document or the patent disclosure, as it appears in the patent and trademark office patent files or records, but reserves all copyrights in other respects anyway. TECHNICAL FIELD
[0005] The following generally relates to the fields of data processing and device architectures. Specifically, a processor-memory architecture and method are disclosed for converting a memory array into a matrix configuration for matrix transformation and performing spatial diversity matrix calculations, such as those matrix calculations for multiple-input / multiple-output (MIMO) and massive MIMO applications. BACKGROUND OF THE INVENTION
[0006] Memory devices are widely used to store information in various electronic devices such as computers, wireless communication devices, cameras, digital displays, etc. Information is stored by programming different states of the memory device. For example, a binary device has two states often represented by logic "1" or logic "0". To access the stored information, the memory device can read (or sense) the stored state in the memory device. To store information, the memory device can write (or program) the state in the memory device. So-called volatile memory devices may require power to maintain this stored information, while non-volatile memory devices can store information persistently even after the memory device itself has been power cycled, for example. Different memory manufacturing methods and configurations achieve different capabilities. For example, dynamic random access memory (DRAM) provides high-density volatile storage devices at low cost. Preliminary research involves resistive random access memory (ReRAM), which promises non-volatile performance similar to DRAM.
[0007] Processor devices are typically used in combination with memory devices to perform a wide variety of different tasks and functions. During operation, the processor executes computer-readable instructions (commonly referred to as "software") from the memory. The computer-readable instructions define basic arithmetic, logic, control, input / output (I / O) operations, and so on. As is well known in the computing field, relatively basic computer-readable instructions can perform a wide variety of complex behaviors when combined in sequence. Processors tend to emphasize different circuit configurations and manufacturing technologies than memory devices. For example, processing performance generally involves clock rate, so most processor manufacturing methods and configurations emphasize very high-speed transistor switching structures and the like.
[0008] Over time, the speed and power consumption of both processors and memories have increased. Generally, these improvements are the result of shrinking device sizes due to the physical limitations of telecommunications signals being restricted by the dielectric and distance of the transmission medium. As previously mentioned, most processors and memories are manufactured using different materials and technologies. Therefore, although processors and memories continue to improve, the physical interface between the processor and the memory is a "bottleneck" for the overall system performance. More directly, regardless of how fast a processor or memory can work individually, the performance of the combined processor and memory system is limited by the transfer rate allowed by the interface. This phenomenon has several common names, such as, "processor-memory wall", "von Neumann Bottleneck Effect", and so on.
[0009] Spatial diversity is a term commonly used to refer to systems (such as wireless antenna systems) with multiple spatially separated or distinct elements. Such systems offer improvements especially in terms of diversity, power, and beamforming gain compared to single-piece antenna solutions. Common examples of spatial diversity systems are multiple-input multiple-output (MIMO) systems, single-input multiple-output (SIMO) systems, and multiple-input single-output (MISO) systems. Spatial diversity techniques are used in, for example, 3GPP technologies (such as LTE / LTE-A and 5G New Radio), as well as other applications.
[0010] Taking MIMO (commonly referred to as SU-MIMO or single-user MIMO) as an example, MIMO technology requires performing multiple matrix calculations in pre-coding computations, for example, for supporting transmissions from a base station to a mobile device.
[0011] Similarly, MU-MIMO (multi-user) MIMO - where multiple users share the connection bandwidth, for example, in IEEE Std. 802.11ac and 802.11ax technologies - relies on the widespread use of channel matrix calculations.
[0012] In addition, so-called "massive MIMO" or "mMIMO" systems of the type utilized in 5G NR technology similarly require matrix operations to support, for example, beamforming and beam steering; however, these systems are much larger in scale compared to systems associated with traditional MIMO applications and must be executed fast enough and efficiently enough to support, in particular, the ultra-low latency guarantees associated with the 5G NR standard.
[0013] Accordingly, there is a significant need for improved methods and apparatus for performing fast and efficient computations or manipulations of matrices, such as those used in MIMO, MU-MIMO, or massive MIMO applications. SUMMARY OF THE INVENTION
[0014] The present disclosure particularly provides methods and apparatus for converting a memory array into a matrix configuration for matrix transformation and performing matrix operations therein.
[0015] In one aspect of the present disclosure, a non-transitory computer-readable medium is disclosed. In one exemplary embodiment, the non-transitory computer-readable medium includes: at least one memory cell array, wherein each memory cell in the at least one memory cell array is configured to store a digital value as an analog value in an analog medium; at least one memory sensing component, wherein the at least one memory sensing component is configured to read the analog value of a first memory cell as a first digital value; and logic. In a variant, the memory cell is configured to store the analog value as an impedance of a conductance in the memory cell.
[0016] In one exemplary embodiment, each of the at least one memory cell arrays includes a plurality of sub-arrays, and the non-transitory computer-readable device is configured to perform matrix and vector calculations in an analog manner, wherein the matrix and / or vector values include real numbers, imaginary numbers, and / or complex numbers, and wherein the plurality of sub-arrays are implemented by complex matrix operations.
[0017] In a variant, the plurality of sub-arrays includes a stack of at least four sub-arrays.
[0018] In another variant, individual sub-arrays correspond to the positive real, negative real, positive imaginary, and negative imaginary parts of matrix coefficients.
[0019] In another variant, at least some of the memory cell sub-arrays are configured to operate in parallel with each other. In one such implementation, at least some of the memory cell sub-arrays are connected to a single memory sensing component. In another such implementation, individual memory cell sub-arrays are connected to individual memory sensing components.
[0020] In another embodiment of the computer-readable medium, at least one memory cell array includes two memory cell arrays configured to operate in parallel with each other. In one such variant, the two memory cell arrays are connected to a single memory sensing component. In one implementation thereof, the single memory sensing component includes an arithmetic logic unit (ALU) configured to combine the results of the two memory cell arrays.
[0021] In another implementation, the single memory sensing component includes at least two arithmetic logic units, which are respectively configured to receive values associated with the two arrays and perform separate calculations.
[0022] In another implementation, the single memory sensing component includes three or more ALUs.
[0023] In another such variant, the two memory cell arrays are connected to their own respective memory sensing components.
[0024] In another exemplary embodiment of the computer-readable medium, the logic is further configured to: receive a surjective opcode; operate the memory cell array as a matrix multiplication unit (MMU) based on the matrix transformation opcode; wherein each memory cell of the MMU modifies an analog value in the analog medium according to the matrix transformation opcode and matrix transformation operand; configure the memory sensing component to convert the analog value of the first memory cell into a second digital value according to the matrix transformation opcode and matrix transformation operand; and write the matrix transformation result based on the second digital value in response to reading the matrix transformation operand into the MMU.
[0025] In one variant, the matrix transformation opcode indicates the size of the MMU. In one such implementation, the matrix transformation opcode corresponds to a wireless communication processing operation (e.g., in a MIMO or massive MIMO system such as a 5G NR gNB or UE). In one such configuration, the operation corresponds to a precoding / beamforming operation, and at least some aspect of the size of the MMU corresponds to the size of the precoding / beamforming matrix (which corresponds to the dimensions of the channel information matrix).
[0026] In another configuration, the operation corresponds to a decoding or data estimation or recovery operation.
[0027] In another variant, the matrix transformation operand includes a vector derived from data to be communicated using two or more antennas within a wireless communication system.
[0028] In yet another variant, the matrix transformation operand includes a vector derived from signals received by two or more antennas within a wireless communication system.
[0029] In another variant, the matrix transformation opcode identifies one or more analog values corresponding to one or more memory cells. In one such variant, the one or more analog values corresponding to one or more memory cells are stored within a lookup table (LUT) data structure. In one implementation, the LUT contains a codebook of a predetermined pre-coding matrix, and the one or more analog values include the values / coefficients of the pre-coding matrix. In one of its variants, the one or more analog values corresponding to one or more memory cells are received from a processor device via a non-transitory computer-readable medium.
[0030] In one variant, each memory cell of the MMU includes a resistive random access memory (ReRAM) cell; and each memory cell of the MMU multiplies an analog value in an analog medium according to a matrix transformation opcode and matrix transformation operands.
[0031] In one variant, each memory cell of the MMU further accumulates an analog value in the analog medium and a previous analog value.
[0032] In one variant, a first digital value is represented by a first radix two (2); and a second digital value is represented by a second radix greater than two (2).
[0033] In one aspect of the present disclosure, an apparatus is disclosed. In one embodiment, the apparatus includes a processor coupled to a non-transitory computer-readable medium; wherein the non-transitory computer-readable medium includes one or more instructions that, when executed by the processor, cause the processor to perform the following operations: write a matrix transformation opcode and matrix transformation operands to the non-transitory computer-readable medium; wherein the matrix transformation opcode causes the non-transitory computer-readable medium to operate an array of memory cells as a matrix structure; wherein the matrix transformation operands modify one or more analog values of the matrix structure; and read a matrix transformation result from the matrix structure.
[0034] In one variant, the non-transitory computer-readable medium further includes one or more instructions that, when executed by the processor, cause the processor to perform the following operations: obtain a pre-coding matrix or obtain a pre-coding matrix index or address within a lookup table (LUT); and obtain data to be communicated by a transmitter; wherein the matrix transformation operands include the pre-coding matrix or the pre-coding matrix index; and wherein the matrix transformation result includes a pre-coding operation that transforms the data into a transmit vector.
[0035] In another variant, the non - transitory computer - readable medium includes one or more instructions that, when executed by the processor, cause the processor to perform the following operations: obtain a data recovery matrix or a data recovery matrix index / address within a look - up table (LUT); and obtain a vector corresponding to a signal received by a receiver; wherein the matrix transformation operand includes the data recovery matrix or the data recovery matrix index. In one embodiment, the matrix transformation result includes an estimate of the original data communicated from a transmitter to a receiver.
[0036] In one variant, the matrix transformation opcode causes the non - transitory computer - readable medium to operate another memory cell array as another matrix structure; and the matrix transformation result associated with the matrix structure and another matrix transformation result associated with the other matrix structure are logically combined.
[0037] In one variant, one or more analog values of the matrix structure are stored within a look - up table (LUT) data structure. In one of its embodiments, at least some of the values of the matrix structure are stored in a portion of the memory cells not configured as a matrix structure. In another embodiment, one or more analog values of the matrix structure are provided from the processor to the non - transitory computer - readable medium.
[0038] In one aspect of the present disclosure, a method for performing transform matrix operations is disclosed. In one embodiment, the method includes: receiving a matrix transformation opcode; configuring at least one memory cell array of a memory as a matrix structure based on the matrix transformation opcode; configuring a memory sensing component based on the matrix transformation opcode; and writing a matrix transformation result from the memory sensing component in response to reading a matrix transformation operand into the matrix structure.
[0039] In one variant, the matrix transformation operand corresponds to a vector derived from data to be communicated by a transmitter using two or more antennas, and configuring the matrix structure includes configuring at least one memory cell array with values of a pre - coding matrix. The matrix transformation result includes a vector corresponding to a pre - coded signal to be transmitted through a wireless communication channel (e.g., within a MIMO system) corresponding to the pre - coding matrix.
[0040] In another variant of the method, the matrix transformation operand corresponds to a vector derived from a signal received by a receiver using two or more antennas, and configuring the matrix structure includes configuring at least one memory cell array with values of a data recovery matrix (corresponding to the pre - coding matrix), and the matrix result includes a vector corresponding to the recovered / estimated data being communicated to the receiver.
[0041] In one embodiment, a matrix structure is configured with the values of four matrices: a first matrix of positive real values of the matrix, a second matrix of negative real values of the matrix, a third matrix of positive imaginary values of the matrix, and a fourth matrix of negative imaginary values of the matrix. In a particular configuration thereof, at least one memory cell array comprises two arrays, each configured with the values of the four matrices.
[0042] In another configuration, the two arrays are configured to operate in parallel with each other. In one variant, the two arrays are connected to a single memory sensing component. The single memory sensing component inputs digital values (obtained using analog-to-digital conversion ADC) corresponding to the results of the two arrays into two separate arithmetic logic units (ALUs) and then combines the results of the two ALUs in a third ALU.
[0043] In an alternative configuration, the single memory sensing component inputs digital values corresponding to the results of the two arrays into one ALU, and the one ALU has logic to appropriately combine the results. The two arrays can be connected, for example, to two separate memory sensing components (each having its own ADC and ALU), and the results of the two memory sensing components can be further combined in another ALU.
[0044] In another embodiment, the method comprises continuously reading multiple matrix transform operands into / through the matrix structure and reconfiguring the memory sensing component between at least some of the runs. In one variant, the multiple matrix transform operands comprise a first operand having real coefficients of a vector and a second operand having imaginary coefficients of a vector, and the method comprises configuring the memory sensing component with a first configuration; running the first operand through the matrix structure; reconfiguring the memory sensing component to account for the imaginary coefficients of the second operand; and running the second operand through the matrix structure.
[0045] In one embodiment of the foregoing embodiments, configuring and reconfiguring the memory sensing component comprises configuring and reconfiguring one or more arithmetic logic units (ALUs) within the memory sensing component.
[0046] In another variant, configuring the memory cell array comprises connecting a plurality of word lines and a plurality of bit lines corresponding to the row size and column size associated with the matrix structure.
[0047] In another variant, the method further comprises determining the row size and column size from the matrix transform opcode.
[0048] In another variant, configuring the memory cell array comprises setting one or more analog values of the matrix structure based on a look-up table (LUT) data structure.
[0049] In yet another variant, the method includes identifying an item from a LUT data structure based on a matrix transform opcode.
[0050] In another variant, a memory sensing component is configured to have a radix greater than two (2) for matrix transform results.
[0051] In another aspect of the present disclosure, a device configured to configure a memory device into a matrix configuration is described. In one embodiment, the device includes: a memory; a processor configured to access the memory; and preprocessor logic configured to allocate one or more memory portions to be used as a matrix configuration.
[0052] In yet another aspect of the present disclosure, a computerized image processing apparatus device configured to dynamically configure a memory into a matrix configuration is disclosed. In one embodiment, the computerized image processing apparatus includes: a camera interface; a digital processor device in data communication with the camera interface; and a memory in data communication with the digital processor device and including at least one computer program.
[0053] In another aspect of the present disclosure, a computerized video processing apparatus device configured to dynamically configure a memory into a matrix configuration is disclosed. In one embodiment, the computerized video processing apparatus includes: a camera interface; a digital processor device in data communication with the camera interface; and a memory in data communication with the digital processor device and including at least one computer program.
[0054] In yet another aspect of the present disclosure, a computerized wireless access node device configured to dynamically configure a memory into a matrix configuration is disclosed. In one embodiment, the computerized wireless access node includes: a wireless interface configured to transmit and receive RF waveforms in a spectral portion; a digital processor device in data communication with the wireless interface; and a memory in data communication with the digital processor device and including at least one computer program.
[0055] In one variant, the computerized wireless access node device is configured as a 3GPP eNB or gNB.
[0056] In another variant, the computerized wireless access node device is configured as an IEEE standard 802.11 wireless access point (AP) capable of performing diversity processing to support a wireless channel.
[0057] In another aspect of the present disclosure, a computerized device is disclosed. In one embodiment, the computerized device includes a 3GPP LTE or 5G NR compliant UE (User Equipment) having diversity processing matrix logic, which has been configured according to one or more of the foregoing methods or apparatuses. In one variant, the UE includes a 5G NR millimeter wave system having massive MIMO capabilities.
[0058] In another embodiment, the computerized device is configured to be capable of performing diversity processing to support an IEEE standard 802.11 compliant STA (Station) for one or more wireless channels.
[0059] In an additional aspect of the present disclosure, a computer-readable device is described. In one embodiment, the device includes a storage medium configured to store one or more computer programs in or in combination with a characterized memory. In one embodiment, the device includes a program memory, HDD, or SDD on a computerized controller device. In another embodiment, the device includes a program memory, HDD, or SSD on a computerized access node.
[0060] These and other aspects should become apparent when considered in light of the disclosure provided herein. Description of the Drawings
[0061] Figure 1A A diagrammatic illustration of a processor-memory architecture and associated matrix operations.
[0062] Figure 1B A diagrammatic illustration of a processor-PIM architecture and associated matrix operations.
[0063] Figure 2 A logic block diagram of an exemplary embodiment of a memory device according to various principles of the present disclosure.
[0064] Figure 3 An exemplary side-by-side illustration of a first memory device configuration and a second memory device configuration.
[0065] Figure 4A A diagrammatic illustration of matrix operations involving positive and negative values performed according to the principles of the present disclosure.
[0066] Figure 4B A diagrammatic illustration of an embodiment of matrix operations involving complex (real and imaginary) values performed according to the principles of the present disclosure.
[0067] Figure 4C A diagrammatic illustration of another embodiment of matrix operations involving complex values performed according to the principles of the present disclosure.
[0068] Figure 5A Graphical depiction of an exemplary 2x2 MIMO system that can utilize the methods and apparatuses of the present disclosure.
[0069] Figure 5B Block diagram of a memory device configured to perform matrix operations for a system according to an aspect of the present disclosure Figure 5A for.
[0070] Figure 5C Block diagram of a first implementation of a memory device according to an aspect of the present disclosure Figure 5B for.
[0071] Figure 5D Block diagram of a second exemplary implementation of a memory device according to an aspect of the present disclosure Figure 5B for.
[0072] Figure 6A Graphical depiction of a MIMO system where channel (pre-coding) information is communicated to the transmitter and the receiver utilizes channel decoding.
[0073] Figure 6B Block diagram of a memory device configured to perform transmitter matrix pre-coding operations for a system according to an aspect of the present disclosure Figure 6A for.
[0074] Figure 6C Block diagram of a memory device configured to perform receiver matrix operations for a system according to an aspect of the present disclosure Figure 6A for.
[0075] Figure 6D Graphical depiction of an exemplary 5G NR MIMO system that can utilize the methods and apparatuses of the present disclosure.
[0076] Figure 7A Graphical depiction of a MIMO system using a codebook of predefined matrices.
[0077] Figure 7B Block diagram of a memory device configured to perform transmitter matrix pre-coding operations for a system according to an aspect of the present disclosure Figure 6A for.
[0078] Figure 8A Logical block diagram of an exemplary implementation of a processor-memory architecture.
[0079] Figure 8B Ladder diagram of an exemplary embodiment illustrating the performance of a set of matrix operations according to the principles of the present disclosure.
[0080] Figure 8C Ladder diagram of another exemplary embodiment illustrating the performance of a set of matrix operations according to the principles of the present disclosure.
[0081] Figure 9 A logic flow diagram of an exemplary method for converting a memory array into a matrix configuration and performing matrix operations therein.
[0082] Figure 10 A logic flow diagram of an exemplary embodiment of a generalized method for converting a memory array into a matrix configuration and performing pre-decoded matrix operations therein.
[0083] Figure 10A For Figure 10 A logic flow diagram of an exemplary method for performing pre-decoded matrix operations using a matrix memory configuration according to the method of
[0084] Figure 10B For Figure 10 A logic flow diagram of another exemplary method for performing pre-decoded matrix operations using a matrix memory configuration according to the method of
[0085] All figures Copyright © 2019 Micron Technology, Inc. All rights reserved. DETAILED DESCRIPTION
[0086] Reference is now made to the drawings, where like numerals refer to like parts throughout.
[0087] As used herein, the term "application" generally refers to, but is not limited to, an executable software unit that implements a particular functionality or topic. The topic of an application can vary widely in any number of disciplines and functions (e.g., on-demand content management, e-commerce transactions, brokerage affairs, home entertainment, calculators, etc.), and an application can have more than one topic. An executable software unit typically runs in a predefined environment; for example, the unit can include a downloadable application that runs within an operating system environment.
[0088] As used herein, the terms "channel decoding" and "decoding" can generally refer to, but are not limited to, both or the calculation of channel pre-decoding and channel decoding operations, such as those used in diversity determination associated with a wireless channel.
[0089] As used herein, the term "computer program" or "software" means any sequence or set of human- or machine-recognizable steps that perform a function. Such a program can be virtually represented in any programming language or environment, including, for example, C / C++, Fortran, COBOL, PASCAL, assembly language, markup languages (e.g., HTML, SGML, XML, VoXML), etc., and object-oriented environments such as the Common Object Request Broker Architecture (CORBA), Java TM(including J2ME, Java Bean, etc.), Register Transfer Language (RTL), Very High Speed Integrated Circuit (VHSIC), Hardware Description Language (VHDL), Verilog, etc.
[0090] As used herein, the term "decentralized" or "distributed" refers to, but is not limited to, a configuration or network architecture involving multiple computerized devices that are capable of performing data communication with each other without requiring a given device to communicate through a designated (e.g., central) network entity (e.g., a server device). For example, a decentralized network enables direct peer-to-peer data communication between multiple UEs (e.g., wireless user devices) that make up the network.
[0091] As used herein, the term "Distributed Unit" (DU) refers to, but is not limited to, a distributed logical node within a wireless network infrastructure. For example, a DU may be implemented as a next-generation Node B (gNB) DU (gNB-DU) controlled by the gNB CU described above. One gNB-DU may support one or more cells; a given cell is supported by one gNB-DU.
[0092] As used herein, the term "diversity" non-restrictively includes spatial diversity (e.g., MIMO, SIMO, MISO, MU-MIMO, SU-MIMO, and massive MIMO); multiplexing (e.g., spatial multiplexing of two or more information channels via spatially diverse antenna elements); beamforming (e.g., 2D and 3D spatial beamforming), etc.
[0093] As used herein, the term "Internet (Internet / internet)" is used interchangeably to refer to inter-networks, including but not limited to the Internet. Other common examples include, but are not limited to: a network of external servers, "cloud" entities (e.g., a memory or storage device that is not local to a device and is typically accessible via a network connection at any time), service nodes, access points, controller devices, client devices, etc. The 5G service core network and network components (e.g., DUs, CUs, gNB small cells or femtocells, external nodes with 5G capabilities) residing in the backhaul, fronthaul, crosshaul, or at the "edge" close to residences, companies, and other occupied areas may be included in the "Internet".
[0094] As used herein, the term "LTE" refers to, but is not limited to, any variant or version of the Long-Term Evolution wireless communication standard when applicable, including Long-Term Evolution in unlicensed spectrum (LTE-U), Long-Term Evolution with Licensed-Assisted Access (LTE-LAA), LTE-Advanced (LTE-A), and 4G / 4.5G LTE.
[0095] As used herein, the terms "5G" and "New Radio (NR)" refer to, but are not limited to, devices, methods, or systems that are compatible with 3GPP Release 15 and any modifications, subsequent releases, or amendments or supplements thereto (whether licensed or unlicensed) for New Radio technology.
[0096] As used herein, the term "memory" includes any type of integrated circuit or other storage device suitable for storing digital data, including but not limited to random access memory (RAM), pseudo-static RAM (PSRAM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM) including double data rate (DDR) type memories and graphics DDR (GDDR) and variations thereof, ferroelectric RAM (FeRAM), magnetic RAM (MRAM), resistive RAM (ReRAM), read-only memory (ROM), programmable ROM (PROM), electrically erasable PROM (EEPROM or 2 EPROM), DDR / 2 SDRAM, EDO / FPMS, latency-reduced DRAM (RLDRAM), static RAM (SRAM), "flash" memory (e.g., NAND / NOR), phase change memory (PCM), 3D cross-point memory (3DXpoint), stacked memories such as HBM / HBM2, and magnetic resistive RAM (MRAM), such as spin torque transfer RAM (STTRAM).
[0097] As used herein, the terms "microprocessor" and "processor" or "digital processor" generally refer to any and all types of digital processing devices, including but not limited to digital signal processors (DSPs), reduced instruction set computers (RISC), general purpose processors (GPPs), microprocessors, gate arrays (e.g., FPGAs), PLDs, reconfigurable computing fabrics (RCFs), array processors, secure microprocessors, and application specific integrated circuits (ASICs). Such digital processors may be contained on a single monolithic IC die or distributed across multiple components.
[0098] As used herein, the term "server" refers to any form of computerized component, system, or entity that is adapted to provide data, files, applications, content, or other services over a computer network to one or more other devices or entities.
[0099] As used herein, the term "storage device" refers to, but is not limited to, a computer hard drive (e.g., a hard disk drive (HDD), a solid state drive (SDD)), a flash drive, a DVR device, a memory, a RAID device or array, an optical medium (e.g., a CD-ROM, a laserdisc, a Blu-ray, etc.), or any other device or medium capable of storing content or other information, including semiconductor devices (e.g., those semiconductor devices described herein as memories) capable of maintaining data in the absence of power. Common examples of memory devices for storage include, but are not limited to: ReRAM, DRAM (e.g., SDRAM, DDR SDRAM, DDR2 SDRAM, DDR3 SDRAM, DDR4 SDRAM, GDDR, RLDRAM, LPDRAM, etc.), DRAM modules (e.g., RDIMM, VLP RDIMM, UDIMM, VLP UDIMM, SODIMM, SORDIMM, mini DIMM, VLP mini DIMM, LRDIMM, NVDIMM, etc.), managed NAND, NAND flash (e.g., SLC NAND, MLC NAND, TLS NAND, serial NAND, 3D NAND, etc.), NOR flash (e.g., parallel NOR, serial NOR, etc.), multi-chip packages, hybrid memory cubes, memory cards, solid state storage devices (SSS), and any number of other memory devices.
[0100] As used herein, the term "Wi-Fi" refers to, but is not limited to, any variant of the IEEE standard 802.11 or related standards when applicable, including 802.11a / b / g / n / s / v / ac / ad / av / ax / ba or 802.11-2012 / 2013, 802.11-2016, and Wi-Fi Direct (specifically including the "Peer-to-Peer (P2P) Wi-Fi Specification", which is incorporated herein by reference in its entirety).
[0101] As used herein, the term "wireless" means any wireless signal, data, communication, or other interface, including but not limited to Wi-Fi, Bluetooth / BLE, 3G / 4G / 4.5G / 5G (3GPP / 3GPP2), HSDPA / HSUPA, TDMA, CBRS, CDMA (e.g., IS-95A, WCDMA, etc.), FHSS, DSSS, GSM, PAN / 802.15, WiMAX (802.16), 802.20, Z-Wave, narrowband / FDMA, OFDM, PCS / DCS, LTE / LTE-A / LTE-U / LTE-LAA, analog cellular, CDPD, satellite systems, millimeter wave or microwave systems, acoustic, and infrared (i.e., IrDA).
[0102] As used herein, the term "xNB" refers to any 3GPP compliant node, including but not limited to eNB (eUTRAN) and gNB (5G NR).
[0103] Overview
[0104] The foregoing "processor-memory wall" performance limitation can be extraordinary where the processor-memory architecture repeats similar operations within a large data set. In these cases, the processor-memory architecture must iteratively transfer, manipulate, and store each element of the data set individually. For example, a 4×4 (sixteen (16) elements) matrix multiplication takes four (4) times as long as a 2×2 (four (4) elements) matrix multiplication. In other words, matrix operations scale exponentially with matrix size.
[0105] Various embodiments of the present disclosure relate to converting a memory array into a matrix configuration for matrix transformation and performing matrix operations therein.
[0106] Matrix transformations are commonly used in a number of different applications and can consume a disproportionate amount of processing and / or memory bandwidth. For example, various communication techniques, such as for wireless systems, use matrix multiplication for applications such as beamforming / precoding and / or diversity applications, such as massive multiple-input multiple-output (MIMO) or MU-MIMO channel processing.
[0107] The exemplary embodiments described herein perform matrix transformation within a memory device that includes a matrix configuration and a matrix multiplication unit (MMU). In one exemplary embodiment, the matrix configuration uses a "crossbar" configuration of resistive elements. Each resistive element stores an impedance level representing a corresponding matrix coefficient value. Electrical signals representing an input vector as analog voltages can be applied to drive the crossbar connectivity. The resulting signals can be converted from analog voltages to digital values by the MMU to obtain a vector-matrix product. In some cases, the MMU can additionally perform various other logical operations in the digital domain.
[0108] Unlike existing solutions that iteratively traverse each element of a matrix to compute element values, the crossbar matrix configuration described below computes multiple elements of a matrix "atomically" (i.e., in a single processing cycle). For example, at least a portion of a vector-matrix product can be computed in parallel. The computation based on the "atomicity" of the matrix configuration results in a significant processing improvement over iterative alternatives. Specifically, while iterative techniques scale with matrix size, the atomic matrix configuration computation is independent of matrix dimensions. In other words, an N×N vector-matrix product can be completed in a single atomic instruction.
[0109] MIMO and large-scale MIMO channel decoding techniques can use a codebook such as a predefined matrix and / or a matrix with a known structure and weighting. In a large-scale MIMO system, a single base station can use hundreds or even thousands of antennas, so large-scale MIMO matrix operations may involve matrices and vectors of extremely large sizes. Matrix multiplication operations are required to, for example, transform a channel information matrix, perform pre-decoding / beamforming, and perform data recovery operations.
[0110] Incidentally, practical limitations on component manufacturing limit the capabilities of each element within an individual memory device. For example, most memory arrays are only designed to distinguish between two (2) states (logical "1", logical "0"). While existing memory sensing components can be scaled to distinguish higher levels of precision (e.g., four (4) states, eight (8) states, etc.), it may be impractical to increase the precision of the memory sensing components to support the precision required for larger transforms commonly used for, e.g., video compression, mathematical transforms, etc.
[0111] To achieve these purposes, various embodiments of the present disclosure logically combine one or more matrix organizations and / or MMUs to provide a higher degree of precision and / or processing complexity than might otherwise be possible. In one such embodiment, a first matrix organization and / or MMU can be used to compute a positive vector matrix product, and a second matrix organization and / or MMU can be used to compute a negative vector matrix product. The positive and negative vector matrix products can be summed to determine a net vector matrix product. Similarly, in some embodiments, multiple (e.g., four) matrix organizations can be combined to compute a vector matrix product, where the matrices include positive and negative complex numbers. Additionally, matrix multiplication operations can be ordered or parallelized according to any number of design considerations.
[0112] In view of the present disclosure, other instances of logical matrix operations can be replaced by equivalent results (e.g., decomposition, common matrix multiplication, etc.).
[0113] Certain applications can save a significant amount of power by turning off system components when not in use. However, sleep procedures typically require the processor and / or memory to shuttle data from operational volatile memory to non-volatile memory so that the data is not lost during a power outage. A wake-up procedure is also necessary to retrieve the stored information from non-volatile memory. Shuttling data back and forth between memories is an inefficient use of processor-memory bandwidth. Accordingly, various embodiments disclosed herein take full advantage of the "non-volatile" nature of matrix organizations. In such embodiments, a matrix organization can retain its matrix coefficient values even when the memory does not have power. More directly, the non-volatile nature of the matrix organization enables the processor and memory to transition into a sleep / low-power mode or perform other tasks without disturbing data from volatile memory to non-volatile memory (and vice versa).
[0114] Given the content of the present disclosure, those of ordinary skill in the art will readily appreciate various other combinations and / or variations of the foregoing.
[0115] Detailed Description of Exemplary Embodiments
[0116] Exemplary embodiments of the apparatus and method of the present disclosure will now be described in detail. Although these exemplary embodiments are described in the context of prior specific processor and / or memory configurations, the general principles and advantages of the present disclosure can be extended to other types of processor and / or memory technologies, and thus the following is exemplary in nature only.
[0117] It should also be understood that although typically described in the context of consumer devices (e.g., within a cellular phone, wireless LAN AP or STA, and / or network base station (such as a 3GPP eNB or gNB or even a femtocell / HNB)), the present disclosure can be readily applied to other types of devices, including, for example, server devices, Internet of Things (IoT) devices, and / or for use by individuals, companies or even governments, such as users other than the prohibited "current" users (e.g., U.S. DoD, etc.). Other applications are possible.
[0118] Those of ordinary skill in the art will immediately recognize other features and advantages of the present disclosure upon reference to the accompanying drawings and the detailed description of the exemplary embodiments given below.
[0119] Processor Memory Architecture -
[0120] Figure 1A Illustrate a common processor memory architecture 100 that is useful for illustrating matrix operations. As Figure 1A shown, a processor 102 is connected to a memory 104 via an interface 106. In an illustrative example, the processor multiplies the elements of an input vector a by a matrix M to compute a vector-matrix product b. Mathematically, the input vector a is considered a single-column matrix with the number of elements equal to the number of rows in the matrix M.
[0121] To compute the first element of the vector-matrix product b 0 the processor must iterate through each permutation of the elements of the input vector a for each element within a row of the matrix M. During the first iteration, the first element of the input vector a is read, the current value of the vector-matrix product b 0 is read, and the corresponding matrix coefficient value M 0 is read. The three (3) read values are used in a multiply-accumulate operation to produce an "intermediate" vector-matrix product b 0,0 . Specifically, the multiply-accumulate operation computes: (a 0 ·M 0 ) + b 0,0 ) + b 0, and write the result value back to b 0 . It should be noted that b 0 is the "intermediate value". After the first iteration but before the second iteration, the intermediate value b 0 may not correspond to the final value of the vector matrix product b 0 .
[0122] During the second iteration, read the second element of the input vector a 1 , retrieve the previously calculated intermediate value b 0 , and read the second matrix coefficient value M 1,0 . Use the three (3) read values in a multiply-accumulate operation to produce the first element of the vector matrix product b 0 . The second iteration completes the calculation of b 0 .
[0123] Although not explicitly shown, the iterative process described above is also performed to produce the second element of the vector matrix product b 1 . Additionally, although the foregoing example is a 2×2 vector matrix product, the techniques described therein are generally extended to support vector matrix calculations of any size. For example, a 3×3 vector matrix product calculation iterates over the input vector for three (3) elements in each of the three (3) rows of the matrix; thus, nine (9) iterations are required. A 1024×1024 matrix operation (which is not uncommon for many applications) would require more than a million iterations. More directly, the foregoing iterative process scales exponentially with the change in matrix size.
[0124] Although the foregoing discussion is presented in the context of vector matrix products, one of ordinary skill in the relevant art will readily understand that matrix matrix products can be performed as a series of vector matrix products. For example, calculate the first vector matrix product corresponding to the first single-column matrix of the input vector, calculate the second vector matrix product corresponding to the second single-column matrix of the input vector, and so on. Thus, a 2×2 matrix matrix product will require two (2) vector matrix calculations (i.e., 2×4 = 8 in total), and a 3×3 matrix matrix product will require three (3) vector matrix calculations (i.e., 3×9 = 27 in total).
[0125] One of ordinary skill in the relevant art will readily understand Figure 1AEach iteration of the process described in is hampered by the bandwidth limitation of interface 106 (“processor-memory wall”). Although the processor and the memory may have internal buses with extremely high bandwidths, the processor-memory system can communicate only as fast as interface 106 can support the telecommunications signaling (based on the dielectric properties and transmission distance (about 1 to 2 centimeters) of the material (usually copper) used in interface 106). In addition, interface 106 may also include a variety of additional signal conditioning, amplification, noise correction, error correction, parity calculations, and / or other interface-based logic that further reduces transaction time.
[0126] A common way to improve the performance of matrix operations is to perform the matrix operations within the local processor cache. Unfortunately, the local processor cache occupies processor die space and has a much higher per-bit manufacturing cost compared to, for example, similar memory devices. Therefore, the size of the local cache of a processor is typically much smaller (e.g., a few megabytes) than its memory (which can be several gigabytes). In practical terms, the smaller local cache is a hard limit on the maximum amount of matrix operations that can be performed locally within the processor. As another drawback, larger matrix operations result in poor cache utilization because only one row and one column are accessed at a time (e.g., for a 1024×1024 vector matrix product, only 1 / 1024 of the cache is in use during a single iteration). Therefore, while the processor cache implementation may be acceptable for small matrices, this technique becomes increasingly undesirable as the complexity of the matrix operations grows.
[0127] Another common approach is the so-called processor-in-memory (PIM). Figure 1B Illustrate one such processor PIM architecture 150. As shown in the figure, processor 152 is connected to memory 154 via interface 156. Memory 154 further includes PIM 162 and memory array 164; PIM 162 is tightly coupled to memory array 164 via internal interface 166.
[0128] Similar to the process described above Figure 1A in Figure 1B the processor-PIM architecture 150 multiplies the elements of the input vector a by the matrix M to compute the vector matrix product b. However, PIM 162 reads, multiplies and accumulates, and writes to memory 164 internally via internal interface 166. Internal interface 166 is much shorter than external interface 156; additionally, internal interface 166 can operate natively without, for example, signal conditioning, amplification, noise correction, error correction, parity calculations, etc.
[0129] Although the processor-PIM architecture 150 provides significant performance improvements compared to, for example, the processor-memory architecture 100, the processor-PIM architecture 150 may have other drawbacks. For example, the manufacturing technology ("silicon process") is substantially different between the processor and the memory device because each silicon process is optimized for different design criteria. For example, the processor silicon process may use a finer transistor structure than the memory silicon process, and the finer transistor structure provides faster switching (which improves performance) but suffers from more leakage (which is undesirable for memory retention). Thus, fabricating the PIM 162 and the memory array 164 on the same die results in at least one of them being implemented in a sub-optimal silicon process. Alternatively, the PIM 162 and the memory array 164 may be implemented in separate dies and bonded together; die-to-die communication generally increases the manufacturing cost and complexity and may suffer from various other impairments (e.g., introduced through process cracks, etc.).
[0130] In addition, those of ordinary skill in the art will readily understand that the PIM 162 and the memory array 164 are "hardened" components; the PIM 162 cannot store data and the memory 164 cannot perform computations. The reality is that once the memory 154 is fabricated, it cannot be changed to, for example, store more data and / or improve / reduce PIM performance / power consumption. Such memory devices are typically customized for their applications; both their design cost and modification cost are high, and in many cases, they are "proprietary" and / or customer / manufacturer specific. In addition, since technology changes at an extremely rapid pace, these devices quickly become obsolete.
[0131] For various reasons, there is a need for improved solutions for matrix operations within a processor and / or memory. Ideally, such solutions would implement matrix operations within a memory device in a manner that minimizes the performance bottleneck of the processor-memory wall. In addition, such solutions should be flexible enough to accommodate a variety of different matrix operations and / or matrix sizes.
[0132] Exemplary memory device-
[0133] Figure 2FIG. 0 is a logic block diagram of an exemplary embodiment of a memory device 200 fabricated in accordance with various principles of the present disclosure. The memory device 200 may include a plurality of partitioned memory cell arrays 220. In some embodiments, each of the partitioned memory cell arrays 220 may be partitioned during device fabrication. In other embodiments, the partitioned memory cell arrays 220 may be partitioned dynamically (i.e., after device fabrication time). Each of the memory cell arrays 220 may include a plurality of groups, each group including a plurality of word lines, a plurality of bit lines, and a plurality of memory cells disposed at intersections of, for example, the plurality of word lines and the plurality of bit lines. Selection of the word lines may be performed by a row decoder 216, and selection of the bit lines may be performed by a column decoder 218.
[0134] The plurality of external terminals included in the semiconductor device 200 may include address terminals 260, command terminals 262, clock terminals 264, data terminals 240, and power terminals 250. An address signal and a group address signal may be supplied to the address terminals 260. The address signal and the group address signal supplied to the address terminals 260 are transmitted to an address decoder 204 via an address input circuit 202. The address decoder 204 receives, for example, the address signal and supplies a decoded row address signal to the row decoder 216 and supplies a decoded column address signal to the column decoder 218. The address decoder 204 may also receive the group address signal and supply the group address signal to the row decoder 216 and the column decoder 218.
[0135] Command signals are supplied to a command input circuit 206 at the command terminals 262. The command terminals 262 may include one or more discrete signals, e.g., row address strobe (RAS), column address strobe (CAS), read / write (R / W). The command signals input to the command terminals 262 are provided to a command decoder 208 via the command input circuit 206. The command decoder 208 may decode the command signals 262 to generate various control signals. For example, RAS may be asserted to specify the row at which data is to be read / written, and CAS may be asserted to specify the column at which data is to be read / written. In some variations, the R / W command signal determines whether the content of the data terminals 240 is written to the memory cells 220 or read therefrom.
[0136] During a read operation, read data may be output from the data terminal 240 to the outside via a read / write amplifier 222 and an input / output circuit 224. Similarly, when a write command is issued and the write command is supplied to the row address and the column address in a timely manner, a write data command may be supplied to the data terminal 240. The write data command may be supplied to a given memory cell array 220 via the input / output circuit 224 and the read / write amplifier 222 and written to the memory cells specified by the row address and the column address. According to some embodiments, the input / output circuit 224 may include an input buffer.
[0137] The clock terminal 264 may be supplied with an external clock signal for synchronous operation. In one variant, the clock signal is a single-ended signal; in other variants, the external clock signals may be complementary to each other (differential signals) and supplied to the clock input circuit 210. The clock input circuit 210 receives the external clock signal and conditions the clock signal to ensure that the resulting internal clock signal has sufficient amplitude and / or frequency for subsequent locked loop operation. The conditioned internal clock signal supplied to the feedback mechanism (internal clock generator 212) provides a stable clock for the internal memory logic. Common examples of internal clock generation logic 212 include, but are not limited to: digital or analog phase-locked loop (PLL), delay-locked loop (DLL), and / or frequency-locked loop (FLL) operation.
[0138] In an alternative variant (not shown), the memory 200 may rely on an external clock (i.e., it does not have an internal clock of its own). For example, a phase-controlled clock signal may be supplied externally to the input / output (IO) circuit 224. This external clock may be used to clock in the write data and clock out the read data. In such a variant, the IO circuit 224 supplies the clock signal to each of the corresponding logic blocks (e.g., address input circuit 202, address decoder 204, command input circuit 206, command decoder 208, etc.).
[0139] A power potential may be supplied to the power terminal 250. In some variants (not shown), these power potentials may be supplied via the input / output (I / O) circuit 224. In some embodiments, the power potentials may be isolated from the I / O circuit 224 so that power noise generated by the IO circuit 224 does not propagate to other circuit blocks. These power potentials are conditioned by an internal power circuit 230. For example, the internal power circuit 230 may generate various internal potentials, such as removing noise and / or parasitic activity, as well as step-up or step-down potentials supplied from the power potential. The internal potentials may be used for, for example, the address circuits (202, 204), command circuits (206, 208), row and column decoders (216, 218), RW amplifier 222, and / or any of a variety of other circuit blocks.
[0140] When the internal power circuit 230 can adequately supply the internal voltage for the power-on sequence, the power-on reset circuit (PON) 228 provides a power-on signal. The temperature sensor 226 may sense the temperature of the semiconductor device 200 and provide a temperature signal; the temperature of the semiconductor device 200 may affect some memory operations.
[0141] In one exemplary embodiment, the memory array 220 can be controlled via one or more configuration registers. In other words, these configuration registers are described in more detail herein for their use in selectively configuring one or more memory arrays 220 into one or more matrix configurations and / or matrix multiplication units (MMUs). In other words, the configuration registers can enable the memory cell architecture within the memory array to dynamically change, for example, its structure, operation, and functionality simultaneously. These and other variations will be apparent to those of ordinary skill in the art in the context of the present disclosure.
[0142] Figure 3 A more detailed side-by-side description of the memory array and matrix configuration circuitry is provided. Figure 3 Both the memory array and matrix configuration circuitry use the same memory cell array, where each memory cell is composed of a resistive element 302 coupled to a word line 304 and a bit line 306. In the first configuration 300, the memory array circuit is configured to operate as a row decoder 316, a column decoder 318, and a memory cell array 320. In the second configuration 350, the matrix configuration circuit is configured to operate as a row driver 317, a matrix multiplication unit (MMU) 319, and an analog crossbar configuration (matrix configuration) 321. In one exemplary embodiment, a look-up table (LUT) and associated logic 315 can be used to store and configure different matrix multiplication unit coefficient values.
[0143] In one exemplary embodiment of the present disclosure, the memory array 320 is composed of resistive random access memory (ReRAM). ReRAM is a non-volatile memory that changes the resistance of a memory cell across a dielectric solid-state material, sometimes referred to as a "memristor". Current ReRAM technology can be implemented in two-dimensional (2D) layers or three-dimensional (3D) stacks of layers; however, higher-order dimensions can be used for future iterations. The complementary metal-oxide-semiconductor (CMOS) compatibility of crossbar ReRAM technology enables the integration of both logic (data processing) and memory (storage devices) on a single chip. Crossbar ReRAM arrays can be formed in a one-transistor / one-resistor (1T1R) configuration and / or in a configuration with one transistor driving n resistive memory cells (1TNR) and other possible configurations.
[0144] Multiple inorganic and organic material systems can achieve thermal and / or ionic resistance switching. In several embodiments, such systems include: phase change chalcogenides (e.g., Ge 2 Sb 2 Te 5 、AgInSbTe and others); binary transition metal oxides (e.g., NiO, TiO 2 and others); perovskites (e.g., Sr(ZR)TrO 3, PCMO, and others); solid-state electrolytes (e.g., GeS, GeSe, SiO x , Cu 2 S, and others); organic charge transfer complexes (e.g., Cu tetracyanoquinodimethane (TCNQ), and others); organic charge acceptor systems (e.g., Al amino-dicyanoimidazole (AIDCN), and others); and / or 2D (laminated) insulating materials (e.g., hexagonal BN, and others); and other possible systems for resistive switching.
[0145] In the illustrated embodiment, the resistive element 302 is a non-linear passive two-terminal electrical component that can change its resistance based on the history of the current application (e.g., hysteresis or memory). In at least one exemplary embodiment, the resistive element 302 can form or break conductive filaments in response to the application of different polarities of current to the first terminal (connected to the word line 304) and the second terminal (connected to the bit line 306). The presence or absence of a conductive filament between the two terminals changes the conductance between the terminals. Although the present invention is presented in the context of a resistive element, those of ordinary skill in the relevant art will readily understand that the principles described herein can be implemented in any circuit characterized by a variable impedance (e.g., resistance and / or reactance). The variable impedance can be implemented by various linear and / or non-linear elements (e.g., resistors, capacitors, inductors, diodes, transistors, thyristors, etc.).
[0146] For illustrative purposes, the operation of the memory array 320 in the first configuration 300 is briefly outlined. During operation in the first configuration, a memory "write" can be achieved by applying a current to the memory cells corresponding to the rows and columns of the memory array. The row decoder 316 can selectively drive various row terminals to select a specific row of the memory array circuit 320. The column decoder 318 can selectively sense / drive various column terminals to "read" and / or "write" the corresponding memory cells, which are uniquely identified by the selected row and column (as Figure 3 emphasized by the thicker line width and blackened cells in the figure). As mentioned above, the application of the current causes the formation (or damage) of conductive filaments in the dielectric solid-state material. In one such case, the low-resistance state (on state) is used to represent logic "1", and the high-resistance state (off state) is used to represent logic "0". To switch a ReRAM cell, a first current having a specific polarity, magnitude, and duration is applied to the dielectric solid-state material. Subsequently, a memory "read" can be achieved by applying a second current to the resistive element and sensing whether the resistive element is in the on state or the off state based on the corresponding impedance. The memory read may or may not be destructive (e.g., the second current may or may not be sufficient to form or break a conductive filament).
[0147] Those of ordinary skill in the relevant art will readily understand that the foregoing discussion of the memory array 320 in the first configuration 300 is consistent with existing memory operations according to, for example, ReRAM memory technology. In contrast, the second configuration 350 uses memory cells as an analog crossbar configuration (matrix configuration) 321 to perform matrix multiplication operations. Although Figure 3 the exemplary implementation of
[0148] corresponds to a 2×4 matrix multiplication unit (MMU), other variations can be substituted for equivalent results. For example, matrices of any larger size (e.g., 3×3, 4×4, 8×8, etc.) can be implemented (subject to the precision achieved by the digital-to-analog conversion (DAC) 308 and analog-to-digital conversion (ADC) 310 components). Figure 3 In the operation of the analog crossbar configuration (matrix configuration) 321, each of the row terminals is simultaneously driven by an analog input signal, and each of the column terminals is simultaneously sensed for an analog output, which is the analog sum of the voltage potentials of the corresponding resistive elements across each row / column combination. It is worth noting that in the second configuration 350, all row and column terminals associated with matrix multiplication are active (as
[0149] emphasized by the thicker line widths and blackened cells in
[0150] ). In other words, the ReRAM crossbar configuration (matrix configuration) 321 uses the matrix configuration to perform "analog computing" that calculates vector matrix products (or scalar matrix products, matrix matrix products, etc.).
[0151] Note that Figure 4AAn illustrative example where a simple "butterfly" calculation 400 can be performed via a 2×4 matrix configuration. Although the conductance can be increased or decreased, the conductance cannot be "negative". Therefore, it may be necessary to perform subtraction in the digital domain. The butterfly operation is described in the following multiplication of matrix (M) and vector (a) (Equation 1):
[0152] Equation 1:
[0153] This simple FFT butterfly 400 of Equation 1 can be decomposed into two different matrices representing positive and negative coefficients (Equations 2 and 3):
[0154] Equation 2:
[0155] Equation 3:
[0156] Equations 2 and 3 can be implemented as analog calculations using a matrix configuration circuit. After the calculation, the resulting analog values can be converted back to the digital domain via the aforementioned ADC. Existing ALU operations can be used to perform subtraction in the digital domain (Equation 4):
[0157] Equation 4:
[0158] In other words, as illustrated in Figure 4A a 2×2 matrix can be further subdivided into a 2×2 positive matrix and a 2×2 negative matrix. The ALU can add / subtract the results of the 2×2 positive matrix and the 2×2 negative matrix to produce a single 2×2 matrix.
[0159] Techniques similar to those described in Figure 4A can be applied to matrices containing imaginary or complex numbers. For example, Figure 4B illustrates an embodiment of calculation 410 (described in Equation 5 below), where a 2×2 matrix M is multiplied by a 2×1 vector a, and where matrix M can contain real, imaginary, complex, positive, and negative numbers. In this example, it is assumed that vector a contains only real numbers.
[0160] Equation 5:
[0161] The matrix multiplication operation 410 can be performed, for example, as in Figure 4B, is implemented using a 2×8 memory fabric array by decomposing the matrix M into four separate matrices represented in the memory array 412 as four 2×2 memory subarrays: i) the first subarray 402 is programmed with real positive values of M, ii) the second subarray 404 is programmed with real negative values of M, iii) the third subarray 406 is programmed with imaginary positive values of M, and iv) the fourth subarray 408 is programmed with imaginary negative values of M. The analog value representing the vector a can be used to drive the entire 2×8 memory fabric, and the analog results of the four separate matrix-vector multiplication operations are obtained from the four separate corresponding subarrays 402, 404, 406, 408. Exemplary four matrix-vector multiplication operations (Equations 6 to 9) are described below.
[0162] Equation 6:
[0163] Equation 7:
[0164] Equation 8:
[0165] Equation 9:
[0166] The simulation results in equations 6 to 9 are obtained from four subarrays and using ADC ( Figure 4B After conversion from the analog-to-digital domain (not shown in FIG. 10 ), an arithmetic logic unit (ALU) may be used to appropriately combine the results in the digital domain (i.e., multiplying the results of the two imaginary coefficient subarrays 406 / 408 by j and -j, respectively; and multiplying the result of the real negative coefficient subarray 404 by -1; and adding all the results together) to obtain the correct final result (see Equation 10).
[0167] Equation 10:
[0168] Thus, similar to the way two 2x2 memory subarrays can be used to represent a 2x2 matrix with real (positive and negative) values (as with respect to Figure 4A Described), four 2×2 memory subarrays can be used to represent 2×2 matrices using real, imaginary, and complex (positive and negative) values. It should be noted that in an exemplary configuration of the apparatus, the matrix-vector multiplication operation above advantageously occurs in a single processing cycle.
[0169] In another embodiment, if Figure 4C As shown in , both the matrix M and the vector a in the matrix-vector multiplication operation 420 can have complex values. Figure 4B The described operation 410 can be applied separately to the real coefficients of vector a and the imaginary coefficients of vector a. The real coefficients of vector a are used to drive the memory array 412A, and the four results of the memory array 412A are digitally combined by the ALU 414A (in an operation such as relative to Figure 4B the described operation).
[0170] In a separate operation, the imaginary coefficients of vector a are used to drive the memory matrix configuration array 412B, and the results are digitally combined by another ALU 414B. In a second operation involving the memory array 412B, the ALU 414B first multiplies the results of the four sub-arrays ( Figure 4B 402, 404, 406, 408 in) by j, -j, 1, and -1 respectively. The results of the first and second ALUs can then be added together in another ALU 416 to obtain the final matrix vector product.
[0171] It should be noted that although Figure 4C the internal details of the arrays 412A and 412B are not shown in, in one embodiment, each of the arrays 412A, 412B includes four sub-arrays described with respect to the memory array 412 relative to Figure 4B . The first memory configuration unit 412A and the second memory configuration unit 412B can be advantageously used to perform the first and second operations in parallel.
[0172] Alternatively, in another exemplary embodiment, the same memory configuration unit 412 ( Figure 4B ) can be used to perform the first and second operations one after another. However, the ALU that combines the results in one configuration applies different types of calculations to the results of the two operations (e.g., multiplying the four results by 1, -1, j, and -j for the first operation, and multiplying the four results by j, -j, 1, and -1 for the second operation).
[0173] As discussed above, a memory configuration array combined with an arithmetic logic unit (ALU) can be used to compute matrix operations of various complexities or sizes, including supporting massive MIMO operations (which can be one or more orders of magnitude larger than corresponding traditional or non-massive spatial diversity calculations).
[0174] In addition, those of ordinary skill in the relevant art will readily appreciate the wide variety of variations and / or capabilities implemented by the ALU. For example, the ALU can provide arithmetic operations (e.g., addition, subtraction, addition with carry, subtraction with borrow, negation, increment, decrement, transfer, etc.), bitwise operations (e.g., AND, OR, XOR, complement), shift operations (e.g., arithmetic shift, logical shift, rotate, rotate through carry, etc.) to implement, for example, multi-precision arithmetic, complex number operations, and / or any extended MMU capabilities to any degree of precision, size, and / or complexity. As used herein, the terms "digital" and / or "logic" within the context of computing refer to processing logic that uses quantized values (e.g., "0" and "1") to represent symbolic values (e.g., "on state", "off state"). In contrast, the term "analog" within the context of computing refers to processing logic that performs computations using continuous variable aspects of physical signaling phenomena such as electrical, chemical, and / or mechanical quantities. Various embodiments of the present disclosure may represent analog input and / or output signals as continuous electrical signals. For example, a voltage potential may have different possible values (e.g., any value between a minimum voltage (0V) and a maximum voltage (1.8V), etc.). A digital-to-analog converter (DAC), an analog-to-digital converter (ADC), an arithmetic logic unit (ALU), and / or variable gain amplification / attenuation may be employed to perform a combination of analog computations using digital components.
[0175] Return reference Figure 3 , in order to configure the memory cells into the crossbar configuration (matrix configuration) 321 of the second configuration 350, each of the resistive elements can be written with a corresponding matrix coefficient value. Different from the first configuration 300, the second configuration 350 can write different degrees of impedance (representing coefficient values) into each ReRAM cell using an electric current amount selected to set the polarity, magnitude, and duration of the specific conductance. In other words, by forming / destroying conductive filaments with different conductivities, multiple different conductivity states can be established. For example, applying a first magnitude can result in a first conductance, applying a second magnitude can result in a second conductance, applying the first magnitude for a longer duration can result in a third conductance, etc. Any permutation of the foregoing write parameters can be substituted for an equivalent result. More directly, different conductances can use multiple states (e.g., three (3), four (4), eight (8), etc.) to represent one or more continuous value ranges (e.g., [0, 0.33, 0.66, 1], [0, 0.25, 0.50, 0.75, 1], [0, 0.125, 0.250, …, 1], etc.), rather than using two (2) resistance states (on state, off state) to represent two (2) digital states (logic "1", logic "0").
[0176] In one embodiment of the present disclosure, matrix coefficient values are stored in advance in a look-up table (LUT) and configured by associated control logic 315. During an initial configuration phase, matrix fabric 321 is written with matrix coefficient values from the LUT via control logic 315. Those of ordinary skill in the art will readily appreciate that certain memory technologies can also implement a write-once-use-many operation. For example, although forming (or breaking) a conductive filament in a ReRAM cell may require a specific duration, magnitude, polarity, and / or direction of current; subsequent use of the memory cell can be repeated multiple times (as long as the conductive filament is not substantially formed or broken within its useful life). In other words, subsequent use of the same matrix fabric 321 configuration can be used to share the initial configuration time.
[0177] In addition, certain memory technologies (such as ReRAM) are non-volatile. Thus, after the matrix fabric circuit is programmed, it can enter a low-power state (or even power down) to save power when not in use. In some cases, the non-volatility of the matrix fabric can be used to further improve power consumption. Specifically, unlike the prior art where matrix coefficient values can be reloaded from non-volatile memory for subsequent processing, an exemplary matrix fabric can store matrix coefficient values even when the memory device is powered down. For subsequent wake-up, the matrix fabric can be used directly.
[0178] In one exemplary embodiment, matrix coefficient values can be derived based on the nature of the matrix operation. For example, coefficients for certain matrix operations can be derived in advance based on "size" (or other structure-defining parameters) and stored in the LUT. As just two such examples, the Fast Fourier Transform (Equation 11) and the Discrete Cosine Transform (DCT) (Equation 12) are reproduced below:
[0179] Equation 11:
[0180] Equation 12:
[0181] As can be determined mathematically from the foregoing equations, the matrix coefficient values (also referred to as "twiddle factors") are determined based on the size of the transform. For example, the coefficients for an 8-point FFT are: etc. In other words, after the size of the FFT is known, the values can be empirically set (where k is 0, 1, 2, 3... 7). In fact, the coefficients for a larger FFT contain the coefficients for a smaller FFT. For example, a 64-point FFT has 64 coefficient values, which include all 32 coefficients used in a 32-point FFT, and all 16 coefficients used in a 16-point FFT, etc. More directly, a single LUT can contain all the coefficients to support any number of different transforms.
[0182] In another exemplary embodiment, matrix coefficient values may be stored in advance. For example, the coefficients for certain matrix multiplication operations may be known or otherwise defined by, for example, an application or a user. For example, image processing computations, such as those described in the co-owned and co-pending U.S. Patent Application No. 16 / 002,644, filed on June 7, 2018, and titled "Image Processor Formed in a Memory Cell Array," incorporated hereinabove by reference, may define various different matrix coefficient values in order to affect, for example, defect correction, color interpolation, white balance, color adjustment, gamma luminance, contrast adjustment, color conversion, downsampling, and / or other image signal processing operations.
[0183] In another example, the coefficients for certain matrix multiplication operations may be determined or otherwise defined by, for example, user considerations, environmental considerations, other devices, and / or other network entities. For example, wireless devices typically experience different multipath effects that can interfere with operations. Various embodiments of the present disclosure determine the multipath effects and use matrix multiplication to correct them. In some cases, a wireless device may calculate each of the independent different channel effects based on the degradation of known signaling. The difference between the desired reference channel signal and the actual reference channel signal may be used to determine the noise effects it experiences (e.g., attenuation, reflection, scattering, and / or other noise effects in a specific frequency range).
[0184] In other embodiments, a wireless device may be instructed to use a predetermined "codebook" of beamforming / precoding configurations. The codebook of beamforming / precoding coefficients may be less precise but may be preferred for other reasons (e.g., speed, simplicity, etc.).
[0185] As previously discussed, diversity techniques (commonly implemented as multiple-input / multiple-output (MIMO) or MU-MIMO systems) use, for example, multiple transmit and multiple receive antennas to convey data associated with the same wireless channel (and, in the case of MU-MIMO, for multiple users). Multiple antenna systems can improve the reliability and gain of signals by simultaneously transmitting the same signal / symbol via several diverse channel streams (signal diversity), thus ensuring that the signal is properly conveyed. Additionally, MIMO systems can simultaneously transmit different parts of a signal (e.g., consecutive symbols of one data stream) or different signals via several different channel streams (via spatial multiplexing), thus increasing the rate of data communication. The demand for high data rates and good channel quality has made MIMO essential for current broadband wireless communication.
[0186] Large-scale or massive MIMO systems use a very large number of antennas (e.g., hundreds or even thousands) all of which require coordination in order to obtain the best possible channel quality and throughput.
[0187] Traditional (MIMO and MU-MIMO) and massive MIMO systems require various matrix multiplication operations using matrices on the order of the number of receive and transmit antennas. This is because massive MIMO systems use the large number of antennas (and thus are represented by large matrices), which are particularly vulnerable to the "processor-memory wall" performance limitations described previously herein.
[0188] According to the present disclosure, some of the matrix operations that can be performed within a working MIMO system and examples of memory bank architectures that can be implemented to perform these operations are discussed below. However, those of ordinary skill in the art will recognize, in view of the present disclosure, that the following examples are in no way to be considered limiting, since the principles, apparatus, and methods of the present disclosure can be extended to other types of operations (matrix or otherwise) and applications.
[0189] Incidentally, the channel in a exemplary MIMO system with N transmit antennas and M receive antennas can be expressed as an M×N channel matrix H, where the M rows represent the receive antennas, the N columns represent the transmit antennas, and the individual H coefficients h ij represent the individual radio signal paths. The channel characteristics can be estimated at the receiver by, for example, transmitting known training / pilot signals from each of the transmitters, noting the signals received by each of the receivers, and then calculating all the signal paths based on the known and received signals.
[0190] The transmitter can be a user device (e.g., a cellular phone), a wireless network base station (e.g., an eNB or gNB or a femtocell / HNB, or a Wi-Fi AP), a wireless node, or any other device that uses multiple transmit antennas to transmit signals. The receiver can similarly be a user device, a network base station, or any device that uses multiple receive antennas to receive signals. It should be understood that although described in terms of an M×N channel matrix of multiple transceiver antenna elements, the exemplary embodiments described herein can be readily extended to other applications, such as where multiple transmit antenna elements (N) are used with a single receive antenna on a user device or UE (including the case of serving multiple user devices), and multiple individual transmitters are combined with a single receiver having multiple (M) receive antenna elements.
[0191] At a single point in time, a MIMO system can be characterized as: b = Ha (ignoring noise), where the transmitted signal is represented by the vector a, the received signal is represented by the vector b, and the channel is represented by the matrix H. To successfully communicate data within a MIMO system, the receiver needs to be able to use the received signal b to derive the transmitted signal a (or the data conveyed by the signal under consideration). In some restricted cases, this can be done by simply inverting the channel matrix H to obtain H -1 , and simply multiplying the received vector by H -1This is achieved in other embodiments by applying a pre - decoding matrix to the data at the transmitter based on the channel matrix to obtain an appropriate transmitted signal. The receiver can then perform decoding or data estimation / recovery matrix operations to obtain / estimate the originally transmitted data.
[0192] Figure 5A Describe an embodiment of a system 550 having a transmitter 502 in communication with a receiver 504, the transmitter 502 having two antennas 502 - 1, 502 - 2, and the receiver 504 having two receiving antennas 504 - 1, 504 - 2. The transmitter uses its two antennas to send data in the form of a transmit vector a=(a1,a2), as Figure 5A shown. Signals from the two antennas pass through a channel represented by a matrix H and are received as symbols b1 and b2 by the two receiver antennas 504 - 1, 504 - 2, representing the values of a size - 2 receive vector b. The four different paths taken by the signals are represented by the four coefficients of the matrix H. The received vector equation for system 550 is shown below (Equation 13) as (b = Ha).
[0193] Equation 13:
[0194] To enable the receiver to obtain the communicated data, the received vector equation can be inverted, and a matrix - vector multiplication operation (Equation 14) can be used to calculate the transmit vector a. This matrix operation can be performed at the receiver using, for example, the memory matrix organization devices and techniques described above.
[0195] Equation 14: a = H -1 b
[0196] Figures 5B to 5D Describe an exemplary embodiment of a matrix organization and MMU circuit 500 that can be used to perform the matrix operation of Equation 14 on the received vector b to obtain the transmit vector a. This is described in more detail in the following Figure 5C and 5D description of the matrix organization 521. Figure 5B The matrix organization 521 is described in more detail in the following description.
[0197] In the memory system 500, the matrix organization 521 is programmed using control logic 515 having the values of the H -1 matrix. It should be noted that, as described above, the receiver can in some cases be configured to use training signals to determine the value of the channel matrix H and then calculate the inverse of the matrix H using known mathematical techniques. After the inverse matrix H -1 is calculated, it can be stored in a look - up table (LUT) within the memory array 522, which can be accessed by the control logic 515.
[0198] AsFigure 5B shows that the matrix operation of Equation 14 is performed by using the memory structure 521 of the vector b drive system 500. That is, the individual values of vector b are converted into analog voltages / currents by DAC 508, which are then input to H -1 Individual rows (lines) of programmed memory matrix configuration 521.
[0199] The result of the memory fabric 521 is then output as an analog voltage or current into the MMU 519, which converts the result into a digital value via its ADC 510. The digital value is then used by the MMU's arithmetic logic unit (ALU) 512 to obtain the entire value of vector a.
[0200] Figure 5C and 5D Description Figure 5B Details of two specific implementations 500A, 500B of the memory device 500 shown in . In some configurations, a 2×2H -1 The matrix may be represented within the memory matrix fabric 521 as a 2×8 memory array having four 2×2 memory sub-arrays 521A, 521B, 521C, 521D, each representing a different aspect of the matrix (as discussed earlier with respect to Figure 4B Described).
[0201] Figure 5C System 500A is illustrated in which a single 2×8 memory array within matrix fabric 521 is driven by two inputs (representing two values of vector b). To accommodate the possible complex values of vector b, Figure 5C The matrix fabric 521 in is driven successively in one approach (i) first by the real values of vector b and (ii) then by the imaginary values of vector b. The first result is temporarily stored in the memory array 522 and then appropriately combined with the second result using the ALU 519. The ALU 519 processes the result of the memory fabric 521 differently depending on whether the memory fabric 521 is driven by the real or imaginary values of vector b (i.e., multiplying the result from submatrix 521A by 1 or j). This is achieved by, for example, the control logic 515 reconfiguring the MMU between passes or runs so that the ALU 519 has the appropriate logic.
[0202] Figure 5D Another embodiment of system 500B is described, in which memory fabric 521 includes two identical 2×8 memory fabric arrays 521-1 and 521-2. The first 2×8 memory array 521-1 is driven by the real values of vector b, and in parallel, the second 2×8 memory array 521-2 is driven by the imaginary values of vector b. In Figure 5DIn an embodiment, the two parallel memory fabric portions 521-1, 521-2 each have their own row drivers. The control logic 515 (i) inputs the real values of the vector b into the first row driver, and (ii) inputs the imaginary values of the vector b into the second row driver. After the two portions of the memory fabric 521 have been driven by the values of the vector b, the results of the two memory arrays can be appropriately combined within a single MMU 519.
[0203] Figure 5D An embodiment includes an MMU 519 having one ADC 510 and three ALUs 512-1, 512-2, 514, but this configuration is illustrative only and not restrictive.
[0204] After the analog signals received from the memory arrays 521-1 and 521-2 are converted to digital values by the ADC 510, four digital values corresponding to the result of the first memory array 521-1 are input into the first ALU 512-1, and four digital values corresponding to the result of the second memory array 521-2 are input into the second ALU 512-2. The two ALUs respectively combine their input values and transfer their results into the third ALU 514, which performs a simple addition operation. It should be noted that, as explained with respect to Figure 4C The first ALU 512-1 and the second ALU 512-2 can be programmed with slightly different logic (e.g., the results of the two left sub-arrays of the first memory array 521-1 are designated as real numbers by the first ALU 512-1, but the results of the two left sub-arrays of the second memory array 521-2 are designated as imaginary numbers by the second ALU 512-2).
[0205] In another exemplary embodiment, the MMU 519 can have a single ALU that obtains all eight results of the two memory fabric arrays 521-1, 521-2 and appropriately combines the results. In other words, the logic of the three ALUs described above can be combined within a single unit.
[0206] In yet another embodiment, the MMU 519 can be replaced by two separate MMUs. That is, the two memory arrays 521-1 output the results into corresponding MMUs, and the corresponding MMUs each have an ADC and an ALU. The results of the two different MMUs can be combined in a third ALU outside of any MMU.
[0207] Figure 6ADescribe an exemplary system 600 with this higher degree of complexity. As shown, system 600 includes a transmitter 602 having N antennas (602-1 to 602-N), and a receiver having M antennas (604-1 to 604-M), where N does not have to be equal to M. The receiver 604 uses training signals to estimate the channel matrix H. In massive MIMO (see, for example Figure 6D ), the number of transmitter and receiver antennas (N and M) can be in the hundreds or thousands (including distribution over multiple discrete devices 606-1... 606-n), and correspondingly, the channel matrix H can have a similarly large size.
[0208] In the case where the matrix H is non-invertible and large, it can be decomposed using singular value decomposition (SVD), using methods well known in the art, and rewritten as: H = UΣV*, where U and V* are the left and right unitary matrices of H, respectively, and Σ is a diagonal matrix of the singular values of H. The received signal vector b (ignoring noise) can be written as: b = Ha = UΣV*a. The data d to be transmitted (a matrix of data symbols) can be transformed into the transmitted signal using Equation 15 (which can be regarded as a type of pre-coding).
[0209] Equation 15: a = Ud
[0210] At the receiver, the data can be decoded / recovered (estimated) from the received signal b using Equation 16. In Figure 6A embodiments, after calculating the channel matrix, the receiver transmits the channel information to the transmitter. The transmitter 602 can use the received channel information to perform a pre-coding operation 606 (Equation 15) to transform the data d into the transmit vector a. The individual values of each vector a are transmitted simultaneously using the individual transmit antennas (602-1 to 602-N). The receiver 604 obtains the received vector b, and uses receiver operations 608 to estimate the original data (Equation 16).
[0211] Equation 16: d = V*b
[0212] Figure 6B Describe an embodiment of a memory matrix organization system and MMU circuit 626 that can be used at the transmitter to perform the pre-coding matrix operation of Equation 15 as described with respect to Figure 6A . In one of its configurations, the control logic 615 obtains the pre-coding matrix U, and programs the matrix memory organization 621 using the values of the matrix U. The memory array (not shown) composed of individual sub-arrays within the matrix organization 621 can be programmed with the positive real coefficients, negative real coefficients, positive imaginary coefficients, and negative imaginary coefficients of the matrix U (using the same as Figure 5CSimilar techniques). In one embodiment, matrix fabric 621 may include two memory arrays that are identically programmed with the values of matrix U, each capable of being driven in parallel with each other (similar to Figure 5D configuration). The values of pre-decoded matrix U are obtained as needed, for example, from a data source external to transmitter memory system 626.
[0213] As Figure 6B shown, individual vectors d (e.g., in-phase and quadrature samples) representing the data to be conveyed are obtained from the row driver by control logic 615 and fed into the row driver. Although vector d is not permanently stored inside memory system 626 (since the data is constantly changing), a buffer of the data may be stored inside memory array 622 to avoid frequent use of the memory-processor interface (which may have a significant penalty in terms of access latency, for example).
[0214] In one embodiment, control logic 615 feeds the real values of each vector d into row driver 617, followed by the imaginary values of the vector (similar to Figure 5C structure). In another embodiment, row driver 617 includes two row drivers connected to the respective memory arrays of memory fabric 621, and control logic feeds the real values and the imaginary values of each vector d in the first row driver into the second row driver (similar to Figure 5D structure).
[0215] Row driver 617 (or drivers) uses DAC 608 to convert the individual digital values of vector d into analog input signals. These converted analog input signals are then used to simultaneously drive all row terminals of matrix fabric 621, which correspond to the memory cells programmed with the values of matrix U. The matrix fabric result is output to MMU 619, which converts the result into digital values via its ADC 610. The digital values are then used by the arithmetic logic unit (ALU) 612 of MMU to obtain the individual values of output vector a.
[0216] Figure 6C illustrates what can be used at the receiver to, as with respect to Figure 6AAn example of a memory matrix organization system and an MMU circuit 650 that perform the received matrix operation (data decoding / estimation) of Equation 16 as described. The control logic 635 in this system 650 obtains the received matrix V*, and uses its value for the program memory organization 641. The control logic 635 then uses the row driver 637 (or two row drivers) to convert the real and imaginary values of the received data vector b into analog signals, and propagates those signals through the programmed matrix organization 641. The matrix organization outputs the analog results corresponding to the part of the data vector d into the MMU 639, and the MMU 639 uses its ADC 630 and ALU 632 to convert and appropriately combine the signals to finally obtain the value of the vector d. It should be noted that this vector d calculated at the receiver is an estimated value of the original data and not an exact copy. Nevertheless, the memory system 650 uses the values of the individual received signals (vector b) to determine the data communicated by the transmitter in a limited number of (processing cycles). Therefore, after the memory organization 641 has been programmed, the communicated data can be decoded or recovered very quickly from the fast stream of received signals.
[0217] Figure 6B and 6C The implementation details of the memory device structure described in may be similar to the structure described above with respect to Figure 5C or 5D.
[0218] Figures 6A to 6C A MIMO system is illustrated where the outgoing signal can be pre-coded using a pre-coding matrix U derived from a known channel matrix H. However, since the channel matrix H is calculated at the receiver, this requires the complete H matrix to be transmitted from the receiver to the transmitter. The transmission of the H matrix can result in a large overhead. To avoid this overhead, modern wireless communication systems use a codebook of predefined pre-coding matrices W, which can be pre-computed for various different antenna configurations and can be easily accessed (e.g., stored in) by all the transmitting and receiving units within a telecommunication network.
[0219] Figure 6D Another system 670 is illustrated, and the methods and devices of the present disclosure (including Figures 6A to 6C those methods and devices) can be used with system 670, where a 5G NR gNB (e.g., DU part) 672 utilizes a plurality of transmit antenna elements 602-1 to 602-n, and transmits signals to a plurality of receivers 606-1 to 606-n (UE 1 to UE n)。In such cases, the channel matrix can become extremely large (including levels such as massive MIMO (mMIMO) as previously described). For example, in various millimeter-wave and sub-6 GHz applications (such as supporting 64 antenna elements in 3GPP Release 15), mMIMO plays an increasingly important role in supporting high-throughput, low-latency communication, and as the number of supported antenna elements increases, the channel matrix will correspondingly become larger. Therefore, Figure 6D the value of "n" in
[0220] Figure 7A Another embodiment of the MIMO system 700 (substantially similar to Figure 6A system 600) is described, which has a transmitter 702 and a receiver. The transmitter 702 has N antennas (702-1 to 702-N), and the receiver has M antennas (704-1 to 704-M), where N does not have to be equal to M. However, the receiver in system 700 does not send the calculated channel matrix H to the transmitter 702. In fact, based on the calculated channel matrix H, the receiver selects a precoding matrix W from a codebook and sends the index corresponding to the precoding matrix in the codebook to the transmitter. The transmitter uses the precoding matrix W to transform the data symbol d into a transmitted signal corresponding to the vector a (Equation 17 below), rather than using the precoding matrix U. The rest of the operation of the system is generally similar to Figure 6A the operation of system 600 described in
[0221] Equation 17: a = Wd
[0222] Figure 7B An embodiment of the memory organization system and MMU circuit 709 that can be used at the transmitter to perform the precoding matrix operation of Equation 17 is described.
[0223] In Figure 7B the embodiment, the control logic 715 obtains the precoding matrix W and uses the values of the matrix W to program the matrix memory organization 721 (including two identical memory arrays 721-1, 721-2, each having four memory sub-arrays 721A, 721B, 721C, 721D). In one embodiment, the values of the precoding matrix W can be input into the memory system 709 from an external source as needed. In another embodiment, the codebook of different precoding matrices is located inside a lookup table (LUT) in the memory system 709 (e.g., in the memory array 722). In the latter case, the control logic 715 can receive the index or address of a specific precoding matrix W from the processor and extract the value of the precoding matrix W from the memory array 722.
[0224] As described, with the memory configuration 721 programmed with the values of the pre-coding matrix W, the control logic 715 inputs the real and imaginary coefficients of the vector d into the first row driver 717-1 and the second row driver 717-2, respectively. The row drivers use their DACs 708 to convert their given values into analog signals and use the analog signals to drive the input terminals of their corresponding memory configuration arrays 721-1, 721-2. The memory configurations 721-1, 721-2 output the results in the form of analog signals to the MMU 719. In another variant, the memory configuration arrays 721-1, 721-2 output the results to their respective MMUs.
[0225] The MMU 719 transforms the analog signals into the digital domain via the ADC 710 and uses the ALUs 712-1, 712-2, 714 to combine the digital values to obtain the value of the vector a.
[0226] In another variant, the MMU 719 has a single ALU that performs the combination. It should be noted that the memory system 709 has a structure similar to the structure described with respect to Figure 5D However, Figure 5D the memory structure described in
[0227] Figures 5A to 7B or its variants can be similarly applied to the pre-coding of data signals using the pre-coding matrix W.
[0228] However, it will also be appreciated that the ability to perform matrix operations using only one or a few atomic operations regardless of the size of the one or more matrices involved (e.g., even if the operations do not scale exponentially with matrix size) can be particularly applicable to techniques implementing large-scale MIMO, which can potentially involve a very large number of receive and / or transmit antennas (and thus result in matrices of very large size, where the benefits of atomic operations are thus advantageously multiplied). As previously mentioned, in some embodiments, the matrix coefficient values may be pre-stored in a look-up table (LUT) and configured by associated control logic. In one exemplary embodiment, the matrix configuration may be configured via dedicated hardware logic. While the present disclosure is presented in the context of internal control logic, external implementations may substitute equivalent results. For example, in other embodiments, the logic includes an in-memory processor (PIM), which may set the matrix coefficient values in a series of reads and writes based on the LUT values. In still other instances, for example, an external processor may perform the LUT and / or logic functions.
[0229] Figure 8A is a logic block diagram of an exemplary implementation of a processor-memory architecture 800 according to various principles described herein. As Figure 8A shown, a processor 802 is coupled to a memory 804; the memory includes a look-up table (LUT) 806, control logic 808, a matrix configuration, and a corresponding matrix multiplication unit (MMU) 810, and a memory array 812.
[0230] In one embodiment, the LUT 806 stores a plurality of matrix value coefficients, dimensions, and / or other parameters associated with different matrix operations. In one exemplary embodiment, the matrix coefficients are pre-coded. The LUT 806 may include various channel matrix (or pre-coded matrix) codebooks that may be predefined and / or empirically determined based on radio channel measurements.
[0231] In one embodiment, the control logic 808 controls the operations of the matrix configuration and the MMU 810 based on instructions received from the processor 802. In one exemplary embodiment, the control logic 808 may form / destroy conductive filaments of different conductivities in each of the memory cells of the matrix configuration according to the aforementioned matrix dimensions and / or matrix value coefficients provided by the LUT 806. Additionally, the control logic 808 may configure the corresponding MMU to perform any additional arithmetic and / or logical manipulations of the matrix configuration. Further, the control logic 808 may select one or more digital vectors to drive the matrix configuration and select one or more digital vectors to store the logical output of the MMU.
[0232] In the field of processing, "instructions" typically contain different types of "instruction syllables": for example, operation codes, operands, and / or other associated data structures (e.g., registers, scalars, vectors).
[0233] As used herein, the term "operation code" (opcode) refers to an instruction that can be interpreted by processor logic, memory logic, or other logic circuits to implement an operation. More directly, the opcode identifies the operation to be performed on one or more operands (inputs) to produce one or more results (outputs). Both the operands and the results can be implemented as data structures. Common examples of data structures include, but are not limited to: scalars, vectors, arrays, lists, records, joints, objects, graphs, trees, and / or any number of other forms of data. Some data structures can contain reference data (data that "points to" other data) either wholly or in part. Common examples of reference data structures include, for example, pointers, indexes, and / or descriptors.
[0234] In one exemplary embodiment, the opcode can identify one or more of the following: matrix operations, the dimensions of matrix operations, and / or the rows and / or columns of memory cells. In one such variant, the operands are decoded identifiers that specify one or more digital vectors to be operated on. For example, an instruction to perform a pre-decoding operation on an incoming data vector using a specific pre-decoded matrix W and store the result in an output digital vector might contain the opcode and operands: PRECODE($input, $output, $pmi), where: PRECODE identifies the natural operation, $input identifies the base address of the input digital vector, $output identifies the base address of the output digital vector, and $PMI identifies the pre-decoding matrix index corresponding to the base address of the pre-decoding matrix used in the operation. In another such example, a 64-point FFT can be broken into two different atomic operations: for example, PRECODE($address, $row, $col), which converts a memory array into rows by column matrix organization at $address, and MULT($address, $input, $output), which stores the vector matrix product and matrix organization of $input from $address to $output.
[0235] Similar logic applies to, for example, the DECODE operation used to support wireless channel diversity processing or other such applications.
[0236] Although Figure 8ADescribe an instruction interface that is functionally separate and different from the input / output (I / O) memory interface, but this is by no means necessary, and various configurations can be used in accordance with the present disclosure. For example, in one such embodiment, the instruction interface can be physically different (e.g., having different pins and / or connectivity). In other embodiments, the instruction interface can be multiplexed with the I / O memory interface (e.g., sharing the same control signaling and address and / or data buses in different communication modes). In still other embodiments, the instruction interface can be virtually accessible via the I / O memory interface (e.g., as registers located within an address space addressable via the I / O interface). Additionally, all or part of the memory device 804 or its logic can be integrated with the processor 802, for example, via an SoC configuration. Given the content of the present disclosure, additional other variations can be introduced by those of ordinary skill in the art.
[0237] In one embodiment, the matrix fabric and the MMU 810 are tightly coupled to the memory array 812 to read and write digital vectors (operands). In one exemplary embodiment, the operands are identified for dedicated data transfer hardware (e.g., direct memory access (DMA)) into and out of the matrix fabric and the MMU 810. In one exemplary variation, the digital vectors of data can have any size and are not limited by the processor word length. For example, the operands can specify N-bit (e.g., 2, 4, 8, 16,... etc.) operands. In other embodiments, the DMA logic 808 may use an existing memory row / column bus interface to read / write the matrix fabric 810. In still other embodiments, the DMA logic 808 can use the existing address / data and read / write control signaling within the internal memory interface to read / write the matrix fabric 810.
[0238] Figure 8B Describe a set of exemplary matrix operations 850 within the context of the exemplary embodiment 800 described in Figure 8A As shown therein, the processor 802 writes instructions to the memory 804 via the interface 807, the instructions specifying an opcode (e.g., characterized by matrix M x,y ), and operands (e.g., digital vectors a, b).
[0239] The control logic 808 determines whether the matrix fabric and / or the matrix multiplication unit (MMU) should be configured / reconfigured. For example, portions of the memory array are converted into one or more matrix fabrics and weighted using the associated matrix coefficient values defined by matrix M x,y . The digital-to-analog (DAC) row drivers and analog-to-digital (ADC) sense amplifiers associated with the matrix fabric may need to be adjusted for dynamic range and / or amplification. Additionally, one or more MMU ALU components can be coupled to one or more matrix fabrics.
[0240] When the matrix fabric and / or matrix multiplication unit (MMU) is appropriately configured, the input operand a is read via a digital-to-analog (DAC) and applied to the matrix fabric M x,y for analog computing. The analog results may alternatively be converted using an analog-to-digital (ADC) conversion for subsequent logical manipulation by the MMU ALU. The output is written to the output operand b.
[0241] Figure 8C Illustrates a set of alternative exemplary matrix operations 860 within the context of the exemplary embodiment 800 described in Figure 8A . In contrast to the diagrams of Figure 8B , the system of Figure 8C uses explicit instructions to convert a memory array into a matrix fabric. Providing a higher degree of atomicity in the instruction behavior enables various related benefits, including, for example, pipeline design and / or reduced instruction set complexity.
[0242] More directly, when the matrix fabric contains the appropriate matrix value coefficients M x,y , matrix operations can be effectively repeated. For example, after weighting the matrix fabric with the coefficients of the pre-coded matrix W, matrix-vector multiplication pre-coding operations can be quickly performed on multiple data vectors arranged in a column (illustrated in Figure 7B and Equation 17).
[0243] In another example, image processing computations, such as those described in the commonly owned and co-pending U.S. Patent Application No. 16 / 002,644, filed June 7, 2018, and titled "Image Processor Formed in a Memory Cell Array," incorporated above, may configure multiple matrix fabrics and MMU processing elements for pipelining, such as defect correction, color interpolation, white balance, color adjustment, gamma brightness, contrast adjustment, color conversion, downsampling, and / or other image signal processing operations. Each of the pipeline stages can be configured once and repeatedly used for each pixel (or pixel group) of the image. For example, the white balance pipeline stage can operate on each pixel of the data using the same matrix fabric with matrix coefficient values set for white balance; the color adjustment pipeline stage can operate on each pixel of the data using the same matrix fabric with matrix coefficient values set for color adjustment, etc. In another such example, the first stage of a 64-point FFT can be handled in thirty-two (32) atomic MMU computations (thirty-two (32) 2-point FFTs) using the same FFT "twiddle factors" (described above).
[0244] In addition, those of ordinary skill in the relevant art will further understand that some matrix configurations may have additional diversity and / or be used beyond their initial configuration. For example, as described above, a 64-point FFT has 64 coefficient values, which include all 32 coefficients used in a 32-point FFT. Thus, a matrix configuration configured for 64-point operations can be reused for 32-point operations, where the 32-point input operand a is appropriately applied to the appropriate rows of the 64-point FFT matrix configuration. Similarly, the FFT rotation factors are a superset of the discrete cosine transform (DCT) rotation factors, and thus, the FFT matrix configuration can also be used (with appropriate application of the input operand a) to compute DCT results.
[0245] In view of the present disclosure, additional arrangements and / or variations of the foregoing examples will be apparent to those of ordinary skill in the relevant art.
[0246] Method-
[0247] Now referring Figure 9 to, a logical flow diagram of an exemplary method 900 of converting a memory array into a matrix configuration for matrix transformation and performing matrix operations therein is presented.
[0248] At step 902 of method 900, the memory device receives one or more instructions. In one embodiment, the memory device receives instructions from a processor. In one such variation, the processor is an application processor (AP) commonly used in consumer electronic devices. In other such variations, the processor is a baseband processor (BB) commonly used in wireless devices.
[0249] Incidentally, a so-called "application processor" is a processor configured to execute an operating system (OS) and one or more applications, firmware, and / or software. The term "operating system" refers to software that controls and manages access to hardware. The OS typically supports processing functions such as task scheduling, application execution, input and output management, memory management, security, and peripheral access.
[0250] A so-called "baseband processor" is a processor configured to communicate with a wireless network via a communication protocol stack. The term "communication protocol stack" refers to software and hardware components that control and manage access to wireless network resources. The communication protocol stack typically includes, but is not limited to: physical layer protocols, data link layer protocols, media access control protocols, network and / or transport protocols, and the like.
[0251] Other peripheral and / or co-processor configurations may be similarly substituted for equivalent results. For example, server devices typically include multiple processors that share common memory resources. Similarly, many common device architectures pair a general-purpose processor with a dedicated co-processor and shared memory resources (such as a graphics engine or a digital signal processor (DSP)). Common examples of such processors include, but are not limited to: graphics processing unit (GPU), video processing unit (VPU), tensor processing unit (TPU), neural network processing unit (NPU), digital signal processor (DSP), image signal processor (ISP). In other embodiments, the memory device receives instructions from an application specific integrated circuit (ASIC) or other forms of processing logic, such as a field programmable gate array (FPGA), programmable logic device (PLD), camera sensor, wireless baseband processor, millimeter wave processor, or modem, and / or a media codec (e.g., image, video, audio, and / or any combination thereof).
[0252] In one exemplary embodiment, the memory device is a resistive random access memory (ReRAM) arranged in a "crossbar" row-column configuration. Although the various embodiments described herein assume a specific memory technology and a specific memory structure, those of ordinary skill in the relevant art will readily appreciate, in view of the present disclosure, that the principles described herein can be extended widely to other technologies and / or structures. For example, certain programmable logic structures (e.g., commonly used in field programmable gate arrays (FPGA) and programmable logic devices (PLD)) may have characteristics similar to those of memories in terms of capabilities and topologies. Similarly, certain processors and / or other memory technologies may change resistance, capacitance, and / or inductance; in such cases, different impedance characteristics can be used to perform analog computations. Additionally, although the "crossbar"-based construction provides a physical structure that is well-suited for two-dimensional (2D) matrix structures, other topologies may be well-suited for higher-order mathematical operations (e.g., matrix-matrix products via three-dimensional (3D) memory stacks, etc.).
[0253] In one exemplary embodiment, the memory device further includes a controller. The controller receives one or more instructions and parses each instruction into one or more instruction components (commonly also referred to as "instruction syllables"). In one exemplary embodiment, an instruction syllable includes at least one opcode and one or more operands. For example, an instruction may be parsed into an opcode, a first source operand, and a destination operand. Other common examples of instruction components may include, but are not limited to: a second source operand (for binary operations), a shift amount, an absolute / relative address, a register (or other reference to a data structure), the closest data structure (i.e., the data structure provided within the instruction itself), a subordinate function, and / or a branch / link value (e.g., executed depending on whether the instruction completes or fails).
[0254] In one embodiment, each received instruction corresponds to an atomic memory controller operation. As used herein, an "atomic" instruction is an instruction that is completed within a single access cycle. Conversely, a "non-atomic" instruction is an instruction that may or may not be completed within a single access cycle. Although a non-atomic instruction may be completed in a single cycle, it must be treated as non-atomic to prevent data race conditions. A race condition occurs when data being accessed (read or written) by a processor instruction can be accessed by another processor instruction before the first processor instruction has had a chance to complete; a race condition can unpredictably result in data read / write errors. In other words, an atomic instruction guarantees that data cannot be observed in an incomplete state.
[0255] In an exemplary embodiment, an atomic instruction may identify a portion of a memory array that is to be transformed into a matrix configuration. In some cases, an atomic instruction may identify properties of the matrix configuration. For example, an atomic instruction may identify a portion of the memory array based on, e.g., a location within the memory array (e.g., via an offset, row, column), size (number of rows, number of columns, and / or other dimensional parameters), granularity (e.g., precision and / or sensitivity). Notably, an atomic instruction may provide very fine-grained control via memory device operations; this can be desirable where memory device operations can be optimized in view of various application-specific considerations.
[0256] In other embodiments, a non-atomic instruction may specify a portion of a memory array that is to be transformed into a matrix configuration. For example, a non-atomic instruction may specify various requirements and / or constraints regarding the matrix configuration. The memory controller may internally allocate resources to accommodate the requirements and / or constraints. In some cases, the memory controller may additionally increase and / or decrease the priority of an instruction based on current memory usage, memory resources, controller bandwidth, and / or other considerations. Such embodiments may be particularly useful where memory device management would otherwise impose an unnecessary burden on the processor.
[0257] In one embodiment, an instruction specifies a matrix operation. In one such variant, the matrix operation may be a vector-matrix product. In another variant, the matrix operation may be a matrix-matrix product. In view of the present disclosure, additional other variants may be substituted by one of ordinary skill in the relevant art. Such variants may include, e.g., a scalar-matrix product, a higher-order matrix product, and / or other transformations including, e.g., linear shift, rotation, reflection, and translation.
[0258] As used herein, terms such as "transformation (transform / transformation)" refer to a mathematical operation that converts an input from a first domain to a second domain. The transformation can be "injective" (each element of the first domain has a unique element in the second domain), "surjective" (each element of the second domain has a unique element in the first domain), or "bijective" (a unique one-to-one mapping of elements from the first domain to the second domain).
[0259] More complex mathematically defined transformations that are often used in the computing field include Fourier transforms (and their derivatives, such as the discrete cosine transform (DCT)), Hilbert transforms, Laplace transforms, and Legendre transforms. In an exemplary embodiment of the present disclosure, the matrix coefficient values for a mathematically defined transformation can be pre-computed and stored in a look-up table (LUT) or other data structure. For example, the twiddle factors for a fast Fourier transform (FFT) and / or DCT can be computed and stored in a LUT. In other embodiments, the matrix coefficient values for a mathematically defined transformation can be computed by a memory controller during (or in preparation for) a matrix configuration transformation process.
[0260] Other transformations may not be based on their mathematical definitions, but can actually be defined based on, for example, an application, another device, and / or a network entity. Such transformations can be commonly used in diversity calculations, encryption, decryption, geometric modeling, mathematical modeling, neural networks, network management, and / or other graph theory based on applications. For example, a wireless network can use a codebook of predetermined antenna weighting matrices to signal the most commonly used beamforming / precoding configurations. In other instances, certain types of encryption can achieve agreement and / or negotiation between different encryption matrices. In such embodiments, the codebook or matrix coefficient values can be pre-agreed, exchanged out-of-band, in-band, or even arbitrarily determined or negotiated.
[0261] In view of the content of the present disclosure, empirically determined transformations can also be substituted for equivalent results. For example, empirically derived transformations that are often used in the computing field include wireless channel decoding, image signal processing, and / or environmental effects of other mathematical modeling. For example, a multipath wireless environment can be characterized by measuring the channel effects of, for example, a reference signal. The resulting channel matrix can be used to constructively interfere with signal reception (e.g., increase signal strength) while destructively interfering with interference (e.g., reduce noise). Similarly, an image with a skewed color tone can be evaluated for overall color balance and corrected mathematically. In some cases, an image can be intentionally skewed based on, for example, user input in order to impose an aesthetic "warm feeling" on the image.
[0262] Various embodiments of the present disclosure may implement "unary" operations within a memory device. Other embodiments may implement "binary" or even higher-order "N-ary" matrix operations. As used herein, the terms "unary," "binary," and "N-ary" refer to operations that respectively take one, two, or N input data structures. In some embodiments, binary and / or N-ary operations may be broken down into one or more unary matrix in-place operators. As used herein, an "in-place" operator is one in which the matrix operation that stores or translates it produces its own state (e.g., its own matrix coefficient values). For example, a binary operation may be decomposed into two (2) unary operations; a first in-place unary operation is performed (the result is stored "in-place"). Subsequently, a second unary operation may be performed on the matrix configuration to obtain a binary result (e.g., a multiply-accumulate operation).
[0263] Still other embodiments may serialize and / or parallelize matrix operations based on various considerations. For example, sequentially-dependent operations may be performed in a "serial" pipeline. For example, image processing computations, such as those described in the commonly-owned and co-pending U.S. Patent Application No. 16 / 002,644, filed on June 7, 2018, and entitled "Image Processor Formed in a Memory Cell Array," incorporated hereinabove by reference, may configure multiple matrix configurations and MMU processing elements for pipelining, e.g., defect correction, color interpolation, white balance, color adjustment, gamma brightness, contrast adjustment, color conversion, downsampling, etc. Pipelining may generally produce extremely high throughput data using minimal matrix configuration resources. Conversely, independent operations may be performed "in parallel" using separate resources. For example, thirty-two (32) separate matrix configuration operations configured as 2-point FFTs may be used to handle the first stage of a 64-point FFT. Additionally, in a separate matrix configuration memory array, some operations may be deconstructed such that two or more parts of the operation are performed in parallel, where the results are then combined. Highly parallelized operations may greatly reduce latency; however, the overall memory configuration resource utilization may be extremely high.
[0264] In one exemplary embodiment, instructions are received from a processor via a dedicated interface. The dedicated interface may be particularly useful in cases where the matrix computing configuration is treated as a coprocessor or a hardware accelerator. It is noteworthy that the dedicated interface does not require arbitration and can operate at extremely high speeds (in some cases, at the native processor speed). In other embodiments, instructions are received via a shared interface.
[0265] The shared interface can be multiplexed in terms of time, resources (e.g., lanes, channels, etc.), or in other ways with other simultaneously active memory interface functionality. Common examples of other memory interface functionality include, but are not limited to: data input / output, memory configuration, processor-in-memory (PIM) communication, direct memory access, and / or any other form of blocking memory access. In some variations, the shared interface can include one or more queuing and / or pipelining mechanisms. For example, some memory technologies can implement a pipelined interface to maximize memory throughput.
[0266] In some embodiments, the instructions can be received from any entity capable of accessing the memory interface. For example, a camera coprocessor (image signal processor (ISP)) may be able to communicate directly with the memory device to, for example, write captured data. In certain implementations, the camera coprocessor may be able to offload its processing tasks to the matrix fabric of the memory device. For example, the ISP can accelerate / offload / parallelize, e.g., color interpolation, white balance, color correction, color conversion, etc. In other instances, a baseband coprocessor (BB) may be able to communicate directly with the memory device to, for example, read / write data for transactions via a network interface. The BB processor may be able to offload, e.g., FFT / IFFT, channel estimation, beamforming / precoding calculations, and / or any number of other networking tasks to the matrix fabric of the memory device. Similarly, video and / or audio codecs typically utilize DCT / IDCT transforms and would benefit from matrix fabric operations. Given the content of this disclosure, additional other variations of the foregoing will be readily apparent to those of ordinary skill in the relevant art.
[0267] Various embodiments of the present disclosure can support queues of multiple instructions. In one exemplary embodiment, matrix operations can be queued together. For example, multiple vector matrix multiplications can be queued together to achieve matrix multiplication. Similarly, as described above, higher-order transforms (e.g., FFT1024) can be achieved by queuing multiple iterations of lower-order constituent transforms (e.g., FFT512, etc.). In yet another example, the ISP processing of an image can include multiple iterations within an iterative space (each iteration can be pre-queued). Given the content of this disclosure, additional other queuing schemes can be readily substituted by those of ordinary skill in the relevant art with equivalent results.
[0268] In some cases, matrix operations can be cascaded together to achieve higher levels of matrix operations. For example, a higher-order FFT (e.g., 1024×1024) can be decomposed into multiple iterations of lower-level FFTs (e.g., four (4) iterations of 512×512 FFTs, sixteen (16) iterations of 256×256 FFTs, etc.). In other instances, an N-point DFT of any size (e.g., which is not a power of two) can be implemented by cascading DFTs of other sizes. Still other examples of cascading and / or chaining matrix transforms can be substituted for equivalent results, and the foregoing are merely illustrative.
[0269] As previously mentioned, the non-volatility of ReRAM naturally retains the memory contents even when the ReRAM is not powered. Accordingly, certain variations of the processor-memory architecture can enable one or more processors to power the memory independently. In some cases, the processor can power the memory when the processor is inactive (e.g., keep the memory active while the processor is in low power). Independent power management of the memory can be particularly useful for performing matrix operations in the memory, for example, even when the processor is in a sleep state. For example, the memory can receive multiple instructions for execution; the processor can transition to a sleep mode until the multiple instructions have been completed. Still other embodiments can use the non-volatile nature of ReRAM to retain the memory contents when the memory is powered down; for example, certain video and / or image processing computations can be maintained within ReRAM during inactivity.
[0270] At step 904 of method 900, a memory array (or a portion thereof) can be converted into a matrix configuration based on an instruction. As used herein, the term "matrix configuration" refers to multiple memory cells having configurable impedances that, when driven by an input vector, yield an output vector and / or matrix. In one embodiment, the matrix configuration can be associated with a portion of the memory map. In some such variations, the portion is configurable in terms of its size and / or location. For example, a configurable memory register can determine whether the bank is configured as a memory or as a matrix configuration. In other variations, the matrix configuration can reuse and / or even preclude memory interface operations. For example, a memory device can allow the memory interface to potentially operate based on GPIO (e.g., in one configuration, the pins of the memory interface can be selectively operated as ADDR / DATA during normal operation, or as, for example, FFT16 during matrix operations).
[0271] In one embodiment, the instruction identifies a matrix configuration characterized by structurally defined coefficients. In one exemplary embodiment, the matrix configuration contains coefficients for a structurally defined matrix operation. For example, a matrix configuration for an 8×8 FFT is an 8×8 matrix configuration that has been pre-filled with structurally defined coefficients for the FFT. In some variations, the matrix configuration may be pre-filled with coefficients of a particular sign (positive, negative) or a particular radix (most significant bit, least significant bit, or middle bit).
[0272] As used herein, the term "structurally defined coefficients" refers to the fact that the coefficients of a matrix multiplication are defined by the matrix structure (e.g., the size of the matrix) rather than the nature of the operation (e.g., multiplying an operand). For example, a structurally defined matrix operation may be identified by row and column designations (e.g., 8×8, 16×16, 32×32, 64×64, 128×128, 256×256, etc.). Although the foregoing discussion is presented in the context of full-rank matrix operations, a deficient matrix operator may be substituted with equivalent results. For example, a matrix operation may have asymmetric columns and / or rows (e.g., 8×16, 16×8, etc.). In fact, many vector-based operations may be considered rows with a single column or columns with a single row (e.g., 8×1, 1×8).
[0273] In some hybrid hardware / software embodiments, control logic (e.g., a memory controller, a processor, a PIM, etc.) may determine whether resources exist to provide the matrix configuration. In one such embodiment, a matrix operation may be evaluated by a pre-processor to determine whether it should be handled within software or within a dedicated matrix configuration. For example, if the existing memory and / or matrix configuration usage consumes all memory device resources, then the matrix operation may need to be handled within software rather than via the matrix configuration. In these cases, the instruction may not fully return (thereby causing a traditional matrix operation via a processor instruction). In another such instance, configuring a temporary matrix configuration to handle simple matrix operations may result in such few returns that the matrix operation should be handled within software.
[0274] Various considerations can be used to determine whether a matrix configuration should be used. For example, memory management may allocate portions of a memory array for the memory and / or matrix configuration. In some implementations, portions of the memory array may be statically allocated. Static allocation may preferably reduce memory management overhead and / or simplify operational overhead (wear leveling, etc.). In other implementations, portions of the memory array may be dynamically allocated. For example, wear leveling may be necessary to ensure that the performance of the memory degrades evenly (rather than wearing out high-usage areas). Still other variations may statically and / or dynamically allocate different portions; for example, subsets of the memory and / or matrix configuration portions may be dynamically and / or statically allocated.
[0275] Incidentally, wear leveling of memory cells can be performed in any discrete amount of memory (e.g., memory banks, memory bank chunks, etc.). Wear leveling matrix fabrications can use similar techniques; for example, in one variant, the wear leveling matrix fabrication portion may require a whole matrix fabrication aggregate movement (a crossbar structure cannot move in pieces). Alternatively, the wear leveling matrix fabrication portion can be performed by first decomposing the matrix fabrication into constituent matrix calculations and diverging the constituent matrix calculations to other locations. More directly, matrix fabrication wear leveling can indirectly benefit from the "logical" matrix manipulations used in other matrix operations (e.g., decomposition, concatenation, parallelization, etc.). Specifically, decomposing a matrix fabrication into its constituent matrix fabrications can achieve better wear leveling management with only slightly more complex operations (e.g., an additional step of logical combination via the MMU).
[0276] In one exemplary embodiment, the transformation includes reconfiguring a row decoder to operate as a matrix fabrication driver that variably drives multiple rows of a memory array. In one variant, the row driver converts a digital value into an analog signal. In one variant, the digital-to-analog conversion includes changing the conductance associated with a memory cell according to a matrix coefficient value. Additionally, the transformation can include reconfiguring a column decoder to perform analog decoding. In one variant, the column decoder is reconfigured to sense an analog signal of a column corresponding to different conductance cells, where the different conductance cells are driven by corresponding rows of different signaling. The column decoder converts the analog signal into a digital value.
[0277] Although the foregoing construction is presented in one particular row-column configuration, other implementations can substitute equivalent results. For example, a column driver can convert a digital value into an analog signal, and a row decoder can convert the analog signal into a digital value. In another such example, a three-dimensional (3D) row-column-depth memory can implement 2D matrices (e.g., row driver / column decoder, row driver / depth decoder, column driver / depth decoder, etc.) and / or 3D matrix arrangements (e.g., row driver / column decoder-driver / depth decoder) in any arrangement.
[0278] In one exemplary embodiment, the matrix coefficient value corresponds to a structurally determined value. The structurally determined value can be based on the nature of the operation. For example, a fast Fourier transform (FFT) of a vector of length N (where N is a power of 2) can be performed using a (2×2) FFT butterfly operation or some higher-order butterfly (e.g., 4×4, 8×8, 16×16, etc.). It is
[0279] noted that the intermediate constituent FFT butterfly operation weights are defined as a function of the unit circle (e.g., ), where both n and k are determined according to the FFT vector length N; in other words, the FFT butterfly weighting operation is structurally defined according to the length vector of length N. In fact, various different transforms are similar in this regard. For example, both the discrete Fourier transform (DFT) and the discrete cosine transform (DCT) use structurally defined coefficients.
[0280] In an exemplary embodiment, the matrix configuration itself has a structurally determined size. The structurally determined size can be based on the nature of the operation; for example, ISP white balance processing can use a 3×3 matrix (corresponding to different values of red (R), green (G), blue (B), luminance (Y), red chrominance (Cr), blue chrominance (Cb), etc.). In another such example, the channel matrix estimation and / or beamforming codebook are typically defined according to the number of multiple input multiple output (MIMO) paths. For example, a 2×2 MIMO channel has a corresponding 2×2 channel matrix and a corresponding 2×2 beamforming / precoding weighting (a 100×200 MIMO channel has a corresponding 100×200 channel matrix, etc.). Given the content of the present disclosure, various other structurally defined values and / or sizes useful for matrix operations can be substituted by those of ordinary skill in the relevant art.
[0281] Some variants may further subdivide the matrix coefficient values to handle manipulations that may not be handled otherwise. In these cases, the matrix configuration may include only a portion of the matrix coefficient values (to perform only a portion of the matrix operation). For example, performing signed operations and / or higher levels of radix calculations may require levels of manufacturing tolerances that are extremely expensive. Signed matrix operations can be divided into positive matrix operations and negative matrix operations (which are later summed by a matrix multiplication unit (MMU) described elsewhere herein). Signed composite matrix operations can be divided into positive real, negative real, positive imaginary, and negative imaginary operations (which are appropriately combined by one or more MMUs). Similarly, high-radix matrix operations can be divided into, for example, the most significant bit (MSB) portion, the least significant bit (LSB) portion, and / or any intermediate bits (which can be shifted and summed by the aforementioned MMU). Given the content of the present disclosure, those of ordinary skill will readily understand additional other variants.
[0282] In an exemplary embodiment, the matrix coefficient values are determined in advance and stored in a look-up table for later reference. For example, both the structurally determined size and the matrix operations with structurally determined values can be stored in advance. Only as one such example, an eight (8)-element FFT has a structurally determined size (8×8) and structurally determined values (e.g., etc.). The FFT8 instruction can cause the configuration of an 8×8 matrix configuration that is pre-filled with the corresponding FFT8 structurally determined values.
[0283] As another such example, antenna beamforming / precoding coefficients are typically pre-defined in a codebook; the wireless network can identify the corresponding index within the codebook to configure antenna beamforming / precoding. For example, a MIMO codebook can identify possible configurations for a 4×4 MIMO system; during operation, the selected configuration can be retrieved from the codebook based on the index therein.
[0284] Although the foregoing examples are presented in the context of dimensions and / or values defined structurally, other embodiments can use dimensions and / or values defined based on one or more other system parameters. For example, low-power operation may require a smaller granularity. Similarly, as previously mentioned, various processing considerations can be weighted in a manner that favors (or obviates) performing matrix operations within a matrix configuration. Additionally, matrix operations can affect other memory considerations, including but not limited to: wear leveling, memory bandwidth, memory embedded processing bandwidth power consumption, row-column and / or depth decoding complexity, etc. The foregoing is merely illustrative, and one of ordinary skill in the art, considering the present disclosure, can substitute various other considerations.
[0285] At step 906 of method 900, one or more matrix multiplication units can be configured based on the instructions. As previously mentioned, certain matrix configurations can implement logic (mathematical identities) to handle single-level matrix operations. However, multiple levels of matrix configurations can be cascaded together to achieve more complex matrix operations. In one exemplary embodiment, a first matrix is used to calculate the positive product of a matrix operation, and a second matrix is used to calculate the negative product of the matrix operation. The resulting positive and negative products can be compiled within the MMU to provide signed matrix multiplication.
[0286] In another exemplary embodiment, assuming the matrix has complex values and the input vector has only real values, a first matrix is used to calculate the positive real product of the matrix operation, a second matrix is used to calculate the negative real product, a third matrix is used to calculate the positive imaginary product, and a fourth matrix is used to calculate the negative imaginary product of the matrix operation. The four results can be compiled within the MMU to obtain a matrix-vector product.
[0287] In another exemplary embodiment, assuming both the matrix and the vector can have complex values, the matrix-vector product can be calculated as follows: by inputting the real coefficients of the vector into the four memory matrices of the previous example, inputting the imaginary coefficients of the vector into four equivalent (or identical) memory matrices, and combining the eight results within one or more MMUs to obtain the matrix-vector product.
[0288] In one other exemplary embodiment, a first matrix is used to calculate the first radix portion of a matrix operation, and a second matrix is used to calculate the second radix portion of the matrix operation. The resulting radix portions can be shifted and / or summed within the MMU to provide a larger radix product.
[0289] Incidentally, logical matrix operations differ from analog matrix operations. An exemplary matrix configuration converts analog voltages or currents into digital values read by a matrix multiplication unit (MMU). Logical operations can manipulate the digital values via mathematical properties (e.g., via matrix factorization, etc.), which cannot be done with analog voltages or currents in this way.
[0290] More generally, groups of matrices can be used to perform different logical manipulations. For example, a matrix can be decomposed or factored into one or more component matrices. Similarly, multiple component matrices can be aggregated or combined into a single matrix. Additionally, a matrix can be expanded in rows and / or columns to produce a larger (but rank-equivalent) deficient matrix. Such logic can be used to implement multiple higher-order matrix operations. For example, multiplying two matrices together can be decomposed into multiple vector-matrix multiplications. These vector-matrix multiplications can be further implemented as multiply-accumulate logic within the matrix multiplication unit (MMU). In other words, even non-unary operations can be treated as a series of segmented unary matrix operations. More generally, one of ordinary skill in the relevant art will readily understand that any matrix operation that can be represented fully or in part as a unary operation can greatly benefit from the various principles described herein.
[0291] Various embodiments of the present disclosure use a matrix multiplication unit (MMU) as the glue logic between multiple component matrix configurations. Additionally, MMU operations can be selectively switched to different rows and / or columns for connectivity. Not all matrix configurations can be used simultaneously; thus, depending on current processing and / or memory usage, matrix configurations can be selectively connected to the MMU. For example, a single MMU can be dynamically connected to different matrix configurations.
[0292] In some embodiments, control logic (e.g., a memory controller, a processor, a PIM, etc.) can determine whether resources exist to provide MMU manipulation, e.g., within a column decoder or elsewhere. For example, the current MMU load can be evaluated by a pre-processor to determine whether the MMU can be overloaded. It is noted that the MMU is mainly used for logical manipulation, so any processing entity with equivalent logical functionality can contribute to the MMU's tasks. For example, a processor-in-memory (PIM) can offload MMU manipulation. Similarly, matrix configuration results can be provided directly to a host processor, which can perform logical manipulation in software.
[0293] More generally, various embodiments of the present disclosure contemplate sharing MMU logic among multiple different matrix fabrications. The sharing can be based on, for example, a time-sharing scheme. For example, the MMU can be allocated to a first matrix fabrication during one time slot and to a second matrix fabrication during another time slot. In other words, unlike the physical structure of the matrix fabrication, which is statically allocated during the duration of matrix operations, the MMU performs logical operations that can be scheduled, segmented, allocated, reserved, and / or partitioned in any number of ways. More generally, various embodiments of the matrix fabrication are based on memory and non-volatility. Thus, the matrix fabrication can be preconfigured and read when needed; the non-volatile nature ensures that the matrix fabrication retains its content even if, for example, the memory device is powered off without the need for processing overhead.
[0294] If both the matrix fabrication and the corresponding matrix multiplication unit (MMU) are successfully transformed and configured, then at step 908 of method 900, the matrix fabrication is driven based on instructions, and at step 910, one or more matrix multiplication units are utilized to calculate a logical result. In one embodiment, one or more operands are converted into electrical signals for analog computation via the matrix fabrication. The analog computation is generated by driving electrical signals through matrix fabrication elements; for example, a voltage drop is a function of the coefficients of the matrix fabrication. The analog computation result is sensed and converted back to a digital domain signal. Subsequently, one or more matrix multiplication units (MMUs) are utilized to manipulate one or more digital domain values to produce a logical result.
[0295] Now referring Figure 10 , a exemplary embodiment of a Figure 9 method applicable to channel pre-coding is shown and described. At step 1002, a memory device receives one or more pre-coding related instructions from, for example, a wireless baseband or an antenna processor, and an array of communication data to be transformed by the memory device. In some variations, the control logic of the memory device is configured to divide the data into several smaller arrays (vectors) and temporarily store the arrays in an internal memory array (such as Figures 8A to 8C array 812).
[0296] At step 1004 of method 1000, a portion of the memory array is converted into a matrix fabrication configured to perform pre-coding matrix operations. As discussed in more detail below with respect to Figure 10A and 10B , in some embodiments, two separate portions of the memory fabrication, each sized four times the size of the pre-coding matrix, are identically configured with the pre-coding matrix W (divided into sub-arrays that hold the respective positive real, negative real, positive imaginary, and negative imaginary values of W).
[0297] At step 1006, one or more matrix multiplication units may be configured to compile, for example, matrix-vector products from (i) the positive real, negative real, positive imaginary, and negative imaginary programming portions of a first array and (ii) the positive real, negative real, positive imaginary, and negative imaginary programming portions of a second array (relative to Figure 7B description).
[0298] At step 1008 of method 1000, matrix configuration is driven based on pre-decoded instructions, and at step 1010, one or more matrix multiplication units are used to compute a logical result. In one embodiment, one or more operands are converted into electrical signals for analog computation via the matrix configuration. As discussed above, the analog computation is produced by driving electrical signals through matrix configuration elements, and subsequently, one or more matrix multiplication units (MMUs) manipulate one or more digital domain values to produce a logical result.
[0299] Now referring to Figure 10A and 10B , methods 1020 and 1040 are used to perform pre-decoding operations on data vector d.
[0300] At Figure 10A step 1022, the memory device receives an instruction specifying a pre-decoded matrix index (or a specific base address). The control logic of the memory device (see, for example, Figure 8A control logic 808 of the device of
[0301] uses the matrix index to look up the value of the appropriate pre-decoded matrix in a look-up table (LUT) (such as a look-up table located inside memory array 812).
[0302] At step 1024, the pre-decoded matrix W is determined, and in step 1026, the memory configuration matrix is configured or programmed with the determined pre-decoded matrix from step 1024.
[0303] In step 1028 of method 1020, the data vector d is obtained sequentially from the memory array and is individually used to drive the matrix configuration configured with the values of the pre-decoded matrix W. Figure 7B description).
[0304] Finally, in step 1032 of method 1020, the computed vector a is stored in, for example, a temporary memory location for later export from the memory device.
[0305] Figure 10B An alternative method is shown. At Figure 10B step 1042, the memory device receives an instruction containing the value of the pre-decoded matrix W, as well as the associated data vector d.
[0306] At step 1044, the memory configuration matrix is configured or programmed with the received pre-decoded matrix from step 1042.
[0307] In step 1046 of method 1040, the data vector d obtained in step 1042 is individually used to drive a matrix configuration configured with the received pre-decoded matrix W.
[0308] In step 1048, the previously configured MMU obtains the result of the matrix configuration and calculates the transmit vector a (using the method described above with respect to Figure 7B ).
[0309] Finally, in step 1050 of method 1040, the calculated vector a is transmitted out of the memory device (e.g., once it is calculated).
[0310] It will be appreciated that, in view of the present disclosure, the foregoing methods may be readily used by one of ordinary skill in the art in conjunction with Figures 5B to 5D the fabric architectures of 6B to 6C and 7B.
[0311] It will be appreciated that while certain aspects of the present disclosure are described in terms of a particular order of steps of a method, these descriptions merely illustrate the broader methods in the present disclosure and may be modified according to the needs of a particular application. In some cases, certain steps may become unnecessary or optional. Additionally, certain steps or functionality may be added to the disclosed embodiments, or the order of execution of two or more steps may be permuted. Further, two or more features from the methods may be combined. All such variations are considered to be encompassed within the disclosure herein and claimed.
[0312] The description presented herein in conjunction with the figures describes example configurations and does not represent all examples that may be implemented or that are within the scope of the claims. The term "exemplary" as used herein means "serving as an example, instance, or illustration" and is not "preferred" or "advantageous" over other examples. For the purpose of providing an understanding of the described techniques, the detailed description includes specific details. However, the techniques may be practiced without these specific details. In some instances, well-known structures and devices are shown in block diagram form to avoid obscuring the concepts of the described examples.
[0313] Any of a variety of different arts and techniques may be used to represent the information and signals described herein. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.
[0314] While the foregoing detailed description has shown, described, and pointed out the novel features of the present disclosure as applicable to various embodiments, it should be understood that those skilled in the art may make various omissions, substitutions, and changes in the form and details of the illustrated apparatus or method without departing from the present disclosure. This specification is in no way limiting, but rather should be regarded as an illustration of the general principles of the present disclosure. The scope of the present disclosure should be determined with reference to the claims.
[0315] It should be further understood that while certain steps and aspects of the various methods and devices described herein may be performed by humans, the aspects and individual methods and devices disclosed are generally computerized / implemented by a computer. Computerized devices and methods are necessary for implementing these aspects in their entirety for any number of reasons including, but not limited to, commercial viability, practicality, and even feasibility (i.e., certain steps / procedures cannot be performed by humans in any practicable manner).
[0316] The functions described herein may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored as one or more instructions or code on a computer-readable device (e.g., a storage medium) or transmitted via a computer-readable device. Computer-readable media includes both non-transitory computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. Non-transitory storage media may be any available media that can be accessed by a general purpose or special purpose computer. Also, any connection may be properly termed a computer-readable media. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, disk and disc include CD, laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically with a laser. Combinations of the above are also included within the scope of computer-readable media.
Claims
1. A non-transitory computer-readable medium comprising: a memory cell array, wherein each memory cell in the memory cell array is configured to store a digital value as an analog value in an analog medium; a memory sensing component, wherein the memory sensing component is configured to read the analog value of a first memory cell as a first digital value; Logic configured to: Receive decoding matrix operation code; operating the memory cell array as a matrix multiplication unit MMU based on the decoded matrix operation operation code; wherein each memory unit of the MMU modifies the analog value in the analog medium according to the decode matrix operation operation code and at least one decode operand; configuring the memory sensing component to convert the analog value of the first memory cell into a second digital value according to a matrix transformation operation code and the at least one decoding operand; and In response to reading the at least one decode operand into the MMU, a matrix transformation result is written based on the second digital value.
2. The non-transitory computer-readable medium of claim 1, wherein the matrix transformation opcode indicates a size of the MMU.
3. The non-transitory computer-readable medium of claim 2, wherein: The matrix transformation operation code includes at least one vector corresponding to the communication data; The decoding matrix operation code includes a wireless channel pre-decoding matrix operation code; and The at least one decoding operand includes a wireless channel pre-decoding operand.
4. The non-transitory computer-readable medium of claim 1, wherein the coding matrix operation spans at least one other MMU.
5. The non-transitory computer-readable medium of claim 1, wherein the coding matrix operation opcode identifies one or more analog values corresponding to one or more memory cells.
6. The non-transitory computer-readable medium of claim 5, wherein: The one or more analog values corresponding to the one or more memory cells are stored in a lookup table (LUT) data structure; and The LUT is located in the memory cell array and includes a codebook of decoding matrices.
7. The non-transitory computer-readable medium of claim 1 , wherein each memory cell of the MMU comprises a Resistive Random Access Memory (ReRAM) cell; and Wherein each memory unit of the MMU multiplies the analog values in the analog medium according to the decoding matrix operation operation code and at least one decoding operand.
8. The non-transitory computer-readable medium of claim 7, wherein each memory unit of the MMU further accumulates the analog value and a previous analog value in the analog medium.
9. The non-transitory computer-readable medium of claim 1, wherein the first digital value is characterized by a first base of two (2); and Wherein the second digital value is characterized by a second base number greater than two (2).
10. An apparatus for performing a matrix transformation operation, comprising: a processor device coupled to a non-transitory computer-readable medium; wherein the non-transitory computer-readable medium comprises one or more instructions that, when executed by the processor device, cause the apparatus to: Writing MIMO (Multiple Input Multiple Output) processing matrix operation operation codes and matrix transformation operands into the non-transitory computer readable medium; wherein the MIMO processing matrix operation operation code causes the non-transitory computer-readable medium to operate at least one memory cell array into a matrix structure; wherein each memory unit modifies one or more analog values of the matrix structure according to the matrix transformation operand and the matrix operation operation code; and Read the matrix transformation result from the matrix structure.
11. The apparatus of claim 10, wherein the non-transitory computer-readable medium further comprises one or more instructions that, when executed by the processor device, cause the processor to: Obtaining information related to a pre-coding matrix; and obtaining one or more vectors derived from a communication data stream; Wherein the matrix transform operands include the one or more vectors, and the matrix transform results include one or more pre-coded transmit vectors.
12. The device of claim 10, wherein the non-transitory computer-readable medium further comprises one or more instructions that, when executed by the processor device, cause the device to: Obtaining information related to the MIMO data recovery matrix; and obtaining one or more vectors corresponding to received signal data; The matrix transformation operands include the one or more vectors, and the matrix transformation results include one or more restored data values.
13. The device of claim 10, wherein the non-transitory computer-readable medium further comprises one or more instructions that, when executed by the processor device, cause the device to: receiving data comprising one or more vectors derived from a matrix of data symbols; Wherein the matrix transform operands include the one or more vectors, and the matrix transform results include one or more transmit vectors corresponding to signals transmitted via two or more antennas of the device.
14. The device of claim 10, wherein the MIMO processing matrix operation opcode causes the non-transitory computer-readable medium to operate another array of memory cells into another matrix structure; and The matrix transformation result associated with the matrix structure and another matrix transformation result associated with the another matrix structure are logically combined.
15. The apparatus according to claim 10, wherein: The one or more analog values of the matrix structure are stored in a look-up table (LUT) data structure; and The LUT data structure includes a codebook of pre-coding matrices.
16. A method for performing a matrix transformation operation, comprising: Receive spatial diversity processing matrix operation code; Based on the spatial diversity processing matrix operation operation code, configuring a memory cell array of a memory into a matrix structure; configuring a memory sensing component based on the spatial diversity processing matrix operation operation code; and writing a matrix transformation result from the memory sensing component in response to reading a matrix transformation operand into the matrix structure, Each memory unit modifies one or more analog values of the matrix structure according to the matrix transformation operand and the spatial diversity processing matrix operation operation code.
17. The method of claim 16, wherein configuring the memory cell array comprises connecting a plurality of word lines and a plurality of bit lines corresponding to row and column dimensions associated with the matrix structure.
18. The method of claim 17, further comprising determining the row size and the column size from the spatial diversity processing matrix operation opcode.
19. The method of claim 16, wherein configuring the memory cell array comprises setting the one or more analog values of the matrix structure based on a look-up table (LUT) data structure.
20. The method of claim 19, further comprising identifying an entry from the LUT data structure based on the spatial diversity processing matrix operation opcode.
21. The method of claim 20, wherein the spatial diversity processing matrix operation comprises a precoding matrix index, and wherein identifying the entry from the LUT data structure comprises finding a precoding matrix corresponding to the precoding matrix index within the LUT data structure using the precoding matrix index.
22. The method of claim 16, wherein the memory sensing component is configured such that a matrix transformation result has a cardinality greater than two (2).
23. The method of claim 16, wherein the spatial diversity processing matrix operation opcode comprises matrix value information, and wherein configuring the memory cell array comprises setting the one or more analog values of the matrix structure based on the matrix value information.
Citation Information
Patent Citations
Image processor formed in an array of memory cells
US10440341B1
Methods and apparatus for routine based fog networking
US11013043B2
Methods and apparatus for incentivizing participation in fog networks
US11184446B2
Methods and apparatus for performing matrix transformations within a memory array
US12118056B2
Methods and apparatus for characterizing memory devices
US20200264688A1