Extended memory architecture
By expanding the memory architecture and utilizing multiple computing devices and communication subsystems, the problem of increased time and resource consumption during data transmission and processing of memory devices is solved, achieving more efficient data transmission and processing and improving the performance of the computing system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MICRON TECHNOLOGY INC
- Filing Date
- 2021-06-25
- Publication Date
- 2026-04-24
AI Technical Summary
Existing memory devices suffer from increased time consumption and resource consumption during data transfer and processing, especially when performing memory operations, and these problems become more pronounced as storage capacity increases.
By employing an extended memory architecture, and through the combination of multiple computing devices and communication subsystems, it enables efficient data transfer and processing between computing devices and memory devices, reduces the number of function calls and commands, allows extended memory operations to be performed within the computing device, and reduces data movement and processing time.
It improves the performance of computing systems, reduces processing time and resource consumption, supports a wider range of memory operations, mitigates or eliminates the effects of locking or mutual exclusion operations, and enables more efficient data transfer and processing.
Smart Images

Figure CN113851168B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to semiconductor memories and methods, and more specifically, to apparatus, systems and methods for expanding memory architectures. Background Technology
[0002] Memory devices are typically provided as internal semiconductor integrated circuits in computers or other electronic systems. Many different types of memory exist, including volatile and non-volatile memory. Volatile memory may require power to maintain its data (e.g., host data, error data, etc.) and includes random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), synchronous dynamic random access memory (SDRAM), and thyristor random access memory (TRAM), etc. Non-volatile memory provides persistent data by retaining the stored data when no power is supplied and can include NAND flash memory, NOR flash memory, and resistive variable memory, such as phase-change random access memory (PCRAM), resistive random access memory (RRAM), and magnetoresistive random access memory (MRAM), such as spin torque transfer random access memory (STT RAM), etc.
[0003] A memory device may be coupled to a host computer (e.g., a host computing device) to store data, commands, and / or instructions for use by the host computer or electronic system while it is operating. For example, data, commands, and / or instructions may be transferred between the host computer and the memory device during operation of the computing or other electronic system. Summary of the Invention
[0004] On one hand, this disclosure relates to an extended memory architecture device comprising: a plurality of computing devices, each comprising: a processing unit configured to perform operations on data blocks; and a memory array configured as a cache for each corresponding processing unit; a first communication subsystem coupled to a host and each of the plurality of computing devices; and a plurality of second communication subsystems coupled to each of the plurality of computing devices, wherein each of the plurality of second communication subsystems is coupled to at least one hardware accelerator; wherein each of the plurality of computing devices is configured to: receive a request to perform an operation from the host; and send a command to perform at least a portion of the operation via one of the plurality of second communication subsystems to the at least one hardware accelerator; and receive a result of performing the operation from the at least one hardware accelerator.
[0005] In another aspect, this disclosure relates to a system with an extended memory architecture, comprising: a plurality of computing devices, each including: a processing unit configured to perform operations on data blocks; and a memory array configured as a cache for each corresponding processing unit; a first communication subsystem coupled to a host and each of the plurality of communication subsystems; a plurality of second communication subsystems coupled to each of the plurality of computing devices, wherein each of the plurality of second communication subsystems is coupled to: at least one hardware accelerator; and at least one internal SRAM; and a non-volatile memory device; wherein each of the plurality of computing devices is configured to: receive a request to perform an operation from the host; and send a command to perform at least a portion of the operation via one of the plurality of second communication subsystems to the at least one hardware accelerator or the at least one internal SRAM; and receive a result of performing the operation from the at least one hardware accelerator or the at least one internal SRAM.
[0006] In another aspect, this disclosure relates to a method involving an extended memory architecture, comprising: transmitting a command from a host to at least one of a plurality of computing devices via a first communication subsystem; transmitting a data block associated with the command from a memory device to the at least one of the plurality of computing devices via a second communication subsystem, wherein: the first communication subsystem is coupled to the host and the at least one of the plurality of computing devices; and the second communication subsystem is coupled to the at least one of the plurality of computing devices and the memory device; in response to the receipt of the command and the data block, the at least one computing device performs an operation using the data block to reduce the data size from a first size to a second size by the at least one computing device; and transmitting the reduced-size data block to the host via the first communication subsystem. Attached Figure Description
[0007] Figure 1 This is a functional block diagram of a computing system comprising a device including a first communication subsystem, a second plurality of communication subsystems, and a plurality of memory devices, according to several embodiments of the present disclosure.
[0008] Figure 2 This is yet another functional block diagram of a computing system comprising a device including a first plurality of communication subsystems, a second plurality of communication subsystems, and a plurality of memory devices, according to several embodiments of the present disclosure.
[0009] Figure 3 This is yet another functional block diagram of a computing system comprising a device including a computing core, multiple communication subsystems, and multiple memory devices, according to several embodiments of the present disclosure.
[0010] Figure 4 This is a functional block diagram of a device in the form of a computing core including several ports, according to several embodiments of the present disclosure.
[0011] Figure 5 This is a flowchart illustrating example methods corresponding to an extended memory architecture according to several embodiments of the present disclosure. Detailed Implementation
[0012] This describes systems, apparatuses, and methods for performing extended memory operations in relation to extended memory communication subsystems. An example apparatus may include a plurality of computing devices coupled to each other. Each of the plurality of computing devices may include a processing unit configured to perform operations on data blocks in response to the receipt of data blocks. Each of the plurality of computing devices may further include a memory array configured as a cache for the processing unit. The example apparatus may further include a first plurality of communication subsystems coupled to the plurality of computing devices and a host, and a second plurality of communication subsystems coupled to the plurality of computing devices and hardware accelerators, SRAM, or additional components. The first and second plurality of communication subsystems are configured to request and / or transmit data blocks.
[0013] Extended memory architectures can transmit instructions that execute operations specified by a single address and operands, and can be executed by a computing device including processing units and memory resources. The computing device can perform extended memory operations on data streaming through the computing device without receiving intervention commands. In an example, the computing device is configured to receive commands that execute operations including performing operations on data using the processing units of the computing device and to determine that operands corresponding to said operations are stored in memory resources. The computing device can further perform operations using the operands stored in the memory resources.
[0014] The computing device may be a RISC-V application processor core capable of supporting a full-featured operating system (such as Linux). This particular core may be used in conjunction with, for example, Internet of Things (IoT) nodes and gateways, storage devices, and / or networks. The core may be coupled to several ports, such as memory ports, system ports, peripheral ports, and / or front-end ports. As an example, the memory port may communicate with a memory device, the system port may communicate with an on-chip accelerator or "fast" SRAM, the peripheral port may communicate with an off-chip serial port, and / or the front-end port may communicate with a host interface, as will be discussed below. Figure 4 Further description.
[0015] In this manner, a first communication subsystem can be used to guide data from a specific port (e.g., a memory port of a computing device) through a communication subsystem (e.g., a multiplexer that selects that specific memory port) and transmit the data through an additional communication subsystem (e.g., an interface of an AXI interconnect interface) to a memory controller, which can then transfer the data to a memory device (e.g., DDR memory, 3-D cross-point memory, NAND memory, etc.). In an example, the AXI interconnect interface may conform to... of AXI version 4 specification includes a subset of the AXI4-Lite control register interface.
[0016] As used herein, "extended memory operation" refers to a memory operation that can be specified by a single address (e.g., a memory address) and operands (e.g., a 64-bit operand). Operands may be represented as multiple bits (e.g., a bit string or a sequence of bits). However, embodiments are not limited to operations specified by 64-bit operands, and operations may be specified by operands larger than 64 bits (e.g., 128 bits) or smaller than 64 bits (e.g., 32 bits). As described herein, the effective address space accessible to perform extended memory operations is the size of the memory devices or file systems accessible to the host computing system or storage controller.
[0017] Extended memory operations may include processing devices (e.g., processing devices consisting of cores 110, 210, 310, 410, or...). Figure 4 The core computing device (specifically shown as 410) executes instructions and / or operations. An example of a core may include a reduced instruction set computing device or other hardware processing device that executes instructions to perform various computing tasks. In some embodiments, performing extended memory operations may include retrieving data and / or instructions stored in memory resources of the computing device, performing operations within the computing device 110 (e.g., without transferring data or instructions to external circuitry), and storing the results of the extended memory operations in memory resources or secondary storage of the computing device 110 (e.g., stored in, for example, the memory resources described herein). Figure 1 (In the memory devices 116-1 and 116-2 described herein).
[0018] Non-limiting examples of extended memory operations may include floating-point addition, 32-bit complex number operations, square root address (SQRT(addr)) operations, conversion operations (e.g., conversion between floating-point and integer formats and / or between floating-point and assumed formats), normalizing data to a fixed format, absolute value operations, etc. In some embodiments, extended memory operations may include in-situ update operations performed by a computing device (e.g., where the result of the extended memory operation is stored at an address where the operand used to perform the extended memory operation was stored before the extended memory operation was performed) and operations where previously stored data is used to determine new data (e.g., where the operand stored at a specific address is used to generate new data that overwrites the specific address of the stored operand).
[0019] Therefore, in some embodiments, the execution of extended memory operations can mitigate or eliminate locking or mutual exclusion operations because the extended memory operations can be performed within a computing device, which reduces contention between multi-threaded executions. Reducing or eliminating locking or mutual exclusion operations on threads during the execution of extended memory operations can result in, for example, performance improvements in the computing system, because the extended memory operations can be performed in parallel within the same computing device or in parallel across two or more computing devices that communicate with each other. Additionally, in some embodiments, the extended memory operations described herein can mitigate or eliminate locking or mutual exclusion operations when the results of the extended memory operations are transferred from the computing device that performed the operation to the host.
[0020] Memory devices can be used to store important or critical data in computing devices and can transfer such data between hosts associated with the computing devices via at least one extended memory architecture. However, as the size and amount of data stored in memory devices increase, transferring data to and from hosts can become time-consuming and resource-intensive. For example, when a host requests to perform a memory operation using a large block of data, the amount of time and / or resources consumed in fulfilling the request can increase proportionally to the size and / or amount of data associated with the block.
[0021] These effects can become more pronounced as the storage capacity of memory devices increases, because more and more data can be stored in the memory devices and thus made available for memory operations. Furthermore, because data can be processed (e.g., memory operations can be performed on the data), the amount of data that can be processed also increases as the amount of data that can be stored in the memory devices increases. This can lead to increased processing time and / or increased consumption of processing resources, which may be exacerbated when performing certain types of memory operations. To mitigate these and other problems, the embodiments described herein allow the use of memory devices, one or more computing devices and / or memory arrays, and a first plurality of communication subsystems (e.g., PCIe interfaces, PCIe XDMA interfaces, AXI interconnect interfaces, etc.) and a second plurality of subsystems (e.g., interfaces, such as AXI interconnects) to perform extended memory operations to more efficiently transfer data from computing devices to memory devices and / or from computing devices to a host, and vice versa.
[0022] In some embodiments, data can be transmitted to multiple memory devices via these communication subsystems, bypassing multiple computing devices. In some embodiments, data can be transmitted via these communication subsystems, passing through at least one of the multiple computing devices. Depending on the route of data transmission, each of the interfaces may have a unique speed. As will be further described below, when bypassing multiple computing devices, data can be transmitted at a higher rate than when data is passed through at least one of the multiple computing devices.
[0023] In some methods, performing a memory operation may require multiple clock cycles and / or multiple function calls to the memory of the computing system (e.g., memory devices and / or memory arrays). In contrast, the embodiments described herein allow for extended memory operations, where the memory operation is performed using a single function call or command. For example, the embodiments described herein allow for the use of fewer function calls or commands compared to methods where at least one command and / or function call is used to load data to be operated on and then at least one subsequent function call or command is used to store the data already operated on. Furthermore, the computing device of the computing system may receive requests to perform memory operations via a first communication subsystem (e.g., a PCIe interface, a multiplexer, an on-chip control network, etc.) and / or a second communication subsystem (e.g., an interface, an interconnect of an AXI interconnect, etc.) and may receive data blocks from the memory device via the first and second communication subsystems for performing the requested memory operation. While the first and second communication subsystems have been described in conjunction, the embodiments are not limited thereto. As an example, the request for data and / or the receipt of data blocks may be made solely through the first communication subsystem or solely through the second communication subsystem.
[0024] By reducing the number of function calls and / or commands used to perform memory operations, the amount of time consumed and / or computational resources consumed in performing such operations can be reduced compared to methods that require multiple function calls and / or commands to perform memory operations. Furthermore, the embodiments described herein can reduce data movement within memory devices and / or memory arrays because data does not need to be loaded into a specific location before performing memory operations. This can reduce processing time compared to some methods, especially in cases where a large amount of data undergoes memory operations.
[0025] Furthermore, the extended memory operations described herein allow for a much larger set of type fields compared to some other methods. For example, an instruction executed by a host to request the use of data in a memory device (e.g., a memory subsystem) to perform an operation may include type, address, and data fields. The instruction may be sent to at least one of a plurality of computing devices via a first communication subsystem (e.g., a multiplexer) and a second communication subsystem (e.g., an interface), and data may be transferred from the memory device via the first and / or second communication subsystems. The type field may correspond to the specific operation requested, the address may correspond to the address where the data to be used to perform the operation is stored, and the data field may correspond to the data to be used to perform the operation (e.g., operands). In some methods, the type field may be limited to reads and / or writes of varying sizes and some simple integer accumulation operations. In contrast, the embodiments described herein allow for the use of a wider range of type fields because the effective address space available when performing extended memory operations can correspond to the size of the memory device. By expanding the address space available for performing operations, the embodiments described herein therefore allow for a wider range of type fields, and thus, a wider range of memory operations can be performed than in methods that do not allow for an effective address space that is the size of the memory device.
[0026] In the following detailed description of this disclosure, reference is made to the accompanying drawings, which form part of this disclosure and illustrate by way of how the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice embodiments of the disclosure, and it should be understood that other embodiments may be utilized and process, electrical, and / or structural changes may be made without departing from the scope of this disclosure.
[0027] As used herein, identifiers such as “X,” “Y,” “N,” “M,” “A,” “B,” “C,” “D,” etc., specifically relating to reference numerals in figures, indicate the number of such particular features that may be included. It should also be understood that the terminology used herein is for the purpose of depicting particular embodiments only and is not intended to be limiting. As used herein, the singular forms “a / an” and “described” may include both singular and plural references unless the context explicitly indicates otherwise. Furthermore, “a plurality of…,” “at least one of…,” and “one or more of…” (e.g., a plurality of memory stores) may refer to one or more memory stores, while “a plurality of…” is intended to refer to more than one such thing. Additionally, the word “can / may” is used throughout this application in an permissive sense (i.e., possible, capable) rather than a mandatory sense (i.e., required). The term “comprising” and its derivatives mean “including (but not limited to).” Depending on the context, the term "coupled / coupling" means a direct or indirect physical connection or connection used for access to and movement (transmission) of commands and / or data. Depending on the context, the terms "data" and "data value" are used interchangeably and may have the same meaning herein.
[0028] The figures in this document follow a numbering convention, where the first digit or the first few digits correspond to the figure number, and the remaining digits identify the elements or components in the figure. Similar elements or components between different figures can be identified by using similar numbers. For example, 104 can be cited. Figure 1 Component "04" in the text, and similar components in Figure 2 The reference numeral 204 may be used. A group or number of similar elements or components are generally referred to herein by a single element number. For example, multiple reference elements 106-1, 106-2, 106-3 may be collectively referred to as 106. It should be understood that the elements shown in the various embodiments herein may be added, interchanged, and / or eliminated to provide several additional embodiments of this disclosure. In addition, the proportions and / or relative sizes of the elements provided in the figures are intended to illustrate certain embodiments of this disclosure and should not be construed as limiting.
[0029] Figure 1 This is a functional block diagram of a computing system 100 comprising a device 104 including a first communication subsystem 108, a plurality of second communication subsystems 106, and a plurality of memory devices 116, according to several embodiments of the present disclosure. As used herein, "device" may refer to (but is not limited to) any of a variety of structures or combinations thereof, such as (for example) a circuit or circuit system, one or more dies, one or more modules, one or more devices, or one or more systems. Figure 1In the embodiments described herein, memory devices 116-1, ..., 116-N may include one or more memory modules (e.g., double data rate (DDR) memory, three-dimensional (3D) cross-point memory, NAND memory, single in-line memory module, dual in-line memory module, etc.). Memory devices 161-1, ..., 116-N may include volatile memory and / or non-volatile memory. In several embodiments, memory devices 116-1, ..., 116-N may include multi-chip devices. Multi-chip devices may include several different memory types and / or memory modules. For example, a memory system may include non-volatile or volatile memory on any type of module.
[0030] Memory devices 116-1, ..., 116-N may provide main memory for computing system 100 or may be used as additional memory or storage devices for the entire computing system 100. Each memory device 116-1, ..., 116-N may include one or more arrays of memory cells, such as volatile and / or non-volatile memory cells. The array may be, for example, a flash array with a NAND architecture. Embodiments are not limited to a specific type of memory device. For example, memory devices may include RAM, ROM, DRAM, SDRAM, PCRAM, RRAM, and flash memory, etc.
[0031] In embodiments where memory devices 116-1, ..., 116-N include non-volatile memory, memory devices 116-1, ..., 116-N may be flash memory devices, such as NAND or NOR flash memory devices. However, embodiments are not limited thereto, and memory devices 116-1, ..., 116-N may include: other non-volatile memory devices, such as non-volatile random access memory devices (e.g., NVRAM, ReRAM, FeRAM, MRAM, PCM); "emerging" memory devices, such as 3-D cross-gate (3D XP) memory devices; or combinations thereof. 3D XP non-volatile memory arrays can perform bit storage based on volume resistance variations combined with stacked cross-gate format data access arrays. Furthermore, compared to many flash-based memories, 3D XP non-volatile memories can perform in-situ write operations, where non-volatile memory cells can be programmed without previously erasing them.
[0032] like Figure 1As described herein, multiple computing devices 110-1, 110-2 (hereinafter collectively referred to as multiple computing devices 110) may be coupled to a communication subsystem (e.g., a Peripheral Component Interconnect High Speed (PCIe) interface, a PCIe XDMA interface, etc.) 108. The communication subsystem 108 may include circuitry and / or logic configured to allocate and deallocate resources from host 102 to computing devices 110 or allocate and deallocate resources to host 102 during the execution of operations described herein. For example, the circuitry and / or logic may communicate data requests to computing devices 110 or allocate and / or deallocate resources to computing devices 110 during the execution of extended memory operations described herein.
[0033] Communication subsystem 108 may be directly coupled to at least one 106-1 of a plurality of communication subsystems 106 (e.g., interfaces of interconnect interfaces 106-1, 106-2, 106-3 (hereinafter collectively referred to as the plurality of communication subsystems 106)). Each of the plurality of communication subsystems 106 may be coupled to a corresponding one of controller 112, accelerator 114, SRAM (e.g., fast SRAM) 118, and peripheral components 120. In one example, the first of the plurality of communication subsystems 106, 106-1, may be coupled to controller 112. In this example, interface 106-1 may be a memory interface. Controller 112 may be coupled to a plurality of memory devices 116-1, ..., 116-N via a plurality of channels 107-1, ..., 107-N.
[0034] Secondly, in this example and as in Figure 1 As described, the second of the multiple communication subsystems 106, 106-2, can be coupled to the accelerator 114 and the SRAM 118. The on-chip accelerator 114 can be used to perform several hypothetical operations and / or to communicate with the internal SRAM on a field-programmable gate array (FPGA) containing the described components. As an example, the components of device 104 can be on an FPGA.
[0035] Furthermore, in this example, a third party 106-3 among the multiple communication subsystems 106 can be coupled to a peripheral component 120. The peripheral component 120 can be either a general-purpose input / output (GPID) LED or a universal asynchronous receiver / transmitter (UART). The GPID LED can be further coupled to additional LEDs, and the UART can be further coupled to a serial port. The multiple communication subsystems 106 can be coupled to each corresponding component via several AXI buses. The third party 106-3 among the multiple communication subsystems 106 can be used to transmit data off-chip via the peripheral component 120 or the off-chip serial port 118.
[0036] Host 102 may be a host system, such as a personal laptop computer, desktop computer, digital camera, smartphone, memory card reader, and / or Internet of Things-enabled device, as well as various other types of host, and may include memory access means, such as a processor (or processing device). Those skilled in the art will understand that "processor" can mean one or more processors, such as a parallel processing system, several coprocessors, etc. Host 102 may include a system motherboard and / or backplane and may include several processing resources (e.g., one or more processors, microprocessors, or some other type of control circuitry). In some embodiments, the host may include a host controller 101, which may be configured to control at least some operations of host 102 by, for example, generating and transmitting commands to the host controller to cause operations such as extended memory operations. Host controller 101 may include circuitry (e.g., hardware) configurable to control at least some operations of host 102. For example, host controller 101 may be an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other combinations of circuitry and / or logic configured to control at least some operations of host 102. The host 102 can communicate with the host interface 108 via channels 103 / 105.
[0037] System 100 may include a single integrated circuit, or a host 102, a communication subsystem 108, multiple communication subsystems 106, a controller 112, an on-chip accelerator 114, an SRAM 118, peripheral components 120, and / or memory devices 116-1, ..., 116-N may be on the same integrated circuit. System 100 may be, for example, a server system and / or a high-performance computing (HPC) system and / or a portion thereof. Although Figure 1 The examples shown illustrate systems with a von Neumann architecture, but embodiments of this disclosure can be implemented in non-von Neumann architectures, which may not include one or more components typically associated with a von Neumann architecture (e.g., CPU, ALU, etc.).
[0038] Controller 112 may be configured to request data blocks from one or more of memory devices 116-1, ..., 116-N and cause cores 110-1, ..., 110-N (which may be referred to herein as "computing devices") to perform operations (e.g., extended memory operations) on said data blocks. Operations may be performed to evaluate functionality that can be specified by a single address associated with the data block and one or more operands. Controller 112 may be further configured to cause the results of extended memory operations to be stored in one or more of the computing devices 110-1, ..., 110-N and / or transmitted to channels (e.g., communication paths 103 and / or 105) and / or host 102 via multiple communication subsystems 106.
[0039] In some embodiments, multiple communication subsystems 106 may request remote commands, start DMA commands, send read / write locations, and / or send start function execution commands to one of multiple computing devices 110. In some embodiments, multiple communication subsystems 106 may request data blocks to be copied from a buffer of computing device 110 to a buffer of memory controller 112 or memory device 116. Conversely, one of the multiple communication subsystems 106 may request data blocks to be copied from a buffer of memory controller 112 or memory device 116 to a buffer of computing device 110. Multiple communication subsystems 106 may request data blocks to be copied from a buffer of host 102 to computing device 110, or vice versa. Multiple communication subsystems 106 may request data blocks to be copied from a buffer of memory controller 112 or memory device 116 to a buffer of host 102. Conversely, multiple communication subsystems 106 may request data blocks to be copied from a buffer of host 102 to a buffer of memory controller 112 or memory device 116. Furthermore, in some embodiments, multiple communication subsystems 106 may request to execute commands from a host on computing device 110. Multiple communication subsystems 106 may request to execute commands from computing device 110 on additional computing device 110. Multiple communication subsystems 106 may request to execute commands from memory controller 112 on computing device 110. In some embodiments, multiple communication subsystems 106 may include at least a portion of a controller (not described).
[0040] In some embodiments, multiple communication subsystems 106 may transfer data blocks (e.g., direct memory access (DMA) data blocks) from computing device 110 to media device 116 (via memory controller 112), or vice versa. Multiple communication subsystems 106 may also transfer data blocks (e.g., DMA blocks) from computing device 110 to host 102, or vice versa. Furthermore, multiple communication subsystems 106 may transfer data blocks (e.g., DMA blocks) from host 102 to media device 116, or vice versa. In some embodiments, a plurality of communication subsystems 106 may receive output (e.g., data to which extended memory operations have been performed) from computing devices 110-1, ..., 110-N and transmit the output from computing devices 110-1, ..., 110-N to controller 115 and / or host 102 of device 104, and vice versa. For example, the plurality of communication subsystems 106 may be configured to receive data that has been operated on by extended memory operations of computing devices 110-1, ..., 110-N and transmit data corresponding to the results of the extended memory operations to controller 115 and / or host 102. In some embodiments, the plurality of communication subsystems 106 may include at least a portion of controller 115. For example, the plurality of communication subsystems 106 may include a circuitry including controller 115 or a portion thereof.
[0041] The memory controller 112 may be a "standard" or "dummy" memory controller. For example, the memory controller 112 may be configured to perform simple operations on memory devices 116-1, ..., 116-N, such as copying, writing, reading, error correction, etc. However, in some embodiments, the memory controller 112 does not perform processing (e.g., data manipulation operations) on the data associated with memory devices 116-1, ..., 116-N. For example, the memory controller 112 may cause read and / or write operations to read data from or write data to memory devices 116-1, ..., 116-N via communication paths 107-1, ..., 107-N, but the memory controller 112 may not process the data read from or written to memory devices 116-1, ..., 116-N. In some embodiments, the memory controller 112 may be a non-volatile memory controller, although embodiments are not limited thereto.
[0042] In some embodiments, a first AXI bus that couples communication subsystem 108 to a first 106-1 of a plurality of communication subsystems 106 is an AXI bus capable of transmitting data faster than a second AXI bus that couples communication subsystem 108 to computing device 110-1. For example, the first AXI bus can transmit at a rate of 300 MHz, while the second AXI bus can transmit at a rate of 100 MHz. Furthermore, the first AXI bus can be an AXI bus capable of transmitting data faster than a third AXI bus that couples computing device 110-1 to one of the plurality of communication subsystems 106.
[0043] Figure 1 Embodiments may include additional circuitry not described so as not to obscure the embodiments of this disclosure. For example, device 104 may include an address circuitry that latches address signals provided via I / O connections through an I / O circuitry. The address signals may be received and decoded by row decoders and column decoders to access memory devices 116-1, ..., 116-N. Those skilled in the art will understand that the number of address input connections may depend on the density and architecture of memory devices 116-1, ..., 116-N.
[0044] In some embodiments, extended memory operations can be used Figure 1 The computing system 100 shown executes by selectively storing or mapping data (e.g., files) into computing device 110. Data may be selectively stored in the address space of computing memory. In some embodiments, data may be selectively stored or mapped into computing device 110 in response to commands received from host 102. In embodiments where commands are received from host 102, commands may be transmitted to computing device 110 via interfaces associated with host 102 (e.g., communication paths 103 and / or 105) and via communication subsystems and multiple communication subsystems 108 and 106. Interfaces 103 / 105, communication subsystem 108, and multiple communication subsystems 106 may be peripheral component interconnect high-speed (PCIe) buses, double data rate (DDR) interfaces, interconnect interfaces (e.g., AXI interconnect interfaces), multiplexers / muxes, or other suitable interfaces or buses. However, embodiments are not limited to this.
[0045] In this scenario, where data (e.g., data to be used to perform extended memory operations) is mapped to a non-limiting instance in computing device 110, host controller 101 can transmit commands to computing device 110 to initiate the execution of extended memory operations using the data mapped to computing device 110. In some embodiments, host controller 101 can look up an address (e.g., a physical address) corresponding to the data mapped to computing device 110 and determine, based on that address, which computing device (e.g., computing device 110-1) the address (and therefore the data) is mapped to. Commands can then be transmitted to the computing device (e.g., computing device 110-1) containing the address (and therefore the data).
[0046] In some embodiments, the data may be 64-bit operands, although embodiments are not limited to operands of a specific size or length. In embodiments where the data is a 64-bit operand, once the host controller 101 transmits a command to initiate the execution of an extended memory operation to the correct computing device (e.g., computing device 110-1) based on the address of the stored data, the computing device (e.g., computing device 110-1) can use the data to perform the extended memory operation.
[0047] In some embodiments, computing device 110 may be individually addressed across a contiguous address space, which facilitates the execution of the extended memory operations described herein. That is, the location where data is stored or the address to which data is mapped is unique for all computing devices 110 such that when host controller 101 looks up an address, the address corresponds to a location in a particular computing device (e.g., computing device 110-1).
[0048] For example, a first computing device 110-1 may have a first set of addresses associated with it, a second computing device 110-2 may have a second set of addresses associated with it, a third computing device 110-3 may have a third set of addresses associated with it, and so on, up to an nth computing device (e.g., computing device 110-N) that may have an nth set of addresses associated with it. That is, the first computing device 110-1 may have a set of addresses from 0000000 to 0999999, the second computing device 110-2 may have a set of addresses from 1000000 to 1999999, the third computing device 110-3 may have a set of addresses from 2000000 to 2999999, and so on. It should be understood that these address numbers are illustrative and not limiting, and may depend on the architecture and / or size (e.g., storage capacity) of the computing device 110.
[0049] As a non-limiting example of an extended memory operation including a floating-point addition-accumulation operation, computing device 110 may treat the destination address as a floating-point number, add the floating-point number to an argument stored at the address of computing device 110, and store the result back to the original address. For example, when host controller 101 (or device controller 115 (not shown)) initiates the execution of a floating-point addition-accumulation extended memory operation, the address of computing device 110 that the host looks up (e.g., the address to which data in the computing device is mapped) may be treated as a floating-point number, and the data stored at said address may be treated as an operand for performing the extended memory operation. In response to receiving a command to initiate the extended memory operation, computing device 110 to which data (e.g., an operand in this example) is mapped may perform an addition operation to add the data to the address (e.g., the numerical value of the address) and store the result of the addition back to the original address of computing device 110.
[0050] As described above, in some embodiments, the execution of such extended memory operations may only require transferring a single command (e.g., a request command) from host 102 (e.g., from host controller 101) to memory device 104 or from controller 115 to computing device 110. Compared to some prior methods, this can reduce, for example, the amount of time required for multiple commands to traverse interfaces 103, 105 and / or the amount of time required to move data (e.g., operands) from one address to another within computing device 110, as well as the amount of time spent performing the operation.
[0051] Furthermore, the execution of extended memory operations according to this disclosure can further reduce the amount of processing power or processing time because, compared to methods where operands must be retrieved and loaded from different locations before the operation is executed, data mapped to the computing device 110 in which the extended memory operation is performed can be used as operands for the extended memory operation, and / or the address to which the data is mapped can be used as an operand for the extended memory operation. That is, at least because the embodiments herein allow skipping operand loading, the performance of the computing system 100 can be improved compared to methods that load operands and then store the results of operations performed between operands.
[0052] Furthermore, in some embodiments, because extended memory operations can be performed within computing device 110 using addresses and data stored in those addresses, and in some embodiments, because the results of extended memory operations can be stored back in the original addresses, locking or mutex operations can be relaxed or eliminated during the execution of extended memory operations. Reducing or eliminating locking or mutex operations on threads during the execution of extended memory operations can lead to improved performance of computing system 100 because extended memory operations can be performed in parallel within the same computing device 110 or across two or more computing devices 110.
[0053] In some embodiments, an effective mapping of data in computing device 110 may include a base address, a segment size, and / or a length. The base address may correspond to the address in computing device 110 where the data mapping is stored. The segment size may correspond to the amount of data that computing system 100 can process (e.g., in bytes), and the length may correspond to the number of bits corresponding to the data. It should be noted that in some embodiments, data stored in computing device 110 may not be cached on host 102. For example, an extended memory operation may be performed entirely within computing device 110 without hindering or otherwise interfering with the transfer of data to or from host 102 during the execution of the extended memory operation.
[0054] In a non-limiting instance where the base address is 4096, the segment size is 1024, and the length is 16,385, the mapped address 7234 can be in the third segment, which can correspond to a third computing device among multiple computing devices 110 (e.g., Figure 2 (Computing device 210-3 in the example). In this example, host 102 and / or communication subsystem 108 and multiple communication subsystems 106 can forward commands (e.g., requests) to perform extended memory operations to third computing device 210-3. Third computing device 210-3 can determine whether data is stored in a mapped address in the memory of third computing device 210-3. If data is stored in a mapped address (e.g., an address in third computing device 210-3), then third computing device 210-3 can use that data to perform the requested extended memory operation and can store the result of the extended memory operation back to the address where the data was originally stored.
[0055] In some embodiments, the computing device 110 containing data requested for the execution of an extended memory operation may be determined by the host controller 101 and / or the communication subsystem 108 and / or multiple communication subsystems 106. For example, a portion of the total address space available to all computing devices 110 may be allocated to each respective computing device. Therefore, the host controller 101 and / or the communication subsystem 108 and / or multiple communication subsystems 106 may have information corresponding to which portions of the total address space correspond to which computing devices 110 and thus can direct the relevant computing devices 110 to perform extended memory operations. In some embodiments, the host controller 101 and / or the second communication subsystem 106 may store addresses (or address ranges) corresponding to the respective computing devices 110 in a data structure (e.g., a table) and direct the execution of extended memory operations to the computing device 110 based on the addresses stored in the data structure.
[0056] However, the embodiments are not limited to this, and in some embodiments, the host controller 101 and / or multiple communication subsystems 106 may determine the size of memory resources (e.g., the amount of data) and determine which computing devices 110 store data to be used for performing extended memory operations based on the size of the memory resources associated with each computing device 110 and the total address space available to all computing devices 110. In embodiments where the host controller 101 and / or multiple communication subsystems 106 determine the computing devices 110 storing data to be used for performing extended memory operations based on the total address space available to all computing devices 110 and the amount of memory resources available to each computing device 110, it is possible to perform extended memory operations across multiple non-overlapping portions of the computing device memory resources.
[0057] Continuing with the above example, if no data is found at the requested address, then the third computing device 210-3 may, as described herein, [further details regarding the process]. Figure 2 The request for data is described in more detail, and an extended memory operation is performed once the data is loaded into the address of the third computing device 210-3. In some embodiments, once the extended memory operation is completed by the computing device (e.g., the third computing device 210-3 in this example), and / or the host 102 may be notified and / or the result of the extended memory operation may be transmitted to the memory device 116 and / or the host 102.
[0058] In some embodiments, memory controller 112 may be configured to retrieve data blocks from memory devices 116-1, ..., 116-N coupled to device 104 in response to a request from a controller or host 102 of device 104. Memory controller 112 may then cause the data blocks to be transferred to computing devices 110-1, ..., 110-N and / or the device controller. Similarly, memory controller 112 may be configured to receive data blocks from computing devices 110 and / or controller 115. Memory controller 112 may then cause the data blocks to be transferred to memory devices 116 coupled to memory controller 104.
[0059] The data block size may be approximately 4 kilobytes (although embodiments are not limited to this specific size) and may be processed by the computing devices 110-1, ..., 110-N in a streaming manner in response to one or more commands generated by the controller 115 and / or the host and sent via the second communication subsystem 106. In some embodiments, the data block may be a 32-bit, 64-bit, 128-bit, or other data word or data block, and / or the data block may correspond to an operand to be used to perform extended memory operations.
[0060] For example, such as regarding Figure 2 In a more detailed description, because computing device 110 can perform an extended memory operation (e.g., a procedure) on a second data block in response to completing an extended memory operation on a previous data block, data blocks can be continuously streamed through computing device 110 as data blocks are processed by computing device 110. In some embodiments, data blocks can be processed through computing device 110 in a streaming manner without intervention commands from controller and / or host 102. That is, in some embodiments, controller 115 (or host 102) can issue commands causing computing device 110 to process data blocks it receives, and data blocks subsequently received by computing device 110 can be processed without additional commands from controller.
[0061] In some embodiments, processing a data block may include performing extended memory operations using the data block. For example, computing devices 110-1, ..., 110-N may perform extended memory operations on the data block in response to commands from a controller via a plurality of communication subsystems 106 to evaluate one or more functions, remove unwanted data, extract relevant data, or otherwise combine extended memory operations with the execution of the data block.
[0062] In a non-limiting instance where data (e.g., data to be used to perform extended memory operations) is mapped to one or more of computing devices 110, the controller can transmit a command to computing device 110 to initiate the execution of extended memory operations using the data mapped to computing device 110. In some embodiments, the controller 115 can look up an address (e.g., a physical address) corresponding to the data mapped to computing device 110 and determine, based on the address, which computing device (e.g., computing device 110-1) the address (and therefore the data) is mapped to. The command can then be transmitted to the computing device (e.g., computing device 110-1) containing the address (and therefore the data). In some embodiments, the command can be transmitted to the computing device (e.g., computing device 110-1) via a second communication subsystem 106.
[0063] The controller 115 (or host) may be further configured to send commands to the computing device 110 to allocate and / or deallocate resources available to the computing device 110 for performing extended memory operations using data blocks. In some embodiments, allocating and / or deallocating resources available to the computing device 110 may involve selectively enabling some computing devices 110 while selectively deactivating others. For example, if fewer computing devices 110 than the total number of computing devices 110 are required to process data blocks, the controller 115 may send commands to the computing devices 110 intended for processing the data blocks to enable only those computing devices 110 that are desired to process the data blocks.
[0064] In some embodiments, controller 115 may be further configured to send commands to synchronize the execution of operations (e.g., memory expansion operations) performed by computing device 110. For example, controller 115 (and / or host) may send a command to first computing device 110-1 to perform a first memory expansion operation, and controller 115 (or host) may send a command to second computing device 110-2 to perform a second memory expansion operation using a second computing device. Synchronizing the execution of operations (e.g., memory expansion operations) performed by computing device 110 via controller 115 may further include causing computing device 110 to perform specific operations at a specific time or in a specific order.
[0065] As described above, data generated by the execution of extended memory operations may be stored in the computing device 110 at the original address where the data was stored before the execution of the extended memory operations. However, in some embodiments, the data generated by the execution of extended memory operations may be converted into a logical record after the execution of the extended memory operations. A logical record may include a data record independent of its physical location. For example, a logical record may be a data record pointing to an address (e.g., location) in at least one of the computing devices 110 that stores physical data corresponding to the execution of the extended memory operations.
[0066] In some embodiments, the result of an extended memory operation may be stored in the same address in the computing device memory as the address where the data was stored prior to the execution of the extended memory operation. However, embodiments are not limited to this, and the result of the extended memory operation may be stored in the same address in the computing device memory as the address where the data was stored prior to the execution of the extended memory operation. In some embodiments, logical records may point to these address locations such that the result of the extended memory operation is accessible from the computing device 110 and transferred to a circuit system outside the computing device 110 (e.g., to a host computer).
[0067] In some embodiments, controller 115 may receive data blocks directly from memory controller 112 and / or may send data blocks directly to memory controller 112. This allows controller 115 to transfer data blocks that have not been processed by computing device 110 (e.g., data blocks not used to perform extended memory operations) to memory controller 112 and receive said data blocks from memory controller 112.
[0068] For example, if controller 115 receives an unprocessed data block from a memory device coupled to memory controller 104, the unprocessed data block may be transferred to memory controller 112, which in turn may transfer the unprocessed data block to the memory device coupled to memory controller 104.
[0069] Similarly, if the host requests an unprocessed (e.g., complete) data block (e.g., a data block that has not been processed by the computing device 110), the memory controller 112 may cause the unprocessed data block to be transmitted to the controller 115, which may then transmit the unprocessed data block to the host.
[0070] Figure 2This is a functional block diagram of a computing system 200 comprising a device 204 including a first plurality of communication subsystems 208, a second plurality of communication subsystems 206, and a plurality of memory devices 216, according to several embodiments of the present disclosure. As used herein, "device" may refer to (but is not limited to) any of a variety of structures or combinations thereof, such as (for example) a circuit or circuit system, one or more dies, one or more modules, one or more devices, or one or more systems. Figure 2 In the embodiments described herein, memory devices 216-1, ..., 216-N may include one or more memory modules (e.g., double data rate (DDR) memory, three-dimensional (3D) cross-point memory, NAND memory, single in-line memory module, dual in-line memory module, etc.). Memory devices 216-1, ..., 216-N may include volatile memory and / or non-volatile memory. In several embodiments, memory devices 216-1, ..., 216-N may include multi-chip devices. Multi-chip devices may include several different memory types and / or memory modules. For example, a memory system may include non-volatile or volatile memory on any type of module.
[0071] Memory devices 216-1, ..., 216-N may provide main memory for computing system 200 or may be used as additional memory or storage devices for the entire computing system 100. Each memory device 216-1, ..., 216-N may include one or more arrays of memory cells, such as volatile and / or non-volatile memory cells. The array may be, for example, a flash array with a NAND architecture. Embodiments are not limited to a specific type of memory device. For example, memory devices may include RAM, ROM, DRAM, SDRAM, PCRAM, RRAM, and flash memory, etc.
[0072] In embodiments where memory devices 216-1, ..., 216-N include non-volatile memory, memory devices 216-1, ..., 216-N may be flash memory devices, such as NAND or NOR flash memory devices. However, embodiments are not limited thereto, and memory devices 216-1, ..., 216-N may include: other non-volatile memory devices, such as non-volatile random access memory devices (e.g., NVRAM, ReRAM, FeRAM, MRAM, PCM); "emerging" memory devices, such as 3-D cross-gate (3D XP) memory devices; or combinations thereof. 3D XP non-volatile memory arrays can perform bit storage based on volume resistance variations combined with stackable cross-gate format data access arrays. Furthermore, compared to many flash-based memories, 3D XP non-volatile memories can perform in-situ write operations, where non-volatile memory cells can be programmed without previously erasing them.
[0073] like Figure 2 The description indicates that host 202 may include host controller 201. Host 102 may communicate via channels 203 / 205 with a first 208-1 (“IF” 208-1) in a first plurality of communication subsystems. IF 208-1 may be a PCIe interface. IF 208-1 may be coupled to a second 208-2 (“IF” 208-2) in the first plurality of communication subsystems 208. IF 208-2 may be a PCIe XDMA interface. IF 208-2 may be coupled to a third 208-3 (“IF” 208-3) in the first plurality of communication subsystems 208. IF 208-3 may be coupled to each of a plurality of computing devices 210.
[0074] Furthermore, IF 208-2 may be coupled to a fourth entity 208-4 (“IF” 208-4) in the first plurality of communication subsystems 208. IF 208-4 may be a message passing interface (MPI). For example, host 202 may send a message that is received by IF 208-4 and held by IF 208-4 until computing device 210 or additional interface 208 retrieves the message to determine subsequent actions. Possible subsequent actions may include performing a specific function on a particular computing device 210, setting a reset vector for external interface 231, or reading / modifying the location of SRAM 233. In an alternative example, computing device 210 may write messages received by IF 208-4 for access by host 202. The host controller 201 can read messages from IF 208-4 and transfer data to or from a location in a device (e.g., SRAM 233, registers or memory device 216 in computing device 210, and / or host memory (e.g., registers, cache, or main memory)). IF 208-4 may also include host registers and / or reset vectors for controlling the selection of an external interface (e.g., interface 231). In at least one instance, external interface 231 may be JTAG interface 231, and IF 208-4 may be used for JTAG selection.
[0075] like Figure 2As described herein, multiple computing devices 210-1, 210-2, 210-3, 210-4, and 210-5 (hereinafter collectively referred to as multiple computing devices 210) may be coupled to SRAM 233. The multiple computing devices 210 may be coupled to SRAM 233 via a bus matrix. Furthermore, the multiple computing devices 210 may be coupled to additional multiple communication subsystems (e.g., multiplexers) 235-1, 235-2, and 235-3. The first multiple communication subsystem 208 and / or the additional multiple communication subsystems 235 may include circuitry and / or logic configured to allocate and deallocate resources to computing devices 210 during the execution of the operations described herein. For example, the circuitry and / or logic may allocate and / or deallocate resources to computing devices 210 during the execution of extended memory operations described herein.
[0076] like Figure 2 The text is incomplete and contains numerous errors. A proper translation is not possible without the full context. Figure 1 In contrast, multiple computing devices 210-1, 210-2, 210-3, 210-4, and 210-5 (hereinafter collectively referred to as multiple computing devices 210) may be coupled to SRAM 233. Furthermore, each of the multiple computing devices 210 may be coupled to additional communication subsystems (e.g., multiplexers) 235-1 (via SRAM 233), 235-2, and 235-3. The additional communication subsystem 235 may include circuitry and / or logic configured to allocate and deallocate resources to computing devices 210 during the execution of the operations described herein. For example, the circuitry and / or logic may allocate and / or deallocate resources to computing devices 210 during the execution of the extended memory operations described herein. Although the examples described above include SRAM coupled to each of the computing devices (e.g., in SRAM 233), the additional communication subsystems 235 may be coupled to SRAM 233. Figure 2 The cache (e.g., SRAM) may be contained within each of the computing devices, but is not limited to this. For example, the cache (e.g., SRAM) may be located in multiple locations, such as outside of device 204, inside device 204, etc.
[0077] Additional communication subsystems 235 may be coupled to second communication subsystems (e.g., interfaces of interconnect interfaces) 206-1, 206-2, 206-3, 206-4 (collectively referred to below as second communication subsystems 206). Each of the second communication subsystems 206 may be coupled to a corresponding one of the controller 212, accelerator 214, multiple SRAMs 218-1, 218-2, and peripheral components 221. In one example, the second communication subsystems 206 may be coupled to corresponding controllers 212, accelerator 214, multiple SRAMs 218, and / or peripheral components 221 via several AXI buses.
[0078] As described, the first of the second plurality of communication subsystems 206, 206-1, may be coupled to a controller (e.g., a memory controller) 212. The controller 212 may be coupled to several memory devices 216-1, ..., 216-N via several channels 207-1, ..., 207-N. The second of the second plurality of communication subsystems 206, 206-2, may be coupled to an accelerator 214 and several SRAMs 218-1, 218-2. The accelerator 214 may be coupled to a logic circuit system 213. The logic circuit system 213 may be on the same field-programmable gate array (FPGA) as the computing device 210, the first plurality of communication subsystems 208, and the second plurality of communication subsystems 206. The logic circuit system 213 may include on-chip accelerators for performing several hypothetical operations and / or for communicating with internal SRAMs (218) on the FPGA. The third of the second plurality of communication subsystems 206, 206-3, may be used to transmit data off-chip via peripheral components 221.
[0079] In some embodiments, a first plurality of AXI buses that couple IF 208-3 to a plurality of computing devices 210, couple the plurality of computing devices 210 to an additional plurality of communication subsystems 235, and couple a second plurality of communication subsystems 206 to a controller, accelerator 214, SRAM 218, or peripheral component 221 can use a faster AXI bus transmission speed than a second plurality of AXI buses that couple IF 208-2 to IF 208-3 and IF 208-4. As an example, the first plurality of AXI buses may have a transmission rate such as 100MHz in the range of 50 to 150MHz, and the second plurality of AXI buses may have a transmission rate such as 250MHz in the range of 150 to 275MHz. A third AXI bus may couple IF 208-3 to communication subsystem 206-1 and may have a faster transmission rate than the first or second plurality of AXI buses. As an example, the third AXI bus may have a transmission rate such as 300MHz in the range of 250 to 350MHz.
[0080] Figure 3 This is a functional block diagram of a computing system 300 comprising a device 304 including multiple communication subsystems 306 and multiple memory devices 316, according to several embodiments of the present disclosure. As used herein, "device" may refer to (but is not limited to) any of a variety of structures or combinations thereof, such as (for example) a circuit or circuit system, one or more dies, one or more modules, one or more devices, or one or more systems. Figure 3In the embodiments described herein, memory devices 316-1, ..., 316-N may include one or more memory modules (e.g., double data rate (DDR) memory, three-dimensional (3D) cross-point memory, NAND memory, single in-line memory module, dual in-line memory module, etc.). Memory devices 316-1, ..., 316-N may include volatile memory and / or non-volatile memory. In several embodiments, memory devices 316-1, ..., 316-N may include multi-chip devices. Multi-chip devices may include several different memory types and / or memory modules. For example, a memory system may include non-volatile or volatile memory on any type of module.
[0081] like Figure 3 As described herein, device 304 may include a computing device (e.g., a computing core). In some embodiments, device 304 may be an FPGA. Figure 1 and 2 In contrast, each port of computing device 310 can be directly coupled to multiple communication subsystems 306 (as an example, without coupling via a set of additional communication subsystems (e.g., communication subsystems 108 and 208) that can be multiplexers). Computing device 310 can be connected to multiple communication subsystems 306 via corresponding ports, including a memory port (“MemPort”) 311-1, a system port (“SystemPort”) 311-2, a peripheral port (“PeriphPort”) 311-3, and a front port (“FrontPort”) 311-4.
[0082] Memory port 311-1 can be directly coupled to communication subsystem 306-1, which is explicitly designated to receive data from memory ports and transmit the data to memory controller 312. System port 311-2 can be directly coupled to communication subsystem 306-2, which is explicitly designated to receive data from system port 311-2 and transmit the data to accelerator (e.g., on-chip accelerator) 314, whereby accelerator 314 can transmit the data to additional logic circuitry system 313. Peripheral port 311-3 can be directly coupled to communication subsystem 306-3, which is explicitly designated to receive data from peripheral port 311-3 and transmit the data to serial port 318. Front port 311-4 can be directly coupled to communication subsystem 306-4, which is explicitly designated to receive data from front port 311-4 and transmit the data to host interface 320 and subsequently to host 302 via channels 303 and / or 305. In this embodiment, a multiplexer can be used instead of a direct connection between the port and the communication subsystem for data transmission.
[0083] In some embodiments, the communication subsystem 306 may facilitate visibility between corresponding address spaces of the computing device 310. For example, the computing device 310 may store data in its memory resources in response to the receipt of data and / or files. The computing device may associate addresses (e.g., physical addresses) corresponding to locations within the memory resources of the computing device 310 where the data is stored. Additionally, the computing device 310 may resolve (e.g., break down) the addresses associated with the data into logical blocks.
[0084] In some embodiments, the zeroth logical block associated with the data may be transferred to a processing device (e.g., a Reduced Instruction Set Computing (RISC) device). A particular computing device (e.g., computing devices 110, 210, 310) may be configured to recognize a particular set of logical addresses that can be accessed by that computing device (e.g., 210-2), while other computing devices (e.g., computing devices 210-3, 210-4, etc., respectively) may be configured to recognize arrays of different logical addresses that can be accessed by those computing devices 110, 210, 310. Alternatively, a first computing device (e.g., computing device 210-2) may access a first set of logical addresses associated with that computing device 210-2, and a second computing device (e.g., computing device 210-3) may access a second set of logical addresses associated with it, and so on.
[0085] If data corresponding to a second set of logical addresses (e.g., logical addresses accessible by the second computing device 210-3) is needed at the first computing device (e.g., computing device 210-2), then the communication subsystem 306 can facilitate communication between the first computing device (e.g., computing device 210-2) and the second computing device (e.g., computing device 210-3) to allow the first computing device (e.g., computing device 210-2) to access the data corresponding to the second set of logical addresses (e.g., a set of logical addresses accessible by the second computing device 210-3). That is, the communication subsystem 308 can facilitate communication between the computing device 310 (e.g., 210-1) and additional computing devices (e.g., computing devices 210-2, 210-3, 210-4) to allow the address spaces of the computing devices to be visible to each other.
[0086] In some embodiments, communication between computing devices 110, 210, and 310 to facilitate address visibility may include receiving a message requesting access to data corresponding to a second set of logical addresses by an event queue of a first computing device (e.g., computing device 210-1), loading the requested data into the memory resources of the first computing device, and transmitting the requested data to a message buffer. Once the data is buffered by the message buffer, the data can be transmitted to a second computing device (e.g., computing device 210-2) via communication subsystem 310.
[0087] For example, during the execution of an extended memory operation, controllers 115, 215, 315 and / or the first computing device (e.g., computing device 210-1) may determine that a host command (e.g., from the host (e.g.)) is executed. Figure 1 The address specified in the command generated by the host 102 (described herein) to initiate the execution of an extended memory operation corresponds to a location in the memory resources of a second computing device (e.g., computing device 210-2) among a plurality of computing devices (110, 210). In this case, the computing device command may be generated and sent from controllers 115, 215, 315 and / or the first computing device 210-1 to the second computing device 210-2 to initiate the execution of an extended memory operation using operands stored in the memory resources of the second computing device 210-2 at the address specified by the computing device command.
[0088] In response to receiving a command from the computing device, the second computing device 210-2 can perform an extended memory operation using operands stored in the memory resources of the second computing device 210-2 at the address specified by the computing device command. This reduces command traffic between the host and the memory controller and / or computing devices 210, 310, because the host does not need to generate additional commands to trigger the execution of the extended memory operation. This increases the overall performance of the computing system by, for example, reducing the time associated with transmitting commands to and from the host.
[0089] In some embodiments, controllers 115, 215, 315 may determine that performing an extended memory operation may involve performing multiple sub-operations. For example, an extended memory operation may be parsed or decomposed into two or more sub-operations that can be performed as part of performing the entire extended memory operation. In this case, controllers 115, 215, 315 and / or communication subsystems 106, 108, 206, 208, 308 may utilize the address visibility described above to facilitate the execution of sub-operations by various computing devices 110, 210, 310. In response to the completion of a sub-operation, controllers 115, 215, 315 may cause the results of the sub-operations to be merged into a single result corresponding to the result of the extended memory operation.
[0090] In other embodiments, an application requesting data stored in computing devices 110, 210, 310 can know which computing devices 110, 210, 310 contain the requested data (e.g., it may have information corresponding to which computing devices 110, 210, 310 contain the requested data). In this example, the application may request data from the relevant computing devices 110, 210, 310, and / or addresses may be loaded into multiple computing devices 110, 120, 130 via communication subsystems 108, 106, 208, 206, 308 and accessible to the application that can request the data.
[0091] Controllers 115, 215, and 315 may be discrete circuit systems physically separate from communication subsystems 108, 106, 208, 206, and 308, and each may be provided as one or more integrated circuits that allow communication between computing devices 110, 210, and 310, memory controllers 112, 212, and 312, and / or controllers 115, 215, and 315. Non-limiting examples of communication subsystems 108, 106, 208, 206, and 308 may include XBAR or other communication subsystems that allow interconnection and / or interoperability between controllers 115, 215, and 315, computing devices 110, 210, and 310, and / or memory controllers 112, 212, and 312.
[0092] As described above, in response to controllers 115, 215, 315, communication subsystems 108, 106, 208, 206, 308 and / or the host (e.g. Figure 1 The host 102 described herein receives commands and can perform extended memory operations using data stored in computing devices 110, 210, 310 and / or from data blocks streaming through computing devices 110, 210, 310.
[0093] Figure 4 This is a functional block diagram of a computing core 410 comprising several ports 411-1, 411-2, 411-3, and 411-4, according to several embodiments of the present disclosure. The computing core 410 may include a memory management unit (MMU) 420, a physical memory protection (PMP) unit 422, and a cache 424.
[0094] The MMU 420 refers to a computer hardware component used for memory and cache operations associated with the processor. The MMU 420 can be responsible for memory management and is integrated into the processor, or in some instances, it can be located on a separate integrated circuit (IC) chip. The MMU 420 can be used for hardware memory management, which may include monitoring and regulating the processor's use of random access memory (RAM) and cache memory. The MMU 420 can be used for operating system (OS) memory management, ensuring sufficient memory resources are available for the objects and data structures of each running program. The MMU 420 can be used for application memory management, allocating the required or used memory for each individual program and then reclaiming the freed memory space when operations are complete or space becomes available.
[0095] In one embodiment, physical memory can be protected using PMP unit 422 to restrict memory access and isolate processes from each other. PMP unit 422 can be used to set memory access privileges (read, write, execute) for specified memory regions. PMP unit 422 can support eight regions using a minimum region size of 4 bytes. In some instances, PMP unit 422 can only be programmed in a privileged mode called M mode (or machine mode). PMP unit 422 can enforce permission for U mode access. However, permission for M mode can be further enforced by locking regions. Cache 424 can be an SRAM cache, a 3D crosspoint cache, etc. Cache 424 can contain 8KB, 16KB, 32KB, etc., and can include error correction coding (ECC).
[0096] In one embodiment, the computing core 410 may further include multiple ports, including a memory port 411-1, a system port 411-2, a peripheral port 411-3, and a front-end port 411-4. The memory port 411-1 may be directly coupled to a communication subsystem (such as a communication subsystem explicitly designated to receive data from the memory port 411-1) that is intended to receive data from the memory port 411-1. Figure 3 (See description). System port 411-2 can be directly coupled to a communication subsystem explicitly designated to receive data from system port 411-2. Data through system port 411-2 can be transmitted to an accelerator (e.g., an on-chip accelerator). Peripheral port 411-3 can be directly coupled to a communication subsystem explicitly designated to receive data from peripheral port 411-3, and this data can ultimately be transmitted to the serial port. Front port 411-4 can be directly coupled to a communication subsystem explicitly designated to receive data from front port 411-4, and this data can ultimately be transmitted to the host interface and subsequently to the host.
[0097] The compute core 410 can be a cache-consistent 64-bit RISC-V processor with full Linux capabilities. In some instances, memory port 411-1, system port 411-2, and peripheral port 411-3 can be outgoing ports, and front port 411-4 can be incoming ports. Instances of compute core 410 can include the U54-MC compute core. The compute core 410 can include an instruction memory system, instruction fetch unit, execution pipeline unit, data memory system, and support global, software, and timer interrupts. The instruction memory system can include a 16-kilobyte (KiB) bidirectional set-associative instruction cache. The access latency for all blocks in the instruction memory system can be one clock cycle. The instruction cache may not be consistent with the remainder of the platform memory system. Writes to the instruction memory can be synchronized with the instruction fetch system by executing the FENCE.I instruction. The instruction cache can have a line size of 64 bytes, and cache line filling can trigger burst accesses outside the compute core 410.
[0098] The instruction fetch unit may include branch prediction hardware to improve processor core performance. The branch predictor may include: a 28-entry Branch Target Buffer (BTB) that predicts the target of the branch taken; a 512-entry Branch History Table (BHT) that predicts the direction of conditional branches; and a 6-entry Return Address Stack (RAS) that predicts the target returned by the procedure. The branch predictor may have a one-cycle delay so that correctly predicted control flow instructions do not incur penalties. Incorrectly predicted control flow instructions may incur a three-cycle penalty.
[0099] The execution pipeline unit can be a single-issue, in-order pipeline. The pipeline can consist of five stages: instruction fetch, instruction decoding and register fetch, execution, data memory access, and register write-back. The pipeline can have a peak execution rate of one instruction per clock cycle and can be completely bypassed, allowing most instructions to have a one-cycle result delay. The pipeline can be interlocked with write-before-read and write-after-write hazards, allowing instructions to be scheduled to avoid pauses.
[0100] In one embodiment, the data memory system may include a DTIM interface that supports up to 8 KiB. The access latency from a core to its own DTIM can be two clock cycles for all words and three clock cycles for a smaller number. Memory requests from one core to any other core's DTIM may be less efficient than memory requests from one core to its own DTIM. Unaligned access is not supported in hardware and can lead to pitfalls that allow software emulation.
[0101] In some embodiments, the computing core 410 may include a floating-point unit (FPU) that provides full hardware support for the IEEE 754-2008 floating-point standard for 32-bit single-precision and 64-bit double-precision algorithms. The FPU may include fully pipelined product and fusion units, iterative division and square root units, a magnitude comparator, and a floating-point to integer conversion unit, providing full hardware support for subnormal and IEEE default values.
[0102] Figure 5This is a flowchart illustrating an example method 528 corresponding to an extended memory architecture according to several embodiments of the present disclosure. At block 530, method 528 may include transmitting a command from a host to at least one of a plurality of computing devices via a first communication subsystem and a second communication subsystem. The first communication subsystem may be coupled to the host, and the second communication subsystem may be coupled to both the computing device and the memory device. The transmission of the command may be in response to a request to receive a block of transmitted data for performing an operation associated with the command. In some embodiments, receiving a command to initiate execution of an operation may include receiving an address corresponding to a memory location in a specific computing device where operands corresponding to the execution of the operation are stored. For example, as described above, the address may be an address in a memory portion where data that will be used as operands in the execution of the operation is stored.
[0103] In block 532, method 528 may include transferring a data block associated with a command from a memory device to at least one of a plurality of computing devices via a second communication subsystem.
[0104] In block 534, method 528 may include, in response to the receipt of a data block, at least one of a plurality of computing devices performing an operation using the data block to reduce the data size from a first size to a second size by at least one of the plurality of computing devices. The execution of the operation may be initiated by a controller. The controller may be similar to that described herein. Figures 1 to 3 The controllers 115, 215, and 315 are described herein. In some embodiments, performing an operation may include performing an extended memory operation, as described herein. The operation may further include performing the operation by a specific computing device without receiving a host command from a host that can be coupled to the controller. In response to the completion of the operation, method 528 may include sending a notification to a host that can be coupled to the controller.
[0105] In block 536, method 528 may include transferring a reduced-size data block to a host via a first communication subsystem. Method 528 may further include using an additional controller (e.g., a memory controller) to cause the data block to be transferred from a memory device to a second plurality of communication subsystems. Method 528 may further include allocating resources corresponding to appropriate computing devices among the plurality of computing devices via the first and second plurality of communication subsystems to perform operations on the data block.
[0106] In some embodiments, the command to initiate the execution of the operation may include an address corresponding to a location in the memory array of a particular computing device, and method 528 may include storing the result of the operation in the address corresponding to the location of the particular computing device. For example, method 528 may include storing the result of the operation in the address corresponding to the memory location in the particular computing device, wherein operands corresponding to the execution of the operation are stored prior to the execution of the memory expansion operation. That is, in some embodiments, the result of the operation may be stored in the same address location of the computing device, wherein data used as operands for the operation are stored prior to the execution of the operation.
[0107] In some embodiments, method 528 may include determining, by the controller, that an operand corresponding to the execution of an operation is not stored in a particular computing device. In response to this determination, method 528 may further include determining, by the controller, that an operand corresponding to the execution of an operation is stored in a memory device coupled to a plurality of computing devices. Method 528 may further include: retrieving, from the memory device, an operand corresponding to the execution of an operation; causing, an operand corresponding to the execution of an operation to be stored in at least one of the plurality of computing devices; and / or initiating the execution of an operation using at least one computing device. The memory device may be similar to... Figure 1 The memory device 116 described herein.
[0108] In some embodiments, method 528 may further include: determining that a portion of the operation will perform at least one sub-operation; sending a command to a computing device different from a specific computing device to cause the execution of the sub-operation; and / or using a computing device different from a specific computing device to perform the sub-operation as a portion of the operation's execution. For example, in some embodiments, it may be determined that the operation will be decomposed into multiple sub-operations, and the controller may, as a portion of the operation's execution, cause different computing devices to perform different sub-operations. In some embodiments, the controller may, as a portion of the operation's execution, communicate with first and second plurality of communication subsystems (e.g., as described herein). Figures 1 to 3 The 108, 106, 208 and 308 described herein coordinately assign sub-operations to two or more of the computing devices.
[0109] Although specific embodiments have been illustrated and described herein, those skilled in the art will understand that arrangements that achieve the same computational results may be used instead of the specific embodiments shown. This disclosure is intended to cover adjustments or variations of one or more embodiments of this disclosure. It should be understood that the foregoing description is illustrative and non-limiting. Those skilled in the art will understand, upon reviewing the foregoing description, combinations of the foregoing embodiments and other embodiments not explicitly described herein. The scope of one or more embodiments of the invention includes other applications in which the foregoing structures and processes are used. Therefore, the scope of one or more embodiments of this disclosure should be determined with reference to the appended claims and the full scope of the equivalents entitled to by such claims.
[0110] In the foregoing specific embodiments, for the purpose of simplifying this disclosure, some features are grouped together in a single embodiment. This approach of the disclosure should not be interpreted as reflecting an intention that the disclosed embodiments of the disclosure must use more features than expressly stated in each claim. Rather, as reflected in the appended claims, the subject matter of the invention lies in fewer than all features of a single disclosed embodiment. Therefore, the appended claims are hereby incorporated into the detailed description, wherein each claim is an independent embodiment.
Claims
1. A device for an extended memory architecture, comprising: Multiple computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410), each comprising: Processing unit, configured to perform operations on data blocks; and A memory array configured as a cache for each corresponding processing unit; A first communication subsystem (108) coupled to each of the host (102, 202, 302) and the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410); and A plurality of second communication subsystems coupled to each of the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410), wherein each of the plurality of second communication subsystems is coupled to at least one hardware accelerator (114, 214, 314); Each of the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410) is configured to: Receive a request to perform an extended memory operation from the host (102, 202, 302); and At least a first portion of the extended memory operation is performed within each of the plurality of computing devices without transferring data associated with the first portion to the outside of the respective computing device; Commands for performing at least a second portion of the extended memory operation are sent to the at least one hardware accelerator (114, 214, 314) via one of the plurality of second communication subsystems; and Receive the result of the at least second part of the extended memory operation from the at least one hardware accelerator (114, 214, 314).
2. The device according to claim 1, wherein: The plurality of second communication subsystems include a plurality of interconnect interfaces; and The first communication subsystem (108) is a high-speed PCIe interface for interconnecting peripheral components.
3. The device of claim 1, wherein the plurality of second communication subsystems includes a controller, and the controller is coupled to a memory device (116-1, 116-N, 216-1, 216-N, 316-1, 316-N).
4. The device of claim 3, wherein the memory device (116-1, 116-N, 216-1, 216-N, 316-1, 316-N) comprises at least one of double data rate DDR memory, three-dimensional 3D cross-point memory, NAND memory, or any combination thereof.
5. The device according to claim 1, wherein the at least one hardware accelerator (114, 214, 314): It is an on-chip accelerator; Coupled to the static random access device SRAM; and It is coupled to an arithmetic logic unit (ALU), which is configured to perform arithmetic operations or logical operations or both.
6. The device of claim 1, wherein the processing unit of each of the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410) is configured with a reduced instruction set architecture.
7. The device according to any one of claims 1 to 6, wherein the extended memory operation performed on the data block includes one of the following: Floating-point accumulation operation; 32-bit complex number operations; Square root address (SQRT(addr)) operation; and Conversion operations to convert between floating-point and integer formats.
8. The device according to any one of claims 1 to 6, wherein: The at least one hardware accelerator (114, 214, 314) is configured to perform the extended memory operation by accessing non-volatile memory devices (116-1, 116-N, 216-1, 216-N, 316-1, 316-N) coupled to the plurality of second communication subsystems; and The at least one hardware accelerator (114, 214, 314) is configured to send a request to cause an additional hardware accelerator (114, 214, 314) to perform an additional portion of the extended memory operation.
9. A system with an extended memory architecture, comprising: Multiple computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410), each comprising: Processing unit, configured to perform operations on data blocks; and A memory array configured as a cache for each corresponding processing unit; A first communication subsystem (108) is coupled to the host (102, 202, 302) and each of the plurality of computing devices; A plurality of second communication subsystems coupled to each of the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410), wherein each of the plurality of second communication subsystems is coupled to: At least one hardware accelerator (114, 214, 314); and At least one internal SRAM (223); and Non-volatile memory devices (116-1, 116-N, 216-1, 216-N, 316-1, 316-N); Each of the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410) is configured to: Receive a request to perform an extended memory operation from the host (102, 202, 302); and At least a first portion of the extended memory operation is performed within each of the plurality of computing devices without transferring data associated with the at least first portion to the outside of the respective computing device; Commands performing at least a second portion of the extended memory operation are sent via one of the plurality of second communication subsystems to the at least one hardware accelerator (114, 214, 314) or the at least one internal SRAM; and Receive the result of the at least second part of performing the extended memory operation from the at least one hardware accelerator (114, 214, 314) or the at least one internal SRAM (223).
10. The system of claim 9, wherein the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410), the first communication subsystem, and the plurality of second communication subsystems are configured on a field-programmable gate array (FPGA), and the non-volatile memory devices (116-1, 116-N, 216-1, 216-N, 316-1, 316-N) are external to the FPGA.
11. The system of claim 9, wherein each of the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410) comprises: Memory ports (311-1, 411-1); System ports (311-2, 411-2); Peripheral ports (311-3, 411-3); and Front ports (311-4, 411-4).
12. The system according to claim 11, wherein: The memory ports (311-1, 411-1) are coupled to the non-volatile memory devices (116-1, 116-N, 216-1, 216-N, 316-1, 316-N) via the first communication subsystem and via at least one of the plurality of second communication subsystems; and The system ports (311-2, 411-2) are coupled to the at least one hardware accelerator (114, 214, 314) via the first communication subsystem and via at least one of the plurality of second communication subsystems; and The peripheral ports (311-3, 411-3) are coupled to the external serial port via the first communication subsystem and at least one of the plurality of second communication subsystems.
13. The system according to any one of claims 9 to 12, wherein: The first communication subsystem is directly coupled to at least one of the plurality of second communication subsystems; and At least one of the plurality of second communication subsystems is configured to transmit the data block from the non-volatile memory device (116-1, 116-N, 216-1, 216-N, 316-1, 316-N) to the first communication subsystem and the host (102, 202, 302), wherein the transmission of the data block bypasses the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410).
14. The system of claim 13, wherein the AXI interconnect that directly couples the first communication subsystem to at least one of the plurality of second communication subsystems is a faster AXI interconnect than the AXI interconnect that couples the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410) to the first communication subsystem and at least one of the plurality of second communication subsystems.
15. A method relating to an extended memory architecture, comprising: Commands are transmitted from the host (102, 202, 302) to at least one of a plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410) via the first communication subsystem (108); The data block associated with the command is transferred from the memory devices (116-1, 116-N, 216-1, 216-N, 316-1, 316-N) to at least one of the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410) via the second communication subsystem, wherein: The first communication subsystem (108) is coupled to the host (102, 202, 302) and at least one of the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410); and The second communication subsystem is coupled to at least one of the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410) and the memory devices (116-1, 116-N, 216-1, 216-N, 316-1, 316-N); In response to the command and the receipt of the data block, the at least one computing device (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410) performs an extended memory operation using the data block to reduce the data size from a first size to a second size, wherein the extended memory operation is performed without transferring the data block outside the at least one computing device, and the at least one computing device is a Reduced Instruction Set Computing (RISC) device; and The reduced-size data blocks are transmitted to the host (102, 202, 302) via the first communication subsystem (108).
16. The method of claim 15, wherein the reduced-size data block is transmitted to the host (102, 202, 302) via a PCIe interface coupled to the first communication subsystem (108).
17. The method of claim 15, further comprising: The memory controller is used to cause the data blocks to be transferred from the memory devices (116-1, 116-N, 216-1, 216-N, 316-1, 316-N) to the second communication subsystem and subsequently to the first communication subsystem (108), wherein the data blocks bypass the plurality of computing devices (110-1, 110-2, 210-1, 210-2, 210-3, 210-4, 210-5, 310, 410); and This includes performing at least one of the following via the memory controller: Read operations associated with the memory devices (116-1, 116-N, 216-1, 216-N, 316-1, 316-N); Copying operations associated with the memory devices (116-1, 116-N, 216-1, 216-N, 316-1, 316-N); Error correction operations associated with the memory devices (116-1, 116-N, 216-1, 216-N, 316-1, 316-N); or Its combination.
18. The method according to any one of claims 15 to 17, further comprising using a memory controller to cause the reduced-size data block to be transferred to the memory device (116-1, 116-N, 216-1, 216-N, 316-1, 316-N); wherein a particular set of logical addresses associated with the memory device corresponds to the at least one computing device and not to at least one other computing device among the plurality of computing devices.
Citation Information
Patent Citations
Hardware accelerators and methods for offload operations
US20180095750A1
Signal processing device and method
US20180196668A1
Hybrid logical to physical address translation for non-volatile storage devices with integrated compute module
US20180293174A1
Efficient and reliable message channel between a host system and an integrated circuit acceleration system
US20190306055A1