Processing-in-memory architecture for memory bound compute workloads
The PiM architecture addresses the challenges of executing memory-bound and compute-bound workloads by enabling advanced debugging, access control, and data refresh operations within memory devices, thereby enhancing workload execution efficiency.
Patent Information
- Application Number
- PCT/US2024/059607
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-12-11
- Publication Date
- 2025-06-19
AI Technical Summary
Existing computing systems face challenges in efficiently executing memory-bound and compute-bound workloads due to limitations in debugging, instruction/command monitoring, access control, and data refresh operations within memory devices.
The implementation of a processing-in-memory (PiM) architecture that enables debugging procedures, instruction/command monitoring, access control, and optimized data refresh operations by generating data and control signaling at the system-on-chip (SoC) to communicate with the PiM architecture through the memory device.
This approach enhances the execution efficiency of memory-bound and compute-bound workloads by improving debugging capabilities, ensuring secure access control, and optimizing data refresh operations within the memory device.
Smart Images

Figure US2024059607_19062025_PF_FP_ABST
Abstract
Description
PROCES SING-IN-MEMORY ARCHITECTURE FOR MEMORY BOUND COMPUTE WORKLOADSFIELD
[0001] This specification generally relates to an architecture for executing computations in a memory device coupled to a system-on-chip (“SoC”).BACKGROUND
[0002] Modem computing systems often incorporate a wide variety of computational processing units that each offer different computing capabilities and tradeoffs. Efficient execution of a given compute job often involves parsing computations into meaningful workloads or sub-tasks that are mapped to available processor cores of a computing system. The computations may be parsed and mapped based on suitability criteria, such as processor capability’, performance, and power. Generally, this overall process of allocating portions of a computation to appropriate processor resources is referred to as heterogeneous compute.
[0003] At least one processing unit of the computing system can be an Intellectual Property block (“IP block”) that executes a respective portion of a computational operation for different multimedia use cases. An example use case can involve processing image or speech data captured respectively by a camera or microphone on the mobile device. The SoC can use a heterogeneous compute operation to process input samples derived from the image data, the speech data, or both. An example step in the heterogeneous computing operation can include providing input samples from the image data to a host processor, such as a neural network processor or machine-learning (ML) engine, of the SoC to generate an inference output.SUMMARY
[0004] This specification describes techniques for implementing debugging and instruction / command monitoring, access control procedures, and data refresh operations in a processing-in-memory architecture (“PiM architecture”) used for memory-bounded and compute-bounded workloads that are executed in a memory device. The techniques include generating data and control signaling at a system-on-chip (SoC) that communicates with the PiM architecture by way of the memory device. The data and control signaling are routed to, and processed by, compute elements of the PiM architecture to trigger execution of certain data processing and computing operations in the memory device.
[0005] The PiM architecture can include discrete processors, processor units, register devices, buffers, etc. The PiM architecture is integrated in an example memory- device, such as a dynamic random-access memory (DRAM) or Double Data Rate (DDR) synchronous DRAM (SDRAM). The memory device may be external (or internal) to the SoC and configured for processing-in-memory- operations and compute-in-memory operations (“CiM operations”) using its multiple PiM compute elements. The memory device cooperates with the SoC to perform computations across one or more memory dies of the memory device. For example, a PiM architecture in the memory device configures, triggers, and executes computations within a computing unit of the PiM architecture based on the data and control signaling generated at the SoC.
[0006] Debugging: With regard to debugging and instruction monitoring, the SoC uses the disclosed techniques to implement and execute debugging procedures to determine a status of computations, and related operations, that are executed in the memory- device using the compute elements of the PiM architecture. Executing the debugging procedures includes identifying and mitigating errors that occur during one or more processing-in-memory (or computational) operations. The SoC can trigger execution of a compute operation in the memory device, e g., using the PiM architecture, based on a request (or command) signal generated by a memory- controller of the SoC. The request signal is passed to a compute element of the PiM architecture in the memory device to trigger execution of instructions in an instruction buffer of the PiM. For example, the request signal can be passed to a processor of the PiM architecture, a mode register of the PiM architecture, or both.
[0007] The processor uses the request signal to execute compute instructions and generate a response signal that indicates a status of operations that are performed responsive to the request signal. The operations include ML and / or neural network computations involving data (or operands), such as inputs and weights / parameters, that are stored in memory banks of the memory device, passed to the memory device from the SoC, or both. The status can indicate success or failure of a discrete step or instruction in a processing (or compute) operation for a ML or neural network inference task or workload performed based on the request. For example, the response signal indicates whether an instruction was executed successfully or whether an error occurred that indicates a requested action failed to execute. In some implementations, an element or device of the PiM architecture generates a flag that is passed to the SoC to indicate success or failure of an instruction.
[0008] Access control: With regard to access control, the SoC uses the disclosed techniques to implement and execute access control procedures that limit a range of memoryaddresses and memory cell locations of the memory device that the PiM architecture is permitted to access. The SoC determines a memory access range and generates a command that conveys the memory' access range to the memory device. The command is used to configure memory access controls at the memory device based on the memory access range specified by the command. For example, the command can trigger or configure access controls that restrict a PiM compute element to accessing only a particular portion of memory in the memory device.
[0009] The command is part of an access control procedure for limiting an ability of compute elements in a PiM architecture to access or modify data at specific locations of the memory device. For example, the access control procedure can restrict a mode register or process unit in the PiM architecture to accessing only specific rows and / or banks of the memory device that contain data and instructions required for compute operations that are assigned to a particular memory die. The command can define (e.g., “carve out”) specific rows and / or banks of memory using a physical address range for a contiguous region of the memory device.
[0010] Refresh operations: With regard to refresh optimization, the SoC uses the disclosed techniques to implement and execute data refresh procedures to optimize when and how refresh signals are issued across rows of memory' cells in a memory' device, such as a DRAM or related integrated memory device. A memory controller on the SoC can be configured to issue a threshold number of refresh commands to ensure compliance with a refresh requirement associated with the memory' device. For example, the refresh requirement for the memory device can specify' a minimum refresh interval at which memory' cells of the memory' device are required to receive a refresh (e.g., a refresh signal). The refresh operation is performed to preserve the integrity of data stored in the memory cells at each row of the memory device.
[0011] The memory controller may be integrated on the SoC and operable to interact with elements of the PiM architecture to manage and monitor refresh operations initiated at the memory device. The memory controller is configured to optimize execution of refresh operations against memory’ banks of the memory device with reference to ML and other computing operations that are executed at the PiM architecture of the memory device. In some implementations, the memory controller optimizes the timing, sequencing, and / or scheduling of when refresh operations are executed at the memory device. For example, the memory controller 105 optimizes timing and scheduling of a refresh to minimize disruptionof ML inference computations that are executed by the PiM architecture using data obtained from memory banks that require periodic refresh.
[0012] The different elements and processing devices of the PiM architecture are leveraged to optimize implementing refresh procedures in the memory device. For example, compute elements such as a mode register and a process unit of the PiM architecture are used to proactively (or dynamically) trigger a refresh of memory cells at each of the rows of the memory device. More specifically, the refresh procedures are optimized by proactively (or dynamically) triggering the refresh of the memon cells at a frequency that exceeds a minimum refresh requirement for refreshing rows of the memon' device. For example, an internal controller of the PiM architecture can proactively (or dynamically) execute refreshes at a frequency that is based on a clock cycle generated by the memory controller, rather than based on predefined refresh rate as articulated in standards document or by a vendor.
[0013] One aspect of the subject matter described in this specification can be embodied in a computer-implemented method that includes generating a first signal at a system-on-chip (SoC) and, based on the first signal, triggering execution of a compute operation in a memory device coupled to the SoC. The method includes receiving, from the memory device, a second signal generated in, or with reference to, a memory die of the memory device after triggering execution of the compute operation in the memory device; and determining a status of the compute operation based on the second signal.
[0014] These and other implementations can each optionally include one or more of the following features. For example, in some implementations, the first signal is a read request signal that embeds a PiM execute command (or instruction) using a subset of bits in the read request signal; and the second signal is a read response signal. In some implementations, the first signal is a read request signal generated by a memon' controller of the SoC; and the second signal is a read response signal generated by a processor in a PiM architecture of the memory device. In some implementations, the first signal is a read request signal generated by a memory controller of the SoC; and the second signal is a read response signal generated by a processor in a PiM architecture of the memory device.
[0015] The compute operation can include a single instruction and triggenng execution of a compute operation in the memory device comprises triggering execution of a single instruction stored in an instruction buffer of the PiM architecture in the memory' device. The compute operation can include multiple instructions and triggering execution of a compute operation in the memory device includes: triggering synchronous execution of the multipleinstructions in the memory device; or triggering asynchronous execution of the multiple instructions in the memory device.
[0016] Determining the status of the compute operation can include: determining whether an error associated with the compute operation occurred in the memory' after execution of the compute operation is triggered in the memory device. In some implementations, the method further includes: generating a configuration signal at the SoC; and triggering selection of a particular mode in the memory’ device based on the configuration signal. The first signal and the configuration signal can be the same signal. The particular mode that is selected based on the configuration signal can be a mode used by the PiM architecture during execution of the compute operation in the memory device.
[0017] In some implementations, triggering selection of a particular mode in the memory device includes: triggering selection of a first error-capture mode where the compute operation is paused or stopped in the PiM architecture in response to detection of an error associated with the compute operation after execution of the compute operation is triggered in the memory device. Triggering selection of a particular mode in the memory device can include: triggering selection of a second, different error-capture mode where execution of the compute operation in the memory continues uninterrupted in response to detection of an error associated w ith the compute operation after execution of the compute operation is triggered in the memory device.
[0018] In some implementations, the first error-capture mode and the second, different error-capture mode each cause the memory device to capture error syndrome data in response to detection of an error associated w ith the compute operation after execution of the compute operation is triggered in the memory device. The error syndrome data can include information about an instruction and a data flow operation being executed when the error occurs in the memory after execution of the compute operation is triggered in the memory device; and the error syndrome data is received by the SoC in response to a request that is transmitted to the memory device from the SoC. The memory' device that executes the compute operation includes a PiM architecture that has a processor unit and a mode register coupled to the processor unit.
[0019] Triggering selection of a particular mode in the memory’ device can include: triggering selection of the particular mode by the mode register. In some implementations, receiving the second signal generated with reference to the memory die of the memory device includes: receiving the second signal from the memory device using a threshold number of input / output (IO) pins coupled to the memory. Receiving the second signal generated withreference to the memory die of the memory device can include: receiving an indicator of success or failure, the indicator being specific to an instruction for execution by the processor unit; and receiving an identifier of the mode register coupled to the processor unit used for executing the instruction. In some implementations, the compute operation that is triggered for execution in the memory device is a portion of a compute for a machine-learning inference workload.
[0020] One aspect of the subject matter described in this specification can be embodied in an SoC that includes a processing device; and a non-transitory machine-readable storage device for storing instructions that are executable by the processing device to cause performance of operations that include: i) generating a first signal at the SoC; ii) based on the first signal, triggering execution of a compute operation in a memory device coupled to the SoC; iii) receiving, from the memory device, a second signal generated with reference to a memory die of the memory device after triggering execution of the compute operation in the memory7; and iv) determining a status of the compute operation based on the second signal.
[0021] One aspect of the subject matter described in this specification can be embodied in a computer-implemented method that includes: i) determining, at an SoC, a memory access range; ii) generating a write command based on the memory access range; iii) transmitting the write command to a memory7device using a control channel that couples the memory7device and the SoC; and iv) configuring an access control at the memory device based on the memory access range specified by the write command. The access control limits a PiM compute element of the memory device to accessing only a particular portion of memory or memoty address locations in the memory device.
[0022] These and other implementations can each optionally include one or more of the following features. For example, in some implementations, configuring an access control at the memory device includes: configuring a row-level memory7access constraint at the memory device, where the row-level memory access constraint limits the PiM compute element to accessing a particular range of rows at the memory7device.
[0023] Configuring an access control at the memory7device can include: configuring, based on the write command, the access control by writing the memory access range to a configuration space of a PiM architecture in the memory device. In some implementations, configuring an access control at the memory7device includes: defining a portion of memory, memory addresses, or memory locations in the memory device as the only portion of memory that is accessible to the PiM compute element of the memory device.
[0024] The memory access range can include a physical address range and the method includes: writing the physical address range to a range register in a memory controller of the SoC; and translating the physical address range to row, column, bank addresses of the memory device. In some implementations, generating the write command includes: generating the write command using the memory7controller of the SoC. Determining a memory access range can include: determining the memory access range using a trusted agent of a central processing unit (CPU) on the system-on-chip (SoC). The trusted agents of the CPU can include: a kernel of the CPU; and a hypervisor of the CPU.
[0025] One aspect of the subject matter described in this specification can be embodied in an SoC that includes a processing device; and a non-transitory machine-readable storage device for storing instructions that are executable by the processing device to cause performance of operations that include: i) determining, at an SoC, a memory access range; ii) generating a write command based on the memory access range; iii) transmitting the write command to a memory device using a control channel that couples the memory7device and the SoC; and iv) configuring an access control at the memory device based on the memory access range specified by the write command. The access control limits a PiM compute element of the memory7device to accessing only a particular portion of memory or memory address locations in the memory7device.
[0026] One aspect of the subject matter described in this specification can be embodied in a computer-implemented method that includes: i) determining, at an SoC, a refresh requirement for refreshing rows of a memory device coupled to the SoC; ii) generating, at the SoC, one or more refresh commands based on the refresh requirement; iii) transmitting the one or more refresh commands to multiple PiM compute elements of the memory7device coupled to the SoC; and iv) based on the one or more refresh commands transmitted to the multiple PiM compute elements, proactively (or dynamically) triggering a refresh of memory cells at each of the rows of the memory device.
[0027] These and other implementations can each optionally include one or more of the following features. For example, in some implementations, proactively triggering a refresh of memory cells at each of the rows of the memory device includes: proactively (or dynamically) triggering the refresh of the memory7cells at a frequency that exceeds a minimum refresh requirement for refreshing rows of the memory7device. The refresh of the memory cells at each of the rows can be proactively (or dynamically) triggered at the frequency that exceeds the minimum refresh requirement using the multiple PiM compute elements of the memory device. In some implementations, transmitting the one or morerefresh command signals includes: transmitting the one or more refresh command signals over one or more clock cycles generated by a memory controller of the SoC.
[0028] One aspect of the subject matter described in this specification can be embodied in an SoC that includes a processing device; and a non-transitory machine-readable storage device for storing instructions that are executable by the processing device to cause performance of operations that include: i) determining, at an SoC, a refresh requirement for refreshing rows of a memory device coupled to the SoC; ii) generating, at the SoC. one or more refresh commands based on the refresh requirement; iii) transmitting the one or more refresh commands to multiple PiM compute elements of the memory device coupled to the SoC; and iv) based on the one or more refresh commands transmitted to the multiple PiM compute elements, proactively (or dynamically) triggering a refresh of memory cells at each of the rows of the memory device.
[0029] Other implementations of this and other aspects include corresponding systems, apparatus, and computer programs, configured to perform the actions of the methods, encoded on computer storage devices. A system of one or more computers can be so configured by virtue of software, firmware, hardware, or a combination of them installed on the system that in operation causes the system to perform the actions. One or more computer programs can be so configured by virtue of having instructions that, when executed by a data processing apparatus, cause the apparatus to perform the actions.
[0030] The details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other potential features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Fig. 1 is a block diagram of an example computing system with at least one SoC.
[0032] Fig. 2 shows an example processor-in-memory architecture used for debugging processes.
[0033] Fig. 3 shows an example processor-in-memory architecture used for access control processes.
[0034] Fig. 4 shows an example processor-in-memory architecture used for memory refresh processes.
[0035] Fig. 5 is an example process used to implement a debugging procedure.
[0036] Fig. 6 is an example process used to implement an access control procedure.
[0037] Fig. 7 is an example process used to implement a memory refresh procedure.
[0038] Fig. 8 shows an example memory controller and a processor-in-memory controller that cooperate to optimize refresh operations of a memory device.
[0039] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION
[0040] Fig. 1 is a block diagram of an example computing system 100 that includes a system-on-chip 102 (“SoC 102”). The SoC 102 includes a central processing unit 104 (“CPU 104”), a memory controller 105, a shared memory' 106 (“memory7106”), a resource manager 108, and an IP / circuit block 110. In some implementations, system 100 can include multiple SoCs and any descriptions for the SoC 102 will apply equally to each of the multiple SoCs that may be included at system 100.
[0041] The CPU 104 can be a general-purpose CPU (e.g., a single or multi-core CPU). The CPU 104 generates one or more indicators, such as an app-launch indicator or a function call that is triggered in response to executing or launching an application at a user device.For example, the application can be a camera application that uses an imaging sensor to generate image data or a gaming application that requires substantial memory and graphics processing resources to render graphical content of the game. The CPU 104 also generates one or more application values, such as pixel values or frame rate. The application values may be associated with a function call, may be descriptive of an event that occurs during execution of the application, or both.
[0042] The memory 106 is a system memory, shared memory7, or both. In the example of Fig. 1, memory 106 is depicted external to circuit block 110. However, memory 106 can include portions of memory that are: i) specific to circuit block 110, ii) external to circuit block 110, or iii) both. The memory 106 can be random access memory of the SoC 102. such as static random access memory (SRAM), dynamic random access memory (DRAM), a synchronous DRAM (SDRAM), or double data rate (DDR) SDRAM.
[0043] In some implementations, aspects of memory7106 are configured as a shared scratchpad memory that supports parallel access of its memory resources by two or more processors of the circuit 110. The memory 106 can also include various other types of memory, such as high bandwidth memory7(HBM), narrow memory7(e.g., for storing 8-bit values), wide memory (e.g., for storing 16-bit or 32-bit values), etc.
[0044] The resource manager 108 is implemented in hardware and software. Aspects of the resource manager 108 can be also implemented as firmware of the SoC 102 or firmware of a device of the SoC 102, such as a DRAM memory device or the CPU 104. The resource manager 108 is a processor-in-memory (PiM) resource manager (“PiM resource manager 108”) that includes control logic implemented in hardware, software, or both. For example, the PiM resource manager 108 can include resources such as flip-flops, registers, buffers, etc. that are implemented in hardware and control logic (e.g., programmed code) that is implemented in software.
[0045] The circuit block 110 generally includes individual IP devices such as processors, processor cores, or special-purpose processing devices. For example, the circuit block 110 can include an image signal processor (ISP) 112, a tensor processing unit (TPU) 114, a digital signal processor (DSP) 116, and a graphics processing unit (GPU) 118. The circuit block 110 is referred to alternatively as an IP block 110, where the IP block can include one or more proprietary hardware elements. For example, each of the ISP 112, TPU 114, DSP 116, and GPU 118 can be a respective proprietary’ IP block (or IP device) of a particular entity or device manufacturer.
[0046] One or more aspects of the PiM resource manager 108 can be implemented as a software routine (or module) of the CPU 104, which uses one or more hardware resources of the CPU 104, such as registers, buffers, etc. The CPU 104 can be configured as an instruction and vector data processing engine that processes data obtained from a system memory of the SoC 102, such as memory 106. In some implementations, each processor (e.g., ISP 112, DSP 116, TPU 114, GPU 118) of the SoC 102 includes multiple cores and the CPU 104 and / or the PiM resource manager 108 can generate control signaling to manage and distribute memory intensive compute operations to a memory device 122 (e.g., DRAM) to minimize the processing load at each core of the processors. The control signaling is routed at system 100 using an example bus 120 of the SoC 102. The control signaling can include commands, requests, data, instructions, or combination of these.
[0047] The PiM resource manager 108 cooperates with the CPU 104. memory controller 105 and storage controller 107 to dynamically control and manage one or more compute-inmemory (CIM) operations. In some implementations, the CIM operations are executed at the SoC 102 in support of a heterogeneous compute operation between two or more processing units that are included among the IP block 110, the CPU 104, or both. More specifically, the PiM resource manager 108 is configured to generate control signaling and use one or morediscrete signal values of the control signaling to manage and execute debug operations, access control operations, and memory refresh operations at the memory device 122.
[0048] The system 100 includes an example memory device 122. The memory device 122 can include multiple memory dies. For example, the memory device 122 can include N memory die, where N is an integer greater than 1. The memory device 122 can be a dynamic random-access memory (DRAM) or Double Data Rate (DDR) synchronous DRAM (SDRAM). The memory device 122 is configured to perform or support various types of PiM operations, CiM operations, and memory-near-computing operations ('‘MnC operations”). The memory device 122 performs or supports these operations using its multiple PiM compute elements, which are described below with reference to Figs. 2-4.
[0049] The SoC 102 cooperates with the memory device 122 to perform computations across one or more memory die of the memory device 122. The computations can be for operations or workloads that involve one or more of the processors at IP block 110. Additionally, the computations can be for a heterogenous operation that spans multiple processors of IP block 110, multiple IP blocks 110, or both. In at least one example the memory device 122 may be external to the SoC 102, whereas in another example the memory device 122 may be internal to the SoC 102.
[0050] In the example of Fig. 1, system 100 and the SoC 102 is an integrated circuit of an example user / client device 130, consumer electronic device, or mobile device, where each of these devices can include items such as a smartphone 130a. tablet 130b, laptop 130c. smartwatch or wearable device 130d. The devices 130 may also include other items such as an eNotebook, Netbook, smart speaker, or mobile computer. In some implementations, the system 100 and the SoC 102 are integrated circuits of a desktop computer, network server, or related cloud-based asset.
[0051] Fig. 2 shows an example processor-in-memory' architecture 200 for debugging processes performed using the SoC 102, the memory device 122, or both. In the example of Fig. 2, the memory device 122 includes a first memory7die-1 that includes one or more memory arrays 210-1 and a second memory die-2 that includes one or more memory arrays 210-2.
[0052] The PiM architecture 200 includes multiple compute elements. For example, the PiM architecture 200 includes mode registers 204-1, 204-2 at memory' die-1 and memory' die- 2, respectively. The PiM architecture 200 further includes process units 206-1. 206-2 at memory die-1 and memory die-2, respectively. Each process unit 206-1. 206-2 can include a processor, a processor unit, or a processor core, such as a CPU. Each process unit 206-1,206-2 can also include an example computation unit such as an arithmetic logic unit (ALU) or multiply-accumulate cell (MAC). The PiM architecture 200, and other PiM architectures disclosed herein, can include multiple control and status registers (CSR) that represent auxiliary registers that are used for reading status and changing configurations and modes of operation in a PiM architecture.
[0053] In some implementations, the PiM architecture 200 is included in the memory’ device 122 as multiple discrete integrated circuits, where each integrated circuit is local to a given memory die (e.g., die-1 and die-2) and interacts or communicates with arrays of memory’ cells at that memory' die. For example, the PiM architecture 200 can include compute elements that are replicated and distributed across each of the memory' die in the memory device 122. In some other implementations, the PiM architecture 200 is included in the memory device 122 as a single integrated circuit that interacts or communicates with each memory die of the memory device 122, including the arrays of memory cells at each memory die.
[0054] The PiM operations can include standard CPU functions, whereas the CiM operations and MnC operations can include standard arithmetic operations, such as computations normally performed by an ALU or MAC. The CiM operations and MnC operations can also include computational functions of a TPU 114, such as multiplication and addition operations for matrix math, vector computations, linear algebra, and dot-product accumulations. In some implementations, each of the PiM operations, CiM operations, and MnC operations are performed in support of machine-learning computations, neural network computations, or both.
[0055] The PiM architecture 200 can include a register or other portion of memory for storing error data 220-1, 220-2 for a respective memory die. Each of the error data 220-1, 220-2 describes an error syndrome for an error that occurred during a compute operation at a corresponding memory' die of the memory7device 122. The mode registers 204-1, 204-2 are used to control or trigger selection of a particular mode in a PiM architecture, such as an error-capture mode, continuous / uninterrupted execution mode, etc.
[0056] In some implementations, a particular mode is selected based on bit values of the mode registers 204-1, 204-2. For example, to trigger or select one or more error capture modes, a single bit or a sequence of bits are defined for use in the mode register. The bit values of a mode register 204-1, 204-2 are set by commands / instructions received from the SoC 102. local commands / instructions executed by a processing unit of the PiM architecture 200, or both. For example, the PiM mode registers values can be programmed using a moderegister write command (MRW), whereas other PiM registers, such as the PiM CSRs, can be programmed using a PiM control register write command.
[0057] A subset of bits in an example read request signal can be used to trigger an error capture mode of the PiM architecture 200. In some implementations, a particular error capture mode is triggered and / or configured in the PiM architecture 200 when a command from the SoC 102 triggers a batch execution of multiple PiM instructions in the memory device 122. In this (and other) implementation(s), a first error-capture mode of the PiM architecture is configured to pause execution of a PiM operation in response to detection or occurrence of an error, whereas a second, different error-capture mode of the PiM architecture is configured to continue execution of a PiM operation in response to detection or occurrence of an error. Additional aspects of the error-capture mode(s) are described in more detail below with reference to the example process of Fig. 5.
[0058] Fig. 3 shows an example processor-in-memory architecture 300 for access control processes performed using the SoC 102, the memory7device 122, or both. The PiM architecture 300 can be an extension of PiM architecture 200. In the example of Fig. 3, the memory controller 105 of the SoC 102 includes a range register 304 and the memory device 122 includes a first row limiter 310-1 and a second row limiter 310-2. The range register 304 is configured to store address values that correspond to physical memory7locations of the memory device 122. Each of the first row limiter 310-1 and the second row limiter 310-2 can be implemented in hardware, software, or both. In some implementations, the first and second row limiters 310-1 , 310-2 define a configuration space of the PiM architecture.
[0059] The SoC 102 determines a memory access range and the memory7controller 105 generates a command that conveys the memory access range to the memory7device 122. The memory access range can be represented as a physical address range. The SoC 102 determines the memory access range using a trusted agent of CPU 104. The trusted agent(s) can be a kernel of the CPU 104, a hypervisor of the CPU 104, or both. The memory access range is represented or expressed by the CPU 104 as a range of physical addresses and the SoC 102 writes the physical address range to a range register 304 in the memory controller 105.
[0060] The memory controller 105 uses the range register 304 to maintain a range of physical addresses for memory7locations that are accessible by compute elements of the PiM architecture of the memory7device 122. The memory controller 105 uses the range of physical addresses stored in the range register 304 to generate and transmit a command, such as an access control message, to compute elements of a PiM architecture in the memorydevice 122. The command can be a write command used to configure memory access controls at the memory device 122 based on the memory access range specified by the command. In some implementations, the memory device 122 is a DRAM device that includes one or more configuration mode registers. In this implementation the write command can be an MR write command corresponding to an MR transaction to configure a certain access control mode. For example, the access control mode is configured based on bit values set in a configuration mode register of the memory device 122.
[0061] The write command defines a range of physical addresses that are accessible by a PiM system of a memory' die in the memory' device 122. The PiM system includes, or is based on, a corresponding PiM architecture 300. In some implementations, the memory controller 105 is configured to translate the physical address range to row, column, bank addresses of the memory' device 122 and then generate the write command to convey the row, column, bank addresses to the memory device 122. The row limiter can use the row, column, bank addresses to generate control signals that define or “can e ouf ’ a portion of memory' as the only portion of memory accessible to PiM compute elements of the memory device. The portion of memory can be a set of rows in a corresponding DRAM bank of the memory device 122.
[0062] Each row limiter 310-1, 310-2 can interact with a respective process unit 206-1, 206-2 to configure memory access controls at the memory device 122 based on the row, column, bank addresses indicated by the write command. For example, the row limiter 310-1 can generate row limit signals based on a memory’ access range specified by commands executed at the process unit 206-1 of corresponding memory' die-1. Similarly, the row limiter 310-2 can generate row limit signals based on a memory' access range specified by commands executed at the process unit 206-2 of corresponding memory die-2.
[0063] In some implementations, compute elements such as the row limiter 310-1, 310-2 and respective process unit 206-1, 206-2 of PiM architecture are configured for row-level constraint. For example, the PiM system can be configured to access any column and / or any bank within the memory device 122, but is constrained to accessing only a subset of rows within a bank. Compute elements of the PiM architecture 300 can be configured to implement differing types of access constraints using various compute elements of the PiM architecture 300. In some implementations, memory' access controls are configured by writing the memory' access range in a mode register of a configuration space in the memorydevice 122 based on the write command.
[0064] Fig. 4 shows an example processor-in-memory architecture 400 used for memory refresh processes performed using the SoC 102. the memory device 122, or both. The PiM architecture 400 can be an extension of PiM architecture 200 and PiM architecture 300. In the example of Fig. 4, the SoC 102 communicates with compute elements of the PiM architecture 400 to optimize refresh procedures executed at the memory device 122. For example, compute elements such as a mode register 204- 1 , 204-2 and a respective process unit 206-1. 206-2 are used to proactively (or dynamically) trigger a refresh of memory cells at each of the rows across each DRAM bank of memory die-1.
[0065] In some implementations, the refresh procedures are optimized by proactively (or dynamically) triggering the refresh of those memory cells at a frequency that exceeds a minimum refresh requirement for refreshing rows of the memory device 122. For example, an internal controller of the process unit 204-1 can proactively or dynamically execute refreshes at a frequency or refresh rate that is based on a clock cycle generated by the memory7controller 105. For example, an internal controller of the PiM architecture can proactively execute refreshes at a 16 millisecond (ms) interval, rather than the minimum 32 ms interval that may be required for the memory device. The refresh operation is represented by a refresh signal triggered by a refresh command processed at the PiM architecture of the memory7device 122. Refresh procedures and refresh optimization are described in more detail below with reference to the example of Fig. 8.
[0066] For clarity, the features, devices, and / or compute elements that are described with reference to PiM architectures 200, 300, 400 can correspond to a single PiM architecture or can be integrated into a single PiM architecture. Additionally, each of PiM architecture 200, 300, 400 can include more or fewer process / ing units, mode registers, and / or row limiters, irrespective of the quantities of these elements that are shown in the corresponding figures that include examples of these elements.
[0067] Fig. 5 is an example process 500 used to execute a debugging procedure. Process 500 is implemented or executed at system 100 using at least the SoC 102 and memory device 122 described above. Hence, descriptions of process 500 will reference the above-mentioned computing resources of system 100. In some examples, the steps or actions of process 500 are enabled by programmed software instructions, firmware instructions, or both. Each type of instruction may be stored in a non- transitory machine-readable storage device and is executable by one or more of the processors or other resources described in this specification.
[0068] Referring again to process 500, the SoC 102 generates a first signal at the SoC (502). For example, the first signal can be a read request signal generated by the memorycontroller 105. In some implementations, in addition to its primary purpose of requesting, reading, or obtaining data from the memory device 122, a read request signal (and its associated data structure of bits) can be re-purposed to provide a secondary purpose of triggering execution of a compute operation in a PiM module of the memory device 122.
[0069] The SoC 102 triggers execution of a compute operation in a memory device coupled to the SoC 102 based on the first signal (504). For example, the compute operation can be a CiM operation where neural network inputs stored at the memory device 122 are retrieved and used as operands for a multiplication or addition operation performed at an ALU, MAC, or compute unit of a PiM architecture 200, 300, 400 in the memory device 122.
[0070] In some implementations, the first signal is an execute signal defined based on an execute protocol that is managed and / or controlled using the CPU 104, the PiM resource manager 108, or both. In a related implementation, the read (or read request) signal is defined based on a structure of bits, such as 16-bits, 32-bits, etc. The SoC 102 can be configured to “piggyback’" an example execute command onto an existing read (or write) signal transmitted to the PiM architecture 200, 300, 400 of the memory device 122. More specifically, an example 32-bit read (or write) request signal can include one or more unused bits (e g., one or two bits) and the SoC 102 can “piggyback” an execute command onto a read / write signal using the one or more unused bits of that signal.
[0071] For example, the SoC 102 can: i) define a HIGH bit value (binary “1"’) to indicate the initial position in a sequence of an unused bits of a 32-bit read / write signal, ii) embed the execute command in the read / write signal by setting the HIGH bit value at that initial / particular bit position for the sequence of unused bits, and iii) use the HIGH bit value as a trigger / command for executing a compute operation at the memory device 122. The SoC 102, a host device (e.g., TPU 114), or the memory controller 105 can also embed other execute or config commands in the read / write signal by setting bit values of other bits in the sequence to a particular HIGH or LOW bit value (binary “0”).
[0072] In some cases, the read request signal specifies an actual read address, that specifies a memory’ location from which data is to be obtained or read, as well as additional operation trigger bits, whereas the read response signal returns the read or requested data as well as error / status information indicating the execution status of the PiM command that was embedded in the prior read request signal. The execute command can be embedded in a read / write signal that is transmitted from the SoC 102 directly to a PiM element in the memory device 122. In some other cases, the execute command can be embedded in a read / write signal that is transmitted to a non-PiM element of the memory device 122.Embedding execute commands in signals transmited to non-PiM elements is useful for minimizing wasted cycles and interruptions at the PiM architecture of memory device 122, which conserves and / or maximizes clock cycle usage across the system 100.
[0073] To further enhance efficiencies, including clock usage across system 100, some (or all) read / write signals and individual execute commands can be configured to trigger execution of multiple instructions, operations, and / or compute operations at a PiM architecture of the memory device 122. For example, a command or instruction from the SoC 102 can be passed to memory device 122 to trigger execution of N (e g., multiple) instructions in an instruction register of a PiM architecture 200 in the memory device 122. In this and other examples, N can be an integer greater than 1. The command / instruction sent from the SoC 122 can be configured to define the number of instructions, for example, by embedding or defining a value for N in the command / instruction.
[0074] In some implementations, the PiM resource manager 108 is configured to determine a threshold number of commands that are required to ensure a PiM architecture can continuously execute a subset of computations required for a given task. This continuous execution of computations can correspond to. or be enabled by. the synchronous execution of instructions described below. The PiM resource manager 108 can determine the threshold number of commands based on an execution bandwidth of a particular PiM architecture at the memory device 122. In some implementations, the PiM execution bandwidth is defined by a minimum number of commands / instructions necessary to ensure the PiM architecture can continuously execute computations required for a given workload.
[0075] The SoC 102 is configured to provide a subset of instructions for buffering at an example instruction buffer of the PiM architecture. In addition to the execute protocol described above, the SoC 102 can also include an instruction protocol that is managed and / or controlled using the CPU 104, the PiM resource manager 108, or both. The instruction protocol can define an instruction signal used to pass a subset of instructions for buffering at an instruction buffer of a PiM architecture in the memory device 122. For instruction tracking, a mode register of the PiM architecture 200 can be configured to function as a program counter that tracks instruction execution, delays, and / or stalls.
[0076] In some implementations, the SoC 102 is configured to construct a subset of commands that include opcodes for encoding specific types of ML / NN computations that a PiM module will perform on a particular sequence of data. The SoC 102 can also embed address configuration commands in a read request signal or simply define a new or unique 16-bit, 32-bit, or 64-bit communication / command signal for configuring a PiM / CiMoperation in the PiM architecture 200. For example, a new 32-bit command signal can indicate a range of memory addresses that store data (e.g., operands) for a particular PiM / CiM operation.
[0077] An example compute operation performed at the PiM architecture 200 can include or span multiple instructions. For a subset of instructions that are buffered in a PiM module of the PiM architecture 200, the SoC 102 and / or PiM module can trigger synchronous execution of the instructions to perform a compute operation. Alternatively, for a subset of instructions that are buffered in a PiM module of the PiM architecture 200, the SoC 102 and / or PiM module can also trigger asynchronous execution of the instructions to perform a given compute operation. In some implementations, aspects of an instruction protocol and an execute protocol are integrated as a single protocol. For example, the SoC 102 can be configured to integrate or combine features for defining processing instructions and execute commands into a single protocol.
[0078] The SoC 102 receives a second signal from the memory (506). The second signal may be a read response signal generated at, or with reference to, a memory die of the memory device 122 after execution of the compute operation is triggered in the memory device. In some implementations, the second signal (e.g., a read response signal) conveys or includes data describing one or more errors that occurred during the compute operation. The SoC 102 determines the status of the compute operation based on the second signal (508). The system 100 can convey success / failure information and certain status information using the second signal generated at, or with reference to, the memory device 122.
[0079] For example, receiving the second signal (e.g., read response or other type of signal) includes receiving an indicator of success or failure of a particular instruction executed by the processing unit 206-1, 206-2. The indicator is specific to an instruction for execution by the processing unit and includes information about the execution status of the instruction. In some implementations, the second signal is a unique response signal, distinct from a read response signal, that is generated within the PiM architecture 200 using a process unit 206-1, 206-2. In some cases, a subset of bits can be generated using a device or computing element of the PiM architecture 200. where the subset of bits may be included in. or appended to, a data structure corresponding to the second / response signal. The subset of bits can encode a parameter that indicates whether the particular instruction was executed successfully, failed to execute, or stalled. The parameter can also indicate that execution of the particular instruction was delayed.
[0080] In some implementations, receiving the second signal includes receiving an identifier of a particular mode register ("‘mode register ID ’’) that is coupled to a processing unit 206-1, 206-2 used to execute the instruction. The mode registers 204-1, 204-2 can specify, to a particular processing unit or PiM control logic, the quantity of compute instructions the processing unit 206-1, 206-2 is assigned to execute. In other words, a mode register of the PiM architecture 200 is configured to tell the PiM control logic how many instructions its processing unit will execute for a given workload. Additionally, the mode register can also include error mode config values that specify the error capture mode for a given workload. In some implementations, the mode registers 204-1, 204-2 are configured to capture or track a program counter’s location within a list of instructions being executed within a PiM architecture 200.
[0081] The read response signal, e.g., the second signal, can be transmitted using one or more data lines (e.g., D / Q lines). The system 100 can be configured to re-use data lines that are typically reserved for providing or obtaining data stored in the memory device 122. In some implementations, the system 100 includes or selects a particular PiM mode, e.g., using the mode registers 204-1, 204-2. where a particular number of data lines are (or will be) used to indicate success or failure of a given operation or instruction execution. For example, the system 100 in this particular PiM mode, the system 100 can toggle a particular number of data lines needed to indicate success or failure of an operation or instruction.
[0082] The data lines represent a communication interface between the SoC 102 and the memory device 122, where the interface includes 16 lines. In general, data tines can be referred to alternatively as D / Q lines or I / O lines. In some implementations, the SoC 102 is configured to receive the response signal (e.g., a second signal) from the memory device 122 using a threshold number of input / output (IO) pins that are coupled to the memory device 122. For example, to indicate whether a PiM instruction executed successfully, the system 100 can be configured to toggle just one of the D / Q tines instead of all 16 tines. For example, when indicating successful or failed execution of an instruction at the memory7device 122, the system 100 can toggle a single (or less than all) D / Q line(s) as a power gating or power control mechanism of the system. In some implementations, the system 100 can use short / er data bursts to convey certain status information.
[0083] Errors that occur during a compute operation can be captured based on an errorcapture mode of a PiM architecture in memory7device 122. A particular error-capture mode defines a manner in which errors that occur during execution of a memory-bound compute operation are captured as error data 220-1, 220-2 by the PiM architecture. The particularmode can be selected based on a configuration signal passed to the PiM architecture from the SoC 102. For example, a resource of the SoC 102 can generate a configuration signal at the SoC and transmit the configuration signal to the PiM architecture of the memory device 122.
[0084] In some implementations, the configuration signal is passed to a PiM architecture as a read / write request signal or some other signal. For example, the configuration signal can be embedded in an existing signal (e.g., a read / write signal) using one or more unused bits of that existing signal. Rather than embedding configuration information into an existing read / write signal, the SoC 102 can be also configured to include a separate / unique configuration signal to convey compute configuration settings to the PiM architecture 200. The SoC 102 is configured to trigger selection of a particular mode in the memory device 122 based on respective values of one or more bits that represent the configuration signal.
[0085] An example error-capture mode can be configured such that, after execution of a compute operation is triggered at a PiM architecture of the memory device 122, the compute operation is automatically paused or stopped in the memory device 122 in response to detection of an error associated with the compute operation. For example, the compute operation can be paused or stopped using a processor unit 206-1, 206-2, or related control logic of the PiM architecture. In another example, a second, different error-capture mode can be configured such that, after execution of a compute operation is triggered in the memory' device 122, execution of the compute operation occurs uninterrupted in response to detection of an error associated with the compute operation.
[0086] During this continuous-execution-error-capture mode, the error capture and detection logic of the PiM architecture can be configured with an efficiency bias, for example, to conserve memory resources and capture cycles of the PiM architecture. The efficiency bias can be configured such that, after detecting a first instance of a given error ty pe, the PiM architecture discards capturing a second instance of the same (or substantially similar) error type. In some implementations, the error capture and detection logic of the PiM architecture includes an overflow flag associated with the given error type. The overflow flag is used to indicate two or more detected instances or occurrences of this given error type.
[0087] For each error-capture mode, the PiM architecture of the memory device 122 is configured to capture error syndrome data during execution of a compute operation and in response to detection of an error associated with the compute operation. The error syndrome data includes information about an instruction and / or a data flow operation that was, or is being, executed when the error occurs during the computing operation. In someimplementations, the error syndrome data is generated based on a stack or memory dump action performed to recover instruction and / or data flow information for a subsequent debugging operation.
[0088] The error syndrome data can be configured to capture one or more of: an index or address of the instruction that caused a particular error to occur, the reason for the error (e.g., undefined instruction, access out-of-bound, etc.), and the memory’ address that the instruction was attempting to access. The error syndrome data, e.g., error data 220-1, 220-2, can be captured and / or stored in an example error register of the PiM architecture, in the mode registers 204-1, 204 of the PiM architecture, or both.
[0089] In some implementations, the mode registers 204-1, 204 can be configured as an overflow resource for the error registers of a PiM architecture 200, 300, 400, to capture error syndrome data that exceeds a storage capacity of the error registers. The SoC 102 is configured to retrieve the error syndrome data from the error register, for example, using a read / write response signal as described above. For example, error syndrome data can be sent or transmitted to the SoC 102 from the memory device 122 upon request from the SoC 102. In some implementations, error syndrome data, or portions thereof, can be embedded into a response signal using one or more unused bits of the response signal.
[0090] Fig. 6 is an example process 600 for executing an access control procedure. Process 600 is also implemented or executed at system 100 using at least the SoC 102 and memory device 122 described above. Hence, descriptions of process 600 will reference the above-mentioned computing resources of system 100. In some examples, the steps or actions of process 600 are enabled by programmed software instructions, firmware instructions, or both. Each ty pe of instruction may be stored in a non-transitory machine-readable storage device and is executable by one or more of the processors or other resources described in this specification.
[0091] Referring again to process 600, the SoC 102 determines a memory access range (602). For example, the CPU 104 or PiM resource manager 108 can dynamically determine the memory access range based on the t pe of PiM operation or the t pe of CiM operation being performed at the memory device 122. In some implementations, the CPU 104 or PiM resource manager 108 can dynamically determine the refresh requirement based on the type of memory’ device 122 that is coupled to the SoC 102.
[0092] The SoC 102 generates a write command based on the memory access range (604). For example, the SoC 102 can use the CPU 104 and / or the PiM resource manager 108 to generate control signals that cause the memory controller 105 to generate a writecommand. In some implementations, the memory controller 105 is configured to read a memory access range (e.g., an address range) specified in its range register 304 and then generate the write command to pass or transmit the address range to a row limiter 310-1, 310- 2 in PiM architecture 300.
[0093] The SoC 102 transmits the write command to a memory device 122 using a control channel that couples the memory’ device 122 and the SoC (606). In some implementations, the memory controller 105 interprets or translates the address range to one or more row, bank, columns of the memory device 122 and passes the address range to the row limiter 310-1, 310-2 using the corresponding row, bank, columns for the address range.
[0094] The system 100 configures access controls at the memory device 122 based on the write command (608). More specifically, the SoC 102 interacts with the memory’ device 122 to configure an access control at the memory device 122 based on the memory access range specified by the write command. In some implementations, the PiM architecture includes a configuration space comprising the row limiters 310-1, 310-2 and one or more configuration mode registers. The system 100 can configure access controls at the memory device 122 by writing the memory access range to a configuration space of a PiM architecture based on the write command. For example, a bit sequence(s) of the write command, e.g., an MR write command, can configure a particular access control mode by setting certain values of a configuration mode register.
[0095] In some implementations, the SoC 102 configures an access control at the memory device 122 by configuring a row-level memory' access constraint at a particular memory’ die of the memory’ device 122. The row-level memory’ access constraint limits a PiM compute element to accessing a particular range of rows at a given bank of the memory die. The SoC 102 configures the row-level memory access constraint in coordination with the PiM architecture, for example, by triggering a certain access control configuration using bits, binary values, or signal values of the write command.
[0096] The access control limits a PiM compute element of the memory' device 122 to accessing only a particular portion of the memory cells or memory arrays in the memory device (610). For example, a bit sequence (s) of the write command can convey a particular address range that is used by’ the row limiters 310-1, 310-2 to constrain (or limit) access to a particular row, column, bank of the memory’ device 122, as indicated by the write command.
[0097] Fig. 7 is an example process 700 for executing a memory refresh procedure. Like other processes described herein, process 700 is also implemented or executed at system 100 using at least the SoC 102 and memory device 122 described above. Hence, descriptions ofprocess 700 will reference the above-mentioned computing resources of system 100. In some examples, the steps or actions of process 700 are enabled by programmed software instructions, firmware instructions, or both. Each type of instruction may be stored in a non- transitory machine-readable storage device and is executable by one or more of the processors or other resources described in this specification.
[0098] Referring again to process 700, the SoC 102 determines a refresh requirement for refreshing rows of a memory device coupled to the SoC (702). For example, the CPU 104 or PiM resource manager 108 can dynamically determine the refresh requirement based on the ty pe of memory device 122 that is coupled to the SoC 102, such as w hether the memory device 122 is a particular ty pe of DRAM or DDR SRAM. The CPU 104 or PiM resource manager 108 can also dynamically determine the refresh requirement based on attributes of the memory cells in the memory device 122. In some implementations, the refresh requirement is predetermined and stored locally at the SoC 102 (e.g., in firmware).
[0099] The SoC 102 generates one or more refresh commands based on the refresh requirement (704). The SoC 102 transmits the one or more refresh commands to multiple processing-in-memory (PiM) compute elements of the memory device coupled to the SoC (706). For example, the SoC 102 transmits the one or more refresh commands to at least the memory7arrays 210-1, 220-2 of the memory' device 122. The SoC 102 proactively triggers a refresh of memory cells at one or more of the rows of the memory device (708). As indicated above, a refresh is an example signaling operation performed against certain memory cells or rows of memory device 122 to preserve the integrity of data stored in the memory cells at each row of the memory device.
[0100] The system 100 proactively optimizes refresh procedures by, for example, triggering the refresh of those memory cells at a frequency that exceeds a minimum refresh requirement for refreshing row s of the memory device 122. In some implementations, system 100 is configured to optimize: i) a timing of issuing and / or executing refresh commands and ii) a total number of refresh commands by tracking memory' access patterns of a PiM command or operation that is executed at the PiM architecture. For example, the system 100 can determine a refresh schedule that meets or exceeds a minimum refresh requirement of the memory device 122 by analyzing past, ongoing, or immediate future PiM operations and the corresponding memory’ access patterns of these operations.
[0101] The SoC 102 triggers the refresh based on the one or more refresh commands that are transmitted to the compute elements of PiM architecture. For example, compute elements such as a mode register 204- 1 , 204-2 and a respective process unit 206-1 , 206-2 can be usedto proactively trigger a refresh of memory cells at each of the rows across each DRAM bank of memory die-1.
[0102] In some implementations, for a given PiM operation, a refresh logic of the SoC 102 can configure a PiM execute command (e.g., a read request signal) to include a refresh signal. The refresh signal can be passed to the memory device 122 to cause corresponding operand banks (odd / even) accessed during a PiM operation to receive a refresh signal after execution of a PiM operation involving those operand banks. For example, the refresh operations can be performed against at least two rows of a memory die and after executing all compute operations for a given task or workload. In other implementations, a PiM command that originates and is executed locally in the PiM architecture, e.g., via a process unit 206-1, 206-2, can be configured to embed a refresh command to trigger a refresh of the operand banks.
[0103] The respective steps of process 500, process 600, or process 700 can be performed at a hardware integrated circuit as part of a larger compute operation to generate a machinelearning (ML) output, including an output for a neural network layer of a neural network that implements one or more ML models. For example, the output can be a portion of a computation for a ML task or inference workload to generate an image processing, speech processing, or image recognition output.
[0104] As indicated above, a portion of the integrated circuit can include a specialpurpose neural network processor or hardware ML accelerator configured to accelerate computations for generating different types of data processing outputs. In some implementations, one or more of the PiM operations, CiM operations, or MnC operations are performed by the memory device 122 to support or enable accelerating computations for generating different types of data processing outputs.
[0105] Fig. 8 shows an example refresh system 800 that includes the memory controller 105, a host processor 802, and the PiM architecture 400. The refresh system 800 is a subsystem of the system 100. In the example of Fig. 8, the host processor 802 can be a specialpurpose integrated circuit, such as TPU 114, and includes an example PiM controller 814 that cooperates with the memory controller 105 to optimize refresh operations of a memorydevice. In some implementations, the TPU 114 is configured to execute an AI / ML model, such as a large language model (LLM). The TPU 114 uses the PiM controller 814 and memory controller 105 to offload a portion of the compute for executing an LLM workload to the PiM architecture 400.
[0106] The PiM controller 814 generates configuration information, compute commands, and PiM execute instructions 806 for offloading compute tasks to the PiM architecture 400. The PiM controller 814 cooperates with the memory controller 105 to transmit the configuration information, compute commands, and PiM execute instructions 806 (collectively “PiM execute instructions 806"’) to the PiM architecture 400. The PiM execute instructions 806 can include information such as operation types, operands, and execution modes that are related to a PiM internal datapath and control. The memory controller 105 uses this information generated by the host processor 802 and PiM controller 814 to infer a subset of future memory7banks that the host processor 802 will likely access to execute a portion of its ML workload at the PiM architecture 400.
[0107] For example, the memory controller 105 includes refresh optimization logic 804 that configures and executes data refresh procedures to optimize when and how refresh signals 810 are issued across rows of memory7banks 830 in the memory device. The refresh optimization logic 804 is configured to schedule refresh intervals of the memory banks based on the operation types and compute durations indicated by the PiM execute instructions 806. In particular, the refresh optimization logic 804 can determine, for different workloads, an optimal refresh interval 820 that maximizes total PiM compute time (e.g., by minimizing PiM compute downtime), which leads to enhanced overall performance at the PiM architecture 400.
[0108] The memory controller 105 leverages its timing and scheduling module 808 to optimize a timing, sequencing, and / or scheduling of when refresh operations are executed at memory banks 830. For example, the refresh optimization logic 804 uses constraints enforced by its timing and scheduling module 808 as guidelines to optimize its refresh control behavior. In some implementations, the memory controller 105 transmits refresh control signals 810 that can establish a refresh configuration (e.g., refresh_config) at the PiM architecture 400. The refresh control signals 810 can be generated based on the timing and scheduling module 808 to optimize the timing, sequencing, and / or scheduling of when refresh operations are executed at memory7banks 830.
[0109] For example, the memory controller 105 optimizes timing and scheduling of refreshes to minimize disruption of ML inference computations at PiM modules that process data obtained from memory banks that are subject to periodic refresh. The PiM controller 814 cooperates with the refresh optimization logic 804 to evaluate performance and power impacts for different refresh rates and determine a refresh scheme that minimizes the performance and power overhead for a given refresh operation. In some implementations,each of the PiM controller 814 and refresh optimization logic 804 is configured to exploit known windows of time to optimize refresh execution. For example, the PiM controller 814 can identify time windows that coincide with a natural pause (e.g.. bank closing / opening) in an ML inference computation to optimize execution of per bank refresh operations.
[0110] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory program carrier for execution by, or to control the operation of, data processing apparatus.
[0111] Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0112] The term “computing system” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g.. code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0113] A computer program (which may also be referred to or described as a program, software, a software application, a module, a software module, a script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0114] A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one ormore scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication netw ork.
[0115] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e g., an FPGA (field programmable gate array), an ASIC (application specific integrated circuit), or a GPGPU (General purpose graphics processing unit).
[0116] Computers suitable for the execution of a computer program include, by way of example, can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. Some elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g.. magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0117] Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by. or incorporated in. special purpose logic circuitry.
[0118] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with auser as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user's client device in response to requests received from the web browser.
[0119] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), e.g., the Internet.
[0120] The computing system can include clients and servers. A client and server are generally remote from each other and ty pically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0121] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0122] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achievedesirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0123] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.
Claims
What is claimed is:
1. A computer-implemented method comprising: generating a first signal at a system-on-chip (SoC); based on the first signal, triggering execution of a compute operation in a memory7device coupled to the SoC: receiving, from the memory device, a second signal generated with reference to a memory die of the memory device after triggering execution of the compute operation in the memory' device; and determining a status of the compute operation based on the second signal.
2. The method of claim 1 , wherein: the first signal is a read request signal that embeds a processor-in-memory (“PiM”) execute command using a subset of bits in the read request signal; and the second signal is a read response signal.
3. The method of claim 1 or 2, wherein: the first signal is a read request signal generated by a memory controller of the SoC; and the second signal is a read response signal generated by a processor in a PiM architecture of the memory7device.
4. The method of any one of claims 1 to 3, wherein the compute operation comprises a single instruction and triggering execution of a compute operation in the memory device comprises: triggering execution of a single instruction stored in an instruction buffer of the PiM architecture in the memory device.
5. The method of any one of claims 1 to 3. wherein the compute operation comprises a plurality of instructions and triggering execution of a compute operation in the memory device comprises: triggering synchronous execution of the plurality7of instructions in the memory device; ortriggering asynchronous execution of the plurality of instructions in the memory' device.
6. The method of any preceding claim, wherein determining the status of the compute operation comprises: determining whether an error associated with the compute operation occurred in the memory after execution of the compute operation is triggered in the memory device.
7. The method of any preceding claim, further comprising: generating a configuration signal at the SoC; and triggering selection of a particular mode in the memory device based on the configuration signal.
8. The method of claim 7, wherein the first signal and the configuration signal are the same signal.
9. The method of claim 7, wherein the particular mode that is selected based on the configuration signal is a mode used by the PiM architecture during execution of the compute operation in the memory device.
10. The method of claim 7, wherein triggering selection of a particular mode in the memory device comprises: triggering selection of a first error-capture mode where the compute operation is paused or stopped in the memory in response to detection of an error associated with the compute operation after execution of the compute operation is triggered in the memory device.
11. The method of claim 10, wherein triggering selection of a particular mode in the memory device comprises: triggering selection of a second, different error-capture mode where execution of the compute operation in the memory continues uninterrupted in response to detection of an error associated with the compute operation after execution of the compute operation is triggered in the memory device.
12. The method of claim 11, wherein the first error-capture mode and the second, different error-capture mode each cause the memory’ device to capture error syndrome data in response to detection of an error associated with the compute operation after execution of the compute operation is triggered in the memory device.
13. The method of claim 12, wherein: the error syndrome data comprises information about an instruction and a data flow operation being executed when the error occurs in the memory after execution of the compute operation is triggered in the memory' device; and the error syndrome data is received by the SoC in response to a request that is transmitted to the memory device from the SoC.
14. The method of claim 12 or 13, wherein the memory’ device that executes the compute operation comprises: a processor-in-memory architecture having a processor unit and a mode register coupled to the processor unit.
15. The method of claim 14, wherein triggering selection of a particular mode in the memory device comprises: triggering selection of the particular mode by the mode register.
16. The method of claim 14 or 15, wherein receiving the second signal generated with reference to the memory die of the memory device comprises: receiving the second signal from the memory device using a threshold number of input / output (IO) pins coupled to the memory.
17. The method of claim 16, wherein receiving the second signal generated with reference to the memory die of the memory device comprises: receiving an indicator of success or failure, the indicator being specific to an instruction for execution by the processor unit; and receiving an identifier of the mode register coupled to the processor unit used for executing the instruction.
18. The method of any preceding claim, wherein the compute operation that is triggered for execution in the memory device is a portion of a compute for a machine-learning inference workload.
19. A system-on-chip (SoC) comprising: a processing device; and a non-transitory machine-readable storage device for storing instructions that are executable by the processing device to cause performance of operations comprising: generating a first signal at the SoC; based on the first signal, triggering execution of a compute operation in a memory device coupled to the SoC; receiving, from the memory device, a second signal generated with reference to a memory die of the memory device after triggering execution of the compute operation in the memory ; and determining a status of the compute operation based on the second signal.
20. A non-transitory machine-readable storage device storing instructions that are executable by a processing device to cause performance of operations comprising: generating a first signal at the SoC; based on the first signal, triggering execution of a compute operation in a memory device coupled to the SoC; receiving, from the memory device, a second signal generated with reference to a memory' die of the memory' device after triggering execution of the compute operation in the memory; and determining a status of the compute operation based on the second signal.
21. A computer-implemented method comprising: determining, at a system-on-chip (SoC), a memory’ access range; generating a write command based on the memory’ access range; transmitting the write command to a memory device using a control channel that couples the memory' device and the SoC; and configuring an access control at the memory device based on the memory access range specified by the write command,wherein the access control limits a processing-in-memory (PiM) compute element of the memory device to accessing only a particular portion of memory in the memory device.
22. The method of claim 21, wherein configuring an access control at the memory device comprises: configuring a row-level memory access constraint at the memory device, wherein the row-level memory access constraint limits the PiM compute element to accessing a particular range of rows at the memory device.
23. The method of claim 21 or 22, wherein configuring an access control at the memorydevice comprises: configuring, based on the write command, the access control by writing the memory access range to a configuration space of a PiM architecture in the memory device.
24. The method of any one of claims 21 to 23. wherein configuring an access control at the memory device comprises: defining a portion of memory in the memory device as the only portion of memory that is accessible to the PiM compute element of the memory- device.
25. The method of any one of claims 21 to 24. wherein the memory access range comprises a physical address range and the method comprises: writing the physical address range to a range register in a memory controller of theSoC; and translating the physical address range to row, column, bank addresses of the memory device.
26. The method of claim 25, wherein generating the write command comprises: generating the write command using the memory controller of the SoC.
27. The method of any one of claims 21 to 26, wherein determining a memory access range comprises: determining the memory access range using a trusted agent of a central processing unit (CPU) on the system-on-chip (SoC).
28. The method of any one of claims 21 to 27, wherein trusted agents of the CPU comprise: a kernel of the CPU; and a hypervisor of the CPU.
29. A system-on-chip (SoC) comprising: a processing device; and a non-transitory machine-readable storage device for storing instructions that are executable by the processing device to cause performance of operations comprising: determining, at a system-on-chip (SoC), a memory access range; generating a write command based on the memory access range; transmitting the write command to a memory device using a control channel that couples the memory device and the SoC; and configuring an access control at the memory device based on the memory access range specified by the write command, wherein the access control limits a processing-in-memory (PiM) compute element of the memory device to accessing only a particular portion of memory in the memory device.
30. A non-transitory machine-readable storage device for storing instructions that are executable by a processing device to cause performance of operations comprising: determining, at a system-on-chip (SoC), a memory access range; generating a write command based on the memoi ' access range; transmitting the write command to a memory device using a control channel that couples the memory device and the SoC; and configuring an access control at the memory device based on the memoi ' access range specified by the write command, wherein the access control limits a processing-in-memory (PiM) compute element of the memory device to accessing only a particular portion of memory in the memory device.
31. A computer-implemented method comprising: determining, at a system-on-chip (SoC), a refresh requirement for refreshing rows of a memory device coupled to the SoC;generating, at the SoC, one or more refresh commands based on the refresh requirement; transmitting the one or more refresh commands to a plurality of processing-in- memory (PiM) compute elements of the memory device coupled to the SoC; and based on the one or more refresh commands transmitted to the plurality of PiM compute elements, proactively triggering a refresh of memory cells at each of the rows of the memory device.
32. The method of claim 31, wherein proactively triggering a refresh of memory cells at each of the rows of the memory device comprises: proactively triggering the refresh of the memory cells at a frequency that exceeds a minimum refresh requirement for refreshing rows of the memory device.
33. The method of claim 32, wherein the refresh of the memory cells at each of the rows is proactively triggered at the frequency that exceeds the minimum refresh requirement using the plurality of PiM compute elements of the memory device.
34. The method of any one of claims 31 to 33, wherein transmitting the one or more refresh command signals comprises: transmitting the one or more refresh command signals over one or more clock cycles generated by a memory controller of the SoC.
35. A system-on-chip (SoC) comprising: a processing device; and a non-transitory machine-readable storage device for storing instructions that are executable by the processing device to cause performance of operations comprising: determining, at a system-on-chip (SoC), a refresh requirement for refreshing rows of a memory device coupled to the SoC; generating, at the SoC. one or more refresh commands based on the refresh requirement; transmitting the one or more refresh commands to a plurality7of processing-in- memory (PiM) compute elements of the memory' device coupled to the SoC; andbased on the one or more refresh commands transmitted to the plurality of PiM compute elements, proactively triggering a refresh of memory cells at each of the rows of the memory device.
Citation Information
Patent Citations
Memory having internal processors and methods of controlling memory access
US20110093665A1
Apparatuses and methods for in-memory operations
US20190294380A1
Detecting execution hazards in offloaded operations
US20220318085A1