Bank group data aggregation for boosting processor-in-memory data access performance

By placing PiM blocks externally to memory bank groups and using interleave configuration logic, the architecture boosts data access performance by enabling faster and more efficient execution of memory-bound workloads.

WO2025147537A1PCT designated stage expired Publication Date: 2025-07-10GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/010121
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-02
Filing Date
2025-01-02
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing memory device architectures that co-locate Processor-in-Memory (PiM) blocks with memory banks constrain data access performance due to long delay timing constraints, limiting the speed and frequency of executing read commands.

Method used

The PiM blocks are placed externally to the bank groups, allowing for a short delay timing constraint that enables faster data access speeds and frequency by leveraging a data interconnection and interleave configuration logic for concurrent routing of data across bank groups.

Benefits of technology

This configuration enhances data access performance by executing data operations at higher frequencies, optimizing memory-bound workloads and improving overall computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025010121_10072025_PF_FP_ABST
    Figure US2025010121_10072025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computational instructions / programs encoded on non-transitory computer-readable media, are disclosed for boosting PiM data access performance using cross bank group data aggregation. A system determines an interleave configuration between a PiM block and bank groups that are external to the PiM block in a memory device of an integrated circuit. The system provides a respective operand from the bank groups to the PiM block based on the interleave configuration and a timing constraint for executing successive read commands against different bank groups of the memory device that are external to the PiM block. The system triggers execution of computations for a memory-bound workload at the memory device. The computations are executed at the PiM block using the respective operands provided from the bank groups.
Need to check novelty before this filing date? Find Prior Art

Description

BANK GROUP DATA AGGREGATION FOR BOOSTING PROCESSOR-IN-MEMORY DATA ACCESS PERFORMANCEFIELD

[0001] This specification generally relates to an architecture for executing computations in a memory device coupled to a system-on-chip (“SoC”).BACKGROUND

[0002] This specification generally relates to memory devices used to execute computations.

[0003] Modem computing systems often incorporate a wide variety of computational processing units that each offer different computing capabilities and tradeoffs. Efficient execution of a given compute job often involves parsing computations into meaningful workloads or sub-tasks of a workload that are mapped to available processor cores of a computing system. The computations may be parsed and mapped based on suitability criteria, such as processor capability, performance, and power. Generally, this overall process of allocating portions of a computation to appropriate processor resources is referred to as heterogeneous compute.

[0004] At least one processing unit of the computing system can be an Intellectual Property block (“IP block") that executes a respective portion of a computational operation for different multimedia workloads. An example use case can involve processing image or speech data captured respectively by a camera or microphone on the mobile device. The SoC can use a heterogeneous compute operation to process input samples derived from the image data, the speech data, or both. An example step in the heterogeneous computing operation can include providing input samples from the image data to a host processor, such as a neural network processor or machine-learning (ML) engine, of the SoC to generate an inference output.SUMMARY

[0005] This specification describes hardware and software techniques that use bank group and cross bank group data aggregation concepts to boost or accelerate data access performance in a processor-in-memory architecture (“PiM architecture'’) used for memory - bounded and compute-bounded workloads in a memory device. The PiM architecture defines one or more PiM blocks of the memory device and each PiM block includes compute elements, such as a processor unit, mode registers, and one or more arithmetic logic units(ALUs). For example, the PiM block can include discrete processors, processor units, register devices, buffers, etc. that cooperate to form one or more PiM compute elements.

[0006] The techniques include generating data and control signaling at a system-on-chip (“SoC”) that communicates with the PiM architecture and memory device by way of a memory controller of the SoC. The data and control signaling generated at the SoC are passed to, and processed by, compute elements of the PiM blocks within a given PiM architecture. For example, the data and control signaling are routed to. and processed by, compute elements of the PiM architecture to trigger execution of certain data processing and computing operations based on bank group and cross bank group data aggregation executed in the memory device. The control signals can also be generated locally at the memory' device, which may be external to the SoC.

[0007] The PiM architecture is integrated in an example memory device, such as a dynamic random-access memory' (DRAM) or Double Data Rate (DDR) synchronous DRAM (SDRAM). The memory' device is configured for processor-in-memory (“PiM”) operations and compute-in-memory operations (“CiM operations”) by way of the PiM architecture and its multiple PiM compute elements. The memory device can be external (or internal) to the SoC. In some implementations, the SoC is integrated (or co-located) with the memory' device, for example, as distinct circuit die(s) co-located in a single integrated circuit package.

[0008] Prior design approaches for memory devices co-locate the PiM blocks with memory banks of a bank group, which limits data access performance and the speed or frequency of executing read commands. For example, co-locating PiM blocks with memory banks of a bank group constrains the data access frequency, speed, and bandwidth of the PiM block based on a long delay timing constraint (e.g., TCCD_L) for executing successive read commands against the banks of the same bank group. In contrast to these prior approaches, the techniques disclosed in this specification provide a unique hardware architecture that places one or more PiM blocks external to the bank groups. This unique placement of the PiM block relative to a bank group allows for boosting and / or accelerating data access performance by executing data access operations based on a short delay timing constraint (e g.. TCCD S).

[0009] For example, when the PiM blocks are placed external to the bank groups, the short delay timing constraint is leveraged to realize faster data access speeds and frequency for executing successive read commands. This is because the short delay constraint imposes a shorter time delay for successive read commands across bank groups relative to the long delay constraint, which imposes a longer time delay for successive read commands acrossbanks of a bank group. For example, the time delay imposed by the short delay constraint can be ! the time delay imposed by the long delay constraint.

[0010] The hardware architecture includes a data interconnection that: i) couples one or more PiM blocks and one or more bank groups of the memory device; and ii) permits concurrent routing of data from two or more memory banks of a bank group to the one or more PiM blocks. The unique hardware architecture also includes interleave configuration logic that manages the routing of the data to the PiM block based on a particular interleave configuration. The interleave configuration logic is software programmable and can be dynamically programmed to implement a range of configurations for interleaving portions of data or operands routed to the PiM blocks of the memory device.

[0011] The SoC can receive a request to execute memory-bound ML workloads and cooperates with the memory device to perform computations locally at the memory device to execute some (or all) of the memory-bound workload within the memory device. As an example, a memory-bound workload can include computational tasks where the overall execution time of the task is primarily determined by the speed of memory access, meaning that a process for workload execution can spend most of its time reading and writing data to / from memory rather than performing complex calculations. This makes the memory bandwidth and optimized data access speeds useful for enhancing workload performance.

[0012] The SoC can also generate control signals to trigger execution of computations for the memory-bound workload at a particular PiM block using data values or operands that are accessed and / or provided at a higher frequency from distinct bank groups based on the short delay timing constraint. Additionally, the SoC can generate control signals to define a certain interleave configuration such that specific data values are routed to a given PiM block of the memory device. In this manner, the memory device leverages the unique hardware architecture to boost or accelerate data access performance at the memory device, which also improves the overall timing for executing the memory -bound workload.

[0013] One aspect of the subject matter described in this specification can be embodied in a method implemented using an integrated circuit having an SoC and a memory device coupled to the SoC. The method includes determining an interleave configuration between a PiM block and bank groups that are external to the PiM block in the memory device; providing, to the PiM block, a respective operand from the bank groups based on the interleave configuration and a timing constraint for executing successive read commands against distinct bank groups of the memory device that are external to the PiM block; and triggering execution of computations for a memory-bound workload, wherein thecomputations are executed at the PiM block using the respective operands from the bank groups.

[0014] These and other implementations can each optionally include one or more of the following features. For example, in some implementations, providing the respective operands from the bank groups includes generating one or more read commands to read the respective operands from the bank groups; reading data representing the respective operands from the distinct bank groups based on the one or more read commands; and providing the respective operands to the PiM block in response to reading the data representing the respective operands.

[0015] Providing the respective operands from the bank groups includes providing, to the PiM block, a first operand from a first bank group based on the interleave configuration; and providing, to the PiM block, a second operand from a second, different bank group based on the interleave configuration and the timing constraint. The timing constraint can be a minimum number of clock cycles between executing successive read commands against the first bank group and the second, different bank group.

[0016] In some implementations, the timing constraint includes fewer clock cycles than a second, different timing constraint that is a minimum number of clock cycles between executing successive read commands against different memoiy banks in any bank group of the memory device. Providing the respective operands from the bank groups includes generating control signals that cause data to be transmitted concurrently from two or more distinct memory banks of the first bank group; and based on the control signals, concurrently transmitting the data from the two or more distinct memoi7banks of the first bank group.

[0017] In some implementations, concurrently transmitting the data from the two or more distinct memory banks of the first bank group includes: concurrently transmitting the data using a data interconnection between the first bank group and the PiM block that is external to the first bank group. In some implementations, a bandwidth of the data interconnection is sized and configured to permit a concurrent routing of a threshold number of bytes from the two or more distinct memoi7banks of the first bank group.

[0018] Another aspect of the subject matter described in this specification can be embodied in a method implemented using an integrated circuit that includes an SoC and a memoiy device coupled to the SoC. The method includes determining a mapping of resources in the memory device to one or more PiM blocks of a PiM architecture of the memory device and based on the mapping, generating a control signal that indicates an interleave configuration between a PiM block and one or more resources in the memorydevice. The method includes providing, to the PiM block, an operand from a first resource of the memory device based on the interleave configuration and a first timing constraint for executing successive read commands against distinct resources of the memory device; and triggering execution of computations for a memory -bound workload, wherein the computations are executed at the PiM block using the operand from the first resource.

[0019] These and other implementations can each optionally include one or more of the following features. For example, in some implementations, determining the mapping of resources includes determining, at the SoC, the mapping of resources using a PiM resource manager of the SoC. Generating the control signal includes generating the control signal at the SoC based on the mapping of resources determined using the PiM resource manager of the SoC. Determining the mapping of resources can include determining, at the memory device, the mapping of resources using interleave configuration logic of the memory device.

[0020] In some implementations, the interleave configuration logic is software programmable. Generating the control signal includes generating the control signal at the memory device based on the mapping of resources determined using the interleave configuration logic of the memory device. The method can include transmitting the control signal within the memory device; and configuring, at the memory device, the interleave configuration based on the control signal.

[0021] Another aspect of the subject matter described in this specification can be embodied in an integrated circuit for a memory device. The integrated circuit includes a PiM block; at least two distinct memory bank groups that are external to the PiM block, each memory bank group having two or more distinct memory banks; and a data interconnection that is configured to: i) couple the PiM block and at least a first bank group of the two distinct bank groups; and ii) permit concurrent routing of data from two or more distinct memory’ banks of the first bank group to the PiM block. The integrated circuit also includes interleave configuration logic configured to manage a routing of the data to the PiM block based on an interleave configuration.

[0022] Other implementations of this and other aspects include corresponding systems, apparatus, and computer programs, configured to perform the actions of the methods, encoded on computer storage devices. A system of one or more computers can be so configured by virtue of software, firmware, hardware, or a combination of them installed on the system that in operation causes the system to perform the actions. One or more computer programs can be so configured by virtue of having instructions that, when executed by a data processing apparatus, cause the apparatus to perform the actions.

[0023] The details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other potential features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Fig. 1 is a block diagram of an example computing system with at least one SoC.

[0025] Fig. 2 shows an example PiM architecture with corresponding compute elements.

[0026] Fig. 3 show s a first example of PiM blocks and corresponding bank groups.

[0027] Fig. 4 shows a second example of PiM blocks and corresponding bank groups.

[0028] Figs. 5A & 5B each show' an example PiM architecture for boosting PiM data access performance.

[0029] Fig. 6 is an example process for boosting PiM data access performance.

[0030] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0031] Fig. 1 is a block diagram of an example computing system 100 that includes a system-on-chip 102 (“SoC 102”). The SoC 102 includes a central processing unit 104 (“CPU 104”), a memory controller 105, a shared memory' 106 (“memory7106”), a resource manager 108, and an IP / circuit block 110. In some implementations, system 100 can include multiple SoCs and any descriptions for the SoC 102 will apply equally to each of the multiple SoCs that may be included at system 100.

[0032] The CPU 104 can be a general-purpose CPU (e.g., a single or multi-core CPU). The CPU 104 generates one or more indicators, such as an app-launch indicator or a function call that is triggered in response to executing or launching an application at a user device.For example, the application can be a camera application that uses an imaging sensor to generate image data or a gaming application that requires substantial memory and graphics processing resources to render graphical content of the game. The CPU 104 also generates one or more application values, such as pixel values or frame rate. The application values may be associated with a function call, may be descriptive of an event that occurs during execution of the application, or both.

[0033] The memory 106 is a system memory, shared memory', or both. In the example of Fig. 1, memory 106 is depicted external to circuit block 110. However, memory 106 caninclude portions of memory that are: i) specific to circuit block 110, ii) external to circuit block 110, or iii) both. The memory 106 can be random access memory of the SoC 102. such as static random-access memory (SRAM), dynamic random-access memory (DRAM), a synchronous DRAM (SDRAM), or double data rate (DDR) SDRAM.

[0034] In some implementations, aspects of memory 106 are configured as a shared scratchpad memory’ that supports parallel access of its memory resources by two or more processors of the circuit 110. The memory 106 can also include various other types of memory, such as high bandwidth memory (HBM), narrow memory (e.g., for storing 8-bit values), wide memory (e.g., for storing 16-bit or 32-bit values), etc.

[0035] The resource manager 108 is implemented in hardware and software. Aspects of the resource manager 108 can be also implemented as firmware of the SoC 102 or firmware of a device of the SoC 102, such as a DRAM memory device or the CPU 104. The resource manager 108 is a processor-in-memory (PiM) resource manager (“PiM resource manager 108”) that includes control logic implemented in hardware, software, or both. For example, the PiM resource manager 108 can include resources such as flip-flops, registers, buffers, etc. that are implemented in hardware and control logic (e.g., programmed code) that is implemented in software.

[0036] The circuit block 110 generally includes individual IP devices such as processors, processor cores, or special-purpose processing devices. For example, the circuit block 110 can include an image signal processor (ISP) 112, a tensor processing unit (TPU) 114, a digital signal processor (DSP) 1 16, and a graphics processing unit (GPU) 118. The circuit block 110 is referred to alternatively as an IP block 110, where the IP block can include one or more proprietary hardware elements. For example, each of the ISP 112, TPU 114, DSP 116, and GPU 118 can be a respective proprietary’ IP block (or IP device) of a particular entity or device manufacturer.

[0037] One or more aspects of the PiM resource manager 108 can be implemented as a software routine (or module) of the CPU 104, which uses one or more hardw are resources of the CPU 104, such as registers, buffers, etc. The CPU 104 can be configured as an instruction and vector data processing engine that processes data obtained from a system memory of the SoC 102, such as memory 106. In some implementations, each processor (e.g., ISP 112, DSP 116, TPU 114, GPU 118) of the SoC 102 includes multiple cores and the CPU 104 and / or the PiM resource manager 108 can generate control signaling 124 to manage and distribute memory intensive compute operations to a memory device 122 (e.g.. DRAM) to minimize the processing load at each core of the processors. The control signaling 124 isrouted at system 100 using an example bus 120 of the SoC 102. The control signaling 124 can include commands, requests, data, instructions, or combination of these.

[0038] The PiM resource manager 108 cooperates with the CPU 104, memory controller 105 and storage controller 107 to dynamically control and manage one or more compute-inmemory (CIM) operations. In some implementations, the CIM operations are executed at the SoC 102 in support of a heterogeneous compute operation between two or more processing units that are included among the IP block 1 10, the CPU 104. or both. More specifically, the PiM resource manager 108 is configured to generate control signaling 124 and use one or more discrete signal values of the control signaling 124 to manage and boost data access operations at the memory' device 122.

[0039] The system 100 includes an example memory device 122. The memory device 122 can include multiple memory dies. For example, the memory device 122 can include N memory dies, where N is an integer greater than 1. In the example of Fig. The memory device 122 can be a dy namic random-access memory (DRAM) or Double Data Rate (DDR) synchronous DRAM (SDRAM). The memory’ device 122 is configured to perform or support various types of PiM operations. CiM operations, and memory -near-computing operations C’MnC operations”). The memory device 122 performs or supports these operations using its multiple PiM compute elements, which are described below with reference to Figs. 2-4.

[0040] The SoC 102 cooperates with the memory device 122 to perform computations across one or more bank groups of the memory device 122. The computations can be for operations or workloads that involve one or more of the processors at IP block 1 10. Additionally, the computations can be for a heterogenous operation that spans multiple processors of IP block 110, multiple IP blocks 110, or both. In at least one example the memory device 122 may be external to the SoC 102, whereas in another example the memory device 122 may be internal to the SoC 102.

[0041] In the example of Fig. 1, system 100 and the SoC 102 is an integrated circuit of an example user / client device 130, consumer electronic device, or mobile device, where each of these devices can include items such as a smartphone 130a, tablet 130b, laptop 130c. smartwatch or wearable device 130d. The devices 130 may also include other items such as an eNotebook, Netbook, smart speaker, or mobile computer. In some implementations, the system 100 and the SoC 102 are integrated circuits of a desktop computer, network server, or related cloud-based asset.

[0042] Fig. 2 shows an example processor-in-memory (PiM) architecture 200 for boosting PiM data access performance based on control signals generated using the SoC 102,the memory device 122, or both. As noted above, the memory device 122 can include N memory dies, where A is an integer greater than 1.

[0043] In the example of Fig. 2, the memory' device 122 includes a first memory die-1, a memory die 210-1, with a first bank group that has multiple memory banks, where each memory7bank includes one or more memory arrays and a second memory7die-2, a memorydie 210-2, with a second bank group that has multiple memory banks, where each memory bank includes one or more memory arrays. In some implementations, the PiM architecture 200 includes multiple bank groups, multiple memory die, or both. For example, a single memory' die can include multiple bank groups and / or multiple bank groups can be distributed across multiple memory die.

[0044] The PiM architecture 200 includes multiple PiM blocks, where each PiM block includes multiple compute elements. For example, a first PiM block of PiM architecture 200 includes mode register 204-1 and process unit 206-1, whereas a second, different PiM block of PiM architecture 200 includes mode register 204-2 and process unit 206-2. Each process unit 206-1, 206-2 can include a processor, a processor unit, or a processor core, such as a CPU. Each process unit 206-1, 206-2 can also include an example computation unit such as an arithmetic logic unit (ALU) or multiply-accumulate cell (MAC). The PiM architecture 200, and other PiM architectures disclosed herein, can include multiple control and status register (CSR) that represent auxiliary registers that are used for reading status and changing configurations and modes of operation in a PiM architecture.

[0045] In some implementations, the PiM architecture 200 is included in the memory device 122 as multiple discrete integrated circuits, where each integrated circuit is local to a given memory die (e.g., die-1 and die-2) and interacts or communicates with arrays of memory cells at that memory7die. For example, the PiM architecture 200 can include compute elements that are replicated and distributed across each of the memory die in the memory device 122. In some other implementations, the PiM architecture 200 is included in the memory7device 122 as a single integrated circuit that interacts or communicates with each memory die of the memory7device 122, including the arrays of memory cells at each memory die. 210-1. 210-2.

[0046] To boost PiM data access performance as described in this specification, the PiM blocks or process units in the PiM architecture 200 are located within the memory' device 122 but outside of a section of the memory device 122 that includes the bank groups. The section may be defined as a discrete memory die or defined in some other way (e.g., a portion of a memory die). Irrespective of the hardware configuration or layout of PiM architecture 200,the PiM blocks are sufficiently external to the bank groups such that the PiM blocks communicate with the bank groups based on a particular timing constraint that can be leveraged to boost PiM data access performance with cross bank group data aggregation.

[0047] In some implementations, the PiM architecture 200 can be sufficiently external to, or separate from, the bank groups through a combination of physical distance, communication protocols, or data connections that enable different PiM blocks of a PiM architecture 200 to execute data access operations across different bank groups of the memory device 122. A data channel / interconnection and corresponding interleave configuration logic can be established between the individual memory arrays of a memory bank and a corresponding process unit and / or mode register of a PiM block. Additional details of this implementation are described below with reference to the examples of Figs. 5 A and 5B.

[0048] The PiM operations can include standard CPU functions, whereas the CiM operations and MnC operations can include standard arithmetic operations, such as computations normally performed by an ALU or MAC. The CiM operations and MnC operations can also include computational functions of a TPU 114, such as multiplication and addition operations for matrix math, vector computations, linear algebra, and dot-product accumulations. In some implementations, each of the PiM operations, CiM operations, and MnC operations are performed in support of machine-learning computations, neural network computations, or both.

[0049] The PiM architecture 200 can include a register or other portion of memory for storing error data for a respective memory' die or group of memoiy banks. For example, the error data can describe an error that occurred during a compute operation at a corresponding PiM block of the memory device 122. In some implementations, the register or other portion of memory is used to store configuration information, or associated instructions, for configuring aspects of a PiM block, or respective memory' die, group of memory banks, or a combination of these. For example, the mode registers 204-1, 204-2 can be used to control or trigger selection of a particular mode in a PiM architecture, such as an error-capture mode, interleave configuration mode. etc.

[0050] In some implementations, a particular mode is selected based on bit values of the mode registers 204-1, 204-2. For example, a single bit (or a sequence of bits) can be defined for use in the mode register to trigger or select one or more interleave configuration modes. The bit values of a mode register 204-1. 204-2 are set by commands / instructions received from the SoC 102, local commands / instructions executed by a processing unit of the PiMarchitecture 200, or both. For example, the PiM mode registers values can be programed using a mode register write command (MRW), whereas other PiM registers, such as the PiM CSRs, can be programmed using a PiM control register write command. A subset of bits in an example control signal can be used to trigger certain interleave configuration modes and bank group (or cross bank group) data aggregation modes / settings of the PiM architecture 200.

[0051] The control signal can be a dedicated PiM config control signal generated at the SoC 102. Additionally, or alternatively, the control signal can be a read request signal generated by the memory' controller 105 of the SoC 102. For example, in addition to its primary purpose of requesting, reading, or obtaining data from the memory device 122, a read request signal (and its associated data structure of bits) can be re-purposed to provide a secondary purpose of triggering execution of a compute operation in a PiM module of the memory device 122. Based on the control signaling, a particular interleave configuration mode or data aggregation mode is triggered and / or configured in the PiM architecture 200. In some cases, this triggering or configuring of certain modes in the PiM architecture 200 occurs when a command from the SoC 102 triggers a batch execution of multiple PiM instructions in the memory device 122.

[0052] A frequency of the memory' device 122 can be measured in megahertz (MHz) and indicates a number of cycles the memory device 122 (e.g., DRAM or SDRAM) can perform in one second, representing the number of data transfers that can occur between, for example, a PiM block and bank group (or memory banks) of the memory device 122 in a given time frame. Generally, higher DRAM frequencies result in faster data access and better system performance.

[0053] For clarity, the memory’ device 122 can have an example structure multiple memory banks, where each memory bank includes one or more memory arrays, and each memory array includes a set of memory cells that arranged, for example, in a row, column format. In general, a memory’ cell can be the smallest unit of memory', storing, for example, a single bit of data, whereas a memory’ array can be a collection of multiple memory cells that are organized together to store larger amounts of data. In some implementations, the memory array can be defined as a two-dimensional array' of memory' cells where data is stored and later retrieved or read and routed to a PiM block 212-1, 212-2, for computational processing at the PiM block using arithmetic circuitry of the PiM block 212-1, 212-2.

[0054] The memory banks are arranged into different groupings of memory banks, which are referred to herein as bank groups. In some examples, the example structure of thememory cells, memory arrays, memory banks, bank groups, and a memory comprising multiple bank groups, is implemented across one or more circuit dies. Each bank group can represent a separate memory entity / unit that is configured such that a row or column cycle memory access operation can be completed within that bank group without impacting operations occurring in another bank group.

[0055] These cells, arrays, banks, and groups of the memory device 122 represent resources (or memory resources) of the memory device 122. The control logic of the system 100 is configured to determine a mapping of resources in the memory device 122 to one or more PiM blocks of a PiM architecture of the memory device. Based on this mapping, configuration logic of the system 100 can generate control signals that indicate an interleave configuration between a PiM block and one or more resources in the memory’ device. This is described in more detail below.

[0056] Using the disclosed techniques, data access operations are optimized by reading memory' cells of bank groups at a frequency that exceeds the read command frequency afforded by a long delay timing constraint, such as the TCCD_L timing constraint. As noted above, the innovative techniques disclosed in this specification leverage unique placement of the PiM blocks relative to a bank group entity to boost and / or accelerate data access performance by executing data access operations based on a short delay timing constraint, such as the TCCD_S. As used in this document, the CCD can refer to column-to-column delay, or command-to-command delay on the column side. The “_S” refers to “short”. whereas the “_L” refers to “long.”

[0057] An internal controller of the PiM block (1), 212-1, or PiM block (2), (212-2), can execute successive read commands at a frequency that is based on a clock cycle generated by the memory device 122. The internal controller can operate based on a particular clock frequency, e.g., a 200 or 800 MHz clock or 1000 MHz clock. Other clocks frequency are also within the scope of this disclosure. In some implementations, the bank group defines tCCD Parameters differently between the same bank group and a different bank group.

[0058] An internal controller of the PiM architecture can leverage the TCCD_S timing constraint to execute cross bank group accesses at, for example, nanosecond (ns) intervals, rather than the slower intervals associated with the TCCD_L timing constraint. For example, the time delay imposed by a short delay timing constraint (TCCD_S) can be ! the time delay imposed by the long delay constraint (TCCD L). In view- of this, the disclosed techniques for cross bank group data access leverages the TCCD S timing constraint to implement higher frequency data access operations (relative to TCCD_L) by reducing the access timingby ! based on the relationship: (TCCD_S = 'A * TCCD_L). The data access operation can be represented by a read command signal triggered by instructions processed at the PiM architecture of the memory device 122.

[0059] In some implementations, the memory device 122 includes a data rate of 6400 Megabits per second (Mbps), representing an example amount of data the memory device 122 can transfer each second. In this example, a delay imposed by the TCCD_S and TCCD_L timing constraints can be defined using clock cycles (tex). For example, the TCCD S constraint can impose a four (4 tex) delay, whereas the TCCD_L imposes a constraint that is greater the TCCD S. For example, the TCCD L constraint can be an eight (8 tcK.) cycle delay period. In this example, if the PiM block operates based on a 200Mhz clock, then data rate: 6400Mbps WCK: * / 2Data rate (6400Mhz) CK:lA WCK (3200Mhz) tex: 1.25ns 4 tCK (4 * 1.25ns) period, and 5ns 200Mhz. For example, when staying within the same bank group at 1,600 Mbps, the tCCD_L constraint requires more than four clocks. In these examples, CK stands for clock and WCK stands for write clock, where each corresponds to a clock frequency of the memory device 122.

[0060] Fig. 3 shows a first example PiM architecture 300. In the example of Fig. 3. the architecture 300 includes bank groups 302 and 304, each including at least two memory banks (310-1 and 310-2, and 312-1 and 312-2, respectively). The architecture 300 also includes at least two PiM blocks 306-1 and 306-2. The PiM architecture 300 represents a design option where the PiM blocks 306-1, 306-2 are constrained to operating within, and executing read / write operations against, a specific bank group. For example, as shown PiM block 306-1 is located within bank group 302 and is constrained to executing its data access operations within bank group 302, whereas PiM blocks 306-2 is located within bank group 304 and is constrained to executing its data access operations within bank group 304.

[0061] In PiM architecture 300, data is read by the PiM blocks from each of the memory banks one array / bank at a time. For example, during an operation, PiM block 212-1 can be directed to read data from each of memory banks 310-1 and 310-2 within bank group 302. In this example operation, the PiM block 306-1 performs its read operation based on a timing constraint and corresponding frequency. For example, the timing constraint can be based on the TCCD L, as defined by the Joint Electron Device Engineering Council (JEDEC) specification for executing successive read operations against the memory banks of a bank group. Operating at TCCD L, the PiM block first reads 32-bytes (32B) of data from memory bank 310- 1 , and once complete, then reads out 32B data from memory bank 310-2 . In thisarchitecture, similar operations occur in bank group 304 using PiM block 306-2 and memorybanks 312-1 and 312-2.

[0062] Fig. 4 shows a second example PiM architecture 400. In the example of Fig. 4, the architecture 400 includes bank groups 402 and 404, which each include two memory banks (310-1 and 310-2, and 312-1 and 312-2, respectively). In contrast with Fig. 3, the architecture 400 includes four PiM blocks - 206-1, 206-2, 206-3, and 206-4 - two of which operate within bank group 402 and two of which operate within bank group 404.

[0063] The operations of architecture 400 are similar to architecture 300 in Fig. 3, in that 32B of data is read from a memory' bank by a PiM block based on the TCCD L timing constraint and associated operating frequency. In contrast with Fig. 3, because architecture 400 includes two PiM blocks per memory bank group, data can be read from multiple memory banks simultaneously. For example, in bank group 402, PiM block 212-1 can read data from memory' bank 310-1 at the same time PiM block 212-2 reads data from memory bank 310-2. As a result, assuming the same operation in Figs. 3 and 4, the operation can execute faster in architecture 400 because multiple memory banks can be read simultaneously by a greater number of PiM blocks. However, the additional PiM blocks 206-2 and 206-4 of architecture 400 result in a larger circuit footprint when architecture 400 is compared to architecture 300.

[0064] The PiM architecture 400 can be an extension of PiM architecture 300, but a more limited design relative to the PiM architecture 200. In the example of Fig. 4. the SoC 102 communicates with compute elements of the PiM architecture 400 to configure data access at the memory' device 122. For example, compute elements such as a mode register 204-1, 204- 2 and a respective process unit 206-1, 206-2 are used to trigger or define a data access configuration for executing read commands against memory cells along rows of each DRAM bank of memory die- 1 or die-2.

[0065] Figs. 5A and 5B each show an example PiM architecture for boosting PiM data access performance. Fig. 5A shows an example of the PiM architecture 500 performing an operation where each PiM block reads data from multiple memory arrays / banks within the same bank group, whereas Fig. 5B shows a ■■cross-bank” configuration in which the PiM blocks 212-1, 212-2 read data from memory' banks 310-1, 312-1 in different bank groups 502, 504, respectively.

[0066] In the examples of Figs. 5A and 5B, the architecture 500 includes two PiM blocks 212-1 and 212-2. two bank groups 502 and 504 each having two memory banks (310-1, 310- 2, 312-1, and 312-2), interleave configuration logic 508, and a data channel 510. Notably, inarchitecture 500 the PiM blocks 212-1 and 212-2 are located outside of the bank groups 502 and 504. While only certain numbers of PiM blocks, bank groups, and memory banks are depicted, architecture 500 is scalable and can be adapted to operate with different arrangements and numbers of these components.

[0067] As discussed above in Figs. 3 and 4, bank groups 502 and 504 operate at a set frequency corresponding, for example, to a long delay timing constraint such as the TCCD_L timing constraint of the JEDEC specification. By locating or placing the PiM blocks 212-1 and 212-2 outside of the bank groups 502 and 504, the PiM blocks can be operated at a higher frequency (e.g., TCCD S of the JEDEC specification - double the frequency of TCCD_L). This higher operating frequency allows PiM blocks 212-1 and 212-2 to read data from memory arrays of the memory banks 310, 312 at a faster rate, for example, twice the normal rate as when PiM blocks are located within the bank groups.

[0068] Fig. 5A show-s an example operation where PiM blocks 212-1 and 212-2 read data from multiple memory banks within the same bank group 502 and 504. As shown in Fig. 5A, PiM block 212-1 reads data from memory banks 310-1 and 310-2 from bank group 504. Similar operations occur for PiM block 212-2. which reads data from memory banks 312-1 and 312-2 from bank group 502.

[0069] Given that PiM blocks 212-1 and 212-2 are reading data at twice the rate when compared to Figs. 3 and 4, the data channel 504A can be adjusted to account for the higher data rate. For example, as shown in Fig. 5 A, the data channel 510 is sized to be 32 bytes (32B) for each memory bank in the bank group (64 bytes (64B) total). The size of the data channel 510 can be scaled appropriately depending on the number of memory banks in a particular bank group that will be read simultaneously.

[0070] Additionally, architecture 500 includes interleave configuration logic 508 that can control the operation of PiM blocks 212-1 and 212-2 to account for data dependency either between memory banks or across bank groups. For example, should data within memory bank 310-1 depend on data from memory7bank 310-2, interleave configuration logic 508 can cause the PiM block 212-1 to read the data from memory bank 310-2 first. Similarly, if data from bank group 502 is required to read data from bank group 504. interleave configuration logic 508 can be configured such that PiM blocks 212-1 and 212-2 read data stored at bank group 502 first. The operations of interleave configuration logic 508 can be dynamic and adjust to data dependency requirements as operations are assigned to PiM blocks 212-1 and 212-2.

[0071] Fig. 5B shows an example operation where PiM blocks 212-1 and 212-2 read data from multiple memory banks within the different bank groups 502 and 504 (a "cross-bank" configuration). In this example operation, PiM block 212-1 reads data from memory bank 310-1 of bank group 502 and memoiy bank 312-1 of bank group 504. Similar operations occur with PiM 212-2. As described above in Fig. 5A, the data channel 504B is sized to account for the number of memory’ banks that will be read simultaneously (e.g., 32 bytes per memory bank). Additionally, as discussed with reference to Fig. 5A. the interleave configuration logic 508 resolves any data dependency’ issues that arise during operation of the PiMs.

[0072] Because the PiM blocks 212-1 and 212-2 shown in Figs. 5A and 5B are located outside of the bank groups 502 and 504. operating frequency limitations for data access within the same bank group do not apply (e.g., TCCD L of the JEDEC specification). Accordingly, faster operating frequencies can be used (e.g., TCCD_S), which provides increased processing performance for the same physical footprint. For example, assuming an identical operation, and that TCCD S is tyvice the speed of TCCD L, the operation shown in Fig. 4 can be completed in the same amount of time using architecture 500 of Figs. 5A or 5B. Notably, architecture 500 requires two fewer PiM blocks than architecture 400, and therefore can achieve the same performance with a smaller physical footprint.

[0073] Fig. 6 is an example process for boosting PiM data access performance in a memory device 122. Process 600 is implemented or executed at system 100 using at least the SoC 102 and memory device 122 described above. Hence, descriptions of process 600 will reference the above-mentioned computing resources of system 100. In some examples, the steps or actions of process 600 are enabled by programmed software instructions, firmware instructions, or both. Each type of instruction may be stored in a non-transitory machine- readable storage device and is executable by one or more of the processors or other resources described in this specification.

[0074] Referring again to process 600, the system 100 determines an interleave configuration between a processor-in-memory ("PiM") block and at least tyvo distinct bank groups that are external to the PiM block in the memory device (602). For example, the CPU 104 or PiM resource manager 108 can dynamically determine the interleave configuration based on the ty pe of PiM operation or the ty pe of CiM operation being performed at the memory device 122. The CPU 104 or PiM resource manager 108 can dynamically determine the interleave configuration based on the type of memory device 122 coupled to the SoC 102, the quantity of memory banks in each bank group of the memory device 122, a layout of datavalues or operands across the bank groups of the memory device 122, or a combination of these. In some implementations, the memory device 122 includes interleave configuration logic that is software programmable. In these implementations, an interleave configuration can be dynamically configured and / or reconfigured at the memory device 122 using the software programmable aspects of interleave configuration logic.

[0075] The SoC 102 can determine a mapping of bank groups to PiM blocks of the memory device 122. The mapping can be used to define interleave configurations and data access requirements for rows or columns of banks in a particular bank group of a memory device coupled to the SoC. In some implementations, the SoC 102 defines an interleave configuration at the memory device 122 by determining a mapping between a PiM block 212- 1, 212-2 and a row or column in a memory bank of a bank group 502, 504 on a particular memory die of the memory device 122.

[0076] The row- or column-level mapping can be indicated or referenced by an interleave configuration that causes data values stored in a row or column of a bank to be interleaved among two or more PiM blocks when the data values are transmitted to the PiM blocks via data interconnection 510. In some implementations, a bandwidth of the data interconnection 510 is sized and configured to permit a concurrent routing of a threshold number of bytes from the two or more distinct memory' banks of the first bank group. For example, threshold number of bytes can be 16B. 32B, 64B, or 128B, where “B"’ is bytes. In some implementations, the threshold can be more than or fewer than 32B.

[0077] The system 1 0 provides a respective operand or data value from the bank groups to the PiM block based on the interleave configuration and a timing constraint for executing successive read commands against distinct bank groups of the memory device that are external to the PiM block (604). For example, the memory device 122 can provide data values from the two different bank groups, transmit the data values via the data interconnection 510, and interleave respective subsets of data values passed to PiM blocks 212-1, 212-2 based on the interleave configuration. In some implementations, the SoC 102 uses the CPU 104 and / or the PiM resource manager 108 to generate and pass control signals to the memory device 122 that define an interleave configuration and trigger data access operations that cause distinct bank groups of the memory device 122 to provide data values to one or more PiM blocks 212-1, 212-2 based on the interleave configuration logic 508.

[0078] The system 100 triggers execution of computations for a memory-bound workload (606). As noted above, a memory-bound workload can include computational tasks where the overall execution time of the task is primarily determined by the speed of memory access.In this regard, the processes workload execution may be required to spend substantial time reading and writing data from memory’ banks of the memory device, rather than performing complex calculations. This timing overhead causes the memory bandwidth and memory access timing a non-trivial (e.g., important) factor in overall performance of the workload.

[0079] The computations are executed at the PiM block using the respective data values or operands provided to the PiM block from the one or more bank groups. For example, the SoC 102 can trigger execution of the computations by transmitting control signals (or commands) to the memory’ device 122 using a control channel that couples the memory device 122 and the SoC. In some implementations, the SoC 102 triggers execution of a compute operation in the memory device 122 based on a request / command signal generated by the memory controller 105 of the SoC 102.

[0080] The request / command signal is passed to a compute element of the PiM block 212-1, 212-2 in the memory device 122 to trigger execution of instructions in an instruction buffer of a particular PiM block 212-1, 212-2. For example, the request signal can be passed to a processor of PiM block 212-1, a mode register of PiM block 212-2, or both. The processor (or mode register) uses or reads the request signal to execute compute instructions and generate a result of computations performed based on the instructions. The results can be passed back to the bank groups for storing in corresponding memory banks of the bank groups.

[0081] The respective steps of process 600 can be performed at a hardware integrated circuit as part of a larger compute operation to generate a machine-learning (ML) output, including an output for a neural network layer of a neural network that implements one or more ML models. For example, the output can be a portion of a computation for a ML task or inference workload to generate an image processing, speech processing, or image recognition output.

[0082] As indicated above, a portion of the integrated circuit can include a specialpurpose neural network processor or hardware ML accelerator configured to accelerate computations for generating different types of data processing outputs. In some implementations, one or more of the PiM operations, CiM operations, or MnC operations are performed by the memory device 122 to support or enable accelerating computations for generating different types of data processing outputs.

[0083] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed inthis specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory program carrier for execution by, or to control the operation of, data processing apparatus.

[0084] Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0085] The term '‘computing system” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0086] A computer program (which may also be referred to or described as a program, software, a software application, a module, a software module, a script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0087] A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication netw ork.

[0088] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to performfunctions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array), an ASIC (application specific integrated circuit), or a GPGPU (General purpose graphics processing unit).

[0089] Computers suitable for the execution of a computer program include, by way of example, can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random-access memory or both. Some elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.

[0090] Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory7, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. The processor and the memory' can be supplemented by, or incorporated in, special purpose logic circuitry'.

[0091] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by' which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user's client device in response to requests received from the web browser.

[0092] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), e.g., the Internet.

[0093] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0094] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0095] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0096] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.

Claims

What is claimed is:

1. A method implemented using an integrated circuit comprising a System-on-Chip (“SoC") and a memory device coupled to the SoC, the method comprising: determining an interleave configuration between a processor-in-memory (“PiM”) block and bank groups that are external to the PiM block in the memory device; providing, to the PiM block, a respective operand from the each of a plurality of bank groups based on the interleave configuration and a timing constraint for executing successive read commands against distinct bank groups of the memory device that are external to the PiM block; and triggering execution of computations for a memory -bound workload, wherein the computations are executed at the PiM block using the respective operands from the bank groups.

2. The method of claim 1, wherein providing the respective operands from the bank groups comprises: generating one or more read commands to read the respective operands from the bank groups: reading data representing the respective operands from the distinct bank groups based on the one or more read commands; and providing the respective operands to the PiM block in response to reading the data representing the respective operands.

3. The method of claim 2, wherein providing the respective operands from the bank groups comprises: providing, to the PiM block, a first operand from a first bank group based on the interleave configuration; and providing, to the PiM block, a second operand from a second, different bank group based on the interleave configuration and the timing constraint.

4. The method of claim 3, wherein the timing constraint is a minimum number of clock cycles between executing successive read commands against the first bank group and the second, different bank group.

5. The method of claim 3 or 4, wherein the timing constraint comprises fewer clock cycles than a second, different timing constraint that is a minimum number of clock cycles between executing successive read commands against different memory banks in any bank group of the memory device.

6. The method of claim 5, wherein providing the respective operands from the bank groups comprises: generating control signals that cause data to be transmitted concurrently from two or more distinct memory banks of the first bank group; and based on the control signals, concurrently transmitting the data from the two or more distinct memory banks of the first bank group.

7. The method of claim 6, wherein concurrently transmitting the data from the two or more distinct memory banks of the first bank group comprises: concurrently transmitting the data using a data interconnection between the first bank group and the PiM block that is external to the first bank group.

8. The method of claim 7, wherein a bandw idth of the data interconnection is sized and configured to permit a concurrent routing of a threshold number of bytes from the two or more distinct memory banks of the first bank group.

9. A method implemented using an integrated circuit comprising a System-on-Chip (“SoC”) and a memory device, the method comprising: determining a mapping of resources in the memory device to one or more processorin-memory (“PiM”) blocks of the memory device; based on the mapping, generating a control signal that indicates an interleave configuration between a PiM block and one or more resources in the memory device; providing, to the PiM block, an operand from a first resource of the memory device based on the interleave configuration and a first timing constraint for executing successive read commands against distinct resources of the memory' device; and triggering execution of computations for a memory -bound workload, wherein the computations are executed at the PiM block using the operand from the first resource.

10. The method of claim 9, wherein determining the mapping of resources comprises: determining, at the SoC, the mapping of resources using a PiM resource manager of the SoC.

11. The method of claim 10, wherein generating the control signal comprises: generating the control signal at the SoC based on the mapping of resources determined using the PiM resource manager of the SoC.

12. The method of claim 9, wherein determining the mapping of resources comprises: determining, at the memory' device, the mapping of resources using interleave configuration logic of the memory device.

13. The method of claim 12, wherein the interleave configuration logic is software programmable.

14. The method of claim 12 or 13. wherein generating the control signal comprises: generating the control signal at the memory device based on the mapping of resources determined using the interleave configuration logic of the memory' device.

15. The method of any one of claims 9 to 14, further comprising: transmitting the control signal within the memory device; and configuring, at the memory' device, the interleave configuration based on the control signal.

16. An integrated circuit for a memory device comprising: a processor-in-memory f‘PiM’?) block; at least two distinct bank groups that are external to the PiM block, each bank group comprising two or more distinct memory banks; a data interconnection that is configured to: i) couple the PiM block and at least a first bank group of the two distinct bank groups; and ii) permit concurrent routing of data from two or more distinct memory banks of the first bank group to the PiM block; andinterleave configuration logic configured to manage a routing of the data to the PiM block based on an interleave configuration.