Heterogeneous execution method and system for master control and near memory accelerator in near memory computing architecture

By adopting explicit synchronization mechanism, offload-scheduling-return mechanism and time-series optimization strategies in the near-memory computing architecture, the access conflicts between the master control and the near-memory accelerator and the low local utilization rate of the row are solved, and the efficient coordination and multi-level collaborative computing scenarios are achieved.

CN120104528AActive Publication Date: 2025-06-06SHANGHAI JIAOTONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510222347.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-06
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In the near-memory computing architecture, access conflicts between the master and the near-memory accelerator, low local utilization of rows and limited scalability lead to increased memory access latency and power consumption, making it difficult to support multi-level collaborative computing scenarios.

Method used

Through the explicit synchronization mechanism, the unload-scheduling-return mechanism and the timely optimization strategy, efficient coordination between the master control and the near-memory accelerator is achieved. The specific method includes real-time update of the memory status of the near-memory accelerator memory controller in host mode, offloading the memory access request to the near-memory accelerator memory controller through an offload mechanism, and returning data through time division multiplexing interrupts and adaptive batch processing mechanisms.

Benefits of technology

Eliminate asynchronous conflicts, maximize the local utilization of rows, improve memory bandwidth utilization, support multi-level collaborative computing scenarios, and reduce transformation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104528A_ABST
    Figure CN120104528A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of communication between hosts and near-memory accelerators, and discloses a heterogeneous execution method and system for a master control and a near-memory accelerator in a near-memory computing architecture, which comprises three mechanisms of dominant synchronization, unloading-scheduling-return and implicit synchronization, and comprises the following steps: in a host mode, a Host MC directly sends a DDR command to a memory device, and bypasses the command to an NMA MC at the same time; updating the memory state recorded in the NMA MC; in a simultaneous memory access mode and a near memory mode, a Host MC unloads a memory access request to an NMA MC, the NMA MC schedules an access request unloaded by a host and an NMA local access request according to a switching-recovery method, correctness and line locality optimization of a memory state are ensured, and the NMA MC returns data to the Host MC through time division multiplexing interruption and a self-adaptive batch processing mechanism; and when the simultaneous memory access mode and the near memory mode are ended, the NMA MC actively adjusts the memory state according to the recorded memory bank Bank state change instruction. The correctness, the consistency and the high efficiency are realized when the Host MC and the NMA MC are accessed and stored at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of communication between a host and a near memory accelerator, and in particular to a method and system for heterogeneous execution of a master control and a near memory accelerator in a near memory computing architecture. Background Art

[0002] In recent years, with the rapid development of artificial intelligence, big data and high-performance computing, traditional computing architecture faces severe challenges of memory wall. In the traditional storage architecture based on dual inline memory module (DIMM), as shown in the figure Figure 1 As shown, the host manages multiple independent channels (Channel) through the memory controller (Memory Controller, MC), and each channel is connected to several DIMM modules. The DIMM is composed of multiple memory ranks (Ranks), and each Rank is further divided into a memory bank group (Bank Group, BG) and a memory bank (Bank). The master controller needs to send activation (ACT), read / write (RD / WR) and precharge (PRE) instructions through complex timing constraints (such as row activation delay tRCD, precharge time tRP, etc.) to access the target data. In order to break through this limitation, the Near Memory Processing (NMP) architecture came into being, as shown in the attached figure Figure 2 This architecture embeds computing units (such as Near Memory Accelerator, NMA) near the memory and processes data directly in the memory, significantly reducing the data transmission distance and power consumption, and improving the overall system performance.

[0003] However, the inherent defects of the traditional architecture are reflected in the following aspects: (1) Asynchronous execution conflicts. When the host controller and NMA initiate memory access at the same time, the two memory controllers (Host MC and NMA MC) lack a synchronization mechanism, which may cause memory bus conflicts and thus cause DRAM errors; (2) Low row locality utilization. The traditional MC scheduling strategy does not fully optimize the row hit rate of adjacent accesses, resulting in frequent row switching operations, increased latency and power consumption; (3) Limited scalability. The existing DIMM-NMP architecture usually adopts a single-level (such as Rank-level) NMA integration method, which is difficult to support multi-level collaborative computing scenarios.

[0004] On this basis, in order to alleviate the access conflict between the master control and NMA, the industry has proposed two mainstream solutions, as shown in the attached figure Figure 3 However, they all have significant limitations: (1) Shielding Access: This solution isolates the access paths of the master and NMA through a hardware multiplexer (MUX) to ensure that both have exclusive use of the DIMM. For example, the master configures the MUX state through the software development kit (SDK) before the NMA executes, and switches back to the master mode after the NMA completes the calculation. However, its core defects include: loss of storage capacity and bandwidth, requiring some DIMMs to be reserved for NMA exclusive use, resulting in a decrease in the overall storage capacity and available bandwidth of the system; limited flexibility, strict requirements on data layout, memory address mapping and dynamic allocation strategies, and difficulty in adapting to dynamic load scenarios.

[0005] (2) Opportunistic Access: This scheme allows the NMA MC to track the master access status in real time and initiate local memory operations whenever possible. It relies on the NMA workload to have a predefined deterministic pattern and avoid master access starvation through probabilistic arbitration. However, its key problems are: it has a narrow scope of application and only supports specific loads with fixed computing patterns and predictable delays (such as regular matrix operations), and cannot adapt to heterogeneous computing needs; it lacks row locality, and opportunistic access makes it difficult to take advantage of the locality of consecutive row addresses, and frequent row switching leads to performance degradation; it has high hardware complexity and requires high-precision state tracking logic to be implemented in the NMA MC, increasing design costs and power consumption.

[0006] Therefore, how to achieve efficient collaboration between the host and the near memory accelerator (NMA) in the near memory computing architecture, especially the synchronization and consistency management of memory access, has become a core technical problem that needs to be solved urgently. Summary of the invention

[0007] The purpose of the present invention is to solve the shortcomings of the above-mentioned prior art and to provide a heterogeneous execution method and system for the master control and near memory accelerator in a near memory computing architecture, which realizes efficient collaboration between the master control and NMA through explicit synchronization mechanism, offload-scheduling-return mechanism and timing optimization strategy.

[0008] On the one hand, a method for heterogeneous execution of a master controller and a near-memory accelerator in a near-memory computing architecture is provided, comprising the following steps: S1: In the host mode, the host memory controller directly sends a DDR command to the memory device, and bypasses the command to the near memory accelerator memory controller, and updates the memory status recorded in the near memory accelerator memory controller in real time; S2: In the simultaneous memory access mode and the near memory mode, the host memory controller offloads the memory access request to the near memory accelerator memory controller through the offload mechanism. The near memory accelerator memory controller schedules the host offloaded access request and the near memory accelerator local access request according to the preset switching-recovery method to ensure the correctness of the memory state and the optimization of row locality. The near memory accelerator memory controller returns the data to the host memory controller through time-division multiplexing interrupts and adaptive batch processing mechanisms. S3: When the simultaneous memory access mode and the near memory mode end, the near memory accelerator memory controller actively adjusts the memory state according to the recorded memory bank Bank state change instruction to ensure consistency with the host memory controller state for switching the host mode.

[0009] Further, in step S1, the real-time updating of the memory status recorded in the near memory accelerator memory controller includes: The DDR command is transmitted via the command / address bus, and the near memory accelerator memory controller processes the received host command as a command issued by itself, and updates the memory status recorded in the near memory accelerator memory controller, so that it can always track the latest memory status.

[0010] Further, in step S2, the host memory controller offloading the memory access request to the near memory accelerator memory controller includes: New memory commands are introduced to optimize the timing of offload operations. The host memory controller sends offload access requests to the near memory accelerator memory controller through the command / address bus, and various delays are omitted through timing optimization.

[0011] Preferably, the timing optimization further comprises: The row activation delay tRCD and the timing constraint tRTRs are omitted in the unloaded read command RDO. Meanwhile, the physical execution of the row activation instruction ACT is omitted for the unloaded read command RDO, and only its logic state is recorded. A read operation is inserted in the offloaded write command WRO to omit the write-to-read delay tWTR, the read-to-write delay tRTW, and the write delay tWL to maximize bus utilization.

[0012] Further, in step S2, the near memory accelerator memory controller schedules the host offload access request and the near memory accelerator local access request according to the preset switching-restoring method, including: In the near memory accelerator memory controller, an independent command queue CMD FIFO and request queue REQ FIFO are set for each memory bank Bank, and the host command and the near memory accelerator request are isolated in the command queue CMD FIFO and the request queue REQ FIFO respectively, and the queue is switched only when one of the following conditions is met, including: There are no remaining commands or requests in the current queue; Issue a pre-charge instruction PRE command; By recording the last BSC command from the command queue CMD FIFO and the request queue REQ FIFO, it is determined whether recovery is required, and the corresponding BSC command is issued when the switch is restored to ensure the consistency of the memory state after the switch.

[0013] Further, in step S2, the time division multiplexing interruption includes: An interrupt-based signal is used to allow the near memory accelerator memory controller to notify the host memory controller, and the effective period of each memory rank sending an interrupt signal is limited by time-division multiplexing the interrupt signal to reduce polling delay.

[0014] Preferably, in step S2, the adaptive batch processing mechanism further comprises: Limit interrupts to be issued in the order in which memory access is offloaded, and when the host memory controller requires the near memory accelerator memory controller to return data, the data is collectively returned in batches, and the batch size is equal to the number of interrupts for offloaded reading, wherein, Disabling subsequent interrupts from triggering when earlier interrupts have not completed; Multiple read access interrupts are merged into batches, and data is returned in batches through a single return instruction RT command. The delay of the RT command is dynamically adjusted by the batch size and is executed in priority to other memory access commands.

[0015] On the other hand, a heterogeneous execution system of a master control and a near memory accelerator in a near memory computing architecture for implementing the above method is provided, comprising a host memory controller and a near memory accelerator memory controller, wherein the host memory controller comprises an unload table, a return unit and a command encoding module, and the near memory accelerator memory controller comprises a bypass path, a configurable FIFO queue and a switch-recovery unit, wherein: The unloading table is used to record the number of unloaded commands for each memory bank Bank, and suspend the generation of new commands when the preset threshold is reached; The return unit is used to manage the timing and priority of the returned data and is configured to dynamically adjust the burst length according to the interrupt count; The command encoding module encodes the return instruction RT command by using the reserved field, and distinguishes the uninstallation command by multiplexing the DDR command signal; The bypass path is used to directly transmit the host memory controller's command to the memory device in the host mode and update the local memory status at the same time; The configurable FIFO queue is used to dynamically allocate FIFO queue capacity in concurrent mode, respectively store host offload access requests and near memory accelerator local access requests, and implement queue switching through an arbitrator; The switch-recovery unit is used to record BSC commands and generate recovery instructions to ensure the correctness of memory state and mode switching.

[0016] Furthermore, the return unit further includes a physical address mapping module and a priority controller, wherein: The physical address mapping module is used to associate interrupt signals with corresponding memory accesses; The priority controller is used to set the issuing priority of the return instruction RT command to be higher than other DDR commands; The capacity allocation of the configurable FIFO queue is achieved in the following way: In the near memory mode, the near memory accelerator memory controller uses the entire request queue REQ FIFO to store the near memory accelerator local access requests, and in the concurrent mode, uses 50% of the capacity of the FIFO queue to store the host offload commands; The alternating scheduling of host offload commands and local access requests is achieved through an arbiter in each hybrid FIFA unit.

[0017] In addition, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method for heterogeneous execution of a master control and a near memory accelerator in a near memory computing architecture described in any one of the above is implemented.

[0018] At the same time, an electronic device is provided, including: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the heterogeneous execution method of the master control and the near memory accelerator in the near memory computing architecture described in any of the above items.

[0019] Compared with the prior art, the present invention has the following beneficial effects: The present invention directly transmits the host memory controller's command to the memory device in the host mode through a bypass path, updates the local memory status at the same time, synchronizes the memory status of the host memory controller and the memory controller of the near memory accelerator in real time, and eliminates asynchronous conflicts; The present invention optimizes queue scheduling based on a switch-recovery method. In a near-memory accelerator memory controller, an independent command queue CMD FIFO and a request queue REQ FIFO are set for each memory bank Bank, and host commands and near-memory accelerator requests are isolated in the command queue CMD FIFO and the request queue REQ FIFO, respectively, to maximize row locality. The present invention optimizes the timing, schedules more offload commands, omits redundant delay constraints in the offload commands, and significantly improves bandwidth utilization; The present invention introduces a new memory command in the return mechanism for batch data return, and uses a time division multiplexing method for interrupt notification, reuses the existing DDR command code and bus interface, and reduces the transformation cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 A schematic diagram of a storage architecture based on a dual in-line memory module DIMM according to the present invention; Figure 2 A schematic diagram of a near-storage computing architecture of the present invention; Figure 3 A schematic diagram comparing the shielded access scheme and the opportunistic access scheme of the prior art; Figure 4 It is a flow chart of a method for heterogeneous execution of a master control and a near memory accelerator in a near memory computing architecture of the present invention; Figure 5 A schematic diagram of the mechanism structure overview of a method for heterogeneous execution of a master controller and a near-memory accelerator in a near-memory computing architecture of the present invention; Figure 6 This is an example diagram of timing optimization in an unloading mechanism of the present invention; Figure 7 This is a switching-recovery example diagram in a memory bank Bank scheduling mechanism of the present invention; Figure 8 This is an example diagram of a timing design of a return mechanism of the present invention; Fig. 9 The present invention is a block diagram of the heterogeneous execution structure of the master control and the near memory accelerator in the near memory computing architecture. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0022] Near Memory Processing (NMP) architecture is a new computing architecture that integrates near memory accelerators (NMA) near the storage unit of the memory. In traditional computing architecture, memory and processor are usually separate entities, and data transmission needs to be carried out through buses and other means; while NMP can process data in memory, reducing the latency and power consumption of data transmission and improving the overall computing performance.

[0023] The present invention is directed to a storage architecture based on dual inline memory modules (DIMM). Figure 1 The figure shows two channels in a storage architecture with DIMM as memory. Generally speaking, the host has multiple independent channels, and each channel is managed by a memory controller (MC) for storage access. One or more DIMMs are connected in the channel through a memory bus, where the memory bus includes a bus (C / A Bus) for transmitting command / address (C / A) and a data bus (Data Bus) for transmitting data. Each DIMM consists of one or two memory ranks, and each rank is further composed of multiple collaborative dynamic random access memory chips (Dynamic Random Access Memory Chip, DRAM Chip), which receive instructions and addresses at the same time, and provide or receive part of the data to form the complete data transmitted on the Data Bus. DRAM Chip can be further divided into a memory bank group (Bank Group, BG) and a memory bank. Bank consists of a two-dimensional storage array, including rows and columns. Since DRAM chips are in a collaborative relationship, accessing a bank of a certain rank actually accesses the corresponding bank in all DRAM chips of that rank at the same time.

[0024] When the Host wants to access the data of a row in a Bank of a Rank, it needs to send an Activation (ACT) instruction (containing the address of BG, Bank, Row) to this Rank to notify the Bank to store the data of this row in the row buffer, and then send a Read (RD) instruction or a Write (WR) instruction (containing the address of Column) to transfer data. If it is an RD instruction, the data at the corresponding address in the Row Buffer will be transferred to the Host MC via the Data Bus; if it is a WR instruction, the Host MC will transfer the data to the corresponding address of the Row Buffer through the Host MC and replace the original data. When the Host no longer needs the data in the Row Buffer or wants to access other rows of data in the Bank, the Host MC needs to send a Precharge (PRE) instruction to store the data in the Row Buffer back to the corresponding row of the Bank. The timing constraints between instructions are defined in the specification and need to be strictly followed to prevent physical errors.

[0025] The ACT and PRE commands will change the bank status to open and close respectively. They are commands that change the bank status and we call them Bank state change (BSC) commands. The MC needs to record (1) the bank status and (2) the storage bus status in real time to (1) calculate the time to send the next instruction and (2) generate the corresponding DDR commands from the requests. If the information it records is incorrect, a memory access error will occur, such as sending the wrong instruction at the wrong time.

[0026] In addition, the Near Memory Processing (NMP) architecture is a new computing architecture that integrates Near Memory Accelerators (NMA) near the storage unit of the memory. In traditional computing architecture, memory and processor are usually separate entities, and data transmission needs to be carried out through buses and other means; while NMP can process data in memory, reducing the latency and power consumption of data transmission and improving overall computing performance.

[0027] In the existing DIMM-based NMP architecture, NMA is generally integrated into the buffer chip next to the rank (rank level), next to the BG (BG level) or next to the bank (bank level). The present invention is aimed at rank-level DIMM-NMP, such as Figure 2 It is worth mentioning that there is generally only one level of NMA, and the present invention is only for the NMP architecture with only one level of NMA. The NMA architecture generally consists of a computing unit (Compute Unit) and a local storage controller (Local MC).

[0028] The present invention ensures the real-time consistency of the memory status of the Host MC and the NMA MC, avoids conflicts and errors, supports concurrent access of the host and the NMA, adapts to diversified computing scenarios, reduces the row switching frequency through intelligent scheduling, improves the memory bandwidth utilization, is compatible with the existing DIMM interface standards, and avoids disruptive modifications to the hardware architecture.

[0029] The specific implementation of the present invention is described below with reference to the accompanying drawings and embodiments.

[0030] First embodiment like Figure 4 and Figure 5 As shown, this embodiment provides a method for heterogeneous execution of a master control and a near memory accelerator in a near memory computing architecture, which generally includes three mechanisms: explicit synchronization, offloading-scheduling-returning, and implicit synchronization. The specific technical solution includes the following steps: S1 explicit synchronization mechanism: In Host Mode, the host memory controller (Host MC) directly sends DDR commands to the memory device, and bypasses the commands to the Near Memory Accelerator Memory Controller (NMA MC) to update the memory status recorded in the NMA MC in real time; S2 offload-scheduling-return mechanism: In concurrent memory access mode (Concurrent Mode) and near memory mode (NMA Mode), Host MC offloads memory access requests to NMA MC through the offload mechanism. NMA MC schedules host offload access requests and NMA local access requests according to the preset switch-recovery (SR) method to ensure the correctness of memory status and row locality optimization. NMA MC returns data to Host MC through time-division multiplexing interrupts and adaptive batch processing mechanism; S3 implicit synchronization mechanism: At the end of the simultaneous memory access mode and the near memory mode, the NMA MC actively adjusts the memory state according to the recorded memory bank Bank state change instruction to ensure that it is consistent with the Host MC state for switching the Host Mode.

[0031] The explicit synchronization mechanism of step S1 is to prepare for switching to NMA Mode or Concurrent Mode, and the real-time updating of the memory status recorded in the NMA MC further includes: The DDR command is transmitted via a command / address (C / A) bus, and the NMA MC processes the received host command as a command issued by itself, and updates the memory status recorded in the NMA MC, so that it can always track the latest memory status.

[0032] In this process, data is directly transferred between the host and the memory device through the data bus, bypassing the data buffer of the NMA.

[0033] In the unloading-scheduling-returning mechanism in step S2, the host memory controller unloading the memory access request to the near memory accelerator memory controller further includes: New memory commands are introduced to optimize the timing of offload operations. The host memory controller sends offload access requests to the near memory accelerator memory controller through the command / address bus, and various delays are omitted through timing optimization.

[0034] Specifically, the row activation delay tRCD and the timing constraint tRTRs are omitted in the unloaded read command RDO, and at the same time, the physical execution of the row activation instruction ACT is omitted for the unloaded read command RDO, and only its logic state is recorded; Insert read operations in the offloaded write command WRO to omit the write-to-read delay tWTR, read-to-write delay tRTW, and write delay tWL to maximize bus utilization In this embodiment, first, Figure 6 As shown in (a), the timing constraints of the continuous RD (RDO) command need to maintain: tBL, for the occupancy of the data bus (DQ bus); tCCD_s or tCCD_l, for the BG and Bank status ( Figure 6 Not shown); tRTRs, used for switching between ranks.

[0035] In this case, one of the NMA memory controllers (eg, NMA-1) will be idle, resulting in bandwidth waste.

[0036] Based on the following two points, we can omit the corresponding timing constraints: The uninstalled RD command does not actually need to read back data from the Bank; tRCD (corresponding to the delay of opening the bank row) can also be omitted because the unloaded ACT command does not actually change the bank state.

[0037] like Figure 6 As shown in (b), this more efficient offloading method enables the NMA end to obtain higher aggregate bandwidth.

[0038] Secondly, if Figure 6As shown in (c), similar bandwidth waste can be found when the WR command is offloaded. However, the timing constraints cannot be omitted in this case because writing data will occupy the data bus after issuing the WR command. To optimize this situation, if there is a read request (such as an RDO command) waiting for arbitration, we insert a read access between the offloaded write operations. This actually omits: tWTR (write-to-read delay) and tRTW (read-to-write delay) during the offload process. At the same time, the write delay tWL can also be omitted.

[0039] like Figure 6 As shown in (d), by scheduling more offload commands, the NMA memory controller can achieve higher bandwidth utilization. It should be noted that when a row switch occurs between WRO and RDO, PREO and ACTO commands are required ( Figure 6 not shown).

[0040] Next, the row locality between offload commands can be naturally preserved in the FIFO structure. Therefore, host commands and NMA requests are isolated in the CMD FIFO and REQ FIFO of each bank in the NMA MC, respectively. However, arbitrary switching between these FIFOs may significantly reduce bandwidth utilization because they usually share little row locality. In addition, offload commands generated based on the possibly incorrect memory state in the HostMC may conflict with the actual bank state if issued directly by the NMA MC. To address these challenges, we introduce the Switch-Recovery (SR) method.

[0041] In step S2, the NMA MC schedules the host offload access request and the NMA local access request according to a preset switch-recovery (SR) method, including: In NMA MC, an independent command queue CMD FIFO and request queue REQFIFO are set for each memory bank Bank, and host commands and near memory accelerator requests are isolated in the command queue CMD FIFO and request queue REQFIFO respectively. The queue is switched only when one of the following conditions is met, including: There are no remaining commands or requests in the current queue; Issue a precharge instruction PRE command.

[0042] Both conditions indicate that row locality can no longer be exploited in the FIFO. The latter condition further ensures that there is no overhead in switching. In addition, since the MC issues a PRE command when row hits reach a predetermined threshold, this approach can prevent starvation in extreme cases such as continuous access to the same row.

[0043] Then, the last BSC command from the command queue CMD FIFO and the request queue REQ FIFO is recorded to determine whether recovery is required, and the corresponding BSC command is issued when the switch is restored to ensure the consistency of the memory state after the switch.

[0044] In this embodiment, Figure 7 Take a Bank in for example: In phase ①, a PRE is sent from the CMD FIFO to update the BSC CMD and Bank status recorded in the SR unit; In phase ②, due to the issuance of PRE, MC switches to REQ FIFO. Then it generates one ACT and three RD / WR to complete all requests; In phase ③, the NMA MC switches and performs recovery. The SR unit checks the switch recovery table to issue the corresponding BSC command; In stage ④, the commands in the CMD FIFO can be issued safely; In stage ⑤, when the last PRE is issued, N2H implicit synchronization is automatically completed.

[0045] In addition, in step S2, the time division multiplexing interruption includes: Use interrupt-based signals to let NMA MC notify Host MC, and limit the effective period of interrupt signals issued by each memory column Rank by time-division multiplexing to reduce polling delay.

[0046] The adaptive batch processing mechanism further comprises: Limit interrupts to be issued in the order in which memory access is offloaded, and when the Host MC requires the NMA MC to return data, return data collectively in batches, where the batch size is equal to the number of interrupts for offloaded reads, where Disabling subsequent interrupts from triggering when earlier interrupts have not completed; Multiple read access interrupts are merged into batches, and data is returned in batches through a single return instruction RT command. The delay of the RT command is dynamically adjusted by the batch size and is executed in priority to other memory access commands.

[0047] Specifically, similar to the existing NMP design, we use an interrupt-based signal (such as a one-bit alert_n signal) to let the NMA MC notify the Host MC. However, this signal is shared by multiple ranks like the C / A bus and DQ bus. Once the alert_n signal is issued, the Host MC needs to poll all ranks to check the source of the signal. This polling delay (up to 50-80 cycles) is acceptable in accelerated designs (such as execution completion notification), but it is unacceptable for memory access.

[0048] To address this challenge, we propose a simple but effective method: time-division multiplexing the interrupt signal and limiting the effective period of each Rank to send the alert_n signal. Since the interface clocks of the Host MC and DIMM are synchronized during the calibration phase, the Host MC can determine the source Rank based on the period when the alert_n signal arrives. This method only introduces an average delay of NRank / 2 cycles while completely eliminating the polling overhead.

[0049] A key issue is that a single-bit signal cannot inform the Host MC which offloaded memory accesses have completed, given the out-of-order execution between banks. Therefore, we restrict interrupts to be issued in the order in which memory accesses are offloaded. If the interrupt for an earlier access has not been issued, the interrupt for the newer access cannot be issued even if it has completed. When the Host MC asks the NMA MC to return data, multiple interrupts may have been accepted. Therefore, the read data can be returned collectively as a batch with a batch size equal to the number of interrupts for offloaded reads.

[0050] We further design the timing constraints of RT: unlike existing commands, its latency is not fixed but determined by the batch size.

[0051] like Figure 8 As shown in the figure, the first RT command sent to NMA-0, the return data will take up the data bus tBL × 2 time. The second RT command sent to NMA-1 takes tBL time. We use the NMA index time to replace the tRL delay because the return data comes from the NMA buffer instead of the memory bank. RT can be regarded as one or more special RD commands that do not access the memory bank. In this way, it can coexist with the existing DDR commands on the memory bus without violating the protocol.

[0052] Second embodiment This embodiment provides a heterogeneous execution system of a master controller and a near memory accelerator in a near memory computing architecture, such as Fig. 9 As shown, it includes a host memory controller Host MC and a near memory accelerator memory controller NMA MC, the Host MC includes an unload table, a return unit and a command encoding module, the NMA MC includes a bypass path, a configurable FIFO queue and a switch-recovery unit, wherein, The offload table (OT) is used to record the number of offloaded commands for each memory bank Bank, and suspend the generation of new commands when the preset threshold is reached; The return unit (RU) is used to manage the timing and priority of the returned data and is configured to dynamically adjust the burst length according to the interrupt count; The command encoding module encodes the return instruction RT command by using the reserved field, and distinguishes the unload command by multiplexing the DDR command signal.

[0053] Specifically, in this embodiment, we use a unified design to implement concurrent mode and NMA mode. Each memory column Rank uses a one-bit register to indicate whether the mode is activated. The register can also be set through an interrupt. An offload table (Offload Table, OT) is introduced to record the number of commands that have been offloaded from each Bank.

[0054] When the number of commands reaches the threshold, OT prevents the generation of new commands ( Fig. 9 ①). The threshold of concurrent mode is set to the NMAFIFO size, and the threshold of NMA mode is set to 0.

[0055] Commands generated from the request buffer bypass the scheduler ( Fig. 9 ②) to omit the corresponding timing constraints. After unloading, the command will update the FSMs, scheduler and arbiter as normal ( Fig. 9 ③).

[0056] Data return processing includes: The return unit (RU) sends RT to obtain return data; The burst length is equal to the interrupt count from the read access ( Fig. 9 ④); RT only needs to update the bus occupancy recorded in the arbiter; The issuance priority of RT is set higher than other DDR commands to ensure low access latency.

[0057] The return unit records the information of offload memory access in each Rank order, including: The recorded information field includes a physical address, a read / write flag, and a command quantity and encoding method thereof; Operational mechanisms include: These two-bit registers are used to update the OT in real time when new accesses are added or old accesses are deleted. Fig. 9 ⑤); For write access, it is deleted when the corresponding interrupt is received; For read access, it is necessary to wait for interrupts and return data at the same time. At the same time, when the address and data are valid, they can be inserted into the read data buffer and then respond to higher-level memory requests.

[0058] The command encoding module uses the current C / A signal interface to encode the new command ( Fig. 9 ⑥). As shown in the following table: In the NMA MC design, the bypass path is used to pass commands from the host memory controller directly to the memory device in host mode while updating the local memory status.

[0059] Specifically, in host mode: Host commands are sent directly to the memory bank via the bypass path. Fig. 9 ⑦); These commands are further transmitted to FSMs, schedulers, and arbiters to update memory states for explicit synchronization; Similarly, write data or read data is transferred directly between memory banks, bypassing the data buffer.

[0060] The configurable FIFO queue is used to dynamically allocate FIFO queue capacity in concurrent mode, respectively store host offload access requests and near memory accelerator local access requests, and implement queue switching through an arbitrator.

[0061] In this embodiment, in concurrent mode: The offloaded commands are inserted into the CMD FIFO via the CMD decoder and distributor ( Fig. 9 ⑧); AsyncDIMM does not introduce a separate dedicated FIFO, but rebuilds the existing REQ FIFO to be configurable.

[0062] The switch-recovery unit is used to record BSC commands and generate recovery instructions to ensure the correctness of memory state and mode switching.

[0063] Specific implementation: In NMA mode: NMA MC uses the entire REQ FIFO to store NMA requests; In concurrent mode: half of the FIFO capacity is reused for host commands ( Fig. 9 ⑨). Since the command width is smaller than the request width, there is no need to modify the FIFO width. A host-NMA arbiter is introduced for each hybrid FIFO unit to arbitrate between the two FIFOs. The SR unit records the BSC command and checks whether a command needs to be generated to restore the Bank status. The host command is sent directly to the scheduler without checking the Bank status through FSMs. This design avoids modifying the logic in FSMs to support more diverse inputs.

[0064] The return unit in NMA MC is similar to Host MC, with the addition of a return unit to record the offload access status ( Fig. 9 ⑩). When the earliest access is completed, an interrupt is sent to the Host MC.

[0065] Wherein, the return unit further includes a physical address mapping module and a priority controller, wherein, The physical address mapping module is used to associate interrupt signals with corresponding memory accesses; The priority controller is used to set the issuing priority of the return instruction RT command to be higher than other DDR commands; The capacity allocation of the configurable FIFO queue is achieved in the following way: In the near memory mode, the near memory accelerator memory controller uses the entire request queue REQ FIFO to store the near memory accelerator local access requests, and in the concurrent mode, uses 50% of the capacity of the FIFO queue to store the host offload commands; The alternating scheduling of host offload commands and local access requests is achieved through an arbiter in each hybrid FIFA unit.

[0066] Finally, it should be noted that the above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

[0067] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A method for heterogeneous execution of a master controller and a near-memory accelerator in a near-memory computing architecture, characterized in that: The steps include: S1: In the host mode, the host memory controller directly sends a DDR command to the memory device, and bypasses the command to the near memory accelerator memory controller, and updates the memory status recorded in the near memory accelerator memory controller in real time; S2: In the simultaneous memory access mode and the near memory mode, the host memory controller offloads the memory access request to the near memory accelerator memory controller through the offload mechanism. The near memory accelerator memory controller schedules the host offloaded access request and the near memory accelerator local access request according to the preset switching-recovery method to ensure the correctness of the memory state and the optimization of row locality. The near memory accelerator memory controller returns the data to the host memory controller through time-division multiplexing interrupts and adaptive batch processing mechanisms. S3: When the simultaneous memory access mode and the near memory mode end, the near memory accelerator memory controller actively adjusts the memory state according to the recorded memory bank Bank state change instruction to ensure consistency with the host memory controller state for switching the host mode.

2. The method for heterogeneous execution of a master controller and a near memory accelerator in a near memory computing architecture according to claim 1, characterized in that: In step S1, the real-time updating of the memory status recorded in the near memory accelerator memory controller further includes: The DDR command is transmitted via the command / address bus, and the near memory accelerator memory controller processes the received host command as a command issued by itself, and updates the memory status recorded in the near memory accelerator memory controller, so that it can always track the latest memory status.

3. The method for heterogeneous execution of a master controller and a near memory accelerator in a near memory computing architecture according to claim 1, characterized in that: In step S2, the host memory controller offloading the memory access request to the near memory accelerator memory controller further includes: New memory commands are introduced to optimize the timing of offload operations. The host memory controller sends offload access requests to the near memory accelerator memory controller through the command / address bus, and various delays are omitted through timing optimization.

4. The method for heterogeneous execution of a master controller and a near memory accelerator in a near memory computing architecture according to claim 3, characterized in that: The timing optimization further includes: The row activation delay tRCD and the timing constraint tRTRs are omitted in the unloaded read command RDO. Meanwhile, the physical execution of the row activation instruction ACT is omitted for the unloaded read command RDO, and only its logic state is recorded. A read operation is inserted in the offloaded write command WRO to omit the write-to-read delay tWTR, the read-to-write delay tRTW, and the write delay tWL to maximize bus utilization.

5. The method for heterogeneous execution of a master controller and a near memory accelerator in a near memory computing architecture according to claim 1, characterized in that: In step S2, the near memory accelerator memory controller schedules the host offload access request and the near memory accelerator local access request according to the preset switching-restoring method, further comprising: In the near memory accelerator memory controller, an independent command queue CMD FIFO and request queue REQ FIFO are set for each memory bank Bank, and the host command and the near memory accelerator request are isolated in the command queue CMD FIFO and the request queue REQ FIFO respectively, and the queue is switched only when one of the following conditions is met, including: There are no remaining commands or requests in the current queue; Issue a pre-charge instruction PRE command; By recording the last BSC command from the command queue CMD FIFO and the request queue REQ FIFO, it is determined whether recovery is required, and the corresponding BSC command is issued when the switch is restored to ensure the consistency of the memory state after the switch.

6. The method for heterogeneous execution of a master controller and a near memory accelerator in a near memory computing architecture according to claim 1, characterized in that: In step S2, the time division multiplexing interruption further includes: An interrupt-based signal is used to allow the near memory accelerator memory controller to notify the host memory controller, and the effective period of each memory rank sending an interrupt signal is limited by time-division multiplexing the interrupt signal to reduce polling delay.

7. The method for heterogeneous execution of a master controller and a near memory accelerator in a near memory computing architecture according to claim 1, characterized in that: In step S2, the adaptive batch processing mechanism further comprises: Limit interrupts to be issued in the order in which memory access is offloaded, and when the host memory controller requires the near memory accelerator memory controller to return data, the data is collectively returned in batches, and the batch size is equal to the number of interrupts for offloaded reading, wherein, Disabling subsequent interrupts from triggering when earlier interrupts have not completed; Multiple read access interrupts are merged into batches, and data is returned in batches through a single return instruction RT command. The delay of the RT command is dynamically adjusted by the batch size and is executed in priority to other memory access commands.

8. A heterogeneous execution system of a master control and a near memory accelerator in a near memory computing architecture for implementing the method as claimed in claims 1 to 7, characterized in that: It includes a host memory controller and a near memory accelerator memory controller, wherein the host memory controller includes an unload table, a return unit and a command encoding module, and the near memory accelerator memory controller includes a bypass path, a configurable FIFO queue and a switch-recovery unit, wherein: The unloading table is used to record the number of unloaded commands for each memory bank Bank, and suspend the generation of new commands when the preset threshold is reached; The return unit is used to manage the timing and priority of the returned data and is configured to dynamically adjust the burst length according to the interrupt count; The command encoding module encodes the return instruction RT command by using the reserved field, and distinguishes the uninstallation command by multiplexing the DDR command signal; The bypass path is used to directly transmit the host memory controller's command to the memory device in the host mode and update the local memory status at the same time; The configurable FIFO queue is used to dynamically allocate FIFO queue capacity in concurrent mode, respectively store host offload access requests and near memory accelerator local access requests, and implement queue switching through an arbitrator; The switch-recovery unit is used to record BSC commands and generate recovery instructions to ensure the correctness of memory state and mode switching.

9. The heterogeneous execution system of the master control and near memory accelerator in the near memory computing architecture according to claim 8, characterized in that: The return unit further includes a physical address mapping module and a priority controller, wherein: The physical address mapping module is used to associate interrupt signals with corresponding memory accesses; The priority controller is used to set the issuing priority of the return instruction RT command to be higher than other DDR commands; The capacity allocation of the configurable FIFO queue is achieved in the following way: In the near memory mode, the near memory accelerator memory controller uses the entire request queue REQ FIFO to store the near memory accelerator local access requests, and in the concurrent mode, uses 50% of the capacity of the FIFO queue to store the host offload commands; The alternating scheduling of host offload commands and local access requests is achieved through an arbiter in each hybrid FIFA unit.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, the method for heterogeneous execution of the master control and the near memory accelerator in the near memory computing architecture described in any one of claims 1-7 is implemented.

11. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the heterogeneous execution method of the master control and the near memory accelerator in the near memory computing architecture as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Neural network accelerator based on time domain in-memory calculation and acceleration method

    CN112580793A

  • Architecture and design of a storage device controller for hyperscale infrastructure

    US20210278998A1

  • Memory disaggregation method, computing system implementing the method

    US20240012684A1