On-chip intelligent memory controller and control method for heterogeneous computing systems
By introducing an on-chip intelligent memory controller (SMC) into a heterogeneous computing system, receiving instructions from the CPU and performing data processing and transmission, the problem that the on-chip storage hierarchy in the prior art cannot be actively operated, and the collaborative work efficiency of the heterogeneous computing system is improved.
Patent Information
- Application Number
- CN202510144465.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-02-10
AI Technical Summary
In the existing heterogeneous computing system, the on-chip storage hierarchy cannot actively perform data processing operations, and the existing cache consistency protocol is not enough to support complex data operations, and it is necessary to modify the functions of the heterogeneous computing unit, which is risky and costly.
An on-chip intelligent memory controller (SMC) is set up in a heterogeneous computing system. The SMC is coupled to the heterogeneous computing unit through a cache consistency structure bus, receives instructions from the CPU and parses it, and actively performs data processing and data transfer or configuration to reduce the interaction between the CPU and other heterogeneous computing units.
It reduces the workload of the CPU, reduces the number of interactions between the CPU and other heterogeneous computing units, and improves the collaborative working efficiency between each computing unit in the heterogeneous computing system.
Smart Images

Figure CN119669121B_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to the field of heterogeneous computing, and more particularly to an on-chip intelligent memory controller and control method for a heterogeneous computing system. Background Art
[0002] In a heterogeneous computing architecture, efficient data management and operations are required to ensure the high-efficiency collaborative work of multiple different types of heterogeneous computing units (e.g., central processing unit (CPU), graphics processing unit (GPU), neural network processing unit (NPU), or domain-specific accelerator (DSA), etc.). Existing heterogeneous computing systems usually only passively provide the data access capabilities of the on-chip storage hierarchy (e.g., L1 / L2 / L3 caches, synchronous random access memory (SRAM), double data rate synchronous dynamic random access memory (DDR SDRAM, simply referred to as DDR), etc.), but these storage hierarchies cannot actively initiate and execute data processing operations based on instructions. In addition, existing cache coherence protocols are mainly designed for cache operations and are insufficient to support complex data operations in a heterogeneous computing system. Existing solutions all require modifying the functions of the heterogeneous computing units in the heterogeneous computing system, which is very risky and costly for these computing units. Summary of the Invention
[0003] In view of the problems described above, this application provides an on-chip intelligent memory controller and control method for a heterogeneous computing system.
[0004] According to one aspect of this application, there is provided an on-chip intelligent memory controller for a heterogeneous computing system, the heterogeneous computing system including a central processing unit and one or more heterogeneous computing units, the on-chip intelligent memory controller including an instruction interface module, a cache coherence processing module, and a data processing module, wherein: The instruction interface module is configured to: receive an instruction from the central processing unit, parse the instruction and determine whether the instruction is associated with cacheable data, when it is determined that the instruction is associated with cacheable data, send the instruction to the cache coherence processing module, and when it is determined that the instruction is associated with non-cacheable data, send the instruction to the data processing module; The cache coherence processing module is configured to: perform cache coherence processing on the instruction and send the instruction to the data processing module after performing the cache coherence processing; The data processing module is configured to: execute the instruction to perform corresponding data processing and send the data processing result to the instruction interface module to transfer or configure data associated with the instruction to one or more heterogeneous computing units.
[0005] According to another aspect of the present application, there is provided an intelligent memory control method for a heterogeneous computing system, the heterogeneous computing system including a central processing unit and one or more heterogeneous computing units, the intelligent memory control method including the following operations performed by an on-chip intelligent memory controller disposed in the heterogeneous computing system: receiving an instruction from the central processing unit for parsing; determining whether the instruction is associated with cacheable data; when it is determined that the instruction is associated with cacheable data, performing cache coherence processing for the instruction, and after performing the cache coherence processing, executing the instruction for corresponding data processing; when it is determined that the instruction is associated with non-cacheable data, executing the instruction for corresponding data processing; and based on the data processing result obtained by executing the instruction, transmitting or configuring data associated with the instruction to one or more heterogeneous computing units.
[0006] According to still another aspect of the present application, there is provided a heterogeneous computing system including the on-chip intelligent memory controller as described above.
[0007] According to the on-chip intelligent memory controller and control method for a heterogeneous computing system of the present application, an instruction can be received from the CPU in the heterogeneous computing system by a specially provided on-chip intelligent memory controller, parsed and executed to actively perform corresponding data processing and data transmission or configuration, so as to offload the instruction or data transmission or configuration operation that originally needs to be performed between the CPU and other heterogeneous computing units (for example, GPU, NPU or DSA) to the on-chip intelligent memory controller for execution, reducing the workload of the CPU and the interaction times between the CPU and other heterogeneous computing units, and improving the cooperation efficiency among the computing units in the heterogeneous computing system. Description of the Drawings
[0008] The present application can be better understood from the following description of the specific embodiments of the present application in conjunction with the accompanying drawings, wherein:
[0009] Figure 1 A schematic logic block diagram of an exemplary heterogeneous computing system including an on-chip intelligent memory controller according to some embodiments of the present application is shown;
[0010] Figure 2 A schematic logic block diagram of an on-chip intelligent memory controller for a heterogeneous computing system according to some embodiments of the present application is shown;
[0011] Figure 3 A schematic operation flowchart of an on-chip intelligent memory controller for a heterogeneous computing system according to some embodiments of the present application is shown;
[0012] Figure 4Shows an exemplary interaction process when executing instructions associated with cacheable data in a heterogeneous computing system according to some embodiments of the present application;
[0013] Figure 5 Shows an exemplary interaction process when executing instructions associated with non-cacheable data in a heterogeneous computing system according to some embodiments of the present application;
[0014] Figure 6 Shows examples of the load-reserved (LR) instruction and the conditional store (SC) instruction in the fifth-generation Reduced Instruction Set Computer Architecture (RISC-V) instruction set for implementing atomic operations;
[0015] Figure 7 Shows an example of an atomic operation implemented by an LR and SC instruction pair and an intelligent memory control instruction defined according to the RISC-V instruction format in a heterogeneous computing system according to some embodiments of the present application;
[0016] Figure 8 Shows an example of an intelligent memory control instruction defined according to the RISC-V instruction format according to some embodiments of the present application;
[0017] Figure 9 Shows an example of the atomic memory operation (AMO) instruction in the RISC-V instruction set for implementing atomic operations;
[0018] Figure 10 Shows an example of an atomic operation when executing an AMO instruction in a heterogeneous computing system according to some embodiments of the present application;
[0019] Figure 11 Shows an example of the interaction process for implementing the configuration operation of the CPU to the registers of the GPU in a traditional heterogeneous computing system;
[0020] Figure 12 Shows an example of the interaction process for implementing the configuration operation of the CPU to the registers of the GPU in a heterogeneous computing system according to some embodiments of the present application. Detailed implementation manners
[0021] Aspects and exemplary embodiments of the present application will be described in detail below. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to those skilled in the art that the present application may be practiced without some of these specific details. The following description of the embodiments is merely provided to better understand the present application by showing examples of the present application. The present application is in no way limited to any specific configuration set forth below, but covers any modification, replacement, and improvement of elements, components, and algorithms without departing from the spirit of the present application. Well-known structures and techniques are not shown in the drawings and the following description so as not to unnecessarily obscure the present application.
[0022] In addition, various operations will be described as multiple discrete operations in a manner that is most helpful in understanding the illustrative embodiments; however, the order of description should not be construed as implying that these operations must be order-dependent. In particular, these operations need not be performed in the order presented.
[0023] References are made herein to "embodiments", "some embodiments", etc. It should be understood, however, that the features recited in the various embodiments are not necessarily applicable only to that embodiment, but may be used in other embodiments. Features in one embodiment may be applied to another embodiment, or may be included in another embodiment. In addition, unless the context otherwise requires, the terms "comprising", "having", and "including" are synonyms. The phrases "A or B" and "A / B" mean "(A), (B), or (A and B)".
[0024] In traditional heterogeneous computing systems, the on-chip memory hierarchy (e.g., L1 / L2 / L3 caches, SRAM, DDR, etc.) typically only provides passive data access capabilities and cannot initiate and execute data operations based on instructions. Additionally, existing cache coherence protocols are mainly designed for cache operations and are insufficient to support complex data operations in heterogeneous computing systems. Existing solutions all require modifying the functionality of the computing units in the heterogeneous computing system, which is very risky and costly for these computing units.
[0025] According to an embodiment of the present application, an on-chip Smart Memory Controller (SMC) is specially provided in a heterogeneous computing system. The SMC is coupled to each heterogeneous computing unit (such as a CPU, GPU, NPU, DSA, etc.) in the heterogeneous computing system through a cache coherence architecture bus, for example. The SMC can receive instructions from the CPU, parse and execute the instructions to actively perform corresponding data processing, data transfer, or configuration, so as to offload the instruction or data transfer or configuration operations that originally need to be executed between the CPU and other heterogeneous computing units to the on-chip Smart Memory Controller for execution, reducing the workload of the CPU and the interaction times between the CPU and other heterogeneous computing units, and thus improving the cooperation efficiency among the computing units in the heterogeneous computing system.
[0026] On the other hand, the on-chip SMC according to an embodiment of the present application can perform cache coherence processing on instructions associated with cacheable data based on an existing cache coherence protocol, and then execute the instructions to perform corresponding data processing and subsequent data transfer or configuration operations, or can directly execute instructions associated with non-cacheable data to perform corresponding data processing and subsequent data transfer or configuration operations. The on-chip SMC can support the transfer or configuration of both cacheable data and non-cacheable data, such that it is not necessary to modify the functions of the computing units in the heterogeneous computing system to support complex data operations in the heterogeneous computing system.
[0027] Figure 1 FIG. shows a schematic logic block diagram of an exemplary heterogeneous computing system including an on-chip Smart Memory Controller (SMC) according to some embodiments of the present application. As Figure 1 shown, the heterogeneous computing system includes various heterogeneous computing units (such as a CPU subsystem, a GPU subsystem, an NPU subsystem) in a traditional heterogeneous computing system, a storage hierarchy (such as L2, L3 caches, DDR, a system-level buffer (SLB)), a cache coherence architecture bus, and a PCIe interface module for communicating with peripheral devices, etc. In addition to these traditional components in the traditional heterogeneous computing system, the heterogeneous computing system according to an embodiment of the present application further includes a specially provided on-chip SMC. As Figure 1 shown, the on-chip SMC and the SLB together can be referred to as an on-chip memory controller subsystem. The on-chip SMC can be coupled to each heterogeneous computing unit and each storage hierarchy in the heterogeneous computing system through a cache coherence architecture bus. The structure and working process of the on-chip SMC will be described in detail below with reference to Figures 2 to 5 FIG.
[0028] Figure 2 FIG. shows a schematic logic block diagram of an on-chip SMC for a heterogeneous computing system according to some embodiments of the present application. As Figure 2As shown, the on-chip SMC can be coupled to a system-level buffer (SLB, typically SRAM) to form an on-chip memory controller subsystem as shown in Figure 1 . The on-chip SMC includes three modules: an instruction interface module, a cache coherence module, and a data processing module. The instruction interface module is used to receive instructions from the CPU, parse the instructions, and determine whether the instructions are associated with cacheable data. When it is determined that the instructions are associated with cacheable data, the instructions are sent to the cache coherence processing module, and when it is determined that the instructions are associated with non-cacheable data, the instructions are sent to the data processing module. The cache coherence processing module is used to perform cache coherence processing on the instructions and send the instructions to the data processing module after performing the cache coherence processing. The data processing module is used to execute the instructions to perform corresponding data processing and send the data processing results to the instruction interface module to transfer or configure data associated with the instructions to one or more heterogeneous computing units. The working process of the on-chip SMC will be described in detail below with reference to Figure 3 , and the exemplary interaction process between the on-chip SMC and other components in the heterogeneous computing system will be described in detail with reference to Figure 4 and Figure 5 .
[0029] Figure 3 shows a schematic operation flowchart of an on-chip SMC for a heterogeneous computing system according to some embodiments of the present application. As shown in Figure 3 , the instruction interface module can receive instructions from the central processing unit, parse the instructions, and then determine whether the instructions are associated with cacheable data. When it is determined that the instructions are associated with cacheable data, the instruction interface module can send the instructions to the cache coherence processing module for cache coherence processing. After performing the cache coherence processing, the cache coherence processing module sends the instructions to the data processing module; when it is determined that the instructions are associated with non-cacheable data, the instruction interface module can directly send the instructions to the data processing module. The data processing module can execute the instructions to perform corresponding data processing and send the data processing results to the instruction interface module to transfer or configure data associated with the instructions to one or more heterogeneous computing units through the instruction interface module. When the SMC finishes executing the instructions, it can send an instruction execution completion confirmation to the central processing unit. In addition, before executing the instructions, the data processing module can determine whether the data associated with the instructions needs to be processed in the SLB, and when it is determined that it is necessary, it can instruct the SLB to perform the corresponding data processing first. Additionally, according to some embodiments of the present application, when performing cache coherence processing, the cache coherence processing module may also need to determine whether the data associated with the instructions needs to be processed in the SLB, and when it is determined that it is necessary, it can instruct the SLB to perform the corresponding data processing.
[0030] Figure 4An exemplary interaction process is shown when executing instructions associated with cacheable data in a heterogeneous computing system according to some embodiments of the present application. As Figure 4 shown, RN-F0, RN-F1, and RN-F2 represent three coherent request nodes. For example, RN-F0 may represent a CPU core in a CPU subsystem in a heterogeneous computing system, and this CPU core sends instructions associated with cacheable data (e.g., RISC-V instructions) to the SMC for data transfer or configuration to the GPU or NPU. When the SMC executes instructions associated with cacheable data, cache coherence processing needs to be performed first. For example, the SMC may need to interact with two related CPU cores (e.g., represented by RN-F1 and RN-F2) for cache coherence processing, such as processing based on the snoop coherence protocol. Then, the SMC can parse and execute the RISC-V instructions. When the SMC executes the RISC-V instructions, it may also be necessary to interact with the SLB or DDR to perform memory access operations. After the SMC finishes executing the RISC-V instructions, it can send an instruction execution completion confirmation to RN-F0 and perform data transfer or configuration associated with the RISC-V instructions to the heterogeneous computing unit (GPU or NPU).
[0031] Figure 5 An exemplary interaction process is shown when executing instructions associated with non-cacheable data in a heterogeneous computing system according to some embodiments of the present application. Similar to Figure 4 this, the SMC can parse and execute the RISC-V instructions from RN-F0, and when executing the RISC-V instructions, it may also be necessary to interact with the SLB or DDR to perform memory access operations. After the SMC finishes executing the RISC-V instructions, it can send an instruction execution completion confirmation to RN-F0 and perform data transfer or configuration associated with the RISC-V instructions to heterogeneous computing units such as the GPU or NPU. Figure 5 The difference from Figure 4 this is only that the RSIC-V instructions sent by RN-F0 to the SMC are associated with non-cacheable data, so the SMC does not need to perform cache coherence processing and can directly execute the RSIC-V instructions.
[0032] Generally speaking, according to an embodiment of the present application, a on-chip SMC for a heterogeneous computing system and a corresponding intelligent memory control method are provided. The intelligent memory control method includes the following operations performed by the on-chip SMC: receiving an instruction from the CPU for parsing; determining whether the instruction is associated with cacheable data; when it is determined that the instruction is associated with cacheable data, performing cache coherence processing for the instruction, and after performing the cache coherence processing, executing the instruction for corresponding data processing; when it is determined that the instruction is associated with non-cacheable data, executing the instruction for corresponding data processing; and based on the data processing result obtained by executing the instruction, transferring or configuring data associated with the instruction to one or more heterogeneous computing units. By using the on-chip SMC and the corresponding intelligent memory control method according to the embodiments of the present application, the operations of transferring or configuring instructions or data that originally need to be performed between the CPU and other heterogeneous computing units can be offloaded to the SMC for execution, reducing the workload of the CPU and the number of interactions between the CPU and other heterogeneous computing units, thereby improving the efficiency of collaborative work among the computing units in the heterogeneous computing system.
[0033] According to an embodiment of the present application, the CPU can operate based on various instruction sets, especially can operate based on the currently widely used RSIC-V instruction set (including its extended instruction set). Correspondingly, according to an embodiment of the present application, the SMC can support the execution of various RSIC-V instructions in the RSIC-V instruction set (including its extended instruction set). In some embodiments of the present application, the SMC can support atomic operations implemented using load-reserved (LR) instructions and store-conditional (SC) instructions.
[0034] Figure 6 Examples of the LR instruction and the SC instruction in the RISC-V instruction set for implementing atomic operations are shown. Figure 7 Examples of executing atomic operations implemented by the LR and SC instruction pair and intelligent memory control instructions defined in the RISC-V instruction format in a heterogeneous computing system according to some embodiments of the present application are shown.
[0035] The LR instruction is used to load the value of a memory address and "reserve" this address, that is, for a period of time afterwards, access to this address by other processors will be blocked until the paired SC instruction is executed. For example, assume that in a traditional heterogeneous computing system, the CPU needs to perform N interactions with the GPU between the LR instruction and the SC instruction pair to complete a certain atomic operation. Then, in general, the CPU needs to perform 2 + N interactions with the GPU to complete this atomic operation. In contrast, in the heterogeneous computing system according to an embodiment of the present application, the CPU will only need to perform 3 interactions with the SMC (including LR, SC, and SMC_OP) to complete this atomic operation, and subsequent operations will be completed by the SMC interacting with the GPU. In this way, the ratio of the number of interactions required by the CPU when using the SMC to the number of interactions required by the CPU when not using the SMC is 3 / (2 + N). Generally, the more data that needs to be transferred or configured, the more the bandwidth efficiency is improved. Taking a GPU with 8 GPU cores or 128 vector engines as an example, this ratio is 30%, which means that when using the SMC, the number of interactions of the CPU and the related power consumption can be reduced by 70%. When the GPU has more cores, the number of interactions of the CPU and the related power consumption will be reduced more. For example, when the GPU has 16 GPU cores, the number of interactions of the CPU and the related power consumption can be reduced by 83.3%.
[0036] Figure 8 Shows examples of intelligent memory control instructions defined according to the RISC-V instruction format in some embodiments of the present application. As Figure 8 shown, multiple SMC operations can be defined according to the RISC-V instruction format, and their operation codes can be represented by SMO.
[0037] In addition to the LR / SC instruction pair, the RISC-V instruction set also provides atomic memory operation (AMO) instructions to implement more direct and efficient atomic operations. Figure 9 Shows examples of atomic memory operation (AMO) instructions in the RISC-V instruction set for implementing atomic operations. Figure 10 Shows examples of atomic operations of executing AMO instructions in a heterogeneous computing system according to some embodiments of the present application. As Figure 10The AMO instructions shown can be executed in the SMC. That is, data processing, data transfer, or configuration related to the AMO instructions can all be completed through the interaction between the SMC and heterogeneous computing units such as GPUs or NPUs, which can greatly reduce the number of interactions between the CPU and these heterogeneous computing units. In addition, when the SMC executes these AMO instructions, it is not restricted by the CPU cache line, and the granularity of data operations can be extended from the cache line to a page or even a data block of any size. Moreover, since the SMC is closer to the data storage location (such as the SLB) compared to the CPU core, executing these AMO instructions in the SMC may also reduce data operation time and latency.
[0038] In addition to the atomic operation instructions in the above RISC-V instruction set, the heterogeneous computing system including an intelligent memory controller according to some embodiments of the present application can also support other new extended intelligent memory control instructions for memory operations defined in the RISC-V instruction format. For example, it can support the configuration operation of registers in the GPU by the CPU or the weight replication operation of the NPU by the CPU, etc. Hereinafter, taking the configuration operation of registers in the GPU by the CPU as an example, the improvement of the heterogeneous computing system according to the embodiments of the present application compared to the traditional heterogeneous computing system will be further described.
[0039] Figure 11 An example of the interaction process for implementing the configuration operation of registers in the GPU by the CPU in a traditional heterogeneous computing system is shown. Figure 12 An example of the interaction process for implementing the configuration operation of registers in the GPU by the CPU in a heterogeneous computing system according to some embodiments of the present application is shown.
[0040] As Figure 11As shown, the CPU needs to interact with the GPU N times to complete the configuration of N registers in the GPU. In the heterogeneous computing system according to an embodiment of the present application, the CPU only needs to send an SMC instruction indicating a register configuration operation to the SMC, and the SMC will parse the instruction and perform subsequent relevant data processing and register configuration operations. In other words, after receiving the SMC instruction from the CPU, the SMC will take over as the host for the data processing and register configuration operations associated with the SMC instruction. Therefore, the number of interactions required by the CPU will be reduced from N to 2, and the ratio of the number of interactions required by the CPU with the use of the SMC to the number of interactions required by the CPU without the use of the SMC is 2 / N. The more registers the CPU needs to configure, the more the bandwidth efficiency is improved. Similarly, taking a GPU with 8 GPU cores or 128 vector engines as an example, this ratio is 25%, which means that with the use of the SMC, the number of interactions and related power consumption of the CPU can be reduced by 75%. When the GPU has more cores, the number of interactions and related power consumption of the CPU will be reduced more. For example, when the GPU has 16 GPU cores, the number of interactions and related power consumption of the CPU can be reduced by 87.5%.
[0041] In addition, the heterogeneous computing system including the intelligent memory controller according to an embodiment of the present application can also support various other data operation instructions as needed. For example, a broadcast data (BDC) instruction for simultaneously sending the same data from the SRAM to multiple destinations; a multicast data (MCD) instruction, similar to the BDC instruction, but allowing the data to be sent to a specified group of computing units; a data query (DQY) instruction allowing querying the attributes of data stored in the SRAM, such as checking the data type, size, or content-based query; a checksum and hash (CSH) instruction for directly generating checksum and hash values from the data stored in the SRAM; and so on.
[0042] Generally, in the heterogeneous computing system according to an embodiment of the present application, by setting up the on-chip SMC to offload the instruction or data transfer or configuration operations that originally needed to be performed between the CPU and other heterogeneous computing units to the on-chip SMC for execution, the workload of the CPU and the number of interactions between the CPU and other heterogeneous computing units are reduced, and the efficiency of collaborative work between the computing units in the heterogeneous computing system is improved. In addition, the on-chip SMC can support the transfer or configuration of both cacheable data and non-cacheable data at the same time, so that it is not necessary to modify the functions of the computing units in the heterogeneous computing system to support complex data operations in the heterogeneous computing system.
[0043] The following paragraphs describe examples of various embodiments of the present application.
[0044] Example 1 includes an on-chip intelligent memory controller for a heterogeneous computing system. The heterogeneous computing system includes a central processing unit and one or more heterogeneous computing units. The on-chip intelligent memory controller includes an instruction interface module, a cache coherence processing module, and a data processing module, where: The instruction interface module is configured to: receive instructions from the central processing unit, parse the instructions and determine whether the instructions are associated with cacheable data. When it is determined that the instructions are associated with cacheable data, send the instructions to the cache coherence processing module. And when it is determined that the instructions are associated with non-cacheable data, send the instructions to the data processing module; The cache coherence processing module is configured to: perform cache coherence processing on the instructions and send the instructions to the data processing module after performing the cache coherence processing; The data processing module is configured to: execute the instructions to perform corresponding data processing and send the data processing result to the instruction interface module to transfer or configure data associated with the instructions to one or more heterogeneous computing units.
[0045] Example 2 includes the on-chip intelligent memory controller according to Example 1, where the on-chip intelligent memory controller is coupled to a system-level buffer in the heterogeneous computing system. And before executing the instructions, the data processing module is further configured to: determine whether the data associated with the instructions needs to be processed in the system-level buffer. And when it is determined that it is necessary, instruct the system-level buffer to perform corresponding data processing.
[0046] Example 3 includes the on-chip intelligent memory controller according to Example 1 or 2, where when performing cache coherence processing, the cache coherence processing module is further configured to: determine whether the data associated with the instructions needs to be processed in the system-level buffer. And when it is determined that it is necessary, instruct the system-level buffer to perform corresponding data processing.
[0047] Example 4 includes the on-chip intelligent memory controller according to any one of Examples 1 to 3, where the central processing unit includes one or more central processing unit cores based on the RISC-V instruction set architecture.
[0048] Example 5 includes the on-chip intelligent memory controller according to any one of Examples 1 to 4, where the instructions include at least one of the following items: intelligent memory control instructions executed between a load-reserved LR and a conditional-store SC instruction pair defined in the RISC-V instruction format; atomic memory operation AMO instructions in the RISC-V instruction set; or newly extended intelligent memory control instructions defined in the RISC-V instruction format.
[0049] Example 6 includes the on-chip intelligent memory controller according to any one of Examples 1 to 5, where the one or more heterogeneous computing units include one or more of the following: a graphics processing unit, a neural network processing unit, or a domain-specific accelerator.
[0050] Example 7 includes the on-chip intelligent memory controller according to any one of Examples 1 to 6, wherein the instruction interface module is coupled to the central processing unit and one or more heterogeneous computing units through a cache coherence architecture bus.
[0051] Example 8 includes an intelligent memory control method for a heterogeneous computing system, the heterogeneous computing system including a central processing unit and one or more heterogeneous computing units, the intelligent memory control method including the following operations performed by an on-chip intelligent memory controller provided in the heterogeneous computing system: receiving an instruction from the central processing unit for parsing; determining whether the instruction is associated with cacheable data; when it is determined that the instruction is associated with cacheable data, performing cache coherence processing for the instruction, and after performing the cache coherence processing, executing the instruction for corresponding data processing; when it is determined that the instruction is associated with non-cacheable data, executing the instruction for corresponding data processing; and based on the data processing result obtained by executing the instruction, transferring or configuring data associated with the instruction to one or more heterogeneous computing units.
[0052] Example 9 includes the intelligent memory control method according to Example 8, wherein the on-chip intelligent memory controller is coupled to a system-level buffer in the heterogeneous computing system, and before executing the instruction, the intelligent memory control method further includes: determining whether the data associated with the instruction needs to be processed in the system-level buffer, and when it is determined that it is needed, instructing the system-level buffer to perform corresponding data processing.
[0053] Example 10 includes the intelligent memory control method according to Example 8 or 9, wherein the intelligent memory control method further includes: when performing cache coherence processing, determining whether the data associated with the instruction needs to be processed in the system-level buffer, and when it is determined that it is needed, instructing the system-level buffer to perform corresponding data processing.
[0054] Example 11 includes the intelligent memory control method according to any one of Examples 8 to 10, wherein the central processing unit includes one or more central processing unit cores based on the RISC-V instruction set architecture.
[0055] Example 12 includes the intelligent memory control method according to any one of Examples 8 to 11, wherein the instruction includes at least one of the following items: an intelligent memory control instruction executed between a load-reserved LR and a conditional-store SC instruction pair defined in the RISC-V instruction format; an atomic memory operation AMO instruction in the RISC-V instruction set; or a newly extended intelligent memory control instruction defined in the RISC-V instruction format.
[0056] Example 13 includes the intelligent memory control method according to any one of Examples 8 to 12, wherein the one or more heterogeneous computing units include one or more of the following: a graphics processing unit, a neural network processing unit, or a domain-specific accelerator.
[0057] Example 14 includes a heterogeneous computing system including an on-chip intelligent memory controller according to any one of Examples 1 to 7.
[0058] Although certain embodiments have been illustrated and described herein for purposes of description, various alternative and / or equivalent embodiments or implementations that accomplish the same purpose may be substituted for the embodiments shown and described without departing from the scope of the present disclosure. This application is intended to cover any adaptations or variations of the embodiments discussed herein. Accordingly, the embodiments described herein are clearly limited only by the appended claims and their equivalents.
Claims
1. An on-chip intelligent memory controller for a heterogeneous computing system, wherein, The heterogeneous computing system includes a central processing unit and one or more heterogeneous computing units. The on-chip intelligent memory controller is coupled to the central processing unit and the one or more heterogeneous computing units through a cache coherence architecture bus. And the on-chip intelligent memory controller includes an instruction interface module, a cache coherence processing module, and a data processing module, wherein: The instruction interface module is configured to: receive an instruction from the central processing unit, parse the instruction and determine whether the instruction is associated with cacheable data. When it is determined that the instruction is associated with cacheable data, send the instruction to the cache coherence processing module. And when it is determined that the instruction is associated with non-cacheable data, send the instruction to the data processing module; The cache coherence processing module is configured to: perform cache coherence processing on the instruction and send the instruction to the data processing module after performing the cache coherence processing; The data processing module is configured to: execute the instruction to perform corresponding data processing and send the data processing result to the instruction interface module to transfer or configure data associated with the instruction to the one or more heterogeneous computing units.
2. The on-chip intelligent memory controller according to claim 1, wherein, The on-chip intelligent memory controller is coupled to a system-level buffer in the heterogeneous computing system. And, before executing the instruction, the data processing module is further configured to: determine whether the data associated with the instruction needs to be processed in the system-level buffer, and when it is determined that it is needed, instruct the system-level buffer to perform corresponding data processing.
3. The on-chip intelligent memory controller according to claim 2, wherein, When performing the cache coherence processing, the cache coherence processing module is further configured to: determine whether the data associated with the instruction needs to be processed in the system-level buffer, and when it is determined that it is needed, instruct the system-level buffer to perform corresponding data processing.
4. The on-chip intelligent memory controller according to claim 1, wherein, The central processing unit includes one or more central processing unit cores based on the RISC-V instruction set architecture.
5. The on-chip intelligent memory controller according to claim 1, wherein, The instruction includes at least one of the following items: An intelligent memory control instruction executed between a load-reserved LR and a conditional store SC instruction pair defined according to the RISC-V instruction format; An atomic memory operation AMO instruction in the RISC-V instruction set; or A newly extended intelligent memory control instruction defined according to the RISC-V instruction format.
6. The on-chip intelligent memory controller according to claim 1, wherein, The one or more heterogeneous computing units include one or more of the following: a graphics processing unit, a neural network processing unit, or a domain-specific accelerator.
7. An intelligent memory control method for a heterogeneous computing system, the heterogeneous computing system including a central processing unit and one or more heterogeneous computing units and an on-chip intelligent memory controller coupled to the central processing unit and the one or more heterogeneous computing units through a cache coherence architecture bus. The intelligent memory control method includes the following operations performed by the on-chip intelligent memory controller: Receive an instruction from the central processing unit for parsing; Determine whether the instruction is associated with cacheable data; When it is determined that the instruction is associated with cacheable data, cache coherence processing is performed for the instruction, and after the cache coherence processing is performed, the instruction is executed to perform corresponding data processing; When it is determined that the instruction is associated with non-cacheable data, the instruction is executed to perform corresponding data processing; And Based on the data processing result obtained by executing the instruction, data associated with the instruction is transferred or configured to the one or more heterogeneous computing units.
8. The intelligent memory control method according to claim 7, wherein, The on-chip intelligent memory controller is coupled to a system-level buffer in the heterogeneous computing system. Before executing the instruction, the intelligent memory control method further includes: Determining whether data associated with the instruction needs to be processed in the system-level buffer, and When it is determined that it is needed, instructing the system-level buffer to perform corresponding data processing.
9. The intelligent memory control method according to claim 8, wherein, The intelligent memory control method further includes: when performing the cache coherence processing, determining whether data associated with the instruction needs to be processed in the system-level buffer, and when it is determined that it is needed, instructing the system-level buffer to perform corresponding data processing.
10. The intelligent memory control method according to claim 9, wherein, The central processing unit includes one or more central processing unit cores based on the RISC-V instruction set architecture.
11. The intelligent memory control method according to claim 7, wherein, The instruction includes at least one of the following items: An intelligent memory control instruction executed between a load-reserved LR and a conditional-store SC instruction pair defined according to the RISC-V instruction format; An atomic memory operation AMO instruction in the RISC-V instruction set; or A newly extended intelligent memory control instruction defined according to the RISC-V instruction format.
12. The intelligent memory control method according to any one of claims 7 to 11, wherein, The one or more heterogeneous computing units include one or more of the following: a graphics processing unit, a neural network processing unit, or a domain-specific accelerator.
13. A heterogeneous computing system, including the on-chip intelligent memory controller according to any one of claims 1 to 6.
Citation Information
Patent Citations
Heterogeneous computing system and operating method thereof
CN110083547A
Program store compare handling between instruction and operand caches
US6865645B1