A method for accelerating core request completion in a shared memory system
By using the NoLocalProbe and SrcInf flags in a shared memory system, the completion message is returned to the source core directly when the Home agent receives a write access request, which solves the memory access delay problem caused by cache consistency in a multi-core processor system and achieves improvement in system performance.
Patent Information
- Application Number
- CN202211240878.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-11
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-10-11
AI Technical Summary
In multi-core processor systems, cache consistency issues lead to long memory access delays, affecting system performance.
By introducing the ‘No Local Cache Consistency Flag’ (NoLocalProbe) and ‘Core Request Information’ (SrcInf) in the shared memory system, the completion message is returned directly to the source core when the Home agent receives a write access request, reducing the communication steps between the source core, the Cache agent and the Home agent.
Without violating the storage consistency sequence model, this method optimizes the interaction between the source core, the Cache agent and the Home agent, reducing the average memory access delay of a single-channel multi-core or multi-core shared memory system.
Smart Images

Figure CN115543201B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of a method for processing a core memory access request in a shared memory system, and more particularly to a method for accelerating the completion of a core request in a shared memory system. (This patent is applicable to both single-channel systems and multi-channel systems, and the emphasis on multi-channel will not be emphasized in the following.) Background Art
[0002] In a multi-core processor system, the speed difference between the processor and the memory leads to the emergence of the "memory wall" problem. The introduction of the cache storage system alleviates the speed mismatch problem between the processor and the memory. In most multi-core processors, each core is set up with one or more levels of private cache (Cache), and multiple cores access the shared last level cache (LLC) through the on-chip network (Network On Chip, NoC). Since the program has temporal locality, the processor core creates a copy of the data in the local cache after obtaining the required data, regardless of whether the same data exists in the cache of other cores. This gives rise to the cache coherence problem (Cache Coherence), the root cause of which is: in a multi-core processor system, there may be multiple copies of the same data in different caches, affecting the correctness of program execution. In this way, a cache coherence protocol is needed to manage multiple copies of shared data.
[0003] Cache consistency has multiple definitions. The definition given by Gharachorloo et al. is:
[0004] (1) Each write operation is visible to all cores;
[0005] (2) Write operations are sequential, that is, all cores observe the same memory access sequence. Therefore, cache consistency requires that write operations must eventually be broadcast to all participating processors. At the same time, write operations to the same address observed by the participating cores must be performed in the same order.
[0006] In addition, when writing a parallel program, we hope to establish an order model between reading and writing. According to this model, the programmer can infer the execution results of the program and its correctness. This model is memory consistency (Memory Coherence).
[0007] A complete consistency model includes two aspects: cache consistency and storage consistency, and the two are complementary: cache consistency defines the behavior of read and write operations on the same storage address, while the storage consistency model defines the read and write behavior of accessing all storage addresses. In a shared storage space, multiple processes perform concurrent read and write operations on the same or different units of storage, and each process sees an order in which these operations are completed. A storage consistency model specifies several constraints on this order. Summary of the invention
[0008] The technical problem to be solved by the present invention is to provide a method for accelerating the completion of core requests in a shared memory system, thereby reducing the average memory access delay of a single-channel multi-core or multi-channel multi-core shared memory system.
[0009] The technical solution adopted by the present invention to solve the technical problem is: to provide a method for accelerating the completion of core requests in a shared memory system, comprising:
[0010] The shared memory system includes a request source CPU and other CPUs, each of the CPUs includes n cores, m LLC bodies and k main memory controllers MC, each of the LLC bodies is managed by a Cache agent, each of the main memory controllers MC is managed by a Home agent, and the n cores and m LLCs, and the m LLC bodies and k main memory controllers MCs communicate via an on-chip network;
[0011] Methods for accelerating core request completion in a shared memory system include:
[0012] Step (1): core #i in the request source CPU performs write access to address A, and sends the write access request to the LLC body and cache agent corresponding to address A in the request source CPU;
[0013] Step (2): The Cache agent receives the write access request of core #i and queries the LLC directory corresponding to address A. If other cores of the requesting source CPU in the LLC directory have a copy of address A, an invalidation probe request is initiated to the core having a copy of address A, and a global write request is sent to the Home agent corresponding to the target main memory controller MC at the same time. The global write request carries the no local cache consistency flag NoLocalProbe and the core #i request information SrcInf.
[0014] Determine whether the core #i request will cause a Probe operation on other cores of the request source CPU. If the core #i request will not cause an invalid Probe request operation on other cores of the request source CPU, set NoLocalProbe=1, otherwise NoLocalProbe=0;
[0015] Step (3): When NoLocalProbe=1, the Home agent of other CPUs receives the write access request of core #i, saves the no local cache consistency flag NoLocalProbe and the core #i request information SrcInf in the suspended buffer of the Home agent, and initiates an access request to the main memory controller MC; wherein, if it is a multi-way CPU system, the data consistency between multiple CPUs is checked when the main memory controller MC accesses the request. It should be noted that if the source core accesses the main memory space managed by the local CPU, the Home agent requests the source CPU, otherwise it requests the other CPUs;
[0016] Step (4): After the Home agent completes the data consistency operation between multiple CPUs caused by the write access request of core #i, the Home agent queries the suspended buffer of the Home agent; if NoLocalProbe=1, the Home agent returns a core request completion message to core #i according to the core #i request information SrcInf, and sends a global consistency completion message to the Cache agent at the same time;
[0017] Step (5): After receiving the global consistency completion message, the Cache agent no longer sends a core request completion message to core #i.
[0018] The core #i request information SrcInf in step (2) includes the requested core CPU number, the core number in the CPU and the core request suspension buffer number information, ensuring that the Home agent can return a completion message to the core #i through the core #i request information SrcInf.
[0019] The step (3) checks the data consistency between multiple CPUs when the main memory controller MC makes an access request, specifically: if other CPUs also have copies of address A, then the copies of address A in other CPUs are invalidated.
[0020] The step (4) also includes: if NoLocalProbe=0, the Home agent only sends a global consistency completion message to the Cache agent.
[0021] The global consistency completion message in step (4) includes: the Home agent has sent a Cmpl message flag CmplSended to core #i.
[0022] In step (5), after the Cache agent receives the global consistency completion message from the Home agent, it determines whether to return the core request completion message to core #i based on the CmplSended flag. If CmplSended=0, the core request completion message is returned to core #i, otherwise it is not returned.
[0023] Beneficial Effects
[0024] Due to the adoption of the above-mentioned technical solution, the present invention has the following advantages and positive effects compared with the prior art: In order to solve the problem of slow submission speed of core request completion message, when Local CA determines that the request source CPU has no core with a data copy, the present invention allows HA to directly return the completion message to the request source core. Without violating the storage consistency order model, this method reduces the average memory access delay of a single-channel multi-core or multi-channel multi-core shared memory system by optimizing the interaction mode between the source core, CA and HA. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic diagram of a conventional execution process of a core memory access request in a multi-channel multi-core memory system according to an embodiment of the present invention;
[0026] Figure 2 It is a schematic diagram of an optimized execution process of a core memory access request in a multi-channel multi-core memory system according to an embodiment of the present invention;
[0027] Figure 3 is a conventional message flow chart of a core write class request in a multi-channel multi-core memory system according to an embodiment of the present invention;
[0028] Figure 4 This is an optimized message flow chart of a core write class request in an implementation mode of the present invention that does not cause local consistency. DETAILED DESCRIPTION
[0029] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall within the scope limited by the appended claims of the application equally.
[0030] The embodiment of the present invention relates to a method for accelerating the completion of core requests in a shared memory system. First, an overview of a multi-core shared memory system is given: Generally, a multi-core shared memory system can be abstractly described as follows: a processor (i.e., CPU) includes n cores, m LLC bodies, and k main memory controllers (Memory Control, MC), which communicate through one or more on-chip networks (see Figure 1 ). To facilitate the subsequent description, the following definitions are proposed:
[0031] 1) Cache Agent (CA): The management component of the processor's last-level cache, responsible for ensuring data consistency between the private caches and LLCs of multiple cores on the chip, and interacting with the Home agent. CA records the distribution of data copies of all cores and LLCs of the request source CPU in a directory manner. When a core initiates a write operation, the write invalidation strategy is used to clear the private copies of other cores.
[0032] 2) Home Agent (HA): The management component of the main memory space of this processor, responsible for the read and write operations of the main memory. For multi-way systems, it is also necessary to ensure data consistency between multiple CPU LLCs and main memory. Protocols such as directory or broadcast monitoring can be used to solve the cache consistency problem between LLCs.
[0033] The following uses an application scenario as an example to describe the core memory access process of a traditional multi-core shared memory system. Multiple cores read address A and save copies of data A in their own private caches. At this time, the core cache and LLC are in invalid (Invalid) or clean sharing state (Share or Forward). Then, a core writes to address A. The general request processing process is as follows (see Figure 3 ):
[0034] 1) At time T1, core #i performs a write access to address A. The write request is sent to the LLC body and CA (defined as Local CA) corresponding to address A in the request source CPU;
[0035] 2) At time T2, Local CA queries the LLC directory corresponding to address A. If the request source CPU has cores with copies, an invalidation probe request is sent to these cores. If the LLC directory is not writable, a global write request is also required to be sent to the HA corresponding to the target main memory. In general, the global request carries the information of the Local CA agent, but not the source core information;
[0036] 3) At time T3, HA receives a global write request. If it finds that other CPUs also have copies of data A, it needs to invalidate the copies of data A in other CPUs.
[0037] 4) At time T4, HA completes data consistency between CPUs and returns a "global consistency completion message" to Local CA;
[0038] 5) At time T5, Local CA receives the "global consistency completion message" and collects the "invalidation replies" returned by all cores in the request source CPU that have copies of A, and returns a completion message (Complete message) to the request source core #i.
[0039] 6) At time T6, after core #i receives the completion message of the write request at address A, it releases the corresponding request resources.
[0040] The storage consistency order rule requires that the request source core and Local CA must be cleared of all cache copies before subsequent access to address A is allowed.
[0041] In this process, the communication between the source core, Local CA and HA requires 4 steps, and the memory access delay is relatively long. Regardless of whether other cores of the CPU where the source core is located have data copies, the global consistency completion message is first transmitted to the Local CA. After CA receives all the global completion messages and local Probe replies, it returns the request completion message to the source core.
[0042] Add "No Local Cache Coherence Flag" (NoLocalProbe) and core request information (SrcInf, including CPU number, core number, core request ID number, etc.) to the global memory access request sent by the Cache agent to the Home agent. HA can use this information to directly return a completion message to the source core. Add "HA has returned a completion message flag to the request source" (CmplSended) to the global coherence completion message returned by the Home agent to the Cache agent.
[0043] The following is a detailed introduction to this implementation method:
[0044] The following takes the above core #i making a write request to address A as an example to introduce the write request processing flow of this implementation mode (see Figure 2 and Figure 4 ):
[0045] 1) At time T1, core #i performs a write access to address A. The write request is sent to the LLC body and CA (defined as Local CA) corresponding to address A in the request source CPU;
[0046] 2) At time T2, Local CA queries the LLC directory corresponding to address A. If the request source CPU has a core with a copy, it sends an invalid probe request to these cores; at the same time, it sends a global write request to HA, carrying NoLocalProbe and request source information. If the request source CPU has a core with a copy, NoLocalProbe = '0', and the subsequent process is the same as Figure 3 ; Otherwise, NoLocalProbe = '1', and the subsequent process is as follows;
[0047] 3) At time T3, HA receives the global write request and saves NoLocalProbe and SrcInf in the HA request suspension buffer. If other CPUs also have copies of A data, the copies of A of other CPUs need to be invalidated;
[0048] 4) At time T4, HA completes data consistency between processors. Since NoLocalProbe = '1', HA returns a completion message to the source core according to the SrcInf information, and returns a "global consistency completion message" (CmplSended = '1') to the Local CA;
[0049] 5) At time T5a, core #i receives the completion message of the write request at address A and releases the corresponding request resources.
[0050] At time T5b, Local CA receives the "global consistency completion message", and because CmplSended = '1', it no longer returns a completion message to the source core #i.
[0051] In this process, the communication between the source core, Local CA and HA only requires three steps. Figure 3 The traditional process reduces one step, eliminates the on-chip network transmission delay of the global consistency completion message from HA to CA and the delay from CA processing the global consistency completion message to generating the core completion message, shortening the average access delay of single-channel or multi-channel systems.
[0052] The foregoing description of specific exemplary embodiments of the present invention is for the purpose of illustration and demonstration. These descriptions are not intended to limit the present invention to the precise form disclosed, and it is clear that many changes and variations can be made based on the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present invention and its practical application, so that those skilled in the art can realize and utilize various different exemplary embodiments of the present invention and various different selections and changes. The scope of the present invention is intended to be limited by the claims and their equivalents.
Claims
1. A method for accelerating the completion of core requests in a shared memory system, It is characterized in that The shared memory system includes a request source CPU and other CPUs, each of the CPUs includes n cores, m LLC bodies and k main memory controllers MC, each of the LLC bodies is managed by a Cache agent, each of the main memory controllers MC is managed by a Home agent, and the n cores and m LLCs, and the m LLC bodies and k main memory controllers MCs communicate via an on-chip network; Methods for accelerating core request completion in a shared memory system include: Step (1): core #i in the request source CPU performs write access to address A, and sends the write access request to the LLC body and cache agent corresponding to address A in the request source CPU; Step (2): The Cache agent receives the write access request of core #i and queries the LLC directory corresponding to address A. If other cores of the requesting source CPU in the LLC directory have a copy of address A, an invalidation probe request is initiated to the core having a copy of address A, and a global write request is sent to the Home agent corresponding to the target main memory controller MC at the same time. The global write request carries the no local cache consistency flag NoLocalProbe and the core #i request information SrcInf. Determine whether the core #i request will cause a Probe operation on other cores of the request source CPU. If the core #i request will not cause an invalid Probe request operation on other cores of the request source CPU, set NoLocalProbe=1, otherwise NoLocalProbe=0; Step (3): When NoLocalProbe=1, the Home agent receives a write access request from core #i, saves the no-local-cache-consistency flag NoLocalProbe and the core #i request information SrcInf in the suspended buffer of the Home agent, and initiates a main memory controller MC access request; wherein, when the main memory controller MC accesses the request, the data consistency between the multiple CPUs is checked; Step (4): After the Home agent completes the data consistency operation between multiple CPUs caused by the write access request of core #i, the Home agent queries the suspended buffer of the Home agent; if NoLocalProbe=1, the Home agent returns a core request completion message to core #i according to the core #i request information SrcInf, and sends a global consistency completion message to the Cache agent at the same time; Step (5): After receiving the global consistency completion message, the Cache agent no longer sends a core request completion message to core #i.
2. The method for accelerating the completion of core requests in a shared memory system according to claim 1, It is characterized in that The core #i request information SrcInf in step (2) includes the requested core CPU number, the core number in the CPU and the core request suspension buffer number information, ensuring that the Home agent can return a completion message to the core #i through the core #i request information SrcInf.
3. The method for accelerating the completion of core requests in a shared memory system according to claim 1, It is characterized in that The step (3) checks the data consistency between multiple CPUs when the main memory controller MC makes an access request, specifically: if other CPUs also have copies of address A, then the copies of address A in other CPUs are invalidated.
4. The method for accelerating the completion of core requests in a shared memory system according to claim 1, It is characterized in that The step (4) also includes: if NoLocalProbe=0, the Home agent only sends a global consistency completion message to the Cache agent.
5. The method for accelerating the completion of core requests in a shared memory system according to claim 1, It is characterized in that The global consistency completion message in step (4) includes: the Home agent has sent a Cmpl message flag CmplSended to core #i.
6. The method for accelerating the completion of core requests in a shared memory system according to claim 5, It is characterized in that In step (5), after the Cache agent receives the global consistency completion message from the Home agent, it determines whether to return the core request completion message to core #i based on the CmplSended flag. If CmplSended=0, the core request completion message is returned to core #i, otherwise it is not returned.
Citation Information
Patent Citations
Processing method and system for supporting software and hardware data consistency in multi-core DSP (Digital Signal Processing)
CN105718242A
Optimized caching agent with integrated directory cache
US20180189180A1