A processor and method for performing hierarchical cache system writeback and invalidation

By introducing cache retrieval and invalid instructions of the specified core layer into the processor, the problem of difficulty in effectively managing multi-core processor cache in the prior art is solved, and the precise retrieval and invalidation of the specified core cache is achieved, which improves the efficiency of the processor.

CN115098409BActive Publication Date: 2025-06-10VIA ALLIANCE SEMICON CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210644875.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-08
Publication Date
2025-06-10
Estimated Expiration
2042-06-08

AI Technical Summary

Technical Problem

In the prior art, when managing a hierarchical cache system in a multi-core processor, it is difficult to effectively reclaim and invalidate the cache content of the core, resulting in the execution speed of other cores being affected or redundant reclaim and invalid operation.

Method used

An instruction set architecture (ISA) instruction is proposed, allowing the processor to re-cache content with the specified core layer as its target and invalidate it. This instruction is converted into multiple micro-instructions through the decoder, using the memory sequence cache area and cache table to realize memory retrieval and invalid operations on the cache of the specified hierarchy.

Benefits of technology

It realizes accurate retrieval and invalidation of cached content in the specified core, avoids unnecessary retrieval of cached content in other cores, improves processor efficiency and reduces redundant operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115098409B_ABST
    Figure CN115098409B_ABST
Patent Text Reader

Abstract

The present invention provides a processor and a method for performing hierarchical cache system writeback and invalidation. The processor has a function of specifying a core hierarchy for performing hierarchical cache system writeback and invalidation. In response to an instruction of an instruction set architecture that targets the cache content of a specified hierarchy within the current core and performs writeback and invalidation on a hierarchical cache system, a decoder of a first core on the processor converts a plurality of microinstructions based on the microcode stored in a microcode memory. According to these microinstructions, a specified request is passed through a memory order buffer of the first core to the hierarchical cache system, enabling the hierarchical cache system to identify the cache lines involved in the specified hierarchy of the first core, write back the cache lines in a modified state to a system memory, and invalidate the identified cache lines completely from the hierarchical cache system regardless of whether they are in the modified state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the management technology of a hierarchical cache system of a processor. Background Art

[0002] In a computer system, the memory can be hierarchical. The upper-level memory has a higher speed, lower latency, but a smaller capacity. The memory hierarchy of most computer systems has the following four levels (sorted from upper to lower): registers; cache memory; system memory (main memory, such as DRAM); and disks (SSD or HD).

[0003] In particular, the cache memory can also be hierarchically designed, arranged from fast to slow in terms of access, including: the first-level cache (L1), the second-level cache (L2), and the third-level cache (L3, also known as the last-level cache memory, last level cache, abbreviated as LLC). Thus, the management of such a hierarchical cache system will significantly affect the system performance.

[0004] How to effectively manage a hierarchical cache system is an important issue in processor design. Summary of the Invention

[0005] This application proposes a management technology for a hierarchical cache system.

[0006] A processor implemented according to an embodiment of the present application includes a first core and a last-level cache memory coupled to the first core. The first core includes a microcode memory, a decoder, a memory ordering buffer (MOB), a first-level cache memory (L1), and a second-level cache memory (L2). The first-level cache memory, the second-level cache memory, the last-level cache memory, and the on-core cache memories of other cores sharing the last-level cache memory form a hierarchical cache system. In response to an instruction of an Instruction Set Architecture (ISA) that targets the cache content of a specified level in the current core and performs write-back and invalidation on the hierarchical cache system, the decoder converts into multiple microinstructions based on a microcode stored in the microcode memory. According to these microinstructions, a specified request is passed through the memory ordering buffer to the hierarchical cache system, causing the hierarchical cache system to identify the cache lines involved in the specified level of the first core, write back the cache lines in a modified state to a system memory, and invalidate the identified cache lines completely from the hierarchical cache system regardless of whether they are in the modified state.

[0007] In one embodiment, the instruction targets the cache content of the first-level cache memory of the current core and performs write-back and invalidation on the hierarchical cache system.

[0008] In one embodiment, the instruction targets the cache content of the first-level and second-level cache memories of the current core and performs write-back and invalidation on the hierarchical cache system.

[0009] The present application also provides a method for specifying a cache level within a core for hierarchical cache system writeback and invalidation, including: providing an instruction of an instruction set architecture to a first core of a processor, where the instruction targets the cache content at a specified level within the current core and performs writeback and invalidation on a hierarchical cache system, where the hierarchical cache system includes a first-level cache memory of the first core and a second-level cache memory, and includes an LLC cache memory of the processor, and also includes intra-core cache memories of other cores in the processor that share the LLC cache memory; causing a decoder of the first core to convert the instruction into multiple microinstructions based on a microcode stored in a microcode memory; and according to the microinstructions, passing a specified request through a memory order buffer of the first core to the hierarchical cache system, causing the hierarchical cache system to identify the cache lines involved in the specified level of the first core, write back the cache lines in a modified state to a system memory, and invalidate the identified cache lines completely from the hierarchical cache system regardless of whether they are in the modified state.

[0010] Specific embodiments are given below and, in conjunction with the accompanying drawings, the content of the present invention is described in detail. Description of the Drawings

[0011] Figure 1 According to an embodiment of the present application, a diagram of a multi-core processor 100 includes four cores core_1, core_2, core_3, and core_4;

[0012] Figure 2 It is a block diagram that illustrates a processor 200 and a core core_1 thereon according to an embodiment of the present invention;

[0013] Figure 3A It illustrates an embodiment of an intra-core cache table 222, where the columns represent the states of the memory addresses of each cache line in the first-level cache memory L1 of the core core_1;

[0014] Figure 3B It illustrates an embodiment of a snooping table 224, where the columns represent the states of the memory addresses of each cache line in the LLC / L3 cache memory and different cores core_1 to core_4;

[0015] Figure 4 It is a flowchart that illustrates the implementation of the instruction L1_WBINVD (or other instructions with the same function) by looking up the snooping table 224;

[0016] Figure 5is a flowchart illustrating the implementation of the instruction L1_WBINVD (or other instructions with the same function) by looking up the cache table 222 in the core cache;

[0017] Figure 6A and Figure 6B illustrates the effect of the instruction L1_WBINVD (or other instructions with the same function) of the present application;

[0018] Figure 7 is a flowchart illustrating the implementation of the instruction CORE_WBINVD (or other instructions with the same function) by looking up the snooping table 224; and

[0019] Figure 8 accompanied by Figure 6A illustrates the effect of the instruction CORE_WBINVD (or other instructions with the same function) of the present application Detailed implementation manners

[0020] The following description lists various embodiments of the present invention. The following description introduces the basic concepts of the present invention and is not intended to limit the content of the present invention. The actual scope of the invention should be defined by the claims.

[0021] The present application particularly discusses the writing back and invalidation of a hierarchical cache system. One traditional technique performs the writing back and invalidation of cache contents at the granularity of the last-level cache memory (LLC) shared by multiple cores; among all the cache lines involved in the last-level cache memory (LLC), the modified cache lines must be first written back to the system memory, and regardless of whether they are modified or not, all are invalidated from the hierarchical cache system (including the on-core cache memories of all cores and the last-level block fetch memory LLC shared by all cores). Another traditional technique performs the writing back and invalidation of cache contents at the granularity of a cache line; processing a single cache line at a time. If this cache line is modified, it must be written back to the system memory, and regardless of whether it is modified or not, this cache line is invalidated from the hierarchical cache system (including the on-core cache memories of all cores and the last-level block fetch memory LLC shared by all cores). However, the traditional processing granularity is not suitable for all applications.

[0022] For example, in a multi-core processor, when a cache vulnerability occurs, the software has a need to save back and invalidate the cache content of the on-core cache memory of the current core; in fact, this situation does not require saving back the cache content exclusive to other cores, but traditional technologies do not have corresponding solutions. If the cache content is saved back and invalidated at the granularity of the last-level cache memory (LLC) shared by multiple cores, the execution speed of other cores will be affected. If the cache content is saved back and invalidated at the granularity of a cache line, complex logical operations must be performed by the software to determine which cache lines to specify for processing; the software may make incorrect judgments, and redundant save-back and invalidation may still occur.

[0023] The solution provided by this application can target the cache content of the specified hierarchy of the current core, that is, target the cache content of the first-level cache memory L1 of the current core, or target the cache content of the first-level and second-level cache memories L1 and L2 of the current core, for saving back and invalidation. In this way, the performance of other cores will not be affected, and redundant cache line save-back and invalidation can be avoided.

[0024] Figure 1 According to an embodiment of this application, a diagram of a multi-core processor 100 is shown, including four cores core_1, core_2, core_3, and core_4. Core core_1 can simply specify to target the cache content of its own first-level cache memory L1 for saving back and invalidation; among them, the cache lines in the M (modified state) will be saved back to the system memory first, and regardless of whether the state is M, all cache lines involved in the first-level cache memory L1 of core core_1 are completely invalidated from the hierarchical cache system (including the first-level and second-level cache memories L1 and L2 of all cores core_1 to core_4, and the last-level cache memory LLC shared by all cores core_1 to core_4). In addition, core core_1 can specify to target the cache content of its own first-level and second-level cache memories L1 and L2 for saving back and invalidation, that is, save back and invalidate the entire on-core cache memory; among them, the cache lines in the M state will be saved back to the system memory first, and regardless of whether the state is M, all cache lines involved in the first-level and second-level cache memories L1 and L2 of core core_1 are completely invalidated from the hierarchical cache system (including the first-level and second-level cache memories L1 and L2 of all cores core_1 to core_4, and the last-level cache memory LLC shared by all cores core_1 to core_4). Other cores core_2 to core_4 can also have the above functions of core core_1.

[0025] The processor disclosed in the present invention can provide instructions of an Instruction Set Architecture (ISA) for the aforementioned functions (writing back and invalidating the cache content at the specified level of the current core). The instruction set architecture supported by the processor is not limited, and it can be the x86 architecture, the Advanced RISC Machine (ARM) architecture, or others.

[0026] In one embodiment, the present invention discloses a processor, in which an instruction of an Instruction Set Architecture (ISA) (hereinafter labeled as L1_WBINVD) is provided, so that the core executing this instruction L1_WBINVD targets the cache content of the first-level cache memory L1 it owns, and makes it write back and invalidate from the hierarchical cache system; among them, the cache line with the state of M (updated) will first write back to the system memory and then invalidate this cache line.

[0027] In another embodiment, the present invention discloses a processor, in which an instruction of an Instruction Set Architecture (ISA) (hereinafter labeled as CORE_WBINVD) is provided, so that the core executing this instruction CORE_WBINVD targets the cache content of the first-level and second-level cache memories L1 and L2 it owns, and makes it write back and invalidate from the hierarchical cache system; among them, the cache line with the state of M (updated) will first write back to the system memory and then invalidate this cache line.

[0028] In another embodiment, the present invention discloses a processor, in which an instruction of an Instruction Set Architecture (ISA) (hereinafter labeled as Li_WBINVD) is provided, and the functions of the aforementioned instruction L1_WBINVD or instruction CORE_WBINVD are implemented in a way of setting operands. If the operand is set to specify the first-level cache memory L1, then this instruction Li_WBINVD corresponds to the aforementioned instruction L1_WBINVD. If the operand is set to specify the cache memories L1 and L2 inside the core, then this instruction Li_WBINVD corresponds to the aforementioned instruction CORE_WBINVD. In program writing, other instructions can be used to fill registers / memories / immediate values before the instruction Li_WBINVD to complete the setting of this operand. There are also other processor implementations that provide an instruction of an Instruction Set Architecture (ISA), although it implements more complex functions, but includes the functions of the above instruction L1_WBINVD or instruction CORE_WBINVD; such instructions also fall within the scope of this application.

[0029] In one implementation, the present invention has a design in the microcode (ucode) of the processor corresponding to these instructions (such as, L1_WBINVD, CORE_WBINVD, Li_WBINVD, or others), and corresponding modifications can be made to the hardware of the processor.

[0030] Figure 2 is a block diagram that illustrates a processor 200 and a core core_1 thereon according to an embodiment of the present invention. The illustrated hierarchical cache system Cache_sys includes a first-level cache memory L1, a second-level cache memory L2, and a last-level cache memory (LLC) L3. The first-level and second-level cache memories L1 and L2 are on-core cache memories of the core core_1. In a multi-core processor, the last-level cache memory (LLC) L3 is shared by multiple cores (such as Figure 1 ), and the hierarchical cache system Cache_sys further includes on-core cache memories of other cores (such as Figure 1 the first-level and second-level cache memories L1 and L2 of other cores core_2...core_4).

[0031] After a piece of instruction is loaded from a system memory 202 to an instruction cache 204 via a bus Bus, it is delivered to a decoder 206 for decoding. The decoder 206 includes an instruction buffer (abbreviated as XIB) 208 and an instruction translator (abbreviated as XLATE) 210. The instruction buffer (XIB) 208 identifies and segments out the instructions proposed in this application (such as, L1_WBINVD, CORE_WBINVD, Li_WBINVD, or others), and the instruction translator (XLATE) 210 translates the instruction (such as, L1_WBINVD, CORE_WBINVD, Li_WBINVD, or others) into multiple microinstructions recognizable by pipeline hardware based on microcode (stored in a microcode memory). The core core_1 operation register renaming module (abbreviated as Rename) 212 processes these microinstructions and operates a reservation station (abbreviated as RS) 214 to issue the renamed microinstructions out of order to an execution unit (EU) 216. Through a memory ordering buffer (abbreviated as MOB) 218, it targets the cache content of the specified hierarchy of the core core_1 (simply specifying L1, or specifying the entire on-chip cache memory including L1 and L2), and writes back and invalidates it from the entire hierarchical cache system Cache_sys (including L3, the L1 and L2 of the core core_1, and the on-chip cache memories of other cores). The out-of-order executed microinstructions will wait in a re-order buffer (abbreviated as ROB) 220 for sequential retirement.

[0032] Based on the above hardware operations, the microinstructions decoded from the instruction L1_WBINVD target the cache content of the first-level cache memory L1 of the core core_1 through the memory ordering buffer (MOB) 218 for write-back and invalidation; among them, the cache lines with the state of M (updated) will be written back to the system memory 202 via the bus Bus, and all cache lines involved in the first-level cache memory L1 will be completely invalidated from the hierarchical cache system Cache_sys (including L3, the L1 and L2 of the core core_1, and the on-chip cache memories of other cores), regardless of whether the state is M (regardless of whether there is an update).

[0033] Based on the above hardware actions, the microinstructions decoded from the instruction CORE_WBINVD target the cache contents of the first-level and second-level cache memories L1 and L2 (i.e., the on-core cache memories of core_1) of core_1 through the memory order buffer (MOB) 218 for write-back and invalidation; among them, the cache lines with the state of M (updated) will be written back to the system memory 202 via the bus Bus, and regardless of whether the state is M (regardless of whether there is an update), all cache lines involved in the first-level and second-level cache memories L1 and L2 are completely invalidated from the hierarchical cache system Cache_sys. All cache lines completely from the hierarchical cache system Cache_sys include L3, L1 and L2 of core_1, and the on-core cache memories of other cores.

[0034] Based on the above hardware actions, the microinstructions decoded from the instruction Li_WBINVD target the cache contents of the specified hierarchy through the memory order buffer (MOB) 218 according to the specification of the operand for write-back and invalidation; among them, the cache lines with the state of M (updated) will be written back to the system memory 202 via the bus Bus, and regardless of whether the state is M (regardless of whether there is an update), all cache lines involved in the specified hierarchy are completely invalidated from the hierarchical cache system Cache_sys (including L3, L1 and L2 of core_1, and the on-core cache memories of other cores).

[0035] The write-back and invalidation targeting the cache contents of the specified hierarchy within the core in this application may include a look-up table design. The table can be maintained in the internal storage space of the core - as shown in the figure, the on-core cache table 222 planned in the hierarchical cache system Cache_sys specifically maintains the cache status of the on-core cache memories (L1 and L2) of core_1. The table can also be maintained in the storage space outside the core, such as the snooping table 224 shown in the figure, which maintains the cache status for the complete hierarchical cache system Cache_sys (including the on-core cache memories of all cores and the last-level cache memory shared by these cores).

[0036] Figure 3A An implementation of the on-core cache table 222 is illustrated. The columns represent the states of the memory addresses of each cache line in the first-level cache memory L1 of core_1. The M state represents the modified state, and I represents the invalid state. Additionally, the E state can represent the exclusive state, and S represents the shared state for multiple cores. In particular, other cores (such as Figure 1The cores (core_2...core_4) may also have an on-core cache table maintained for their own first-level cache memory L1.

[0037] Figure 3B An implementation of the snoop table 224 is illustrated. The columns represent the memory addresses of each cache line in the last-level cache memory LLC / L3 and the states of different cores core_1 to core_4. The M state represents the modified state, the I represents the invalid state, the E state represents the exclusive state, and the S represents the shared state among multiple cores. For example, the box 302 shows that in core_1, there is a cache line with the memory address 0x800F00 that is shared among multiple cores, and a cache line with the memory address 0x801000 in the cache with the M state.

[0038] The following details how the microinstructions decoded from the instruction L1_WBINVD (or CORE_WBINVD, or Li_WBINVD, or other instructions with the same function according to the present application) operate the hardware and implement the instruction function using the table lookup technique.

[0039] First, discuss the instruction L1_WBINVD (or other instructions with the same function). It targets only the cache content of the first-level cache memory L1 of the current core for write-back and invalidation.

[0040] Figure 4 The figure is a flowchart illustrating the implementation of the instruction L1_WBINVD (or other instructions with the same function) using the snoop table 224; the microinstructions decoded from this instruction include a specified request L1_WBINVD_req. The following discussion refers to Figure 2 For illustration, the core that initiates the instruction is Figure 2 core_1.

[0041] In step S402, the memory order buffer (MOB) 218 receives the specified request L1_WBINVD_req and passes the specified request L1_WBINVD_req to the first-level cache memory L1 of core_1.

[0042] In step S404, in response to the specified request L1_WBINVD_req, the first-level cache memory L1 of core_1 returns the memory address to the memory order buffer (MOB) 218; the returned memory address corresponds to the cache line cached in the first-level cache memory L1 of core_1.

[0043] Step S406: The Memory Order Buffer (MOB) 218 pairs each memory address returned in step S304 with a write-back and invalidate request WB_req and passes it to the snoop table 224 for table lookup.

[0044] The snoop table 224 (such as Figure 3B ) records the cache status of the entire hierarchical cache system Cache_sys. Therefore, after each memory address returned in step S404 is paired with a write-back and invalidate request WB_req and queries the snoop table 224 in step S406, the cache status of the queried memory address in the hierarchical cache system Cache_sys can be obtained. Step S408 then pairs each memory address with a snoop request snoop_req and passes it to the hierarchical cache system Cache_sys according to the lookup result of the snoop table 224.

[0045] Step S410: The hierarchical cache system Cache_sys loads the cache line in the cache being snooped that is in the modified state (M state) onto a bus Bus, and invalidates the entire cache line being snooped from the hierarchical cache system Cache_sys (including the on-core cache memories of all cores and the last-level cache memory L3 shared by all cores), regardless of whether the cache line is in the modified state.

[0046] Step S412: Write the cache line loaded onto the bus Bus in step S410 to the system memory 202.

[0047] After sorting out, Figure 4 first identify the cache lines in the first-level cache memory L1 of core core_1, and then query the snoop table 224 that records the cache status of the entire hierarchical cache system Cache_sys one by one, so that the snoop request snoop_req sent to the hierarchical cache system Cache_sys targets the cache content of the first-level cache memory L1 of core core_1, realizing write-back and invalidate targeting the cache content of the first-level cache memory L1 of core core_1. In particular, the above steps may vary slightly in the order for performance considerations.

[0048] Figure 5 is a flowchart illustrating the implementation of the instruction L1_WBINVD (or other instructions with the same function) through table lookup in the on-core cache table 222. Similarly, Figure 4 , the microinstructions decoded from this instruction include a specified request L1_WBINVD_req, and the following discussion is also based on Figure 2 for explanation. The core that initiates the instruction is Figure 2 core core_1.

[0049] Step S502, the memory order buffer (MOB) 218 receives the specified request L1_WBINVD_req and sends the specified request L1_WBINVD_req to the hierarchical cache system Cache_sys.

[0050] Step S504, in response to the specified request L1_WBINVD_req, the hierarchical cache system Cache_sys looks up the on-chip cache table 222. The on-chip cache table 222 (such as Figure 3A ) records the cache status of the first-level cache memory L1 of the core core_1. Step S504 identifies the cache lines in the first-level cache memory L1 of the core core_1 that are in the modified state (M state) by looking up the on-chip cache table 222, and loads them onto the bus Bus.

[0051] Step S506 further identifies all the cache lines cached in the first-level cache memory L1 of the core core_1 according to the lookup result of the on-chip cache table 222, and invalidates all the cache lines from the hierarchical cache system Cache_sys (including the on-chip cache memories of all cores and the last-level cache memory L3 shared by all cores).

[0052] Step S508, writes the cache lines loaded onto the bus Bus in step S504 to the system memory 202.

[0053] After sorting out, Figure 5 it is known from the record made by the on-chip cache table 222 specifically for the cache status of the first-level cache memory L1 of the core core_1 about the cache content of the first-level cache memory L1 of the core core_1, and the write-back and invalidation targeting the cache content of the first-level cache memory L1 of the core core_1 is realized.

[0054] Figure 6A And Figure 6B Illustrate the effect of the instruction L1_WBINVD (or other instructions with the same function) of the present application by way of example. Figure 6A Mark the modified (M state) cache content with a slash. After the core core_1 with the first-level and second-level cache memories L1 and L2 executes the instruction L1_WBINVD (or other instructions with the same function) of the present application, such as Figure 6BAs shown, the modified cache lines of the first-level cache memory L1 are written back to the system memory 202, and all the original cache lines of the first-level cache memory L1 are completely invalidated from the entire hierarchical cache system Cache_sys (including L1, L2, L3, and other cores not shown). In the figure, the third-level cache memory L3 is in an inclusive form of the upper-level cache content, while the second-level cache memory L2 is in a non-inclusive non-exclusive (NINE) form of the upper-level cache content. However, the instruction L1_WBINVD (or other instructions with the same function) of the present application is not limited to such a hierarchical cache system architecture and can be applied to various hierarchical architectures.

[0055] Next, discuss the instruction CORE_WBINVD (or other instructions with the same function), which targets the cache content of the first-level and second-level cache memories L1 and L2 (the complete on-core cache memory) of the current core for write-back and invalidation.

[0056] Figure 7 The figure is a flowchart illustrating the implementation of the instruction CORE_WBINVD (or other instructions with the same function) by looking up the snoop table 224; the microinstructions decoded from this instruction include a specified request CORE_WBINVD_req. The following discussion is based on Figure 2 For illustration, the core that issues the instruction is Figure 2 core_1.

[0057] In step S702, the memory order buffer (MOB) 218 receives the specified request CORE_WBINVD_req and sends the specified request CORE_WBINVD_req to the last-level cache memory L3.

[0058] In step S704, in response to the source of the specified request CORE_WBINVD_req (core core_1), the last-level cache memory L3 determines that the write-back and invalidation target is the cache content of all the on-core cache memories (L1 and L2) of the core core_1 and looks up the snoop table 224. For example, for the core core_1 that issues the instruction, query Figure 3B field 302 of the snoop table 224 to know which cache lines are cached in the first-level and second-level cache memories L1 and L2 of the core core_1 and which of these cache lines are modified (status M). Based on field 302 of the table 224 in Figure 3, step S704 also includes loading the cache lines corresponding to the memory addresses in the modified state (M state) in the core core_1 onto a bus Bus.

[0059] Step S706: The last-level cache memory L3 issues a snoop request snoop_req to invalidate the cache lines in the hierarchical cache system Cache_sys that belong to the write-back and invalidation targets determined in step S704 from the hierarchical cache system Cache_sys (including the on-core cache memories of all cores and the last-level cache memory L3 shared by all cores).

[0060] Step S708: Write the cache line loaded onto the bus Bus in step S706 to the system memory 202.

[0061] After sorting,[[]] Figure 7 By using the records made in the snoop table 224 specifically for the on-core cache memories of each core core_1~core_4, the cache content of the on-core cache memory (L1 and L2) of the initiating core core_1 is obtained, and write-back and invalidation targeting the cache content of the on-core cache memory of core_1 is realized.

[0062] Supplemented by Figure 6A , Figure 8 The effects of the instruction CORE_WBINVD (or other instructions with the same function) of the present application are illustrated by examples. As Figure 8 shown, the modified cache lines of the first-level and second-level cache memories L1 and L2 are all written back to the system memory 202, and all the original cache lines of the first-level and second-level cache memories L1 and L2 are completely invalidated from the entire hierarchical cache system Cache_sys (including L1, L2, L3, and other cores not shown).

[0063] Any multi-core processor that uses ISA instructions, combined with hardware and microcode design, to specify the on-core hierarchy (simply specify the first-level cache memory L1, or specify the first-level and second-level cache memories L1 and L1) for write-back and invalidation of the hierarchical cache system belongs to the technical scope of the present application.

[0064] Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Those skilled in the art can make some modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention is defined by the claims.

Claims

1. A processor for performing cache system writeback and invalidation at a specified hierarchy within its core, the processor comprises: A first core, including a decoder, a memory order buffer, and a first-level cache memory; and A last-level cache memory, coupled to the first core; wherein: According to an instruction of an instruction set architecture that targets the cache content of the specified hierarchy of the first core and performs writeback and invalidation on the hierarchical cache system, the decoder converts into multiple micro-instructions, wherein the hierarchical cache system includes the first-level cache memory, the last-level cache memory, and the on-core cache memories of other cores sharing the last-level cache memory; According to the micro-instructions, a specified request is passed through the memory order buffer to the hierarchical cache system, causing the hierarchical cache system to identify the cache lines involved in the specified hierarchy of the first core, write back the cache lines in the modified state to the system memory, and invalidate the identified cache lines from the hierarchical cache system.

2. The processor according to claim 1, wherein: The instruction specifies to target the cache content of the first-level cache memory of the first core and perform writeback and invalidation on the hierarchical cache system.

3. The processor according to claim 2, wherein: The specified request passed through the memory order buffer to the hierarchical cache system drives the first-level cache memory of the first core to return a memory address to the memory order buffer, and the returned memory address corresponds to the cache line cached by the first-level cache memory of the first core.

4. The processor according to claim 3, further comprises: A snooping table, located in the storage space outside the first core, listing the states of each memory address in different cores, and the states include the modified state, the exclusive state, the shared state, and the invalid state.

5. The processor according to claim 4, wherein: The memory order buffer pairs each memory address returned by the first-level cache memory of the first core with a writeback and invalidation request and passes it to the snooping table for look-up; Each memory address is paired with a snooping request and passed to the hierarchical cache system according to the look-up result of the snooping table; The hierarchical cache system loads the cache lines in the modified state among the snooped cache lines onto the bus, invalidates the snooped cache lines from the hierarchical cache system, wherein the bus is configured for communication between the system memory and the processor; and The bus writes the loaded cache lines into the system memory.

6. The processor according to claim 2, further comprises: An on-core cache table, located in the storage space of the first core and belonging to the hierarchical cache system, listing the states of each memory address in the first-level cache memory of the first core, and the states include the modified state, the exclusive state, the shared state, and the invalid state.

7. The processor according to claim 6, wherein: In accordance with the specified request, the hierarchical cache system looks up the in-core cache table, identifies the cache lines in the first-level cache memory of the first core that are in the modified state, and loads the cache lines in the modified state onto the bus, where the bus is configured for communication between the system memory and the processor; The hierarchical cache system also looks up the in-core cache table to identify all the cache lines cached in the first-level cache memory of the first core, and invalidates all the identified cache lines in the first-level cache memory from the hierarchical cache system; and The bus writes the loaded cache lines to the system memory.

8. The processor according to claim 1, wherein the processor further includes a second-level cache memory, the hierarchical cache system further includes the second-level cache memory, and the instruction specifies to target the cache contents of the first-level and second-level cache memories of the first core and perform write-back and invalidation on the hierarchical cache system.

9. The processor according to claim 1, further including a microcode memory, and the decoder converts the instruction into multiple microinstructions according to the microcode stored in the microcode memory.

10. The processor according to claim 2, further including: A snooping table, located in the storage space outside the first core, listing the states of multiple memory addresses representing different cache lines in different cores, where the states include the modified state, the exclusive state, the shared state, and the invalid state.

11. The processor according to claim 10, wherein: The specified request is passed to the last-level cache memory through the memory order buffer; In response to the first core as the source of the specified request, the last-level cache memory determines that the write-back and invalidation target is the cache contents of all the in-core cache memories of the first core, looks up the snooping table, and loads the cache lines corresponding to the memory addresses in the modified state in the first core onto the bus, where the bus is configured for communication between the system memory and the processor; The last-level cache memory issues a snooping request to invalidate the cache contents of all the in-core cache memories of the first core, which is the write-back and invalidation target, from the hierarchical cache system; and The bus writes the loaded cache lines to the system memory.

12. A method for performing write-back and invalidation on a hierarchical cache system at a specified hierarchy within a core of a processor, the method includes: Providing an instruction of the instruction set architecture to the first core of the processor, the instruction targeting the cache contents of the specified hierarchy of the first core and performing write-back and invalidation on the hierarchical cache system, where the decoder converts the instruction into multiple microinstructions, and the hierarchical cache system includes a first-level cache memory, a last-level cache memory, and the in-core cache memories of other cores sharing the last-level cache memory; and According to the micro-instruction, pass the specified request through the memory order buffer of the first core to the hierarchical cache system, so that the hierarchical cache system identifies the cache lines involved in the specified level of the first core, write back the cache lines in the modified state to the system memory, and invalidate the identified cache lines from the hierarchical cache system.

13. The method according to claim 12, wherein: The instruction specifies to target the cache content of the first-level cache memory of the first core and perform write-back and invalidation on the hierarchical cache system.

14. The method according to claim 13, wherein: The specified request passed to the hierarchical cache system through the memory order buffer drives the first-level cache memory of the first core to return the memory address to the memory order buffer, and the returned memory address corresponds to the cache line cached by the first-level cache memory of the first core.

15. The method according to claim 14, further comprising: Let the snooping table be carried in the storage space outside the first core, listing the states of each memory address in different cores, and the states include the modified state, exclusive state, shared state, and invalid state.

16. The method according to claim 15, wherein: The memory order buffer combines each memory address returned by the first-level cache memory of the first core with the write-back and invalidation requests and passes them to the snooping table for look-up; Each of the memory addresses is combined with a snooping request and passed to the hierarchical cache system according to the look-up result of the snooping table; In the hierarchical cache system, for the cache lines being snooped, the cache lines in the modified state are loaded onto the bus, and the cache lines being snooped are invalidated from the hierarchical cache system, wherein the bus is configured for communication between the system memory and the processor; and The bus writes the loaded cache lines to the system memory.

17. The method according to claim 13, further comprising: Let the in-core cache table be carried in the storage space of the first core and belong to the hierarchical cache system, listing the states of each memory address in the first-level cache memory of the first core, and the states include the modified state, exclusive state, shared state, and invalid state.

18. The method according to claim 17, wherein: According to the specified request, the hierarchical cache system looks up the in-core cache table to identify the cache lines in the modified state in the first-level cache memory of the first core, and loads the cache lines in the modified state onto the bus, wherein the bus is configured for communication between the system memory and the processor; The hierarchical cache system also looks up the in-core cache table to identify all the cache lines cached by the first-level cache memory of the first core, and completely invalidates the identified all cache lines of the first-level cache memory from the hierarchical cache system; and The bus writes the loaded cache lines to the system memory.

19. The method according to claim 12, wherein the processor further includes a second-level cache memory, the hierarchical cache system further includes the second-level cache memory, and the instruction specifies to target the cache content of the first-level cache memory of the first core and the cache content of the second-level cache memory, and perform write-back and invalidation on the hierarchical cache system.

20. The method according to claim 12, further comprising causing the decoder of the first core to convert the instruction into a plurality of micro-instructions based on the microcode stored in the microcode memory.

21. The method according to claim 12, wherein: the instruction specifies to target the cache content of the first-level cache memory of the first core and perform write-back and invalidation on the hierarchical cache system.

22. The method according to claim 21, further comprising: causing the snooping table to be loaded in a storage space outside the first core, listing the states of multiple memory addresses representing different cache lines in different cores, and the states include the modified state, the exclusive state, the shared state, and the invalid state.

23. The method according to claim 22, wherein: the specified request is passed to the last-level cache memory through the memory order buffer; in response to the first core where the specified request originates, the last-level cache memory determines that the write-back and invalidation target is the cache content of all the on-core cache memories of the first core, looks up the snooping table, and loads the cache line corresponding to the memory address in the modified state in the first core onto the bus, wherein the bus is configured for communication between the system memory and the processor; the last-level cache memory also issues a snooping request to invalidate the cache content of all the on-core cache memories of the first core that is the determined write-back and invalidation target from the hierarchical cache system; and the bus writes the loaded cache line into the system memory.

Citation Information

Patent Citations

  • Computer system and method for specifying key for cache write-back and invalidation

    CN114064517A