Managing multiple invalidations of a cache memory

By partitioning the local cache into multiple banks and using a directory-based cache coherence protocol to consolidate multiple invalidation commands, the method addresses latency issues in multi-core processor systems, enabling efficient processing of simultaneous invalidation commands and reducing system latency.

FR3157585A1Active Publication Date: 2025-06-27KALRAY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
FR2023015090
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-27
Estimated Expiration
2043-12-22

AI Technical Summary

Technical Problem

In multi-core processor systems with shared memory, processing multiple simultaneous cache line invalidation commands can lead to latency issues, as each memory bank may send invalidation commands simultaneously, overwhelming the cache controller.

Method used

The proposed method involves partitioning the local cache into multiple banks, each capable of processing a respective invalidation command, and using a directory-based cache coherence protocol to manage multiple invalidation commands by consolidating them into a single command that can be processed simultaneously by the cache.

Benefits of technology

This approach allows the cache controller to process all incoming invalidation commands within a single clock cycle, reducing latency and improving system performance by ensuring that memory banks can acknowledge write accesses without delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a processor comprising multiple cores (10) each comprising a local cache (L1); multiple memory banks (12) forming a memory shared by the multiple cores; a directory-based cache coherence protocol manager (DIR), comprising for each core a circuit (16) for consolidating multiple cache line invalidation commands (INVAL) received from the different memory banks. The consolidation circuit comprises a counter of received invalidation commands (PARCNT); a selection circuit (32) configured to, depending on whether the count (N) of received invalidation commands is equal to 1 or greater, transmit to the cache the single invalidation command received or an invalidation command consolidating the cache lines identified in the received invalidation commands, in a format usable by the cache to simultaneously invalidate the identified cache lines. Figure for abstract: Fig. 3
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Management of multiple invalidations of a cache memory Technical field

[0001] The invention relates to the management of cache coherence in a multi-core processor system whose cores can simultaneously access different banks forming a shared memory. Background

[0002] There are several cache coherence protocols designed to ensure, in a multi-core system, that the local cache of each processor core reflects the data updated in shared memory by the other cores.

[0003] A commonly used operation in coherence protocols is cache line invalidation. At any given time, there may be duplication of the same memory block, or cache line, in different local caches. If a core writes to memory at an address corresponding to this cache line, the copies in the other cores become obsolete.

[0004] To account for this, when a core writes to memory, the cache coherence protocol sends an invalidation command to the other cores that own the corresponding cache line. This invalidation command tells the local cache controllers to invalidate their local copies. A core accessing a cache line thus invalidated in its local cache will have to retrieve the updated line from memory again.

[0005] To manage the sending of invalidation commands, the coherence protocol can be based on a directory in which the memory controller records the cache lines used and the cores that have them in their caches.

[0006] When the shared memory is structured into several banks simultaneously accessible by all the cores, each bank is likely to send a simultaneous invalidation command to the cores concerned. Thus, a given core can receive several simultaneous invalidation commands. The processing of these multiple commands by the cache controller can cause a latency of several cycles during which the memory banks having sent these invalidation commands cannot normally acknowledge the write accesses having triggered these sendings.

[0007] To avoid such latency, techniques have been proposed that allow the cache controller to process multiple invalidation commands simultaneously. An example of such a technique is described in US Pat. No. 6,701,417, which uses a protocol based on directories and write-through caches. More speci Specifically, this patent proposes to partition the local cache into several banks, each of which is designed to process a respective invalidation command. Thus, the cache controller can simultaneously process at most as many invalidation commands as there are banks. Summary

[0008] A method for managing cache coherence in a multi-core processor system is generally provided, where each core has a respective cache and has access to multiple banks of memory shared between the cores, the method comprising the steps of managing a directory in each memory bank for implementing directory-based cache coherence; accessing a current memory address for writing by a core; searching the directory managing the current address for cores that have a cache line corresponding to the current address; sending to the cores identified by the directory respective cache line invalidation commands, the commands including the memory address of the cache line; and for each core, serving multiple invalidation commands received from different memory banks.The multiple invalidation command service includes the steps of counting the number of invalidation commands received since a last clock cycle; when the count of received invalidation commands is one, transmitting the received invalidation command to the cache; and when the count of received invalidation commands is greater than one, transmitting to the cache a single command consolidating the cache lines identified in the received invalidation commands, in a format usable by the cache to simultaneously invalidate the identified cache lines.

[0009] Each core may have a respective multi-way set-associative cache, the method then comprising the steps of recording in the directories the ways in which the caches store the cache lines; transmitting the ways in the invalidation commands sent to the cores; and including in the consolidated invalidation command, for each invalidation command received, a pair of coordinates including a set index, extracted from the memory address, and the way.

[0010] The method may comprise the step of responding by the cache to the consolidated invalidation command by simultaneously invalidating each cache line located at the intersection of the set and the path determined by a respective pair of coordinates.

[0011] The service of the multiple invalidation commands may comprise the steps of forming a bit mask comprising a bit at 1 at the positions identified by set indices extracted from the memory addresses of the invalidation commands received; when the count of invalidation commands received is greater than a threshold, include the bitmask in the consolidated invalidation command; and have the cache respond to the consolidated command by simultaneously invalidating the cache sets marked in the bitmask.

[0012] The directory may record the cores and ways for each cache line as a combined bitmask marking the cores having the line in their cache and the ways in which the line is present in those cores' caches, the way information transmitted in the invalidation commands then including the bits of the bitmask identifying the ways.

[0013] The consolidated invalidation commands may be configured to carry a field comprising a fixed number of bits, a first part of which identifies a type of command among an original invalidation command, a consolidated command with coordinates, and a consolidated command with bit mask, and a second part defines for the respective types: the address of the cache line, the coordinate pairs, and the bit mask.

[0014] Also provided is a processor comprising multiple cores each comprising a local cache; multiple memory banks forming a memory shared by the multiple cores; a directory-based cache coherence protocol manager, comprising for each core a circuit for consolidating multiple cache line invalidation commands received from the different memory banks. The consolidation circuit comprises a counter of received invalidation commands; a selection circuit configured to, depending on whether the count of received invalidation commands is equal to 1 or greater, transmit to the cache the single received invalidation command or an invalidation command consolidating the cache lines identified in the received invalidation commands, in a format usable by the cache to simultaneously invalidate the identified cache lines.

[0015] Each core may have a respective multi-way set-associative cache, the processor further comprising a directory associated with each bank, configured to record with each cache line, the ways in which the cache line is present in the different cores and to include the ways in the invalidation commands sent to the cores; and the consolidation circuit configured to include in the consolidated invalidation command, for each cache line invalidation command received, a pair of coordinates including a set index, extracted from a memory address of the cache line, and the way.

[0016] The consolidation circuit may be configured to include in the consolidated invalidation command, when the count of received invalidation commands is greater than a threshold, a bit mask marking cache sets to be invalidated, where each set includes the cache line identified by a respective received invalidation command.

[0017] Alternatively, in a multi-core processor system, where each core has a respective multi-way set-associative cache, and has access to multiple banks of a memory shared between the cores, a method for managing cache coherence comprises the steps of managing a directory in each memory bank for implementing directory-based cache coherence; accessing a current memory address for writing by a core; searching the directory managing the current address for cores that have a cache line corresponding to the current address; sending to the cores identified by the directory respective cache line invalidation commands, the commands including the memory address of the cache line; for each core, serving multiple invalidation commands received from different memory banks;registering in directories the ways in which the caches store the cache lines; transmitting the ways in invalidation commands sent to the cores; serving the multiple invalidation commands received by a core by transmitting to the core's cache a single consolidated command including, for each invalidation command received, a pair of coordinates including a set index, extracted from the memory address of the cache line, and the way; and responding by a cache to a consolidated invalidation command by simultaneously invalidating each cache line located at the intersection of the set and the way determined by a respective pair of coordinates.

[0018] The multiple invalidation command service may include the steps of counting the number of invalidation commands received since a last clock cycle; forming a bit mask including a 1 bit at the positions identified by the set indices; when the count of received invalidation commands is greater than a threshold, including the bit mask in the consolidated invalidation command in place of the coordinate pairs; and responding by the cache to the consolidated command by simultaneously invalidating the cache sets marked in the bit mask. Summary description of the drawings

[0019] Embodiments will be set out in the following description, given without limitation in relation to the attached figures among which:

[0020] [Fig. 1] partially depicts one embodiment of a multi-core processor architecture configured to implement a directory-based cache coherence protocol;

[0021] [Fig.2A] shows three types of consolidated invalidation commands, according to one embodiment, allowing a cache controller to simultaneously invalidate multiple cache lines;

[0022] [Fig.2B] represents an example of recording in a directory used for a cache coherence protocol; and

[0023] [Fig.3] is a block diagram of a logic circuit configured to generate the consolidated commands of [Fig.2A]. Detailed description

[0024] The cache coherence protocol described in patent US6701417 requires a particular cache structure and the number of invalidation commands that can be processed simultaneously is limited to the number of cache banks, in practice four.

[0025] A cache coherence protocol is described below with which the cache controller can simultaneously process all invalidation commands that may arrive at it, at the cost of a compromise in performance based on a threshold number of commands processed simultaneously. In the worst case, the compromise is, for a multi-way set-associative cache, the coarse invalidation of an entire set of cache lines, instead of invalidating a single cache line.

[0026] The concept of simultaneous processing of several invalidation commands relates to a clock cycle of the cores. In practice, the processing is carried out asynchronously by combinational logic circuits which produce the desired operations in a time that is more or less long depending on the case, but which remains less than one clock period.

[0027] Furthermore, the protocol is adaptive by operating a selection between several modes depending on the number of invalidation commands to be processed. In the first mode, selected for a single command to be processed, a conventional invalidation command can be used, precise, making it possible to identify the cache line concerned and to invalidate it only if it is still present in the cache, namely listed in the cache tag memory.

[0028] In a coarse mode, selectable for any number of commands to be processed, a consolidated command is produced which simultaneously invalidates the sets containing the affected cache lines.

[0029] A less brutal intermediate mode is preferably used up to a certain threshold of number of commands to be processed, producing a consolidated invalidation command which identifies each cache line by a pair of coordinates (set containing the line, path containing the line). The cache controller then simultaneously invalidates all the lines identified by these pairs of coordinates. The presence of these lines in the cache is then not verified, which constitutes a small price to pay for the possibility of simultaneously processing these several invalidation commands. In practice, the invalidation of a line absent from the cache causes the invalidation of a memory location intended to contain a cache line. This location may be empty, in which case it is already invalid and the operation has no effect. The location may also contain a new cache line, in which case that new line is invalidated "for nothing", but this simply incurs subsequent latency to reload the line when the core attempts to access it.

[0030] The number of invalidations that can be processed in this mode depends on the size of a parameter of the consolidated command, which size conditions the number of pairs that can be transmitted.

[0031] The effectiveness of this protocol, in terms of granularity in the selection of cache lines to be invalidated, is related to the frequency of cases where bursts of multiple simultaneous invalidation commands occur, and to the number of commands to be processed in each burst. It turns out in practice that cases where a single invalidation command is to be processed are more frequent than bursts of several commands. In addition, bursts with a small number of commands (for example between 2 and 4) are more frequent than those with a larger number of commands.

[0032] [Fig. 1] partially represents a multi-core processor architecture usable for implementing such a cache coherence protocol on the basis of directories. Two processor cores 10 have simultaneous read and write access RW to several banks 12 (of which only three are represented) of a shared memory through a matrix switch ("crossbar") managed by arbiters ARBITER. Each core includes a local level 1 cache LL

[0033] Cache coherence is managed by a DIR directory which may be centralized or distributed, as shown, in each memory bank 12. The directories are configured to record the cache lines currently in use in each core. Each DIR directory associated with a bank lists only the cache lines corresponding to the physical addresses assigned to the bank. The directories communicate with the cores through respective so-called "master-sides" 16, which may be part of the L1 caches or the fabric switch. The master-sides are connected to the DIR directories by point-to-point signaling links, shown in dotted lines.

[0034] Typically, a directory receives a new record each time a core misses a read in its cache and retrieves the cache line from the corresponding bank. The new record identifies the cache line and the core. A record is kept in the directory until that cache line is invalidated. Thus, the directory information may be out of date in the meantime.

[0035] Directory records are generally stored in associative memory indexed by the physical addresses of the cache lines.

[0036] When a given core accesses a memory line subject to coherence in writing, this event is signaled to the directories. The directory that is responsible for managing the coherence of this address then issues an invalidation INVAL to the other cores registered as owners of this cache line.

[0037] In order to manage multiple invalidations simultaneously according to the adaptive mode previously mentioned, each master side comprises a consolidation circuit CONSLD receiving the individual invalidation commands coming from all the directories. From the individual invalidation commands, this circuit produces a consolidated invalidation command C-INVAL to the corresponding L1 cache memory.

[0038] As previously indicated, the consolidated command can be of two, or preferably three types: PRECISE, ARRAY_OF_SLOTS, and MASK_OF_SETS.

[0039] [Fig.2A] illustrates an example of formats of the three types of consolidated commands, assuming for example that the physical memory addresses are 40 bits and the cache is an associative cache with 64 sets of 8 ways whose lines comprise 64 bytes (i.e. a cache of 32 KB).

[0040] Each consolidated command consists of two bits of type identifier IType[65:64] followed by a 64-bit field IData[63:0] whose function depends on the type. A binary identifier 00 indicates that there are no invalidations to process.

[0041] The binary identifier 01 corresponds to the PRECISE type. This consolidated command is produced when a single invalidation command is pending since the last clock cycle. It provides in its IData field what the single incoming invalidation command normally provides, namely the physical address @PHY[39:0] used for memory access. In principle, a certain number of low-order bits of the physical address, called the offset and defining the position of the data in a cache line, are not useful, so that here we can remove 6 low-order bits and keep only @PHY[39:6]. In the truncated address, the 6 low-order bits [11:6] constitute what is called the index, and are used to identify the set containing the cache line.The remaining high-order bits [39:12] are the so-called tag, which is used to query the cache's tag memory, which identifies, among other things, the path containing the cache line. If the tag memory does not provide a corresponding value, the cache line was evicted after the repository issued the invalidation command.

[0042] A consolidated command of type PRECISE is treated by the cache controller as a classic invalidation command, by invalidating the unequivocally identified cache line.

[0043] The binary identifier 10 corresponds to the type ARRAY_OF_SLOTS. This consolidated command is produced when two or more invalidation commands are received simultaneously in the same clock cycle. The number of commands that can be processed simultaneously is limited by the size of the IData field, here 64 bits. This IData field In this example, there are four 14-bit data items used to identify four cache lines to be invalidated. The 56 bits corresponding to these data items can be right-aligned in the IData field. Each 14-bit data item corresponds to a pair of coordinates, namely in this example a 6-bit index IDX[5:0] and an 8-bit channel W[7:0]. The cache line to be invalidated is therefore the one located at the intersection of the set identified by the IDX index and the channel identified by the W field.

[0044] For reasons discussed later, the W field is in this example an 8-bit mask, identifying one to eight channels by the position of a bit at 1. Alternatively, the channel number can be coded on 3 bits, in which case the IData field could contain 7 pairs of coordinates.

[0045] The IDX indexes in this example correspond to the bits [11:6] of the respective physical addresses provided by the incoming individual invalidation commands. If fewer than four invalidation commands are to be processed, the W masks associated with the missing commands can be set to 0, indicating that they will be ignored.

[0046] As for the values ​​of the W fields identifying the ways, the DIR directories are configured to further record the ways in which the lines are stored in the caches, and also broadcast these ways in the invalidation commands. These ways broadcast by the invalidation commands are then included in the W fields of the consolidated invalidation command of type ARRAY_OF_SLOTS.

[0047] At the DIR level, ideally each cache line is recorded with the cores that own it and, for each core, the path in which it is located. This can represent a significant number of bits to manipulate for each line in a system with a large number of cores. For example, for 16 cores with 8-way caches, we would need 16 3-bit fields, or 48 bits, to encode one path out of 8 for each of the 16 cores.

[0048] [Fig.2B] illustrates an embodiment of directory registration that reduces complexity and hardware cost, with the trade-off of reducing the accuracy of path identification. In each directory, a core mask CORE[15:0] having one bit for each core and a path mask W[7:0] having one bit for each path are maintained for each cache line registered @PHY[39:6]. This reduces the number of bits to 24 for a 16-core system with 8-way caches. It is then the path mask W of this registration that is transmitted in the corresponding invalidation command.

[0049] With such way masks, if multiple cores use the same cache line but that line is stored in a different way in each core, the way mask for that cache line would have multiple bits set to 1.

[0050] The cache controller is configured to process such a consolidated command in reading each pair of coordinates (IDX, W) and, for each pair, invalidating the cache lines located at the intersection of the set identified by IDX and the marked ways in W. If the way mask W contains more than one 1 bit, additional cache lines, if they exist in the cache, will be unnecessarily invalidated, but this is the price to pay for simplifying the directory structure. In practice, situations where more than one way is marked in a way mask are rare.

[0051] An invalidation of a cache line by the cache controller can be carried out conventionally by toggling a validity flag in a matrix of flip-flops representing the intersections of the sets and the cache paths. Each of these flip-flops is individually accessible by the logic circuits of the controller, so that these circuits can be configured to simultaneously toggle any number of flags. Once a flag is thus changed to the "invalid" state, a subsequent read access by the core to this location fails after checking the flag ("cache miss") and is diverted to the shared memory to reload an updated cache line.

[0052] The binary identifier 11 corresponds to the MASK_OF_SETS type. This consolidated command is produced when the number of invalidations to be processed is greater than the number of commands that can be processed by a consolidated command of the ARRAY_OF_SLOTS type, namely 4 in this example. The IData field is then a mask of sets S-MASK[63:0] where each bit at 1 indicates that the set corresponding to the position of the bit in the mask must be invalidated. Thus, the size of the IData field is at least equal to the number of sets in the cache, 64 in this example.

[0053] To generate the S-MASK mask, the index fields of the physical addresses of the pending invalidation commands are used, namely the @PHY[11:6] values. The @PHY[11:6] values ​​determine the bit positions to be set to 1 in the S-MASK mask. This S-MASK mask can then be used by the cache controller to simultaneously invalidate all the sets marked in this mask.

[0054] [Fig. 3] is a block diagram of a logic circuit configured to generate the consolidated commands, including the contents of the IData field of each type of consolidated command. Such a circuit can be integrated into each master side, at the CONSLD block of [Fig. 1]. The circuit operates on all point-to-point connections between the master side and the memory bank directories.

[0055] From each of the memory banks, 16 in this example, the circuit receives a bundle of INVAL_i conductors used to transmit an invalidation command. One conductor carries an EN_i flag indicating the presence of an invalidation command, the following conductors carry the address of the cache line to be invalidated @PHY[39:6], and the last conductors carry the channel mask W[7:0]. The flags EN_i are supplied to individual inputs of a parallel counter PARCNT which provides at any time the number of active commands N. The concatenation of these flags forms a 16-bit active command mask EN[15:0]. The count N controls two four-way multiplexers 30, 32 which produce respectively the IType and IData parameters of the consolidated invalidation command C-INVAL.

[0056] When N = 0, multiplexer 30 selects the binary value 00 for IType, and multiplexer 32 selects the value 0 for IData.

[0057] When N = 1, the multiplexer 30 selects the binary value 01 for IType, and the multiplexer 32 selects the output of a combinational logic circuit CL0 for IData.

[0058] The CL0 circuit receives all the cache line addresses @PHY[39:6] and transmits the one for which the EN_i flag is active.

[0059] When 1 < N < 4, the multiplexer 30 selects the binary value 10 for IType, and the multiplexer 32 selects the output of a combinational logic circuit CL1 for IData.

[0060] The circuit CL1 receives all the IDX index values ​​contained in the least significant bits of the cache line addresses, namely the bits @PHY[11:6] in this example, and also receives the channel masks W[7:0]. The circuit forms the pairs of values ​​(IDX index, channel mask W) for only the values ​​which have an active EN_i flag, and positions them in 56 bits used to form the IData field.

[0061] When N > 4, the multiplexer 30 selects the binary value 11 for IType, and the multiplexer 32 selects the output of a combinational logic circuit CL2 for IData.

[0062] The CL2 circuit receives all the IDX index values ​​contained in the cache line addresses, namely the @PHY[11:6] bits in this example. The circuit forms a 64-bit mask by setting to 1 all the bits at the positions determined by the indices whose EN_i flag is active.

[0063] According to an alternative for the consolidated invalidation command of type PRECISE, since the directories store and transmit the channel information W, the circuit CL0 can be designed simply to extract the pair IDX, W from the single invalidation command received, in fact forming a special case of a command of type ARRAY_OF_SLOTS. This saves one clock cycle required to consult the cache tag memory, at the cost of the risk of invalidating "for nothing" an evicted cache line or additional cache lines if the mask W marks several channels. Such an alternative is in fact implemented by omitting the circuit CL0 and using the circuit CL1 for counts N between 1 and 4.

Claims

Claims

1. A method for managing cache coherence in a multi-core processor system, where each core (10) has a respective cache (L1) and has access to several banks (12) of a memory shared between the cores, the method comprising the following steps: managing a directory (DIR) in each memory bank for implementing directory-based cache coherence; writing to a current memory address by a core; searching the directory managing the current address for cores that have a cache line corresponding to the current address; sending to the cores identified by the directory respective invalidation (INVAL) commands of the cache line, the commands including the memory address of the cache line; and for each core, serving multiple invalidation commands received from different memory banks;characterized in that the step of serving the multiple invalidation commands comprises the following steps: counting the number (N) of invalidation commands received since a last clock cycle; when the count of invalidation commands received is equal to one, transmitting to the cache the invalidation command (PRECISE) received; and when the count of invalidation commands received is greater than one, transmitting to the cache a single command (C-INVAL) consolidating the cache lines identified in the invalidation commands received, in a format usable by the cache to simultaneously invalidate the identified cache lines.;

2. The method of claim 1, wherein each core has a respective multi-way set-associative cache, the method further comprising the steps of: recording in directories (DIR) the ways (W) in which the caches store the cache lines; transmitting the ways in the invalidation commands sent to the cores; and including in the consolidated invalidation command, for each invalidation command received, a pair of coordinates including a set index (IDX), extracted from the memory address, and the way (W).

3. A method according to claim 2, comprising the step of responding by the cache to the consolidated invalidation command by simultaneously invalidating each cache line located at the intersection of the set and the path determined by a respective pair of coordinates.

4. The method of claim 2, wherein the step of serving the multiple invalidation commands comprises the steps of: forming a bit mask (MASK_OF_SETS) comprising a 1 bit at positions identified by set indices (IDX) extracted from the memory addresses of the received invalidation commands; when the count of received invalidation commands is greater than a threshold, including the bit mask in the consolidated invalidation command; and responding by the cache to the consolidated command by simultaneously invalidating the cache sets marked in the bit mask.

5. The method of claim 2, wherein the directory records the cores and ways for each cache line as a combined bit mask marking the cores having the line in their cache and the ways in which the line is present in those cores' caches, and the way information transmitted in the invalidation commands includes the bits of the bit mask identifying the ways.

6. The method of claim 2, wherein the consolidated invalidation commands are configured to carry a field comprising a fixed number of bits of which: a first part (IType) identifies a type of command among an original invalidation command, a consolidated command with coordinates, and a consolidated command with bit mask, and a second part (IData) defines for the respective types: the address of the cache line, the coordinate pairs, and the bit mask.

7. A processor comprising: multiple cores (10) each comprising a local cache (Ll); multiple memory banks (12) forming a memory shared by the multiple cores; a directory-based cache coherence protocol manager (DIR), comprising for each core a circuit (16) for consolidating multiple cache line invalidation commands (INVAL) received from the different memory banks, the consolidation circuit comprising: a counter of received invalidation commands (PARCNT); a selection circuit (32) configured to, depending on whether the count (N) of invalidation commands received is equal to 1 or greater, transmit to the cache the single invalidation command received or an invalidation command consolidating the cache lines identified in the invalidation commands received, in a format usable by the cache to simultaneously invalidate the identified cache lines.

8. The processor of claim 7, wherein each core has a respective multi-way set-associative cache, the processor further comprising: a directory (DIR) associated with each bank, configured to record with each cache line, the ways in which the cache line is present in the different cores and to include the ways (W) in the invalidation commands sent to the cores; and the consolidation circuit configured to include in the consolidated invalidation command, for each cache line invalidation command received, a pair of coordinates including a set index (IDX), extracted from a memory address of the cache line, and the way (W).

9. The processor of claim 8, wherein the consolidation circuit is configured to include in the consolidated invalidation command, when the count of received invalidation commands is greater than a threshold, a bit mask marking cache sets to be invalidated, where each set includes the cache line identified by a respective received invalidation command.

10. A method for managing cache coherence in a multi-core processor system, where each core (10) has a respective multi-way set-associative cache (L1), and has access to several banks (12) of a memory shared between the cores, the method comprising the steps of: managing a directory (DIR) in each memory bank for implementing directory-based cache coherence; writing to a current memory address by a core; searching the directory managing the current address for cores that have a cache line corresponding to the current address; sending to the cores identified by the directory respective cache line invalidation (INVAL) commands, the commands including the memory address of the cache line; and for each core, serving multiple received invalidation commands from different memory banks; characterized in that it comprises the following steps: register in directories (DIR) the ways (W) in which the caches store the cache lines; transmit the paths in the disable commands sent to the cores; and serving the multiple invalidation commands received by a core by transmitting to the core's cache a single consolidated command (C-INVAL) including, for each invalidation command received, a pair of coordinates including a set index (IDX), extracted from the memory address of the cache line, and the path (W); and responding with a cache to a consolidated invalidation command by simultaneously invalidating each cache line located at the intersection of the set and the path determined by a respective pair of coordinates.

11. The method of claim 10, wherein the step of serving the multiple disable commands comprises the steps of: counting the number (N) of disable commands received since a last clock cycle; form a bit mask (MASK_OF_SETS) comprising a bit at 1 at the positions identified by the set indices (IDX); when the count of received invalidation commands is greater than a threshold, include the bit mask in the consolidated invalidation command instead of the coordinate pairs; and respond with the cache to the consolidated command by simultaneously invalidating the cache sets marked in the bitmask.

Citation Information

Patent Citations

  • Data Coherency Manager with Mapping Between Physical and Virtual Address Spaces

    US20190266092A1

  • Method and apparatus for supporting multiple cache line invalidations per cycle

    US6701417B2