Integrated circuit and method of performing cache management operations

CN115994113BActive Publication Date: 2026-09-08ANDES TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111634472.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-10-18
Filing Date
2021-12-29
Publication Date
2026-09-08
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

无论如何,一个计算核无法通过TileLink互连总线管理其他计算核的不同数据高速缓存之间的数据一致性

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115994113B_ABST
    Figure CN115994113B_ABST
Patent Text Reader

Abstract

The present application provides an integrated circuit and a method for performing cache management operations. The integrated circuit includes a master interface, a slave interface, and a link. The link is connected between the master interface and the slave interface, wherein the link includes an A channel, a B channel, a C channel, a D channel, and an E channel. The A channel can transmit cache management operation information from the master interface to the slave interface, wherein the cache management operation information is used to manage data coherency between different data caches. The D channel can transmit cache management operation response information from the slave interface to the master interface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an integrated circuit, and more particularly to a method for performing cache management operations in an integrated circuit. Background Technology

[0002] A system-on-chip (SoC) is a semiconductor technology used to integrate multiple complex, multifunctional systems onto a single chip. Silicon intellectual property (IP) embedded within an SoC can be interconnected via an interconnect bus. For example, TileLink is an interconnect bus designed for SoCs. The Level 1 cache (L1 cache, such as a data cache) of different compute cores can be connected to the Level 2 cache (L2 cache) via an interconnect bus conforming to the TileLink specification. However, a single compute core cannot manage data consistency between different data caches of other compute cores via the TileLink interconnect bus. Summary of the Invention

[0003] This invention provides an integrated circuit and a method for performing cache management operation (CMO) so that a main interface (e.g., the main interface of a first computing core) performs cache management operation on the data cache of other computing cores.

[0004] In an embodiment of the present invention, the integrated circuit includes a link and a master interface, wherein the link connects the master interface and the slave interface, and the link includes channels A, B, C, D, and E. The method for performing cache management operations includes: sending AcquireBlock information conforming to the TileLink specification from the master interface to the slave interface via channel A; receiving GrantData or Grant information conforming to the TileLink specification from the slave interface via channel D; sending GrantAck information conforming to the TileLink specification from the master interface to the slave interface via channel E; sending ReleaseData or Release information conforming to the TileLink specification from the master interface to the slave interface via channel C; and receiving GrantData or Release information conforming to the TileLink specification from the slave interface via channel D. The TileLink specification includes the following: ReleaseAck information; ProbeBlock or ProbePerm information compliant with the TileLink specification sent by the slave interface via channel B; ProbeAckData or ProbeAck information compliant with the TileLink specification sent by the master interface to the slave interface via channel C; cache management operation information sent by the master interface to the slave interface via channel A, where cache management operation information is used to manage data consistency between different data caches; and cache management operation response information received by the master interface from the slave interface via channel D. Note that the cache management operation information and cache management operation response information are not part of the TileLink specification.

[0005] In one embodiment of the present invention, the aforementioned integrated circuit includes a master interface, a slave interface, and a link. The link connects the master interface and the slave interface, and includes channels A, B, C, D, and E. Channel A is used to transmit AcquireBlock information conforming to the TileLink specification sent by the master interface to the slave interface. Channel D is used to transmit GrantData information or Grant information conforming to the TileLink specification sent by the slave interface to the master interface. Channel E is used to transmit GrantAck information conforming to the TileLink specification sent by the master interface to the slave interface. Channel C is used to transmit ReleaseData information or Release information conforming to the TileLink specification sent by the master interface to the slave interface. Channel D is further used to transmit ReleaseAck information conforming to the TileLink specification sent by the slave interface to the master interface. Channel B is used to transmit ProbeBlock information or ProbePerm information conforming to the TileLink specification sent by the slave interface to the master interface. Channel C is used to transmit ProbeAckData or ProbeAck information conforming to the TileLink specification sent by the master interface to the slave interface. Channel A is used to transmit cache management operation information from the master interface to the slave interface, where cache management operation information is used to manage data consistency between different data caches. Channel D is used to transmit cache management operation response information from the slave interface to the master interface. Cache management operation information and cache management operation response information are non-TileLink specified information.

[0006] Based on the above, in some embodiments, the first computing core of the integrated circuit can send cache management operation information to the second-level cache via a TileLink-compliant A channel. The second-level cache can execute the cache management operation information from the first computing core, and in response to the execution of the cache management operation information, send probe information (e.g., ProbeBlock or ProbePerm information) to other computing cores (the second computing core), thereby managing data consistency among all first-level caches (e.g., the data caches of the first and second computing cores). Therefore, the first computing core can perform cache management operations on the data caches of other computing cores by sending cache management operation information. Attached Figure Description

[0007] Figure 1 This is a schematic diagram of a circuit block of an integrated circuit according to an embodiment of the present invention;

[0008] Figure 2This is a flowchart illustrating a method for performing cache management operations in an integrated circuit according to an embodiment of the present invention.

[0009] Figure 3 This is a circuit block diagram of an integrated circuit according to one embodiment;

[0010] Figure 4 This is a schematic diagram of the circuit block for managing the computing core, as shown in one embodiment.

[0011] Figure 5 This is a schematic diagram illustrating the operation of computing and verifying the management of data consistency in its own data cache, according to one embodiment.

[0012] Figure 6 This is a schematic diagram illustrating the operation of a computing core managing data consistency with the data cache of other computing cores, according to one embodiment.

[0013] Figure 7 The diagram shows a circuit block diagram of a second-level cache according to one embodiment.

[0014] Explanation of reference numerals in the attached figures

[0015] 100, 300: Integrated Circuits

[0016] 110: Main Interface

[0017] 120: Subordinate Interface

[0018] 310: Consistent Interconnect Circuit

[0019] 320: Memory-mapped input / output circuit

[0020] 330: Level 2 Cache

[0021] 331: Memory Intermediate Circuit

[0022] 332: Register Intermediate Circuit

[0023] 333: Processing Register Circuit

[0024] 334: Control Circuit

[0025] 335: Static Random Access Memory (SRAM)

[0026] CHA: Channel A

[0027] CHB:B Channel

[0028] CHC:C channel

[0029] CHD:D Channel

[0030] CHE:E channel

[0031] CMO_M: CMO Information

[0032] CMOA_M:CMOAck information

[0033] CORE1, CORE2, CORE3, CORE4: Calculation cores

[0034] D$: Data Cache

[0035] I$: Instruction Cache

[0036] LINK1, LINK31, LINK32, LINK33, LINK34, LINK35, LINK61: Links

[0037] P_M: Detection Information

[0038] PA_M: Probe response information

[0039] R_M: Release information

[0040] RA_M: Release response information

[0041] S210, S220: Steps Detailed Implementation

[0042] Reference will now be made in detail to exemplary embodiments of the invention, examples of which are illustrated in the accompanying drawings. Wherever possible, the same element references are used in the drawings and description to denote the same or similar parts.

[0043] The term "coupled (or connected)" as used throughout this specification (including the claims) may refer to any direct or indirect means of connection. For example, if the text describes a first device coupled (or connected) to a second device, it should be interpreted as the first device being directly connected to the second device, or the first device being indirectly connected to the second device through other devices or some means of connection. The terms "first," "second," etc., used throughout this specification (including the claims) are used to name components or distinguish different embodiments or scopes, and are not intended to limit the upper or lower limit of the number of components, nor to limit the order of components. Furthermore, wherever possible, components / components / steps using the same reference numerals in the drawings and embodiments represent the same or similar parts. Components / components / steps using the same reference numerals or the same terms in different embodiments may be referred to mutually in the relevant descriptions.

[0044] Figure 1 This is a schematic diagram of a circuit block of an integrated circuit 100 according to an embodiment of the present invention. Figure 1 The integrated circuit 100 shown includes a master interface 110, a slave interface 120, and a link LINK1. Link LINK1 connects the master interface 110 and the slave interface 120. Figure 1 In the illustrated embodiment, link LINK1 includes channel A (CHA), channel B (CHB), channel C (CHC), channel D (CHD), and channel E (CHE). Channels A (CHA), B (CHB), C (CHC), D (CHD), and E (CHE) are link channels conforming to the TileLink specification. TileLink is a well-known technical specification and will not be elaborated upon here.

[0045] The master interface 110 can transmit TileLink-compliant messages to the slave interface 120 via channel A (CHA), channel C (CHC), and / or channel E (CHE), while the slave interface 120 can transmit TileLink-compliant messages to the master interface 110 via channel B (CHB) and / or channel D (CHD). For example, the master interface 110 can send a TileLink-compliant AcquireBlock message to the slave interface 120 via channel A (CHA), then the master interface 110 can receive TileLink-compliant GrantData (or Grant information) message sent by the slave interface 120 via channel D (CHD), and finally the master interface 110 can send a TileLink-compliant GrantAck message to the slave interface 120 via channel E (CHE). For example, the master interface 110 can send ReleaseData (or Release information) conforming to the TileLink specification to the slave interface 120 via channel C (CHC). Then, the master interface 110 can receive ReleaseAck information conforming to the TileLink specification sent by the slave interface 120 via channel D (CHD). For another example, the master interface 110 can receive ProbeBlock (or ProbePerm) information conforming to the TileLink specification sent by the slave interface 120 via channel B (CHB). Then, the master interface 110 can send ProbeAckData (or ProbeAck) information conforming to the TileLink specification to the slave interface 120 via channel C (CHC).

[0046] The AcquireBlock, GrantData, Grant, GrantAck, ReleaseData, Release, ReleaseAck, ProbeBlock, ProbePerm, ProbeAckData, and ProbeAck information mentioned above are all defined by the TileLink specification and will not be elaborated upon here. Furthermore, unlike the TileLink specification, channel A (CHA) can transmit the cache management operation message (CMO) from the master interface 110 to the slave interface 120, and channel D (CHD) can transmit the cache management operation acknowledgement message (CMOAck) from the slave interface 120 to the master interface 110. CMO and CMOAck information are custom information (not standardized by TileLink).

[0047] Figure 1 The master interface 110, slave interface 120, and link LINK1 shown can be applied to any transmission interface of the integrated circuit 100. For example, in some embodiments, the master interface 110 can be configured in the data cache (level 1 cache, L1 cache) of any of the multiple computing cores of the integrated circuit 100, while the slave interface 120 can be configured in a level 2 cache (L2 cache) of the integrated circuit 100. CMO information is used to manage data consistency between different data caches of different computing cores. For example, one computing core of the integrated circuit 100 (first computing core, not shown) Figure 1 It is possible to modify another (or more) computing cores (second computing cores, not shown) by transmitting CMO information. Figure 1 The permissions of the cache lines of the data cache are set to manage data consistency between different data caches.

[0048] Figure 2 This is a flowchart illustrating a method for performing cache management operations in an integrated circuit according to an embodiment of the present invention. Please refer to... Figure 1 and Figure 2 In step S210, the master interface 110 can send CMO information to the slave interface 120 via channel A (CHA). In step S220, the master interface 110 can receive the CMOAck information returned by the slave interface 120 via channel D (CHD).

[0049] Figure 3 This is a circuit block diagram of an integrated circuit 300 according to one embodiment. According to the actual design, Figure 3 The integrated circuit 300 shown can be referenced. Figure 1 The relevant description of the integrated circuit 100 shown, and / or Figure 1 The integrated circuit 100 shown can be referenced. Figure 3 The relevant description of the integrated circuit 300 shown is provided below. Figure 3 In the illustrated embodiment, the integrated circuit 300 includes multiple computing cores, for example... Figure 3 The calculation cores shown are CORE1, CORE2, CORE3, and CORE4. Figure 3 The symbol "I$" indicates the instruction cache, while Figure 3 The symbol “D$” represents the data cache. The data cache D$ of computing cores CORE1 to CORE4 can be regarded as the first-level cache (L1-cache).

[0050] Integrated circuit 300 also includes coherent interconnect circuitry 310, memory-mapped input / output (MMIO) circuitry 320, and a level 2 cache (L2 cache) 330. Depending on the actual design, the MMIO circuitry 320 can be any MMIO circuit, such as an existing MMIO circuit or another MMIO circuit.

[0051] The coherent interconnect circuit 310 is connected to the data cache D$ of computing cores CORE1 to CORE4 via links LINK31, LINK32, LINK33, and LINK34. The coherent interconnect circuit 310 is also connected to the second-level cache 330 via link LINK35. Depending on the actual design, the coherent interconnect circuit 310 can be any interconnect circuit, such as existing interconnect circuits or other interconnect circuits. Any of links LINK31 to LINK35 can be referenced... Figure 1 The relevant descriptions of link LINK1 shown are omitted here. One computing core of integrated circuit 300 (e.g., computing core CORE1) can change the cache line permissions of the data cache D$ of other computing cores (e.g., computing cores CORE2 to CORE4) of integrated circuit 300 by transmitting CM0 information, so as to manage data consistency between different data caches D$.

[0052] In actual operation, any one of the compute cores CORE1 to CORE4 can execute cache management operation instructions. These cache management operation instructions can manage data consistency between different data caches D$. This embodiment does not limit the syntax of the cache management operation instructions. For example, the cache management operation instructions can include the following pseudocode. Here, "L" is the label, "CMO.VAR." indicates a cache management operation, "cmo_specifier" indicates the operation type, "rd" indicates the destination register (used to store cache line addresses), "rs1" indicates the source register (used to store the current cache line address), and "rs2" indicates the boundary register (used to store the end cache line address). "bne" indicates a branch or jump, that is, when the address of rd has not yet reached the address of rs2, it returns to the label "L" to continue execution. Therefore, for cache lines within a specified range (from the cache line address stored in register rs1 to the cache line address stored in register rs2) in different data caches D$, the cache management operation instructions can manage the data consistency of one or more cache lines within the specified range.

[0053] L:CMO.VAR.<cmo_specifier> rd,rs1,rs2

[0054] bne rd,rs2,L

[0055] Based on the operation class "cmo_specifier", cache management operation instructions can perform at least one of the following operations: write-back or invalidate at least one cache line for each of different data caches D$ of different compute cores. The operation class "cmo_specifier" can be defined according to the actual design. For example, the operation class "cmo_specifier" can be "INVAL" (invalidate operation class), "WBINVAL" (write-back and invalidate operation class), or "WB" (write-back operation class). When the operation class "cmo_specifier" is "INVAL", the cache management operation instruction can invalidate one or more cache lines within a specified range. When the operation class "cmo_specifier" is "WBINVAL", the cache management operation instruction can write one or more cache lines within a specified range back to the second-level cache 330, and then invalidate the cache lines within the specified range. When the operation category "cmo_specifier" is "WB", the cache management operation instruction can write one or more cache lines within the specified range back to the second-level cache 330.

[0056] For example, compute core CORE1 can execute the cache management operation instructions. In response to the execution of these instructions, compute core CORE1 can send CMO information from the master interface of its data cache D$ to the slave interface of the consistency interconnect circuit 310. The consistency interconnect circuit 310 can then forward the CMO information from compute core CORE1 to the slave interface of the second-level cache 330. Based on the CMO information from compute core CORE1, the second-level cache 330 can change the permissions of the corresponding cache lines of the data caches D$ of compute cores CORE2 to CORE4 to manage data consistency between different data caches D$.

[0057] In detail, assuming that compute core CORE1 executes the aforementioned cache management operation instructions, compute core CORE1 can not only manage the data consistency of its own data cache D$, but also manage the data consistency of the data caches D$ of other compute cores (e.g., compute cores CORE2 to CORE4). Details regarding compute core CORE1's management of data consistency for its own data cache D$ will be provided in [reference needed]. Figure 5 This will be explained in more detail. Regarding the management of data consistency by the CORE1 compute core over the data cache D$ of other compute cores, the specifics will be provided in [reference needed]. Figure 6 Please provide an explanation.

[0058] Figure 4 As shown in one embodiment, Figure 3 The diagram shows a circuit block representation of the CORE1 core. Based on the actual design, Figure 3 The compute cores CORE2, CORE3, and CORE4 shown can be referenced. Figure 4 The relevant explanations for calculating core CORE1 shown are analogous, so they will not be repeated here. Figure 4The compute core CORE1 shown includes an instruction fetch unit (IFU) 410, an instruction cache unit (ICU) 420, an instruction pipeline (IPIPE) 430, a load / store unit (LSU) 440, and a data cache unit (DCU) 450. IFU 410 sends a fetch request to ICU 420. ICU 420 can forward the fetch request to the consistent interconnect circuit 310 via a TileLink-compliant link to fetch the instruction. IFU 410 sends the fetched instruction INST to IIPE 430. IIPE 430 decodes and executes the fetched instruction INST. When the fetched instruction INST is a memory access instruction, IIPE 430 sends a load / store request to LSU 440.

[0059] When a load request is a non-cacheable request, the LSU 440 can forward it to the Consistency Interconnect Circuit 310 via a TileLink-compliant link to retrieve the data. When a load request is a cacheable request, the LSU 440 can send it to the DCU 450. The DCU 450 can then forward the cacheable request to the Consistency Interconnect Circuit 310 via link LINK 31 to retrieve the data. Specifically, the DCU 450's master interface IFm sends a cacheable request to the slave interfaces IFs of the Consistency Interconnect Circuit 310, and then the slave interfaces IFs of the Consistency Interconnect Circuit 310 send the data back to the DCU 450's master interface IFm. Figure 4 The main interface IFm, link LINK31, and slave interface IFs shown can be referenced. Figure 1 The descriptions of the main interface 110, link LINK1 and slave interface 120 shown are analogous and will not be repeated here.

[0060] Figure 5 This is a schematic diagram illustrating the operation of the computing core CORE1 in managing data consistency of its own data cache D$, according to one embodiment. Figure 5 The vertical axis shown represents time. Please refer to... Figure 3 and Figure 5Assuming that the compute core CORE1 executes the cache management operation instruction, the master interface of the compute core CORE1's data cache D$ can send a release message R_M to the slave interface of the second-level cache 330 via the C channel CHC. By sending the release message R_M, the compute core CORE1's data cache D$ can automatically downgrade the permission of the target cache line. Therefore, the compute core CORE1 can manage the data consistency of its own data cache D$.

[0061] For example, suppose the operation class cmo_specifier of the cache management operation instruction is "WBINVAL". When the target cache line is in the "M state" (Modified state, meaning the target cache line only exists in the data cache D$ of compute core CORE1, and the target cache line is dirty), Figure 5 The release information R_M shown can be ReleaseData information conforming to the TileLink specification. Therefore, compute core CORE1 can write the dirty data of the target cache line back to the second-level cache 330 and reduce the permissions of the target cache line from "T" permissions (Tip permissions, i.e., the main interface has a copy of the content of the target cache line, and the main interface has read and write permissions for the target cache line) to "N" permissions (Nothing permissions, i.e., the main interface does not have the target cache line). Based on the ReleaseData information, the state of the target cache line changes from "M state" to "I state" (Invalid state, i.e., the target cache line does not exist in the data cache D$ of compute core CORE1). In the case where the state of the target cache line is "E state" (Exclusive state, i.e., the target cache line only exists in the data cache D$ of compute core CORE1, and the target cache line is clean), Figure 5 The release information R_M shown can be a Release message conforming to the TileLink specification. Therefore, compute core CORE1 can downgrade the target cache line's permissions from "T" to "N". Based on the Release information, the target cache line's state changes from "E state" to "I state". When the target cache line's state is "S state" (Shared state, meaning the target cache line exists in compute core CORE1's data cache D$, and the target cache line may share data cache D$ with other compute cores), Figure 5The release information R_M shown can be Release information. Therefore, the compute core CORE1 can downgrade the target cache line's permissions from "B" permissions (Branch permissions, meaning the main interface has a copy of the target cache line's content and read-only permissions) to "N" permissions. Based on the Release information, the target cache line's state changes from "S state" to "I state". "M state", "E state", "S state", and "I state" are the four states specified by the well-known MESI protocol, and therefore will not be elaborated upon here. "N" permissions, "B" permissions, and "T" permissions are three well-known cache line permission definitions, and therefore will not be elaborated upon here.

[0062] For another example, suppose the operation category cmo_specifier of the cache management operation instruction is "INVAL". When the target cache line is in the "M state", Figure 5 The release information R_M shown can be a Release message conforming to the TileLink specification, therefore, the compute core CORE1 can downgrade the target cache line's privileges from "T" to "N". Based on the Release information, the target cache line's state changes from "M state" to "I state". When the target cache line is in "E state", Figure 5 The release information R_M shown can be Release information, therefore, the compute core CORE1 can downgrade the target cache line's privileges from "T" to "N". Based on the Release information, the target cache line's state changes from "E state" to "I state". When the target cache line's state is "S state", Figure 5 The release information R_M shown can be a Release message, so the compute core CORE1 can downgrade the target cache line's privileges from "B" to "N". Based on the Release message, the target cache line's state changes from "S" to "I".

[0063] For example, suppose the operation class cmo_specifier of the cache management operation instruction is "WB". When the target cache line is in the "M state", Figure 5 The release information R_M shown can be ReleaseData information conforming to the TileLink specification. Therefore, the compute core CORE1 can write the dirty data of the target cache line back to the second-level cache 330 and reduce the privilege of the target cache line from "T" to "T" (that is, the privilege of the target cache line is maintained at "T"). Based on the Release information, the state of the target cache line changes from "M state" to "E state".

[0064] After the L2 cache 330 completes the execution of the release message R_M, the slave interface of the L2 cache 330 can send the release acknowledgment message RA_M to the master interface of the data cache D$ of the compute core CORE1 via the D channel CHD. For example, Figure 5 The release information R_M shown can be a ReleaseAck message conforming to the TileLink specification.

[0065] Figure 6 This is a schematic diagram illustrating the operation of computing core CORE1 in managing data consistency with the data cache D$ of other computing cores, according to one embodiment. Figure 6 The vertical axis shown represents time. Please refer to... Figure 3 and Figure 6 Assuming that compute core CORE1 executes the cache management operation instruction, the master interface of compute core CORE1's data cache D$ can send a cache management operation (CMO) message CMO_M to the slave interface of the second-level cache 330 via channel A CHA. By sending the CMO message CMO_M, compute core CORE1's data cache D$ can change the permissions of the target cache lines of other compute cores (e.g., compute cores CORE2 to CORE4)'s data cache D$ to manage data consistency among compute cores CORE1 to CORE4's data cache D$.

[0066] In detail, the second-level cache 330 receives and executes the CMO information CMO_M from compute core CORE1. In response to the execution of CMO_M, the second-level cache 330 sends probe information P_M via channel B (CHB) to the data caches D$ of other compute cores (e.g., compute cores CORE2 to CORE4) to manage data consistency between the data caches D$ of compute core CORE1 and CORE2 to CORE4. Depending on the actual operation, the probe information P_M can be a ProbeBlock message or a ProbePerm message conforming to the TileLink specification. The probe information P_M can change the permissions of the target cache line of the data cache D$ of the compute core (e.g., compute cores CORE2 to CORE4). In response to the execution of probe information P_M, compute cores CORE2 to CORE4 send back probe response information PA_M to the second-level cache 330 via channel C (CHC). Based on the probe response information PA_M transmitted by the computing cores CORE2 to CORE4, the second-level cache 330 can send the CMOAck information CMOA_M to the computing core CORE1 through the D channel CHD.

[0067] Therefore, compute core CORE1 can manage data consistency for the data cache D$ of other compute cores (e.g., compute cores CORE2 to CORE4). The following will refer to... Figure 6 The protocol shown illustrates several specific examples of how compute core CORE1 changes the target cache line permissions for other compute cores CORE2 to CORE4.

[0068] For example, suppose the operation class cmo_specifier of the cache management operation instruction is "WBINVAL". When the target cache line is in the "M state", Figure 6 The CMO information shown, CMO_M, can be CMOBlock information (CMO information with cached data), while Figure 6 The probe information P_M shown can be ProbeBlock information conforming to the TileLink specification. Therefore, the second-level cache 330 can provide the data of the target cache line to the compute cores CORE2 to CORE4. Based on the ProbeBlock information, compute cores CORE2 to CORE4 send back ProbeAckData information (probe acknowledgment information PA_M) conforming to the TileLink specification to the second-level cache 330 via the C channel CHC. Therefore, the privilege of the target cache line in the data cache D$ of compute cores CORE2 to CORE4 can be reduced from "T" privilege to "N" privilege. Based on the CMOBlock information, the state of the target cache line changes from "M state" to "I state". When the state of the target cache line is "E state", Figure 6 The CMO information shown, CMO_M, can be CMOBlock information. Figure 6 The probe information P_M shown can be ProbeBlock information, while Figure 6 The probe response information PA_M shown can be a ProbeAck message conforming to the TileLink specification. Therefore, the privileges of the target cache line in the data cache D$ of compute cores CORE2 to CORE4 can be reduced from "T" privileges to "N" privileges. Based on the CMOBlock information, the state of the target cache line changes from "E state" to "I state". When the target cache line is in "S state", Figure 6 The CMO information shown, CMO_M, can be CMOBlock information. Figure 6 The probe information P_M shown can be ProbeBlock information, while Figure 6The probe response information PA_M shown can be ProbeAck information. Therefore, the privilege of the target cache line in the data cache D$ of compute cores CORE2 to CORE4 can be reduced from "B" privilege to "N" privilege. Based on the CMOBlock information, the state of the target cache line changes from "S state" to "I state". When the state of the target cache line is "I state", Figure 6 The CMO information shown, CMO_M, can be CMOBlock information. Figure 6 The probe information P_M shown can be ProbeBlock information, while Figure 6 The probe response information PA_M shown can be a ProbeAck message. Therefore, the access permission of the target cache line in the data cache D$ of compute cores CORE2 to CORE4 can be maintained at "N" permission. Based on the CMOBlock information, the state of the target cache line is maintained at "I".

[0069] For another example, suppose the operation category cmo_specifier of the cache management operation instruction is "INVAL". When the target cache line is in the "M state", Figure 6 The CMO information shown, CMO_M, can be CMOPerm information (CMO information without cached data), while... Figure 6 The probe information P_M shown can be ProbePerm information conforming to the TileLink specification. Based on the ProbePerm information, compute cores CORE2 to CORE4 send back ProbeAck information (probe acknowledgment information PA_M) conforming to the TileLink specification to the second-level cache 330 via the C channel CHC. Therefore, the privilege of the target cache line in the data cache D$ of compute cores CORE2 to CORE4 can be reduced from "T" privilege to "N" privilege. Based on the CMOPerm information, the state of the target cache line changes from "M state" to "I state". When the state of the target cache line is "E state", Figure 6 The CMO information shown, CMO_M, can be CMOPerm information. Figure 6 The probe information P_M shown can be ProbePerm information, while Figure 6 The probe response information PA_M shown can be ProbeAck information. Therefore, the privileges of the target cache line in the data cache D$ of compute cores CORE2 to CORE4 can be reduced from "T" privileges to "N" privileges. Based on the CMOperm information, the state of the target cache line changes from "E state" to "I state". When the target cache line is in the "S state", Figure 6 The CMO information shown, CMO_M, can be CMOPerm information. Figure 6 The probe information P_M shown can be ProbePerm information, while Figure 6 The probe response information PA_M shown can be ProbeAck information. Therefore, the privileges of the target cache line in the data cache D$ of compute cores CORE2 to CORE4 can be reduced from "B" privileges to "N" privileges. Based on the CMOperm information, the state of the target cache line changes from "S state" to "I state". When the target cache line is in the "I state", Figure 6 The CMO information shown, CMO_M, can be CMOPerm information. Figure 6 The probe information P_M shown can be ProbePerm information, while Figure 6 The probe response information PA_M shown can be a ProbeAck message. Therefore, the access permission of the target cache line in the data cache D$ of CORE2 to CORE4 can be maintained at "N" permission. Based on the CMOperm information, the state of the target cache line is maintained at "I".

[0070] For example, suppose the operation class cmo_specifier of the cache management operation instruction is "WB". When the target cache line is in the "M state", Figure 6 The CMO information shown, CMO_M, can be CMOBlock information. Figure 6 The probe information P_M shown can be ProbeBlock information conforming to the TileLink specification, while Figure 6 The probe response information PA_M shown can be ProbeAckData information conforming to the TileLink specification. Therefore, the privileges of the target cache line in the data cache D$ of compute cores CORE2 to CORE4 can be maintained at "T" privileges. Based on the CMOBlock information, the state of the target cache line changes from "M state" to "E state". When the target cache line is in the "E state", Figure 6 The CMO information shown, CMO_M, can be CMOBlock information. Figure 6 The probe information P_M shown can be ProbeBlock information, while Figure 6 The probe response information PA_M shown can be a ProbeAck message conforming to the TileLink specification. Therefore, the privileges of the target cache line in the data cache D$ of compute cores CORE2 to CORE4 can be maintained at "T". Based on the CMOBlock information, the state of the target cache line is maintained at "E state". When the state of the target cache line is "S state", Figure 6 The CMO information shown, CMO_M, can be CMOBlock information. Figure 6 The probe information P_M shown can be ProbeBlock information, while Figure 6 The probe response information PA_M shown can be ProbeAck information. Therefore, the permission of the target cache line in the data cache D$ of compute cores CORE2 to CORE4 can be maintained at "B". Based on the CMOBlock information, the state of the target cache line is maintained at "S state". When the state of the target cache line is "I state", Figure 6 The CMO information shown, CMO_M, can be CMOBlock information. Figure 6 The probe information P_M shown can be ProbeBlock information, while Figure 6 The probe response information PA_M shown can be a ProbeAck message. Therefore, the access permission of the target cache line in the data cache D$ of compute cores CORE2 to CORE4 can be maintained at "N" permission. Based on the CMOBlock information, the state of the target cache line is maintained at "I".

[0071] In summary, the first computing core (e.g., computing core CORE1) of the integrated circuit 300 described in the above embodiments can send non-TileLink specification CMO information CMO_M to the second-level cache 330 via a TileLink specification-compliant A channel. The second-level cache 330 can execute the CMO information CMO_M from the first computing core, and in response to the execution of the CMO information CMO_M, send probe information (e.g., ProbeBlock information or ProbePerm information) to other computing cores (second computing cores, such as computing cores CORE2 to CORE4), thereby managing data consistency among all first-level caches (e.g., the data caches D$ of computing cores CORE1 to CORE4). Therefore, any one of computing cores CORE1 to CORE4 can perform cache management operations on the data caches D$ of other computing cores by sending the CMO information CMO_M.

[0072] Figure 1 The link LINK1 shown is not limited to applications in [specific applications]. Figure 3 The links shown are LINK31 to LINK35. According to actual design, in some embodiments, Figure 1 The illustrated link LINK1 can be applied to the internal circuitry of the second-level cache 330. For example, Figure 7 This is a circuit block diagram of a second-level cache 330 according to one embodiment. Please refer to... Figure 3 and Figure 7 . Figure 7The second-level cache 330 shown includes a memory agent circuit 331, a register agent circuit 332, a handling register circuit 333, a control circuit 334, and static random access memory (SRAM) 335.

[0073] Any of the compute cores CORE1 to CORE4 can transmit any information defined by the TileLink specification to the memory intermediary circuit 331 through the consistent interconnect circuit 310. The master interface of the memory intermediary circuit 331 can transmit information defined by the TileLink specification to the slave interface of the processing register circuit 333 through link LINK61. Link LINK61 can be referenced... Figure 1 The relevant descriptions of link LINK1 shown are omitted here.

[0074] The register intermediary circuit 332 has multiple registers. For example, according to a practical design, the register intermediary circuit 332 has a command register and an access line register corresponding to each of the computing cores CORE1 to CORE4. Any one of the command registers is used to record the command issued by a corresponding computing core, while the access line register is used to record the address / index of the target cache line issued by a corresponding computing core. Furthermore, according to a practical design, the register intermediary circuit 332 may also be configured with a cache control status register. Based on the contents of these registers, the register intermediary circuit 332 can generate CMO information for the memory intermediary circuit 331, and the main interface of the memory intermediary circuit 331 can transfer the CMO information to the main interface of the processing register circuit 333 via link LINK61. The CMO information can be deduced by analogy with the relevant descriptions of the foregoing embodiments, and therefore will not be repeated here.

[0075] The processing register circuit 333 is also known as the handling register file (HRF). The processing register circuit 333 can notify the control circuit 334 to perform corresponding operations based on information from link 61, thereby accessing SRAM 335. For example, link 61 can transmit information defined by the TileLink specification to the processing register circuit 333, and the processing register circuit 333 can execute the TileLink information to access SRAM 335 through the control circuit 334. When the register intermediary circuit 332 transmits CMO information to the processing register circuit 333 through the memory intermediary circuit 331 and link 61, the processing register circuit 333 can respond to the execution of the CMO information by changing the permissions of the cache lines of SRAM 335 to manage data consistency between the data cache D$ of computing cores CORE1 to CORE4 and SRAM 335.

[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for performing cache management operations in an integrated circuit, the integrated circuit including a link and a master interface, the link being connected between the master interface and a slave interface, the link including channel A, channel B, channel C, channel D, and channel E, characterized in that, The method includes: The master interface sends an AcquireBlock message conforming to the TileLink specification to the slave interface via the A channel; The master interface receives GrantData or Grant information conforming to the TileLink specification sent by the slave interface through the D channel; The master interface sends GrantAck information conforming to the TileLink specification to the slave interface via the E channel; The master interface sends ReleaseData or Release information conforming to the TileLink specification to the slave interface via the C channel; The master interface receives ReleaseAck information conforming to the TileLink specification sent by the slave interface through the D channel; The master interface receives ProbeBlock or ProbePerm information conforming to the TileLink specification from the slave interface via the B channel; The master interface sends ProbeAckData or ProbeAck information conforming to the TileLink specification to the slave interface via the C channel; The master interface sends cache management operation information to the slave interface via channel A, wherein the cache management operation information is used to manage data consistency between different data caches; and The master interface receives cache management operation response information from the slave interface via the D channel, wherein the cache management operation information and the cache management operation response information are unstandardized TileLink information, and The master interface and the slave interface are respectively configured in the first computing core of the integrated circuit and the second-level cache of the integrated circuit, or the master interface and the slave interface are respectively configured in the memory intermediary circuit of the second-level cache and the processing register circuit of the second-level cache.

2. The method according to claim 1, characterized in that, The method further includes: The first computing core executes cache management operation instructions, wherein the cache management operation instructions are used to manage data consistency between different data caches; and In response to the execution of the cache management operation instruction, the cache management operation information is sent from the master interface of the first computing core to the slave interface of the second-level cache.

3. The method according to claim 2, characterized in that, The cache management operation instructions are used to perform at least one of write-back and invalidation on at least one cache line of each of the different data caches of different computing cores.

4. The method according to claim 1, characterized in that, The method further includes: The cache management operation information from the first computing core is received and executed by the second-level cache; The second-level cache sends the ProbeBlock information or the ProbePerm information to at least one second computing core in response to the execution of the cache management operation information, so as to manage the data consistency between the data cache of the first computing core and at least one data cache of the at least one second computing core; The second-level cache receives the ProbeAckData information or the ProbeAck information from the at least one second computing core; and The second-level cache sends the cache management operation response information to the first computing core.

5. The method according to claim 4, characterized in that, The method further includes: The cache management operation instructions are executed by the first computing core; The master interface of the data cache of the first computing core sends the cache management operation information to the slave interface of the second-level cache through the A channel to change the permissions of the target cache line of the at least one data cache of the at least one second computing core.

6. The method according to claim 5, characterized in that, The cache management operation information is CMOBlock information, and the second-level cache sends the ProbeBlock information to the at least one second computing core in response to the execution of the cache management operation information.

7. The method according to claim 5, characterized in that, The cache management operation information is CMOPerm information, and the second-level cache sends the ProbePerm information to the at least one second computing core in response to the execution of the cache management operation information.

8. The method according to claim 5, characterized in that, The operation category of the cache management operation instruction is write-back and invalidation operation.

9. The method according to claim 5, characterized in that, The operation category of the cache management operation instruction is the invalidation operation category.

10. The method according to claim 5, characterized in that, The operation category of the cache management operation instruction is the write-back operation category.

11. The method according to claim 1, characterized in that, The second-level cache also includes register intermediary circuitry, control circuitry, and static random access memory, and the method further includes: One of the multiple computing cores transmits information defined by the TileLink specification to the memory intermediary circuit via a consistent interconnect circuit; The information is transmitted from the master interface of the memory intermediary circuit to the slave interface of the processing register circuit via the link; The register intermediary circuit generates the cache management operation information and sends it to the memory intermediary circuit. The master interface of the memory intermediary circuit transfers the cache management operation information to the slave interface of the processing register circuit via the link; and The processing register circuitry changes the permissions of a cache line in the static random access memory in response to the execution of the cache management operation information, in order to manage data consistency between the data cache of each of the plurality of computing cores and the static random access memory.

12. An integrated circuit, characterized in that, The integrated circuit includes: Main interface; Slave interface, wherein the master interface and the slave interface are respectively configured in the first computing core of the integrated circuit and the second-level cache of the integrated circuit, or the master interface and the slave interface are respectively configured in the memory intermediary circuit of the second-level cache and the processing register circuit of the second-level cache; and A link, connecting the master interface and the slave interface, includes channels A, B, C, D, and E. Channel A transmits AcquireBlock information conforming to the TileLink specification sent by the master interface to the slave interface. Channel D transmits GrantData or Grant information conforming to the TileLink specification sent by the slave interface to the master interface. Channel E transmits GrantAck information conforming to the TileLink specification sent by the master interface to the slave interface. Channel C transmits ReleaseData or Release information conforming to the TileLink specification sent by the master interface to the slave interface. Channel D further transmits TileLink information conforming to the TileLink specification sent by the slave interface. The releaseAck information conforming to the TileLink specification is transmitted to the master interface. The B channel is used to transmit the ProbeBlock information or ProbePerm information conforming to the TileLink specification sent by the slave interface to the master interface. The C channel is further used to transmit the ProbeAckData information or ProbeAck information conforming to the TileLink specification sent by the master interface to the slave interface. The A channel is further used to transmit the cache management operation information of the master interface to the slave interface. The cache management operation information is used to manage data consistency between different data caches. The D channel is further used to transmit the cache management operation response information of the slave interface to the master interface. The cache management operation information and the cache management operation response information are TileLink non-standard information.

13. The integrated circuit according to claim 12, characterized in that, The first computing core is used to execute cache management operation instructions to manage data consistency between different data caches, wherein the first computing core sends cache management operation information from the master interface of the first computing core to the slave interface of the second-level cache in response to the execution of the cache management operation instructions.

14. The integrated circuit according to claim 13, characterized in that, The cache management operation instructions are used to perform at least one of write-back and invalidation on at least one cache line of each of the different data caches of different computing cores.

15. The integrated circuit according to claim 12, characterized in that, The second-level cache is used to receive and execute cache management operation information from the first computing core, wherein the second-level cache sends the ProbeBlock information or the ProbePerm information to at least one second computing core in response to the execution of the cache management operation information to manage data consistency between the data cache of the first computing core and at least one data cache of the at least one second computing core, the second-level cache receives the ProbeAckData information or the ProbeAck information from the at least one second computing core, and the second-level cache sends the cache management operation response information to the first computing core.

16. The integrated circuit according to claim 15, characterized in that, The first computing core executes cache management operation instructions, and the master interface of the first computing core's data cache sends the cache management operation information to the slave interface of the second-level cache through the A channel to change the permissions of the target cache line of the at least one data cache of the at least one second computing core.

17. The integrated circuit according to claim 16, characterized in that, The cache management operation information is CMOBlock information, and the second-level cache sends the ProbeBlock information to the at least one second computing core in response to the execution of the cache management operation information.

18. The integrated circuit according to claim 16, characterized in that, The cache management operation information is CMOPerm information, and the second-level cache sends the ProbePerm information to the at least one second computing core in response to the execution of the cache management operation information.

19. The integrated circuit according to claim 16, characterized in that, The operation category of the cache management operation instruction is write-back and invalidation operation.

20. The integrated circuit according to claim 16, characterized in that, The operation category of the cache management operation instruction is the invalidation operation category.

21. The integrated circuit according to claim 16, characterized in that, The operation category of the cache management operation instruction is the write-back operation category.

22. The integrated circuit according to claim 12, characterized in that, The second-level cache also includes a register intermediary circuit, a control circuit, and static random access memory (SRAM). One of the multiple computing cores transmits information defined by the TileLink specification to the memory intermediary circuit via a consistent interconnect circuit. The master interface of the memory intermediary circuit transmits the information defined by the TileLink specification to the slave interface of the processing register circuit via the link. The register intermediary circuit generates cache management operation information for the memory intermediary circuit. The master interface of the memory intermediary circuit forwards the cache management operation information to the slave interface of the processing register circuit via the link. The processing register circuit, in response to the execution of the cache management operation information, modifies the permissions of the cache line of the SRAM to manage data consistency between the data cache of each of the multiple computing cores and the SRAM.

Citation Information

Patent Citations

  • Integrated circuit capable of independently operating a plurality of communication channels

    CN101310262A

  • Virtualized caches

    US20210157728A1