Circuit and Method

By introducing master nodes and monitoring filters in the on-chip system, the consistency and serialized management problems in the data transmission protocol are solved, data conflicts and live locks are avoided, and the efficiency and reliability of data transmission are improved.

CN112631988BActive Publication Date: 2025-07-08ARM LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011062537.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-08
Filing Date
2020-09-30
Publication Date
2025-07-08
Estimated Expiration
2040-09-30

AI Technical Summary

Technical Problem

In systems on chip or network on chip systems, existing data transmission protocols such as AMBA CHI protocols are difficult to effectively manage data consistency and serialization, resulting in data processing transaction conflicts and live locks.

Method used

By introducing a master node into the system, it is responsible for serialization and consistency control of data access operations, ensuring the consistency of data written to the memory address with the read data of subsequent access requests, and detecting and avoiding data conflicts through monitoring filters and exclusive monitors.

Benefits of technology

Data consistency and serialization between data processing nodes are realized, and the live lock situation is avoided, and the efficiency and reliability of data transmission are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112631988B_ABST
    Figure CN112631988B_ABST
Patent Text Reader

Abstract

This application relates to circuits and methods. The circuit includes: a set of two or more data processing nodes, each data processing node having a respective storage circuit for holding data; and a master node configured to serialize data access operations and control the consistency between data held by one or more data processing nodes such that data written to a memory address is consistent with data read from that memory address in response to a subsequent access request; a requesting node of the set of data processing nodes is configured to transmit a request to the master node to perform an exclusive access to a given data instance at a given memory address; the master node is configured to transmit information to other data processing nodes of the set of data processing nodes in response to the request to control processing by those other data processing nodes of any further data instances held by those other data processing nodes at the given memory address.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to circuits and methods. Background Art

[0002] For example, in the context of a system-on-chip (SoC) or a network-on-chip (NoC) system, a data transfer protocol can regulate the operation of data transfer between devices or nodes interconnected via an interconnect circuit. An example of such a data transfer protocol is the so-called AMBA (Advanced Microcontroller Bus Architecture) CHI (Coherent Hub Interface) protocol.

[0003] In the CHI protocol, nodes can be classified as request nodes (RNs), home nodes (HNs), or slave nodes (SNs). Nodes can be fully coherent or input / output (I / O) coherent. Fully coherent HNs or RNs (HN-F and RN-F respectively) include coherent cache storage; fully coherent SNs (SN-F) are paired with HN-F. HN-F can manage the coherence and / or serialization of a memory region and can be referred to as an example of a point of coherence (POC) and / or a point of serialization (POS).

[0004] Here, the term "coherent" means that data written by one node to a memory address in a coherent memory system is consistent with data subsequently read by another node from that memory address in the coherent memory system. Thus, the role of the logic associated with the coherence function is to ensure that, before a data processing transaction occurs, the most up-to-date copy is provided. If another node changes its copy, the coherence system will invalidate the other copies, which must then be reacquired if needed. Similarly, if a data processing transaction involves modifying a data item, the coherence logic will avoid conflicts with other existing copies of the data item.

[0005] Serialization involves the order of processing of memory access requests from potentially multiple request nodes and may require different waiting periods to be serviced such that the results of these requests are presented to the request nodes in the correct order and any dependencies between the requests (e.g., data read after being written to the same address) are correctly handled.

[0006] Data access, such as a read request, can be performed via the HN-F, which can service the read request itself (e.g., by accessing a cache memory), or can forward the read request to the SN-F for parsing, e.g., if the desired data item has to be read from main memory or a higher-level cache memory. In such an example, the SN-F can include a dynamic memory controller (DMC) associated with a memory such as a dynamic random access memory (DRAM). In instances where the HN-F itself cannot service the request, the HN-F processes the issuing of a read request to the SN-F. SUMMARY OF THE INVENTION

[0007] In an example arrangement, a circuit is provided that includes:

[0008] a set of two or more data processing nodes, each data processing node having a respective storage circuit for holding data; and

[0009] a master node for serializing data access operations and controlling the consistency between data held by one or more of the data processing nodes such that data written to a memory address is consistent with data read from that memory address in response to a subsequent access request;

[0010] wherein:

[0011] a requesting node of the set of data processing nodes is configured to transmit a request to the master node for exclusive access to a given data instance at a given memory address; and

[0012] the master node is configured to transmit information to other data processing nodes of the set of data processing nodes in response to the request to control the processing by those other data processing nodes of any additional data instances held by those other data processing nodes at the given memory address.

[0013] In another example arrangement, a method is provided that includes:

[0014] holding data by a set of two or more data processing nodes;

[0015] serializing data access operations by a master node;

[0016] controlling, by the master node, the consistency between data held by one or more of the data processing nodes such that data written to a memory address is consistent with data read from that memory address in response to a subsequent access request;

[0017] A requesting node in a set of data processing nodes transmits a request to a master node to perform an exclusive access to a given data instance at a given memory address; and

[0018] In response to the request, the master node transmits information to other data processing nodes in the set of data processing nodes to control the processing by those other data processing nodes of any additional data instances held by those other data processing nodes at the given memory address.

[0019] In another example arrangement, a circuit is provided that includes:

[0020] A set of two or more data processing nodes, each data processing node having a respective storage circuit for holding data; and

[0021] A master node for serializing the execution of data access operations and controlling the consistency between data held by one or more data processing nodes such that data written to a memory address is consistent with data read from that memory address in response to a subsequent access request;

[0022] Wherein:

[0023] The requesting node of the set of data processing nodes is configured to initiate an operation sequence that requests exclusive access to data at a given memory address, the operation sequence including an exclusive store operation to the given memory address; and

[0024] The master node is configured to, prior to executing the exclusive store operation, detect whether the requesting node currently holds a data instance at the given memory address by issuing a request to the requesting node, and selectively execute the exclusive store operation in response to the detection, the request for the requesting node to indicate to the master node whether the requesting node currently holds a data instance at the given memory address.

[0025] In another example arrangement, a method is provided that includes:

[0026] Holding data by a set of two or more data processing nodes;

[0027] Serializing data access operations by a master node;

[0028] Controlling, by the master node, the consistency between data held by one or more data processing nodes such that data written to a memory address is consistent with data read from that memory address in response to a subsequent access request;

[0029] Initiate a sequence of operations that requests exclusive access to data at a given memory address through a requesting node in a set of data processing nodes, the sequence of operations including an exclusive store operation on the given memory address;

[0030] Before performing the exclusive store operation, detect, by a master node, whether the requesting node currently holds a data instance at the given memory address, the detecting step including issuing a request to the requesting node, the request for the requesting node to indicate to the master node whether the requesting node currently holds a data instance at the given memory address; and

[0031] Selectively perform the exclusive store operation in response to the detection.

[0032] In another example arrangement, a circuit is provided that includes:

[0033] A set of two or more data processing nodes, each data processing node having a respective storage circuit for holding data; and

[0034] A master node for serializing the execution of data access operations and controlling the consistency between data held by one or more data processing nodes so that data written to a memory address is consistent with data read from that memory address in response to subsequent access requests;

[0035] Wherein:

[0036] The requesting node of the set of data processing nodes is configured to initiate an exclusive request transaction for a given memory address; and

[0037] The master node is configured to, in response to the exclusive request transaction, detect whether the requesting node stores a data instance at the given memory address when executing the exclusive request transaction, and when it is detected that the requesting node stores a data instance at the given memory address, instruct other data processing nodes in the set of data processing nodes to invalidate any additional data instances at the given memory address.

[0038] In another example arrangement, a method is provided that includes:

[0039] Hold data through a set of two or more data processing nodes;

[0040] Serialize data access operations through a master node;

[0041] Control the consistency between data held by one or more data processing nodes through the master node so that data written to a memory address is consistent with data read from that memory address in response to subsequent access requests;

[0042] Initiate an exclusive request transaction for a given memory address through a requesting node in a set of data processing nodes; and

[0043] Through a master node, in response to the exclusive request transaction, detect whether the requesting node stores a data instance at the given memory address when executing the exclusive request transaction; and

[0044] Through a master node, when it is detected that the requesting node stores a data instance at the given memory address, instruct other data processing nodes in the set of data processing nodes to invalidate any additional data instances at the given memory address.

[0045] Further corresponding aspects and features of the present technology are defined by the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The present technology will be further described only by way of example with reference to its embodiments shown in the accompanying drawings, in which:

[0047] Figure 1 An example circuit is schematically shown;

[0048] Figure 2 A master node is schematically shown;

[0049] Figure 3 A cache memory is schematically shown;

[0050] Figure 4a An exclusive sequence is schematically shown;

[0051] Figure 4b The use of an exclusive monitor is schematically shown;

[0052] Figure 5 and Figure 6 is a schematic flowchart showing a corresponding method;

[0053] Figure 7 is a schematic timeline representation;

[0054] Figures 8 to 10 is a schematic flowchart showing a corresponding method;

[0055] Figure 11 Snooping is schematically shown;

[0056] Figure 12 and Figure 13 is a corresponding schematic timeline representation; and

[0057] Figures 14 to 17 is a schematic flowchart showing a corresponding method. Detailed implementation mode

[0058] Before discussing the embodiments with reference to the accompanying drawings, the following description of the embodiments is provided.

[0059] An example embodiment provides a circuit that includes: a set of two or more data processing nodes, each data processing node having a corresponding storage circuit for holding data; and a master node configured to serialize data access operations and control the consistency between data held by one or more data processing nodes so that data written to a memory address is consistent with data read from that memory address in response to a subsequent access request; wherein: a requesting node of the set of data processing nodes is configured to send a request to the master node for exclusive access to a given data instance at a given memory address; and the master node is configured to send information to other data processing nodes of the set of data processing nodes in response to the request to control the processing of any further data instances held by those other data processing nodes at the given memory address.

[0060] In the example embodiment, in response to a particular requesting node (RN) requesting exclusive access to a given address (e.g., one or more cache lines), the master node or HN may initiate various actions or detections at other nodes that can hold the cache line. The given instance is the instance held at the requesting node. Examples of actions that may be initiated include configuring the master node to initiate transferring data at the given memory address to the requesting node. In other examples, the master node may be configured to control the detection of whether each of the other data processing nodes is currently performing an exclusive access operation for the given memory address. If so, that is, it is detected that one of the other data processing nodes is currently performing an exclusive access operation for the given memory address, then one of the other data processing nodes may be configured to retain the data instance at the given memory address and indicate to the master node that the given memory address has a shared state between one of the other data processing nodes and the requesting node. However, if not, that is, it is detected that one of the other data processing nodes is not currently performing an exclusive access operation for the given memory address, then one of the other data processing nodes may be configured to invalidate the data instance at the given memory address held at one of the other data processing nodes.

[0061] In some examples, the operation of the master node may be assisted by configuring the master node to maintain listening data indicating data instances stored at one or more of the data processing nodes in the set of data processing nodes.

[0062] Another example embodiment provides a method that includes: maintaining data by a set of two or more data processing nodes; serializing data access operations by a master node; controlling consistency between data maintained by one or more data processing nodes by the master node such that data written to a memory address is consistent with data read from that memory address in response to a subsequent access request; transmitting, by a requesting node of the set of data processing, a request to the master node to perform an exclusive access to a given data instance at a given memory address; and transmitting, by the master node in response to the request, information to other data processing nodes in the set of data processing nodes to control processing by those other data processing nodes of any further data instances maintained by those other data processing nodes at the given memory address.

[0063] Another example embodiment provides a circuit that includes: a set of two or more data processing nodes, each data processing node having a respective storage circuit for maintaining data; and a master node for serializing the execution of data access operations and for controlling consistency between data maintained by one or more data processing nodes such that data written to a memory address is consistent with data read from that memory address in response to a subsequent access request; wherein: a requesting node of the set of data processing nodes is configured to initiate an operation sequence that requests an exclusive access to data at a given memory address, the operation sequence including an exclusive store operation to the given memory address; and the master node is configured to, prior to performing the exclusive store operation, detect whether the requesting node currently holds a data instance at the given memory address by issuing a request to the requesting node, and selectively perform the exclusive store operation in response to the detection, the request for the requesting node to indicate to the master node whether the requesting node currently holds a data instance at the given memory address.

[0064] Example embodiments can provide an arrangement by which an exclusive store can be selectively performed based on whether a requesting node still holds a copy of a relevant memory address, such as a cache line, at the time of performing the exclusive store. Using such an arrangement can help avoid so-called live-lock situations due to exclusive contention between processes, cores, or nodes.

[0065] In some examples, the master node is configured to perform an exclusive store operation when the master node detects that the requesting node currently holds a data instance at the given memory address. The operation of the master node can be assisted by configuring the master node to maintain control data that at least partially indicates which data is currently held by the storage circuits of the data processing nodes.

[0066] The above operation sequence can include applying an exclusive label to at least the given memory address and subsequently releasing the exclusive label.

[0067] Another example embodiment provides a method that includes: maintaining data by a set of two or more data processing nodes; serializing data access operations by a master node; controlling consistency among data maintained by one or more data processing nodes by the master node such that data written to a memory address is consistent with data read from the memory address in response to a subsequent access request; initiating, by a requesting node in the set of data processing nodes, an operation sequence that requests exclusive access to data at a given memory address, the operation sequence including an exclusive store operation to the given memory address; detecting, by the master node before performing the exclusive store operation, whether the requesting node currently holds a data instance at the given memory address, the detecting step including issuing a request to the requesting node that requests the requesting node to indicate to the master node whether the requesting node currently holds a data instance at the given memory address; and selectively performing the exclusive store operation in response to the detection.

[0068] Another example embodiment provides a circuit that includes: a set of two or more data processing nodes, each data processing node having a respective storage circuit for maintaining data; and a master node for serializing execution of data access operations and controlling consistency among data maintained by one or more data processing nodes such that data written to a memory address is consistent with data read from the memory address in response to a subsequent access request; wherein: a requesting node of the set of data processing nodes is configured to initiate an exclusive request transaction for a given memory address; and the master node is configured to, in response to the exclusive request transaction, detect whether the requesting node stores a data instance at the given memory address when executing the exclusive request transaction, and when it is detected that the requesting node stores a data instance at the given memory address, instruct other data processing nodes in the set of data processing nodes to invalidate any further data instances at the given memory address.

[0069] In an example embodiment, the master node may instruct other nodes to invalidate their own copies to provide a unique copy of a memory address, such as a cache line, at a node.

[0070] In an example, when it is detected that the requesting node does not store a data instance at the given memory address, the master node is configured not to instruct other data processing nodes in the set of data processing nodes to invalidate any further instances of data at the given memory address.

[0071] In an example, when it is detected that the requesting node does not store a data instance at the given memory address, the master node is configured not to instruct other data processing nodes in the set of data processing nodes to invalidate any additional data instances at the given memory address.

[0072] In some examples, to detect whether a requesting node currently holds a data instance at a given memory address, a request is issued to the requesting node for the requesting node to indicate to the master node whether the requesting node currently holds a data instance at the given memory address. In other examples, the master node may refer to snooping data that indicates data instances at one or more data processing nodes in a set of data processing nodes, particularly in the case of using so-called exact snooping data, that is, the snooping data is configured to indicate data instances at each node in the set of data processing nodes; and the master node is configured to, when executing an exclusive request transaction, detect whether the requesting node stores a data instance at the given memory address by referring to the snooping data.

[0073] Another example embodiment provides a method that includes: maintaining data by a set of two or more data processing nodes; serializing data access operations by a master node; controlling consistency among data maintained by one or more data processing nodes by the master node such that data written to a memory address is consistent with data read from the memory address in response to a subsequent access request; initiating, by a requesting node in the set of data processing nodes, an exclusive request transaction for a given memory address; and by the master node, when executing the exclusive request transaction in response to the exclusive request transaction, detecting whether the requesting node stores a data instance at the given memory address; and by the master node, when detecting that the requesting node stores a data instance at the given memory address, instructing other data processing nodes in the set of data processing nodes to invalidate any further data instances at the given memory address.

[0074] Another example embodiment provides a snooping query or a similar transaction (and / or circuitry for issuing and / or processing such transactions) that returns the current state (such as the coherence state) of a cache line or other data item, but does not return the data or change the state of the cache line or other data item.

[0075] Circuit Overview

[0076] Now referring to the drawings, Figure 1 schematically shows a circuit embodied as a network of devices interconnected by an interconnect 100. The device may be provided as a single integrated circuit, such as a so-called system-on-chip (SoC) or network-on-chip (NoC), or as a plurality of interconnected discrete devices.

[0077] Various so-called nodes are connected via interconnect 100. These nodes include: one or more host nodes (HN) 110 that supervise data consistency within the networking system; one or more slave nodes (SN), such as a higher-level cache memory 120 (the reference to "higher-level" is with respect to the cache memory provided by the requesting node and is described below); main memory 130 and peripheral devices 140. Figure 1 The selection of the slave nodes shown is by way of example, and zero or more slave nodes of each type may be provided.

[0078] In other examples, the functionality of the HN may be provided by the HN circuit 112 of the interconnect 100. For this reason, both the HN 110 and the HN circuit 112 are shown in dashed lines; generally, a single HN is provided for a particular memory region to supervise consistency among the various nodes, but whether the HN functionality is implemented at the interconnect or elsewhere is a matter of design choice. The memory space may be divided among multiple HNs.

[0079] Figure 1 Also shown are a plurality of so-called request nodes (RN) 160, 170, 180 that operate according to the CHI (Coherence Hub Interface) protocol.

[0080] The nodes may be fully coherent or input / output (I / O) coherent. A fully coherent HN or RN (HN-F, RN-F respectively) includes a coherence cache storage device. A fully coherent SN (SN-F) is paired with an HN-F. The HN-F may manage the coherence of a memory region. In this example, the RNs 160 - 180 are fully coherent RNs (RN-F), each RN having an associated cache memory 162, 172, 182 as an example of a storage circuit for holding data at that node.

[0081] In an example arrangement, each of one or more slave nodes may be configured to accept each data transfer to that slave node independent of any other data transfer to that slave node.

[0082] Thus, Figure 1 the arrangement provides an example of a circuit that includes a set of two or more data processing nodes 160, 170, 180, each processing node having a respective storage circuit 162, 172, 182 for holding data; and a host node 110 / 112 for serializing data access operations and controlling the consistency among the data held by one or more data processing nodes such that data written to a memory address is consistent with data read from that memory address in response to a subsequent access request.

[0083] Transactions may be used in Figure 1process data in the arrangement. For any particular transaction, the node that initiates the transaction (e.g., by sending a request message) is referred to as the request node or master node of that transaction. Thus, although Figure 1 shows multiple RNs, for any particular transaction, one of these RNs will be the RN for that transaction. The HN receives request messages from the relevant RNs and processes the execution of the transaction while maintaining consistency among the data held by the corresponding nodes.

[0084] master node

[0085] Figure 2 schematically shows the operation of the master node 110 in its function as a cache coherence controller including a snooping filter.

[0086] The term "snooping filter" is a historical term and is used herein to refer to (e.g.) a control device that may have an associated "directory" that stores information indicating which data is stored in which cache, and the snooping filter itself at least helps to process data access to the cache information, thereby providing a cache coherence function.

[0087] In this embodiment, the cache function controller includes a snooping filter. The snooping filter can provide some or all of the functions related to supervising data access processing on the cache-related system.

[0088] In Figure 2 , the snooping filter 200 having the directory 210 as described above is associated with the controller 205 and the transaction router 220. The transaction router 220 communicates data with one or more RNs having cache memories. Each RN may have an associated agent ( Figure 2 not shown in ) that is responsible for local processing of data read and write operations with respect to that cache memory. The HN itself may have a cache memory 230 under the supervision of the HN (for consistency purposes).

[0089] The snooping filter 200 processes at least a part of the process under which, when Figure 1 any RN of aims to access or modify data stored as a cache line in any cache memory, that node obtains permission to do so. As part of this process, the snooping filter 200 checks whether any other cache memory has a copy of the cache line to be modified. If any other copies exist in other cache memories, these copies need to be cleared and invalidated. If these copies themselves contain modifications to the data stored in the line, then at least in some cases, the controller 205 (or the snooping filter 200) instructs the cache memory to write the line back to the main memory.

[0090] In the case where a node makes a read access to data stored in a cache memory, it is important that the RN requesting the read can access the latest correct version of the cached data. The controller 205 supervises this process so that if another cache has a more recently modified version of the required data, that other cache writes back the modified version and / or forwards a copy of the modified version for caching at the current requesting node.

[0091] The snooping operation or query can also be initiated by Figure 2 the circuitry. This may involve sending a message to the memory address indicated by the directory as being cached by another cache memory and receiving a response from the cache memory that received the message, the response indicating whether that cache memory is actually caching that memory address. Example snooping operations are discussed further below.

[0092] In Figure 1 many practical examples of data processing systems of the type shown, almost all of the checks performed by the snooping filter 200 are likely to miss, that is, they will not expose data replication between multiple caches. However, the checks performed by the snooping filter 200 are necessary for maintaining cache coherence. To improve the efficiency of the process and allow the snooping filter 200 to avoid performing checks that will definitely miss, the snooping filter 200 maintains a directory 210 to indicate to the snooping filter 200 at which cache(s) which data is stored. In some examples, this can allow the snooping filter 200 to reduce the number of snooping operations (by avoiding performing operations in cases where a particular line is not saved in any cache or is only held in the cache corresponding to the node currently accessing that data). In an example, it can also allow data communication associated with snooping operations to be better targeted at the appropriate cache(s) (e.g., unicast or multicast communication) rather than being broadcast to all caches as a snooping operation.

[0093] Accordingly, when a potential snooping operation is initiated, the snooping filter 200 can query the directory 210 to detect whether the information in question is held in one or more caches. If a snooping operation is indeed required to query the current state of data at one or more caches, the snooping filter 200 can perform that query as an appropriate unicast or multicast communication (rather than a broadcast communication).

[0094] In the design of the snooping filter directory 210, various levels of so-called "precision" are available.

[0095] In some examples, an "exact" snooping filter is used, that is, the snooping filter 200 is a so-called inclusive snooping filter, which means that it has a continuous requirement to maintain a complete list of all data held by all caches subject to cache coherence constraints. To do so, the snooping filter 200 (as part of the HN 110 / 112) needs to be notified by an agent associated with the cache memory that a cache insertion has occurred. But to perform this function effectively, the agent should also notify whether the cache line has been evicted from the cache memory (removed from the cache memory), either as a simple deletion or invalidation (in the case of unmodified data), or as a write-back to the main memory (for modified data). Issuing an eviction signal is the responsibility of the corresponding coherence agent. However, some operation protocols associated with multiple caches may suggest issuing an eviction signal but do not necessarily enforce it. In any case, there may be situations where, due to software errors (e.g., in a virtual memory system) or memory errors, the cache memory may not always be able to issue all eviction signals.

[0096] Therefore, from the perspective of processing requirements, directly maintaining an exact snooping filter can be very burdensome. In a large system with many distributed cache memories, the storage size required to provide an exact snooping filter directory is also very important.

[0097] Therefore, in other examples, different levels of imprecision can be used for the snooping filter while still providing some useful functionality for snooping filter operation.

[0098] For example, the snooping filter can maintain a complete record of writes to the cache memory but may not maintain a complete record of cache evictions.

[0099] In other examples, if a cache line is held by no more than a predetermined number of RNs, the snooping filter can indicate the location of the cache line; if the cache line is held by more than that number of RNs, instead of recording the exact location of each data instance, a flag is recorded to indicate that the line is held at (or at least has been stored to) multiple locations.

[0100] It should be noted that in the case of a snooping filter outside the "exact" snooping filter directory, the snooping filter cannot generate a definite answer about where the cache line is currently stored and its state at that location simply by using the information maintained and held at the snooping filter.

[0101] Cache Memory

[0102] Figure 3 As schematically shown, various aspects of the cache memory are, for example, the local cache memory 230 at the HN as described above or one of the cache memories at the RN.

[0103] The controller 300 (acting as an agent as described above) controls writes, reads, and evictions to the memory storage device 310. Associated with each cache line 305 stored in the storage device 310 is a status indication (drawn horizontally to the right of the relevant cache line for illustrative purposes only). The controller 300 can appropriately change and / or report the status of the cache line based on locally executed operations and / or in response to instructions received from the HN. Note that a particular cache memory structure (e.g., set-associative structure) is irrelevant to this discussion.

[0104] Example statuses include invalid (I); unique clean (UC), indicating that this is the only held copy that does not currently need to be written back to main memory; unique dirty (UD) indicating that this is the only held copy that is different from the copy held in main memory and thus needs to be written back to main memory at some point; shared clean (SC), indicating a clean shared copy (held in multiple cache memories); and shared dirty (SD), indicating a shared copy that needs to be written back to main memory at some point. Note that the "dirty" indication does not necessarily mean that the copy is different from main memory, but rather that the RN holding it is responsible for writing it back to memory. Also note that in the case of shared copies, in at least some example protocols, only one of these example protocols is marked as "dirty". Other shared copies can coexist with the SD copy but are in the SC state.

[0105] Exclusive transaction

[0106] The so-called exclusive memory transaction and exclusive sequence will now be described.

[0107] The exclusive sequence formed by an exclusive memory transaction does not prevent other accesses to the memory address or (one or more) cache lines, but allows detection of whether an intervening access has occurred, in which case the exclusive transaction aborts.

[0108] In Figure 4a an overview of the exclusive sequence is schematically shown. The process starts with an exclusive load (sometimes abbreviated as LDREX) 405. At this stage, if the relevant RN does not already have a copy of the data item it needs (e.g., a cache line), it can use Figure 7The techniques in [description] obtain a copy. Among those techniques to be described below, the RN would prefer to obtain a unique copy, but a shared copy can be used (anyway at this stage). At this stage, it is important that, when attempting to obtain a unique copy, the RN does not invalidate the copy of another node if another node has partially passed the exclusive sequence of the line itself. If the RN already has a shared copy, no further action is required in obtaining a copy of the line.

[0109] Then, some processing 415 is performed on the relevant line. The amount of time taken for this processing is not bounded and can be variable. In other words, there is no predefined interval between the exclusive load 405 and the subsequent exclusive store 425.

[0110] If the line is invalidated at the RN between steps 405 and 425, the sequence will fail and may need to be retried.

[0111] At the exclusive store 425, it is necessary to transition to the unique state of the cache line at the RN. Then the exclusive store is implemented locally, thus retaining a unique dirty line at the RN.

[0112] Two other techniques to be described below are relevant here. Figure 12 The techniques in [description] solve a problem if the RN loses the cache line (it will be invalidated) between initiating an exclusive store transaction and the operation being issued for execution by the HN. The Figure 13 The techniques in [description] enhance this approach by returning the data in the shared state if the line has been invalidated.

[0113] As Figure 4b shown, exclusive memory access can be controlled by a so-called exclusive monitor. The exclusive monitor can be considered a simple state machine with only two possible states: "open" and "exclusive".

[0114] By setting the exclusive monitor and then checking its state, a memory transaction can be able to detect whether any other intervening action can access one or more memory addresses covered by the exclusive monitor. In a distributed system such as Figure 1 shown, each RN 400 can have a corresponding "local exclusive monitor" (LEM) 410 associated with the corresponding RN (or actually associated with the processing core within the RN). In some example arrangements, the LEM can be integrated with the load-store unit of the associated core.

[0115] The global exclusive monitor or GEM 420 can be associated with multiple nodes and can be provided, for example, at the master node 430 to track the exclusivity of multiple potential addresses from multiple potential processing elements.

[0116] Some example arrangements employ both local and global exclusive monitors simultaneously.

[0117] In operation, as described above, the exclusive monitor can act as two state machines and can move between the above-mentioned states.

[0118] Figure 5 and Figure 6 A schematic flowchart is provided that shows an example method that can use exclusive monitoring.

[0119] Refer to Figure 5 , to initiate a so-called exclusive load operation, at step 500, one or more exclusive monitors related to the load operation are set (whether LEM or GEM or both), that is, their states are moved to the "exclusive" state or remain in the "exclusive" state. Then, at step 510, the exclusive load is executed.

[0120] Refer to Figure 6 , to perform the corresponding exclusive store operation, at step 600, one or more exclusive monitors related to the exclusive store are checked. If they are still in the "exclusive" state (as detected in step 610), the control is passed to step 620 where the store is executed, and at step 630, the process ends successfully.

[0121] If the answer is no in step 610, the process is aborted at step 640.

[0122] Therefore, using the exclusive monitor allows detection of whether the intermediate process has written back to the relevant address. In this case, the exclusive store itself is aborted.

[0123] The basic operating principle of this type of exclusive access is that when multiple agents compete for exclusive access, the one that succeeds is the "first to store". This characteristic stems from the following observation: an exclusive store following an exclusive load is not necessarily an atomic operation, and in fact, any number of instructions are allowed between the load and store parts.

[0124] If multiple agents start exclusive sequences in an interleaved manner and each start of a sequence blocks the completion of another exclusive sequence, potential problems may occur. In this case, a so-called "livelock" may occur, such that no sequence reaches completion and all agents must ultimately repeat their respective sequences.

[0125] Example 1: Exclusive Load

[0126] In reference Figure 1In a hardware coherence system of the described type, performance may be improved if a particular RN can obtain a data item (e.g., a cache line) in a unique state when starting to execute an exclusive sequence, as this allows an exclusive store operation to be performed locally without the need to issue a transaction to the home node and involves the coherence control arrangement discussed above.

[0127] Reference Figure 7 provides a time-based diagram where time generally proceeds down the page as drawn, and operations are indicated between a requesting node 700, a home node 710, and other RNs 720.

[0128] The RN 700 issues a request here called "Read Preferred Unique" (RPU) 730 to obtain a copy of the cache line in question. This indicates to the HN that, for performance reasons, the RN 700 would prefer to have the particular cache line in a unique state, but the RN will accept the cache line in a shared state if required to satisfy the demands of other exclusive access sequences. The RN 700 issues the RPU as a transaction to the HN 710.

[0129] In this example, a listen-based exclusive contention detection is used to perform the determination Figure 1 as to whether other agents in the overall system are mid-way through their own exclusive sequences. The home node issues a listen preferred unique 740 to the other RNs compared to the other RNs. This query has various functions. It indicates that if the particular cache line specified by the query is not currently in use in an exclusive sequence, then that cache line should be invalidated. However, if the line is currently in use in an exclusive sequence, the receiving RN should retain its copy in response to the listen preferred unique 740. More details will be provided below.

[0130] Thus, step 740 is an example of a detection in which the home node is configured to control whether each of the other data processing nodes is currently performing an exclusive access operation to a given memory address.

[0131] Figure 7 Also shown is a response 750 from another RN to the HN, and a response 760 from the HN to the RN 700.

[0132] In terms of the operation at the home node 710, reference Figure 8Schematic flowchart, where, in response to receiving RPU transaction 730, HN initiates listening preference unique 740 at step 800. At step 810, HN also detects whether the relevant cache line is held at the original RN 700 (although in some examples this test is not required because reading the preference unique context can actually imply that the original RN has no copy). If not, then HN 710 sends a copy to the original RN at step 820 (in this step, when the requesting node does not hold a data instance at a given memory address, the master node is configured to initiate the transfer of the data at the given memory address to the requesting node), and then in any case, HN sends response 760 at step 830.

[0133] Figure 9 Schematically shows the operation of each other RN that receives listening preference unique 740 at step 900.

[0134] At step 910, if the relevant line specified by listening preference unique 740 is held at this receiving RN, then control passes to step 920, where, for example, using a local exclusive monitor associated with the cache line, it is detected whether the address is included in the exclusive sequence of this receiving RN. If the answer is "yes", then at step 930, the local copy is retained and its state is set to "shared". Control then passes to step 940, where response 750 indicating the line state is returned to HN 710.

[0135] Return to step 910. If the relevant line is not held at the receiving RN, then at step 940, a simple response indicating "not held" or "invalid" is provided.

[0136] Return to step 920. If the line is held but not in the exclusive sequence, then at step 950, the local copy is invalidated, and control passes to step 940, where the status ("invalid") is returned as response 750.

[0137] It is always legal for the receiving RN to respond by indicating that it has retained the relevant line and moved to the shared state. If in the process of the exclusive sequence, the receiving RN does not invalidate its own line copy.

[0138] Figure 10 Relates to the operation at RN 700, which issues RPU transaction 730 at step 1000. When response 760 is received, at step 1010, if the response indicates that the line is "shared", then control passes to step 1015A, where sequence-related processing is performed. At step 1020, RN 700 continues to initiate an exclusive transaction via HN 710.

[0139] If the result of step 1010 is negative, control passes to step 1015B, where sequence-related processing is performed. Note that "A" and "B" are used in this way to denote steps 1015A / B to indicate that the processing (not specified in this discussion but could correspond to Figure 4a the processing 415 generally shown in Figure 10 can be the same in either path of

[0140] At step 1025, RN 700 checks again whether the line is still unique (e.g., by using another RPU transaction as described above). If the result is "yes", control passes to step 1030, where RN 700 locally initiates an exclusive sequence using its own unique copy. If the result is "no", control passes to step 1020 above.

[0141] Thus, in these examples, the requesting node 700 of the set of data processing nodes is configured to send a request to the master node for exclusive access to a given data instance at a given memory address (e.g., which can be held by this RN 700); and the master node 710 is configured to send information to the other data processing nodes of the set of data processing nodes in response to the request to control the processing of any further data instances at the given memory address held by these other data processing nodes. When it is detected that one of the other data processing nodes is currently performing an exclusive access operation on the given memory address, one of the other data processing nodes is configured to retain the data instance at the given memory address and indicate to the master node that the given memory address has a shared state between one of the other data processing nodes and the requesting node. When it is detected that one of the other data processing nodes is not currently performing an exclusive access operation on the given memory address, one of the other data processing nodes is configured to invalidate the data instance at the given memory address held at one of the other data processing nodes.

[0142] In a variant of these embodiments, when the snooping filter determines that there is no cache copy, so-called direct memory transfer ("DMT") can be used. This will always return the line in a unique state.

[0143] So-called direct cache transfer ("DCT") is typically only used when the HN can determine that there is only a single cache copy present. Here, a forward snoop requesting a unique copy is only used when the HN can determine that the snoop needs to be sent to a single cache.

[0144] Example 2: Avoiding livelock

[0145] In a further example, note that the exclusive access assumes that when the host or the RN is executing an exclusive sequence, for the store part of the sequence to succeed, the host should check that no other host can perform a store to the relevant line during the entire exclusive load-exclusive store sequence. In some examples, this can be checked by confirming that the line remains allocated within the host at the point where it issues the exclusive store.

[0146] However, although this can solve the situation where another host performs a store from the moment the host completes its exclusive load to the moment it issues the exclusive store, it does not necessarily solve the situation where another host can perform a store between the point where the exclusive store is issued and the point where the exclusive store is scheduled or serialized by the HN.

[0147] If the host can invalidate its cache line by another host while its exclusive store is in progress (resulting in the exclusive store failing), a livelock situation may occur if the transaction can still continue and invalidate the cache line in other caches. In principle, this can happen on multiple hosts, resulting in none of the attempted exclusive transactions succeeding.

[0148] To solve this problem, a snoop query transaction is proposed, which is a snoop operation that returns the current state of the cache line, without the cache line returning data or changing the state of the cache line.

[0149] A potential purpose of the snoop query is to establish the state of the cache line when the snoop filter itself has a certain degree of imprecision (e.g., for any of the reasons discussed above). The snoop query transaction can be used in association with exclusive access because it enables the HN to determine that at the point where the transaction associated with the exclusive store is associated at the master node, the thread with exclusive capabilities still has a copy of the cache line.

[0150] Figure 11 A schematic example of a snoop query 1100 issued by the HN to the RN is provided, which results in the return 1110 of the state of the cache line specified by the snoop query, but no other changes are made to the state or content of the cache line.

[0151] In an example using such a query, refer to Figure 12 the timeline representation.

[0152] The RN 1200 initiates 1220 an exclusive store operation. This requires the HN 1210 to schedule it among other data access operations supervised by the HN 1210.

[0153] As a precursor to a scheduling operation, the HN issues the listen query 1100 to the originating RN as described above, and the originating RN responds 1240 with the state 1110 of the cache line at the RN, confirming that the line remains allocated at the RN 1200 despite any time gap between the initiation of the exclusive operation 1220 and the current time. In response, the HN 1210 initiates the execution 1250 of the exclusive operation. In other words, the listen query 230 does not change the state of the line, but only checks whether the copy is still being held.

[0154] In these examples, the requesting node 1200 of the set of data processing nodes is configured to initiate an operation sequence that requires exclusive access to data at a given memory address, the operation sequence including an exclusive store operation to the given memory address; and the master node 1210 is configured to, prior to executing the exclusive store operation, detect whether the requesting node currently holds a data instance at the given memory address by issuing a request to the requesting node, and selectively execute the exclusive store operation in response to the detection, the request for the requesting node to indicate to the master node whether the requesting node currently holds a data instance at the given memory address.

[0155] In step 1250 above, the master node is configured to execute the exclusive store operation when the master node detects that the requesting node currently holds a data instance at the given memory address.

[0156] Example 3: "Making Reads Uniquely Exclusive"

[0157] Figure 13 is a schematic timeline showing the operations between the RN 1300, the HN 1310, and other RNs 1320 for a so-called Making Reads Uniquely Exclusive (MRUX) transaction 1330.

[0158] The effect of the MRUX transaction is to invalidate other copies of the line or other data item held at the RN 1300. However, in this example, these actions are only performed if the originating RN 1300 still has a copy of its associated line when an invalidate is issued to the other RNs.

[0159] The HN - 1310 can detect whether the RN1300 has retained a copy of the line through various techniques. One is to use the listen as discussed above Figure 11 discussed. Another approach is to use a precise listen filter at the HN1310. In Figure 13In it, the use of snoop 1340 and its response 1350 are shown as examples by dashed lines so that HN 1310 can detect that RN 1300 still has a copy of its own relevant cache line. Assuming this is the case, then in stage 1360, HN 1310 instructs other RNs 1320 to invalidate the copies of their own cache lines. If not, then HN 1310 returns 1370 a shared copy of the relevant cache line. Note that Figure 13 Step 1370 in simplifies the process required for HN 1310 to obtain and return the shared copy, but it is sufficient to illustrate the nature of the operations performed by HN 1310.

[0160] Also note that steps 1360 and 1370 are mutually exclusive; only one of them is allowed to occur.

[0161] In response to step 1360, one or more other RNs 1320 provide response 1380. This allows HN 1310 to provide a "return unique" indication 1390 to RN1300, indicating that the copy held at the RN is now unique.

[0162] Referring to Figure 14 the flowchart in, which shows the operations of HN 1310. At step 1400, HN receives the MRUX transaction 1330. At step 1410, HN detects whether the copy still remains at the original RN 1300, for example, by using snoop and / or by using the content of the snoop filter directory. If the answer is "yes", then at step 1420, HN 1310 initiates the invalidation of other copies of the cache line, but if the answer is "no", then at step 1430, HN initiates the return of a shared copy of the cache line to RN 1300.

[0163] In these examples, the requesting node 1300 of the set of data processing nodes is configured to initiate an exclusive request transaction for a given memory address; and the master node 1310 is configured to detect, in response to the exclusive request transaction, that the requesting node is storing a data instance at the given memory address when executing the exclusive request transaction, and when detecting that the requesting node is storing a data instance at the given memory address, to instruct other data processing nodes 1320 of the set of data processing nodes to invalidate any further data instances at the given memory address.

[0164] When it is detected that the requesting node is not storing a data instance at the given memory address, the master node is configured not to instruct other data processing nodes of the set of data processing nodes to invalidate any further data instances at the given memory address.

[0165] As described above, in the example, either a listen (in which case the master node is configured to detect whether a requesting node currently holds a data instance at a given memory address by issuing a request to the requesting node, the request being for the requesting node to indicate to the master node whether the requesting node currently holds a data instance at the given memory address) or an exact listen filter directory (in which case the master node is configured to maintain listen data indicating data instances stored in a collection of one or more data, the listen data being configured to indicate data instances at each node in a collection of data processing nodes; and the master node is configured to, when executing an exclusive request transaction, detect whether a requesting node stores a data instance at a given memory address by referring to the listen data) can be used.

[0166] Example method

[0167] Figure 15 is a schematic flow chart showing a method that includes:

[0168] (at step 1500) data is held by a collection of two or more data processing nodes;

[0169] (at step 1510) the master node serializes data access operations;

[0170] (at step 1520) the master node controls the consistency between data held by one or more data processing nodes such that data written to a memory address is consistent with data read from that memory address in response to a subsequent access request;

[0171] (at step 1530) a requesting node in the collection of data processing transmits a request to the master node for exclusive access to a given data instance at a given memory address; and

[0172] (at step 1540) in response to the request, the master node transmits information to other data processing nodes in the collection of data processing nodes to control the processing by those other data processing nodes of any further data instances held by those other data processing nodes at the given memory address.

[0173] Figure 16 is a schematic flow chart showing a method that includes:

[0174] (at step 1600) data is held by a collection of two or more data processing nodes;

[0175] (at step 1610) the master node serializes data access operations;

[0176] (At step 1620), the master node controls the consistency among data maintained by one or more data processing nodes, such that data written to a memory address is consistent with data read from that memory address in response to a subsequent access request;

[0177] (At step 1630), a requesting node in a set of data processing nodes initiates an operation sequence that requests exclusive access to data at a given memory address, the operation sequence including an exclusive store operation on the given memory address;

[0178] (At step 1640), the master node detects, before performing the exclusive store operation, whether the requesting node currently holds a data instance at the given memory address, the detection step including sending a request to the requesting node that requests the requesting node to indicate to the master node whether the requesting node currently holds a data instance at the given memory address; and

[0179] Selectively perform the exclusive store operation in response to the detection.

[0180] Figure 17 is a schematic flow chart showing a method that includes:

[0181] (At step 1700), data is maintained by a set of two or more data processing nodes;

[0182] (At step 1710), the master node serializes data access operations;

[0183] (At step 1720), the master node controls the consistency among data maintained by one or more data processing nodes, such that data written to a memory address is consistent with data read from that memory address in response to a subsequent access request;

[0184] (At step 1730), a requesting node in a set of data processing nodes initiates an exclusive request transaction for a given memory address; and

[0185] (At step 1740), the master node detects, in response to the exclusive request transaction, whether the requesting node stores a data instance at the given memory address when performing the exclusive request transaction; and

[0186] (At step 1750), when it is detected that the requesting node stores a data instance at the given memory address, the master node instructs other data processing nodes in the set of data processing nodes to invalidate any further data instances at the given memory address.

[0187] In this application, the phrase "configured to..." means that an element of a device has a configuration capable of performing the defined operation. In this context, "configuration" refers to the arrangement or manner of hardware or software interconnection. For example, a device may have dedicated hardware that provides the defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to" does not imply that the device element needs to be altered in any way to provide the defined operation.

[0188] Although illustrative embodiments of the present technology have been described in detail herein with reference to the accompanying drawings, it should be understood that the present technology is not limited to those exact embodiments, and that various changes, additions, and modifications may be made by those skilled in the art without departing from the scope and spirit of the technology as defined by the appended claims. For example, various combinations of the features of the dependent claims may be made with the features of the independent claims without departing from the scope of the present technology.

Claims

1. A circuit, comprising: A set of two or more data processing nodes, each data processing node having a corresponding storage circuit for holding data; And A master node, the master node being configured to serialize data access operations and control the consistency between data held by one or more data processing nodes, such that data written to a memory address is consistent with data read from that memory address in response to subsequent access requests; Wherein: A requesting node of the set of data processing nodes is configured to transmit a request to the master node to perform an exclusive access to a given data instance at a given memory address; and The master node is configured to transmit information to other data processing nodes of the set of data processing nodes in response to the request, to control the processing by those other data processing nodes of any further data instances held by those other data processing nodes at the given memory address, Wherein the master node is configured to control the detection of whether each of the other data processing nodes is currently performing an exclusive access operation to the given memory address, and Wherein, when it is detected that one of the other data processing nodes is currently performing an exclusive access operation to the given memory address, this one of the other data processing nodes is configured to hold a data instance at the given memory address and indicate to the master node that the given memory address has a shared state between this one of the other data processing nodes and the requesting node.

2. The circuit according to claim 1, wherein, The master node is configured to initiate the transmission of data at the given memory address to the requesting node.

3. The circuit according to claim 1, wherein, When it is detected that one of the other data processing nodes is not currently performing an exclusive access operation to the given memory address, one of the other data processing nodes is configured to invalidate the data instance at the given memory address held by this one of the other data processing nodes.

4. The circuit according to claim 1, wherein, The master node is configured to maintain monitoring data indicating data instances stored at one or more data processing nodes in the set of data processing nodes.

5. A circuit, comprising: A set of two or more data processing nodes, each data processing node having a corresponding storage circuit for holding data; And A master node, the master node being configured to serialize the execution of data access operations and control the consistency between data held by one or more data processing nodes, such that data written to a memory address is consistent with data read from that memory address in response to subsequent access requests; Wherein: A requesting node of the set of data processing nodes is configured to initiate an operation sequence that requests exclusive access to data at a given memory address, the operation sequence including an exclusive store operation to the given memory address. The master node is configured to: before performing the exclusive storage operation, detect whether the requesting node currently holds a data instance at the given memory address by sending a request to the requesting node, and selectively perform the exclusive storage operation in response to the detection, the request being for the requesting node to indicate to the master node whether the requesting node currently holds a data instance at the given memory address; and The master node is configured to perform the exclusive storage operation when the master node detects that the requesting node currently holds a data instance at the given memory address.

6. The circuit according to claim 5, wherein, The master node is configured to maintain control data that at least partially indicates which data is currently held by the storage circuit of the data processing node.

7. The circuit according to claim 5, wherein, The operation sequence includes applying an exclusive tag for at least the given memory address and then releasing the exclusive tag.

8. A circuit, comprising: A set of two or more data processing nodes, each data processing node having a corresponding storage circuit for holding data; And A master node for serializing the execution of data access operations and controlling the consistency between data held by one or more data processing nodes, such that data written to a memory address is consistent with data read from that memory address in response to subsequent access requests; Wherein: The requesting node of the set of data processing nodes is configured to initiate an exclusive request transaction for a given memory address; and The master node is configured to: detect whether the requesting node stores a data instance at the given memory address when executing the exclusive request transaction in response to the exclusive request transaction, and when it is detected that the requesting node stores a data instance at the given memory address, instruct other data processing nodes in the set of data processing nodes to invalidate any further data instances at the given memory address; The master node is configured to detect whether the requesting node currently holds a data instance at the given memory address by sending a request to the requesting node, the request being for the requesting node to indicate to the master node whether the requesting node currently holds a data instance at the given memory address.

9. The circuit according to claim 8, wherein, When it is detected that the requesting node does not store a data instance at the given memory address, the master node is configured not to instruct other data processing nodes in the set of data processing nodes to invalidate any further data instances at the given memory address.

10. The circuit according to claim 8, wherein, The master node is configured to maintain snooping data that indicates data instances stored at one or more data processing nodes in the set of data processing nodes.

11. The circuit according to claim 10, wherein: The snooping data is configured to indicate data instances stored at each data processing node in the set of data processing nodes; and The master node is configured to: when executing the exclusive request transaction, detect whether the requesting node stores a data instance at the given memory address by referring to the listening data.

Citation Information

Patent Citations

  • Data processing

    CN110235113A

  • Method and system for cache coherence in DSM multiprocessor system without growth of the sharing vector

    US20030163543A1