A data access method, interworking system and device
Patent Information
- Application Number
- CN202180098267.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-20
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-05-20
AI Technical Summary
[0057]以上第二方面到第五方面的有益效果,请参见第一方面的有益效果,不重复赘述。
Smart Images

Figure CN117355823B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data storage technology, and in particular to a data access method, interconnection system and apparatus. Background Technology
[0002] To meet increasing system performance demands, performance goals can be achieved by scaling the system. Common system scaling methods include vertical scaling (also known as scale-up) and horizontal scaling (also known as scale-out). Both methods scale the system on a node-by-node basis. A node typically includes processing power, storage capacity, and communication capabilities; a single chip is a typical node.
[0003] Scale-up primarily expands upon the existing infrastructure of the node itself, such as increasing the number of processors on a chip to improve computing power. Figure 1 As shown by the vertical expansion arrows, CPU2 and CPU3 can be added to node 0, which currently only contains CPU 0 and CPU 1, to increase the number of processors; the memory capacity can also be increased to improve storage capabilities, such as... Figure 1 As shown by the vertical scaling arrows, the memory in node 0 can be expanded from 2MB to 4MB; input / output (I / O) interfaces can also be optimized to improve interaction bandwidth, etc. Horizontal scaling (scale-out) expands the system size by increasing the number of nodes, such as... Figure 1 As indicated by the horizontal scaling arrows, the number of nodes increases from two (nodes 0 and 1 in the diagram) to four (nodes 0-3 in the diagram). While maintaining efficient communication between nodes, increasing the number of nodes translates to improved overall system performance. Typical systems employing horizontal scaling (scale-out) include multi-chip interconnect architectures in high-performance computing (HPC). Compared to vertical scaling (scale-up), horizontal scaling (scale-out) offers greater flexibility in system configuration, allowing for the addition or removal of nodes at any time.
[0004] In systems employing a scale-out approach, the efficiency of interactions between nodes significantly impacts system performance. To improve performance, nodes in the system utilize multi-path networking. While multi-path networking can enhance interconnection efficiency, ensuring the ordered processing of data access request messages that require a specific order of execution results while minimizing overhead remains a challenge. Summary of the Invention
[0005] This application provides a data access method, interconnect system, and apparatus for controlling the visible order of execution results of message sequences with minimal overhead for interconnect systems interconnected in a horizontally scaled manner.
[0006] Firstly, a data access method is provided, applied to an interconnected system, the interconnected system comprising at least two nodes interconnected in a horizontal scaling manner, the at least two nodes interconnected in a horizontal scaling manner comprising a first node and a second node, the method comprising:
[0007] The first node receives a message sequence from an external node. The messages in the message sequence are used to request data access to a target address. The target address belongs to the storage space managed by the second node. The external node is a node outside the interconnection system. The external node has stricter constraints on the visible order of execution results than the first node.
[0008] The first node obtains exclusive processing rights for the target address;
[0009] The first node performs data access operations on the target address based on its exclusive processing permissions for the target address, wherein the visible order of the execution results of the data access operations satisfies the execution result order constraint followed by the external node.
[0010] In the above implementation, after receiving a message sequence from an external node (i.e., a node outside the interconnected system), the first node obtains exclusive processing rights (i.e., E-state in storage consistency) for the corresponding target address (i.e., the target address corresponding to each message in the message sequence) from the second node. This allows the exclusive processing rights for the target address to be transferred from the second node to the first node. Since a node possesses exclusive processing rights for the target address corresponding to the message sequence, it means that the node has the ability to control the visible order of the execution results for that message sequence. Therefore, the first node can control the visible order of the execution results of the message sequence. Because the above implementation does not require serial communication, sequence number-based interaction, or sequence number-based reordering mechanisms, the visible order of execution results can be controlled with minimal overhead.
[0011] In one possible implementation, the first node acquiring exclusive processing rights to the target address includes: the first node acquiring exclusive processing rights to the target address from the second node in response to receiving the message sequence.
[0012] In the above implementation, when the first node receives the message sequence, it obtains the exclusive processing permission of the corresponding address based on the target address of the message sequence. This allows for targeted acquisition of exclusive processing permission for the corresponding address, thereby reducing system overhead.
[0013] In one possible implementation, the first node, in response to receiving the message sequence, obtains exclusive processing rights for the target address from the second node, including: after receiving the message sequence, the first node sends a permission acquisition request to the second node, the permission acquisition request carrying a target address corresponding to at least one message in the message sequence; the first node receives a permission acquisition response from the second node, the permission acquisition response indicating that the exclusive processing rights for the target address corresponding to the at least one message be transferred to the first node.
[0014] In the above implementation, the permission acquisition process acquires exclusive processing permissions at a granularity of "at least one target address corresponding to a message." This means the permission acquisition request can carry the target addresses corresponding to the entire message sequence and can exceed the cache line size. In other words, compared to the cache line size, this implementation can use a larger granularity to acquire exclusive processing permissions for target addresses. Since the permission acquisition request can carry the target addresses corresponding to more than one message, or even the target addresses corresponding to all messages in the entire message sequence, system overhead can be reduced and permission acquisition efficiency improved.
[0015] In one possible implementation, after receiving the message sequence, the first node sends an access request to the second node, including: after receiving the message sequence, the first node sends the access request through a first path between itself and the second node, wherein the first path is selected by the first node based on the congestion status of each path between itself and the second node; the first node receives the access response from a second path between itself and the second node, wherein the second path may be the same as or different from the first path.
[0016] In the above implementation, on the one hand, since the first node can select an appropriate path to transmit permission acquisition requests based on the congestion control mechanism, communication efficiency can be improved and bandwidth usage can be reduced.
[0017] In one possible implementation, the exclusive processing rights for the target address are obtained by the first node from the second node before receiving the message sequence.
[0018] In the above implementation, the first node can obtain independent processing permissions for the target address in advance. In this way, when a message sequence is received, it can directly perform data access operations on the target address corresponding to the message sequence based on the independent processing permissions of the target address obtained in advance, thereby reducing the latency of data access operations.
[0019] In one possible implementation, the method further includes: before receiving the message sequence, the first node obtains exclusive processing rights for a specified address range from the second node, wherein the specified address range includes the target address corresponding to the message in the message sequence.
[0020] Optionally, if the first node, after obtaining exclusive processing permission for the specified address range, does not perform data access operations on the first address within the specified address range within a set time period, then the exclusive processing permission for the first address is released, wherein the first address is any address within the specified address range.
[0021] In the above implementation, after the first node obtains independent processing permission for the target address (such as the first address), if it does not perform data access operation on the target address within a set time period, it indicates that the independent processing permission of the target address has expired. In this case, the first node releases the independent processing permission of the target address so that other nodes can obtain the independent processing permission of the target address to perform data access processing.
[0022] In one possible implementation, obtaining exclusive processing rights for a specified address range from the second node includes: the first node determining the specified address range based on the addresses of the storage space managed by the second node from the target addresses corresponding to historical data access operations; and the first node obtaining exclusive processing rights for the specified address range from the second node.
[0023] In the above implementation method, based on statistics, if data access operations on certain addresses have been performed frequently in the past, the probability of data access to those addresses in the future will also be high. Therefore, by adopting the above implementation method, the independent processing rights of the target address obtained by the first node in advance are more likely to be used in the subsequent message sequence processing. Thus, the independent processing rights of the target address can be obtained in advance in a targeted manner, thereby reducing system overhead.
[0024] In one possible implementation, the first node performs data access operations on the target address based on its exclusive processing permissions for the target address, including: the first node obtaining a cache address corresponding to the target address from the second node; and the first node performing data access operations on the cache address corresponding to the target address based on its exclusive processing permissions for the target address.
[0025] In one possible implementation, the first node obtains the cache address corresponding to the target address from the second node, including:
[0026] The first node sends an address acquisition request through a first path between itself and the second node. The address acquisition request carries a first target address, which includes the target address corresponding to at least one message in the message sequence. The first path is selected by the first node based on the congestion status of each path between itself and the second node. The first node receives an address acquisition response from a second path between itself and the second node. The address acquisition response carries a first cached address corresponding to the first target address. The second path may be the same as or different from the first path.
[0027] In the above implementation, since the first node can select a suitable path to transmit the address and obtain the request based on the congestion control mechanism, the communication efficiency can be improved and the bandwidth usage can be reduced.
[0028] In one possible implementation, after the first node performs data access operations on the target address based on the exclusive processing permission of the target address, the method further includes: the first node releasing the exclusive processing permission of the target address.
[0029] In the above implementation, after the first node performs data access on the target address corresponding to the message in the message sequence, it releases the independent processing permission for data access on the corresponding target address, thereby enabling other nodes to obtain independent processing permission for that address so as to perform data access operations on that address.
[0030] In one possible implementation, the method further includes: after the first node obtains exclusive processing permission for the target address, if no data access operation is performed on the target address within a set time period, the exclusive processing permission for the target address is released to avoid the first node occupying the exclusive processing permission of the address for a long time.
[0031] In one possible implementation, the method further includes: before the first node performs data access on the first target address corresponding to the first message in the message sequence, if it receives a permission acquisition request for requesting exclusive processing permission for the first target address, the first node releases the exclusive processing permission for the first target address so that other nodes can perform data access operations on the first target address. The priority of the other nodes may be higher than that of the first node, or the priority of the other nodes' data access request for the first target address may be higher than the priority of the corresponding message in the message sequence received by the first node.
[0032] In one possible implementation, the method further includes: the first node sending the execution result of the data access operation on the target address to the external node according to the constraint of the external node on the visibility order of the execution result.
[0033] In one possible implementation, the method further includes: before the first node accesses data to the first target address corresponding to the first message in the message sequence, if it receives a permission acquisition request for requesting exclusive processing permission of the first target address, then it releases the exclusive processing permission of the first target address.
[0034] In one possible implementation, the first node and the second node are both System-on-a-Chip (SoC) chips, and the external node is an input / output (I / O) device.
[0035] Secondly, an interconnection system is provided, the interconnection system comprising at least two nodes interconnected in a horizontal scaling manner, the at least two nodes interconnected in a horizontal scaling manner comprising a first node and a second node, the first node being used for:
[0036] Receive a message sequence from an external node, wherein the messages in the message sequence are used to request data access to a target address, the target address belongs to the storage space managed by the second node, the external node is a node outside the interconnection system, and the external node has a stricter constraint on the visible order of the data access operation execution results than the first node.
[0037] Obtain exclusive processing rights for the target address;
[0038] Based on the exclusive processing permission of the target address, data access operations are performed on the target address, wherein the visibility order of the execution results of the data access operations satisfies the constraint of the external node on the visibility order of the execution results.
[0039] In one possible implementation, the first node is specifically configured to: in response to receiving the message sequence, obtain exclusive processing rights for the target address from the second node.
[0040] In one possible implementation, the first node is specifically configured to send a permission acquisition request to the second node after receiving the message sequence, the permission acquisition request carrying a target address corresponding to at least one message in the message sequence; the second node is configured to send a permission acquisition response to the first node after receiving the permission acquisition request, the permission acquisition response being used to instruct the exclusive processing permission of the target address corresponding to the at least one message to be transferred to the first node.
[0041] In one possible implementation, the first node is specifically used to send the permission acquisition request through a first path between the first node and the second node after receiving the message sequence, wherein the first path is selected by the first node based on the congestion status of each path between the first node and the second node;
[0042] The second node is specifically used to send the permission acquisition response from a second path between itself and the first node, wherein the second path may be the same as or different from the first path.
[0043] In one possible implementation, the exclusive processing rights for the target address are obtained by the first node from the second node before receiving the message sequence.
[0044] In one possible implementation, the first node is further configured to: obtain from the second node, before receiving the message sequence, operation permission to access data within a specified address range in the storage space managed by the second node, wherein the first address is an address within the specified address range.
[0045] Optionally, if after obtaining the permission to access data at the first address, no data is accessed at the first address within a set time period, the permission to access data at the first address is released.
[0046] In one possible implementation, the first node is specifically used to: determine the specified address range based on the addresses of the storage space managed by the second node from the target addresses corresponding to historical data access operations; and obtain exclusive processing rights for the specified address range from the second node.
[0047] In one possible implementation, the first node is specifically used to: obtain the cache address corresponding to the target address from the second node; and perform data access operations on the cache address corresponding to the target address based on the exclusive processing permission of the target address.
[0048] In one possible implementation, the first node is further configured to: after performing data access operations on the target address based on the exclusive processing permission of the target address, release the exclusive processing permission of the target address.
[0049] In one possible implementation, the first node is further configured to: after obtaining exclusive processing rights for the target address, if no data access operation is performed on the target address within a set time period, release the exclusive processing rights for the target address.
[0050] In one possible implementation, the first node is further configured to: send the execution result of the data access operation on the target address to the external node according to the external node's constraint on the visibility order of the execution result.
[0051] In one possible implementation, the first node is further configured to: before performing data access on the first target address corresponding to the first message in the message sequence, if a permission acquisition request is received requesting exclusive processing permission for the first target address, then release the exclusive processing permission for the first target address.
[0052] In one possible implementation, the second node is configured to: after the first node obtains exclusive processing rights to the first target address, if the first node fails to return the exclusive processing rights to the first target address within a set time period, then regain exclusive processing rights to the first target address, wherein the first target address is the target address corresponding to the first message in the message sequence.
[0053] In one possible implementation, the first node and the second node are both SoC chips, and the external node is an I / O device.
[0054] Thirdly, a SoC chip is provided, comprising: one or more processors and one or more memories; wherein the one or more memories store one or more computer programs, the one or more computer programs including instructions that, when executed by the one or more processors, cause the SoC chip to perform the method as described in any one of the first aspects above.
[0055] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium including a computer program that, when run on a computing device, causes the computing device to perform the method as described in any one of the first aspects above.
[0056] Fifthly, a computer program product is provided, which, when invoked by a computer, causes the computer to perform the method as described in any one of the first aspects above.
[0057] For the beneficial effects in aspects two through five above, please refer to the beneficial effects in aspect one, which will not be repeated here. Attached Figure Description
[0058] Figure 1 This is a schematic diagram illustrating the system expansion methods.
[0059] Figure 2 This is a schematic diagram of a multipath network for a scale-out system.
[0060] Figure 3 A schematic diagram illustrating the sequential processing method used to preserve order in a scale-out system;
[0061] Figure 4 This is a schematic diagram illustrating the use of sequence number preservation in a scale-out system.
[0062] Figure 5 A schematic diagram of the architecture of a scale-out system provided for embodiments of this application;
[0063] Figure 6 A block diagram illustrating the data access process provided in the embodiments of this application;
[0064] Figure 7 This is an interactive diagram illustrating the data access process in an embodiment of this application;
[0065] Figure 8a and Figure 8b These are schematic diagrams illustrating the exclusive processing rights of the first node in this application embodiment for obtaining the target address corresponding to each message based on multiple paths;
[0066] Figure 9a , Figure 9b and Figure 9c These are schematic diagrams illustrating the validity period timing of the exclusive processing permission for the first target address in this embodiment of the application, as well as the processing situation after timeout. Detailed Implementation
[0067] First, let me describe some of the concepts involved in this application:
[0068] Data access operations: also known as data read and write operations, include write operations and read operations. A write operation refers to storing data at a target address, while a read operation refers to retrieving data from the target address.
[0069] Storage consistency refers to the requirement that after hardware performs read and write operations, other nodes or external systems have certain requirements regarding the globally visible order of the results of these read and write operations (whether read and write operations were performed). For example, if a node performs a read operation on two addresses in succession (i.e., performs two read and write operations), or if a node performs two write operations on an address in succession, then if other nodes or external systems not only know that two read and write operations have been performed, but also know (i.e., globally visible) that the order of the results of these two read and write operations matches the software expectation, then the storage consistency requirement is met.
[0070] It should be noted that, for ease of description, the order in which the execution results of read and write operations are globally visible will be referred to as the visible order of execution results in the following text.
[0071] Cache coherence: The processor is a faster operating device than memory. When the processor performs read / write operations on memory, waiting for the operation to complete before processing other tasks will cause processor blocking and reduce efficiency. Therefore, a cache (which is much faster than memory but has a smaller capacity) can be configured for each processor. When the processor writes data to memory, the data can be written to the cache first, allowing other tasks to be processed. Direct memory access (DMA) devices handle the data storage to memory. Similarly, when the processor reads data from memory, the DMA device first stores the data from memory to the cache, and then the processor reads the data from the cache. When different processors perform read / write operations on the same address in memory through their respective caches, there are strict requirements for the execution order of read / write operations. The next read / write operation must be blocked until the previous one is completed to prevent inconsistencies between the data in the cache and memory caused by simultaneous read / write operations.
[0072] Cache-coherent devices adhere to the MESI protocol, which defines four states for a cache line (the smallest unit of cache): Exclusive (E state), Modified (M state), Shared (S state), and Invalid (I state). Specifically, the E state indicates that the cache line is valid, the data in the cache is consistent with the data in memory, and the data exists only in this cache; the M state indicates that the cache line is valid, the data has been modified, the data in the cache is inconsistent with the data in memory, and the data exists only in this cache; the S state indicates that the cache line is valid, the data in the cache is consistent with the data in memory, and the data exists in multiple caches; and the I state indicates that the cache line is invalid.
[0073] Storage consistency models, ranked from strongest to weakest in terms of the required order of execution results visibility, include: sequential consistency (SC), total store order (TSO), and relaxed model (RM). The SC model requires that the order of read and write operations on shared memory in hardware strictly match the order of operations required by software instructions. The TSO model, building upon the SC model, introduces a caching mechanism, relaxing the order constraint on write-read (write-then-read) operations, meaning that read operations can complete before write operations. The RM model is the most relaxed, imposing no order constraint on any read or write operations.
[0074] In systems employing a scale-out approach, to improve the efficiency of interaction between nodes and thus enhance system performance, nodes are networked using a multipath approach, enabling the system to support multipath communication.
[0075] For example, Figure 2 This diagram illustrates a multi-path network of nodes in a scale-out system. A first node (Node0 in the diagram) sends a message sequence to a second node (Node3 in the diagram). The messages in this sequence request data access to a target address within the memory storage space of the second node. Due to the multi-path network, the system can select a path to transmit the messages based on the congestion level (i.e., busyness) of each path. For example, at time T0, the system selects the currently idle direct path from Node0 to Node3 to transmit the message. However, at time T1, because the direct path from Node0 to Node3 is heavily congested and blocked, the system selects a relatively idle path from Node0 to Node1 to Node3 to transmit message 1. This fully utilizes communication resources, achieves congestion control, and improves interconnection efficiency.
[0076] Systems that use a scale-out networking approach (hereinafter referred to as scale-out systems) support out-of-order propagation along the path. However, when interfacing with external systems that require in-order execution results, such as I / O devices, the external system must adhere to certain order constraints (such as the order constraints of the SC model or TSO model). Therefore, for message sequences from the external system, it is necessary to ensure that the visible order of the execution results of the corresponding data access operations within the scale-out system conforms to the corresponding order constraints to meet the expectations of the external system. The scale-out system can return the execution results (i.e., data access responses) of the corresponding data access messages to the external system according to these order constraints.
[0077] To ensure that the visible order of execution results from a message sequence from an external system conforms to the order constraints of that external system, one industry practice is to use serial processing on a single path. The sending node sends messages in sequence, and after sending one message, it waits for the receiving node to process the message before sending the next message.
[0078] For example, Figure 3This diagram illustrates a scale-out system employing a serial processing approach to ensure that the visible order of execution results for a message sequence adheres to order constraints. As shown, in the scale-out system, the first node (Node0 in the diagram) receives a message sequence (message 0, message 1, ..., message N) from an external I / O device and then forwards the messages to the second node (Node1 in the diagram) for processing. The first node (Node0) receives the message sequence from the I / O device, which has strict requirements regarding the visible order of execution results. To ensure that the visible order of the execution results meets the I / O device's requirements, the first node (Node0), as the upstream master node, needs to send the messages sequentially through an out-of-order communication path to the second node (Node1), which is the downstream slave node. The messages are then processed on the second node (Node1) according to the order they were received. In the above process, after the second node (Node1) finishes processing the received message, it needs to complete a handshake with the first node (Node0) as the upstream node (i.e., reply with a response message). Only after receiving the response message will the first node (Node0) send the next message to the second node (Node1) as the downstream node, so as to ensure that the visible order of the execution results of the message sequence is consistent with the message order in the message sequence, thereby satisfying the requirement of the I / O device for the visible order of the execution results of data access operations.
[0079] For example, Figure 3 As shown, the first node (Node0) receives a message sequence (message 0, message 1, ..., message N) from the I / O device. First, it sends message 0 to the second node (Node1) to enable the second node to perform the corresponding data access operation. After receiving the response 0 from the second node (Node1), it sends message 1 to the second node (Node1) to enable the second node to perform the corresponding data access operation. After receiving the response 1, it sends the next message to the second node, and so on, until message N is sent to the second node (Node1) and the second node returns response N. In this way, the first node (Node0) can obtain the globally visible data access operation execution results (response 0, response 1, ..., response N). The order of these execution results is consistent with the message order in the message sequence, which meets the requirement of the I / O device for the visible order of data access operation execution results.
[0080] The above-mentioned serial processing method results in excessively long processing delays, severely impacting bandwidth and leading to low interconnection efficiency because it requires waiting for one message to be processed before sending the next.
[0081] To reduce processing latency, another approach in the industry is to use sequence numbers to guarantee execution order. In this method, a set of sequence numbers is used to mark each message sequence. Messages carrying sequence numbers can be transmitted out of order along various paths, and the corresponding order is restored at downstream nodes according to the sequence numbers in the messages.
[0082] For example, Figure 4 This diagram illustrates a scale-out system that uses sequence numbers to ensure the visible order of execution results for a message sequence adheres to order constraints. As shown, after the first node (Node0) in the scale-out system receives a message sequence (message 0, message 1, ..., message N) from an external I / O device, because the visible order of the execution results of this message sequence must adhere to strict order constraints, the first node (Node0), as the upstream node, assigns a corresponding sequence number to each message before sending it to the second node (Node1), which is the downstream node. Figure 4 The sequence number in the message sequence is consistent with the message sequence number in the message sequence. For example, message 0 is marked with a sequence number of 0 (i.e., Seq Num = 0), message 1 is marked with a sequence number of 1 (i.e., Seq Num = 1), and so on, with message N being marked with a sequence number of N (i.e., Seq Num = N). Messages carrying sequence numbers can be propagated out of order along the path. After receiving the messages, the second node (Node1) uses a mechanism similar to a reorder buffer (ROB) to restore the order based on the sequence numbers carried in the messages. The response messages returned by the second node (Node1) to the first node (Node0) can also be transmitted out of order. The first node (Node0) can determine the order of response messages based on the recorded message sequence number to ensure the visible order of data access operation execution results. For example, response 0 is the data access operation execution result corresponding to message 0, response 1 is the data access operation execution result corresponding to message 0, and response N is the data access operation execution result corresponding to message N. The first node (Node0) can determine the order of response 0 to response N based on the sequence numbers corresponding to messages 0 to N respectively, and then return the responses to the I / O device in order.
[0083] The method of adding sequence numbers described above has high complexity and consumes significant resources because it requires additional sequence numbers to be carried in the messages. This necessitates corresponding processing mechanisms from upstream and downstream nodes, and each process (one process corresponds to one message sequence) between nodes needs to be configured with a set of sequence numbers. Furthermore, it is difficult to extend to multi-path scenarios. For the same process, since each path can be used for communication, each path needs to support a corresponding sequence number. The overhead increases dramatically with the number of paths, which is unacceptable for the system.
[0084] To address the aforementioned issues, embodiments of this application provide a data access method, an interconnection system, and an apparatus. These embodiments can be applied to systems employing scale-out and multi-path networking to achieve order preservation in multi-path scenarios with minimal overhead.
[0085] This application extends the Cache Coherence Domain (CC) range in the Scale-out system by including the boundary node between the Scale-out system and the external system within the CC range of the downstream node. This allows control operations on the visible order of execution results to be migrated from the downstream node to the boundary node. In other words, the visible order of data access operation execution results is controlled at the boundary node. This ensures that data access operations can be processed in parallel within the Scale-out system, guaranteeing bandwidth. Furthermore, it only requires obtaining address operation permissions in the consistency process to migrate the processing order. Compared to the sequence number-based scheme described above, this approach saves system overhead while ensuring that the visible order of data access operation execution results meets the requirements of the external system.
[0086] In this context, the boundary node between the scale-out system and the external system refers to the node in the scale-out system that receives a sequence of messages sent by the external system (such as an external I / O device). For example, Figure 2 , Figure 3 or Figure 4 The first node in the array can be called the boundary node.
[0087] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0088] See Figure 5 This is a schematic diagram of the architecture of the scale-out system in the embodiments of this application.
[0089] As shown in the figure, the system includes a first node (Node0), a second node (Node3), a third node (Node1), and a fourth node (Node2). These nodes are interconnected, forming a multi-path network structure. Taking the path between the first and second nodes as an example, there are three paths between them: Node0-Node3, Node0-Node1-Node3, and Node0-Node2-Node3. It should be noted that the number of nodes in a real-world system can be higher than this. Figure 5 The number of nodes shown may be more or less; however, this application embodiment does not limit the number of nodes in the system.
[0090] A typical application of the above system architecture is an HPC chip interconnect system. In an HPC chip interconnect system, the nodes can be System on Chip (SoC) chips. Nodes in an HPC chip interconnect system can interact with nodes outside the system. These external nodes can be I / O devices. For example, when an I / O device is connected to an SoC chip via the Peripheral Component Interconnect Express (PCIE) standard, the I / O device can be a PCIE board. When an I / O device is connected to an SoC chip via a network transmission protocol, the I / O device can be an Ethernet interface.
[0091] For ease of description, in the following description, the node that interacts with the external system in the above scale-out system is referred to as the boundary node (or upstream node). The boundary node can receive message sequences from the external system. The messages in this message sequence are data access request messages, used to request data access operations on addresses within the storage space managed by the target node (or downstream node) in the scale-out system. It is understood that the storage space managed by the target node refers to the storage space of the memory within the target node. It is also understood that in the above scale-out system, any node can become either a boundary node (or upstream node) or a target node (or downstream node).
[0092] For example, such as Figure 5As shown, in a scenario where the first node receives a message sequence from an external I / O device, this first node is the boundary node. The target address corresponding to the data access request message in this message sequence belongs to the storage space managed by the second node, making the second node the target node for this message sequence. If the second node receives a message sequence from an external I / O device, in this scenario, the second node is also a boundary node. The target address corresponding to the data access request message in this message sequence belongs to the storage space managed by the third node, making the third node the target node for this message sequence.
[0093] Based on the above system architecture, this embodiment of the application extends the cache consistency range (CC range) to migrate the exclusive processing rights (also known as E-state) of the target address from the downstream node to the boundary node (upstream node), thereby migrating the control operations for the visible order of data access operation execution results from the downstream node to the boundary node. Thus, for message sequences received by the boundary node from external systems, since the exclusive processing rights of the target address have been migrated to the boundary node, the boundary node can undertake the control operations for the visible order of data access operation execution results, thereby ensuring that the visible order of execution results meets the requirements of the external system.
[0094] by Figure 5 For example, the first node (Node0) receives a message sequence from an external I / O device. This message sequence includes N data access request messages (Req0, Req1, ..., ReqN as shown in the figure), where N is an integer greater than or equal to 2. These N data access request messages are used to request data access to addresses within the memory space managed by the second node (Node3) in the system. That is, the target addresses corresponding to these N data access request messages belong to the address space of the memory on the second node. The first node (Node0) obtains exclusive processing rights for the corresponding addresses from the second node (Node3), thus migrating the exclusive processing rights for the corresponding addresses from the second node (Node3) to the first node (Node0), as shown in the figure. Originally, the exclusive processing rights for the addresses were only available within the CC range of the second node (Node3) (e.g., ...). Figure 5 As shown in a), after adopting the embodiment of this application, the first node (Node0) can be included in the CC scope (e.g., by migrating the address with exclusive processing rights (i.e., E-state migration)). Figure 5 (as shown in b).
[0095] In this embodiment, the exclusive processing permission of the target address is migrated from the downstream node to the boundary node, so that the boundary node can perform data access operations based on storage consistency requirements and can control the visible order of the data access operation execution results without the need for the downstream node to perform additional order processing. Thus, multi-path networking of the scale-out system can be achieved with less overhead.
[0096] Furthermore, within a scale-out system, inter-node communication can still support multipath communication. Optionally, nodes in the system can select paths based on congestion mechanisms, transmitting messages to other nodes via the selected paths. Messages between nodes can be transmitted out of order via multipath to reduce system latency and improve bandwidth utilization.
[0097] Based on the above architecture, a typical application scenario is the interaction between I / O devices (or I / O systems) and HPC systems. The message sequence sent from I / O devices to the HPC system usually has requirements regarding the visible order of execution results. However, HPC systems use a multi-path scale-out interconnection method, where messages are transmitted out of order along the paths between nodes and can be transmitted via multiple paths. Therefore, for a message sequence sent by an I / O device, after entering the scale-out system, it is necessary to ensure that the visible order of execution results meets the requirements of that I / O device.
[0098] Based on the above system architecture, the following will combine... Figure 6 and Figure 7 This application describes the data access process provided in its embodiments. This process can be applied to an interconnect system, which may include at least two nodes interconnected in a scale-out manner. These at least two scale-out interconnected nodes include a first node and a second node, for example, both the first node and the second node are SoC chips. The following process is described using the first node as the boundary node and the second node as the target node as an example.
[0099] in, Figure 6 This is a general block diagram of the data access process provided in the embodiments of this application. Figure 7This is a schematic diagram illustrating the data access process in a specific application scenario of this application embodiment. In this scenario, a first node in the interconnected system receives a message sequence from an external node (such as an I / O device). This message sequence includes a first message and a second message. The first message is a write request, used to request a write operation to a first target address, and the second message is a write request, used to request a write operation to a second target address. The first and second target addresses are addresses within the storage space managed by the second node in the interconnected system. The visible order of the execution results of this message sequence must adhere to order constraints. For example, the visible order from front to back is: the data access execution result corresponding to the first message, and the data access execution result corresponding to the second message.
[0100] It should be noted that, Figure 7 This description only uses an example of a message sequence containing two write requests. For other data access operation types (such as read requests), or when the message sequence contains more messages, it can be based on... Figure 7 The principle of the process shown is as follows.
[0101] like Figure 6 As shown, the data access process provided in this application embodiment may include the following steps:
[0102] S601: The first node receives a message sequence from an external node.
[0103] The external node is a node outside the interconnected system, such as an I / O device. This external node follows the execution result order constraint. The constraint on the visible order of data access operation execution results of this external node is stricter than that of the first node. For example, this external node follows the execution result visible order requirement required by the SC model, TSO model or other types of storage consistency models.
[0104] This message sequence includes at least two data access request messages, which are used to request data access to a target address. For example, the message sequence may include a write request to write data to the target address, or a read request to read data from the target address. The target address belongs to the storage space managed by the second node in the system; that is, the target address is an address within the storage space managed by the second node. For example, the target address is the physical address of the memory on the second node. Understandably, this message sequence may include a write request to write data to the memory of the second node, and / or a request to read data from the memory of the second node.
[0105] The access or read / write operations involved in the embodiments of this application can support write-write (write first, then write), write-read (write first, then read), read-write (read first, then write), and read-read (read first, then read) operations. All messages in the message sequence can have the same message type, such as all being write requests or all being read requests, or they can be different, such as some messages being write requests and others being read requests. The target addresses corresponding to each message in the message sequence can be the same or different, or some messages can have the same target address while others have different target addresses.
[0106] For example, such as Figure 7 As shown, an external node (such as an I / O device) sends a message sequence to the first node in the interconnected system. This message sequence includes a first message (write request 1) and a second message (write request 2). The first message (write request 1) carries the data to be written (1) and a first destination address, requesting that the data to be written (1) be stored at the first destination address. The second message (write request 2) carries the second destination address of the data to be written (2), requesting that the data to be written (2) be stored at the second destination address. Both the first and second destination addresses belong to the storage space managed by the second node within the system.
[0107] S602: The first node obtains exclusive processing rights for the aforementioned target address.
[0108] In this context, exclusive processing rights for the target address can refer to the E-state in cache consistency, indicating that a node has the right to access data at that address. A node with exclusive processing rights for the target address corresponding to a message sequence (such as the first node in this process) has the ability to process the visible order of execution results for that message sequence. Taking a message sequence that includes a first message and a second message, where the target address corresponding to the first message is the first target address and the target address corresponding to the second message is the second target address, as an example, the first node can obtain exclusive processing rights (i.e., E-state) for both the first and second target addresses from the second node.
[0109] After the first node obtains exclusive processing rights for the target address from the second node, the CC scope is extended from the second node to the first node, enabling the first node to participate in cache consistency management. Other nodes (such as the second node) cannot perform data access operations that require operation rights on the target address. Control over the visible order of execution results is transferred to the first node, that is, the first node controls the visible order of the execution results of data access operations.
[0110] In this embodiment, the first node can acquire exclusive processing rights to the target address using either a first acquisition method or a second acquisition method. The first acquisition method refers to initiating the acquisition process after receiving a message sequence; that is, the message sequence reception operation can trigger the acquisition process. The second acquisition method involves acquiring exclusive processing rights from other nodes (such as the second node) in advance. The first and second acquisition methods are described in detail below.
[0111] First method of acquisition:
[0112] When the first acquisition method is adopted, after the first node receives a message sequence from an external node of the system, it determines that the target address corresponding to the message in the message sequence belongs to the storage space managed by the second node, and then obtains exclusive processing rights for the target address from the second node.
[0113] Optionally, the process of the first node acquiring exclusive processing rights to the target address from the second node may include: after receiving a message sequence from a node outside the system, the first node sends a permission acquisition request to the second node, the permission acquisition request carrying a target address, which includes the target address corresponding to at least one message in the message sequence; the second node sends a permission acquisition response to the first node, the permission acquisition response indicating that the first node is allowed to acquire exclusive processing rights to the corresponding target address. Further, after receiving the permission acquisition response from the second node, the first node may send a confirmation message to the second node to inform the second node that the first node has acquired exclusive processing rights to the corresponding target address.
[0114] In the aforementioned permission acquisition process, the first node can acquire permissions in units of cache lines. Taking a cache line size of 64 bytes as an example, a single permission acquisition request is used to acquire exclusive processing permissions for an address range with a capacity of 64 bytes. In another embodiment, the first node can also acquire permissions at a larger granularity; that is, a single permission acquisition request can acquire exclusive processing permissions for an address range with a capacity greater than the cache line size. For example, exclusive processing permissions can be acquired for an address range of 4KB pages. In this way, a single permission acquisition request can achieve E-state acquisition of the target address corresponding to the entire message sequence, thereby reducing complexity and system overhead.
[0115] In the aforementioned permission acquisition process, there are no strict requirements on the path taken by the first node when sending permission acquisition requests or the order in which they are sent. For example, different permission acquisition requests can be sent from different paths or the same path, and can be sent out of order or in parallel. Similarly, this embodiment of the application does not strictly require the path taken by the second node when sending permission acquisition responses or the order in which they are sent. For example, different permission acquisition responses can be sent from different paths or the same path, and can be sent out of order or in parallel. Furthermore, the paths taken by the permission acquisition requests and their corresponding permission acquisition responses may be the same or different.
[0116] Optionally, in implementation, the first node can select a suitable path to send an access request to the second node based on a congestion control mechanism to reduce latency and improve communication efficiency; the second node can also select a suitable path to send an access response to the first node based on a congestion control mechanism. For example, after receiving a message sequence, the first node sends an access request through the first path between itself and the second node. This first path is selected by the first node based on the congestion status of all paths between it and the second node; it can be the path with the lowest current congestion among all paths between the first and second nodes. The second node can then send a corresponding access response through the second path between itself and the first node. This second path is selected by the second node based on the congestion status of all paths between it and the first node; it can also be the path with the lowest current congestion among all paths between the second and first nodes. The first and second paths may be the same or different.
[0117] For example, with Figure 5 Taking the system architecture shown as an example, Figure 8a and Figure 8b This diagram illustrates the exclusive processing rights of the first node for obtaining the target address corresponding to each message based on multi-path acquisition. For example... Figure 8a As shown, the message sequence includes message 0, message 1, ..., message N. The permission acquisition request (Req0:Get_E) corresponding to message 0 is transmitted through the direct path (Node0-Node3) between the first and second nodes. The permission acquisition request (Req1:Get_E) corresponding to message 1 is transmitted through the indirect path (Node0-Node2-Node3) between the first and second nodes. The permission acquisition request (Req2:Get_E) corresponding to message 2 is transmitted through the indirect path (Node0-Node1-Node3) between the first and second nodes. The first node can select an appropriate path to transmit the permission acquisition requests for each message based on a congestion control mechanism. The permission acquisition requests corresponding to messages 0, 1, and 2 can also be sent in parallel.
[0118] like Figure 8b As shown, the permission acquisition response (Req0:Res) corresponding to the permission acquisition request (Req0:Get_E) is transmitted through the indirect path (Node3-Node1-Node0) between the second node and the first node; the permission acquisition response (Req1:Res) corresponding to the permission acquisition request (Req1:Get_E) is transmitted through the direct path (Node3-Node0) between the second node and the first node; and the permission acquisition response (Req2:Res) corresponding to the permission acquisition request (Req2:Get_E) is transmitted through the indirect path (Node3-Node2-Node0) between the second node and the first node. The second node can select an appropriate path to transmit the permission acquisition responses for each message based on a congestion control mechanism. The permission acquisition responses can also be sent in parallel.
[0119] For example, such as Figure 7 As shown, the first node sends a first permission acquisition request (GET_E1) and a second permission acquisition request (GET_E2) to the second node. Both requests can be GET_E messages. The first permission acquisition request includes a first target address, used to request exclusive processing rights for that target address. The second permission acquisition request includes a second target address, used to request exclusive processing rights for that target address. The order in which the first and second permission acquisition requests are sent from the first node to the second node is not restricted, and the paths traversed by the first and second permission acquisition requests can be paths selected by the first node based on congestion control mechanisms.
[0120] After receiving the first permission acquisition request, the second node sends a first permission acquisition response (RSP1) to the first node. Upon receiving the second permission acquisition request, the second node sends a second permission acquisition response (RSP2) to the first node. The first permission acquisition response (RSP1) indicates that exclusive processing rights for the first target address will be transferred to the first node; the second permission acquisition response (RSP2) indicates that exclusive processing rights for the second target address will be transferred to the first node. The order in which the second node sends the first and second permission acquisition responses to the first node is not restricted, and the path traversed by the first and second permission acquisition responses can be a path selected by the second node based on congestion control mechanisms.
[0121] After receiving the first permission acquisition response, the first node sends a first acknowledgment message (ACK1) to the second node to inform it that it has acquired exclusive processing rights to the first target address. After receiving the second permission acquisition response, the first node sends a second acknowledgment message (ACK2) to the second node to inform it that it has acquired exclusive processing rights to the second target address. The order in which the first and second acknowledgment messages are sent is not restricted, and the paths traversed by the first and second acknowledgment messages can be paths selected by the first node based on congestion control mechanisms.
[0122] Second method of acquisition:
[0123] When using the second acquisition method, the first node can obtain exclusive access to addresses within the storage space managed by other nodes (such as the second node) in advance. For example, the first node can obtain exclusive access to addresses within the storage space managed by the second node according to a set period, a set time, or when idle.
[0124] Considering that the storage space managed by the second node may be large, in order to reduce system overhead and the impact on data storage operations on other nodes, in this embodiment of the application, the first node can determine the addresses belonging to the storage space managed by the second node from the target addresses corresponding to the historical data access operations (i.e., the target addresses corresponding to the historical data access operations), and determine a specified address range accordingly, and obtain exclusive processing rights for the addresses within the specified address range from the second node. The specified address range matches the target addresses involved in the historical data access operations; for example, it can be the same as the target address range involved in the historical data access operations, or it can include the target addresses involved in the historical data access operations, and further expand the address range appropriately.
[0125] Optionally, when using the second acquisition method, the process of the first node acquiring exclusive processing rights to the target address from the second node may include: the first node sending a permission acquisition request to the second node, the permission acquisition request carrying the target address; and the second node sending a permission acquisition response to the first node, the permission acquisition response indicating that the exclusive processing rights to the target address are transferred to the first node. The target address may be an address within the specified address range determined by the above method.
[0126] In the aforementioned permission acquisition process, there are no strict requirements on the path taken by the first node when sending permission acquisition requests or the order in which they are sent. For example, different permission acquisition requests can be sent from different paths or the same path, and can be sent out of order or in parallel. Similarly, this embodiment of the application does not strictly require the path taken by the second node when sending permission acquisition responses or the order in which they are sent. For example, different permission acquisition responses can be sent from different paths or the same path, and can be sent out of order or in parallel. Furthermore, the paths taken by the permission acquisition requests and their corresponding permission acquisition responses may be the same or different.
[0127] In the aforementioned permission acquisition process, the first node can acquire permissions in units of cache lines. Taking a cache line size of 64 bytes as an example, a single permission acquisition request is used to acquire exclusive processing permissions for an address range with a capacity of 64 bytes. In another embodiment, the first node can also acquire permissions with a larger granularity; that is, a single permission acquisition request can acquire exclusive processing permissions for an address range with a capacity greater than the cache line size. For example, exclusive processing permissions for an address range of 4KB pages can be acquired in units of that size. In this way, a single permission acquisition request can achieve E-state acquisition of a larger address range, thereby reducing complexity and system overhead.
[0128] Optionally, in this embodiment of the application, when the first acquisition method, the second acquisition method or other possible acquisition methods are used to obtain the permission, the exclusive processing permission of the migrated address has a validity period. The validity period represents the maximum duration for which the exclusive processing permission of the address is migrated to other nodes (such as the first node mentioned above). The size of the validity period can be preset.
[0129] In practical applications, upstream nodes (such as the first node in this embodiment) may fail, for example, if the first node is hot-swapped and removed from the system. In this case, the first node will be unable to return the exclusive processing rights of the target address obtained from the second node to the second node, causing a consistency error when the system subsequently performs data access operations on that target address. This embodiment solves the above problem by setting a validity period for the permissions. For ease of description, taking the transfer of exclusive processing rights of the first target address as an example, when the exclusive processing rights of the first target address are transferred from the second node to the first node, if the exclusive processing rights of the first target address are not returned to the second node within a set time period (for example, if no message indicating the return of the exclusive processing rights of the first target address is received from the first node within the set time period), then the second node regains the exclusive processing rights of the first target address. The set time period is the duration corresponding to the validity period.
[0130] Taking the transfer of exclusive processing rights for the first target address from the second node to the first node as an example, in another scenario, after the first node acquires exclusive processing rights for the first target address, if it does not perform any data access operations on the first target address within a set time period, it indicates that the exclusive processing rights for data access on the first target address have timed out or expired. The first node will release the exclusive processing rights for the first address, that is, return the exclusive processing rights for the first target address to the second node. For example, it can send a message to the second node to instruct the return of the exclusive processing rights for the first target address.
[0131] Optionally, taking the transfer of exclusive processing rights for the first target address from the second node to the first node as an example, the first node can start a first timer to time the validity period of the exclusive processing rights for the first target address, and the second node can start a second timer to time the validity period of the exclusive processing rights for the first target address. The first node can start the first timer for the exclusive processing rights of the first target address when sending a permission acquisition request to request exclusive processing rights for the first target address, or when receiving a permission acquisition response from the second node in response to the permission acquisition request. Similarly, the second node can start the second timer for the exclusive processing rights of the first target address when receiving the aforementioned permission acquisition request from the first node, or when the second node sends a corresponding permission acquisition response to the first node. At the first node, if the first node performs a data access operation on the first target address during the operation of the first timer, the first timer is released; at the second node, if an instruction to return exclusive processing rights to the first target address is received during the operation of the second timer, the second timer is released.
[0132] Optionally, in some embodiments, the duration of the first timer and the second timer is the same, which is the validity period. In other embodiments, the durations of the first timer and the second timer are different, wherein the duration of the first timer is shorter than the duration of the second timer. Considering that there is a certain delay between the time the permission acquisition response is sent on the second node and the time the permission acquisition response is received on the first node, setting the duration of the first timer to be shorter than the duration of the second timer can ensure the consistency of exclusive processing permissions for the same address on the first node and the second node, thereby avoiding the situation where exclusive processing permissions for the same address are valid on both the first node and the second node.
[0133] Optionally, in other embodiments, after receiving the permission acquisition response from the second node, the first node may also send a confirmation message to the second node to indicate that it has acquired exclusive processing permission for the first target address. Correspondingly, upon receiving the confirmation message from the first node confirming that the first node has acquired exclusive processing permission for the first target address, the second node may start a second timer for the exclusive processing permission of the first target address. This ensures that the validity period of the exclusive processing permission for the first target address is closer for both the first and second nodes, contributing to the consistency of the exclusive processing permission for the first target address between the first and second nodes. In this case, the duration of the first timer and the duration of the second timer can be set to be substantially the same.
[0134] For example, Figure 9a , Figure 9b and Figure 9c The expiration timer for the exclusive processing permission of the first target address and the handling after timeout are shown respectively.
[0135] like Figure 9a As shown, the first node sends a first permission acquisition request (GET_E1) to the second node to request exclusive processing rights to the first target address; after receiving the first permission acquisition request, the second node sends a first permission acquisition response (RSP1) to the first node to transfer the exclusive processing rights to the first target address to the first node, and starts a second timer to count the duration of the transfer of the exclusive processing rights to the first target address, the duration of which is Th; after receiving the first permission acquisition response, the first node starts a first timer to count the duration of the transfer of the exclusive processing rights to the first target address, the duration of which is Tm. <Th。
[0136] At time Tf during the operation of the first and second timers, the first node is removed before the exclusive processing rights of the first target address are returned, and cannot send an instruction to the second node to return the exclusive processing rights of the first target address; on the second node, when the second timer expires, the second node regains the exclusive processing rights of the first target address.
[0137] like Figure 9bAs shown, the first node sends a first permission acquisition request (GET_E1) to the second node to request exclusive processing rights to the first target address; after receiving the first permission acquisition request, the second node sends a first permission acquisition response (RSP1) to the first node to transfer the exclusive processing rights to the first target address to the first node, and starts a second timer to count the duration of the transfer of the exclusive processing rights to the first target address, the duration of which is Th; after receiving the first permission acquisition response, the first node starts a first timer to count the duration of the transfer of the exclusive processing rights to the first target address, the duration of which is Tm. <Th。
[0138] During the first timer's operation, if the first node does not perform any data access operations on the first target address, the first timer expires. When the first timer expires, the first node sends an instruction to the second node to return the exclusive processing rights of the first target address and further releases the first timer. After receiving the instruction, the second node obtains the exclusive processing rights of the first target address and further releases the second timer.
[0139] like Figure 9c As shown, the first node sends a first permission acquisition request (GET_E1) to the second node to request exclusive processing rights to the first target address; after receiving the first permission acquisition request, the second node sends a first permission acquisition response (RSP1) to the first node to transfer the exclusive processing rights to the first target address to the first node, and starts a second timer to count the duration of the transfer of the exclusive processing rights to the first target address, the duration of which is Th; after receiving the first permission acquisition response, the first node starts a first timer to count the duration of the transfer of the exclusive processing rights to the first target address, the duration of which is Tm. <Th。
[0140] At time Tf during the operation of the first and second timers, the first node performs a data access operation on the first target address and sends an instruction to the second node to return the exclusive processing rights of the first target address. Upon receiving this instruction, the second node regains the exclusive processing rights of the first target address. The data access operation performed by the first node on the first target address may succeed or fail. Regardless of the success or failure of the data access operation, the first node can send an instruction to the second node to return the exclusive processing rights of the first target address.
[0141] S603: The first node, based on its exclusive processing rights to the aforementioned target address, performs data access operations on that target address. The visibility order of the execution results of the data access operations satisfies the constraint imposed by the external nodes on the visibility order of the execution results.
[0142] In this step, taking a message sequence including a first message and a second message as an example, after the first node obtains the first target address and the second target address, it can perform data access operations on the first target address based on its exclusive processing rights, and on the second target address based on its exclusive processing rights. The order of performing data access operations on the first target address and the second target address is only required to meet the storage consistency requirement. For example, if the first target address and the second target address are different addresses, data caching operations can be performed on the first target address first based on the first message, and then on the second target address based on the second message; alternatively, data caching operations can be performed on the second target address first based on the second message, and then on the first target address based on the first message; or data caching operations can be performed on the first target address and the second target address in parallel. Furthermore, if the first target address and the second target address are the same address, data caching operations on that target address must be performed first based on the first message, and then on the second message.
[0143] In this step, the visible order of the execution results of data access operations on the first node needs to satisfy the visibility order constraint of the data storage operation execution results that the external nodes adhere to. For example, if the external node requires the execution results to be visible in the following order: from front to back, the execution results corresponding to the first message and the second message, then on the first node, the order of the data access operation execution results is: the execution results of the data access operation on the first target address and the execution results of the data access operation on the second target address, and this order is globally visible.
[0144] In this step, optionally, the first node can obtain the cache address corresponding to the target address from the second node and perform data access on the cache address.
[0145] For example, the first node can send an address retrieval request to the second node, which carries the target address (such as the first target address and / or the second target address). After receiving the address retrieval request, the second node sends an address retrieval response to the first node, which carries the cached address corresponding to the target address.
[0146] Optionally, in some embodiments, the first node may send address retrieval requests corresponding to each message according to the execution result visibility order required by the external node, and may record the sending order of the address retrieval requests as the global visibility order of the execution results. For example, according to the constraint of the external node on the execution result visibility order, the visibility order of the execution results should be consistent with the order of the message sequence. Therefore, the first node sends the address retrieval requests corresponding to the corresponding messages to the second node in the order of the message sequence.
[0147] Optionally, in other embodiments, the first node may select a suitable path to send an address retrieval request to the second node based on a congestion control mechanism to reduce latency and improve communication efficiency; the second node may also select a suitable path to send an address retrieval response to the first node based on a congestion control mechanism. For example, the first node sends an address retrieval request through a first path between itself and the second node, where the first path is selected by the first node based on the congestion status of all paths between it and the second node, and the first path may be the path with the lowest current congestion among all paths between the first and second nodes. The second node may send a corresponding address retrieval response through a second path between itself and the first node, where the second path is selected by the second node based on the congestion status of all paths between it and the first node, and the second path may be the path with the lowest current congestion among all paths between the second and first nodes. The first path and the second path may be the same or different.
[0148] For example, such as Figure 7 As shown, after obtaining exclusive processing rights to the first target address, the first node sends a first address acquisition request to the second node. This request may include the first target address and the data access operation type (in this example, the data access operation type is a write operation) to request the cache address corresponding to the first target address. Similarly, after obtaining exclusive processing rights to the second target address, the first node may send a second address acquisition request to the second node. This request may also include the second target address and the data access operation type (in this example, the data access operation type is a write operation) to request the cache address corresponding to the second target address. Both the first and second address acquisition requests can be writeback messages. The order in which the first and second address acquisition requests are sent from the first node to the second node is not restricted, and the paths traversed by the first and second address acquisition requests can be paths selected by the first node based on congestion control mechanisms.
[0149] After receiving the first address acquisition request, the second node sends a first address acquisition response (RSP3) to the first node, indicating the second cache address corresponding to the first target address. Upon receiving the second address acquisition request, the second node sends a second address acquisition response (RSP4) to the first node, indicating the first cache address corresponding to the second target address. The order in which the second node sends the first and second address acquisition responses is not restricted, and the path traversed by the first and second address acquisition responses can be a path selected by the second node based on congestion control mechanisms.
[0150] The first node obtains the cache address corresponding to the first target address. After writing the data to be written (1) to the first cache address, it can further send an instruction to the second node to return the exclusive processing permission of the first target address. The second node can then regain the exclusive processing permission of the first target address according to this instruction. Similarly, the first node obtains the cache address corresponding to the second target address. After writing the data to be written (2) to the second cache address, it can further send an instruction to the second node to return the exclusive processing permission of the second target address. The second node can then regain the exclusive processing permission of the second target address according to this instruction. Through this process, after completing the data access operation, the first node can return the exclusive processing permission of the corresponding address to the second node, so that the second node or other nodes can perform data access operations on that address.
[0151] Optionally, in the above Figure 6 Following step S603 of the illustrated process, the following steps may also be included:
[0152] S604: The first node sends the execution result of the data access operation on the target address to the external node according to the constraint of the visible order of the execution result of the external node.
[0153] For example, such as Figure 7 As shown, if the external node has strict order constraints on the execution results, the required execution results should be in the following order: from front to back, the execution results corresponding to the first message and the execution results corresponding to the second message. Then, the first node sends the first response and the second response to the external node in sequence, where the first response is the response message to the first message and the second response is the response message to the second message.
[0154] Optionally, in some embodiments of this application, before the first node performs data access on the target address corresponding to a message (such as the first message) in the message sequence, if it receives a request to acquire exclusive processing rights for the first target address (the request is used to request exclusive processing rights for the first target address), then the exclusive processing rights for the first target address are released, so that other nodes can acquire exclusive processing rights for the first target address. The other nodes can be other nodes within the system where the first node resides, excluding the first node. The priority of these other nodes may be higher than that of the first node, or the priority of their data access request for the first target address may be higher than the priority of the corresponding message in the message sequence received by the first node.
[0155] As can be seen from the above description, by migrating the exclusive processing rights (i.e., E-state) of the target addresses corresponding to messages (Req0-ReqN) in the message sequence of the downstream node (second node) to the boundary node (first node), since a node possesses the E-state of the target address corresponding to the message sequence, it means that the node has the ability to control the visible order of the execution results for that message sequence. Therefore, the visible order of the execution results can be controlled at the upstream node, ensuring that the visible order of the message sequence's execution results conforms to the expectations of the external system. In this embodiment, controlling the visible order of the message sequence's execution results at the upstream node eliminates the need for interaction between the upstream and downstream nodes and eliminates the need for a sequence number-based reordering mechanism to ensure the controllability of the visible order of the execution results, thereby reducing system overhead.
[0156] Furthermore, this application also provides a SoC chip, which may include one or more processors and one or more memories; wherein the one or more memories store one or more computer programs, and the one or more computer programs include instructions that, when executed by the one or more processors, cause the SoC chip to perform the methods provided in the foregoing embodiments.
[0157] Furthermore, embodiments of this application also provide a computer-readable storage medium comprising a computer program that, when executed on a computing device, causes the computing device to perform the method provided in the foregoing embodiments.
[0158] Furthermore, this application also provides a computer program product that, when invoked by a computer, causes the computer to execute the method provided in the foregoing embodiments.
[0159] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0160] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0161] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0162] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0163] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0164] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0165] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0166] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A data access method applied to an interconnected system, the interconnected system comprising at least two nodes interconnected in a horizontal scaling manner, the at least two nodes interconnected in a horizontal scaling manner comprising a first node and a second node, characterized in that, The method includes: The first node receives a message sequence from an external node. The messages in the message sequence are used to request data access to a target address. The target address belongs to the storage space managed by the second node. The external node is a node outside the interconnection system. The external node has stricter constraints on the visible order of the data access operation execution results than the first node. The first node obtains exclusive processing rights for the target address; The first node performs data access operations on the target address based on its exclusive processing permissions for the target address, wherein the visibility order of the execution results of the data access operations satisfies the constraint imposed by the external node on the visibility order of the execution results of the data access operations.
2. The method as described in claim 1, characterized in that, The first node obtains exclusive processing rights for the target address, including: In response to receiving the message sequence, the first node obtains exclusive processing rights for the target address from the second node.
3. The method as described in claim 2, characterized in that, In response to receiving the message sequence, the first node obtains exclusive processing rights for the target address from the second node, including: After receiving the message sequence, the first node sends an access request to the second node, the access request carrying the target address corresponding to at least one message in the message sequence; The first node receives a permission acquisition response from the second node, the permission acquisition response being used to instruct the exclusive processing permission of the target address corresponding to the at least one message to be transferred to the first node.
4. The method as described in claim 3, characterized in that, After receiving the message sequence, the first node sends a permission acquisition request to the second node, including: After receiving the message sequence, the first node sends the permission acquisition request through the first path between itself and the second node, wherein the first path is selected by the first node based on the congestion status of each path between itself and the second node; The first node receives the permission acquisition response from a second path between itself and the second node, where the second path may be the same as or different from the first path.
5. The method as described in claim 1, characterized in that, The exclusive processing rights for the target address are obtained by the first node from the second node before receiving the message sequence.
6. The method as described in claim 5, characterized in that, The method further includes: Before receiving the message sequence, the first node obtains exclusive processing rights for a specified address range from the second node, the specified address range including the target address corresponding to the message in the message sequence.
7. The method as described in claim 6, characterized in that, The step of obtaining exclusive processing permissions for a specified address range from the second node includes: The first node determines the specified address range based on the addresses belonging to the storage space managed by the second node from the target addresses corresponding to historical data access operations; The first node obtains exclusive processing rights for the specified address range from the second node.
8. The method according to any one of claims 1-7, characterized in that, The first node, based on its exclusive processing rights to the target address, performs data access operations on the target address, including: The first node obtains the cache address corresponding to the target address from the second node; The first node performs data access operations on the cache address corresponding to the target address based on the exclusive processing permission of the target address.
9. The method according to any one of claims 1-7, characterized in that, After the first node performs data access operations on the target address based on its exclusive processing permissions for the target address, the process further includes: The first node releases the exclusive processing rights of the target address.
10. The method according to any one of claims 1-7, characterized in that, Also includes: After the first node obtains exclusive processing rights to the target address, if no data access operation is performed on the target address within a set time period, the exclusive processing rights to the target address are released.
11. The method according to any one of claims 1-7, characterized in that, Also includes: The first node sends the execution results of the data access operation on the target address to the external node according to the constraint of the visible order of the data access operation execution results of the external node.
12. The method according to any one of claims 1-7, characterized in that, The first node and the second node are both System-on-a-Chip (SoC) chips, and the external node is an input / output (I / O) device.
13. An interconnection system comprising at least two nodes interconnected in a horizontally expanded manner, wherein the at least two nodes interconnected in a horizontally expanded manner comprise a first node and a second node, characterized in that: The first node is used for: Receive a message sequence from an external node, wherein the messages in the message sequence are used to request data access to a target address, the target address belongs to the storage space managed by the second node, the external node is a node outside the interconnection system, and the external node has a stricter constraint on the visible order of the data access operation execution results than the first node. Obtain exclusive processing rights for the target address; Based on the exclusive processing permission of the target address, data access operations are performed on the target address, wherein the visibility order of the execution results of the data access operations satisfies the constraint of the external node on the visibility order of the execution results of the data access operations.
14. The interconnection system as described in claim 13, characterized in that, The first node is specifically used for: In response to receiving the message sequence, exclusive processing rights for the target address are obtained from the second node.
15. The interconnection system as described in claim 14, characterized in that: The first node is specifically used to send an access request to the second node after receiving the message sequence, the access request carrying the target address corresponding to at least one message in the message sequence; The second node is configured to send a permission acquisition response to the first node after receiving the permission acquisition request. The permission acquisition response is configured to indicate that the exclusive processing permission of the target address corresponding to the at least one message is transferred to the first node.
16. The interconnection system as described in claim 15, characterized in that: The first node is specifically used to receive the message sequence and then send the permission acquisition request through a first path between it and the second node, wherein the first path is selected by the first node based on the congestion status of each path between it and the second node; The second node is specifically used to send the permission acquisition response from a second path between itself and the first node, wherein the second path may be the same as or different from the first path.
17. The interconnection system as described in claim 13, characterized in that, The exclusive processing rights for the target address are obtained by the first node from the second node before receiving the message sequence.
18. The interconnection system as claimed in claim 17, characterized in that, The first node is also used for: Before receiving the message sequence, the system obtains data access permissions from the second node for a specified address range within the storage space managed by the second node, wherein the specified address range includes the target address corresponding to the message in the message sequence.
19. The interconnection system as described in claim 18, characterized in that, The first node is specifically used for: The specified address range is determined based on the addresses belonging to the storage space managed by the second node from the target addresses corresponding to historical data access operations. Obtain exclusive processing rights for the specified address range from the second node.
20. The interconnection system according to any one of claims 13-19, characterized in that, The first node is specifically used for: Obtain the cache address corresponding to the target address from the second node; Based on the exclusive processing permission of the target address, data access operations are performed on the cache address corresponding to the target address.
21. The interconnection system according to any one of claims 13-19, characterized in that, The first node is also used for: Based on the exclusive processing permission of the target address, after performing data access operations on the target address, the exclusive processing permission of the target address is released.
22. The interconnection system according to any one of claims 13-19, characterized in that, The first node is also used for: After obtaining exclusive processing rights to the target address, if no data access operation is performed on the target address within a set time period, the exclusive processing rights to the target address are released.
23. The interconnection system according to any one of claims 13-19, characterized in that, The first node is also used for: Based on the constraint of the visible order of the data access operation execution results of the external nodes, the execution results of the data access operation on the target address are sent to the external nodes.
24. The interconnection system according to any one of claims 13-19, characterized in that, The second node is used for: After the first node obtains exclusive processing rights to the first target address, if the first node fails to return the exclusive processing rights to the first target address within a set time period, it will regain exclusive processing rights to the first target address, wherein the first target address is the target address corresponding to the first message in the message sequence.
25. The interconnection system according to any one of claims 13-19, characterized in that, The first node and the second node are both System-on-a-Chip (SoC) chips, and the external node is an input / output (I / O) device.
26. A system-on-a-chip (SoC) chip, characterized in that, include: One or more processors and one or more memories; wherein the one or more memories store one or more computer programs, the one or more computer programs including instructions that, when executed by the one or more processors, cause the SoC chip to perform the method as described in any one of claims 1-12.
27. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program that, when run on a computing device, causes the computing device to perform the method as described in any one of claims 1-12.
28. A computer program product, characterized in that, When the computer program product is invoked by a computer, it causes the computer to perform the method as described in any one of claims 1-12.
Citation Information
Patent Citations
RDMA (Remote Direct Memory Access) communication method between non-tightly-coupled systems of sharing system address space
CN104202391A
Network on chip system and establishment method of network on chip communication link
CN105450555A