Multi-source heterogeneous distributed system, memory access method, and storage medium
By deploying a unified interconnect bus unit and protocol adaptation interface module in a multi-source heterogeneous distributed system, communication problems between different types of devices are solved, memory sharing and flexible topological interconnection are realized, and system performance is improved.
Patent Information
- Application Number
- PCT/CN2024/139303
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-13
- Publication Date
- 2025-07-03
AI Technical Summary
The inability to communicate effectively between different types of processors and accelerators in a multi-source heterogeneous distributed system, resulting in limited memory expansion access and inability to perform optimal performance.
Deploy a unified interconnection bus unit and protocol adaptation interface module on each device to realize the conversion of a consistent protocol interface and a unified interconnection bus protocol for different types of devices, and realize memory sharing and flexible topological interconnection through a unified interconnection bus unit.
It realizes memory sharing between different types of devices, alleviates the problems of memory wall and IO wall, and improves the performance of multi-source heterogeneous distributed systems.
Smart Images

Figure CN2024139303_03072025_PF_FP_ABST
Abstract
Description
A multi-source heterogeneous distributed system, memory access method and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 28, 2023, with application number 202311843443.2, entitled “A Multi-source Heterogeneous Distributed System, Memory Access Method and Storage Medium,” the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present application relates to the field of data storage technology, and in particular to a multi-source heterogeneous distributed system, a memory access method, and a non-volatile readable storage medium. Background Art
[0004] As Moore's Law slows, large-scale parallel processing of large amounts of data becomes increasingly demanding. This requires more than just a single device to quickly complete the task. Instead, it requires combining multiple types of devices for processing, a process known as heterogeneous processing. To ensure efficient communication between these devices, various cache coherent interconnect protocol standards have been proposed.
[0005] Different types of processors use different off-chip cache coherence buses to interconnect with input / output (IO) peripherals. Furthermore, the interconnection bus between processors is only for processors of the same type and cannot be used to interconnect other processors. Because different types of processors use different interconnection buses, intercommunication between different types of processors is impossible. Processors and IO peripherals also use different bus protocols, requiring different processing. Currently, this interconnection topology requires the implementation of multiple types of interconnection buses, resulting in a complex implementation structure. Furthermore, memory between different types of devices cannot be shared, limiting memory expansion access and preventing optimal performance of multi-source heterogeneous distributed systems.
[0006] It can be seen that how to improve the performance of multi-source heterogeneous distributed systems is a problem that technical personnel in this field need to solve. Summary of the Invention
[0007] The purpose of some embodiments of the present application is to provide a multi-source heterogeneous distributed system, a memory access method and a non-volatile readable storage medium, which can solve the problem of poor performance of the multi-source heterogeneous distributed system.
[0008] To solve the above technical problems, some embodiments of the present application provide a multi-source heterogeneous distributed system, comprising at least two processors and at least one accelerator; wherein each processor and accelerator has its own corresponding memory unit; each processor and accelerator is deployed with a unified interconnect bus unit; a protocol adapter interface module is deployed on the unified interconnect bus unit; the protocol adapter interface module is used to implement conversion between the consistency protocol interface of different types of devices and the unified interconnect bus protocol interface;
[0009] A first unified interconnect bus unit is configured to receive a read request sent by a first device through a protocol adapter interface module, set a cache state of a first memory unit of the first device to an invalid state, encapsulate the read request into a request message conforming to a set message format, and send the request to a second unified interconnect bus unit of a second device; wherein the first device and the second device are any two devices in each processor and accelerator;
[0010] The second unified interconnect bus unit is configured to, after receiving the request message, read data from the second memory unit of the second device; set the cache state of the second memory unit to a shared state, and encapsulate the data into a response message that conforms to the message format; and send the response message to the first unified interconnect bus unit;
[0011] The first unified interconnection bus unit is configured to store the data carried in the received response message into the first memory unit; and set the cache state of the first memory unit to a shared state according to the cache state carried in the response message.
[0012] On the one hand, the unified interconnection bus unit includes a protocol layer, an adaptation layer and a physical layer; wherein, the protocol layer includes a protocol adaptation interface module, a request queue management module, a response queue management module, a message parsing and encapsulation module and a clock domain conversion module; wherein, the protocol adaptation interface module is respectively connected to the request queue management module and the response queue management module, and is used to transmit the received request to the request queue management module or the response queue management module according to the request type; the message parsing and encapsulation module is respectively connected to the request queue management module, the response queue management module and the clock domain conversion module, and is used to realize the encapsulation and decapsulation of the message.
[0013] On the one hand, the first protocol adaptation interface module of the first unified interconnection bus unit is used to receive a read request sent by the first device and transmit the read request to the first request queue management module through the internal bus;
[0014] A first request queue management module is configured to receive a read request; set the cache state of the first memory unit of the first device to an invalid state; query the node forwarding table for a second device address that matches the second identifier according to the second identifier of the second device carried in the read request; and transmit the read request and the second device address to the first message parsing and encapsulation module;
[0015] a first message parsing and encapsulating module, configured to encapsulate the received read request according to a message format to obtain a request message; and transmit the request message and the second device address to the first clock domain conversion module;
[0016] A first clock domain conversion module, configured to transmit the request message and the second device address to the hub through the first adaptation layer and the first physical layer;
[0017] The hub is configured to forward the request message to the second device corresponding to the second device address.
[0018] On the one hand, the first protocol adaptation interface module of the first unified interconnection bus unit is used to receive a read request sent by the first device and transmit the read request to the first request queue management module through the internal bus;
[0019] A first request queue management module is configured to receive a read request; set a cache state of a first memory unit of a first device to an invalid state; and transmit the read request to a first message parsing and encapsulation module;
[0020] A first message parsing and encapsulating module is used to encapsulate the received read request according to the message format to obtain a request message; and transmit the request message to the first clock domain conversion module;
[0021] A first clock domain conversion module, configured to transmit the request message to the hub through the first adaptation layer and the first physical layer;
[0022] The hub is configured to query the node forwarding table for a second device address that matches the second identifier of the second device carried in the request message, and forward the request message to the second device corresponding to the second device address.
[0023] On the one hand, the first message parsing and encapsulation module is used to encapsulate the first identifier of the first device, the second identifier of the second device to which the read request points, the message type and message sequence number to which the read request belongs, and the second memory address according to the identifier of the source device, the identifier of the destination device, the message length, the message type, the message sequence number, the memory address, the memory data and the message format of the write enable signal to obtain a request message; wherein the second memory address carries a status identifier that the first memory unit is in an invalid state.
[0024] On the one hand, the second clock domain conversion module of the second unified interconnection bus unit is used to receive the request message transmitted by the hub through the second physical layer and the second adaptation layer, and forward the request message to the second message parsing and encapsulation module;
[0025] A second message parsing and encapsulation module is used to parse the request message to obtain the first identifier, message sequence number and second memory address of the first device; and send the second memory address to the second response queue management module;
[0026] A second response queue management module is configured to read data from the second memory unit according to the second memory address and transmit the data to the second message parsing and encapsulation module; and set the cache state of the second memory unit to a shared state;
[0027] The second message parsing and encapsulation module is configured to receive data; query the node forwarding table for a first device address that matches the first identifier based on the first identifier; encapsulate the second identifier, the first identifier, the length of the data, the message type to which the response request belongs, the message sequence number, the first memory address, the data, and the write enable signal according to the message format to obtain a response message; and transmit the response message and the first device address to the second clock domain conversion module; wherein the first memory address carries a status identifier indicating that the second memory unit is in a shared state;
[0028] A second clock domain conversion module, configured to transmit the response message and the first device address to the hub via the second adaptation layer and the second physical layer;
[0029] The hub is configured to forward the response message to the first device corresponding to the first device address.
[0030] On the one hand, the second clock domain conversion module of the second unified interconnection bus unit is used to receive the request message transmitted by the hub through the second physical layer and the second adaptation layer, and forward the request message to the second message parsing and encapsulation module;
[0031] A second message parsing and encapsulation module is used to parse the request message to obtain the first identifier, message sequence number and second memory address of the first device; and send the second memory address to the second response queue management module;
[0032] A second response queue management module is configured to read data from the second memory unit according to the second memory address and transmit the data to the second message parsing and encapsulation module; the cache state of the second memory unit is set to a shared state;
[0033] The second message parsing and encapsulation module is configured to receive data; query the first memory address that matches the first identifier from the node forwarding table according to the first identifier; encapsulate the second identifier, the first identifier, the length of the data, the message type to which the response request belongs, the message sequence number, the first memory address, the data, and the write enable signal according to the message format to obtain a response message; and transmit the response message to the second clock domain conversion module; wherein the first memory address carries a status identifier indicating that the second memory unit is in a shared state;
[0034] A second clock domain conversion module, configured to transmit the response message to the hub via the second adaptation layer and the second physical layer;
[0035] The hub is configured to query a first device address matched with the first identifier from a node forwarding table according to the first identifier carried in the response message; and forward the response message to a first device corresponding to the first device address.
[0036] On the one hand, the first unified interconnection bus unit is used to encapsulate the first identifier, the second identifier, the message type to which the monitoring belongs, and the message sequence number according to the message format to obtain a monitoring message; and transmit the monitoring message and the second memory address to the first clock domain conversion module;
[0037] A first clock domain conversion module is used to transmit the monitoring message and the second device address to the hub through the first adaptation layer and the first physical layer;
[0038] The hub is configured to forward the monitoring message to the second device corresponding to the second device address.
[0039] On the one hand, the first unified interconnection bus unit is used to encapsulate the first identifier, the second identifier, the message type to which the monitoring belongs, and the message sequence number according to the message format to obtain the monitoring message; and transmit the monitoring message to the first clock domain conversion module;
[0040] A first clock domain conversion module, configured to transmit the monitored message to the hub via the first adaptation layer and the first physical layer;
[0041] The hub is configured to query the node forwarding table for the second device address matched by the second identifier of the second device carried in the monitoring message, and forward the monitoring message to the second device corresponding to the second device address.
[0042] On the one hand, the third unified interconnection bus unit is used to receive a write request sent by a third device, set the cache state of the third memory unit of the third device to an invalid state; encapsulate the data to be written into a write message that conforms to the message format; and send the write message to the first unified interconnection bus unit through the hub;
[0043] The first unified interconnect bus unit is configured to receive a write message fed back by the hub, store the data to be written carried in the write message in the first memory unit; set the cache state of the first memory unit to a unique clean state according to the cache state carried in the write message; and send a monitoring message to the second unified interconnect bus unit through the hub;
[0044] The second unified interconnection bus unit is configured to receive a monitoring message sent by the hub; and set the cache state of the second memory unit to an invalid state according to the invalid state carried in the monitoring message.
[0045] On the one hand, the third protocol adaptation interface module of the third unified interconnection bus unit is used to receive a write request sent by a third device and transmit the write request to the third request queue management module via the internal bus;
[0046] a third request queue management module configured to receive a write request; set the cache state of the third memory unit of the third device to an invalid state; query the node forwarding table for a first device address that matches the first identifier carried in the write request; and transmit the write request and the first device address to the third message parsing and encapsulation module;
[0047] a third message parsing and encapsulating module, configured to encapsulate the received write request according to the message format to obtain a write message; and transmit the write message and the first device address to the third clock domain conversion module;
[0048] a third clock domain conversion module, configured to transmit the write message and the first device address to the hub via the third adaptation layer and the third physical layer;
[0049] The hub is configured to forward the write message to the first device corresponding to the first device address.
[0050] On the one hand, the third protocol adaptation interface module of the third unified interconnection bus unit is used to receive a write request sent by a third device and transmit the write request to the third request queue management module via the internal bus;
[0051] a third request queue management module configured to receive a write request; set the cache state of the third memory unit of the third device to an invalid state; query the node forwarding table for a first device address that matches the first identifier carried in the write request; and transmit the write request and the first device address to the third message parsing and encapsulation module;
[0052] a third message parsing and encapsulating module, configured to encapsulate the received write request according to the message format to obtain a write message; and transmit the write message to the third clock domain conversion module;
[0053] A third clock domain conversion module, configured to transmit the write message to the hub via the third adaptation layer and the third physical layer;
[0054] The hub is configured to query a first device address matched with the first identifier from a node forwarding table according to the first identifier of the first device carried in the write message; and forward the write message to the first device corresponding to the first device address.
[0055] On the one hand, the third message parsing and encapsulation module is used to encapsulate the third identifier of the third device, the first identifier of the first device to which the write request points, the message type and message sequence number to which the write request belongs, the first memory address, and the data to be written according to the message format of the source device identifier, the destination device identifier, the message length, the message type, the message sequence number, the memory address, the memory data, and the write enable signal to obtain a write message; wherein the first memory address carries a status identifier that the third memory unit is in an invalid state.
[0056] On the one hand, both the adaptation layer and the physical layer are provided with a bypass unit; wherein the bypass unit is used to realize a direct connection between the adaptation layer and the physical layer;
[0057] When the first device and the second device are located on the same printed circuit board, the bypass units of the first adaptation layer and the first physical layer of the first device and the second adaptation layer and the second physical layer of the second device are turned on to achieve direct connection between the first device and the second device.
[0058] On the one hand, it also includes a non-volatile readable storage medium that is set independently of each processor and accelerator; a unified interconnection bus unit is deployed on the non-volatile readable storage medium.
[0059] On the one hand, each unified interconnection bus unit is used to determine a matching cache state based on the state of data on its corresponding memory unit; wherein the cache state includes an invalid state, a unique state, and a shared state; the unique state includes a unique clean state, a unique dirty state, a unique clean idle state, and a unique partially dirty state; and the shared state includes a shared clean state and a shared dirty state.
[0060] Some embodiments of the present application further provide a memory access method, including:
[0061] Based on the protocol adapter interface module receiving a read request sent by the first device, setting the cache state of the first memory unit of the first device to an invalid state; wherein the protocol adapter interface module is used to implement conversion between the consistency protocol interface of different types of devices and the first unified interconnect bus protocol interface;
[0062] encapsulate the read request into a first request message that complies with a set message format according to the first identifier of the first device and the second identifier of the second device to which the read request points;
[0063] Sending the first request message to the second unified interconnection bus unit of the second device;
[0064] receiving a response message fed back by the second unified interconnection bus unit, and storing data carried in the response message into the first memory unit;
[0065] According to the cache state carried in the response message, the cache state of the first memory unit is set to a shared state.
[0066] On the one hand, it also includes:
[0067] receiving a second request message sent by the hub; wherein the second request message is transmitted to the hub by the second unified interconnection bus unit of the second device;
[0068] Reading data from the first memory unit according to the first memory address carried in the second request message;
[0069] Setting the cache state of the first memory unit to a shared state, and encapsulating the data into a response message that conforms to a message format;
[0070] The response message is sent to the second device through the hub.
[0071] On the one hand, it also includes:
[0072] receiving a write message fed back by the hub, and storing the data to be written carried in the write message in the first memory unit; wherein the write message is transmitted to the hub by the third unified interconnection bus unit of the third device;
[0073] According to the cache state carried in the write message, the cache state of the first memory unit is set to a unique clean state;
[0074] The monitoring message is sent to the second unified interconnection bus unit through the hub, so that the second unified interconnection bus unit receives the monitoring message sent by the hub; according to the invalid state carried in the monitoring message, the cache state of the second memory unit is set to invalid state.
[0075] Some embodiments of the present application further provide a non-volatile readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned memory access method are implemented.
[0076] As can be seen from the above technical solution, a multi-source heterogeneous distributed system includes at least two processors and at least one accelerator; each processor and accelerator has its own corresponding memory unit; each processor and accelerator is deployed with a unified interconnect bus unit; a protocol adapter interface module is deployed on the unified interconnect bus unit; the protocol adapter interface module is used to implement conversion between the consistency protocol interface of different types of devices and the unified interconnect bus protocol interface. The first device and the second device are any two devices among the processors and accelerators. The first unified interconnect bus unit deployed on the first device is used to receive a read request sent by the first device through the protocol adapter interface module. Because the first memory unit currently does not contain data, the cache state of the first memory unit of the first device can be set to invalid state. The read request is encapsulated into a request message that conforms to a set message format and sent to the second unified interconnect bus unit of the second device. The second unified interconnect bus unit is used to read data from the second memory unit of the second device after receiving the request message. Because data sharing is required, the cache state of the second memory unit can be set to shared state and the data can be encapsulated into a response message that conforms to the message format. The response message is sent to the first unified interconnect bus unit. The first unified interconnection bus unit is used to store the data carried by the received response message into the first memory unit; according to the cache status carried in the response message, the cache status of the first memory unit is set to a shared state. The beneficial effect of the present application is that by deploying a unified interconnection bus unit on each device and deploying a protocol adapter interface module on the unified interconnection bus unit, compatibility with multiple different types of devices can be achieved, so that processors with multiple different types of instruction sets, different types of accelerators, and different memory units can be unified in one system to form a computer system, which can realize memory sharing between different processors and between different processors and accelerators, forming a large memory pool, and alleviating the memory wall and IO wall problems. The unified interconnection bus unit can be used to interconnect different types of devices through different topological forms, support flexible interconnection topology, and can flexibly expand the scale without affecting the existing deployment. By recording the cache status of the memory unit, consistent memory communication between devices in the system is achieved. According to the deployment method of the present application, the performance of the multi-source heterogeneous distributed system is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate some embodiments of the present application, the following will briefly introduce the drawings required for use in some embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0078] FIG1 is a schematic diagram of the structure of a multi-source heterogeneous distributed system provided by some embodiments of the present application;
[0079] FIG2 is a schematic diagram of the structure of a protocol adaptation interface module provided in some embodiments of the present application;
[0080] FIG3 is a schematic diagram of an architecture for implementing cache coherent interconnection between multiple types of processors and multiple types of accelerators according to some embodiments of the present application;
[0081] FIG4 is a block diagram of internal functional modules of a unified interconnect bus unit provided in some embodiments of the present application;
[0082] FIG5 is a flowchart of a memory access method provided in some embodiments of the present application. DETAILED DESCRIPTION
[0083] The following will be combined with the accompanying drawings of some embodiments of the present application to clearly and completely describe the technical solutions of some embodiments of the present application. Obviously, some of the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on some embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of this application.
[0084] The terms "including" and "having," as used in the specification and accompanying drawings of this application, and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements and may include steps or elements that are not listed.
[0085] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0086] With the advent of the artificial intelligence era, the memory wall and input / output (IO) wall problems are becoming increasingly serious. The memory wall problem primarily arises from the inability of memory bandwidth to keep up with the rapidly growing number of central processing unit (CPU) cores, resulting in memory bandwidth becoming a bottleneck. The IO wall problem primarily arises from insufficient memory capacity, which slows access to external storage. How to effectively expand memory to address these issues has become a pressing issue for the industry.
[0087] The most common method of memory expansion is to add memory channels to the processor. Currently, 8 and even 12 channels are supported, but this number cannot be increased indefinitely. Each additional memory channel increases the number of signals, which poses significant challenges to processor power consumption, heat dissipation, packaging, and printed circuit board (PCB) design. Of course, some new non-volatile readable storage media have emerged on the market, but the current Double Data Rate (DDR) interface is not compatible with multiple non-volatile readable storage media.
[0088] The PCI-Express (PCIE) interface is a bus type commonly used by high-performance I / O devices. The PCIE 6.0 standard has been officially released, with a maximum speed of 128 GBps (switching bandwidth). However, the PCIE interface does not support cache coherence transactions and also has issues with memory address space isolation, making it impossible to expand host memory.
[0089] Furthermore, as Moore's Law slows, processing large amounts of data requires massive parallel processing. This can no longer rely on a single, fast processing component. Instead, many types of processors must be combined, creating a phenomenon known as heterogeneous processing. Efficient communication between these devices also requires a new bus protocol.
[0090] Currently, the commonly used off-chip cache coherent buses for interconnecting processors and IO peripherals mainly include the multi-protocol interconnect technology bus (Compute Express Link, CXL), the cache coherent interconnect protocol for accelerators (Cache Coherent Interconnect for Accelerators, CCIX) of the microprocessor (Advanced RISC Machines, ARM), and the bus and its communication protocol (NVlink).
[0091] Interprocessor interconnect buses primarily include the Ultra Path Interconnect (UPI) bus, Advanced Micro Devices (AMD)'s Infinity Fabric (IF) bus, ARM's Corelink series, and custom coherent buses developed by Chinese CPU manufacturers. Interprocessor interconnect buses enable access to memory between different processors and also enable memory expansion.
[0092] Currently, processor bus interconnects and processor-IO peripheral interconnects typically use different protocols. For ease of description, consider three types of processors and three types of accelerators. For ease of description, the three types of processors can be classified as first-, second-, and third-type processors; the three types of accelerators are field programmable gate arrays (FPGAs), graphics processing units (GPUs), and digital shoreline analysis systems (DSASs). First-type processors can use the Ultra Path Interconnect (UPI) interprocessor bus, second-type processors can use the CoreLink bus, and third-type processors can use the Huawei Cache Coherent System (HCCS). First-type processors and FPGAs use the Multiprotocol Interconnect (CXL); second-type processors and GPUs use the Cache Coherent Interconnect (CCIX); and third-type processors and the DSAS use the high-speed Serial Computer Interface eXtension (PCIE) bus.
[0093] The CXL bus is an asymmetric, master-slave bus for off-chip cache coherence interconnects. It uses a master-slave model and cannot directly interconnect masters or slaves, limiting its use cases. Furthermore, using the CXL bus controller with other processors requires purchasing a license (software license agreement or authorization), which is expensive. The CCIX interconnect bus has high latency, complex implementation, and low memory read performance. NVLink is a proprietary protocol that cannot be used by other processors. Interprocessor interconnects are specific to specific processors.
[0094] Server systems are composed of heterogeneous processors and accelerators from multiple sources. Different processors use different interconnect buses, resulting in intercommunication problems. Different bus protocols are also used between processors and between processors and I / O peripherals, requiring separate processing. Current interconnect topologies require multiple interconnect buses, making implementation complex. Furthermore, memory cannot be shared between different server types, limiting access to memory expansion and hindering the optimal performance of multi-source heterogeneous distributed systems.
[0095] Therefore, some embodiments of the present application provide a multi-source heterogeneous distributed system, a memory access method and a non-volatile readable storage medium, each processor and accelerator has its own corresponding memory unit; a unified interconnection bus unit is deployed on each processor and accelerator; a protocol adapter interface module is deployed on the unified interconnection bus unit, which can achieve compatibility with multiple different types of devices. By deploying a unified interconnection bus unit on each device, processors with multiple different types of instruction sets, different types of accelerators, and different memory units can be unified in one system to form a computer system, which can achieve memory sharing between different processors and between different processors and accelerators, forming a large memory pool, alleviating the memory wall and IO wall problems, and greatly improving the performance of the multi-source heterogeneous distributed system. The unified interconnection bus unit can be used to interconnect different types of devices through different topological forms, achieve consistent memory communication between devices in the system, support flexible interconnection topology, and can flexibly expand the scale without affecting the existing deployment.
[0096] Next, a multi-source heterogeneous distributed system provided by some embodiments of the present application is introduced in detail. Figure 1 is a structural diagram of a multi-source heterogeneous distributed system provided by some embodiments of the present application, the system comprising at least two processors 11 and at least one accelerator 12; wherein each processor 11 and accelerator 12 has its own corresponding memory unit; each processor 11 and accelerator 12 is deployed with a unified interconnection bus unit 13; a protocol adapter interface module is deployed on the unified interconnection bus unit 13; the protocol adapter interface module is used to realize a conversion hub between the consistency protocol interface of different types of devices and the unified interconnection bus protocol interface.
[0097] Figure 1 shows a schematic diagram using two processors 11 and one accelerator 12 as an example. The actual number of processors 11 and accelerators 12 can be determined based on the actual business needs of the multi-source heterogeneous distributed system and is not limited here. The two processors 11 can be of the same or different types and are not limited here.
[0098] In some embodiments of the present application, for ease of description, the processor 11 and the accelerator 12 may be collectively referred to as devices. The data processing methods between different devices are similar, and any two devices in each processor 11 and accelerator 12, namely the first device and the second device, are taken as an example for detailed introduction.
[0099] For ease of distinction, the unified cache consistency bus (UCCB) unit 13 deployed on the first device may be referred to as a first unified interconnect bus unit 13, and the unified interconnect bus unit 13 deployed on the second device may be referred to as a second unified interconnect bus unit 13. Since the first unified interconnect bus unit 13 and the second unified interconnect bus unit 13 are deployed on different devices but have the same architecture, they are represented by the same number in some embodiments of the present application.
[0100] The first unified interconnection bus unit 13 is used to receive a read request sent by the first device through the protocol adapter interface module, set the cache status of the first memory unit of the first device to an invalid state; encapsulate the read request into a request message that complies with the set message format; and send the request message to the second unified interconnection bus unit 13 of the second device.
[0101] In actual applications, the first unified interconnection bus unit 13 may encapsulate the read request into a request message that complies with a set message format according to the first identifier of the first device and the second identifier of the second device to which the read request is directed.
[0102] The second unified interconnection bus unit 13 is used to read data from the second memory unit of the second device after receiving the request message; set the cache state of the second memory unit to a shared state, and encapsulate the data into a response message that conforms to the message format; and send the response message to the first unified interconnection bus unit 13.
[0103] In actual applications, the second unified interconnection bus unit 13 may read data from the second memory unit of the second device according to the second memory address carried in the request message.
[0104] The first unified interconnection bus unit 13 is configured to store the data carried in the received response message into the first memory unit; and set the cache state of the first memory unit to a shared state according to the cache state carried in the response message.
[0105] In actual applications, the unified interconnection bus unit 13 deployed on each processor 11 and accelerator 12 can be connected to the hub to support different types of topologies.
[0106] The first unified interconnect bus unit 13 can send the request message to the second unified interconnect bus unit 13 of the second device through the hub; the second unified interconnect bus unit 13 can receive the request message through the hub and send the response message to the first unified interconnect bus unit 13 through the hub.
[0107] In some embodiments of the present application, the protocol adapter interface module can adapt to the consistency protocol interface of different types of devices. Different processors have different external consistency protocol interfaces. By deploying the protocol adapter interface module on the unified interconnect bus unit 13, conversion between the consistency protocol interface of different types of devices and the unified interconnect bus protocol interface, i.e., the UCCB interface, can be achieved.
[0108] Fig. 2 is a structural representation of a protocol adaptation interface module provided by some embodiments of the present application. In Fig. 2, taking the processor that adapts three types as an example, three processor protocol conversion submodules can be deployed in the protocol adaptation interface module. For ease of description, they can be referred to as the first processor protocol conversion submodule, the second processor protocol conversion submodule and the third processor protocol conversion submodule. Each processor protocol conversion submodule is used to adapt the consistency protocol interface of a type of processor. The three processor protocol conversion submodules are connected to a multiplexer (MUX / DEMUX) respectively.
[0109] In actual applications, by modifying the protocol adaptation interface module, such as adding a device protocol conversion submodule in the protocol adaptation interface module, more types of devices can be adapted without modifying other modules.
[0110] Figure 3 is a schematic diagram of an architecture for implementing cache consistency interconnection between multiple types of processors and multiple types of accelerators provided by some embodiments of the present application. For ease of distinction, the multiple types of processors can be referred to as first-type processors, second-type processors, third-type processors, and fourth-type processors, respectively. The multiple types of accelerators are referred to as first-type accelerators, second-type accelerators, and third-type accelerators, respectively. In Figure 3, each processor and each accelerator has its corresponding memory unit. In order to achieve interconnection between the devices, a unified interconnection bus unit is set on each processor and each accelerator. In Figure 3, an independent non-volatile readable storage medium is also provided, which can be used to expand the memory of a multi-source heterogeneous distributed system. A unified interconnection bus unit is also provided on the independent non-volatile readable storage medium, so that interconnection with memory units on other devices can be achieved.
[0111] In some embodiments of the present application, a unified interconnect bus unit 13 may be deployed on each device. The unified interconnect bus unit 13 may include a protocol layer, an adaptation layer, and a physical layer.
[0112] Figure 4 is a block diagram of the internal functional modules of a unified interconnect bus unit provided in some embodiments of the present application. The unified interconnect bus unit may include a protocol layer, an adaptation layer, and a physical layer. The protocol layer may include a protocol adaptation interface module, a request queue management module, a response queue management module, a message parsing and encapsulation module, and a clock domain conversion module. The protocol adaptation interface module is connected to the request queue management module and the response queue management module respectively, and is used to transmit the received request to the request queue management module or the response queue management module according to the request type; the message parsing and encapsulation module is connected to the request queue management module, the response queue management module, and the clock domain conversion module respectively, and is used to implement the encapsulation and decapsulation of the message. The adaptation layer may include a protocol arbitration module, a connection state management module, a cyclic redundancy check code (CRC) control and retransmission module, and a bypass module. The physical layer may include a link control module, a mapping and remapping module for the LAN emulation (Lane) path, a scrambling and descrambling module, and a bypass module.
[0113] The physical layer and adaptation layer can be based on the open (UCIe, Universal Chiplet Interconnect Express) interconnect standard. Designed based on the UCIe interconnect standard, the physical layer and adaptation layer offer lower latency than interconnect buses like PCIE-based CXL and CCIX, and are compatible with multiple protocols. Future protocol optimizations will not affect the physical and adaptation layers. The physical layer can be used to implement functions such as link initialization and training. The adaptation layer can select and arbitrate between multiple protocols and is also responsible for link state management. Bypass modules in the physical and adaptation layers are optional functional modules in both layers.
[0114] In some embodiments of the present application, the bypass unit can be used to achieve direct connection between the adaptation layer and the physical layer; when the first device and the second device are located on the same printed circuit board, the bypass units of the first adaptation layer and the first physical layer of the first device and the second adaptation layer and the second physical layer of the second device are in an open state to achieve direct connection between the first device and the second device.
[0115] In practical applications, when interconnecting homogeneous devices, that is, interconnecting multiple devices on the same PCB motherboard, because the link is known, the link parameters can be statically configured, simplifying the functions of the physical layer and adaptation layer, basically achieving transparent transmission and minimizing latency.
[0116] During the initialization phase, initialization of the physical and adaptation layers includes link training, protocol and parameter negotiation with remote nodes, discovery and enumeration of each node device, and configuration and querying of protocol registers. Protocol layer initialization includes configuring unique node IDs and node forwarding tables to ensure unrestricted forwarding between nodes.
[0117] In some embodiments of the present application, for ease of distinction, the protocol adaptation interface module included in the protocol layer of the first unified interconnection bus unit 13 may be referred to as the first protocol adaptation interface module, the request queue management module may be referred to as the first request queue management module, the response queue management module may be referred to as the first response queue management module, the message parsing and encapsulation module may be referred to as the first message parsing and encapsulation module, and the clock domain conversion module may be referred to as the first clock domain conversion module. Similarly, the protocol adaptation interface module included in the protocol layer of the second unified interconnection bus unit 13 may be referred to as the second protocol adaptation interface module, the request queue management module may be referred to as the second request queue management module, the response queue management module may be referred to as the second response queue management module, the message parsing and encapsulation module may be referred to as the second message parsing and encapsulation module, and the clock domain conversion module may be referred to as the second clock domain conversion module.
[0118] Taking the first device reading data from the second device as an example, the first protocol adaptation interface module of the first unified interconnection bus unit 13 can be used to receive the read request sent by the first device and transmit the read request to the first request queue management module through the internal bus.
[0119] The first request queue management module is used to receive a read request; set the cache status of the first memory unit of the first device to an invalid state; query the second device address matched by the second identifier from the node forwarding table based on the second identifier of the second device carried in the read request; and transmit the read request and the second device address to the first message parsing and encapsulation module.
[0120] In some embodiments of the present application, a device can be considered a node, and the node forwarding table can record the correspondence between the identifier and device address of each device in the multi-source heterogeneous distributed system. The device identifier (ID) is unique and can be used to distinguish different devices.
[0121] The first message parsing and encapsulating module is used to encapsulate the received read request according to the message format to obtain a request message; and transmit the request message and the second device address to the first clock domain conversion module.
[0122] The first clock domain conversion module is configured to transmit the request message and the second device address to the hub through the first adaptation layer and the first physical layer. The hub is configured to forward the request message to the second device corresponding to the second device address.
[0123] In addition to the hub's role as a transit in the above introduction, the hub can also assume the function of device identification. Taking the first device reading data from the second device as an example, the first protocol adapter interface module of the first unified interconnection bus unit 13 is used to receive the read request sent by the first device and transmit the read request to the first request queue management module through the internal bus. The first request queue management module is used to receive the read request; set the cache status of the first memory unit of the first device to an invalid state; and transmit the read request to the first message parsing and encapsulation module. The first message parsing and encapsulation module is used to encapsulate the received read request according to the message format to obtain a request message; and transmit the request message to the first clock domain conversion module. The first clock domain conversion module is used to transmit the request message to the hub through the first adaptation layer and the first physical layer. The hub is used to query the second device address matched by the second identifier from the node forwarding table based on the second identifier of the second device carried in the request message; and forward the request message to the second device corresponding to the second device address.
[0124] The message parsing and encapsulation module may include a pre-set message format. Table 1 is a schematic diagram of a message format provided in some embodiments of the present application. The message in Table 1 contains 8 parts: the source device identifier (source ID), the destination device identifier (destination ID), the message length (LEN), the message type (Type), the message sequence number (Tagid), the memory address (Addr), the memory data (Data), and the write enable signal (be).
[0125] Table 1
[0126] Each interconnected node in a multi-source heterogeneous distributed system has a unique ID number, which is used for routing, supporting a total of 256 nodes. Based on actual application requirements, message types can include read and write request messages, response messages, and monitoring messages. The message sequence number (Tagid) is used for packet loss processing such as retransmission. The memory address is 48 bits in total. The memory data is the actual transmitted number, and its write memory data bit width can be 256 bits or 512 bits. The write enable signal can be written byte by byte, with a bit width of data bit width / 8.
[0127] Taking the first device reading data from the second device as an example, the first message parsing and encapsulation module can encapsulate the first identifier of the first device, the second identifier of the second device pointed to by the read request, the message type and message sequence number to which the read request belongs, and the second memory address according to the identifier of the source device, the identifier of the destination device, the message length, the message type, the message sequence number, the memory address, the memory data and the message format of the write enable signal to obtain a request message; wherein the second memory address can carry a status identifier that the first memory unit is in an invalid state.
[0128] The above introduction is based on the operation process of each module of the first unified interconnection bus unit 13 on the first device when the first device reads data from the second device. Next, the operation process of each module on the second unified interconnection bus unit 13 will be introduced.
[0129] The second clock domain conversion module of the second unified interconnection bus unit 13 is configured to receive the request message transmitted by the hub through the second physical layer and the second adaptation layer, and forward the request message to the second message parsing and encapsulation module.
[0130] The second message parsing and encapsulating module is used to parse the request message to obtain the first identifier, message sequence number and second memory address of the first device; and send the second memory address to the second response queue management module.
[0131] The second response queue management module is used to read data from the second memory unit according to the second memory address and transmit the data to the second message parsing and encapsulation module; and set the cache state of the second memory unit to a shared state.
[0132] The second message parsing and encapsulation module is used to receive data; according to the first identifier, query the first device address matched by the first identifier from the node forwarding table; encapsulate the second identifier, the first identifier, the length of the data, the message type to which the response request belongs, the message sequence number, the first memory address, the data and the write enable signal according to the message format to obtain a response message; transmit the response message and the first device address to the second clock domain conversion module; wherein the first memory address carries a status identifier that the second memory unit is in a shared state.
[0133] The second clock domain conversion module is configured to transmit the response message and the first device address to the hub through the second adaptation layer and the second physical layer. The hub is configured to forward the response message to the first device corresponding to the first device address.
[0134] In addition to the hub's role as a transit point as described above, the hub can also assume the function of device identification. Taking the first device reading data from the second device as an example, the second clock domain conversion module of the second unified interconnection bus unit 13 is used to receive the request message transmitted by the hub through the second physical layer and the second adaptation layer, and forward the request message to the second message parsing and encapsulation module. The second message parsing and encapsulation module is used to parse the request message to obtain the first identifier, message sequence number and second memory address of the first device; and send the second memory address to the second response queue management module. The second response queue management module is used to read data from the second memory unit according to the second memory address, and transmit the data to the second message parsing and encapsulation module; the cache state of the second memory unit is set to a shared state. The second message parsing and encapsulation module is configured to receive data; query the node forwarding table for the first memory address that matches the first identifier based on the first identifier; encapsulate the second identifier, the first identifier, the length of the data, the message type to which the response request belongs, the message sequence number, the first memory address, the data, and the write enable signal according to the message format to obtain a response message; and transmit the response message to the second clock domain conversion module; wherein the first memory address carries a status identifier indicating that the second memory unit is in a shared state. The second clock domain conversion module is configured to transmit the response message to the hub through the second adaptation layer and the second physical layer; the hub is configured to query the node forwarding table for the first device address that matches the first identifier based on the first identifier carried in the response message; and forward the response message to the first device corresponding to the first device address.
[0135] Taking the example of a first device reading data from a second device, after the first device obtains the data transmitted from the second device, the first unified interconnect bus unit 13 can encapsulate the first identifier, the second identifier, the message type to which the monitoring belongs, and the message sequence number according to the message format to obtain a monitoring message; and transmit the monitoring message and the second memory address to the first clock domain conversion module. The first clock domain conversion module is used to transmit the monitoring message and the second device address to the hub through the first adaptation layer and the first physical layer. The hub is used to forward the monitoring message to the second device corresponding to the second device address, thus completing a complete read transaction request.
[0136] When sending a monitoring message, the hub, in addition to acting as a transit agent, can also assume the function of device identification. In actual applications, the first unified interconnection bus unit 13 can encapsulate the first identifier, the second identifier, the message type to which the monitoring belongs, and the message sequence number according to the message format to obtain a monitoring message; and transmit the monitoring message to the first clock domain conversion module. The first clock domain conversion module is used to transmit the monitoring message to the hub through the first adaptation layer and the first physical layer. The hub is used to query the second device address matched by the second identifier from the node forwarding table based on the second identifier of the second device carried in the monitoring message; and forward the monitoring message to the second device corresponding to the second device address.
[0137] In addition to reading memory data, common memory operations also include writing memory data. The following will use the example of a third device writing data to a first device to explain.
[0138] For ease of distinction, the unified interconnect bus unit 13 deployed on the third device may be referred to as a third unified interconnect bus unit 13 .
[0139] The third unified interconnection bus unit 13 can receive a write request sent by a third device, set the cache status of the third memory unit of the third device to an invalid state; encapsulate the data to be written into a write message that conforms to the message format; and send the write message to the first unified interconnection bus unit 13 through the hub.
[0140] The first unified interconnect bus unit 13 is configured to receive write messages fed back by the hub, store the data to be written carried in the write messages in the first memory unit, set the cache state of the first memory unit to a unique clean state based on the cache state carried in the write messages, and send monitoring messages to the second unified interconnect bus unit 13 via the hub. The second unified interconnect bus unit 13 is configured to receive monitoring messages sent by the hub, and set the cache state of the second memory unit to an invalid state based on the invalid state carried in the monitoring messages.
[0141] In some embodiments of the present application, there may be multiple categories of cache status. By setting the cache status, each device can intuitively understand the status of different memory units. For example, the device can determine whether the data in the memory unit is unique and unmodified, or shared and modified.
[0142] In practical applications, "unique" can be used to indicate that data exists only in the current memory unit and does not exist in other memory units. "Shared" can be used to indicate that data can exist in multiple memory units. "Clean" can be used to indicate that the data has not been modified. "Dirty" can be used to indicate that the data has been modified.
[0143] In some embodiments of the present application, the unified interconnect bus unit 13 can determine a matching cache state based on the state of the data on its corresponding memory unit; wherein the cache state may include an invalid state, a unique state, and a shared state; the unique state includes a unique clean state, a unique dirty state, a unique clean idle state, and a unique partially dirty state; the shared state includes a shared clean state and a shared dirty state.
[0144] The cache model defined by the unified interconnect bus unit supports different coherence protocols. Cache states include Unique Clean (UC), which indicates that a cache line is unique and "clean," exists only in the current cache, and has not been modified. Unique Dirty (UD), which indicates that a cache line is unique but "dirty," exists only in the current cache, but the cache line data has been modified and not updated to memory. Shared Clean (SC), which indicates that a cache line is non-unique and "clean," and may exist in other caches, and its data may have been modified, but is still "clean" in the current cache. Shared Dirty (SD), which indicates that a cache line is non-unique and "dirty," and may exist in other caches, but its data has been modified. Invalid, which indicates that a cache line is invalid and not in the cache. Unique Clean Empty (UCE), which indicates that a cache line exists only in the current cache, is unique, and all data bytes are invalid. Unique Dirty Partial (UDP): The state of the cache line is unique but partially "dirty". The cache line only exists in the current cache. The cache line is unique, but only part of the data in the cache line is valid and "dirty".
[0145] Of the seven states mentioned above, the first five are for processors using the ACE (AXI Coherency Extensions) coherence interface, while the full seven states are for processors using the CHI (Coherent Hub Interface) coherence interface. The queue management module internally determines which processor the coherence request originates from and selects the coherence state to use.
[0146] Before writing data, the first memory unit does not contain data, so the cache state is invalid. After the data is written to the first memory unit, the data recorded in the first memory unit is the original data in the third memory unit written to the first memory unit, so the cache state of the first memory unit can be set to the only clean state.
[0147] In the above introduction, the first device reads data from the second device and the third device writes data to the first device, both of which are based on the example of writing data to the same memory address of the first device. Therefore, after reading data from the second device and writing it to the memory location corresponding to the memory address, and then writing data to the same memory location of the first device by the third device, the memory address data of the second device is invalid. Therefore, the second unified interconnection bus unit 13 can set the cache status of the second memory unit to an invalid state after receiving the monitoring message.
[0148] In some embodiments of the present application, for ease of distinction, the protocol adaptation interface module included in the protocol layer on the third unified interconnection bus unit 13 may be referred to as the third protocol adaptation interface module, the request queue management module may be referred to as the third request queue management module, the response queue management module may be referred to as the third response queue management module, the message parsing and encapsulation module may be referred to as the third message parsing and encapsulation module, and the clock domain conversion module may be referred to as the third clock domain conversion module.
[0149] The third protocol adaptation interface module of the third unified interconnection bus unit 13 receives the write request sent by the third device, and transmits the write request to the third request queue management module through the internal bus.
[0150] a third request queue management module configured to receive a write request; set the cache state of the third memory unit of the third device to an invalid state; query the node forwarding table for a first device address that matches the first identifier carried in the write request; and transmit the write request and the first device address to the third message parsing and encapsulation module;
[0151] a third message parsing and encapsulating module, configured to encapsulate the received write request according to the message format to obtain a write message; and transmit the write message and the first device address to the third clock domain conversion module;
[0152] The third clock domain conversion module is used to transmit the write message and the first device address to the hub through the third adaptation layer and the third physical layer; the hub is used to forward the write message to the first device corresponding to the first device address.
[0153] In addition to the hub's role as a transit agent as described above, the hub can also assume a device identification function. Taking the example of a third device writing data to a first device, the third protocol adapter interface module of the third unified interconnect bus unit 13 is used to receive a write request sent by the third device and transmit the write request to the third request queue management module via the internal bus. The third request queue management module is used to receive the write request; set the cache status of the third memory unit of the third device to an invalid state; query the node forwarding table for the first device address that matches the first identifier based on the first identifier of the first device carried in the write request; and transmit the write request and the first device address to the third message parsing and encapsulation module. The third message parsing and encapsulation module is used to encapsulate the received write request according to the message format to obtain a write message; transmit the write message to the third clock domain conversion module; and the third clock domain conversion module is used to transmit the write message to the hub via the third adaptation layer and the third physical layer. The hub is used to query the node forwarding table for the first device address that matches the first identifier based on the first identifier of the first device carried in the write message; and forward the write message to the first device corresponding to the first device address.
[0154] In a specific implementation, the third message parsing and encapsulation module can encapsulate the third identifier of the third device, the first identifier of the first device to which the write request points, the message type and message sequence number to which the write request belongs, the first memory address, and the data to be written according to the identifier of the source device, the identifier of the destination device, the message length, the message type, the message sequence number, the memory address, the memory data, and the message format of the write enable signal to obtain a write message; wherein the first memory address carries a status identifier that the third memory unit is in an invalid state.
[0155] In some embodiments of the present application, each processor 11 may include a domestic processor with a different instruction set; the accelerator 12 includes a graphics processor and a field programmable gate array.
[0156] There are two main types of domestic processors: one uses the open-source and free AMD processor and Reduced Instruction Set Computer (RISC-V) architecture. The other uses the open but licensed ARM architecture. Accelerators can be domestically produced, such as domestic graphics processors and field-programmable gate arrays.
[0157] When the first device is an ARM processor 11, the first device may send a read request to the first unified interconnect bus unit 13 through a CHI interface. When the first device is a RISC-V processor 11, the first device may send a read request to the first unified interconnect bus unit 13 through an ACE interface.
[0158] To expand memory space, in addition to the memory units corresponding to each processor 11 and accelerator 12, a non-volatile readable storage medium independent of each processor 11 and accelerator 12 can also be provided; a unified interconnect bus unit 13 is deployed on the non-volatile readable storage medium. The non-volatile readable storage medium can be persistent memory (Intel Optane Persistent Memory, PMEM), low-power double data rate synchronous dynamic random access memory (Low Power Double Data Rate, LPDDR), solid state drive (SSD), etc.
[0159] The following example illustrates how a server system with multiple processors and heterogeneous accelerators can read and write memory data. For example, a processor of the first type reads data from a processor of the second type. The following example illustrates how the processor of the first type reads data from a processor of the second type.
[0160] (1) The first-class processor wants to read the memory data of the second-class processor. The first-class processor sends a read request to the UCCB through the CHI interface. The protocol adapter interface module of the UCCB processes the read request and sends it to the request queue management module through the internal bus. The cache state of the internal cache state machine is set to invalid. The message parsing and encapsulation module encapsulates the read request into a request message according to the message format corresponding to Table 1. The source ID of the request message is the ID of the first-class processor itself, the destination ID points to the second-class processor, the Type is a read-only transaction request, the Tagid is 1, the address is the memory address of the second-class processor, and the data is 0. It is sent to the hub through the adaptation layer and the physical layer. The hub directly forwards it to the second-class processor.
[0161] (2) The UCCB of the second-class processor receives the read message, changes the internal cache state to shared clean (SC), reads data from the memory and assembles it into a response message of the corresponding message format. The destination ID of the response message is the first-class processor ID, the type is response reply, the tagid is consistent with the read request and is still 1, addr is the memory address of the first-class processor, and data is the memory data. It is sent to the hub, and the hub forwards it to the first-class processor.
[0162] (3) After receiving the response message, the UCCB on the first-class processor updates the cache state to SC and sends a snoop message to the second-class processor. The source ID of the snoop message is the first-class processor's own ID, the destination ID points to the second-class processor, the type is snoop response, and the address and data are both 0. This completes a read transaction request.
[0163] On the basis of the first type processor reading the data of the second type processor, taking the domestic graphics processor writing data to the same memory address of the first type processor as an example, its operation process may include:
[0164] (4) The domestic graphics processor is ready to write data to the same memory address of the first-class processor. The UCCB of the domestic graphics processor initiates a write transaction request, which is encapsulated into a write message. The source ID of the write message is the ID of the domestic graphics processor, the destination ID is the ID of the first-class processor, the Type is a write transaction request, the Tagid is 1, the Tddr is the write request address (the same as the read request address), and the Data is the write memory data. The package is sent to the hub, and the hub sends the write message to the first-class processor according to the device address corresponding to the destination ID. The cache status is set to invalid through the UCCB of the GPU. In some embodiments of the present application, writing and reading belong to different operations. The write and read operations can be distinguished by Type, so the same digital representation can be used for the Tagid of the write and read operations. The Tagid can start from 1 and increase for requests with the same destination ID.
[0165] (5) The UCCB of the first type of processor receives the write message. After parsing, the message parsing and encapsulation module stores the data in the memory of the corresponding address and sets the cache status to unique clean (UC). At the same time, it returns a write response message (the format is the same as the read response) to the domestic graphics processor.
[0166] (6) The memory address data of the second-type processor is invalid. The UCCB of the first-type processor initiates a monitoring message to inform the second-type processor of the invalid state. The message format includes the Type as monitoring request, the Address as write request address, and the Data as invalid indication state.
[0167] (7) The UCCB of the second type of processor receives the monitoring request and sets the internal cache state to invalid. At this point, one consistent read request and one consistent write request have been completed.
[0168] The server system can include domestic processors with different instruction sets and different types of domestic accelerators. Both the domestic processors and accelerators include corresponding domestic memory units. In practical applications, the server system can also include independent non-volatile readable storage media, which can utilize domestic PMEM. Devices are interconnected via the UCCB interconnect bus. The interconnection topology can support a fully interconnected mesh topology, a crossbar topology, and other topologies. This allows for interconnection between processors, between accelerators, and between processors and accelerators, supporting system-wide sharing of memory resources. Within this topology, any number of UCCB-capable devices can be added.
[0169] Through the UCCB set up in this application, processors with multiple different types of instruction sets, different types of accelerators, and different memory units can be unified into one system to form a computer system. This can realize the sharing of local memory on the processor and remote memory on the accelerator to form a large memory pool, while alleviating the memory wall and IO wall.
[0170] As can be seen from the above technical solution, a multi-source heterogeneous distributed system includes at least two processors and at least one accelerator; each processor and accelerator has its own corresponding memory unit; each processor and accelerator is deployed with a unified interconnect bus unit; a protocol adapter interface module is deployed on the unified interconnect bus unit; the protocol adapter interface module is used to implement conversion between the consistency protocol interface of different types of devices and the unified interconnect bus protocol interface. The first device and the second device are any two devices among the processors and accelerators. The first unified interconnect bus unit deployed on the first device is used to receive a read request sent by the first device through the protocol adapter interface module. Because the first memory unit currently does not contain data, the cache state of the first memory unit of the first device can be set to invalid state. The read request is encapsulated into a request message that conforms to a set message format and sent to the second unified interconnect bus unit of the second device. The second unified interconnect bus unit is used to read data from the second memory unit of the second device after receiving the request message. Because data sharing is required, the cache state of the second memory unit can be set to shared state and the data can be encapsulated into a response message that conforms to the message format. The response message is sent to the first unified interconnect bus unit. The first unified interconnection bus unit is used to store the data carried by the received response message into the first memory unit; according to the cache status carried in the response message, the cache status of the first memory unit is set to a shared state. The beneficial effect of the present application is that by deploying a unified interconnection bus unit on each device and deploying a protocol adapter interface module on the unified interconnection bus unit, compatibility with multiple different types of devices can be achieved, so that processors with multiple different types of instruction sets, different types of accelerators, and different memory units can be unified in one system to form a computer system, which can realize memory sharing between different processors and between different processors and accelerators, forming a large memory pool, and alleviating the memory wall and IO wall problems. The unified interconnection bus unit can be used to interconnect different types of devices through different topological forms, support flexible interconnection topology, and can flexibly expand the scale without affecting the existing deployment. By recording the cache status of the memory unit, consistent memory communication between devices in the system is achieved. According to the deployment method of the present application, the performance of the multi-source heterogeneous distributed system is greatly improved.
[0171] FIG5 is a flowchart of a memory access method provided in some embodiments of the present application, including:
[0172] S501: Based on a protocol adaptation interface module receiving a read request sent by a first device, a cache state of a first memory unit of the first device is set to an invalid state.
[0173] The protocol adaptation interface module is used to realize the conversion between the consistency protocol interface of different types of devices and the unified interconnection bus protocol interface.
[0174] S502: Encapsulate the read request into a first request message that complies with a set message format according to the first identifier of the first device and the second identifier of the second device to which the read request points.
[0175] S503: Send the first request message to the second unified interconnection bus unit of the second device.
[0176] S504: Receive a response message fed back by the second unified interconnection bus unit, and store the data carried in the response message into the first memory unit.
[0177] S505: According to the cache status carried in the response message, the cache status of the first memory unit is set to a shared state.
[0178] In some embodiments, it further includes:
[0179] receiving a second request message sent by the hub; wherein the second request message is transmitted to the hub by the second unified interconnection bus unit of the second device;
[0180] Reading data from the first memory unit according to the first memory address carried in the second request message;
[0181] Setting the cache state of the first memory unit to a shared state, and encapsulating the data into a response message that conforms to a message format;
[0182] The response message is sent to the second device through the hub.
[0183] In some embodiments, it further includes:
[0184] receiving a write message fed back by the hub, and storing the data to be written carried in the write message in the first memory unit; wherein the write message is transmitted to the hub by the third unified interconnection bus unit of the third device;
[0185] According to the cache state carried in the write message, the cache state of the first memory unit is set to a unique clean state;
[0186] The monitoring message is sent to the second unified interconnection bus unit through the hub, so that the second unified interconnection bus unit receives the monitoring message sent by the hub; according to the invalid state carried in the monitoring message, the cache state of the second memory unit is set to invalid state.
[0187] For the description of features in some embodiments corresponding to FIG5 , reference can be made to the relevant description of some embodiments corresponding to FIG1 , and they will not be detailed here.
[0188] It can be seen from the above technical solution that the unified interconnection bus unit deployed on the first device can receive the read request sent by the first device based on the protocol adapter interface module, and set the cache status of the first memory unit of the first device to an invalid state. According to the first identifier of the first device and the second identifier of the second device pointed to by the read request, the read request is encapsulated into a first request message that complies with the set message format. The first request message is sent to the second unified interconnection bus unit of the second device, so that the second unified interconnection bus unit can read data from the second memory unit of the second device according to the second memory address carried in the request message; the cache status of the second memory unit is set to a shared state, and the data is encapsulated into a response message that complies with the message format; the response message is sent to the first unified interconnection bus unit. The first unified interconnection bus unit receives the response message fed back by the second unified interconnection bus unit, and stores the data carried in the response message to the first memory unit. According to the cache status carried in the response message, the cache status of the first memory unit is set to a shared state. The beneficial effect of the present application is that by deploying a unified interconnection bus unit on each device and deploying a protocol adapter interface module on the unified interconnection bus unit, compatibility with multiple different types of devices can be achieved, so that processors with multiple different types of instruction sets, different types of accelerators, and different memory units can be unified in one system to form a computer system, which can realize memory sharing between different processors and between different processors and accelerators, forming a large memory pool, and alleviating the problems of memory walls and IO walls. The unified interconnection bus unit can be used to interconnect different types of devices through different topological forms, support flexible interconnection topologies, and can be flexibly expanded without affecting existing deployments. By recording the cache status of the memory unit, consistent memory communication between devices in the system is achieved. According to the deployment method of the present application, the performance of multi-source heterogeneous distributed systems is greatly improved.
[0189] Some embodiments of the present application further provide a computer non-volatile readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned memory access method are implemented.
[0190] The above describes in detail a multi-source heterogeneous distributed system, memory access method, and computer non-volatile readable storage medium provided by some embodiments of the present application. The various embodiments in the specification are described in a progressive manner. Some embodiments focus on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the devices disclosed in some embodiments, since they correspond to the methods disclosed in some embodiments, the description is relatively simple. For the relevant parts, please refer to the method section.
[0191] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with some of the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0192] The above is a detailed introduction to a multi-source heterogeneous distributed system, a memory access method, and a computer non-volatile readable storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The descriptions of some of the above embodiments are only used to help understand the method of the present application and its core ideas. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A multi-source heterogeneous distributed system, characterized in that, It includes at least two processors and at least one accelerator; wherein, each of the processors and the accelerator has its corresponding memory unit; on each of the processors and the accelerator, a unified interconnection bus unit is deployed, and a protocol adaptation interface module is deployed on the unified interconnection bus unit; the protocol adaptation interface module is configured to implement the conversion between the consistency protocol interfaces of different types of devices and the unified interconnection bus protocol interface; The first unified interconnection bus unit is configured to receive a read request sent by a first device through the protocol adaptation interface module, set the cache state of the first memory unit of the first device to an invalid state; encapsulate the read request into a request message conforming to a set message format and send it to the second unified interconnection bus unit of a second device; wherein, the first device and the second device are any two devices among the processors and the accelerator; The second unified interconnection bus unit is configured to, after receiving the request message, read data from the second memory unit of the second device; set the cache state of the second memory unit to a shared state, and encapsulate the data into a response message conforming to the message format; send the response message to the first unified interconnection bus unit; The first unified interconnection bus unit is configured to store the data carried in the received response message into the first memory unit; set the cache state of the first memory unit to a shared state according to the cache state carried in the response message.
2. The multi-source heterogeneous distributed system according to claim 1, wherein The unified interconnection bus unit includes a protocol layer, an adaptation layer, and a physical layer; wherein, the protocol layer includes the protocol adaptation interface module, a request queue management module, a response queue management module, a message parsing and encapsulation module, and a clock domain conversion module; wherein, the protocol adaptation interface module is respectively connected to the request queue management module and the response queue management module, and is configured to transmit the received request to the request queue management module or the response queue management module according to the request type; the message parsing and encapsulation module is respectively connected to the request queue management module, the response queue management module, and the clock domain conversion module, and is configured to implement the encapsulation and decapsulation of messages.
3. The multi-source heterogeneous distributed system according to claim 2, wherein The first protocol adaptation interface module of the first unified interconnection bus unit is configured to receive the read request sent by the first device and transmit the read request to the first request queue management module through an internal bus; The first request queue management module is configured to receive the read request; Set the cache state of the first memory unit of the first device to an invalid state; Query the second device address matched by the second identifier from the node forwarding table according to the second identifier of the second device carried in the read request; transmit the read request and the second device address to the first message parsing and encapsulation module; The first message parsing and encapsulation module is configured to encapsulate the received read request according to the message format to obtain the request message; Transmit the request message and the second device address to the first clock domain conversion module; The first clock domain conversion module is configured to transmit the request message and the second device address to a hub through a first adaptation layer and a first physical layer; The hub is configured to forward the request message to the second device corresponding to the second device address.
4. The multi-source heterogeneous distributed system according to claim 2, characterized in that, The first protocol adaptation interface module of the first unified interconnection bus unit is configured to receive a read request sent by the first device and transmit the read request to a first request queue management module through an internal bus; The first request queue management module is configured to receive the read request; Set the cache status of the first memory unit of the first device to an invalid status; Transmit the read request to a first message parsing and encapsulation module; The first message parsing and encapsulation module is configured to encapsulate the received read request according to the message format to obtain the request message; Transmit the request message to the first clock domain conversion module; The first clock domain conversion module is configured to transmit the request message to a hub through a first adaptation layer and a first physical layer; The hub is configured to query, from a node forwarding table, a second device address matching the second identifier carried in the request message; and forward the request message to the second device corresponding to the second device address.
5. The multi-source heterogeneous distributed system according to claim 3 or 4, characterized in that The first message parsing and encapsulation module is configured to encapsulate the first identifier of the first device, the second identifier of the second device pointed to by the read request, the message type and message sequence number to which the read request belongs, and a second memory address according to a message format of an identifier of a source device, an identifier of a destination device, a message length, a message type, a message sequence number, a memory address, memory data, and a write enable signal to obtain the request message; wherein, a status identifier indicating that the first memory unit is in an invalid status is carried in the second memory address.
6. The multi-source heterogeneous distributed system according to claim 5, characterized in that The second clock domain conversion module of the second unified interconnection bus unit is configured to receive the request message transmitted by the hub through a second physical layer and a second adaptation layer, and forward the request message to a second message parsing and encapsulation module; The second message parsing and encapsulation module is configured to parse the request message to obtain the first identifier of the first device, the message sequence number, and the second memory address; Send the second memory address to a second response queue management module; The second response queue management module is configured to read data from the second memory unit according to the second memory address and transmit the data to the second message parsing and encapsulation module; Set the cache status of the second memory unit to a shared status; The second message parsing and encapsulation module is configured to receive the data; query the first device address matching the first identifier from the node forwarding table according to the first identifier; encapsulate the second identifier, the first identifier, the length of the data, the message type to which the response request belongs, the message sequence number, the first memory address, the data, and the write enable signal according to the message format to obtain the response message; transmit the response message and the first device address to the second clock domain conversion module; wherein, the first memory address carries a status identifier indicating that the second memory unit is in a shared state; The second clock domain conversion module is configured to transmit the response message and the first device address to the hub through the second adaptation layer and the second physical layer; The hub is configured to forward the response message to the first device corresponding to the first device address; 7. The multi-source heterogeneous distributed system according to claim 5, wherein The second clock domain conversion module of the second unified interconnection bus unit is configured to receive the request message transmitted by the hub through the second physical layer and the second adaptation layer, and forward the request message to the second message parsing and encapsulation module; The second message parsing and encapsulation module is configured to parse the request message to obtain the first identifier of the first device, the message sequence number, and the second memory address; Send the second memory address to the second response queue management module; The second response queue management module is configured to read data from the second memory unit according to the second memory address, and transmit the data to the second message parsing and encapsulation module; Set the cache status of the second memory unit to the shared state; The second message parsing and encapsulation module is configured to receive the data; Query the first memory address matching the first identifier from the node forwarding table according to the first identifier; encapsulate the second identifier, the first identifier, the length of the data, the message type to which the response request belongs, the message sequence number, the first memory address, the data, and the write enable signal according to the message format to obtain the response message; transmit the response message to the second clock domain conversion module; wherein, the first memory address carries a status identifier indicating that the second memory unit is in a shared state; The second clock domain conversion module is configured to transmit the response message to the hub through the second adaptation layer and the second physical layer; The hub is configured to query the first device address matching the first identifier from the node forwarding table according to the first identifier carried in the response message; Forward the response message to the first device corresponding to the first device address; 8. The multi-source heterogeneous distributed system according to claim 5, wherein The first unified interconnection bus unit is configured to encapsulate the first identifier, the second identifier, the message type to which the listening belongs, and the message sequence number according to the message format to obtain a listening message; Transmit the listening message and the second memory address to the first clock domain conversion module; The first clock domain conversion module is configured to transmit the monitoring message and the second device address to the hub through the first adaptation layer and the first physical layer; The hub is configured to forward the monitoring message to the second device corresponding to the second device address.
9. The multi-source heterogeneous distributed system according to claim 5, characterized in that The first unified interconnection bus unit is configured to encapsulate the first identifier, the second identifier, the message type to which the monitoring belongs, and the message sequence number according to the message format to obtain a monitoring message; Transmit the monitoring message to the first clock domain conversion module; The first clock domain conversion module is configured to transmit the monitoring message to the hub through the first adaptation layer and the first physical layer; The hub is configured to query, from the node forwarding table, the second device address that matches the second identifier carried in the monitoring message; forward the monitoring message to the second device corresponding to the second device address.
10. The multi-source heterogeneous distributed system according to claim 2, characterized in that, The third unified interconnection bus unit is configured to receive a write request sent by a third device, set the cache status of the third memory unit of the third device to an invalid status; encapsulate the data to be written into a write message that conforms to the message format; send the write message to the first unified interconnection bus unit through the hub; The first unified interconnection bus unit is configured to receive the write message fed back by the hub and store the data to be written carried in the write message into the first memory unit; According to the cache status carried in the write message, set the cache status of the first memory unit to the only clean status; Send a monitoring message to the second unified interconnection bus unit through the hub; The second unified interconnection bus unit is configured to receive the monitoring message sent by the hub; According to the invalid status carried in the monitoring message, set the cache status of the second memory unit to an invalid status.
11. The multi-source heterogeneous distributed system according to claim 10, wherein The third protocol adaptation interface module of the third unified interconnection bus unit is configured to receive a write request sent by the third device and transmit the write request to the third request queue management module through an internal bus; The third request queue management module is configured to receive the write request; Set the cache status of the third memory unit of the third device to an invalid status; According to the first identifier of the first device carried in the write request, query the first device address that matches the first identifier from the node forwarding table; transmit the write request and the first device address to the third message parsing and encapsulation module; The third message parsing and encapsulation module is configured to encapsulate the received write request according to the message format to obtain the write message; transmit the write message and the first device address to the third clock domain conversion module; The third clock domain conversion module is configured to transmit the write message and the first device address to the hub through the third adaptation layer and the third physical layer; The hub is configured to forward the write message to the first device corresponding to the first device address.
12. The multi-source heterogeneous distributed system according to claim 10, wherein The third protocol adaptation interface module of the third unified interconnection bus unit is configured to receive a write request sent by the third device and transmit the write request to the third request queue management module through an internal bus; The third request queue management module is configured to receive the write request; Set the cache status of the third memory unit of the third device to an invalid status; Query the first device address matching the first identifier from the node forwarding table according to the first identifier of the first device carried in the write request; transmit the write request and the first device address to the third message parsing and encapsulation module; The third message parsing and encapsulation module is configured to encapsulate the received write request according to the message format to obtain the write message; transmit the write message to the third clock domain conversion module; The third clock domain conversion module is configured to transmit the write message to the hub through the third adaptation layer and the third physical layer; The hub is configured to query the first device address matching the first identifier from the node forwarding table according to the first identifier of the first device carried in the write message; Forward the write message to the first device corresponding to the first device address.
13. The multi-source heterogeneous distributed system according to claim 11 or 12, characterized in that, The third message parsing and encapsulation module is configured to encapsulate the third identifier of the third device, the first identifier of the first device pointed to by the write request, the message type and message sequence number to which the write request belongs, the first memory address, and the data to be written according to the message format of the identifier of the source device, the identifier of the destination device, the message length, the message type, the message sequence number, the memory address, the memory data, and the write enable signal to obtain the write message; wherein, the first memory address carries a status identifier indicating that the third memory unit is in an invalid status.
14. The multi-source heterogeneous distributed system according to claim 2, wherein Both the adaptation layer and the physical layer are provided with bypass units; wherein, the bypass units are configured to implement direct connection between the adaptation layer and the physical layer; When the first device and the second device are located on the same printed circuit board, the bypass units of the first adaptation layer and the first physical layer of the first device and the second adaptation layer and the second physical layer of the second device are in an open state to implement direct connection between the first device and the second device.
15. The multi-source heterogeneous distributed system according to claim 1, wherein It further includes a non-volatile readable storage medium independently provided from each of the processors and the accelerators; the unified interconnection bus unit is deployed on the non-volatile readable storage medium.
16. The multi-source heterogeneous distributed system according to claim 1, wherein Each of the unified interconnection bus units is configured to determine a matching cache status according to the status of data on its corresponding memory unit; wherein, the cache status includes an invalid status, a unique status, and a shared status; the unique status includes a unique clean status, a unique dirty status, a unique clean idle status, and a unique partially dirty status; the shared status includes a shared clean status and a shared dirty status.
17. A memory access method, characterized in that, Includes: The protocol adaptation interface module receives a read request sent by a first device and sets the cache status of a first memory unit of the first device to an invalid status; wherein, the protocol adaptation interface module is configured to implement the conversion between the consistency protocol interfaces of different types of devices and a first unified interconnection bus protocol interface. According to the first identifier of the first device and the second identifier of a second device pointed to by the read request, encapsulate the read request into a first request message conforming to a set message format. Send the first request message to a second unified interconnection bus unit of the second device. Receive a response message fed back by the second unified interconnection bus unit and store the data carried in the response message into the first memory unit. According to the cache status carried in the response message, set the cache status of the first memory unit to a shared status.
18. The memory access method according to claim 17, wherein Further includes: Receive a second request message sent by a hub; wherein, the second request message is transmitted from a second unified interconnection bus unit of a second device to the hub. Read data from the first memory unit according to the first memory address carried in the second request message. Set the cache status of the first memory unit to a shared status and encapsulate the data into a response message conforming to the message format. Send the response message to the second device through the hub.
19. The memory access method according to claim 17, wherein Further includes: Receive a write message fed back by the hub and store the data to be written carried in the write message into the first memory unit; wherein, the write message is transmitted from a third unified interconnection bus unit of a third device to the hub. According to the cache status carried in the write message, set the cache status of the first memory unit to a unique clean status. Send a monitoring message to the second unified interconnection bus unit through the hub so that the second unified interconnection bus unit can receive the monitoring message sent by the hub; according to the invalid status carried in the monitoring message, set the cache status of a second memory unit to an invalid status.
20. A computer non-volatile readable storage medium, characterized in that, A computer non - volatile readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the memory access method according to any one of claims 17 to 19.
Citation Information
Patent Citations
Handling cache write-back and cache eviction for cache coherence
CN104520824A
Memory interconnection method, system and related device
CN114567683A
High-speed communication method and device of heterogeneous equipment and heterogeneous communication system
CN116886751A
Multi-source heterogeneous distributed system, memory access method and storage medium
CN117806553A
High performance interconnect coherence protocol
US20140115268A1
Cited By
Data exchange architecture, system and method capable of realizing cache consistency
CN120508413A
Data exchange architecture, system and method enabling cache coherency
CN120508413B
UCIe design module test method and device, electronic equipment and storage medium
CN121118781A
Cross-protocol interaction method, electronic device, readable medium and program product
CN121367739A
Method and device for converting AXI protocol into CHI protocol
CN121579408A