Data transmission method and device, heterogeneous system, and coherent interconnect processing component
By dynamically sensing the topology through consistent interconnect processing devices, masking the differences in interconnect protocols between devices, establishing memory-consistent interconnect communication links and allocating cache space, the problem of interconnect complexity in heterogeneous systems is solved, and the distributed computing performance of heterogeneous systems is improved.
Patent Information
- Application Number
- PCT/CN2025/098342
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-17
- Filing Date
- 2025-05-30
- Publication Date
- 2026-01-22
AI Technical Summary
In heterogeneous systems, different types of devices use different interconnection protocols, which makes interconnection complex and difficult to achieve consistency, thus affecting the performance of distributed computing.
A consistent interconnect processing device is provided, including a consistent interconnect interface, a consistent interconnect controller, and a dynamic address configuration engine. It dynamically senses the topology, shields the differences in interconnect protocols between devices, establishes a memory consistent interconnect communication link, allocates cache space, and realizes cache consistent transaction processing.
It achieves memory-consistent interconnection between devices in heterogeneous systems, reduces the number of accesses to local and remote memory, improves access performance, simplifies topology deployment configuration, and enhances distributed computing performance.
Smart Images

Figure CN2025098342_22012026_PF_FP_ABST
Abstract
Description
Data transmission methods, devices, heterogeneous systems, and consistent interconnect processing devices
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202410954443.8, filed on July 17, 2024, entitled "Data Transmission Method, Apparatus, Heterogeneous System and Consistent Interconnection Processing Device", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of data storage technology, and in particular to data transmission methods, devices, heterogeneous systems, and consistent interconnect processing devices. Background Technology
[0004] Device memory interconnection in a distributed system refers to the memory sharing and communication mechanism between different physical or logical devices in a distributed computing environment. This interconnection allows nodes in a distributed system to efficiently exchange data and coordinate their work. However, in heterogeneous systems, there are many different types of devices that may use different interconnection protocols. This results in fragmented and inconsistent interconnections between different types of devices in current heterogeneous systems, requiring the maintenance of multiple different interconnection buses and leading to complex management. Summary of the Invention
[0005] The purpose of this application is to provide a data transmission method, device, heterogeneous system, and consistent interconnection processing device for realizing consistent interconnection between different types of devices in a heterogeneous system and improving the performance of the heterogeneous system in performing distributed computing.
[0006] To address the aforementioned technical problems, this application provides a data transmission method applied to a conformance interconnect processing device, comprising:
[0007] Based on the number of other devices in the heterogeneous system that need to establish a memory-coherent interconnect with the device, enable the corresponding coherent interconnect interface on the device, and obtain the memory interconnect parameters of the other devices after initializing the physical communication link between the device and other devices based on the coherent interconnect interface.
[0008] Based on the memory interconnect parameters of the device and the memory interconnect parameters of other devices, establish a memory-consistent interconnect communication link between the device and other devices, and allocate corresponding cache space from the device.
[0009] Based on the memory-consistent interconnect communication link and cache space, cache-consistent transaction processing is performed between the device and other devices to realize memory-consistent interconnect transmission requests between the device and other devices.
[0010] On the one hand, based on the memory interconnect parameters of the device and the memory interconnect parameters of other devices, the corresponding cache space is allocated from the device, including:
[0011] The amount of cache space is determined based on the device type in the memory interconnect parameters, and cache space corresponding to the memory of the device and the memory of other devices is allocated from the device.
[0012] On the other hand, it allocates cache space from the device itself, corresponding to the memory of the device and the memory of other devices, including:
[0013] The size of each cache space is determined based on the device memory size in the memory interconnect parameters, so that the larger the device memory size, the larger the allocated cache space;
[0014] The cache size is determined by the cache space available on the device.
[0015] On the other hand, it allocates cache space from the device itself, corresponding to the memory of the device and the memory of other devices, including:
[0016] The length of the access path from the current device to the target device's memory is determined based on the memory interconnect parameters. The size of each cache space is then determined based on the length of the access path, so that the longer the access path, the larger the allocated cache space.
[0017] The cache size is determined by the cache space available on the device.
[0018] On the other hand, it allocates cache space from the device itself, corresponding to the memory of the device and the memory of other devices, including:
[0019] The length of the access path from the current device to the target device's memory is determined based on the memory interconnect parameters. The size of each cache space is determined based on the length of the access path and the device memory size in the memory interconnect parameters, so that the larger the device memory size, the larger the allocated cache space, and the longer the access path, the larger the allocated cache space.
[0020] The cache size is determined by the cache space available on the device.
[0021] On the other hand, it allocates cache space from the device itself, corresponding to the memory of the device and the memory of other devices, including:
[0022] The first query time for the device to access the target device's memory and the second query time for the device to access the target cache space corresponding to the target device's memory are determined based on the memory interconnect parameters. The size of each cache space is determined based on the balance between the first query time and the second query time when executing the memory consistency interconnect transmission request.
[0023] The cache size is determined by the cache space available on the device.
[0024] On the other hand, based on the memory-consistent interconnect communication link and cache space, cache-consistent transaction processing is performed between the device and other devices to realize memory-consistent interconnect transmission requests between the device and other devices, including:
[0025] Query the target cache space corresponding to the target memory address of the memory consistency interconnect transfer request, and perform the data transfer operation corresponding to the data transfer request based on the target cache space.
[0026] On the other hand, memory interconnect parameters include service type;
[0027] The business type is either multiple processing units performing read / write operations on the device memory or only one processing unit performing read / write operations on the device memory.
[0028] On the other hand, if the memory-consistent interconnect transmission request is a request issued by the device to other devices, cache-consistent transaction processing is performed between the device and other devices based on the memory-consistent interconnect communication link and cache space to realize the memory-consistent interconnect transmission request between the device and other devices, including:
[0029] If the service type in the memory interconnect parameters of the target device is that there are multiple processing units performing read and write operations on the device memory, then a non-cached data transfer request is initiated to the consistency interconnect interface of the target device to perform read and write operations on the target device's memory without changing the cache state of the target device.
[0030] If the service type in the memory interconnect parameters of the target device is that there is only one processing unit performing read and write operations on the device memory, then a cache consistency read request is initiated to the consistency interconnect interface of the target device to perform read and write operations on the cache space and memory of the target device.
[0031] On the other hand, based on the memory-consistent interconnect communication link and cache space, cache-consistent transaction processing is performed between the device and other devices to realize memory-consistent interconnect transmission requests between the device and other devices, including:
[0032] The memory consistency interconnect transmission request is converted into a read / write operation on memory. When a local read / write operation misses the cache, the corresponding consistency interconnect interface is selected based on address arbitration to transmit the memory consistency interconnect transmission request message. After processing the response message, cache consistency processing of the corresponding cache space is performed.
[0033] On the other hand, cache space corresponds to device type;
[0034] The memory consistency interconnect transmission request is converted into a memory read / write operation. When a local read / write operation misses the cache, the corresponding consistency interconnect interface is selected based on address arbitration to transmit the memory consistency interconnect transmission request message. After processing the response message, cache consistency processing for the corresponding cache space is performed, including:
[0035] Based on the target memory address of the memory consistency interconnect transmission request, determine the cache space corresponding to the target memory address locally, and trigger the corresponding cache controller to query the target data from the corresponding cache space;
[0036] If the target data is not found in the corresponding cache space, the corresponding consistency interconnect interface corresponding to the target memory address is determined according to the address arbitration to transmit the memory consistency interconnect transmission request message, and after the response message is processed, the corresponding cache controller is triggered to execute the cache consistency processing of the corresponding cache space.
[0037] On the other hand, based on the memory interconnect parameters of the device and other devices, a memory-consistent interconnect communication link is established between the device and other devices, and corresponding cache space is allocated from the device, including:
[0038] Based on the memory interconnect parameters of the device and the memory interconnect parameters of other devices, a four-state cache coherence protocol is used to establish a memory coherence interconnect communication link between the device and other devices, and corresponding cache space is allocated from the device.
[0039] The four-state cache consistency protocol uses four states—modified, exclusive, shared, and invalid—to represent the cache state.
[0040] On the other hand, based on the memory-consistent interconnect communication link and cache space, cache-consistent transaction processing is performed between the device and other devices to realize memory-consistent interconnect transmission requests between the device and other devices, including:
[0041] The memory coherent interconnect transmission request is converted into a coherent interconnect protocol message.
[0042] The types of consistent interconnection protocol messages include request transaction messages and response transaction messages. The content of a request transaction message includes a request message validity flag, a request transaction type, a physical address of the request transaction, and a unique identifier for the request transaction. The content of a response transaction message includes a response message validity flag, a response transaction type, an identifier for the response transaction, and response return data. The identifier for the response transaction corresponds to the identifier for the request transaction.
[0043] On the other hand, the length of the physical address of the requested transaction is determined based on the scale of the devices in the heterogeneous system.
[0044] On the other hand, enabling the corresponding consistency interconnect interface from the device itself includes:
[0045] A consistent interconnect interface is implemented using the Ethernet physical media access control interface of the device.
[0046] To address the aforementioned technical problems, this application also provides a coherent interconnect processing device, comprising: a coherent interconnect interface, a coherent interconnect controller, and a dynamic address configuration engine;
[0047] The consistency interconnect controller is used to enable the corresponding consistency interconnect interface from the host device based on the number of other devices in the heterogeneous system that need to establish a memory consistency interconnect with the host device, and obtain the memory interconnect parameters of other devices after initializing the physical communication link between the host device and other devices based on the consistency interconnect interface; establish a memory consistency interconnect communication link between the host device and other devices based on the memory interconnect parameters of the host device and the memory interconnect parameters of other devices, and call the dynamic address configuration engine to allocate the corresponding cache space from the host device; and perform cache consistency transaction processing between the host device and other devices based on the memory consistency interconnect communication link and cache space to realize the memory consistency interconnect transmission request between the host device and other devices.
[0048] To address the aforementioned technical problems, this application also provides a heterogeneous system, comprising multiple heterogeneous devices, each of which is equipped with a consistency interconnect processing device;
[0049] The consistency interconnect processing device is used to enable the corresponding consistency interconnect interface from the host device based on the number of other devices in the heterogeneous system that need to establish a memory consistency interconnect with the host device, and obtain the memory interconnect parameters of the other devices after initializing the physical communication link between the host device and other devices based on the consistency interconnect interface; establish the memory consistency interconnect communication link between the host device and other devices based on the memory interconnect parameters of the host device and the memory interconnect parameters of other devices, and allocate the corresponding cache space from the host device; and perform cache consistency transaction processing between the host device and other devices based on the memory consistency interconnect communication link and cache space to realize the memory consistency interconnect transmission request between the host device and other devices.
[0050] To address the aforementioned technical problems, this application also provides a data transmission device, comprising:
[0051] Memory, used to store computer programs;
[0052] A processor is used to execute computer programs, which, when executed by the processor, implement the steps of any of the data transmission methods described above.
[0053] To address the aforementioned technical problems, this application also provides a non-volatile storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of any of the above data transmission methods.
[0054] To address the aforementioned technical problems, this application also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of any of the data transmission methods described above.
[0055] The data transmission method provided in this application has the following advantages: The application, installed on a device, uses a consistent interconnect processing device to establish a memory-consistent interconnect with other devices in a heterogeneous system, based on the number of such devices. It enables the corresponding consistent interconnect interface on the device, initializes the physical communication link between the device and other devices based on the consistent interconnect interface, obtains the memory interconnect parameters of the other devices, establishes a memory-consistent interconnect communication link between the device and other devices according to the memory interconnect parameters of the device and other devices, and allocates corresponding cache space from the device. Based on the memory-consistent interconnect communication link and cache space, it performs cache-consistent transaction processing between the device and other devices to realize memory-consistent interconnect transmission requests between the device and other devices. This achieves a scheme that dynamically senses the topology and automatically builds memory-consistent interconnect communication links between devices in a heterogeneous system. During deployment, no manual configuration based on the device topology is required. This eliminates the problem of different types of devices in a heterogeneous system being unable to establish memory interconnects due to different interconnect protocols. Simultaneously, it achieves cache consistency between devices to reduce the number of accesses to local and remote memory, further improving access performance.
[0056] This application also provides a consistent interconnect processing device, a heterogeneous system, a data transmission device, and a non-volatile storage medium, which have the above-mentioned beneficial effects, and will not be elaborated here. Attached Figure Description
[0057] To more clearly illustrate the technical solutions of the embodiments or related technologies of this application, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0058] Figure 1 is a schematic diagram of the structure of a consistency interconnect processing device provided in an embodiment of this application;
[0059] Figure 2 is a schematic diagram of a consistent interconnect interface provided in an embodiment of this application;
[0060] Figure 3 is a schematic diagram of a consistency interconnect controller provided in an embodiment of this application;
[0061] Figure 4 is an architecture diagram of a heterogeneous system provided in an embodiment of this application;
[0062] Figure 5 is a flowchart of a data transmission method provided in an embodiment of this application;
[0063] Figure 6 is a schematic diagram of a system global address space allocation provided in an embodiment of this application;
[0064] Figure 7 is a flowchart of a specific implementation of S503 in Figure 5 provided in an embodiment of this application;
[0065] Figure 8 is a schematic diagram of the structure of a data transmission device provided in an embodiment of this application. Specific Implementation
[0066] The core of this application is to provide a data transmission method, device, heterogeneous system, and consistent interconnection processing device to achieve consistent interconnection between different types of devices in a heterogeneous system and improve the performance of the heterogeneous system in performing distributed computing.
[0067] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0068] To facilitate understanding of the technical solutions provided in the embodiments of this application, some key terms used in the embodiments of this application will be explained here first.
[0069] Inter-chip interconnects in heterogeneous systems include interconnects between processors, interconnects between processors and accelerators, and interconnects between accelerators. Inter-processor interconnect buses include Ultra Path Interconnect (UPI), Quick Path Interconnect (QPI), Hyper-Transport (HT), CoreLink, and others. Interconnects between processors and accelerators most commonly use the Peripheral Component Interconnect Express (PCIe) bus. In recent years, to address issues such as the "memory wall" and "IO wall," the industry has proposed cache coherent interconnect buses for processors and accelerators, including the Cache Coherent Interconnect for Accelerators (CCIX) standard and the Compute Express Link (CXL) protocol. The main coherent interconnect buses between accelerators include the NVLink bus for interconnecting graphics processing units (GPUs) and the High-speed Custom Communication System (HCCS) bus for interconnecting artificial intelligence processors.
[0070] Cache coherence refers to the spatial coherence of the cache of multi-core processors, the cache of devices, and memory. Device memory with cache coherence can be accessed as unified memory.
[0071] As can be seen, current heterogeneous systems integrate multiple protocol buses. Interconnect protocols between processors are typically defined by the processor manufacturers; interconnects between processors and accelerators currently mostly use the PCIe bus. For interconnect buses with cache coherency features, such as CXL, each device manufacturer has generally defined its own independent protocol bus, which is inconsistent with the interconnect bus between processors; interconnect protocols between accelerators are also defined by the accelerator manufacturers. This results in many different bus protocols. When building a distributed computing topology for a heterogeneous system, it is necessary to first complete the topology deployment, then perform initialization configuration, configure the bus protocol, and then establish data communication. If the topology is changed, re-initialization and configuration are often required, making the process complex and labor-intensive.
[0072] To address the issue of consistent interconnection between different types of devices in heterogeneous systems, this application provides a hardware-based consistent interconnection processing device. This device dynamically senses the topology of the installed devices and configures consistent interconnection between devices based on the topology sensing results. During data transmission, it masks the differences in interconnection protocols between different types of devices, eliminating the problem that different types of devices in heterogeneous systems cannot establish memory interconnections due to different interconnection protocols. Furthermore, by achieving cache consistency between devices, it reduces the number of accesses to local and remote memory, further improving access performance.
[0073] For ease of understanding, the consistency interconnect processing device provided in this application will be described first.
[0074] Figure 1 is a schematic diagram of the structure of a consistency interconnect processing device provided in an embodiment of this application.
[0075] As shown in Figure 1, the conformance interconnect processing device provided in this application embodiment may include: a conformance interconnect interface, a conformance interconnect controller, and a dynamic address configuration engine.
[0076] The consistency interconnect controller is used to enable the corresponding consistency interconnect interface from the host device based on the number of other devices in the heterogeneous system that need to establish a memory consistency interconnect with the host device, and obtain the memory interconnect parameters of other devices after initializing the physical communication link between the host device and other devices based on the consistency interconnect interface; establish a memory consistency interconnect communication link between the host device and other devices based on the memory interconnect parameters of the host device and the memory interconnect parameters of other devices, and call the dynamic address configuration engine to allocate the corresponding cache space from the host device; and perform cache consistency transaction processing between the host device and other devices based on the memory consistency interconnect communication link and cache space to realize the memory consistency interconnect transmission request between the host device and other devices.
[0077] It should be noted that, in the embodiments of this application, the types of "devices" mainly include processors and accelerators. In the hardware environment, these processors and accelerators can be devices installed on the same server or devices installed on different servers.
[0078] The conformance interconnect processing device provided in this application embodiment can be installed as hardware on a device. In one application scenario provided by this application embodiment, the conformance interconnect processing device provided by this application embodiment can be installed on the devices used when building a heterogeneous system. In this case, the conformance interconnect processing device provided by this application embodiment can mask the differences in interconnect protocols defined at the factory, achieving conformance interconnection between heterogeneous devices. In another application scenario provided by this application embodiment, for programmable controllers such as Field Programmable Gate Arrays (FPGAs), the conformance interconnect processing device provided by this application embodiment can be added to enable conformance interconnection with other devices.
[0079] To facilitate understanding, the coherent interconnect processing device provided in this application embodiment will first be described within the context of the device and the heterogeneous system. As shown in Figure 1, the coherent interconnect processing device provided in this application embodiment supports coherent interconnects between accelerators, coherent interconnects between processors and accelerators, and coherent interconnects between processors. It also includes interconnections between the device and other local devices installed on the same server, as well as interconnections between the device and remote devices installed on different servers.
[0080] In this embodiment, a coherent interconnect interface can be implemented based on the hardware resources of the device, and a unified coherent interconnect bus can be used to establish physical communication links between devices. The coherent interconnect interface can dynamically adapt to connect with different types of devices, including coherent interconnects with processors and coherent interconnects with accelerators. It is also used to dynamically adapt to different topologies, obtain device memory interconnect parameters, configure a global address pool, and complete the coherent interconnect protocol processing for global memory in heterogeneous systems. The coherent interconnect interface may include a physical layer, a link layer, and a transport layer.
[0081] On the device, the corresponding coherent interconnect interface can be enabled according to the number of other devices required to establish memory coherent interconnects. In some optional embodiments of this application, such as an accelerator, four coherent interconnect interfaces can be pre-prepared on the accelerator. The four coherent interconnect interfaces can be enabled in one / two / three / four paths according to different topologies to adapt to different topologies. As shown in Figure 1, the accelerator can use two coherent interconnect interfaces to interconnect with two processors, and the other two coherent interconnect interfaces to interconnect with two other accelerators.
[0082] In the consistent interconnect processing device provided in this application embodiment, the consistent interconnect controller manages the cache consistency function, realizing cache consistency management between the caches of various devices in the heterogeneous system and between the system memory. The consistent interconnect controller can be used to receive memory resource handshake information initiated by interconnect devices, extract memory interconnect parameters to complete global memory allocation, and set routing windows; receive access requests initiated by the accelerated application units of the devices, perform cache hit lookups, and initiate different memory transaction requests based on the cache hit status and the lookup address range (specifying which device's memory it is) to complete the consistent lookup function processing.
[0083] In the consistent interconnect processing device provided in this application embodiment, a dynamic address configuration engine is used to configure the local cache space and the mapping mechanism between cache addresses and memory addresses based on the memory interconnect parameters obtained from the consistent interconnect interface. This is used to accelerate application address routing and can be dynamically configured according to the topology to improve cache hit rate. The dynamic address configuration engine can also be used to allocate a global address space for accessing global addresses; the above configuration can be dynamically completed according to the topology.
[0084] In Figure 1, the memory subsystem refers to the device's memory, which is a component of consistent and unified memory in a heterogeneous system.
[0085] In Figure 1, the accelerated application unit is a dynamically configured unit used to implement applications that require accelerated computing, such as deep learning.
[0086] Figure 2 is a schematic diagram of a consistent interconnect interface provided in an embodiment of this application.
[0087] In this embodiment of the application, the consistency interconnect interface is used to implement the cache consistency protocol provided in this embodiment of the application, including the physical layer, the link layer and the transport layer.
[0088] As shown in Figure 2, the physical layer and link layer are used for link management, including link initialization and reliability management. The implementation of these two layers is flexible and can be selected based on the physical link resources of the device board. For example, an Interlaken interface for lightweight packet interconnect can be selected, as can an Ethernet Physical Medium Attachment Media Access Control (PHYMAC) interface, or a Serial Transceiver (SerDes) interface. A high-speed interface type can also be chosen. In some optional embodiments of this application, an Ethernet PHYMAC interface is used, supporting a rate of 56Gbps / lane (channel). Using 8 lanes, the single interface rate can reach 448Gbps.
[0089] In the consistency interconnect interface provided in this application embodiment, the transport layer uses a custom consistency transaction type to process consistency transaction messages. In this application embodiment, the consistency transaction type can be based on the four-state cache consistency (Modify, Exclusive, Shared, Invalid, MESI) protocol specification.
[0090] In a single-core cache, each cache line has two flags: a dirty (or modified) flag and a valid flag. These flags effectively describe the data relationship between the cache and memory (whether the data is valid and whether it has been modified). In a multi-core processor, multiple cores share some data, and the MESI protocol includes a description of the shared state.
[0091] In the MESI protocol, each cache line has four states, represented by two bits: Modify (M state), Exclusive (E state), Shared (S state), and Invalid (I state). The Modify state indicates that the data line is valid, has been modified, and is inconsistent with the data in memory; the data exists only in this cache. The Exclusive state indicates that the data line is valid, matches the data in memory, and exists only in this cache. The Shared state indicates that the data line is valid, matches the data in memory, and exists in multiple caches. The Invalid state indicates that the data line is invalid.
[0092] The transport layer user interface is a custom Cache Coherence Interface (CCI) bus. This application provides a set of reference message formats as follows.
[0093] Table 1 is a request transaction message format table provided in the embodiments of this application. In the consistency interconnection protocol defined in the embodiments of this application, the message format for requesting transactions can be as shown in Table 1. The fields represent the message content type, and the bit width is the data size of the message content.
[0094] Table 1
[0095] Table 2 is a response transaction message format table provided in the embodiments of this application. In the consistency interconnection protocol defined in the embodiments of this application, the message format for response transactions can be as shown in Table 2. The fields represent the message content type, and the bit width is the data size of the message content.
[0096] Table 2
[0097] In this embodiment, the physical address can use 64 bits as the smallest granularity, so the lowest 6 bits are 0, which is the physical address of the requested system memory. The system memory includes the memory of the device where it is located and the memory of remote devices (remote processor memory, remote accelerator memory). Different bit widths of addresses are used according to the system topology connection.
[0098] Depending on the deployment of the accelerated application, sometimes the read / write address is a virtual address. In this case, the consistency interconnect processing device can also be used to convert the virtual address into a physical address and then encapsulate and parse it according to the message format described above.
[0099] The opcodes of request and response transaction messages contain a consistent transaction flow for interacting with other devices' consistent transactions. Table 3 shows the request transaction message types provided in this application embodiment. As shown in Table 3, the consistent interconnection protocol provided in this application embodiment can implement request transaction types including read requests and write requests. Read requests can include read data (RdD), read memory (RdCur), and read permission (RdNoData), while write requests can include invalid, read-then-write (RDITW), write data (WrM), flush specified cache line (CLFlush), and cache flushed (CacheFlushed). Response transaction types include cached data acknowledgment (CaDRsp), memory data acknowledgment (MemDRsp), invalid (InvldRsp), cached data writable (WCaDRsp), memory data writable (WMemDRsp), and write complete (Wrdone). The operation process for each transaction type is described in Table 3. The four states M, E, S, and I are explained in the MESI protocol description above.
[0100] Table 3
[0101] In addition to implementing consensus protocol transaction processing functions, the transport layer can also realize the function of sensing device memory interconnect parameters. Memory interconnect parameters can include device type, device memory size, service type, etc.
[0102] A consensus protocol transport layer control module (Cfg) can be deployed to achieve the function of sensing memory interconnect parameters. The consensus protocol transport layer control module can deploy two sets of registers. One set of registers stores the memory interconnect parameters of the local device, which can be called the local attribute register. The other set of registers stores the memory interconnect parameters of other devices, which can be called the remote attribute register. In the registers, the device type register and the device memory size register have fixed values that vary depending on the board type of the local device, while the service type register can be initialized and configured by the acceleration service after the device powers on, and varies depending on the service type. The value of the remote attribute register can be obtained by the consensus protocol transport layer control module reading the local registers of other devices by sending a message after the consensus interconnect processing device has completed the link initialization configuration.
[0103] Figure 3 is a schematic diagram of the structure of a consistency interconnect controller provided in an embodiment of this application.
[0104] In the consistency interconnect processing device provided in this application embodiment, the consistency interconnect controller is used for managing cache consistency functions, and its implementation module block diagram is shown in Figure 3. The interface processing module deploys the Cache Coherence Interface (CCI) protocol provided in this application embodiment, used to parse or encapsulate the custom consistency bus message format of this application embodiment, including receiving request or response information from the request-response processing module, encapsulating it into a custom message format and sending it to the consistency interconnect interface; receiving remote request or response messages sent by the consistency interconnect interface, parsing them, and sending them to the request-response processing module for processing.
[0105] As described above in the embodiments of this application, a dynamic address configuration engine is used to configure the local cache space and the mapping mechanism between cache addresses and memory addresses based on the memory interconnect parameters obtained from the consistency interconnect interface. In the consistency interconnect controller, the request-response processing module can be used to configure the capacity of the local cache space based on the memory interconnect parameters obtained from the dynamic address configuration engine, and to complete cache consistency transaction-related processing: converting memory read / write requests into memory read / write operations; if a local read / write operation misses the cache, selecting a suitable cache consistency interface path based on address arbitration to send a protocol transaction message; and processing the response message to write data back to the cache or change the state of the cacheline. The request-response processing module can communicate with the memory subsystem and the accelerated application unit through the Advanced eXtensible Interface (AXI) bus.
[0106] In some optional embodiments of this application, a cache space can be deployed locally to establish cache-consistent interconnection with the local device and other devices.
[0107] For ease of management, in some optional embodiments of this application, cache spaces corresponding one-to-one with the memory of the local device and the memory of other devices can be allocated from the local device based on the memory interconnect parameters of the local device and the memory interconnect parameters of other devices. In this case, as shown in Figure 3, the coherent interconnect controller provided in this application embodiment may further include a read / write processing (Cafu_proc) module and a cache controller. Multiple cache spaces can be established locally, corresponding one-to-one with the local device and other devices. A cache controller can be set up for each cache space to manage the cache lookup tasks and cache line status update tasks of that cache space. Alternatively, a host cache controller and an accelerator cache controller can be set up to manage the respective cache spaces of the processor and accelerator.
[0108] The read / write processing module is used to complete the read / write operations of the device's accelerated applications. It determines the corresponding cache space based on the read / write address and triggers the corresponding cache controller to perform cache lookup and cache line status update.
[0109] Therefore, this application proposes a unified cache-coherent interconnect scheme that is topology-dynamically aware and can be used for hardware implementation of interconnections between accelerators, between accelerators and processors, and between processors. The coherent interconnect processing device provided in this application unifies the interface protocols between accelerators and external systems (with remote accelerators and remote processors), eliminating the need for complex conversions between different coherence protocol interfaces. Accelerated applications on accelerators can use this unified coherent interconnect processing device to access not only local device memory but also the memory of remote processors and remote accelerators in a consistent manner, expanding accelerator memory capacity and improving memory performance. This effectively overcomes device performance bottlenecks and significantly improves computing performance in large-scale distributed computing. The coherent interconnect processing device provided in this application can customize various coherence protocol transaction types to implement cache coherence functionality in hardware, eliminating the need for software program development. This not only reduces the programming difficulty for application developers but also improves the memory access performance of applications. Based on this, the consistent interconnect processing device provided in this application can dynamically sense the system topology, dynamically enable the consistent interconnect interface according to different topologies, dynamically allocate global memory space and cache capacity, and dynamically select corresponding cache consistent transactions. It can adapt to various accelerated applications with different computing power and supports changes in various topologies, such as wireless mesh networks (Mesh), Clos networks, and topologies based on the Clos network model such as Fat Tree topology, Crossbar topology, and fully interconnected topology, simplifying deployment difficulty. Combined with the reconfigurable features supported by the FPGA hardware itself, it supports dynamic switching between different accelerated applications and flexibly switches topologies according to the computing power of the accelerated application. Furthermore, through the cache consistency protocol, it supports the unification of memory across devices in heterogeneous systems, enabling accelerated applications to consistently access all memory within the system, greatly improving the acceleration performance of memory-sensitive applications such as deep learning.
[0110] Figure 4 is an architecture diagram of a heterogeneous system provided in an embodiment of this application.
[0111] As shown in Figure 4, the heterogeneous system provided in this embodiment includes multiple heterogeneous devices, each equipped with a consistency interconnect processing device. The consistency interconnect processing device is used to, based on the number of other devices in the heterogeneous system with which it establishes a memory consistency interconnect with its own device, enable the corresponding consistency interconnect interface from its own device, initialize the physical communication link between its own device and other devices based on the consistency interconnect interface, and then obtain the memory interconnect parameters of the other devices. Based on the memory interconnect parameters of its own device and the memory interconnect parameters of other devices, it establishes a memory consistency interconnect communication link between its own device and other devices, and allocates corresponding cache space from its own device. Based on the memory consistency interconnect communication link and the cache space, it performs cache consistency transaction processing between its own device and other devices to realize the memory consistency interconnect transmission request between its own device and other devices.
[0112] Referring to the coherent interconnect processing device provided in the above embodiments of this application, in heterogeneous systems, cache coherent interconnects can be established between accelerators, between processors and accelerators, and between processors based on the coherent interconnect processing device provided in the embodiments of this application. Devices are connected to each other through coherent interconnect interfaces and coherent interconnect buses. The working process is detailed in the following embodiments of this application.
[0113] Based on the above architecture, the data transmission method provided in the embodiments of this application will be described below with reference to the accompanying drawings.
[0114] Figure 5 is a flowchart of a data transmission method provided in an embodiment of this application.
[0115] As shown in Figure 5, the data transmission method provided in this application embodiment for a coherent interconnect processing device includes: S501: based on the number of other devices in the heterogeneous system that need to establish a memory coherent interconnect with the device, enable the corresponding coherent interconnect interface from the device, and obtain the memory interconnect parameters of the other devices after initializing the physical communication link between the device and other devices based on the coherent interconnect interface.
[0116] S502: Based on the memory interconnect parameters of the device and the memory interconnect parameters of other devices, establish a memory consistency interconnect communication link between the device and other devices, and allocate corresponding cache space from the device.
[0117] S503: Based on the memory-consistent interconnect communication link and cache space, perform cache-consistent transaction processing between the device and other devices to realize memory-consistent interconnect transmission requests between the device and other devices.
[0118] It should be noted that the specific implementation of the embodiments of this application can refer to some or all of the descriptions of the consistent interconnect processing device and heterogeneous system provided in the above embodiments.
[0119] In this embodiment, for S501, after the device is powered on, the coherent interconnect processing device provided in this embodiment automatically senses the device topology to determine other devices that need to establish a memory coherent interconnect with the device, and enables the corresponding number of coherent interconnect interfaces. After each device establishes a physical communication link through the coherent interconnect interface and the coherent interconnect bus, memory resource handshake information is transmitted through the coherent interconnect processing device.
[0120] In this embodiment, enabling the corresponding coherent interconnect interface from the device may include implementing the coherent interconnect interface using the device's Ethernet physical media access control interface. In practical applications, other types of interfaces may also be used, depending on the hardware resources provided by the device.
[0121] For the S502, the consistency interconnect processing device establishes a memory consistency interconnect communication link between itself and other devices based on the memory interconnect parameters of its own device and those of other devices. This includes extracting memory interconnect parameters to complete global memory allocation for the heterogeneous system and setting up routes. Based on this, the consistency interconnect processing device allocates cache space from its own device corresponding to the memory of its own device and the memory of other devices, implementing a cache consistency protocol to further improve the data access efficiency of distributed computing applications.
[0122] In some optional embodiments of this application, S502, establishing a memory consistency interconnect communication link between the device and other devices based on the memory interconnect parameters of the device and the memory interconnect parameters of other devices, and allocating corresponding cache space from the device, may include: establishing a memory consistency interconnect communication link between the device and other devices using a four-state cache consistency (Modify, Exclusive, Shared, Invalid, MESI) protocol based on the memory interconnect parameters of the device and the memory interconnect parameters of other devices, and allocating corresponding cache space from the device; wherein, the four-state cache consistency protocol represents the cache state through four states: Modify, Exclusive, Shared, and Invalid.
[0123] For S503, the memory-consistent interconnect communication link and cache space established by the application consistency interconnect processing device realize the distributed computing acceleration task in the heterogeneous system, perform memory consistency and cache consistency transaction processing, shield the differences of the interconnect protocols that come with different devices from the factory, and execute the memory-consistent interconnect transmission request between the device and other devices using the consistency interconnect protocol customized in the embodiment of this application.
[0124] If a four-state cache consistency protocol is adopted, then in S503, cache consistency transaction processing between the local device and other devices is performed based on the memory consistency interconnect communication link and cache space to realize the memory consistency interconnect transmission request between the local device and other devices. This can include: converting the memory consistency interconnect transmission request into a consistency interconnect protocol message; wherein, the consistency interconnect protocol message type includes a request transaction message and a response transaction message; the content of the request transaction message includes a request message validity flag, a request transaction type, a physical address of the request transaction, and a unique identifier of the request transaction; the content of the response transaction message includes a response message validity flag, a response transaction type, an identifier of the response transaction, and response return data; the identifier of the response transaction corresponds to the identifier of the request transaction.
[0125] In this embodiment, the length of the physical address of the requested transaction is determined based on the scale of devices in the heterogeneous system. That is, the more devices there are, the larger the device memory is, and the more different physical addresses need to be maintained.
[0126] The data transmission method provided in this application embodiment applies a consistent interconnect processing device installed on a device to enable the corresponding consistent interconnect interface on the device based on the number of other devices in the heterogeneous system that need to establish memory consistent interconnect with the device. After initializing the physical communication link between the device and other devices based on the consistent interconnect interface, the method obtains the memory interconnect parameters of other devices. Based on the memory interconnect parameters of the device and other devices, it establishes a memory consistent interconnect communication link between the device and other devices and allocates corresponding cache space in the device. Based on the memory consistent interconnect communication link and cache space, it performs cache consistent transaction processing between the device and other devices to realize the memory consistent interconnect transmission request between the device and other devices. This method realizes a scheme that dynamically senses the topology and automatically builds memory consistent interconnect communication links between devices in a heterogeneous system. During deployment, no manual configuration based on the device topology is required. It can eliminate the problem that different types of devices in a heterogeneous system cannot establish memory interconnects due to different interconnect protocols. At the same time, it realizes cache consistency between devices to reduce the number of accesses to local memory and remote memory, further improving access performance.
[0127] Figure 6 is a schematic diagram of a system global address space allocation provided in an embodiment of this application.
[0128] Based on the above embodiments, this application further describes the process of establishing the global address space and cache space of a heterogeneous system.
[0129] In this embodiment, after the device powers on, the consensus interconnect processing device completes the initial connection of the link. The consensus protocol transport layer control module of the consensus interconnect interface then sends messages to other devices to obtain memory interconnect parameters, which are then stored in the local attribute register and the remote attribute register. These memory interconnect parameters may include the device type and the device memory size.
[0130] The dynamic address configuration engine in the coherent interconnect processor reads the memory interconnect parameters obtained from each coherent interconnect interface of the device and provides them to the coherent interconnect controller of the coherent interconnect processor for cache space configuration.
[0131] The coherent interconnect processor configures its global address space based on the memory size of its own device and the memory sizes of other devices. Subsequent accelerated applications use this memory space to select the coherent interconnect interface for data transmission. As shown in Figure 6, taking a heterogeneous system consisting of one processor and four accelerators as an example, assuming the processor has 128GB of memory and each of the three accelerators has 16GB of memory, then the global address space 0 to 3FFFFFFFF corresponds to the memory space of accelerator 1, the global address space 3FFFFFFFF to 7FFFFFFFF corresponds to the memory space of accelerator 2, the global address space 7FFFFFFFF to FFFFFFFFF corresponds to the memory space of accelerator 3, and the global address space FFFFFFFFF to 2FFFFFFFFF corresponds to the memory space of the processor.
[0132] To further improve data access efficiency, a cache space corresponding to the device's memory is established. For ease of management, in this embodiment, S502 allocates a corresponding cache space from the device based on the memory interconnect parameters of the device and the memory interconnect parameters of other devices. This may include: determining the number of cache spaces based on the device type in the memory interconnect parameters, and allocating cache spaces corresponding to the memory of the device and the memory of other devices from the device.
[0133] In some optional embodiments of this application, allocating cache space corresponding to the memory of the local device and the memory of other devices from the local device may include: determining the size of each cache space based on the device memory size in the memory interconnect parameters, so that the larger the device memory size, the larger the allocated cache space; and allocating cache space from the local device according to the size of the cache space. Taking the example shown in Figure 6, the processor cache space can be configured to be 128KB based on the processor memory size, and the accelerator cache space can be configured to be 48KB based on the accelerator memory size (the sum of the memory sizes of three accelerators can be 48GB). When there are multiple processors or multiple accelerators, a one-to-one corresponding cache space can also be configured for each device.
[0134] In some optional embodiments of this application, allocating cache space corresponding to the memory of the local device and the memory of other devices from the local device may also include: determining the length of the access path from the local device to the memory of the target device based on memory interconnect parameters; determining the size of each cache space based on the length of the access path, so that the longer the access path, the larger the allocated cache space; and caching space from the local device based on the size of the cache space. The longer the access path from the local device to the memory of the target device, the longer the time spent accessing the target device memory. In this case, a larger cache space can be allocated locally for such target device memory to improve the hit rate of querying data in the target device memory in the local cache space, thereby reducing the need to access the target device memory with long paths.
[0135] In some optional embodiments of this application, allocating cache space corresponding to the memory of the local device and the memory of other devices from the local device may further include: determining the length of the access path from the local device to the memory of the target device based on memory interconnect parameters; determining the size of each cache space based on the access path length and the device memory size in the memory interconnect parameters, so that the larger the device memory size, the larger the allocated cache space, and the longer the access path length, the larger the allocated cache space; and caching space from the local device based on the size of the cache space. That is, factors such as device memory size and access path length can be comprehensively considered to allocate appropriate cache space for the memory of the local device and the memory of other devices to balance storage resources and access time overhead.
[0136] In practical applications, the larger the cache space, the longer the cache lookup time. Setting the cache space too large may actually prolong the access time. Therefore, in some optional embodiments of this application, allocating cache space from the local device corresponding to the memory of the local device and the memory of other devices may further include: determining the first query time for the local device to access the memory of the target device and the second query time for the local device to access the target cache space corresponding to the memory of the target device based on memory interconnect parameters; determining the size of each cache space based on the balance between the first query time and the second query time when executing a memory consistency interconnect transmission request; and allocating cache space from the local device according to the size of the cache space.
[0137] Based on the above optional implementation methods provided in the embodiments of this application, it can be seen that device type, device memory size, access path length, and access time overhead can all be used as the basis for determining the cache space size corresponding to the device memory. In practical applications, it is not limited to considering only these factors. For example, the cache space size can be determined by combining the access frequency.
[0138] Since corresponding cache spaces are allocated based on device type, in this embodiment, S503 performs cache consistency transaction processing between the device and other devices based on the memory consistency interconnect communication link and cache space to realize the memory consistency interconnect transmission request between the device and other devices. This can include: querying the target cache space corresponding to the target memory address corresponding to the memory consistency interconnect transmission request, and performing the data transmission operation corresponding to the data transmission request based on the target cache space. Referring to the consistency interconnect processing device described in the above embodiments of this application, the read / write processing module in the consistency interconnect controller can determine the corresponding cache space based on the read / write address and trigger the corresponding cache controller to perform cache lookup and cache line status update. When the read / write address is a processor memory address, the processor cache controller is triggered to perform a lookup of the local processor cache space; when the read / write address is an accelerator memory address, the accelerator cache controller is triggered to perform a lookup of the local accelerator cache space. If corresponding cache spaces are allocated for each device, the corresponding cache controller can be triggered to perform a cache hit after the corresponding cache space is found by distinguishing the device memory address.
[0139] Based on the above embodiments, in the embodiments of this application, the memory interconnect parameters may further include service type; the service type is either multiple processing units performing read / write operations on the device memory or only one processing unit performing read / write operations on the device memory.
[0140] Depending on the type of business the device is operating on, different cache consistency implementation mechanisms can be selected. For example, when multiple accelerators are performing parallel acceleration of the model, there will not be multiple accelerators operating on the same memory address at the same time. In this case, caching can be set only on the local device, and the remote device can be exempt from caching. A simplified inter-device cache consistency protocol can be used. However, during operations such as synchronization, multiple devices will operate on the same memory address at the same time, and the caching of the remote accelerator needs to be enabled. In this case, the full inter-device cache consistency protocol transaction processing should be used.
[0141] In this embodiment of the application, if the memory consistency interconnect transmission request is a request issued by the device to other devices, S503 performs cache consistency transaction processing between the device and other devices based on the memory consistency interconnect communication link and cache space to realize the memory consistency interconnect transmission request between the device and other devices. This can include: if the service type in the memory interconnect parameters of the target device is that there are multiple processing units performing read and write operations on the device memory, then an uncached data transmission request is initiated to the consistency interconnect interface of the target device to perform read and write operations on the target device's memory without changing the cache state of the target device; if the service type in the memory interconnect parameters of the target device is that there is only one processing unit performing read and write operations on the device memory, then a cache consistency read request is initiated to the consistency interconnect interface of the target device to perform read and write operations on the target device's cache space and the target device's memory.
[0142] Based on the above embodiments, in this application embodiment, S503 performs cache consistency transaction processing between the local device and other devices based on the memory consistency interconnect communication link and cache space to realize the memory consistency interconnect transmission request between the local device and other devices. This can include: converting the memory consistency interconnect transmission request into a read / write operation on memory, so that when the local read / write fails to hit the cache, the corresponding consistency interconnect interface is selected according to address arbitration to transmit the message of the memory consistency interconnect transmission request, and after completing the processing of the response message, the cache consistency processing of the corresponding cache space is performed.
[0143] In this embodiment, if the cache space is configured to correspond to the device type, the memory consistency interconnect transmission request is converted into a read / write operation on memory. When a local read / write operation misses the cache, the corresponding consistency interconnect interface is selected based on address arbitration to transmit the memory consistency interconnect transmission request message. After processing the response message, cache consistency processing for the corresponding cache space is executed. This can include: determining the cache space corresponding to the target memory address locally based on the target memory address of the memory consistency interconnect transmission request, and triggering the corresponding cache controller to query the target data from the corresponding cache space; if the target data is not found in the corresponding cache space, the consistency interconnect interface corresponding to the target memory address is determined based on address arbitration to transmit the memory consistency interconnect transmission request message, and after processing the response message, the corresponding cache controller is triggered to execute cache consistency processing for the corresponding cache space.
[0144] Based on the above embodiments, taking the heterogeneous system shown in Figure 4 as an example, this application embodiment further illustrates the device memory access process.
[0145] The heterogeneous system shown in Figure 4 includes a processor and three accelerators (accelerator 1, accelerator 2, and accelerator 3), which can be used to accelerate deep learning applications such as convolutional neural networks (CNNs). The devices form a fully interconnected topology as shown in Figure 4.
[0146] Taking the access of accelerator 1 to the memory of accelerator 2 as an example, the data transmission method provided in this application embodiment is illustrated, which mainly includes an initialization process and a consistent access process.
[0147] In the initialization process, as shown in the topology in Figure 4, each device (processor or accelerator) physically supports four coherent interconnect interfaces. After power-on, the physical link layer of the coherent interconnect interfaces performs initial link operations, enabling three coherent interconnect interfaces according to the actual connection topology. The four devices are then interconnected in pairs, forming a fully interconnected topology. It should be noted that the topology can be dynamically changed, and the coherent interconnect interfaces can be dynamically configured according to the actual topology.
[0148] After power-on and completing the initial link connection, the consensus protocol transport layer control module of the consensus interconnect interface sends messages to the other three devices to obtain the device type, memory size, and service type fixed on the other devices' boards. For example, the processor's board type can be recorded as 0, its memory capacity as 128GB, and its service type as 0 (with cache); the device type of accelerator 2 and accelerator 3 can be recorded as 1, its memory capacity as 16GB, and its service type as 1 (without cache). Accelerator 1 stores the obtained memory interconnect parameters in its local registers for other modules to read.
[0149] The dynamic address configuration engine reads the memory interconnect parameters obtained from the coherent interconnect interfaces of the device. It can configure the processor cache space to 128KB according to the processor memory size, configure the accelerator cache space to 48KB according to the accelerator memory size (which can be the sum of the memory sizes of the three accelerators), and configure the cache size according to the memory size to improve the cache hit rate.
[0150] Meanwhile, a global address space is configured based on the device memory size of the four devices. Subsequent accelerated applications use this memory space to select a consistent interconnect interface for transmission, as shown in Figure 6.
[0151] Assuming that the business type of accelerator 2 is asynchronous operation, that is, there is only one processing unit reading and writing to the device memory, accelerator 1 can choose the corresponding transaction type to reduce cache consistency probing operations and speed up data reading.
[0152] Figure 7 is a flowchart of a specific implementation of S503 in Figure 5 provided by an embodiment of this application.
[0153] After completing the above initialization configuration, the process of the application coherence interconnect processing device executing the memory coherence interconnect transmission request for the accelerated application unit can be shown in Figure 7, including: S701: The accelerated application unit of accelerator 1 initiates a request operation for the target read address (e.g., A = 0x403458924).
[0154] S702: The consistency interconnect controller of accelerator 1 identifies that the target read address is located in the address space range of accelerator 2 based on the window position of the target read address in the global address, and searches for the cache space corresponding to accelerator 2 locally.
[0155] S703: The consistency interconnect controller of accelerator 1 determines whether the target data has been found in the cache space corresponding to accelerator 2. If yes, proceed to S704; otherwise, proceed to S705.
[0156] S704: Return the target data from the cache space corresponding to accelerator 2 and end.
[0157] S705: The consistency interconnect controller of accelerator 1 checks the service type of accelerator 2 according to the remote attribute register of the consistency protocol transport layer control module. If it determines that there is only one processing unit in the acceleration application that performs read and write operations on the device memory, it initiates a direct read data request transaction to the consistency interconnect interface of accelerator 2.
[0158] S706: The consensus interconnect interface of accelerator 1 encapsulates the received read memory request transactions into a consensus protocol message of a custom format and sends it to accelerator 2 through the physical communication link.
[0159] S707: The consensus interconnect interface of Accelerator 2 parses the received consensus protocol messages and sends the operation request signal to the local consensus interconnect controller.
[0160] S708: The Consistency Interconnect Controller of Accelerator 2 directly initiates a read request to local memory based on the received read memory request transaction, and sends the read target data to the local Consistency Interconnect interface.
[0161] S709: The consistency interconnect interface of Accelerator 2 encapsulates the target data into a custom consistency protocol message that returns data transactions from the memory of the remote device and sends it to Accelerator 1 via the physical communication link.
[0162] S710: The consensus interconnect interface of Accelerator 1 parses the received consensus protocol messages and extracts the target data to send to the local consensus interconnect controller.
[0163] S711: The consistency interconnect controller of Accelerator 1 stores the target data in the local accelerator cache space and sends the target data to the acceleration application unit to complete the read request processing.
[0164] The process of an accelerator accessing other accelerators or processors is similar.
[0165] It should be noted that in the embodiments of this application, some steps or features may be omitted or not executed. The hardware or software functional modules defined are for ease of explanation and are not the only implementation of the data transmission method provided in the embodiments of this application.
[0166] The foregoing has detailed some embodiments of the data transmission method. Based on this, this application also discloses data transmission devices, non-volatile storage media, and computer program products corresponding to the above methods.
[0167] Figure 8 is a schematic diagram of the structure of a data transmission device provided in an embodiment of this application.
[0168] As shown in FIG8, the data transmission device provided in this application embodiment includes: a memory 810 for storing a computer program 811; and a processor 820 for executing the computer program 811, wherein the computer program 811, when executed by the processor 820, implements the steps of the data transmission method provided in any of the above embodiments.
[0169] The processor 820 may include one or more processing cores, such as a 3-core processor or an 8-core processor. The processor 820 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 820 may also include a main processor and a coprocessor. The main processor, also known as the Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 820 may integrate a Graphics Processing Unit (GPU) responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 820 may also include an Artificial Intelligence (AI) processor for handling computational operations related to machine learning.
[0170] The memory 810 may include one or more non-volatile storage media, which may be non-transitory. The memory 810 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory 810 is used to store at least the following computer program 811, wherein, after being loaded and executed by the processor 820, the computer program 811 is able to implement the relevant steps in the data transmission method disclosed in the foregoing embodiments. In addition, the resources stored in the memory 810 may also include an operating system 812 and data 813, and the storage method may be temporary storage or permanent storage. The operating system 812 may be Windows or other types of operating systems. The data 813 may include, but is not limited to, the data involved in the above methods.
[0171] In some embodiments, the data transmission device may further include a display screen 830, a power supply 840, a communication interface 850, an input / output interface 860, a sensor 870, and a communication bus 880.
[0172] Those skilled in the art will understand that the structure shown in Figure 8 does not constitute a limitation on the data transmission device and may include more or fewer components than shown.
[0173] The data transmission device provided in this application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the steps of the data transmission method provided in the above embodiments, and the effect is the same as above.
[0174] This application provides a non-volatile storage medium storing a computer program thereon, which, when executed by a processor, can implement the steps of the data transmission method provided in any of the above embodiments.
[0175] The non-volatile storage medium may include: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks or optical disks, and other media that can store program code.
[0176] For a description of the non-volatile storage medium provided in the embodiments of this application, please refer to the above method embodiments. The effects it achieves are the same as the data transmission method provided in the embodiments of this application, and will not be repeated here.
[0177] This application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the data transmission method provided in any of the above embodiments.
[0178] For a description of the computer program product provided in the embodiments of this application, please refer to the above method embodiments. The effects it achieves are the same as the data transmission method provided in the embodiments of this application, and will not be repeated here.
[0179] The data transmission method, apparatus, heterogeneous system, and conformance interconnect processing device provided in this application have been described in detail above. The embodiments in the specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus, non-volatile storage medium, and computer program products disclosed in the embodiments, since they correspond to the conformance interconnect processing device and data transmission method disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the descriptions of the conformance interconnect processing device and data transmission method. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.
[0180] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
Claims
1. A data transmission method, characterized by, The application is applied to a consistent interconnection processing device, comprising: Enabling a corresponding consistent interconnection interface from the device according to the number of other devices in a heterogeneous system with which the device needs to establish memory consistent interconnection, and obtaining memory interconnection parameters of the other devices after initializing a physical communication link between the device and the other devices based on the consistent interconnection interface; Establishing a memory consistent interconnection communication link between the device and the other devices according to the memory interconnection parameters of the device and the memory interconnection parameters of the other devices, and allocating corresponding cache spaces from the device; Performing cache consistent transaction processing between the device and the other devices based on the memory consistent interconnection communication link and the cache spaces to realize memory consistent interconnection transmission requests between the device and the other devices.
2. The data transmission method of claim 1, wherein, Allocating corresponding cache spaces from the device according to the memory interconnection parameters of the device and the memory interconnection parameters of the other devices, comprising: Determining the number of cache spaces according to the device type in the memory interconnection parameters, and allocating cache spaces corresponding to the memory of the device and the memory of the other devices from the device.
3. The data transmission method of claim 2, wherein, Allocating cache spaces corresponding to the memory of the device and the memory of the other devices from the device, comprising: Determining the size of each cache space according to the device memory size in the memory interconnection parameters, so that the larger the device memory size, the larger the allocated cache space; Caching the cache spaces from the device according to the size of the cache spaces.
4. The data transmission method of claim 2, wherein, Allocating cache spaces corresponding to the memory of the device and the memory of the other devices from the device, comprising: Determining the length of an access path from the device to the memory of a target device according to the memory interconnection parameters, and determining the size of each cache space according to the length of the access path, so that the longer the access path, the larger the allocated cache space; Caching the cache spaces from the device according to the size of the cache spaces.
5. The data transmission method of claim 2, wherein, Allocating cache spaces corresponding to the memory of the device and the memory of the other devices from the device, comprising: Determining the length of an access path from the device to the memory of a target device according to the memory interconnection parameters, and determining the size of each cache space according to the length of the access path and the device memory size in the memory interconnection parameters, so that the larger the device memory size, the larger the allocated cache space, and the longer the access path, the larger the allocated cache space; Caching the cache spaces from the device according to the size of the cache spaces.
6. The data transmission method of claim 2, wherein, Allocating cache spaces corresponding to the memory of the device and the memory of the other devices from the device, comprising: Determining a first query time for the device to access the memory of a target device and a second query time for the device to access a target cache space corresponding to the memory of the target device according to the memory interconnection parameters, and determining the size of each cache space according to the balanced relationship between the first query time and the second query time when executing the memory consistent interconnection transmission request; Caching the cache spaces from the device according to the size of the cache spaces.
7. The data transmission method of claim 2, wherein, Performing cache consistent transaction processing between the device and the other devices based on the memory consistent interconnection communication link and the cache spaces to realize memory consistent interconnection transmission requests between the device and the other devices, comprising: query a target cache space corresponding to a target memory address corresponding to the memory-coherent interconnect transmission request, and perform a data transmission operation corresponding to the data transmission request based on the target cache space.
8. The data transmission method of claim 1, wherein, The memory interconnection parameter comprises a service type. The service type is a read-write operation of a plurality of processing units on the device memory or a read-write operation of only one processing unit on the device memory.
9. The data transmission method of claim 8, wherein, If the memory-coherent interconnect transmission request is a request sent by the device to other devices, cache-coherent transaction processing of the device and other devices is performed based on the memory-coherent interconnection communication link and the cache space to realize the memory-coherent interconnect transmission request between the device and other devices, comprising: If the service type in the memory interconnection parameter of the target device is a read-write operation of a plurality of processing units on the device memory, a non-cacheable data transmission request is initiated to the coherent interface of the target device to perform the read-write operation on the memory of the target device without changing the cache state of the target device. If the service type in the memory interconnection parameter of the target device is a read-write operation of only one processing unit on the device memory, a cache-coherent read request is initiated to the coherent interface of the target device to perform the read-write operation on the cache space of the target device and the memory of the target device.
10. The data transmission method of claim 1, wherein, Cache-coherent transaction processing of the device and other devices is performed based on the memory-coherent interconnection communication link and the cache space to realize the memory-coherent interconnect transmission request between the device and other devices, comprising: The memory-coherent interconnect transmission request is converted into a read-write operation on the memory, and when the local read-write cache is not hit, the packet of the memory-coherent interconnect transmission request is transmitted to the corresponding coherent interface according to address arbitration, and after the processing of the response packet is completed, cache-coherent processing of the corresponding cache space is performed.
11. The data transmission method of claim 10, wherein, The cache space corresponds to the device type. The memory-coherent interconnect transmission request is converted into a read-write operation on the memory, and when the local read-write cache is not hit, the packet of the memory-coherent interconnect transmission request is transmitted to the corresponding coherent interface according to address arbitration, and after the processing of the response packet is completed, cache-coherent processing of the corresponding cache space is performed, comprising: The cache space corresponding to the target memory address in the memory-coherent interconnect transmission request is determined in the local device, and the corresponding cache controller is triggered to query the target data from the corresponding cache space; If the target data is not queried in the corresponding cache space, the packet of the memory-coherent interconnect transmission request corresponding to the target memory address is transmitted to the coherent interface according to address arbitration, and after the processing of the response packet is completed, the corresponding cache controller is triggered to perform cache-coherent processing of the corresponding cache space.
12. The data transmission method of claim 1, wherein, According to the memory interconnection parameter of the device and the memory interconnection parameter of other devices, a memory-coherent interconnect communication link of the device and other devices is established, and a corresponding cache space is allocated from the device, comprising: According to the memory interconnection parameter of the device and the memory interconnection parameter of other devices, a four-state cache consistency protocol is adopted to establish the memory consistency interconnection communication link of the device and other devices, and corresponding cache space is allocated from the device. The four-state cache consistency protocol represents the cache state through modification, exclusive, shared, and invalid four states.
13. The data transmission method of claim 12, wherein, Based on the memory consistency interconnection communication link and the cache space, cache consistency transaction processing of the device and other devices is performed to realize the memory consistency interconnection transmission request between the device and other devices, including: The memory consistency interconnection transmission request is converted into a consistency interconnection protocol message; The type of the consistency interconnection protocol message includes a request transaction message and a response transaction message; the content of the request transaction message includes a request message valid flag, a request transaction type, a request transaction physical address, and a request transaction unique identifier; the content of the response transaction message includes a response message valid flag, a response transaction type, a response transaction identifier, and response return data; the response transaction identifier corresponds to the request transaction identifier.
14. The data transmission method of claim 13, wherein, The length of the request transaction physical address is determined according to the device scale in the heterogeneous system.
15. The data transmission method of claim 1, wherein, The corresponding consistency interconnection interface of the device is enabled, including: The consistency interconnection interface is realized by the Ethernet physical media access control interface of the device.
16. A coherent interconnect processing device, comprising: Including: a consistency interconnection interface, a consistency interconnection controller, and a dynamic address configuration engine; The consistency interconnection controller is configured to enable the corresponding consistency interconnection interface of the device according to the number of other devices in the heterogeneous system that needs to establish memory consistency interconnection with the device, and to obtain the memory interconnection parameter of other devices after initializing the physical communication link between the device and other devices based on the consistency interconnection interface; according to the memory interconnection parameter of the device and the memory interconnection parameter of other devices, the memory consistency interconnection communication link of the device and other devices is established, and the dynamic address configuration engine is called to allocate corresponding cache space from the device; Based on the memory consistency interconnection communication link and the cache space, cache consistency transaction processing of the device and other devices is performed to realize the memory consistency interconnection transmission request between the device and other devices.
17. A heterogeneous system, comprising: A plurality of heterogeneous devices are included, and each of the heterogeneous devices is installed with a consistency interconnection processing device; The consistency interconnection processing device is configured to enable the corresponding consistency interconnection interface of the device according to the number of other devices in the heterogeneous system that needs to establish memory consistency interconnection with the device, and to obtain the memory interconnection parameter of other devices after initializing the physical communication link between the device and other devices based on the consistency interconnection interface; according to the memory interconnection parameter of the device and the memory interconnection parameter of other devices, the memory consistency interconnection communication link of the device and other devices is established, and corresponding cache space is allocated from the device; Performing cache coherency transactions for the local device and other devices based on the memory coherent interconnect communication links and cache spaces to achieve memory coherent interconnect transfer requests between the local device and other devices.
18. A data transmission device, characterized by Comprising: a memory configured to store a computer program; a processor configured to execute the computer program, the computer program, when executed by the processor, implements the steps of the data transfer method according to any one of claims 1 to 15.
19. A non-transitory storage medium having stored thereon a computer program, characterized in that the computer program, when executed by the processor, implements the steps of the data transfer method according to any one of claims 1 to 15.
20. A computer program product comprising computer programs / instructions, characterized in that, the computer program / instructions, when executed by the processor, implement the steps of the data transfer method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Memory and storage controller with integrated memory coherency interconnect
CN113609033A
Multi-source heterogeneous distributed system, memory access method and storage medium
CN117806553A
Data processing system, method and device, medium and computer program product
CN118113631A
Data transmission method and device, heterogeneous system and consistent interconnection processing device
CN118503195A
Port server for heterogeneous hardware
US20220383173A1
Cited By
Soft and hard heterogeneous data processing system and method for atmospheric detection laser radar
CN122111888A
A dual exchange network algorithm fusion heterogeneous computing method based on a CLOS architecture
CN122247955A