A data processing system, method and connection device

By connecting device management and mapping computing cluster memory addresses, the system adapts ultra-low latency communication protocols to RDMA protocols, solving the problems of speed and flexibility in cross-computing cluster access and improving the efficiency of the data processing system.

CN118939597BActive Publication Date: 2025-12-12HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411002660.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-12-01
Filing Date
2023-03-07
Publication Date
2025-12-12
Estimated Expiration
2043-03-07

AI Technical Summary

Technical Problem

How to adapt ultra-low latency communication protocols to RDMA protocols to improve the speed and flexibility of data access across computing clusters.

Method used

By connecting devices to manage the memory address information of multiple computing clusters, cross-computing cluster memory access is achieved. Ultra-low latency communication protocols and RDMA protocols are used for adaptation. Connecting devices uniformly manage the memory address information of each computing device and establish memory address mapping relationships at the application level.

Benefits of technology

It improves access speed and memory access flexibility across computing clusters, reduces access latency, and enhances the flexibility of applications accessing the memory space of multiple computing clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118939597B_ABST
    Figure CN118939597B_ABST
Patent Text Reader

Abstract

A data processing system, method and connection device are used to adapt a super low latency communication protocol (such as a CXL protocol or a UB protocol) to a RDMA protocol. In the present application, the data processing system comprises a first computing cluster, a second computing cluster and a connection device, wherein the first computing cluster comprises a first computing device, and the second computing cluster comprises a second computing device. Further, the connection device is connected to the first computing device, and is used to manage memory address information of the first computing cluster provided by the first computing device, and memory address information of the second computing cluster provided by the second computing device; the first computing device is used to: receive an access request, the access request being used to access a memory space of the second computing device; find an address of the memory space from the memory address information of the second computing cluster managed by the connection device; and access the memory space in the second computing device according to the address.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the original application with the application number 202310248707.3 and the original filing date of March 7, 2023, and the entire contents of the original application are incorporated herein by reference.

[0002] This application claims priority to the Chinese patent application with the application number 202211531954.6 and the filing date of December 1, 2022, and the title of “a memory management method”, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0003] The present application relates to the technical field of computing, and in particular to a data processing system, method and connection device. BACKGROUND

[0004] Remote direct memory access (RDMA) is conceptually relative to direct memory access (DMA). In DMA technology, an external device (i.e., a peripheral component interconnect express (PCIe) device) can bypass a central processing unit (CPU) to directly access the memory of a host; in RDMA technology, the external device can also bypass the CPU to access the memory of another remote host.

[0005] With the advent of the big data era, data centers not only have to face massive data and users, but also need to provide ultra-fast data communication services to users. The emergence of emerging ultra-low latency communication (such as compute express link (CXL)) protocols brings new possibilities to data communication. The ultra-low latency communication protocol can realize high-speed and efficient interconnection between CPUs and graphics processing units (GPUs), CPUs and field programmable gate arrays (FPGAs), or CPUs and other accelerators.

[0006] However, how to implement the adaptation of the ultra-low latency communication protocol and the RDMA protocol is a technical problem to be solved at present. SUMMARY

[0007] The application provides a data processing system, a method and a connection device, which are used for adapting a super low latency communication protocol (such as a CXL protocol or a unified bus (UB) protocol) and a RDMA protocol.

[0008] In a first aspect, the application provides a data processing system, which comprises a first computing cluster, a second computing cluster and a connection device. The first computing cluster comprises a first computing device, and the second computing cluster comprises a second computing device. Further, the connection device is connected with the first computing device, and is used for managing memory address information of the first computing cluster provided by the first computing device and memory address information of the second computing cluster provided by the second computing device. The first computing device is used for receiving an access request, the access request being used for accessing a memory space of the second computing device, searching for an address of the memory space from the memory address information of the second computing cluster managed by the connection device, and accessing the memory space in the second computing device according to the address. In a possible implementation manner, the data processing system further comprises a network device, and the network device is used for connecting the first computing cluster and the second computing cluster. The network device is, for example, a RDMA network interface controller (RNIC). In a possible implementation manner, the connection device is connected with the first computing device through a super low latency communication protocol (such as a CXL protocol or a UB protocol).

[0009] In the above technical solution, the connection device manages memory address information of computing clusters to which a plurality of computing devices belong and which are respectively provided by the plurality of computing devices. When a computing device connected with the connection device needs to access a memory space in a computing device in another computing cluster across the computing clusters, the address used for accessing the memory space in the computing device in the other computing cluster can be searched from the connection device, and then the memory space is accessed according to the address. In this way, RDMA protocol-based access across the computing clusters is adapted to super low latency communication protocol-based access in the computing clusters.

[0010] In a possible implementation manner, the first computing cluster further comprises a third computing device, and the connection device is used for connecting the first computing device and the third computing device. The first computing device is further used for obtaining memory address information of the third computing device, addressing the memory address information of the third computing device and the memory address information of the first computing device to form the memory address information of the first computing cluster, and sending the memory address information of the first computing cluster to the connection device.

[0011] In the technical solution, the first computing device can obtain the memory address information of other computing devices in the first computing cluster to which the first computing device belongs, and address according to the memory address information of other computing devices in the first computing cluster to which the first computing device belongs and the memory address information of the first computing device, to form the memory address information of the first computing cluster from the perspective of the first computing device. In this way, the first computing device can access the memory space in the first computing cluster based on the ultra-low latency communication protocol, which helps to improve the access speed, i.e., reduce the access latency.

[0012] In a possible implementation, the memory address information of the first computing device includes information of a plurality of memories in the first computing device; and the memory address information of the third computing device includes information of a plurality of memories in the third computing device. In the technical solution, each computing device can include a plurality of types of memories, which helps to improve the flexibility of memory access.

[0013] In a possible implementation, the connection device is further configured to receive the memory address information of the first computing cluster from the first computing device, and / or receive the memory address information of the second computing cluster from the second computing device. In the technical solution, the connection device can receive the memory address information of the computing cluster to which the computing device belongs from the computing device, and then uniformly manage the memory address information from a plurality of computing devices, so as to realize that a certain computing device connected to the connection device can access the memory space in the computing device in another computing cluster across the computing clusters.

[0014] In a possible implementation, the memory address information of the first computing cluster provided by the first computing device is specifically obtained by the first computing device from the operating system layer by addressing the memory of the first computing cluster, and the memory address information of the second computing cluster provided by the second computing device is specifically obtained by the second computing device from the operating system layer by addressing the memory of the second computing cluster. Further, the connection device is further configured to further address the sum of the memory of the first computing cluster and the second computing cluster from the application layer to obtain the memory address information on the application layer, and establish the mapping relationship between the memory address information on the application layer and the memory address information of each computing device on the operating system layer. In this way, the data processing system can provide the lengths of a plurality of computing clusters together to the application, so that the application can access any memory space in the plurality of computing clusters, and the flexibility of application access is improved.

[0015] In a possible implementation, the first computing device receives the access request, and the access request specifically includes a first address of application delivery. When the first computing device searches for the address of the memory space from the memory address information of the second computing cluster managed by the connection device, the first computing device specifically sends a query request to the connection device, and the query request includes the first address. Correspondingly, the first computing device receives a query response from the connection device, and the query response includes a second address. The second address is obtained by the connection device according to a mapping relationship between the memory address information at the application level and the memory address information of the second computing cluster at the operating system level of the second computing device, and the first address.

[0016] In a possible implementation, the first computing device further determines that the first address accessed by the access request is not in the first computing cluster, that is, the first address does not belong to the memory address in the first computing cluster. For example, the first computing device obtains the memory address information at the application level corresponding to the first computing cluster from the connection device, and determines that the first address does not belong to the memory address in the first computing cluster according to the first address and the memory address information at the application level corresponding to the first computing cluster. Further, the first computing device further determines that the first address belongs to the memory address in the second computing device of the second computing cluster, that is, the access request is an access request for accessing the memory space of the second computing device.

[0017] In addition, in another possible implementation, the first computing device further determines that the first address is in the first computing cluster, that is, the first address belongs to the memory address in the first computing cluster, and thus accesses the memory space indicated by the first address in the first computing cluster. For example, the first computing device obtains the memory address information at the application level corresponding to the first computing cluster from the connection device, and determines that the first address belongs to the memory address in the first computing cluster according to the first address and the memory address information at the application level corresponding to the first computing cluster.

[0018] In a second aspect, the present application provides a connection device, which includes: a transceiver module, configured to receive a query request sent by a first computing device, the query request being used to query an address accessed by an access request of an address space of a second computing device of a second computing cluster; and a management module, configured to send an address obtained by the query to the first computing device through the transceiver module according to the query request and memory address information of the second computing cluster. The connection device is connected with the first computing device in a first computing cluster, and the management module is configured to manage memory address information of the first computing cluster provided by the first computing device and memory address information of the second computing cluster provided by the second computing device.

[0019] In a possible implementation, the transceiving module is further configured to receive memory address information of the first computing cluster from the first computing device, and / or receive memory address information of the second computing cluster from the second computing device.

[0020] In a third aspect, the present application provides a computing device, which is specifically a first computing device, comprising: a transceiving module configured to receive an access request, the access request being used to access a memory space of a second computing device, the first computing device being located in a first computing cluster, and the second computing device being located in a second computing cluster; a searching module configured to search for an address of the memory space from a connection device; and an accessing module configured to access the memory space of the second computing device according to the address; wherein the connection device is connected to the first computing device, and the connection device is configured to manage memory address information of the first computing cluster provided by the first computing device, and memory address information of the second computing cluster provided by the second computing device.

[0021] In a possible implementation, the first computing cluster and the second computing cluster are connected through a network device.

[0022] In a possible implementation, the connection device is connected to the first computing device through a CXL protocol or a UB protocol.

[0023] In a possible implementation, the first computing cluster further comprises a third computing device, the connection device is configured to connect the first computing device and the third computing device, the processing module is further configured to obtain memory address information of the third computing device, and the memory address information of the first computing cluster is formed according to the memory address information of the third computing device and the memory address information of the first computing device; and the transceiving module is further configured to send the memory address information of the first computing cluster to the connection device.

[0024] In a possible implementation, the memory address information of the first computing device comprises information of a plurality of memories in the first computing device, and the memory address information of the third computing device comprises information of a plurality of memories in the third computing device.

[0025] In a possible implementation, the processing module is further configured to determine, after the transceiving module receives the access request, that the access request is an access request used to access the memory space of the second computing device, based on the fact that an address accessed by the access request (i.e., a first address) does not belong to the memory address of the first computing cluster.

[0026] In a fourth aspect, the present application provides a data processing method, comprising: receiving, by a first computing device, an access request, the access request being used to access a memory space of a second computing device in a second computing cluster; sending, to a connection device, a query request, the query request being used to query an address accessed by the access request, the first computing device being located in a first computing cluster; receiving, by the connection device, the query request, and sending, to the first computing device, an address obtained by querying according to the query request and memory address information of the second computing cluster; and accessing, by the first computing device, the memory space in the second computing device according to the received address.

[0027] In a possible implementation, the connection device further manages the memory address information of the first computing cluster provided by the first computing device, and the memory address information of the second computing cluster provided by the second computing device.

[0028] In a possible implementation, the connection device is connected to the first computing device through a CXL protocol or a UB protocol.

[0029] In a possible implementation, the first computing cluster further comprises a third computing device, and the connection device is used to connect the first computing device and the third computing device; the first computing device further obtains memory address information of the third computing device; the first computing device addresses the memory address information of the third computing device and the memory address information of the first computing device to form the memory address information of the first computing cluster; and the first computing device sends the memory address information of the first computing cluster to the connection device.

[0030] In a possible implementation, the memory address information of the first computing device comprises information of a plurality of memories in the first computing device; and the memory address information of the third computing device comprises information of a plurality of memories in the third computing device.

[0031] In a possible implementation, after receiving the access request, the first computing device determines that the access request is an access request used to access the memory space of the second computing device based on the fact that the address accessed by the access request does not belong to the memory address of the first computing cluster.

[0032] In a possible implementation, the connection device further receives the memory address information of the first computing cluster from the first computing device, and / or receives the memory address information of the second computing cluster from the second computing device.

[0033] In a fifth aspect, the present application provides a computer readable storage medium, the computer readable storage medium storing a computer program or instructions, when the computer program or instructions are executed by a device, the method in any possible implementation of the fourth aspect is implemented.

[0034] In a sixth aspect, the present application provides a computer program product, which comprises a computer program or instructions, and when the computer program or instructions are executed by a device, the method in any possible implementation manner of the fourth aspect is realized.

[0035] In a seventh aspect, the present application provides a device, comprising a processor and a memory, wherein the memory is configured to store a computer program, and the processor is configured to execute the computer program stored in the memory, so that the device realizes the method in any possible implementation manner of the fourth aspect.

[0036] The technical effects that can be achieved by any one of the second aspect to the seventh aspect can refer to the description of the beneficial effects of the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 A schematic diagram of a host transmitting data to a device through a PCIe protocol is provided in the present application;

[0038] Figure 2 A structural schematic diagram of a data processing system is provided in the present application;

[0039] Figure 3 A schematic diagram of a memory pool of a data processing system is provided in the present application;

[0040] Figure 4 A flowchart of a data processing method is provided in the present application;

[0041] Figure 5 A schematic diagram of a ring synchronization mode is provided in the present application;

[0042] Figure 6 A flowchart of another data processing method is provided in the present application;

[0043] Figure 7 A structural schematic diagram of a data processing apparatus is provided in the present application. DETAILED DESCRIPTION

[0044] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0045] In order to facilitate the explanation of the embodiments of the present application, the professional terms and technologies involved in the embodiments of the present application are explained first.

[0046] 1. The peripheral component interconnect express (PCIe) protocol is a high-speed serial point-to-point double-channel high-bandwidth transmission. Figure 1As a schematic diagram for a host to transmit data to a device through a PCIe protocol, specifically, a central processing unit (CPU) in the host needs to write data to a local memory in the host first, and then set a flag, wherein the flag is used to inform the device that the host has written data to the local memory in the host. Correspondingly, the device queries the flag, determines that the host has written data to the local memory in the host according to the flag, and then reads data from the local memory in the host, and stores the read data to a local memory in the device. The device transmits data to the host through the PCIe protocol in a manner similar to the manner in which the host transmits data to the device through the PCIe protocol, which will not be described again.

[0047] 2. A compute express link (CXL) protocol, the CXL protocol is a new protocol based on the PCIe protocol and optimized for cache and memory, and the CXL protocol runs on a PCIe physical layer.

[0048] The CXL protocol can be divided into the following three protocols:

[0049] The CXL.io protocol is an enumeration configuration protocol, mainly used for discovery and enumeration of devices, reporting errors, etc.

[0050] The CXL.cache protocol enables devices to solve memory consistency and access host memory with low latency. Specifically, the CXL.cache protocol mainly provides the ability to quickly access host memory by maintaining cache consistency on the device side. The protocol allows devices to participate in the consistency cache protocol of the CPU in the host. The device can directly use the host memory and the local cache of the host, and the host can also obtain data from the cache of the device without using memory as an intermediary.

[0051] The CXL.mem protocol is used to enable the host to access device memory as if it were accessing its own local memory. In the CXL.mem protocol, the CPU is used to send requests to the device, and the device is used to reply responses to the CPU. The requests sent by the CPU are divided into data requests and non-data requests, and correspondingly, the responses replied by the device are also divided into data responses and non-data responses.

[0052] 3. A unified bus (UB) protocol, which is a super-low latency communication protocol independently developed by Huawei, is used to provide high-performance interconnection for data centers, and make the interconnection of data centers like a computer.

[0053] 4. Memory pool: A program can pre-allocate a large block of memory from the system to create a memory pool. Memory allocation and deallocation are then performed within this pool. When the memory pool is insufficient, the program requests additional memory from the system.

[0054] 5. The load / store instructions are ARM (Advanced RISC Machine) instructions used for transferring data between registers and memory. Specifically, the load instruction loads data from memory into a register; the store instruction stores data from a register into memory. Since other ARM instructions can only operate on registers, when other ARM instructions need to manipulate data, they must first load the data from memory into a register using the load instruction. Conversely, data from completed execution of other ARM instructions needs to be stored from registers into memory using the store instruction.

[0055] Based on the explanations of the aforementioned technical terms, such as Figure 2 This is a schematic diagram of the structure of a data processing system provided as an example in this application.

[0056] This data processing system comprises multiple computing clusters. Taking any one computing cluster as an example, it includes multiple computing devices (computer nodes). Each computing device is connected to a connection device; that is, the connection device connects any two computing devices among these multiple computing devices, and the two computing devices communicate based on an ultra-low latency communication protocol. Each computing cluster can be considered to correspond to its own connection device, which may be located within or outside the computing cluster. The ultra-low latency communication protocol could be, for example, the CXL protocol or the UB protocol. Alternatively, it can be understood that any computing device within a computing cluster can, based on the ultra-low latency communication protocol, perceive the memory of other computing devices within its cluster, and thus address the memory of all computing devices within its cluster.

[0057] Multiple computing clusters are further networked using the RDMA protocol. Any two computing clusters are connected via network devices, such as RDMA network interface controllers (RNICs). That is, the network devices are used to connect any two computing clusters within the multiple computing clusters. For example, two computing devices located in two different computing clusters each include an RNIC, and these two RNICs can communicate based on the RDMA protocol, thus enabling communication between the computing clusters to which these two computing devices belong.

[0058] Figure 2In the example, the data processing system includes computing devices 1-8, and the computing devices 1-8 respectively include memories 1-8. Accordingly, the memories 1-8 form a memory pool. Figure 2 The number of computing clusters, the number of computing devices in each computing cluster, and the positional relationship between the connection device and the computing cluster do not constitute a limitation on the structure of the data processing system in the present application.

[0059] In the data processing system, the memories in the plurality of computing devices can form a memory pool. In combination with Figure 2 In the example, the data processing system includes computing devices 1-8, and the computing devices 1-8 respectively include memories 1-8. Accordingly, the memories 1-8 form a memory pool.

[0060] Further, taking any one computing device as an example, the computing device includes one or more processors, such as a CPU, a GPU, a neural-network processing unit (NPU), an FPGA, etc. For any one processor in the computing device, the memory pool can include one or more memories in different forms.

[0061] Based on the access speed from fast to slow, four kinds of memory forms are exemplarily provided as follows:

[0062] (1) Device attached memory (DAM): For a CPU, it is, for example, a double data rate synchronous dynamic random access memory (DDR SDRAM) linked to the CPU; for a GPU, it is, for example, a high bandwidth memory (HBM) in the GPU; for an NPU, it is, for example, an HBM and a DDR in the NPU, etc. Exemplarily, the access latency of the device attached memory is less than 100 ns.

[0063] (2) Device local memory (DLM): host memory within the computing device to which the GPU belongs for a GPU; and, for a CPU, an expander memory within the computing device to which the CPU belongs. For example, the access latency of the device local memory is in the range of 100ns to 300ns.

[0064] (3) Small network memory (SNM): memory within the computing cluster. For example, the access latency of the SNM is in the range of 300ns to 600ns. Figure 2 For example, computing device 1 and computing device 2 are within computing cluster 1, and for the CPU in computing device 1, memory 2 in computing device 2 is the SNM of the CPU in computing device 1.

[0065] (4) Large network memory (LNM): memory across the computing clusters. For example, the access latency of the LNM is greater than 600ns. Figure 2 For example, computing device 1 is within computing cluster 1, and computing device 5 is within computing cluster 2, and for the CPU in computing device 1, memory 5 in computing device 5 is the LNM of the CPU in computing device 1.

[0066] For example, the access latency of the SNM is in the range of 300ns to 600ns. Figure 2 , Figure 3 An exemplary diagram of a memory pool of a data processing system is provided. In the diagram, computing device 1 includes processor 11 and processor 12, computing device 2 includes processor 2, and computing device 5 includes processor 5. Further, computing device 1 includes device direct memory (denoted as MEM 11) of processor 11, device direct memory (denoted as MEM 121) of processor 12, and expander memory (denoted as MEM 122) of processor 12. Computing device 2 includes device direct memory (denoted as MEM 2) of processor 2. Computing device 5 includes device direct memory (denoted as MEM 5) of processor 5. From the perspective of processor 11, MEM 11 is device direct memory, MEM 121 and MEM 122 are device local memory, MEM 2 is small network memory, and MEM 5 is large network memory. Processor 11 can access MEM 2 based on a super low latency communication protocol, and access MEM 5 via RDMA, where processor 11 can access MEM 5 based on RNIC 1 in computing device 1 and RNIC 5 in computing device 5 (as described in step 606 below).

[0067] For example, the access latency of the SNM is in the range of 300ns to 600ns. Figure 2and Figure 3 The present application provides a data processing method, which is specifically an addressing method. The addressing method can be executed by a first connection device corresponding to a first computing cluster. The first connection device can uniformly manage memory address information of the first computing cluster and memory address information of a second computing cluster, so as to enable a first computing device in the first computing cluster to access a second computing device in the second computing cluster.

[0068] For example, the first computing cluster can be computing cluster 1 in Figure 2 , and the second computing cluster can be computing cluster 2 in Figure 2 . Correspondingly, the first computing device can be any one of computing device 1 to computing device 4, and the second computing device can be any one of computing device 5 to computing device 8.

[0069] Alternatively, the first computing cluster in the data processing method can be computing cluster 2 in Figure 2 , and the second computing cluster can be computing cluster 1 in Figure 2 . Correspondingly, the first computing device can be any one of computing device 5 to computing device 8, and the second computing device can be any one of computing device 1 to computing device 4.

[0070] Figure 4 The present application provides a flowchart of a data processing method:

[0071] Step 401: A first computing device acquires memory address information of a third computing device.

[0072] The third computing device is any computing device in the first computing cluster except the first computing device. For example, in Figure 2 , the first computing cluster is computing cluster 1, the first computing device is computing device 1, and the third computing device is computing device 2. The first computing device and the third computing device are connected through the first connection device.

[0073] The third computing device can include multiple memories. For example, the third computing device includes DDR SDRAM linked to CPU, HBM in GPU, and extended memory. The memory address information of the third computing device can be information of the multiple memories, such as the size and type of each memory in the multiple memories.

[0074] Step 402: The first computing device addresses to form memory address information (denoted as first memory address information) of the first computing cluster according to the memory address information of the third computing device and the memory address information of the first computing device.

[0075] The first computing device can include a plurality of memories. For example, the first computing device includes DDR SDRAM linked to the CPU, HBM in the GPU chip, and extended memory, etc. The memory address information of the first computing device can be information of the plurality of memories, such as the size, type, etc. of the plurality of memories.

[0076] Specifically, the first computing device addresses according to the memory address information of all computing devices in the first computing cluster (including the memory address information of the third computing device and the memory address information of the first computing device), to form the memory address information of the first computing cluster from the perspective of the first computing device (i.e. the first memory address information). In combination with the above Figure 2 For example, the first computing cluster is computing device 1, which can obtain the memory address information of computing device 1 to computing device 4, and address according to the memory address information of computing device 1 to computing device 4. For example, computing device 1 to computing device 4 each have 25G of memory, i.e. computing cluster 1 has a total of 100G of memory, and then computing device 1 can address the 100G of memory from the perspective of the first computing device.

[0077] It can also be understood that, since the computing devices in the first computing cluster are connected based on the ultra-low latency communication protocol, the first computing device can perceive the memories of the other computing devices in the first computing cluster except itself, and then address the memories of all computing devices in the first computing cluster to obtain the first memory address information. Further, for other computing devices (such as the third computing device) in the first computing cluster, the memories of the other computing devices in the first computing cluster except itself can also be perceived, and then the memories of all computing devices in the first computing cluster are addressed to obtain the memory address information of the first computing cluster from the perspective of the other computing device.

[0078] It should be noted that the two computing devices address from their respective perspectives, and the first memory address information of the first computing cluster obtained by the two computing devices is different, in combination with the above Figure 2 For example, computing device 1 can consider that the address number of its local 25G of memory is from the 1G to the 25G, and consider that the address number of the 25G of memory of computing device 2 is from the 26G to the 50G; and computing device 2 can consider that the address number of its local 25G of memory is from the 1G to the 25G, and consider that the address number of the 25G of memory of computing device 1 is from the 26G to the 50G.

[0079] In step 403, the first computing device sends the first memory address information to the first connection device. Correspondingly, the first connection device receives the first memory address information from the first computing device.

[0080] Further, other computing devices in the first computing cluster can also send their addressed memory address information of the first computing cluster to the first connection device after the addressing. Accordingly, the first connection device can store the memory address information of the first computing cluster from each computing device in the first computing cluster, which is addressed from the perspective of the computing device. In combination Figure 2 For example, the memory address information of the first computing cluster addressed from the computing device 1 to the computing device 4 is denoted as memory address information 1 to memory address information 4, and the first connection device records the memory address information 1 to memory address information 4. In this application, the memory address information of the first computing cluster recorded in the first connection device from the perspective of each computing device in the first computing cluster is referred to as the memory address information set of the first computing cluster.

[0081] Similarly, each computing device in the second computing cluster can also address the address of all computing devices in the second computing cluster to obtain the memory address information of the second computing cluster from the perspective of each computing device in the second computing cluster, and the second connection device records the memory address information of the second computing cluster from the perspective of each computing device in the second computing cluster. For specific implementation, please refer to the description in steps 401 to 403. In combination Figure 2 For example, the memory address information of the second computing cluster addressed from the computing device 5 to the computing device 8 is denoted as memory address information 5 to memory address information 8, and the second connection device records the memory address information 5 to memory address information 8. In this application, the memory address information of the second computing cluster recorded in the second connection device from the perspective of each computing device in the second computing cluster is referred to as the memory address information set of the second computing cluster.

[0082] Further, the first connection device can also obtain the memory address information set of the second computing cluster recorded in the second connection device. In one example, the first connection device is directly connected with the second connection device, and the first connection device can directly synchronize the memory address information set of the second computing cluster recorded in the second connection device from the second connection device. In another example, the first connection device is connected with the first computing device, and the second connection device is connected with the second computing device, and the first computing device is connected with the second computing device through a network device (such as RNIC). For example, the second connection device first synchronizes the memory address information set of the second computing cluster to the second computing device based on the ultra-low latency communication protocol, and the second computing device then synchronizes the memory address information set of the second computing cluster to the first computing device based on the RDMA protocol, and then the first computing device synchronizes the memory address information set of the second computing cluster to the first connection device based on the ultra-low latency communication protocol. In combination Figure 2For example, the first connection device can acquire the memory address information 5 to the memory address information 8 recorded in the second connection device, i.e., the first connection device can manage the memory address information 1 to the memory address information 8.

[0083] In addition, the memory address information in each of the computing devices in the second computing cluster can also be synchronized to the first connection device by the RDMA protocol. In combination with Figure 2 For example, the computing device 5 synchronizes the memory address information 5 to the first connection device, the computing device 6 synchronizes the memory address information 6 to the first connection device, the computing device 7 synchronizes the memory address information 7 to the first connection device, and the computing device 8 synchronizes the memory address information 8 to the first connection device. In this way, the first connection device can manage the memory address information 1 to the memory address information 8.

[0084] It should be noted that the above Figure 4 In the related embodiments, the first computing device addresses the memory of the first computing cluster at the operating system level (or the hardware level). Further, the first connection device can also address the sum of the memories of the plurality of computing clusters in the data processing system to obtain the memory address information at the application level. Then, the first connection device establishes the mapping relationship between the memory address information at the application level and the memory address information of each computing device at the operating system level.

[0085] In combination with Figure 2 For example, Table 1 below is a mapping relationship managed by the first connection device provided in the present application. In Table 1, each of the addresses 1 to n+m corresponds to the same address length, such as 4 kb, 8 kb, etc. The computing cluster 1 corresponds to the addresses 1 to n at the application level, i.e., the length of the addresses 1 to n at the application level is 100G, and n is a positive integer, where n is specifically the ratio of 100G to the address length. The computing cluster 2 corresponds to the addresses n+1 to n+m at the application level, i.e., the length of the addresses n+1 to n+m at the application level is 100G, and m is a positive integer, where m is specifically the ratio of 100G to the address length.

[0086] Further, from the perspective of different computing devices, the same address at the application level corresponds to different addresses at the operating system level. For example, the address 1 at the application level corresponds to the address 1-1 in the computing device 1, to the address 2-1 in the computing device 2, to the address 3-1 in the computing device 3, and to the address 4-1 in the computing device 4. For another example, the address n+1 at the application level corresponds to the address 5-1 in the computing device 5, to the address 6-1 in the computing device 6, to the address 7-1 in the computing device 7, and to the address 8-1 in the computing device 8.

[0087] Table 1

[0088]

[0089] Further, the first connection device can send the first computing cluster's multiple application layer addresses and the length of each address to the first computing device, such as sending the above addresses 1 to n, a total of n addresses, and the length of each address (such as 4kb) to the first computing device. Alternatively, the first connection device can send the first computing cluster's application layer's starting address and total length to the first computing device, such as sending the above address 1 and 100G to the first computing device. In addition, the first connection device can also send the second computing cluster's multiple application layer addresses and the length of each address to the first computing device, or send the second computing cluster's application layer's starting address and total length to the first computing device. The information sent by the first connection device to the first computing device is collectively referred to as the first information in this application. Accordingly, the first computing device can determine whether the memory space indicated by the address to be accessed by the application is located in the first computing cluster according to the first information from the first connection device (the specific implementation can be referred to in the description of step 602 below).

[0090] Similarly, the second connection device can also obtain the memory address information of the first computing cluster recorded in the first connection device from the perspective of each computing device in the first computing cluster, and address the sum of the memories of the multiple computing clusters in the entire data processing system to obtain the memory address information at the application layer, which can be used by the application running on the computing device in the second computing cluster to access the entire data processing system. Further, the second connection device can also send the second computing cluster's multiple application layer addresses and the length of each address to the first computing device, or send the second computing cluster's application layer's starting address and total length to the second computing device. The specific implementation is similar to the above first connection device and will not be described again.

[0091] It should be noted that the data processing system of the present application can have K computing clusters, K being a positive integer, and for a computing cluster, the connection device corresponding to the computing cluster can record a memory address information set of the computing cluster. In order to realize that the connection device corresponding to each computing cluster can obtain the memory address information set in the connection device of all other computing clusters, the present application provides a ring synchronization mode, that is, K connection devices corresponding to K computing clusters can form multiple ring structures, each ring structure includes multiple connection devices capable of communicating with each other based on a super low latency communication protocol, so that the multiple connection devices in the ring structure can exchange their respective memory address information sets based on the super low latency communication protocol. Subsequently, each ring structure in the multiple ring structures further provides a connection device, that is, the connection device provided by each ring structure in the multiple ring structures further forms a new ring structure based on the RDMA protocol, and the multiple connection devices in the new ring structure can exchange the memory address information set of the ring structure in which each connection device is located based on the RDMA protocol.

[0092] As Figure 5 A schematic diagram of a ring synchronization mode provided by the present application is shown in the figure, in which the thin solid circle represents the connection device, and the thick solid circle represents the network device. Among them, connection device 1 to connection device 8 form ring 1, connection device 9 to connection device 16 form ring 2, connection device 17 to connection device 24 form ring 3, connection device 25 to connection device 32 form ring 4, and connection device 1, connection device 9, connection device 17 and connection device 25 form ring 5. For convenience of description, the memory address information set of the computing cluster corresponding to each of connection device 1 to connection device 32 is respectively recorded as memory address information set 1 to memory address information set 32.

[0093] Take ring 1 as an example to explain the ring transmission:

[0094] In the first ring transmission:

[0095] Connection device 1 sends memory address information set 1 to connection device 2, connection device 2 sends memory address information set 2 to connection device 3,..., connection device 7 sends memory address information set 7 to connection device 8, and connection device 8 sends memory address information set 8 to connection device 1.

[0096] In the second ring transmission:

[0097] Connection device 1 sends memory address information set 8 to connection device 2, connection device 2 sends memory address information set 1 to connection device 3,..., connection device 7 sends memory address information set 6 to connection device 8, and connection device 8 sends memory address information set 7 to connection device 1.

[0098] Similarly, each of the connection devices 1 to 8 can obtain the memory address information set 1 to 8.

[0099] Similarly, in the ring 2 corresponding ring transmission, each of the connection devices 9 to 16 can obtain the memory address information set 9 to 16. In the ring 3 corresponding ring transmission, each of the connection devices 17 to 24 can obtain the memory address information set 17 to 24. In the ring 4 corresponding ring transmission, each of the connection devices 25 to 32 can obtain the memory address information set 25 to 32.

[0100] After the ring transmission of the rings 1 to 4 is completed, the connection devices 1, 9, 17 and 25 in the ring 5 further perform ring transmission (in a similar manner to the ring 1). Thus, the connection devices 1, 9, 17 and 25 can obtain the memory address information set 1 to 32. Then, the connection device 1 sends the memory address information set 1 to 32 to each of the connection devices in the ring 1, the connection device 9 sends the memory address information set 1 to 32 to each of the connection devices in the ring 2, the connection device 17 sends the memory address information set 1 to 32 to each of the connection devices in the ring 3, and the connection device 25 sends the memory address information set 1 to 32 to each of the connection devices in the ring 4.

[0101] Thus, the connection device corresponding to each computing cluster can not only manage the memory address information set of the computing cluster, but also manage the memory address information set of other computing clusters. When a computing device needs to access a computing device in other computing clusters than the computing cluster to which the computing device belongs, the computing device can first query the address it needs to access from the connection device corresponding to the computing cluster to which the computing device belongs, and then access according to the queried address.

[0102] As Figure 6 A flowchart of a data processing method is exemplarily provided in the present application, which is a data access method of a first computing device accessing a memory space in a second computing device.

[0103] In step 601, the first computing device receives an access request, wherein the access request includes a first address to be accessed.

[0104] In step 602, the first computing device determines whether the memory space indicated by the first address is located in the first computing cluster. If yes, step 603 is performed, otherwise, steps 604 to 606 are performed.

[0105] In one implementation, the first computing device generates the access request according to the memory address information at the application level, and the access request includes a first address to be accessed by the application, the first address being used to indicate an address space in a certain computing cluster of the data processing system. Figure 2 In the example in Table 1, the total memory of the computing cluster 1 is 100G, and the total memory of the computing cluster 2 is 100G. Therefore, the total memory of the entire data processing system is 200G, and the first computing device can provide 200G of memory space for the application, i.e., the application can access any address space in the 200G.

[0106] In one possible implementation, the first computing device determines, according to the first information, whether the memory space indicated by the first address is located in the first computing cluster. In the example in Table 1, the first computing device is the computing device 1, the first computing cluster is the computing cluster 1, and when the first address is the address 1 at the application level, the computing device 1 determines, according to the first information, that the memory space indicated by the first address is located in the computing cluster 1. When the first address is the address n+1 at the application level, the computing device 1 determines, according to the first information, that the memory space indicated by the first address is not located in the computing cluster 1, and further, the first computing device can determine, according to the first information, that the memory space indicated by the first address is located in the computing cluster 2, i.e., the computing cluster 2 is the second computing cluster.

[0107] In step 603, the first computing device accesses the memory space in the first computing cluster according to the first address.

[0108] The memory space indicated by the first address can be located in the first computing device or in other computing devices in the first computing cluster except the first computing device. In the example in Table 1, the first computing device is the computing device 1, the first computing cluster is the computing cluster 1, and when the first address is the address 1 at the application level, the computing device 1 can access the address 1-1 according to the address 1 at the application level, where the memory space indicated by the address 1-1 is located in the computing device 1. When the first address is the address n at the application level, the computing device 1 can access the address 1-n according to the address n at the application level, where the memory space indicated by the address 1-n is located in the computing device 4, and so on.

[0109] The first computing device accesses the memory space indicated by the first address through the ultra-low latency communication protocol. Specifically, when the access request is a write request, the access request further includes target data to be written to the first address, and the first computing device can store the target data into the memory space indicated by the first address. When the access request is a read request, the first computing device can load the target data from the memory space indicated by the first address.

[0110] At step 604, the first computing device sends a query request to the first connection device, and correspondingly, the first connection device receives the query request from the first computing device, wherein the query request includes the first address.

[0111] At step 605, the first connection device sends a query response to the first computing device, and correspondingly, the first computing device receives the query response from the first connection device, wherein the query response includes the second address.

[0112] Specifically, the first connection device determines the second address corresponding to the first address according to the first address in the query request and the mapping relationship between the memory address information at the application layer and the memory address information at the operating system layer of each computing device, and the second address is used for the first computing device to access the second computing device.

[0113] In combination with the example in Table 1, when the first address is the address n+1 at the application layer, the computing device 1 sends a query request to the first connection device, the first connection device determines that the second address is specifically address 5-1 when starting from the perspective of the computing device 5 according to the address n+1 at the application layer in the query request and the mapping relationship in Table 1, and then sends a query response to the computing device 1, wherein the query response includes the identifier and the address 5-1 of the computing device 5; or the first connection device can also determine that the second address is specifically address 6-1 when starting from the perspective of the computing device 6, and then send a query response to the computing device 1, wherein the query response includes the identifier and the address 6-1 of the computing device 6. It can be understood that the first connection device can determine the second address corresponding to the first address starting from the perspective of multiple different computing devices according to the first address, and further, the first connection device can select a computing device with low computing load, and send the identifier of the computing device with low computing load and the second address corresponding to the computing device to the first computing device.

[0114] At step 606, the first computing device accesses the memory space in the second computing device according to the second address.

[0115] Specifically, the first computing device accesses the second address in the second computing device based on the RDMA protocol.

[0116] When the access is a write request (i.e. RDMA write), the target data to be written into the first address is also included in the access request, the CPU of the second computing device registers the second address in advance to the RNIC of the second computing device and obtains a local key, wherein the local key is used to indicate that the RDMA has the right to operate the second address. Further, the RNIC of the second computing device encapsulates the second address and the local key into a special message and delivers it to the first computing device, that is, the RNIC of the second computing device gives the operating right of the second address to the first computing device. Correspondingly, after the RNIC of the first computing device receives the special message sent by the RNIC of the second computing device, it encapsulates the second address, the local key and the target data stored in the first computing device into an RDMA write and sends it to the second computing device, so that the first computing device can write the target data into the second address of the second computing device.

[0117] When the access is a read request (i.e. RDMA read), the CPU of the second computing device registers the second address in advance to the RNIC of the second computing device and obtains a local key, wherein the local key is used to indicate that the RDMA has the right to operate the second address. Further, the RNIC of the second computing device encapsulates the second address and the local key into a special message and delivers it to the first computing device, that is, the second computing device gives the operating right of the second address to the first computing device. Correspondingly, after the RNIC of the first computing device receives the special message sent by the RNIC of the second computing device, it encapsulates the third address for storing the target data, the second address and the local key into an RDMA read and sends it to the second computing device, so that the first computing device can read the target data from the second address of the second computing device to the third address of the first computing device.

[0118] It should be noted that the first computing device can also monitor the occupancy of each memory in the local. For example, the first computing device can include DDR SDRAM linked to the CPU, HBM in the GPU, and extended memory. The first computing device can monitor that the occupancy of the DDR SDRAM is 70%, the occupancy of the HBM is 50%, and the occupancy of the extended memory is 20%. Similarly, other computing devices can also monitor the occupancy of each memory in the local. Subsequently, the first computing device can obtain the occupancy of each memory of each computing device in the computing cluster to which the first computing device belongs, and obtain the occupancy of each memory of each computing device in other computing clusters. When the first computing device obtains the access request (specifically, the write request) of the application, the first computing device can determine in which memory to write the target data in the write request according to the occupancy of each memory in the entire data processing system. For example, the first computing device can consider the occupancy of each memory and the latency of the first computing device accessing each memory (i.e., the latency of accessing the directly connected memory of the device, the latency of accessing the local memory of the device, the latency of accessing the small network memory, and the latency of accessing the large network memory), and select a memory from multiple memories to write the target data. When the selected memory is in the first computing cluster, the first computing device can store the target data in the memory based on the ultra-low latency communication protocol. When the selected memory is in other computing clusters (such as the second computing cluster), the first computing device can write the target data into the memory based on the RDMA protocol. Further, the implementation of the first computing device obtaining the occupancy of each memory of each computing device in the computing cluster to which the first computing device belongs and the first computing device obtaining the occupancy of each memory of each computing device in other computing clusters can refer to the above Figure 5 In the related embodiments, the thin solid circle represents a computing device, and the thick solid circle represents a network device.

[0119] It should be further noted that the first computing device can also monitor the frequency at which the data stored in each local memory is read by the first computing device, and the same applies to the other computing devices. Then, the first computing device can obtain the frequency at which the data stored in each memory of the other computing devices in the computing cluster to which the first computing device belongs is read by the first computing device, and obtain the frequency at which the data stored in each memory of the computing devices in the other computing cluster is read by the first computing device. When the first computing device determines that the frequency at which a certain data is read by the first computing device is inconsistent with the frequency interval corresponding to the memory of the first computing device, the first computing device can move the data to the memory corresponding to the frequency interval to which the read frequency of the data belongs. It can be understood that each memory corresponds to a different frequency interval, such as the DDR SDRAM of the computing device 1 corresponding to the frequency interval 1, the HBM of the computing device 1 corresponding to the frequency interval 2, the extended memory in the computing device 1 corresponding to the frequency interval 3, the memory in the computing device 2 corresponding to the frequency interval 4, and the memory in the computing device 5 corresponding to the frequency interval 5, wherein the frequency interval 1, the frequency interval 2, the frequency interval 3, the frequency interval 4, and the frequency interval 5 decrease in turn. For example, the first computing device monitors that the data 1 is located in the memory in the computing device 5, and the frequency at which the first computing device reads the data 1 is located in the frequency interval 1, so the first computing device needs to move the data 1 to the DDR SDRAM of the computing device 1; for another example, the first computing device monitors that the data 2 is located in the HBM of the computing device 1, and the frequency at which the first computing device reads the data 2 is located in the frequency interval 4, so the first computing device needs to move the data 2 to the memory of the computing device 2. In this way, the first computing device can place the data with a higher read frequency in the memory with a lower read latency, which helps to improve the efficiency of data reading. Further, the implementation manner of the first computing device obtaining the frequency at which the data stored in each memory of the other computing devices in the computing cluster to which the first computing device belongs is read by the first computing device, and the first computing device obtaining the frequency at which the data stored in each memory of the computing devices in the other computing cluster is read by the first computing device can refer to the above Figure 5 In the related embodiments, the thin solid circle represents a computing device, and the thick solid circle represents a network device.

[0120] Based on the same inventive concept, the application provides a possible data processing system, which comprises: a first computing cluster comprising a first computing device; a second computing cluster comprising a second computing device; a first connection device connected with the first computing device, used for managing memory address information of the first computing cluster provided by the first computing device and memory address information of the second computing cluster provided by the second computing device; the first computing device is used for: receiving an access request, the access request being used for accessing a memory space of the second computing device; searching for an address of the memory space from the memory address information of the second computing cluster managed by the first connection device; and accessing the memory space in the second computing device according to the address.

[0121] In a possible implementation, the data processing system further comprises: a network device, used for connecting the first computing cluster and the second computing cluster.

[0122] In a possible implementation, the first connection device is connected with the first computing device through a CXL protocol or a UB protocol.

[0123] In a possible implementation, the first computing cluster further comprises a third computing device, the first connection device is used for connecting the first computing device and the third computing device; the first computing device is further used for: obtaining memory address information of the third computing device; addressing the memory address information of the third computing device and the memory address information of the first computing device to form the memory address information of the first computing cluster; and sending the memory address information of the first computing cluster to the first connection device.

[0124] In a possible implementation, the memory address information of the first computing device comprises information of a plurality of memories in the first computing device; the memory address information of the third computing device comprises information of a plurality of memories in the third computing device.

[0125] In a possible implementation, the first connection device is further used for: receiving the memory address information of the first computing cluster from the first computing device, and / or receiving the memory address information of the second computing cluster from the second computing device.

[0126] Based on the same inventive concept, as shown in Figure 7 The data processing apparatus 70 provided by the embodiment of the application can be applied to the flowchart shown above and perform the functions of the first computing device or the first connection device in the method embodiment described above. For the convenience of description, Figure 7 Only the main components of the apparatus are shown.

[0127] The data processing apparatus 70 comprises a processor 701, a memory 702 and a communication interface 703. Any two of the processor 701, the memory 702 and the communication interface 703 can be connected through a bus 704.

[0128] The processor 701 can be a central processing unit (CPU) that can be configured to execute program instructions in the memory 702 to implement the operations of the above-described embodiments of the related methods. In addition to the CPU, the processor 701 can also be an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a system on chip (SoC), or a complex programmable logic device (CPLD), a graphics processing unit (GPU), a neural-network processing unit (NPU), etc. Figures 4 to 6

[0129] It should be noted that in actual applications, the number of processors 701 can be multiple, and the multiple processors 701 can include multiple processors of the same type or multiple processors of different types, for example, the multiple processors 701 are multiple CPUs. For another example, the multiple processors 701 include one or more CPUs and one or more GPUs. For another example, the multiple processors 701 include one or more CPUs and one or more NPUs. Or, the multiple processors 701 include one or more CPUs, one or more GPUs, and one or more NPUs, etc. Among them, the processor 701 (such as CPU, NPU, etc.) can include one core or multiple cores.

[0130] The memory 702 refers to a device for storing program instructions and / or data, which can be a memory or a hard disk.

[0131] ​The memory refers to an internal memory directly exchanging data with the processor 701, which can read and write data at any time and at a very high speed, and is used as a temporary data storage for an operating system or other programs running on the processor 701. The memory includes volatile memory such as random access memory (RAM), dynamic random access memory (DRAM), and the like, and can also include non-volatile memory such as storage class memory (SCM) and the like, or a combination of volatile memory and non-volatile memory, and the like. In actual applications, multiple memories can be configured in the data processing apparatus 70, and optionally, the multiple memories can be of different types. The embodiments do not limit the number and types of memories. In addition, the memory can be configured to have a power retention function. The power retention function refers to that when the system is powered off and then powered on again, the data stored in the memory will not be lost. The memory with the power retention function is referred to as non-volatile memory.

[0132] The hard disk includes but is not limited to non-volatile memory such as read-only memory (ROM), a hard disk drive (HDD), or a solid-state disk (SSD), and the like. Unlike the memory, the hard disk has a slower read-write speed and is usually used to store data persistently. In an embodiment, the data, program instructions, and the like in the hard disk need to be loaded into the memory first, and then the processor obtains these data and / or program instructions from the memory.

[0133] The communication interface 703 is configured to communicate with other devices. For example, when the data processing apparatus 70 is a first computing device, the communication interface 703 is configured to receive an access request by the first computing device, send a query request to a connection device, receive a query response from the connection device, and the like; when the data processing apparatus 70 is a connection device, the communication interface 703 is configured to receive a query request from a first computing device by the connection device, and send a query response to the first computing device, and the like.

[0134] Based on the same inventive concept described above, the present application provides a computer-readable storage medium, which stores a computer program or instructions, and when the computer program or instructions are executed by a device, the above-mentioned Figures 4 to 6 The method in the related method embodiments.

[0135] Based on the same inventive concept, the present application provides a computer program product, which comprises computer programs or instructions, and when the computer programs or instructions are executed by a device, the above-mentioned Figures 4 to 6 The method in the related method embodiment.

[0136] Based on the same inventive concept, the present application provides a device, which comprises a processor and a memory connected to the processor, the memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory, so that the device realizes the above-mentioned Figures 4 to 6 The method in the related method embodiment.

[0137] Based on the same inventive concept, the present application provides a data processing apparatus, which can be specifically Figures 4 to 6 The first computing device or the first connection device in the related method embodiment, and correspondingly, when the data processing apparatus is the first computing device, the data processing apparatus can comprise modules (such as transceiving module, searching module and accessing module) for executing the functions of the above-mentioned first computing device; when the data processing apparatus is the first connection device, the data processing apparatus can comprise modules (such as transceiving module and management module) for executing the functions of the above-mentioned first connection device.

[0138] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of these items, including any combination of single item (s) or multiple items (s). For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", wherein a, b, c can be single or multiple. "And / or" describes the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B, which can mean: A exists alone, A and B exist together, and B exists alone, wherein A and B can be singular or plural. In the textual description of the present application, the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship; in the formula of the present application, the character " / ", indicates that the front and rear associated objects are in a "division" relationship.

[0139] It can be understood that various numerical numbers involved in the embodiments of the present application are only for convenient differentiation, and do not limit the scope of the embodiments of the present application. The size of the serial number of the above-mentioned processes does not mean the execution order, and the execution order of the processes should be determined according to its function and inherent logic.

[0140] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. A data processing system, characterized by The application relates to a computer cluster system, comprising: a first computing cluster, comprising a first computing device and a connection device connected with the first computing device through a super low delay communication protocol; a second computing cluster, comprising a second computing device; the second computing cluster is connected with the first computing cluster through a network device and communicates through a remote direct memory access (RDMA) protocol; the connection device acquires memory address information of the first computing cluster according to the super low delay communication protocol, acquires memory address information of the second computing cluster through the RDMA protocol, and addresses a sum of the memory address information of the first computing cluster and the memory address information of the second computing cluster to obtain application layer memory address information; the first computing device is used for: receiving an access request, wherein the access request carries first address information in the application layer memory address information; determining, according to the first address information, that the access request is used for accessing a device in the second computing cluster; determining, through the super low delay communication protocol, second address information of the second computing device corresponding to the first address information from memory address information of the second computing cluster managed by the connection device; accessing, according to the second address information, a memory space in the second computing device through the RDMA protocol.

2. The system of claim 1, wherein, The first computing cluster further comprises a third computing device, and the connection device is used for connecting the first computing device and the third computing device; the first computing device is further used for: acquiring memory address information of the third computing device; addressing the memory address information of the third computing device and the memory address information of the first computing device to form the memory address information of the first computing cluster; sending the memory address information of the first computing cluster to the connection device.

3. The system of claim 2, wherein, The memory address information of the first computing device comprises information of a plurality of memories in the first computing device; the memory address information of the third computing device comprises information of a plurality of memories in the third computing device.

4. The system of claim 1, wherein, The first computing device is further used for: after receiving the access request, determining, based on the address accessed by the access request not belonging to the memory address of the first computing cluster, that the access request is an access request for accessing the memory space of the second computing device.

5. The system of any one of claims 1-4, wherein, The connection device is further used for: receiving the memory address information of the first computing cluster from the first computing device, wherein the memory address information of the first computing cluster comprises first computing cluster memory address information formed by the first computing device addressing all the memory address information of the computing devices in the first computing cluster from the perspective of the first computing device; and / or, receiving the memory address information of the second computing cluster from the second computing device, wherein the memory address information of the second computing cluster comprises second computing cluster memory address information formed by the second computing device addressing all the memory address information of the computing devices in the second computing cluster from the perspective of the second computing device.

6. A data processing method, characterized by, The application relates to a computer cluster system, comprising: A first computing device of a first computing cluster receives an access request, the access request carrying first address information in application-layer memory address information; The first computing device determines, according to the first address information, that the access request is used to access a computing device in a second computing cluster; the second computing cluster is connected to the first computing cluster through a network device and communicates through a remote direct memory access (RDMA) protocol; The first computing device sends a query request to a connection device in the first computing cluster through an ultra-low latency communication protocol, the query request being used to query an address of a memory space of a second computing device in the second computing cluster corresponding to the first address information, wherein the connection device is used to obtain memory address information of the first computing cluster according to the ultra-low latency communication protocol, obtain memory address information of the second computing cluster through the RDMA protocol, and address a sum of the memory address information of the first computing cluster and the memory address information of the second computing cluster to obtain application-layer memory address information; The connection device receives the query request, determines, according to the query request, second address information of the second computing device corresponding to the first address information in the memory address information of the second computing cluster, and sends the second address information to the first computing device; The first computing device accesses the memory space in the second computing device through the RDMA protocol according to the second address information.

7. The method of claim 6, wherein, The first computing cluster further includes a third computing device, and the connection device is used to connect the first computing device and the third computing device; The method further includes: The first computing device obtains memory address information of the third computing device; The first computing device addresses the memory address information of the third computing device and the memory address information of the first computing device to form the memory address information of the first computing cluster; The first computing device sends the memory address information of the first computing cluster to the connection device.

8. The method of claim 7, wherein, The memory address information of the first computing device includes information of a plurality of memories in the first computing device; The memory address information of the third computing device includes information of a plurality of memories in the third computing device.

9. The method of claim 6, wherein, After the first computing device receives the access request, the method further includes: The first computing device determines, based on the address accessed by the access request not belonging to the memory address of the first computing cluster, that the access request is an access request used to access the memory space of the second computing device.

10. The method of any one of claims 6-9, wherein, Further comprising: The connection device receives the memory address information of the first computing cluster from the first computing device, the memory address information of the first computing cluster including first computing cluster memory address information formed by the first computing device addressing all memory address information of computing devices in the first computing cluster from the perspective of the first computing device; And / or, receiving the memory address information of the second computing cluster from the second computing device, the memory address information of the second computing cluster including the memory address information of all computing devices in the second computing cluster according to the second computing device, forming the memory address information of the second computing cluster from the perspective of the second computing device.

11. A connecting device, characterized in that Comprise: A transceiving module, configured to receive a query request sent by a first computing device in a first computing cluster, the query request carrying first address information in memory address information at an application layer, and used to query an address of a memory space of a second computing device in a second computing cluster corresponding to the first address information, the first computing cluster including a connection device, the connection device being configured to obtain memory address information of the first computing cluster according to an ultra-low latency communication protocol, and obtain memory address information of the second computing cluster through an RDMA protocol, and encode a sum of the memory address information of the first computing cluster and the memory address information of the second computing cluster to obtain the memory address information at the application layer; A management module, configured to find the address of the memory space of the second computing device in the memory address information of the second computing cluster according to the query request, and send the queried address to the first computing device.

Citation Information

Patent Citations

  • Data processing method and device based on distributed shared memory system

    CN113590364A

  • Data transmission method, processor system and memory access system

    CN113852656A

  • Storage cluster interconnection method and device, computer equipment and storage medium

    CN114357049A