A data transmission method, apparatus, computing device, and storage medium

By obtaining the shared memory address of the target device in the computing device and using unified addressing technology, the shared memory is directly connected to the PCIe bus, the problem of low data transmission efficiency between GMEM and SMEM is solved, and efficient data transmission is achieved.

CN119645919BActive Publication Date: 2025-06-03ZHEJIANG LAB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510169425.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-03
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

In the prior art, data transmission between devices is inefficient, especially in the asynchronous memory copying process between GMEM and SMEM, data must be transmitted through GMEM, resulting in slow access speed and low transmission efficiency.

Method used

By obtaining the shared memory address of the target device in the computing device and using unified addressing technology, the shared memory is directly connected to the PCIe bus, thereby skipping GMEM relay and directly accessing SMEM through the PCIe bus for data transmission.

Benefits of technology

It simplifies the data transmission process, improves data transmission efficiency, reduces transmission delay, and improves the execution speed of computing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119645919B_ABST
    Figure CN119645919B_ABST
Patent Text Reader

Abstract

The present invention discloses a data transmission method, apparatus, computing device, and storage medium, relating to the field of computer technologies. The method includes: in response to a data transmission request received from a target device, a computing device obtains a shared memory address of the target device from a storage medium, where the shared memory address refers to the unique address of the shared memory in the target device; wherein, the shared memory and the global memory in the target device both adopt a unified addressing technology, so that the shared memory is directly connected to the PCIe bus through the shared memory address. The computing device initiates a data transmission operation to the shared memory address through the PCIe bus. In this way, GMEM and SMEM use a unified address space, and the address of SMEM is exposed to the outside. This means that both the CPU and the RDMA controller can access these two types of memory through a single address space. This application solves the technical problem of low data transmission efficiency in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a data transmission method, apparatus, computing device, and storage medium. Background Art

[0002] Memory areas in a computing system generally include: host memory, global memory (GMEM), shared memory (SMEM), registers, and various caches.

[0003] Internal hardware components in a computing system are connected through PCIe (PCI Express). PCIe is a general serial connection standard suitable for connecting various memory areas for fast data transmission and sharing.

[0004] RDMA (Remote Direct Memory Access) is a high-performance network communication technology that allows nodes in a network to directly access the memory of another node without the intervention of the operating system or CPU, and is particularly suitable for high-performance computing and data center scenarios that require large-scale data exchange.

[0005] GDR (GPU Direct RDMA) is a technology introduced by NVIDIA that allows the GPU to directly exchange data with other devices supporting RDMA without going through the CPU for transfer. However, currently, the memory for GDR operations is the GPU's global memory (GMEM), rather than the thread-block-local shared memory (SMEM). For this reason, NVIDIA introduced a new feature TMA (Tensor Memory Accelerator) in the Hopper architecture for asynchronous memory copying between the GPU's GMEM and SMEM.

[0006] That is to say, in the traditional TMA architecture, data is first transmitted to GMEM and then transferred from GMEM to SMEM through asynchronous memory copying. The access speed of GMEM is slow, and the access speed of SMEM is fast. However, due to the fact that data must be transmitted through GMEM, the data transmission efficiency is low. Summary of the Invention

[0007] Embodiments of this application provide a data transmission method, apparatus, computing device, and storage medium to solve the technical problem of low data transmission efficiency between devices in the prior art.

[0008] To achieve the above object, the embodiments of this application adopt the following technical solutions:

[0009] In a first aspect, an embodiment of the present application provides a data transmission method, which is applied to a computing device. The computing device includes a storage medium. The method includes:

[0010] In response to a data transmission request received from a target device, obtain the shared memory address of the target device from the storage medium. The shared memory address refers to the unique address of the shared memory in the target device. Among them, the shared memory and the global memory in the target device both adopt a unified addressing technology, so that the shared memory address is used for the shared memory to be directly connected to the PCIe bus;

[0011] Initiate a data transmission operation to the shared memory address through the PCIe bus.

[0012] In combination with the first aspect, in a possible design, before obtaining the shared memory address of the target device from the storage medium, the method further includes:

[0013] Obtain the shared memory address provided by the target device through a message request of the PCIe protocol;

[0014] Write the shared memory address into the storage medium.

[0015] In combination with the first aspect, in a possible design, the computing device is a CPU or an RDMA controller, the target device is a computing device card, and the unified addressing technology includes:

[0016] Build a 2 64 -bit byte-sized address space;

[0017] Allocate the upper 16-bit addresses in the address space to multiple nodes in the target cluster, where the address space of each node is 256TB;

[0018] Allocate the address space of each node to the CPU and the computing device card corresponding to the node. The address space of the CPU is 128TB, and the address space of the computing device card is 128TB;

[0019] Among them, each node supports 64 computing device cards, each computing device card is allocated 2TB of address space, and the shared memory and the global memory respectively occupy 1TB of address space in the computing device card.

[0020] In combination with the first aspect, in a possible design, the obtaining the shared memory address provided by the target device through a message request of the PCIe protocol includes:

[0021] Send a message request to the target device; the message request is used to request the shared memory address of the target device;

[0022] Receive response information; the response information indicates the address space of the shared memory of the target device, and the response information is sent by the target device in response to the message request;

[0023] Obtain the shared memory address of the target device from the response information.

[0024] Combined with the first aspect, in a possible design, the target device includes an address register A, a register L, and a comparator. The register L indicates the n high-order bits of the address in the address register A. After initiating a data transfer operation to the shared memory address, the method further includes:

[0025] The target device verifies the shared memory address through the comparator to obtain a verification result;

[0026] If the verification result is true, the target device processes the data; if the verification result is false, the target device discards the data;

[0027] Wherein, the verification result indicates whether the L-n bit address of the shared memory address is the same as the L-n bit address in the address register A.

[0028] Combined with the first aspect, in a possible design, the initiating a data transfer operation to the shared memory address includes:

[0029] When the computing device initiates a read operation instruction to the target device through the PCIe bus, the target device locks the shared memory corresponding to the shared memory address;

[0030] The target device reads the data in the shared memory into the data buffer based on the data start address and data length information in the read operation instruction;

[0031] In the case where all the data is read from the shared memory into the data buffer, the target device releases the shared memory;

[0032] The target device sends the data to the PCIe bus based on the shared memory.

[0033] Combined with the first aspect, in a possible design, the initiating a data transfer operation to the shared memory address includes:

[0034] When the computing device issues a write operation instruction to the target device via the PCIe bus, the target device writes the data into the data buffer based on the data start address and data length information in the write operation instruction;

[0035] When all the data is written into the data buffer, the target device locks the shared memory corresponding to the shared memory address;

[0036] After the data is written from the data buffer into the shared memory, the target device releases the shared memory and sends a signal to the PCIe bus, and the signal indicates that the write operation instruction is successfully executed.

[0037] In a second aspect, an embodiment of the present application provides a data transmission device, and the device includes:

[0038] An address acquisition module, configured to, in response to a data transmission request received from a target device, acquire a shared memory address of the target device from a storage medium of a computing device, where the shared memory address refers to a unique address of a shared memory in the target device, and where the shared memory in the target device and the global memory both adopt a unified addressing technology so that the shared memory address is used for the shared memory to be directly connected to the PCIe bus;

[0039] A data transmission module, configured to initiate a data transmission operation to the shared memory address via the PCIe bus.

[0040] In a third aspect, an embodiment of the present application provides a computing device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the method of the first aspect and its possible design manners.

[0041] In a fourth aspect, an embodiment of the present application provides a storage medium, where a computer program is stored in the storage medium, and where the computer program is configured to execute the method of the first aspect and its possible design manners when running.

[0042] Compared with the prior art, a data transmission method, apparatus, computing device, and storage medium provided by an embodiment of the present application. The computing device responds to a data transmission request received from a target device, and obtains the shared memory address of the target device from the storage medium. The shared memory address refers to the unique address of the shared memory in the target device. Among them, the shared memory and the global memory in the target device both adopt a unified addressing technology, so that the shared memory is directly connected to the PCIe bus through the shared memory address. The computing device initiates a data transmission operation to the shared memory address through the PCIe bus, thereby performing data transmission with the target device. GMEM and SMEM use a unified address space, and the address of SMEM is exposed to the external PCIe bus. In this way, the computing device skips the GMEM transfer in the traditional method and can directly access the shared memory through the PCIe bus and transmit data to SMEM at one time. This means that both the CPU and the RDMA controller can access these two types of memory, GMEM and SMEM, through a single address space. This method solves the problem of address conversion required between GMEM and SMEM in the traditional architecture. Since the SMEM can be directly accessed externally and the read / write speed of SMEM is fast, the data transmission process is simplified and the data transmission efficiency is improved.

[0043] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0045] Figure 1 A hardware structure block diagram of a computing device provided by an embodiment of the present application is shown;

[0046] Figure 2 A schematic diagram of the interconnection relationship between multiple devices provided by an embodiment of the present application;

[0047] Figure 3 A flowchart of a data transmission method provided by an embodiment of the present application is shown;

[0048] Figure 4 A flowchart of a unified addressing technology provided by an embodiment of the present application is shown;

[0049] Figure 5 A flowchart of a method for obtaining a shared memory address provided by an embodiment of the present application is provided;

[0050] Figure 6Shows a flowchart of a method for reading data using a buffer provided by an embodiment of the present application;

[0051] Figure 7 Shows a flowchart of a method for writing data using a buffer provided by an embodiment of the present application;

[0052] Figure 8 Shows a system framework diagram of a computing device based on unified addressing provided by an embodiment of the present application;

[0053] Figure 9 Shows a structural block diagram of a data transmission device provided by an embodiment of the present application. Detailed implementation manners

[0054] For a clearer understanding of the purpose, technical solution and advantages of the present application, the present application will be described and illustrated below with reference to the accompanying drawings and embodiments.

[0055] Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the general meanings understood by those with ordinary skills in the technical field to which the present application belongs. In the present application, words such as "a", "one", "a kind of", "the", "these" and the like do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "comprising", "having" and any variants thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, device, product or equipment including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or equipment. The terms "connected", "coupled" and the like involved in the present application do not limit to physical or mechanical connections, but may include electrical connections, whether directly or indirectly. The "multiple" involved in the present application means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third" and the like involved in the present application are only used to distinguish similar objects and do not represent a specific sorting of the objects.

[0056] With the progress of technology, PCIe, RDMA, and GDR technologies have gradually emerged. PCIe is a general-purpose serial connection standard mainly used for connecting internal hardware components in a computing system, especially for connecting various memory regions for fast data transfer and sharing. RDMA is a high-performance network communication technology that allows nodes in a network to directly access the memory of another node without the intervention of the operating system or CPU, and is particularly suitable for high-performance computing and data center scenarios that require large-scale data exchange. GDR is a technology introduced by NVIDIA that allows the GPU to directly exchange data with other RDMA-enabled devices without going through the CPU for transfer. TMA is used for asynchronous memory copying between the GMEM and SMEM of the GPU.

[0057] Generally speaking, PCIe, RDMA, and GDR are all important technologies aimed at improving communication performance. However, in traditional technologies, the SMEM cannot be directly accessed. During the execution of write operation instructions, data is first transferred to the GMEM and then transferred from the GMEM to the SMEM through asynchronous memory copying. The PCIe bus cannot bypass the GMEM to directly access the SMEM during the asynchronous transfer process, so it not only increases the data transfer latency but also may cause data transfer bottlenecks.

[0058] Based on this, the embodiments of the present application provide a data transfer method applied to a computing device. The computing device can perform data transfer with a target device, obtain the unique address (shared memory address) of the shared memory in the target device, and store the shared memory address in the storage medium of the computing device. Then, when the computing device needs to read and write data to the target device, it can directly obtain the shared memory address from the storage medium to directly access the shared memory of the target device and transfer the data to the shared memory in one go, skipping the GMEM transfer in the traditional method and simplifying the intermediate links of data transfer, thus accelerating the transfer speed.

[0059] This method can be executed on a terminal, a computer, or a similar computing system. Taking the operation on a computing device as an example, Figure 1 shows a hardware structure block diagram of a computing device provided by the embodiments of the present application. As Figure 1 shown, the computing device may include one or more ( Figure 1 only one is shown in the figure) processors 102 and a memory 104 for storing data. The processor 102 may include, but is not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA. The above computing device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above computing device. For example, the computing device may further include moreFigure 1 more or fewer components as shown, or having a configuration different from that Figure 1 shown.

[0060] The memory 104 can be used to store computer programs, for example, software programs and modules of application software. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the above-mentioned method. The memory 104 can be used to store data. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computing device through a network. Examples of the above-mentioned network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.

[0061] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the computing device. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RadioFrequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0062] The embodiments of the present application can be applied to data transmission, and are particularly suitable for scenarios that require a large amount of data transmission, such as large model training, high-performance computing (HPC), machine learning, etc. In these scenarios, the speed and efficiency of data transmission directly affect the completion time and cost of computing tasks. By adopting the technology provided by the embodiments of the present application, the efficiency of data processing can be significantly improved, and the execution of computing tasks can be accelerated.

[0063] Figure 2 is a schematic diagram of the interconnection relationship between multiple devices provided by the embodiments of the present application, as Figure 2 shown, the cluster includes multiple computing nodes, such as computing node 0, computing node 1, computing node 2, and computing node 3. Each computing node is interconnected through a high-speed network (RoCE or IB network). Each computing node includes a CPU and multiple computing device cards, and the devices within the computing node are connected through PCIe, that is, the data between the devices in the computing node is transmitted through the PCIe interface. The computing device can be a CPU or an RDMA controller, and the target device can be a computing device card.

[0064] Figure 3The flowchart of a data transmission method provided by an embodiment of the present application is shown. As Figure 3 shown, the method includes steps S301 to S302.

[0065] Step S301: In response to a data transmission request received from a target device, a computing device obtains the shared memory address of the target device from a storage medium. The shared memory address refers to the unique address of the shared memory in the target device.

[0066] Among them, the shared memory and the global memory in the target device both adopt a unified addressing technology, so that the shared memory is directly connected to the PCIe bus through the shared memory address.

[0067] In the present application, the computing device and the target device refer to devices with a PCIe interface that follows the PCIe protocol. In the traditional technology, only the address of GMEM is exposed to the external PCIe bus. Therefore, the computing device can only access GMEM first through the PCIe bus and then read and write SMEM through GMEM. In this way, although the read and write speed of SMEM is very fast, due to the slow read and write speed of GMEM, the overall data transmission efficiency is low.

[0068] In this step, the unique address (shared memory address) of SMEM is stored in the storage medium by the computing device, so that when the computing device needs data transmission, it can directly access SMEM based on the shared memory address, thus skipping the link of accessing GMEM and directly transmitting all data to SMEM at one time. Since the read and write speed of SMEM is fast, directly connecting SMEM to the PCIe bus can improve the transmission efficiency.

[0069] Before step S301, the method further includes: the computing device obtains the shared memory address provided by the target device through a message request of the PCIe protocol; writes the shared memory address into the storage medium.

[0070] Specifically, before the computing device obtains the shared memory address of the target device, the computing device can also perform data transmission with the target device, such as initiating a message request through the PCIe protocol, so that the target device responds to the message request and sends data to the computing device. In this embodiment, the PCIe bus first accesses GMEM and then reads and writes SMEM through GMEM, so as to realize the transmission of the shared memory address.

[0071] On the one hand, using the PCIe protocol to obtain the shared memory address helps to ensure the compatibility of communication content among different computing devices compared to using other protocols to obtain the shared memory address. That is to say, it helps the computing device and the target device to accurately identify the communication content. For example, the communication content is a message request indicating to obtain the shared memory address or a response message carrying the shared memory address. This not only helps the communication reliability but also ensures the accuracy of the obtained address.

[0072] On the other hand, using the message request mechanism of the PCIe protocol to obtain the shared memory address, since the message request can customize the request content, compared to using other requests of the PCIe protocol to obtain the shared memory address, it not only makes full use of the scalability of the message request mechanism but also ensures that the communication content follows the regulations of the PCIe protocol and the feasibility of the address acquisition operation.

[0073] In some embodiments, the computing device also obtains the global memory address of the target device, and this global memory address refers to the unique address of the GMEM in the target device. That is to say, GMEM and SMEM use a unified address space, and based on this address space, the corresponding global memory or shared memory can be accessed. This design of unified addressing eliminates the need for address conversion between GMEM and SMEM, thus simplifying the data transmission process.

[0074] In some embodiments, both the shared memory and the global memory in the target device adopt the unified addressing technology. Figure 4 The flowchart of a unified addressing technology provided by an embodiment of the present application is shown. As Figure 4 shown, this unified addressing technology includes step S401 to step S403.

[0075] Step S401: Construct an address space with a byte size of 2 64 bits.

[0076] This step uniformly addresses all devices in the cluster according to the 64-bit address length, and the total byte size of the address space is 2 64 bits.

[0077] Step S402: Allocate the upper 16-bit addresses in the address space to multiple nodes in the target cluster, where the address space of each node is 256TB.

[0078] Among them, the upper 16-bit addresses are used for addressing different nodes in the computing cluster, and a total of 65,536 nodes are supported.

[0079] In this step, the upper 16-bit address refers to the highest 16 bits of the address (i.e., from the 49th bit to the 64th bit). These high-order parts represent the part of the address space that is farthest from the processor and are used to allocate independent address spaces for different hardware devices or system resources.

[0080] Step S403: Allocate the address space of each node to the CPU and the computing device card corresponding to the node. The address space of the CPU is 128TB, and the address space of the computing device card is 128TB.

[0081] In this step, the address space of each node is 256TB, which can be evenly allocated to the CPU and the computing device card, 128TB each. Each node can support 64 computing device cards, and there are at most 65,536 nodes, so a total of 4,194,304 computing cards can be supported. The address space of the computing device card on each node is 128TB. The address space of each card is 2TB, of which 1TB is GMEM and 1TB is SMEM.

[0082] Thus, GMEM and SMEM are uniformly addressed, and the global memory address and the shared memory address are obtained respectively. A computing device (such as a CPU) can directly access GMEM based on the global memory address and directly access SMEM based on the shared memory address.

[0083] In the embodiments of the present application, GMEM and SMEM in the computing device can also be uniformly addressed by using the method provided in steps S401 to S403 above, so that the target device can quickly access SMEM in the computing device.

[0084] Figure 5 A flowchart of a method for obtaining a shared memory address provided by the embodiments of the present application is provided. As Figure 5 shown, in combination with the above unified addressing technology, obtaining the shared memory address provided by the target device through the message request of the PCIe protocol may further include steps S501 to S503.

[0085] Step S501: Send a message request to the target device; the message request is used to request to obtain the shared memory address of the target device.

[0086] Step S502: Receive response information; the response information indicates the address space of the shared memory of the target device, and the response information is sent by the target device in response to the message request.

[0087] Step S503: Obtain the shared memory address of the target device from the response information.

[0088] In the embodiments described in steps S501 to S503, after the target device receives a message request, if it supports P2P transmission, it will reply. For example, it will return response information based on the received message request. Conversely, if it does not support P2P transmission, it will not reply.

[0089] Exemplarily, the message request includes "PCIe Requester ID", which contains the identity (ID) of the computing device. Then the target device can return response information according to the identity, that is, return the shared memory address to the device to which the identity indicated by "PCIe Requester ID" belongs.

[0090] The target device can save the shared memory address in a specified field of the response information, so that the computing device can obtain the shared memory address of the target device from the specified field (such as Shared Memory Address) in the response information.

[0091] By addressing the SMEM of the target device through the above step S301 and exposing the addressing to the outside, the outside can directly obtain the shared memory address from the storage medium without first accessing the GMEM and then the SMEM, thereby performing step S302.

[0092] In step S302, the computing device initiates a data transfer operation to the shared memory address, thereby performing data transfer with the target device.

[0093] Among them, the data transfer operation includes data read and write operations, that is, writing data to the SMEM or reading data from the SMEM.

[0094] Since all the SMEMs in each computing device card are separately connected to the PCIe and each SMEM has a unique address, read and write operations can be directly performed through the PCIe.

[0095] In some of the embodiments, the target device verifies the shared memory address. Specifically, the target device includes an address register A, a register L, and a comparator. The register L indicates the n high-order significant bits of the address in the address register A. The comparator performs address validity verification. The address register A and the register L are read and written through the PCIe bus and are two programmable registers. Assume that the input shared memory address is b 63 b 62 b 61 …b 1 b 0 , the value of the address register A is a 63 a 62 a 61 …a 1 a 0 , and the value of the register L is n. If b63 b 62 b 61 …b 64-n equal to a 63 a 62 a 61 …a 64-n If so, the comparator output is true, indicating that the address is selected. Otherwise, the output is false, indicating that the address is not selected. If the verification result is true, the target device processes the data. If the verification result is false, the target device discards the data, thereby improving security. In some embodiments, if the device supports multi-segment separated addresses, multiple unified addressing units need to be added. Each unified addressing unit performs the above comparison step to simultaneously compare and verify the multi-segment separated addresses.

[0096] In the embodiments of the present application, considering the speed difference between the PCIe bus and the internal bus of the computing device card, read / write buffer queues are introduced in each SMEM to accelerate the asynchronous data transfer rate. The read / write buffer queues are implemented by hardware, can receive read / write requests from the PCIe, and can automatically complete the read / write of the SMEM. A locking mechanism is added to each SMEM to lock the corresponding memory area during the read / write of the SMEM, so as to ensure that the read / write of the buffer queue does not conflict with the read / write of the computing kernel.

[0097] Specifically, Figure 6 shows a flowchart of a method for reading data using a buffer provided by an embodiment of the present application. As Figure 6 shown, initiating a data transfer operation to a shared memory address includes steps S601 to S604.

[0098] Step S601: When the computing device initiates a read operation instruction to the target device through the PCIe bus, the target device locks the shared memory corresponding to the shared memory address.

[0099] Step S602: Based on the data start address and data length information in the read operation instruction, the target device reads the data in the shared memory into the data buffer.

[0100] Step S603: When all the data is read from the shared memory into the data buffer, the target device releases the shared memory.

[0101] Step S604: Based on the shared memory, the target device sends data to the PCIe bus.

[0102] In this embodiment, the read / write buffer queue includes a data buffer, a data start address, and a data length. When the PCIe bus initiates a read operation, after receiving the instruction, the read / write buffer queue saves the read data start address and data length information, locks the corresponding SMEM block, reads the data into the data buffer; after completion, releases the locked SMEM block, and sends the corresponding data to the PCIe bus.

[0103] Figure 7 The flowchart of a method for implementing write data using buffering provided by an embodiment of the present application is shown. As Figure 7 shown, initiating a data transfer operation to a shared memory address includes steps S701 to S703.

[0104] Step S701: When the computing device initiates a write operation instruction to the target device through the PCIe bus, the target device writes the data into the data buffer based on the data start address and data length information in the write operation instruction.

[0105] Step S702: In the case where all the data is written into the data buffer, the target device locks the shared memory corresponding to the shared memory address.

[0106] Step S703: After the data is written from the data buffer into the shared memory, the target device releases the shared memory and sends a signal to the PCIe bus, indicating that the write operation instruction is successfully executed.

[0107] In this embodiment, when the PCIe bus initiates a write operation, after receiving the instruction, the read-write buffer queue saves the write data start address and data length information, and the data is written into the data buffer; then the corresponding SMEM block is locked and the data is written. After completion, a signal indicating successful execution is sent to the PCIe bus.

[0108] Thus, by adding a locking mechanism to each SMEM, the corresponding shared memory area is locked during SMEM read and write to ensure that the read and write of the buffer queue do not conflict with the read and write of the computing kernel.

[0109] In the embodiment described in the above steps S301 to S302, the computing device, in response to the received data transfer request from the target device, obtains the shared memory address of the target device from the storage medium. The shared memory address refers to the unique address of the shared memory in the target device. Among them, the shared memory in the target device and the global memory both adopt a unified addressing technology, so that the shared memory is directly connected to the PCIe bus through the shared memory address. The computing device initiates a data transfer operation to the shared memory address through the PCIe bus, thereby performing data transfer with the target device. GMEM and SMEM use a unified address space, and the address of SMEM is exposed to the external PCIe bus. In this way, the computing device skips the GMEM transfer in the traditional method and can directly access the shared memory by the PCIe bus and transfer the data to SMEM at one time. This means that both the CPU and the RDMA controller can access these two types of memory, GMEM and SMEM, through a single address space. This design eliminates the need for transfer and replication operations between GMEM and SMEM in the traditional architecture, thereby simplifying the data transfer process.

[0110] Secondly, by directly connecting the SMEM to the PCIe bus, it allows data to be directly transferred from the source to the SMEM, thus skipping the intermediate step of GMEM. This direct connection significantly reduces the path length of data transfer and improves the efficiency of data transfer.

[0111] Since the data transfer rates of different computing devices may not be consistent, buffer queues are introduced in the bus and SMEM transfers. The buffer queue can temporarily store data until the receiving device is ready to receive the data. This mechanism effectively reduces the transfer waiting time caused by rate mismatches and ensures the continuity and efficiency of data transfer.

[0112] In addition, to ensure that this method can be compatible with traditional computing devices, the computing core can still directly read and write the SMEM and GMEM. This consideration of compatibility enables existing systems to utilize the advantages of this method without large-scale hardware upgrades.

[0113] Generally speaking, the method provided by the embodiments of this application has the advantages of few data copy times, fast transmission speed, low latency, and compatibility with existing systems. In data transfer for large model training, especially in RDMA transfers based on PCIe or high-speed networks, it is a very effective and feasible method.

[0114] The following further illustrates the method provided by the embodiments of this application with a specific example. Figure 8 It shows a system framework diagram of a computing device based on unified addressing provided by the embodiments of this application. As Figure 8 shown, each node includes a CPU, a computing device card, and an RDMA controller (network card). The computing device cards are respectively connected to the CPU and the RDMA controller through PCIe. In the computing device card, the PCIe interface realizes direct access to the global memory (GMEM) and the shared memory (SMEM) through a programmable unified address unit. Among them, each shared memory is connected to the PCIe interface through a buffer queue, and the buffer queue can temporarily store data until the receiving device (CPU, RDMA controller, or SMEM) is ready to receive the data. The shared memory can transfer data with the computing core.

[0115] Figure 9 It shows a structural block diagram of a data transfer device provided by the embodiments of this application. As Figure 9 shown, the device includes:

[0116] An address acquisition module 91, configured to, in response to a data transmission request received from a target device, acquire a shared memory address of the target device from a storage medium of a computing device, where the shared memory address refers to a unique address of a shared memory in the target device, and wherein the shared memory and the global memory in the target device both adopt a unified addressing technology, so that the shared memory address is used for the shared memory to be directly connected to a PCIe bus.

[0117] A data transmission module 92, configured to initiate a data transmission operation to the shared memory address through a PCIe bus.

[0118] In some embodiments, the address acquisition module 91 is further configured to acquire the shared memory address provided by the target device through a message request of a PCIe protocol; and write the shared memory address into the storage medium.

[0119] In some embodiments, the address acquisition module 91 is further configured to send a message request to the target device; the message request is used to request to acquire the shared memory address of the target device; receive a response message; the response message indicates an address space of the shared memory of the target device, and the response message is sent by the target device in response to the message request; and acquire the shared memory address of the target device from the response message.

[0120] In some embodiments, the address acquisition module 91 is further configured to address the shared memory and the global memory in the target device by using a unified addressing technology, and the unified addressing technology includes: constructing an address space with a byte size of 2 64 bits; allocating the upper 16-bit addresses in the address space to multiple nodes in a target cluster, where the address space of each node is 256TB; allocating the address space of each node to a computing device card corresponding to the node, the address space of the CPU is 128TB, and the address space of the computing device card is 128TB; wherein each node supports 64 computing device cards, each computing device card is allocated an address space of 2TB, and the shared memory and the global memory respectively occupy an address space of 1TB in the computing device card.

[0121] In some embodiments, the data transmission module 92 is further configured to enable the target device to verify the shared memory address through a comparator to obtain a verification result; if the verification result is true, the target device processes the data; if the verification result is false, the target device discards the data; wherein the verification result indicates whether the L-n bit address in the shared memory address is the same as the L-n bit address in an address register A.

[0122] In some of these embodiments, the data transmission module 92 is further configured to, when the computing device issues a read operation instruction to the target device via the PCIe bus, cause the target device to lock the shared memory corresponding to the shared memory address; the target device reads the data in the shared memory into the data buffer based on the data start address and data length information in the read operation instruction; in the case where all the data is read from the shared memory into the data buffer, the target device releases the shared memory; and the target device sends the data to the PCIe bus based on the shared memory.

[0123] In some of these embodiments, the data transmission module 92 is further configured to, when the computing device issues a write operation instruction to the target device via the PCIe bus, cause the target device to write the data into the data buffer based on the data start address and data length information in the write operation instruction; in the case where all the data is written into the data buffer, the target device locks the shared memory corresponding to the shared memory address; after the data is written from the data buffer into the shared memory, the target device releases the shared memory and sends a signal to the PCIe bus, where the signal indicates that the write operation instruction is successfully executed.

[0124] It should be noted that the above-mentioned respective modules can be functional modules or program modules, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned respective modules can be located in the same processor; or the above-mentioned respective modules can also be located in different processors in any combined form.

[0125] It should be noted that the specific examples in this embodiment can refer to the examples described in the above-mentioned embodiments and optional implementation manners, and will not be elaborated herein.

[0126] In addition, in combination with the method provided in the above-mentioned embodiments, a storage medium can also be provided in this embodiment to implement it. A computer program is stored on the storage medium; when the computer program is executed by a processor, any one of the data transmission methods in the above-mentioned embodiments is implemented.

[0127] The embodiments of the present application further provide a computer program product, which, when running on a computer, causes the computer to execute each function or step executed by the processor in the above-mentioned method embodiments.

[0128] It should be understood that the specific embodiments described here are only used to explain this application, rather than to limit it. According to the embodiments provided in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0129] Obviously, the accompanying drawings are only some examples or embodiments of the present application. For those of ordinary skill in the art, the present application can also be applied to other similar situations based on these drawings without creative work. Additionally, it can be understood that although the work done during this development process may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be regarded as insufficient disclosure of the present application.

[0130] The term "embodiment" in the present application means that the specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean independence or alternative to other embodiments that are mutually exclusive. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.

[0131] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be construed as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A data transmission method, characterized in that: Applied to a computing device, the computing device comprising a storage medium, the method comprising: In response to a data transmission request received from a target device, obtaining a shared memory address of the target device from the storage medium, wherein the shared memory address refers to a unique address of a shared memory in the target device, wherein the shared memory in the target device and the global memory both adopt a unified addressing technology so that the shared memory is directly connected to a PCIe bus through the shared memory address; When the computing device initiates a read operation instruction to the target device through the PCIe bus, the target device locks the shared memory corresponding to the shared memory address; The target device reads the data in the shared memory into a data buffer area based on the data starting address and data length information in the read operation instruction; In a case where all the data are read from the shared memory to the data buffer area, the target device releases the shared memory; The target device sends the data to the PCIe bus based on the shared memory.

2. The data transmission method according to claim 1, characterized in that: Before obtaining the shared memory address of the target device from the storage medium, the method further includes: Obtain the shared memory address provided by the target device through a message request of the PCIe protocol; The shared memory address is written into the storage medium.

3. The data transmission method according to claim 2, characterized in that: The computing device is a CPU or an RDMA controller, the target device is a computing device card, and the unified addressing technology includes: Build 2 64 Byte-sized address space; Allocating high 16 bits of the address in the address space to multiple nodes in the target cluster, wherein the address space of each node is 256 TB; Allocating the address space of each node to the CPU and computing device card corresponding to the node, the address space of the CPU being 128 TB, and the address space of the computing device card being 128 TB; Each of the nodes supports 64 computing device cards, each of the computing device cards is allocated with 2 TB of address space, and the shared memory and the global memory respectively occupy 1 TB of address space in the computing device card.

4. The data transmission method according to claim 3, characterized in that: The acquiring of the shared memory address provided by the target device through a message request of the PCIe protocol includes: Sending a message request to the target device; the message request is used to request to obtain the shared memory address of the target device; receiving response information; the response information indicates the address space of the shared memory of the target device, and the response information is sent by the target device in response to the message request; The shared memory address of the target device is obtained from the response information.

5. The data transmission method according to claim 3, characterized in that: The target device includes an address register A, a register L, and a comparator, wherein the register L indicates n high-significant bits of the address in the address register A. After initiating a data transmission operation to the shared memory address, the method further includes: The target device verifies the shared memory address through the comparator to obtain a verification result; If the verification result is true, the target device processes the data; if the verification result is false, the target device discards the data; The verification result indicates whether the 63 to (64-n) bit address of the shared memory address is the same as the 63 to (64-n) bit address in the address register A.

6. The data transmission method according to any one of claims 1 to 5, characterized in that: The initiating a data transmission operation to the shared memory address includes: When the computing device initiates a write operation instruction to the target device through the PCIe bus, the target device writes the data into the data buffer area based on the data start address and data length information in the write operation instruction; When all the data are written into the data buffer area, the target device locks the shared memory corresponding to the shared memory address; After the data is written from the data buffer area into the shared memory, the target device releases the shared memory and sends a signal to the PCIe bus, where the signal indicates that the write operation instruction is executed successfully.

7. A data transmission device, characterized in that: The device comprises: An address acquisition module is used to obtain a shared memory address of the target device from a storage medium of the computing device in response to a data transmission request received from the target device, wherein the shared memory address refers to a unique address of the shared memory in the target device, wherein the shared memory in the target device and the global memory both adopt a unified addressing technology so that the shared memory address is used for the shared memory to be directly connected to the PCIe bus; A data transmission module, configured to cause the target device to lock the shared memory corresponding to the shared memory address when the computing device initiates a read operation instruction to the target device through the PCIe bus; The target device reads the data in the shared memory into a data buffer area based on the data starting address and data length information in the read operation instruction; In a case where all the data are read from the shared memory to the data buffer area, the target device releases the shared memory; The target device sends the data to the PCIe bus based on the shared memory.

8. A computing device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the data transmission method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the data transmission method according to any one of claims 1 to 6 when running.

Citation Information

Patent Citations

  • Data processing method and device based on GPU, electronic equipment and medium

    CN119440827A