A data processing method, device, equipment and medium

By allocating bus address space to slave devices in the master field programmable logic gate array device and using high-speed serial computer extended bus standard link for data transmission, the problems of cross-node GPU card and FPGA card communication delay and hardware resource overhead are solved, and efficient cross-node communication between slave devices is achieved.

CN116644010BActive Publication Date: 2025-06-24GUANGDONG INSPUR BIG DATA RES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310686129.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-09
Publication Date
2025-06-24
Estimated Expiration
2043-06-09

AI Technical Summary

Technical Problem

When communicating between cross-node GPU cards and traditional FPGA cards, it is usually necessary to rely on the participation of server CPU, memory, PCIe Chip set, etc., resulting in high communication delays during data transmission and large hardware resource overhead.

Method used

By allocating bus address space to slave devices in the master field programmable logic gate array device and using high-speed serial computer extended bus standard link for data transmission, cross-node communication between slave devices is realized and dependence on servers is reduced.

Benefits of technology

This method effectively reduces cross-node communication delay and hardware resource overhead, enables slave devices to run relatively independently, and reduces the degree of coupling to the server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116644010B_ABST
    Figure CN116644010B_ABST
Patent Text Reader

Abstract

The present invention discloses a data processing method, apparatus, device and medium, which are applied to the technical field of data processing, and include: based on the bus address space allocated to the first slave device, writing the data to be processed into the storage component of the first slave device through a first Peripheral Component Interconnect Express (PCIe) link, so that the first slave device processes the data to be processed to obtain result data; acquiring the result data sent by the first slave device; determining the destination address of the result data, and if the destination address points to the storage space of the second slave device, writing the result data into the storage component of the second slave device through a second PCIe link. In the prior art, there are problems of high communication delay during the data transmission process and large hardware resource overhead. The present invention can reduce the cross-node communication delay and hardware resource overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a data processing method, device, equipment and medium. Background Art

[0002] Currently, when communicating between cross-node GPU (i.e., Graphics Processing Unit) cards and traditional FPGA (i.e., Field Programmable Gate Array) cards, it is usually necessary to rely on the server CPU (i.e., Central Processing Unit), memory, and PCIe (i.e., peripheral component interconnect express) Chip set, etc. That is, it needs to be used in a tightly coupled manner with the server through the PCIe interface bus and it is difficult to run independently of the host CPU. In this way, there are problems of high communication latency in the data transmission process and large hardware resource overhead. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a data processing method, device, equipment and medium, which can reduce cross-node communication latency and hardware resource overhead. The specific solutions are as follows:

[0004] In a first aspect, the present invention discloses a data processing method, which is applied to a master field programmable gate array device and includes:

[0005] Based on the bus address space allocated for the first slave device, write the data to be processed into the storage component of the first slave device through a first peripheral component interconnect express link, so that the first slave device processes the data to be processed to obtain result data;

[0006] Obtain the result data sent by the first slave device;

[0007] Determine the destination address of the result data. If the destination address points to the storage space of the second slave device, write the result data into the storage component of the second slave device through a second peripheral component interconnect express link.

[0008] Optionally, after writing the data to be processed into the storage component of the first slave device based on the bus address space allocated for the first slave device and through a first peripheral component interconnect express link, it further includes:

[0009] Send an interrupt data packet to notify the first slave device to process the data to be processed.

[0010] Optionally, notifying the first slave device to process the data to be processed by the sending interruption data packet includes:

[0011] Writing the interruption data packet into a specified register of the first slave device based on a mapping relationship of a register in the first slave device on a main field programmable gate array device bus, so as to notify the first slave device to process the data to be processed.

[0012] Optionally, it further includes:

[0013] If the destination address points to a storage space of a local device, writing the result data into a local storage component.

[0014] Optionally, it further includes:

[0015] If the destination address points to a local network module, sending the result data to the network module and sending it to a network through the network module.

[0016] Optionally, it further includes:

[0017] If the destination address points to a local acceleration core module, sending the result data to the acceleration core module and invoking a preset processing logic through the acceleration core module to process the result data.

[0018] Optionally, it further includes:

[0019] Determining a data processing type based on a descriptor corresponding to the result data;

[0020] Correspondingly, processing the result data by invoking the preset processing logic through the acceleration core module includes: invoking the preset processing logic corresponding to the data processing type through the acceleration core module to process the result data.

[0021] Optionally, processing the result data by invoking the preset processing logic corresponding to the data processing type through the acceleration core module includes:

[0022] If the data processing type is encryption processing, invoking an encryption processing logic through the acceleration core module to encrypt the result data;

[0023] If the data processing type is compression processing, invoking a compression processing logic through the acceleration core module to compress the result data;

[0024] If the data processing type is first compression processing and then encryption processing, the acceleration core module is used to call the compression processing logic to perform compression processing on the result data to obtain compressed data, and the encryption processing logic is called to perform encryption processing on the compressed data.

[0025] Optionally, after notifying the first slave device by sending an interrupt data packet, it further includes:

[0026] When the notification sent by the first slave device is obtained, the result data is read from the storage component of the first slave device;

[0027] Among them, the result data is stored in the storage component of the first slave device by the first slave device.

[0028] Optionally, after reading the result data from the storage component of the first slave device, it further includes:

[0029] The result data is cached to the storage component of the local device.

[0030] Optionally, it further includes:

[0031] The acceleration core module is used to compress the result data to obtain compressed data, and encrypt the compressed data to obtain encrypted data.

[0032] Optionally, after encrypting the compressed data to obtain encrypted data, it further includes:

[0033] The network module is used to encapsulate the encrypted data and send it to other nodes through the optical network.

[0034] Optionally, reading the result data from the storage component of the first slave device includes:

[0035] The direct data access engine is used to read the result data from the storage component of the first slave device.

[0036] Optionally, it further includes:

[0037] Discover the in-place slave devices and allocate bus address spaces for each in-place slave device to complete the registration of in-place devices;

[0038] Allocate bus address spaces for each local module.

[0039] Optionally, the first slave device is a graphics processing unit acceleration card or a field programmable gate array acceleration card, and the second slave device is a graphics processing unit acceleration card or a field programmable gate array acceleration card.

[0040] Optionally, writing the data to be processed into the storage component of the first slave device based on the bus address space allocated to the first slave device and via a first Peripheral Component Interconnect Express (PCIe) link includes:

[0041] Obtain the data to be processed from a network. If the data to be processed is encrypted data, decrypt it to obtain decrypted data;

[0042] Based on the bus address space allocated to the first slave device and via a first Peripheral Component Interconnect Express (PCIe) link, write the decrypted data into the storage component of the first slave device.

[0043] Optionally, writing the data to be processed into the storage component of the first slave device based on the bus address space allocated to the first slave device and via a first Peripheral Component Interconnect Express (PCIe) link includes:

[0044] Obtain the data to be processed from a network. If the data to be processed is compressed data, decompress it to obtain decompressed data;

[0045] Based on the bus address space allocated to the first slave device and via a first Peripheral Component Interconnect Express (PCIe) link, write the decompressed data into the storage component of the first slave device.

[0046] In a second aspect, the present invention discloses a data processing apparatus, which is applied to a main Field Programmable Gate Array (FPGA) device and includes:

[0047] A data-to-be-processed writing module, configured to write the data to be processed into the storage component of the first slave device based on the bus address space allocated to the first slave device and via a first Peripheral Component Interconnect Express (PCIe) link, so that the first slave device processes the data to be processed to obtain result data;

[0048] A result data obtaining module, configured to obtain the result data sent by the first slave device;

[0049] A result data forwarding module, configured to determine the destination address of the result data. If the destination address points to the storage space of a second slave device, write the result data into the storage component of the second slave device via a second Peripheral Component Interconnect Express (PCIe) link.

[0050] In a third aspect, the present invention discloses an electronic device, including a memory and a processor, where:

[0051] The memory is used to store a computer program;

[0052] The processor is configured to execute the computer program to implement the foregoing data processing method.

[0053] Fourthly, the present invention discloses a computer-readable storage medium for storing a computer program, wherein when the computer program is executed by a processor, the foregoing data processing method is implemented.

[0054] It can be seen that the present invention is applied to a main field programmable gate array device. Based on the bus address space allocated to the first slave device, the data to be processed is written into the storage component of the first slave device through a first high-speed serial computer expansion bus standard link, so that the first slave device processes the data to be processed to obtain result data, and then the result data sent by the first slave device is obtained, and the destination address of the result data is determined. If the destination address points to the storage space of the second slave device, the result data is written into the storage component of the second slave device through a second high-speed serial computer expansion bus standard link. That is to say, in the main field programmable gate array device in the embodiment of the present invention, a bus address space is allocated to the slave device. Based on the bus address space allocated to the first slave device, the data to be processed is written into the storage component of the first slave device through a first high-speed serial computer expansion bus standard link, and the result data obtained by processing the data to be processed sent by the first slave device is obtained. When the destination address of the result data points to the storage space of the second slave device, the result data is written into the storage component of the second slave device through a second high-speed serial computer expansion bus standard link. In this way, through the main field programmable gate array device, a bus address space is allocated to the slave device to realize cross-node communication of the slave device, and the slave device can be a graphics processing unit acceleration card or a field programmable gate array acceleration card.

[0055] The beneficial effect of the present application is that it reduces the coupling between the slave device and the server, and can reduce the cross-node communication delay and the hardware resource overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0057] Figure 1 FIG. is a schematic diagram of GPUDirect P2P data interaction provided by an embodiment of the present invention;

[0058] Figure 2 FIG. is a schematic diagram of GPUDirect RDMA data interaction provided by an embodiment of the present invention;

[0059] Figure 3 FIG. is a flowchart of a data processing method provided by an embodiment of the present invention;

[0060] Figure 4 A schematic diagram of a specific slave device provided by an embodiment of the present invention;

[0061] Figure 5 A schematic diagram of a specific master device provided by an embodiment of the present invention;

[0062] Figure 6 A schematic diagram of the mapping relationship of the hardware storage resources of a slave device GPU card on the 64-bit bus of the master device provided by an embodiment of the present invention;

[0063] Figure 7 A schematic diagram of the mapping relationship of the internal storage resources of a master FPGA device on the bus provided by an embodiment of the present invention;

[0064] Figure 8 A schematic diagram of the mapping relationship of the storage resources of multiple slave devices on the FPGA bus of the master device provided by an embodiment of the present invention;

[0065] Figure 9 A flowchart of device registration provided by an embodiment of the present invention;

[0066] Figure 10 A schematic diagram of the structure of a data processing device provided by an embodiment of the present invention;

[0067] Figure 11 A structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0069] GPUDirect Shared Memory (Direct Shared Memory) enables the GPU and third-party PCIe devices to achieve shared memory through the shared host memory. GPUDirect P2P (GPUDirect Peer-to-Peer) allows two GPU devices under the same PCIe Root Complex domain to directly access each other's GPU video memory. There is no need to copy data between the GPUs through the host memory as an intermediate step. Compared with the GPUDirect Shared Memory solution, this reduces the steps of copying data from the GPU video memory to the host memory and from the host memory to the GPU video memory, reduces the data path latency, and improves the data transfer efficiency. See Figure 1 as shown in Figure 1 Figure [0000154] shows a schematic diagram of GPUDirect P2P data interaction provided by an embodiment of the present invention. However, this solution realizes the interaction of GPU video memory data within a single server node. The solution depends on the participation of the host CPU or within the same PCIe domain to achieve the interaction of GPU video memory data. Among them, the video memory can be GDDR (Graphics Double Data Rate, a high-performance video memory). GPUDirect (Direct) RDMA (Remote Direct Memory Access) is a technology that uses RDMA-related technologies, physical network cards, and transmission networks to directly interact with GPU video memory data between different nodes. This technology solves the problems in the traditional network (such as TCP (Transmission Control Protocol) / IP (Internet Protocol) data transmission process, where data transmission needs to go through links such as the CPU and host cache, resulting in large latency and high CPU occupancy rate), and realizes the function of directly accessing each other's GPU video memory between different nodes through network cards. See Figure 2 as shown in Figure 2 Figure [0000156] shows a schematic diagram of GPUDirect RDMA data interaction provided by an embodiment of the present invention.

[0070] GPUDirect P2P technology is developed and implemented based on GPUs as PCIe devices under the HOST. Communication with other devices relies on the participation of the Host CPU, memory, and PCIe Switch (i.e., switch) system, etc. Devices are tightly coupled to the server CPU, memory, etc. through PCIe, and communication is limited to GPU cards within a single node. When using GPUDirect P2P technology for video memory data exchange between GPUs, it is only limited to direct video memory data interaction between GPUs under the same PCIe domain through the PCIe Chipset. If it crosses the PCIe domains of two CPUs, the CPU and CPU memory are still required to participate in data transmission. When crossing CPUs, the latency of video memory data interaction between GPUs in this solution is still very large, and the CPU overhead is still high. Although GPUDirect RDMA technology uses RDMA technology to solve the communication problem between GPU cards across nodes, it requires high-performance network cards and local server CPUs under the same PCIe domain to run relevant network protocol software to help the GPU complete cross-node data transmission. The GPU and the server are still in a tightly coupled relationship connected through PCIe. The GPU cannot run independently of the server. Cross-node communication between GPUs can only be carried out by connecting the network card to the switch, and the communication network topology is not flexible enough, with low packet forwarding efficiency and large communication latency. When deploying a distributed computing system platform based on these two technologies, a huge number of hardware devices such as host CPUs, memory, and network cards are required, occupying a large amount of rack space in the data center, and the costs of deployment and maintenance are very high.

[0071] Currently, when communicating between cross-node GPU cards and traditional FGPA cards, it usually relies on the participation of the server CPU, memory, and PCIe Chip set, etc. That is, it needs to be tightly coupled to the server through the PCIe interface bus for use and is difficult to run independently of the host CPU. In this way, there are problems of relatively high communication latency and large hardware resource overhead during the data transmission process. For this reason, the present invention provides a method that can reduce cross-node communication latency and hardware resource overhead.

[0072] See Figure 3 As shown, an embodiment of the present invention discloses a data processing method applied to a main field programmable gate array device, including:

[0073] Step S11: Based on the bus address space allocated to the first slave device, write the data to be processed into the storage component of the first slave device through the first Peripheral Component Interconnect Express (PCIe) link, so that the first slave device processes the data to be processed to obtain result data.

[0074] In one embodiment, data to be processed can be obtained from a network. If the data to be processed is encrypted data, it is decrypted to obtain decrypted data. Based on the bus address space allocated to the first slave device, the decrypted data is written into the storage component of the first slave device through a first Peripheral Component Interconnect Express (PCIe) link. In another embodiment, data to be processed can be obtained from a network. If the data to be processed is compressed data, it is decompressed to obtain decompressed data. Based on the bus address space allocated to the first slave device, the decompressed data is written into the storage component of the first slave device through a first PCIe link. In one embodiment, the data to be processed can be obtained and parsed through a network module, specifically a RoCE (RDMA over Converged Ethernet) module. The data to be processed is decrypted and decompressed by an acceleration core module to obtain decompressed data, and then, based on the bus address space allocated to the first slave device, the decompressed data is written into the storage component of the first slave device through a first PCIe link. Among them, the storage component can be GDDR. The data to be processed can be data sent by a remote management and control platform, or data sent by other master devices, or data sent by other slave devices with network sending functions.

[0075] It can be understood that the main field programmable gate array device is the master device, and the master device is a field programmable gate array device. Embodiments of the present invention can discover in-place slave devices and allocate bus address spaces to each in-place slave device to complete in-place device registration. Moreover, the master device allocates bus address spaces to each local module. The slave device allocates bus address spaces to each module of the slave device itself based on the bus address space allocated by the master device. Specifically, different bus address spaces and base addresses of registers and memory can be allocated to each module. The mapping relationship of the storage resources in the slave device on the bus is as follows: starting from the base address obtained from the master device bus, they are the configuration space register, BAR register, and GDDR video memory resources in sequence. The mapping relationship of the internal storage resources of the master device on the bus is as follows: starting from the base address, they are the configuration space register, BAR register space, and GDDR video memory resources in sequence. The registers and GDDR video memory resources correspond to the modules.

[0076] In the embodiments of the present application, the slave devices can all use a general PCIe acceleration card, such as a graphics processing unit (GPU) acceleration card or a field-programmable gate array (FPGA) acceleration card. That is, the first slave device is a GPU acceleration card or an FPGA acceleration card. An FPGA is a highly programmable heterogeneous multi-core chip with rich hardware resources inside, such as lookup tables, registers, DSP (Digital Signal Processing) cores, AI (Artificial Intelligence) cores, PCIe HIP (Hard Intellectual Property), high-speed Serdes (Serializer / Deserializer) and bus interconnection resources, etc. Users can use these resources to implement various data processing engines, complex bus protocols, and network protocols, etc. A GPU is a dedicated graphics processing chip that was originally used for graphics and image processing and is now widely used in the field of AI artificial intelligence computing. It is also an important computing chip. Currently, as a PCIe device, a GPU is plugged into the slot of a data center server and usually communicates with the host and other nodes through PCIe. Although there is a flexible and high-speed NVLink (a bus and its communication protocol) communication technology between GPUs, this technology is limited to GPU communication within a single node, and a GPU cannot be directly connected to the external data center network. RoCE is a reliable transport protocol that uses a converged Ethernet network to achieve direct access to local and remote memory. It can reduce the overhead on the HOST host CPU during the data transfer process, and the CPU only needs to be responsible for the management of the control plane.

[0077] In a specific implementation, the slave device is directly connected to the PCIe slot of the master device through a PCIe link and an SMBUS (System Management Bus) bus. The main FPGA device internally has multiple PCIe HIP hard cores in the PCIe root complex mode, an out-of-band supervision module, a link initialization module, an address allocation and conversion module, a GDDR controller (video memory controller), a DMA (Direct Memory Access) engine, a RoCE protocol stack (network module), a MAC controller (Media Access Controller), and an acceleration core. Inside the PCIe-based GPU slave device chip, there are usually a computing core, a GDDR controller, a PCIe HIP, a BAR (base address register), and a Copy (data transfer) engine, etc. There is also an out-of-band management I2C (Inter-Integrated Circuit, a two-way two-wire synchronous serial bus) channel on the board. For example, see Figure 4 as shown in Figure 4A schematic diagram of a specific slave device provided by an embodiment of the present invention. Refer to Figure 5 as shown Figure 5 A schematic diagram of a specific master device provided by an embodiment of the present invention.

[0078] Among them, the master device of the present invention specifically includes the following modules: The out-of-band supervision module is connected to the slave device through the I2C bus, reads and writes the slave device registers, and obtains the status data of the slave device to monitor and manage the slave device; Two PCIe HIPs work in the PCIe root complex mode and dock with the slave device PCIe HIP. It should be noted that in the figure, PCIe HIP0 (i.e., hard IP 0) and PCIe HIP1 (i.e., hard IP 1) are taken as examples, and the PCIe HIP can be extended according to the actual number of slave devices in place; The link initialization module converts the read and write commands of the configuration space register into standard PCIe TLP (i.e., Transaction Layer Packet) packets to complete the enumeration and configuration space register configuration of the slave device PCIe; The bus address allocation and conversion module connects the control status register interfaces of each module and the GDDR controller and the DMA engine, and allocates different address spaces and base addresses of registers and memory to each module; The network protocol module supporting RoCEv2 has the function of remotely directly transferring memory data with heterogeneous devices in other nodes, and ensures high-bandwidth and reliable network communication between the FPGA card and the remote management control platform and other FPGA cards; The MAC controller module is mainly responsible for receiving and sending network data packets; The GDDR controller is a controller for high-performance video memory, and the video memory it controls has a higher working frequency, less heat generation and smaller volume; The DMA engine receives and processes the load data in the TLP packet from the slave device, caches it to the master device GDDR or sends it to the acceleration core for processing or sends it to other node devices through the RoCE network. Similarly, the data from the master device GDDR, the acceleration core or the RoCE network can also be sent to the slave device through the DMA engine and the PCIe HIP; The acceleration core can be an algorithm core implemented by using some FPGA logic resources, such as encryption and decryption, compression and decompression, etc., and can process the data from the slave device DMA, the RoCE network and the local GDDR. Among them, 1 to 16 represent the transmission paths. For example, 1 is the transmission path between the out-of-band supervision module and the link initialization module.

[0079] Refer to Figure 6 as shown Figure 6 A schematic diagram of the mapping relationship of the hardware storage resources of a slave device GPU card on the 64-bit bus of the master device provided by an embodiment of the present invention. Among them, starting from the base address 0x1000_0000_0000_0000 obtained from the master device bus, they are the configuration space register, the BAR register and the GDDR video memory resources in sequence. Refer toFigure 7 As shown Figure 7 This is a schematic diagram of the mapping relationship of the internal storage resources of the main FPGA device on the bus provided by the embodiment of the present invention. Starting from the base address 0x0, they are the configuration space register, the BAR register space, and the GDDR video memory resources in sequence. See Figure 8 As shown Figure 8 This is a schematic diagram of the mapping relationship of the storage resources of multiple slave devices on the main device FPGA bus. From low to high, they are the resource mapping of the main device locally and the resource mapping of each slave device in sequence.

[0080] Further, see Figure 9 , Figure 9 This is a device registration flowchart provided by the embodiment of the present invention. Taking the slave device GPU card in the node as an example, when the system is powered on and initialized, the slave device registration and the initialization of the slave device PCIe configuration space and the storage resource mapping process are as follows: power on and initialize the main device, triggered by the main device, power on and initialize the slave device. The slave device can be a GPU card or an FPGA card. Determine whether the link with the in-place slave device is successful, that is, whether a physical layer connection is established with the in-place slave device. If not, perform the failure count. If so, perform the configuration space scan and register read and write, and the required memory size of the BAR register can be obtained. Then register the bus address space, and then broadcast the bus address space newly registered by the slave device, that is, the allocated bus address space, write it into the specified space of the mounted GDDR, and update the bus address space for the in-place slave device and the remote platform.

[0081] Further, the embodiment of the present invention can send an interrupt data packet to notify the first slave device to process the to-be-processed data. In a specific implementation manner, the interrupt data packet can be written into the specified register of the first slave device based on the mapping relationship of the register in the first slave device on the main field programmable gate array device bus to notify the first slave device to process the to-be-processed data.

[0082] Step S12: Obtain the result data sent by the first slave device.

[0083] Step S13: Determine the destination address of the result data. If the destination address points to the storage space of the second slave device, write the result data into the storage component of the second slave device through the second high-speed serial computer extension bus standard link. Wherein, the second slave device is a graphics processor acceleration card or a field programmable gate array acceleration card.

[0084] Further, if the destination address points to the storage space of the local device, write the result data into the local storage component.

[0085] In addition, if the destination address points to a local network module, the result data is sent to the network module and then sent to the network through the network module.

[0086] Furthermore, if the destination address points to a local acceleration core module, the result data is sent to the acceleration core module, and the preset processing logic is called by the acceleration core module to process the result data.

[0087] Moreover, in a specific implementation manner, the data processing type can be determined based on the descriptor corresponding to the result data; correspondingly, the step of calling the preset processing logic by the acceleration core module to process the result data includes: calling the preset processing logic corresponding to the data processing type by the acceleration core module to process the result data.

[0088] Wherein, if the data processing type is encryption processing, the encryption processing logic is called by the acceleration core module to encrypt the result data; if the data processing type is compression processing, the compression processing logic is called by the acceleration core module to compress the result data; if the data processing type is first compression processing and then encryption processing, the compression processing logic is called by the acceleration core module to compress the result data to obtain compressed data, and the encryption processing logic is called to encrypt the compressed data.

[0089] That is to say, in the implementation of the present invention, the descriptor describes the next operation for the result data. Based on the descriptor and the destination address, the result data can be forwarded and the next operation can be performed.

[0090] In one implementation manner, after sending an interrupt data packet to notify the first slave device, the embodiment of the present invention can also, when receiving the notification sent by the first slave device, read the result data from the storage component of the first slave device; wherein, the result data is stored in the storage component of the first slave device by the first slave device. That is to say, after the first slave device processes the data to be processed to obtain the result data, the result data is stored in the storage component of the first slave device by the first slave device. Among them, the direct data access engine is used to read the result data from the storage component of the first slave device.

[0091] And, after reading the result data from the storage component of the first slave device, the result data can be cached to the storage component of the local device. Furthermore, the acceleration core module is used to compress the result data to obtain compressed data, and the compressed data is encrypted to obtain encrypted data. Then the network module is used to encapsulate the encrypted data and send it to other nodes through the optical network.

[0092] That is, in a specific embodiment, after obtaining the result data, the slave device can notify the master device to read from the slave device, or send the result data to the master device.

[0093] It can be understood that in the embodiments of the present invention, the master device initializes and registers the PCIe slave device to complete the address mapping of the storage resources. Moreover, an out-of-band supervision module formed through the I2C channel can be used inside the master device to monitor and manage the slave device and the master device. It is converted from the internal bus interface to a standard format TLP packet. In addition, both the slave device and the master device are equipped with DMA controllers. When one party performs a DMA operation, it does not affect the normal DMA operation of the other party; when the master device FPGA communicates with other node devices through the RoCE network, encryption, decryption, compression, and decompression operations are performed locally to ensure data communication security, reduce communication volume, and reduce communication latency; the source or destination of the data processed by the master device DMA engine can be the slave device cache, local cache, or from or sent to the RoCE module and the acceleration core module for further processing, and different descriptors can be used for marking and distinction. Further, combined with Figure 5 , the master device provided by the present invention, wherein the DMA engine, the acceleration core, and the RoCE module support multiple data stream processing combinations. The following further lists several achievable data processing flows:

[0094] 1. After the data to be processed is parsed from the RoCE (network module), it first passes through the acceleration core for decryption and is stored in the GDDR of the master device through path 12. Then it is read from the GDDR and decompressed by the acceleration core. After decompression, it is sent to PCIe HIP0 through the DMA engine and based on paths 11 and 9, and written into the GDDR of the corresponding slave device 0 (GPU card) of PCIe HIP0. An interrupt TLP packet is sent to notify the GPU card to process. Specifically, the DMA engine sends an in-TLP packet, which is packetized by the bus address allocation and conversion module and the link initialization module group in sequence to obtain a write packet for the specified register, and is written into the specified register of the GPU card, and passes through Figure 5 paths 5, 4, and 2 in

[0095] 2. After the slave device 0 GPU card processes the data, it writes to the Doorbell register to notify the DMA engine in the master FPGA device to read the result in its GDDR (passing through Figure 5 paths 2, 4, and 5 in

[0096] 3. The DMA engine can perform the next operation according to the descriptors sent from device 0 (GPU card) and the destination address of the data packet. If the address points to the storage space of device 1, the relevant information in the TLP header is modified and forwarded to device 1 through the DMA engine. Specifically, data is obtained through PCIE HIP0 and forwarded to device 1 through PCIE HIP0 (through paths 9 and 10); if the address points to the storage space of the local device, it is written to the local storage GDDR through the DMA engine (through paths 9 and 11); if the address points to RoCE or the acceleration core, it is sent to the RoCE module through the DMA engine for packet encapsulation and then sent to the network (through paths 9, 13, and 16), or sent to the acceleration core for further processing (through paths 9 and 15).

[0097] It should be noted that in the embodiments of the present invention, all devices can use FPGA acceleration cards, or some can use GPU acceleration cards and some can use FPGA acceleration cards to form a heterogeneous acceleration system; when using a GPU card or an FPGA card, the GPU card or FPGA card is independent of the traditional Host motherboard and communicates directly with the main device FPGA through the PCIe link as an endpoint (slave device), completing the enumeration, configuration, and initialization of the GPU card or FPGA card as a slave device, which can reduce the coupling degree between the GPU card or FPGA card and the server; it does not depend on the HOST's CPU, memory, and PCIe-related chipset, reducing the indirect overhead cost of these resources when users use GPU or FPGA resources; the device-to-device communication in the present invention uses a mature and general PCIe link and direct communication based on the protocol supported by the PCIe link, with high flexibility and lower communication latency; a single main device FPGA in the present invention can support connection with multiple GPU cards or FPGA cards through PCIe, and all the storage resources (in-band BAR registers, in-band caches) on each slave device are mapped on the internal interconnection bus of the main device, enabling the main device to schedule and use the resources of the slave device; the main device FPGA in the present invention can communicate with the slave device through the out-of-band supervision module through the out-of-band I2C channel to monitor and manage the status of the slave device; before the data is sent to the GPU card and after it is output from the GPU card, it can be processed twice (such as encryption / decryption, compression / decompression) in the FPGA card, and then sent to the next-level processing module or sent out through the optical port, improving data transmission security, reducing communication volume, and communication latency; the DMA engine module in the present invention can forward the TLP packet to other slave devices according to the PCIe TLP packet address, or store it in the main device GDDR, or encapsulate it through the RoCE protocol and forward it to other node devices through the optical network.

[0098] In this way, multiple PCIe interfaces in the master device (which can be an FPGA acceleration card) are used to connect slave devices, reducing the coupling degree between the PCIe-based slave devices and the server motherboard. The initialization, registration, and address mapping processes of the slave devices are flexibly completed. The DMA engine is used to connect multiple modules in the master device and the slave devices for flexible data processing, effectively reducing the communication latency between devices, greatly improving the ability to expand communication with adjacent nodes, significantly reducing the deployment cost of heterogeneous systems using PCIe devices, and providing a new computing acceleration platform for distributed applications. It can solve the problems of large communication latency between cross-node GPUs during the AI training process based on GPUs in modern data centers, dependence on the server CPU, memory, and PCIe Chipset, etc., and large hardware resource overhead during the data transmission process.

[0099] See Figure 10 As shown, an embodiment of the present invention discloses a data processing device, which is applied to a master field programmable gate array device and includes:

[0100] A to-be-processed data writing module 11, configured to write to-be-processed data into a storage component of the first slave device based on a bus address space allocated to the first slave device and through a first Peripheral Component Interconnect Express (PCIe) link, so that the first slave device processes the to-be-processed data to obtain result data;

[0101] A result data acquisition module 12, configured to acquire the result data sent by the first slave device;

[0102] A result data forwarding module 13, configured to determine a destination address of the result data. If the destination address points to a storage space of a second slave device, the result data is written into the storage component of the second slave device through a second PCIe link.

[0103] It can be seen that the embodiment of the present invention is applied to the main field programmable gate array device. Based on the bus address space allocated to the first slave device, the data to be processed is written into the storage component of the first slave device through the first high-speed serial computer expansion bus standard link, so that the first slave device processes the data to be processed to obtain result data, and then the result data sent by the first slave device is obtained, and the destination address of the result data is determined. If the destination address points to the storage space of the second slave device, the result data is written into the storage component of the second slave device through the second high-speed serial computer expansion bus standard link. That is, the main field programmable gate array device in the embodiment of the present invention allocates a bus address space for the slave device, based on the bus address space allocated to the first slave device, and writes the data to be processed into the storage component of the first slave device through the first high-speed serial computer expansion bus standard link, and obtains the result data obtained by processing the data to be processed sent by the first slave device. When the destination address of the result data points to the storage space of the second slave device, the result data is written into the storage component of the second slave device through the second high-speed serial computer expansion bus standard link. In this way, through the main field programmable gate array device, a bus address space is allocated for the slave device, cross-node communication of the slave device is realized, the slave device can be a graphics processor acceleration card or a field programmable gate array acceleration card, the coupling between the slave device and the server is reduced, and the cross-node communication delay and hardware resource overhead can be reduced.

[0104] The device further includes an interrupt notification module, configured to send an interrupt data packet to notify the first slave device to process the data to be processed.

[0105] Among them, the interrupt notification module is specifically configured to write the interrupt data packet into a specified register of the first slave device based on the mapping relationship of the register in the first slave device on the main field programmable gate array device bus, so as to notify the first slave device to process the data to be processed.

[0106] Among them, the result data forwarding module 13 is further configured to, if the destination address points to the storage space of the local device, write the result data into the local storage component. If the destination address points to the local network module, send the result data to the network module and send it to the network through the network module. If the destination address points to the local acceleration core module, send the result data to the acceleration core module and call a preset processing logic to process the result data through the acceleration core module.

[0107] Moreover, the result data forwarding module 13 is further configured to determine the data processing type based on the descriptor corresponding to the result data; and call the preset processing logic corresponding to the data processing type through the acceleration core module to process the result data. Specifically, if the data processing type is encryption processing, the encryption processing logic is called through the acceleration core module to encrypt the result data; if the data processing type is compression processing, the compression processing logic is called through the acceleration core module to compress the result data; if the data processing type is first compression processing and then encryption processing, the compression processing logic is called through the acceleration core module to compress the result data to obtain compressed data, and the encryption processing logic is called to encrypt the compressed data.

[0108] Furthermore, the apparatus further includes a data reading module, which is specifically configured to, when receiving the notification sent by the first slave device, read the result data from the storage component of the first slave device; wherein the result data is stored in the storage component of the first slave device by the first slave device.

[0109] In addition, the apparatus further includes a data caching module, which is configured to cache the result data to the storage component of the local device after reading the result data from the storage component of the first slave device. Furthermore, the apparatus further includes: an acceleration core module, which is configured to compress the result data to obtain compressed data, and encrypt the compressed data to obtain encrypted data; a network module, which is configured to encapsulate the encrypted data and send it to other nodes through an optical network.

[0110] Among them, the data reading module is specifically configured to read the result data from the storage component of the first slave device by using a direct data access engine.

[0111] In addition, the apparatus further includes a slave device registration and bus address space allocation module, which is configured to discover the in-place slave devices, allocate bus address spaces for each in-place slave device to complete the in-place device registration; and allocate bus address spaces for each local module.

[0112] The first slave device is a graphics processing unit acceleration card or a field programmable gate array acceleration card, and the second slave device is a graphics processing unit acceleration card or a field programmable gate array acceleration card.

[0113] In one implementation, the data to be processed writing module 11 is specifically configured to obtain the data to be processed from a network. If the data to be processed is encrypted data, it is decrypted to obtain decrypted data; based on the bus address space allocated for the first slave device, the decrypted data is written into the storage component of the first slave device through a first high-speed serial computer expansion bus standard link.

[0114] In another embodiment, the data to be processed writing module 11 is specifically configured to obtain the data to be processed from a network. If the data to be processed is compressed data, it is decompressed to obtain the decompressed data. Based on the bus address space allocated to the first slave device, the decompressed data is written into the storage component of the first slave device through a first Peripheral Component Interconnect Express (PCIe) link.

[0115] See Figure 11 As shown, an embodiment of the present invention discloses an electronic device 20, including a processor 21 and a memory 22. Among them, the memory 22 is used to store a computer program, and the processor 21 is used to execute the computer program, which is the data processing method disclosed in the foregoing embodiment.

[0116] For the specific process of the above data processing method, reference can be made to the corresponding content disclosed in the foregoing embodiment, and details will not be repeated here.

[0117] Moreover, as a carrier for resource storage, the memory 22 can be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc., and the storage method can be temporary storage or permanent storage.

[0118] In addition, the electronic device 20 further includes a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20. The communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of the present invention, and specific limitations are not imposed here. The input / output interface 25 is used to obtain external input data or output data to the outside, and the specific interface type can be selected according to specific application requirements, and specific limitations are not imposed here.

[0119] Furthermore, an embodiment of the present invention also discloses a computer-readable storage medium for storing a computer program. When the computer program is executed by a processor, the data processing method disclosed in the foregoing embodiment is implemented.

[0120] For the specific process of the above data processing method, reference can be made to the corresponding content disclosed in the foregoing embodiment, and details will not be repeated here.

[0121] In this specification, the embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0122] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in software modules executed by a processor, or in a combination of both. The software modules may be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium well-known in the art.

[0123] The above has introduced in detail a data processing method, apparatus, device and medium provided by the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A data processing method, characterized in that, Applied to the main field programmable gate array device, including: Based on the bus address space allocated to the first slave device, and writing the data to be processed into the storage component of the first slave device through the first high-speed serial computer expansion bus standard link, so that the first slave device processes the data to be processed to obtain result data; Obtain the result data sent by the first slave device; Determine the destination address of the result data. If the destination address points to the storage space of the second slave device, write the result data into the storage component of the second slave device through the second high-speed serial computer expansion bus standard link; if the destination address points to the storage space of the local device, write the result data into the local storage component; if the destination address points to the local network module, send the result data to the network module and send it to the network through the network module; if the destination address points to the local acceleration core module, send the result data to the acceleration core module and call the preset processing logic through the acceleration core module to process the result data; Also includes: Discover the in-place slave devices and allocate bus address spaces for each in-place slave device to complete the in-place device registration; and enable the slave device to allocate bus address spaces for each module of the slave device itself based on the allocated bus address space; Allocate bus address spaces for each local module; The first slave device is a graphics processing unit acceleration card or a field programmable gate array acceleration card, and the second slave device is a graphics processing unit acceleration card or a field programmable gate array acceleration card; A single main field programmable gate array device supports connection with multiple graphics processing unit acceleration cards or field programmable gate array cards through PCIe, and maps all the storage resources on each slave device to the internal interconnection bus of the main device.

2. The data processing method according to claim 1, wherein After writing the data to be processed into the storage component of the first slave device based on the bus address space allocated to the first slave device and through the first high-speed serial computer expansion bus standard link, it also includes: Send an interrupt data packet to notify the first slave device to process the data to be processed.

3. The data processing method according to claim 2, characterized in that The sending of the interrupt data packet to notify the first slave device to process the data to be processed includes: Writing the interrupt data packet into the specified register of the first slave device based on the mapping relationship of the register in the first slave device on the bus of the main field programmable gate array device to notify the first slave device to process the data to be processed.

4. The data processing method according to claim 1, wherein Also includes: Determine the data processing type based on the descriptor corresponding to the result data; Correspondingly, the processing of the result data by calling the preset processing logic through the acceleration core module includes: calling the preset processing logic corresponding to the data processing type through the acceleration core module to process the result data.

5. The data processing method according to claim 4, wherein The processing of the result data by calling the preset processing logic corresponding to the data processing type through the acceleration core module includes: If the data processing type is encryption processing, the encryption processing logic is called through the acceleration core module to encrypt the result data; If the data processing type is compression processing, the compression processing logic is called through the acceleration core module to compress the result data; If the data processing type is first compression processing and then encryption processing, the compression processing logic is called through the acceleration core module to compress the result data to obtain compressed data, and the encryption processing logic is called to encrypt the compressed data.

6. The data processing method according to claim 2, wherein After notifying the first slave device by sending an interrupt data packet, it further includes: When the notification sent by the first slave device is obtained, the result data is read from the storage component of the first slave device; Wherein, the result data is stored in the storage component of the first slave device by the first slave device.

7. The data processing method according to claim 6, characterized in that, After reading the result data from the storage component of the first slave device, it further includes: Caching the result data to the storage component of the local device.

8. The data processing method according to claim 7, wherein It further includes: Using the acceleration core module to compress the result data to obtain compressed data, and encrypting the compressed data to obtain encrypted data.

9. The data processing method according to claim 8, wherein After encrypting the compressed data to obtain encrypted data, it further includes: Using the network module to encapsulate the encrypted data and sending it to other nodes through the optical network.

10. The data processing method according to claim 6, wherein The reading the result data from the storage component of the first slave device includes: Using the direct data access engine to read the result data from the storage component of the first slave device.

11. The data processing method according to any one of claims 1 to 10, characterized in that, The writing the data to be processed into the storage component of the first slave device based on the bus address space allocated for the first slave device and through the first high-speed serial computer expansion bus standard link includes: Obtaining the data to be processed from the network, and if the data to be processed is encrypted data, decrypting it to obtain decrypted data; Based on the bus address space allocated for the first slave device and through the first high-speed serial computer expansion bus standard link, writing the decrypted data into the storage component of the first slave device.

12. The data processing method according to any one of claims 1 to 10, characterized in that, The writing the data to be processed into the storage component of the first slave device based on the bus address space allocated for the first slave device and through the first high-speed serial computer expansion bus standard link includes: Obtaining the data to be processed from the network, and if the data to be processed is compressed data, decompressing it to obtain decompressed data; Based on the bus address space allocated for the first slave device and through the first high-speed serial computer expansion bus standard link, writing the decompressed data into the storage component of the first slave device.

13. A data processing device, characterized in that, Applied to the main field programmable gate array device, it includes: A data-to-be-processed writing module, configured to write the data to be processed into the storage component of the first slave device based on the bus address space allocated for the first slave device and through the first high-speed serial computer expansion bus standard link, so that the first slave device processes the data to be processed to obtain result data; A result data acquisition module, configured to acquire the result data sent by the first slave device; A result data forwarding module is configured to determine a destination address of the result data. If the destination address points to the storage space of a second slave device, the result data is written into the storage component of the second slave device through a second Peripheral Component Interconnect Express (PCIe) link. The result data forwarding module is further configured to: if the destination address points to the storage space of a local device, write the result data into a local storage component; if the destination address points to a local network module, send the result data to the network module and then send it to the network through the network module; if the destination address points to a local acceleration core module, send the result data to the acceleration core module and call a preset processing logic through the acceleration core module to process the result data. It further includes: A slave device registration and bus address space allocation module, which is configured to discover the slave devices in place and allocate bus address spaces for each slave device in place to complete the registration of the devices in place; and enable each slave device to allocate bus address spaces for its own modules based on the allocated bus address spaces; allocate bus address spaces for local modules; the first slave device is a graphics processing unit (GPU) acceleration card or a field programmable gate array (FPGA) acceleration card, and the second slave device is a GPU acceleration card or an FPGA acceleration card; a single main FPGA device supports connection with multiple GPU acceleration cards or FPGA cards through PCIe, and maps all the storage resources on each slave device onto the internal interconnect bus of the main device.

14. An electronic device, characterized in that, It includes a memory and a processor, wherein: The memory is used to store a computer program. The processor is configured to execute the computer program to implement the data processing method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, It is used to store a computer program, wherein the computer program, when executed by the processor, implements the data processing method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Method and device for expanding PCIe bus region

    CN104285218A

  • Data processing method and device and medium

    CN114138481A