Data processing method and device for artificial intelligence scene and storage medium

By adopting RDMA and RoCEv2 protocols in private cloud intelligent computing scenarios, the data transmission path between the GPU server and the PFS system is optimized, solving the problems of high CPU consumption and latency caused by traditional Ethernet protocols, and achieving efficient and low-latency data processing.

CN121750404AActive Publication Date: 2026-03-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-27
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In private cloud computing scenarios, data communication between GPU servers and PFS systems relies on traditional Ethernet protocols, resulting in high CPU computing resource consumption and large data transmission latency, which becomes a bottleneck for system performance.

Method used

By employing the RDMA and RoCEv2 communication protocols, and optimizing the data transmission path through virtual switches and intelligent computing gateways, data can be transmitted directly between the GPU server and the PFS system memory, bypassing the operating system kernel. Combined with virtual routers and intelligent computing gateways for packet encapsulation and address mapping, efficient and low-latency data transmission can be achieved.

Benefits of technology

It significantly improves data processing efficiency, reduces CPU utilization and data transmission latency, and enhances the performance and cost-effectiveness of GPU computing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750404A_ABST
    Figure CN121750404A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device for an artificial intelligence scene and a readable storage medium, and relates to the technical field of computers, in particular to the technical field of artificial intelligence, cloud computing and model training and pushing. According to the specific implementation scheme, the method comprises the steps that a graphics processor server in a virtual private cloud network initiates a data processing request for the parallel file storage system through a remote direct memory access protocol; the virtual switch forwards the data processing request to a virtual router through a plurality of physical switches; the data processing request forwarded by the virtual switch carries a value for identifying a remote direct memory access protocol, so that each physical switch preferentially processes the data processing request based on a pre-configured priority strategy; the virtual router repackages the message of the data processing request and forwards the message to the intelligent computing gateway; and the intelligent computing gateway removes the header of the data processing request, performs address mapping, and forwards the mapped message of the data processing request to the parallel file storage system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, specifically to the fields of artificial intelligence, cloud computing, and model training and extension technology, and in particular to a data processing method, apparatus, and storage medium for artificial intelligence scenarios. Background Technology

[0002] In private cloud intelligent computing scenarios, artificial intelligence (AI) training tasks serve as the core driving force. Training and inference between graphics processing unit (GPU) clusters heavily rely on high-performance computing (HPN) networks to achieve efficient data transfer and state synchronization. Meanwhile, while Virtual Private Cloud (VPC) networks offer isolation and flexibility, their resource potential has not been fully realized. Therefore, exploring how to unlock the potential value of VPCs in intelligent computing scenarios and achieve performance complementarity with HPNs will become a key research direction for improving the overall resource utilization of private clouds and reducing the total cost of ownership (TCO). Summary of the Invention

[0003] This disclosure provides a data processing method, apparatus, and storage medium for artificial intelligence scenarios.

[0004] According to one aspect of this disclosure, a data processing method for artificial intelligence scenarios is provided, applicable in model training or model inference scenarios, including:

[0005] In a virtual private cloud network, a graphics processor server initiates a data processing request to a parallel file storage system via a remote direct memory access protocol.

[0006] The virtual switch forwards the data processing request to the virtual router through multiple physical switches; the data processing request forwarded by the virtual switch carries a value identifying the Remote Direct Memory Access Protocol, so that each physical switch can prioritize the data processing request based on a pre-configured priority policy.

[0007] The virtual router re-encapsulates the data processing request message and forwards it to the intelligent computing gateway.

[0008] The intelligent computing gateway removes the header of the data processing request, performs address mapping, and forwards the mapped data processing request message to the parallel file storage system.

[0009] According to another aspect of this disclosure, a data processing system for artificial intelligence scenarios is provided, applied in model training or model inference scenarios, including:

[0010] A graphics processing server is used to initiate data processing requests to a parallel file storage system via a remote direct memory access protocol; the graphics processing server is a graphics processing server in a virtual private cloud network.

[0011] A virtual switch is used to forward the data processing request to a virtual router through multiple physical switches; the data processing request forwarded by the virtual switch carries a value identifying the Remote Direct Memory Access Protocol, so that each physical switch can prioritize the data processing request based on a pre-configured priority policy.

[0012] The virtual router is used to re-encapsulate the data processing request message and forward it to the intelligent computing gateway.

[0013] The intelligent computing gateway is used to remove the header of the data processing request, perform address mapping, and forward the mapped data processing request message to the parallel file storage system.

[0014] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described above and any possible implementations.

[0018] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the methods described above and any possible implementation thereof.

[0019] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the aspects and any possible implementations described above.

[0020] According to the technology disclosed herein, a data processing scheme for a GPU server based on the RDMA protocol to a PFS system can be implemented, and data processing efficiency can be effectively improved.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0022] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0023] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure;

[0024] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure;

[0025] Figure 3 This is a schematic diagram of the message structure of a data processing request initiated by the GPU server in an embodiment of this disclosure;

[0026] Figure 4 This is a schematic diagram of the packet structure processed by the virtual switch in this embodiment of the disclosure;

[0027] Figure 5 This is a schematic diagram of the structure of a packet processed by a virtual router in an embodiment of this disclosure;

[0028] Figure 6 This is a schematic diagram of the message structure processed by the intelligent computing gateway in this embodiment of the present disclosure;

[0029] Figure 7 This is a schematic diagram of the structure of the data return packet generated by the PFS system in this embodiment of the disclosure;

[0030] Figure 8 This is a schematic diagram of the structure of the intelligent computing gateway processing data return packets in an embodiment of this disclosure;

[0031] Figure 9 This is a schematic diagram of the structure of the virtual router processing data return packets in an embodiment of this disclosure;

[0032] Figure 10 This is a schematic diagram according to the third embodiment of the present disclosure;

[0033] Figure 11 These are schematic diagrams based on the third and fourth embodiments of this disclosure;

[0034] Figure 12 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation

[0035] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0036] Obviously, the described embodiments are only some, not all, of the embodiments disclosed herein. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0037] It should be noted that the terminal devices involved in the embodiments of this disclosure may include, but are not limited to, smart devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, and tablet computers; the display devices may include, but are not limited to, personal computers, televisions, and other devices with display functions.

[0038] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0039] In the AI ​​training and inference process, data interaction between GPU servers and Parallel File Storage (PFS) systems occurs at multiple critical stages. To achieve the optimal balance between performance and cost, an architecture can be adopted that deploys the PFS system in a public service area and eliminates the need for a separate network for storage. Based on this, by reusing VPCs and pairing them with high-performance intelligent computing gateways, a smart computing communication acceleration solution that combines low cost and high bandwidth can be built.

[0040] However, in the aforementioned acceleration scheme, data communication between the GPU server and the PFS system is implemented using the Ethernet protocol. However, the Ethernet protocol follows a traditional network communication paradigm, and its data transmission process heavily relies on the operating system's kernel protocol stack for layer-by-layer processing. Specifically, when data interacts between the GPU server and the PFS system, multiple data copy operations between user space and kernel space are triggered. This multi-layered data copy mechanism not only significantly consumes the computing resources of the Central Processing Unit (CPU), increasing the CPU's burden when handling data transmission tasks, but also introduces a large amount of additional processing overhead, thus greatly increasing data transmission latency. This latency becomes a key bottleneck restricting the overall system performance in high-concurrency, high-data-volume GPU computing scenarios.

[0041] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure; as shown Figure 1 As shown, this embodiment provides a data processing method for artificial intelligence scenarios, applied in model training or model inference scenarios, and specifically includes the following steps:

[0042] In S101, the GPU server in the VPC network initiates a data processing request to the PFS system via the Remote Direct Memory Access (RDMA) protocol.

[0043] The data processing requests for artificial intelligence scenarios in this embodiment can include data access requests in model training and inference scenarios. For example, during model training or inference, the GPU server needs to retrieve data from the PFS system; in this case, the initiated data processing request can be a data access request. Additionally, during model training or inference, intermediate computational or inference data typically needs to be stored in the PFS for later use. In this case, the initiated data processing request can be a data storage request. Optionally, the data processing requests in this embodiment can also include other forms of data processing.

[0044] The RDMA protocol technology in this embodiment enables direct data transfer between computer memory, bypassing the operating system kernel mode throughout. Compared with the traditional Transmission Control Protocol / Internet Protocol (TCP / IP) communication mode, it has significant advantages such as low latency, high throughput, and low CPU utilization.

[0045] S102. The virtual switch forwards the data processing request to the virtual router through multiple physical switches. The data processing request forwarded by the virtual switch carries a value that identifies the RDAM protocol, so that each physical switch can prioritize the processing of the data processing request based on a pre-configured priority policy.

[0046] Specifically, each physical switch on the path from the virtual switch to the virtual router can determine that the data processing request is an RDMA-based message based on the value of the RDMA protocol identifier carried in the data processing request. Furthermore, each physical switch is pre-configured with a priority policy that requires RDMA messages to be transmitted first. Thus, upon receiving the data processing request, each physical switch can prioritize its transmission based on the priority policy.

[0047] S103. The virtual router re-encapsulates the data processing request packets and forwards them to the intelligent computing gateway.

[0048] S104. The intelligent computing gateway removes the header of the data processing request, performs address mapping, and forwards the mapped data processing request message to the PFS system.

[0049] To ensure that data processing requests can be correctly transmitted to the PFS system located in the public service area, in this embodiment, the virtual router needs to re-encapsulate the data processing request packets and forward them to the intelligent computing gateway; correspondingly, the intelligent computing gateway needs to remove the header of the data processing request, perform address mapping, and finally accurately forward the mapped packets to the PFS system.

[0050] The data processing method for artificial intelligence scenarios in this embodiment initiates a data processing request to the PFS system through the RDMA protocol by the GPU server in the VPC network. On the GPU server side, the RDMA protocol allows data to bypass the kernel protocol stack of the operating system and be efficiently transferred directly between the memory of the GPU server and the PFS system, thereby avoiding multiple data copies between user space and kernel space and effectively improving data processing efficiency.

[0051] According to authoritative tests and real-world deployment case studies, when using RDMA technology to implement write operations from a GPU server to a PFS system, the number of input / output operations per second (IOPS) is increased by approximately 40% compared to the Ethernet protocol. Simultaneously, data transmission latency is significantly reduced from milliseconds on Ethernet to around 0.2ms. This substantial improvement in latency is crucial for GPU computing tasks, effectively reducing the idle time the GPU spends waiting for data transmission, greatly improving GPU utilization efficiency, and ultimately bringing more significant performance gains and cost-effectiveness to high-performance computing applications.

[0052] This embodiment of the data processing method for artificial intelligence scenarios deeply integrates the advantages of RDMA protocol technology and proposes a set of efficient, lossless, and extremely low-latency RDMA communication solutions that are compatible with GPU servers and PFS systems. This solution can meet the stringent requirements of data transmission performance for AI training and inference scenarios in private clouds.

[0053] Figure 2 This is a schematic diagram based on the second embodiment of the present disclosure; the data processing method for artificial intelligence scenarios in this embodiment, in the above... Figure 1 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 2 As shown, the data processing method for artificial intelligence scenarios in this embodiment may specifically include the following steps:

[0054] In S201, the GPU server in the VPC network initiates a data processing request to the PFS system via the RDMA protocol; when the GPU server initiates a data processing request, a value identifying the RDMA protocol is set in the data processing request.

[0055] This embodiment focuses on how to implement RDMA communication between the GPU server and PFS in the existing architecture.

[0056] RDMA communication protocols are mainly divided into three categories: Infiniband, RoCE, and iWARP. Infiniband is designed specifically for RDMA, relying on hardware mechanisms to ensure data transmission reliability and low latency, but it requires dedicated network cards, switches, and routing equipment, resulting in high hardware costs and a closed ecosystem. RoCE and iWARP implement RDMA functionality based on standard Ethernet. Among them, RoCE has a more mature industry ecosystem and better compatibility due to support from multiple vendors. Considering cost, ecosystem, and performance requirements, this embodiment selects RoCE as the high-performance communication protocol between the GPU server and the PFS system.

[0057] The RoCE protocol includes two versions, v1 and v2: RoCEv1 is based on the Ethernet link layer and only supports communication within the same subnet; RoCEv2 is based on the User Datagram Protocol (UDP) / IP protocol stack and supports cross-subnet routing. Since reusing VPCs and high-performance intelligent computing gateways necessitates cross-subnet communication, RoCEv2 was chosen as the communication protocol.

[0058] Further optionally, in this embodiment, the GPU server in the VPC network initiates a data processing request to the PFS system via the RDMA protocol. This can be achieved by the GPU server in the VPC network using a Virtual Function (VF) unit that passes through the virtual machine to implement RDMA protocol communication and initiate a data processing request to the PFS system. This can effectively improve the efficiency of the data processing request initiated by the GPU server based on the RDMA protocol.

[0059] Optionally, in the scenario of this embodiment, virtualization partitioning technology is used on the same physical machine to generate multiple different VF units. Each VF unit has an independent input / output channel, which can support multiple different users to use different VF units on the same physical machine to communicate with the remote PFS system through the RDMA protocol. This can effectively improve the efficiency of data processing requests initiated by the GPU server based on the RDMA protocol.

[0060] Specifically, when the GPU server initiates a data processing request, it can set a value identifying the RDMA protocol in the Type of Service (TOS) field of the IP header (HEADER; HDR) of the data packet to be transmitted, based on the packet type. For example, Figure 3 This is a schematic diagram of the message structure of a data processing request initiated by the GPU server in an embodiment of this disclosure.

[0061] like Figure 3 As shown, the message structure may include Ethernet (ETH) HDR, IP HDR, UDP port, base transport header (BTH) for InfiniBand (IB) networks, IB payload, invariant cyclic redundancy check (ICRC), and frame check sequence (FCS).

[0062] The ETH HDR includes a source (src) Media Access Control (MAC) address and a destination (dst) MAC address. The IP HDR includes a src IP and a dst IP; the src MAC, dst MAC, src IP, and dst IP can be filled in according to the specific circumstances of the GPU server accessing the PFS system, and the UDPPORT can be 4791. Additionally, the IP HDR includes a TOS field. For example, referring to the above description in this embodiment, when the GPU server initiates a data processing request, it can set the message type in the TOS field to identify the transmission protocol used. For example, in this embodiment, the TOS field is set to indicate the RDMA protocol used.

[0063] Further optionally, in one embodiment of this disclosure, when the GPU server initiates a data processing request, if it detects that the message type of the data packet to be transmitted is a data packet, it sets a first value in the TOS field of IP HDR to identify the current message as a data packet transmitted based on the RDMA protocol.

[0064] If the packet type of the data packet to be transmitted is detected to be a control packet, a second value is set in the TOS field of IP HDR to identify the current packet as a control packet transmitted based on the RDMA protocol; the second value is not equal to the first value.

[0065] For example, the first value can be 40, the second value can be 48, and in practical applications, other values ​​can also be set.

[0066] Optionally, in one application scenario of this disclosure, the TOS field may include 8 bits, specifically the first 6 bits may be set with a value that identifies the protocol type of the current message, such as the value of the RMDA protocol in this embodiment.

[0067] In this embodiment, by adopting the above method, the GPU server can accurately and effectively identify the value of the RDMA protocol in the data processing request when initiating the data processing request.

[0068] S202. The virtual switch receives a data processing request, encapsulates the data processing request, and forwards the encapsulated data processing request to the virtual router through multiple physical switches, so that each physical switch can prioritize processing the data processing request based on a pre-configured priority policy.

[0069] Specifically, when the GPU server initiates access to the PFS system, the virtual switch in the Data Processing Unit (DPU) of the device where the GPU server is located receives the data processing request initiated by the GPU and encapsulates the data packet into a Virtual Extensible Local Area Network (VxLAN) packet format.

[0070] For example, Figure 4 This is a schematic diagram of the packet structure processed by the virtual switch in this embodiment of the present disclosure. The first line is the packet before the virtual switch processes it, and the second line is the packet after the virtual switch processes it. The packet includes the VxLAN VPCVXLAN Network Identifier (VNI) added after encapsulation, the outer (OUT) IP HDR, OUT ETH HDR, and OUT FCS, where VxLAN VPC VNI identifies the VNI to which the VPC belongs.

[0071] In the ETH HDR, the src MAC uses the MAC address of the GPU virtual machine, and the dst MAC uses the MAC address of the gateway. In the IP HDR, the src IP uses the IP address of the GPU server, and the dst IP uses the virtual address of PFS in the cloud. The UDP port can be 4791. Additionally, the IP HDR includes a TOS field. In the OUT IP HDR, the src IP uses the IP address of the DPU network port, and the dst IP uses the IP address of the Virtual Tunnel Endpoint (VTEP) of the virtual router. Specifically, the OUT IP HDR also transparently transmits the packet type value identifying the transport protocol from the TOS field of the inner IP HDR, such as the value identifying the RDMA protocol in this embodiment. This allows the physical switch to directly determine that the data processing request is based on the RDMA protocol based on the RDMA protocol value in the TOS field of the OUT IP HDR without disassembling the packet during subsequent forwarding. In the OUT ETH HDR, the src MAC uses the MAC address of the DPU network port, and the dst MAC uses the MAC address of the uplink switch.

[0072] To improve the efficiency of RDMA data processing requests, priority policies can be pre-configured in each physical switch. Specifically, this priority policy can be based on the value of the protocol type of the packet identified in the TOS field. For example, if the value of the protocol type in the TOS field is in the range A1-A2, the priority is highest; if the value is in the range B1-B2, the priority is next; if the value is in the range C1-C2, the priority is next, and so on. Multiple priority rules can be configured in the priority policy on each physical switch. For example, if the value of the protocol type in the TOS field of a data processing request based on RDMA transmission is in the range A1-A2, such as 40-48, then based on the RDMA value in the TOS field of the current data processing request (e.g., 40 or 48), the data processing request can be transmitted first based on the priority policy, effectively improving the efficiency of RDMA data processing requests in this embodiment.

[0073] In this embodiment, setting the RDMA flag in the first six fields of the TOS field in IP HDR is taken as an example. Furthermore, the set values ​​are transparently transmitted to the outer IP packets so that during transmission, each physical switch can promptly detect that the current packet is an RDMA packet.

[0074] Optionally, in this embodiment, each physical switch can also be configured with a priority-based flow control policy. This policy limits the transmission rate of data. If the receiving and processing capacity of the current physical switch (as a downstream device) cannot match the sending rate of the upstream device, a pause control frame is sent to the upstream device to request it to pause the transmission of data of a specified priority. Transmission resumes after a preset waiting period. This approach provides an effective flow control strategy, thereby significantly improving data transmission efficiency.

[0075] By adopting this method, the transmission rate of traffic can be controlled in a timely and effective manner on each physical switch.

[0076] S203. The virtual router re-encapsulates the data processing request packets and forwards them to the intelligent computing gateway.

[0077] Specifically, before re-encapsulating a data processing request packet, the virtual router first verifies that the destination address of the data processing request is an address within a public service area. Specifically, by matching routing rules and Access Control List (ACL) rules, confirming that the destination address of the data packet is within a public service area effectively ensures the accuracy and validity of data transmission.

[0078] Figure 5 This is a schematic diagram of the structure of packets processed by the virtual router in this embodiment of the disclosure, as shown below. Figure 5 As shown, the first row is a schematic diagram of the packet structure before virtual router encapsulation, and the second row is a schematic diagram of the packet structure after virtual router encapsulation. Figure 5 As shown, when a virtual router re-encapsulates a data processing request packet, it may specifically include: modifying the src MAC address in OUTETH HDR to the MAC address of the virtual router's network interface card; modifying the src IP address in OUT IP HDR to the virtual router's VTEP IP address; and modifying the dst IP address to the intelligent computing gateway's VTEP IP address. In short, at each hop in the path, re-encapsulation requires adaptive modifications to the MAC or IP addresses in OUT ETH HDR, OUT IP HDR, and ETH HDR to ensure the normal transmission of the data processing request. For details, please refer to relevant technologies; they will not be elaborated upon here.

[0079] S204. The intelligent computing gateway removes the header of the data processing request, performs address mapping, and forwards the mapped data packet to the PFS system.

[0080] Specifically, Figure 6 This is a schematic diagram of the message structure processed by the intelligent computing gateway in an embodiment of this disclosure, as shown below. Figure 6 As shown, the intelligent computing gateway removes the VxLAN header and, in order to ensure the correct transmission of data processing requests, adaptively modifies srcAMC in ETH HDR to the MAC address of the intelligent computing gateway's physical port, dst MAC to the MAC address of the uplink switch, srcIP in IP HDR to the virtual IP address of the GPU server in the public service area, and dst IP in IP HDR to the real IP address of the PFS system in the public service area, i.e., the real IP address of the PFS system in the public area.

[0081] It should be noted that in this embodiment, the example given is that the network segment of the PFS system's real IP address in the public service area conflicts with the network segment of the GPU server's real IP address in the public service area. To ensure the correct transmission of data processing requests, such as... Figure 4 and Figure 5As shown, the dst IP in IP HDR uses the virtual address of the PFS system in the cloud, which is unique within the network. However, since this virtual address is not a real address, the intelligent computing gateway needs to modify the dst IP in IP HDR to the real IP address of the PFS system in the public service area to ensure that data packets can be correctly transmitted to the PFS system during processing. Correspondingly, to ensure the uniqueness of the GPU server address, the src IP in IP HDR also needs to be modified to the virtual IP address of the GPU server in the public service area.

[0082] Alternatively, if in the network planning, there is no conflict between the network segment of the real IP address of the PFS system in the public service area and the network segment of the real IP address of the GPU server, then the dst IP in the IP HDR of the data processing request received by the intelligent computing gateway is itself the real IP address of the PFS system in the public area, and the src IP is also the real IP address of the GPU server in the public service area. Therefore, it is not necessary to perform address translation on the dst IP in the IP HDR.

[0083] The S205 and PFS systems receive data packets, further generate data response packets, and send them out.

[0084] Specifically, Figure 7 This is a schematic diagram of the structure of the data return packet generated by the PFS system in an embodiment of this disclosure. For example... Figure 7 As shown, when generating a data response packet, the PFS system can set the scr IP in the IP HDR to the real IP address of the PFS system's own server, set the dst IP to the virtual IP address of the GPU server in the public service area, encapsulate UDP, IB BTH, and payload data, and then send it out. The PFS system's response to the GPU server is the reverse process of the GPU server sending a data request to PFS, as described above. Figure 6 The reverse process of the shown procedure is referenced. Figure 6 The structure shown in the second line generates a data return packet.

[0085] It should be noted that, in this embodiment, when the data processing request is a data access request, the PFS system needs to send back the corresponding data as a response packet after receiving the data access request. If the data processing request is a data storage request, the PFS system needs to send back the storage result as a response packet after receiving the data storage request.

[0086] S206. The intelligent computing gateway receives the data response packet, encapsulates the header, and then sends it out.

[0087] Specifically, the processing of the intelligent computing gateway may include: modifying the destination IP in IP HDR to the GPU server IP address; and adding VxLAN header, UDP header, OUT IP HDR, and OUT ETH HDR in sequence on the outer layer.

[0088] Because the intelligent computing gateway publishes the virtual IP address range of the GPU server in the public service area, data return packets will be directed to the intelligent computing gateway for processing.

[0089] Specifically, after receiving a data packet, the intelligent computing gateway matches it against the destination IP address to determine which VPC's GPU server the destination IP belongs to. If no match is found, the data packet is discarded. Then, based on the configured ACL rules, it determines whether to allow the data packet. If allowed, the data packet is processed; otherwise, it is dropped. In this embodiment, we take the example of matching the destination IP address to a specific GPU server within a VPC and allowing the data packet based on the ACL rules.

[0090] Specifically, Figure 8 This is a schematic diagram of the structure of the intelligent computing gateway processing data return packets in an embodiment of this disclosure. Figure 8 As shown, the specific operations for handling data return packets may include modifying the IP HDR's dstip address to the GPU server's IP address, checking if a conversion rule exists for the source address, and performing the conversion if it does. This also includes recalculating the checksum, etc.

[0091] Furthermore, such as Figure 8 As shown, a VxLAN header, a UDP header, an OUTIPHDR, and an OUTETHHDR are added sequentially to the outer layer of the original data return packet. The VxLAN header sets the VPC vni to the VNI of the VPC where the GPU server resides. The UDP header sets the source port using a random function and the destination port to 4789. The OUTIPHDR sets the scr IP address to the service IP address of the intelligent computing gateway, the dst IP address to the virtual router IP address, and sets the packet length and checksum, etc. The OUTETHHDR sets the dstMAC to the MAC address of the uplink switch and the src MAC to the physical port MAC address of the intelligent computing gateway. After completion, the data return packet is sent out. It should be noted that... Figure 8 The process shown can also be considered as described above in practical applications. Figure 6 For details on the reverse process shown, please refer to the above. Figure 5 The relevant records are shown below.

[0092] S207. The virtual router receives the data report, re-encapsulates it, and sends it out.

[0093] Specifically, after receiving the data response, the virtual router parses the data packet and, based on routing rules, Address Resolution Protocol (ARP) rules, and Forwarding Database (FDB) rules, determines the dst MAC address in the inner ETH HDR and the dst IP address in the OUTIP HDR required for the data packet to be sent to the GPU server. Figure 9 This is a schematic diagram of the structure of the virtual router processing data return packets in an embodiment of this disclosure. Figure 9 As shown, when the virtual router re-encapsulates the data return packet, it includes modifying the scr MAC address in ETH HDR to the MAC address of the GPU server virtual machine, modifying the src IP in OUT IP HDR to the virtual router's service IP address, and setting the dst IP address to the GPU virtual machine's service IP address, which refers to the virtual switch's service IP address on the node where the GPU virtual machine resides. Finally, the data return packet is sent out. It should be noted that... Figure 9 The process shown can also be considered as described above in practical applications. Figure 5 For details on the reverse process shown, please refer to the above. Figure 5 The relevant records are shown below.

[0094] Furthermore, the encapsulation of data return packets by the intelligent computing gateway and the re-encapsulation of data return packets by the virtual router in this embodiment are adaptive encapsulations in network communication, designed to ensure that data return packets are correctly sent to the GPU server. For further details, please refer to relevant prior art.

[0095] After receiving the data packet, the virtual switch on the node where the S208 GPU virtual machine is located decapsulates it and finally sends the data packet to the GPU server.

[0096] Based on the above steps, the GPU server based on the RDMA protocol can access and process data or store data on the PFS system, and the PFS system can respond to the GPU server with data packets. Moreover, by adopting the above steps S205-S208, the accurate and efficient transmission of data packets can also be achieved.

[0097] The data processing method for artificial intelligence scenarios in this embodiment, by adopting the above steps, can ensure the real-time performance and accuracy of data processing between the GPU server and the PFS system based on the RDMA protocol, and can effectively improve the efficiency of data processing.

[0098] Figure 10 This is a schematic diagram based on the third embodiment of this disclosure; the data processing method for artificial intelligence scenarios in this embodiment, in the above... Figure 1 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 10 As shown, the data processing method for artificial intelligence scenarios in this embodiment may specifically include the following steps:

[0099] In S1001, the GPU server in the VPC network initiates a data processing request to the PFS system via the RDMA protocol;

[0100] S1002. The virtual switch receives a data processing request and, based on the message type in the data processing request, sets the value of the RDMA protocol identifier in the TOS field of the IP HDR of the data processing request message and passes it through to the outer IP message.

[0101] For example, in step S1002, the specific implementation of setting the value of the RDMA identifier in the TOS field of the IP header of the data processing request message according to the message type in the data processing request message may include the following steps:

[0102] (1) Detect whether the message type of the data processing request is a data message or a control message. If it is a data message, proceed to step (2); if it is a control message, proceed to step (3).

[0103] (2) Set the first value in the service type field of the IP HDR structure to identify the current packet as a data packet based on RDMA transmission;

[0104] (3) Set a second value in the service type field of the IP HDR structure to identify the current message as a control message based on RDMA transmission; the second value is not equal to the first value;

[0105] For example, the first value can be 40, the second value can be 48, and in practical applications, other values ​​can also be set.

[0106] Specifically, when the GPU server initiates access to PFS, the data packet is encapsulated into a VxLAN packet format on the virtual switch within the DPU on the device where the GPU server resides. Then, based on the data processing request, it can be determined that the data processing request is based on the RDMA protocol, and the TOS field in the IP HDR of the packet is set with a value identifying the RDMA protocol to indicate that the data processing request is a data processing request transmitted based on the RDMA protocol.

[0107] With the above Figure 2 The difference between the embodiments shown is that the above-described embodiments are... Figure 2In the illustrated embodiment, the value identifying the RDMA protocol is set in the data processing request when the GPU server initiates the data processing request. However, in this embodiment, the value identifying the RDMA protocol is set in the data processing request after the virtual switch receives the data processing request. That is, in this embodiment, the message format of the data processing request initiated by the GPU server corresponding to step S1001 does not include... Figure 3 The TOS field in the IP HDR shown is the same as the others. Correspondingly, the packet structure processed by the virtual switch in step S1202 of this embodiment is the same as described above. Figure 4 The second line is the same.

[0108] In this embodiment, by adopting the above method, the virtual switch can accurately and effectively identify the value of the RDMA protocol in the data processing request when initiating the data processing request.

[0109] S1003. The virtual switch forwards the configured data processing request to the virtual router through multiple physical switches, so that each physical switch can prioritize the data processing request based on the pre-configured priority policy.

[0110] Steps S1004-S1009 are the same as steps S203-SS08 above, and will not be repeated here.

[0111] Based on the above, this embodiment is similar to the one described above. Figure 2 The only difference in the illustrated embodiment is that, in this embodiment, the value identifying the RDMA protocol is set in the TOS field of IP HDR at the virtual switch, while... Figure 2 The illustrated embodiment sets the value identifying the RDMA protocol directly in the TOS field of IP HDR when the GPU initiates a data processing request. The remaining steps are implemented in exactly the same way; please refer to the above for details. Figure 2 The relevant descriptions of the embodiments shown will not be repeated here.

[0112] The data processing method for artificial intelligence scenarios in this embodiment, by adopting the above steps, can ensure the real-time performance and accuracy of data processing between the GPU server and the PFS system based on the RDMA protocol, and can effectively improve the efficiency of data processing.

[0113] Additionally, it should be noted that Explicit Congestion Notification (ECN) is a network layer congestion control mechanism. Its core is to transmit network congestion signals to the sender by marking data packets without dropping them, thus solving the high overhead and high latency problems of the traditional "packet loss as congestion signal" mechanism. It is especially suitable for the RoCE network of the intelligent computing center in the embodiments of this disclosure.

[0114] In the high-performance computing and data center environment disclosed herein, RoCEv2 runs on top of UDP. UDP itself is a connectionless and stateless protocol, unlike TCP which has complex acknowledgment and retransmission mechanisms. In UDP, the sender needs to use specific Application Programming Interfaces (APIs) such as the IP_ECNsocket option to detect whether the path supports ECN, and set the ECT code point (ECT(0) or ECT(1)) in the IP header of the sent UDP packet to indicate that the packet supports ECN.

[0115] Based on this, Figure 2 Step S202 of the illustrated embodiment or Figure 10 In step S1003 of the illustrated embodiment, during the process of the virtual switch forwarding data processing requests to the virtual router through multiple physical switches, when each physical switch detects traffic congestion, it marks the ECN identifier in the TOS field of the IP HDR of the congested data processing request packet. Specifically, each physical switch can be configured with a queue to store data processing requests to be processed. Each physical switch can detect the number of data processing requests to be processed in the queue. If it is less than a first preset threshold, the network is considered to be unobstructed. If the number of data processing requests to be processed is greater than or equal to the first preset threshold and less than a second preset threshold, and the second preset threshold is greater than the first preset threshold, congestion can be considered to have occurred, and a preset proportion of data processing requests can be randomly selected and marked with the ECN identifier. If the number of data processing requests processed is greater than or equal to the second preset threshold, the network is considered to be severely congested, and all data processing requests can be marked with the ECN identifier.

[0116] Specifically, referring to the description in the above embodiments, the TOS field may include 8 bits, which can be divided into two parts. As described in the above embodiments, the first 6 bits can be used to identify the type of transmission protocol used by the data packet. As described in the above embodiments, the value of the RDMA protocol used for the data processing request can be identified in the first 6 bits of the TOS field. Further, in this embodiment of the disclosure, the ECN can also be identified in the last two bits of the TOS field.

[0117] By adopting the above technical solution, when a physical switch detects traffic congestion, it can promptly and accurately notify the congestion information by marking an explicit congestion notification flag.

[0118] Correspondingly, after the PFS system receives a data processing request carrying an ECN identifier, it can determine that the network transmitting the data processing request is congested. At this time, it can generate a corresponding Congestion Notification Packet (CNP) and send the CNP to the sender of the data processing request, such as the GPU server in this embodiment, so that the sender can adjust the sending strategy of the data processing request based on the CNP, such as reducing the sending rate.

[0119] For example, the CNP carries key information about the original data stream that caused the congestion, such as the source IP address, destination IP address, transport layer ports (including source and destination port identifiers), and queue pair (QP) information, and sends it back to the sender, such as the GPU server in this embodiment. Correspondingly, after receiving the CNP, the sender can adjust the sending strategy for the data stream based on the key information of the original data stream, such as reducing the sending rate to reduce congestion, thereby effectively improving the data processing efficiency between the GPU server and the PFS system based on the RDMA protocol.

[0120] The CNP transmission method in this embodiment is the same as described above. Figure 2 The data return packet transmission method in steps S205-S208 of the illustrated embodiment is the same. For details, please refer to the relevant descriptions in the above embodiments, which will not be repeated here.

[0121] Figure 11 This is a schematic diagram based on the fourth embodiment of the present disclosure; as shown Figure 11 As shown, the data processing system 1100 for artificial intelligence scenarios in this embodiment is applied in model training or model inference scenarios, and may specifically include:

[0122] Graphics processor server 1101 is used to initiate data processing requests to parallel file storage system 1102 via remote direct memory access protocol; graphics processor server 1101 is a graphics processor server in a virtual private cloud network.

[0123] Virtual switch 1103 is used to forward the data processing request to a virtual router through multiple physical switches; the data processing request forwarded by the virtual switch carries a value identifying the Remote Direct Memory Access Protocol, so that each physical switch can prioritize the processing of the data processing request based on a pre-configured priority policy; virtual router 1104 is used to re-encapsulate the data processing request message and forward it to the intelligent computing gateway.

[0124] The intelligent computing gateway 1105 is used to remove the header of the data processing request, perform address mapping, and forward the mapped data processing request message to the parallel file storage system 1102.

[0125] The data processing system 1100 for artificial intelligence scenarios in this embodiment achieves the same data processing principle and technical effect as the above-mentioned graphics processor server 1101, parallel file storage system 1102, virtual switch 1103, virtual router 1104 and intelligent computing gateway 1105. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.

[0126] Further optionally, in one embodiment of this disclosure, when the graphics processor server 1101 initiates the data processing request, a value identifying the remote direct memory access protocol is set in the data processing request.

[0127] Further optionally, in one embodiment of this disclosure, the graphics processor server 1101 is specifically configured to, when initiating the data processing request, set a value identifying the Remote Direct Memory Access Protocol in the service type field of the Internet Protocol header of the data processing request message, according to the message type of the data packet to be transmitted.

[0128] Further optionally, in one embodiment of this disclosure, the graphics processor server 1101 is specifically used for:

[0129] When initiating the data processing request, if the message type of the data packet to be transmitted is detected to be a data packet, a first value is set in the service type field of the Internet Protocol header to identify the current message as a data packet transmitted based on the Remote Direct Memory Access Protocol.

[0130] If the packet type of the data packet to be transmitted is detected to be a control packet, a second value is set in the service type field of the Internet Protocol header to identify the current packet as a control packet transmitted based on the Remote Direct Memory Access Protocol; the second value is not equal to the first value.

[0131] Further optionally, in another embodiment of this disclosure, the virtual switch 1103 is used for:

[0132] Receive the data processing request, and set a value that identifies the remote direct memory access protocol in the data processing request.

[0133] Further optionally, in another embodiment of this disclosure, the virtual switch 1103 is used for:

[0134] Based on the message type in the data processing request, a value identifying Remote Direct Memory Access Protocol is set in the Service Type field of the Internet Protocol header of the data processing request message, and then passed through to the outer Internet Protocol message.

[0135] Further optionally, in another embodiment of this disclosure, the virtual switch 1103 is used for:

[0136] If the message type in the data processing request is detected to be a data message, a first value is set in the service type field of the Internet Protocol header to identify the current message as a data message transmitted based on the Remote Direct Memory Access Protocol.

[0137] If the message type in the data processing request is detected to be a control message, a second value is set in the service type field of the Internet Protocol header to identify the current message as a control message transmitted based on the Remote Direct Memory Access Protocol; the second value is not equal to the first value.

[0138] Further optionally, in one embodiment of this disclosure, the virtual router 1104 is also configured to:

[0139] The virtual router confirms that the destination address of the data processing request is an address in the public service area.

[0140] Further optionally, in one embodiment of this disclosure, the intelligent computing gateway 1105 is specifically used for:

[0141] Remove the header of the data processing request;

[0142] Modify the source and destination media access control addresses in the Ethernet frame header to the media access control address of the virtual port of the virtual router; modify the source Internet Protocol address in the Internet Protocol header to the virtual Internet Protocol address of the graphics processor server in the public service area.

[0143] Further optionally, in one embodiment of this disclosure, the intelligent computing gateway 1105 is specifically used for:

[0144] If the actual Internet Protocol address of the parallel file storage system conflicts with the Internet Protocol address of the graphics processor server, the intelligent computing gateway will modify the destination Internet Protocol address in the Internet Protocol header to the actual Internet Protocol address of the parallel file storage system in the public service area.

[0145] Further optionally, in one embodiment of this disclosure, the graphics processor server 1101 is specifically used for:

[0146] Based on the virtual functional units inside the direct-access virtual machine, the communication of the remote direct memory access protocol is implemented, and the data processing request to the parallel file storage system is initiated.

[0147] Further optionally, in one embodiment of this disclosure, virtualization partitioning technology is used on the same physical machine to generate multiple different virtual functional units. Each virtual functional unit has an independent input / output channel, which can support multiple different users to use different virtual functional units on the same physical machine to communicate with the remote parallel file storage system through the remote direct memory access protocol.

[0148] Further optionally, in one embodiment of this disclosure, each of the physical switches is configured with a priority-based flow control policy, which is used to limit: if the receiving and processing capacity of the current physical switch, which is a downstream device, cannot match the sending rate of the upstream device, a pause control frame is sent to the upstream device to request the upstream device to pause the transmission of data of a specified priority, and to resume transmission after a preset waiting period.

[0149] Further optionally, in one embodiment of this disclosure, the parallel file storage system 1102 is configured to generate a data response packet based on the received data processing request and send it.

[0150] The intelligent computing gateway 1105 is also used to receive the data return packet, encapsulate the header, and send it out; the virtual router 1104 is also used to receive the data return packet, re-encapsulate it, and send it out.

[0151] The virtual switch 1103 is also used to receive the data return packet, deseal it, and send the data return packet to the graphics processor server 1101; the virtual switch is the virtual switch of the node where the virtual machine of the graphics processor server 1101 is located.

[0152] Further optionally, in one embodiment of this disclosure, when each of the physical switches detects traffic congestion, it marks an explicit congestion notification identifier in the service type field of the Internet Protocol header of the congestion data processing request message.

[0153] Further optionally, in one embodiment of this disclosure, the parallel file storage system 1102 is also used for:

[0154] If a data processing request carrying the explicit congestion notification identifier is received, a congestion notification message is generated.

[0155] The congestion notification message is sent to the graphics processing server that sent the data processing request, so that the graphics processing server can adjust its sending strategy.

[0156] Further optionally, in one embodiment of this disclosure, the parallel file storage system 1102 is specifically used for:

[0157] Based on the source Internet Protocol address, destination Internet Protocol address, source port identifier, destination port identifier, and queue pair information corresponding to the data processing request, a congestion notification message is generated.

[0158] The data processing system 1100 for artificial intelligence scenarios described above achieves the same data processing principle and technical effect as the graphics processor server 1101, parallel file storage system 1102, virtual switch 1103, virtual router 1104, and intelligent computing gateway 1105 described above. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.

[0159] The acquisition, storage, and application of any type of information, such as user personal information, involved in the technical solutions disclosed herein comply with relevant laws and regulations and do not violate public order and good morals.

[0160] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0161] Figure 12 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0162] like Figure 12 As shown, device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1202 or a computer program loaded from storage unit 1208 into random access memory (RAM) 1203. The RAM 1203 may also store various programs and data required for the operation of device 1200. The computing unit 1201, ROM 1202, and RAM 1203 are interconnected via bus 1204. Input / output (I / O) interface 1205 is also connected to bus 1204.

[0163] Multiple components in device 1200 are connected to I / O interface 1205, including: input unit 1206, such as keyboard, mouse, etc.; output unit 1207, such as various types of monitors, speakers, etc.; storage unit 1208, such as disk, optical disk, etc.; and communication unit 1209, such as network card, modem, wireless transceiver, etc. Communication unit 1209 allows device 1200 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0164] The computing unit 1201 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 performs the various methods and processes described above, such as the methods of this disclosure. For example, in some embodiments, the methods of this disclosure can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1200 via ROM 1202 and / or communication unit 1209. When the computer program is loaded into RAM 1203 and executed by the computing unit 1201, one or more steps of the methods of this disclosure described above can be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to perform the methods described above in this disclosure by any other suitable means (e.g., by means of firmware).

[0165] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0166] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0167] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0168] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0169] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0170] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0171] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0172] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A data processing method for artificial intelligence scenarios, applied in model training or model inference scenarios, comprising: In a virtual private cloud network, a graphics processor server initiates a data processing request to a parallel file storage system via a remote direct memory access protocol. The virtual switch forwards the data processing request to the virtual router through multiple physical switches; the data processing request forwarded by the virtual switch carries a value identifying the Remote Direct Memory Access Protocol, so that each physical switch can prioritize the data processing request based on a pre-configured priority policy. The virtual router re-encapsulates the data processing request message and forwards it to the intelligent computing gateway. The intelligent computing gateway removes the header of the data processing request, performs address mapping, and forwards the mapped data processing request message to the parallel file storage system.

2. The method according to claim 1, wherein, When the graphics processor initiates the data processing request, it sets a value that identifies the remote direct memory access protocol in the data processing request.

3. The method according to claim 2, wherein, When the graphics processor initiates the data processing request, it sets a value identifying the Remote Direct Memory Access Protocol in the data processing request, including: When the graphics processor initiates the data processing request, it sets a value identifying the Remote Direct Memory Access Protocol (RDP) in the service type field of the Internet Protocol header of the data processing request message, according to the message type of the data packet to be transmitted.

4. The method according to claim 3, wherein, When the graphics processor initiates the data processing request, it sets a value identifying the Remote Direct Memory Access Protocol (RDP) in the Type of Service field of the Internet Protocol header of the data processing request message, based on the message type of the data packet to be transmitted, including: When the graphics processor initiates the data processing request, if it detects that the message type of the data packet to be transmitted is a data packet, it sets a first value in the service type field of the Internet Protocol header to identify the current message as a data packet transmitted based on the Remote Direct Memory Access Protocol. If the packet type of the data packet to be transmitted is detected to be a control packet, a second value is set in the service type field of the Internet Protocol header to identify the current packet as a control packet transmitted based on the Remote Direct Memory Access Protocol; the second value is not equal to the first value.

5. The method according to claim 1, wherein, Before the virtual switch forwards the data processing request to the virtual router through multiple physical switches, the method further includes: The virtual switch receives the data processing request and sets a value that identifies the Remote Direct Memory Access Protocol in the data processing request.

6. The method according to claim 5, wherein, Setting a value that identifies the Remote Direct Memory Access Protocol in the data processing request includes: Based on the message type in the data processing request, a value identifying Remote Direct Memory Access Protocol is set in the Service Type field of the Internet Protocol header of the data processing request message, and then passed through to the outer Internet Protocol message.

7. The method according to claim 1, wherein, Before the virtual router re-encapsulates the data processing request message, the method further includes: The virtual router confirms that the destination address of the data processing request is an address in the public service area.

8. The method according to claim 1, wherein, The intelligent computing gateway removes the header of the data processing request and performs address mapping, including: The intelligent computing gateway removes the header of the data processing request; The intelligent computing gateway modifies the source and destination media access control addresses in the Ethernet frame header to the media access control address of the virtual port of the virtual router; and modifies the source Internet Protocol address in the Internet Protocol header to the virtual Internet Protocol address of the graphics processor server in the public service area.

9. The method according to claim 8, wherein, The intelligent computing gateway removes the header of the data processing request and performs address mapping, and also includes: If the actual Internet Protocol address of the parallel file storage system conflicts with the Internet Protocol address of the graphics processor server, the intelligent computing gateway will modify the destination Internet Protocol address in the Internet Protocol header to the actual Internet Protocol address of the parallel file storage system in the public service area.

10. The method according to claim 1, wherein, In a virtual private cloud network, a graphics processing unit (GPU) server initiates a data processing request to a parallel file storage system via a remote direct memory access protocol, including: The graphics processor server in the virtual private cloud network communicates with the remote direct memory access protocol based on the virtual functional unit inside the virtual machine, and initiates the data processing request to the parallel file storage system.

11. The method according to claim 10, wherein, The same physical machine employs virtualization partitioning technology to generate multiple different virtual functional units. Each virtual functional unit has an independent input / output channel, enabling multiple different users to use different virtual functional units on the same physical machine to communicate with the remote parallel file storage system via the remote direct memory access protocol.

12. The method according to claim 1, wherein, Each of the physical switches is configured with a priority-based flow control policy. The priority-based flow control policy is used to limit the transmission rate of the upstream device if the receiving and processing capacity of the current physical switch, which is a downstream device, cannot match the transmission rate of the upstream device. The upstream device is then sent a pause control frame to request the upstream device to pause the transmission of data of a specified priority and resume transmission after a preset waiting period.

13. The method according to any one of claims 1-12, wherein, The method further includes: The parallel file storage system generates a data response packet based on the received data processing request and sends it. The intelligent computing gateway receives the data response packet, encapsulates the header, and sends it. The virtual router receives the data return packet, re-encapsulates it, and then sends it out. The virtual switch on the node where the virtual machine of the graphics processor server is located receives the data return packet, deseals it, and sends the data return packet to the graphics processor server.

14. The method according to claim 1, wherein, When each physical switch detects traffic congestion, it marks an explicit congestion notification identifier in the service type field of the Internet Protocol header of the congestion data processing request message.

15. The method according to claim 14, wherein, The method further includes: If the parallel file storage system receives the data processing request carrying the explicit congestion notification identifier, it generates a congestion notification message. The parallel file storage system sends the congestion notification message to the graphics processing server that sent the data processing request, so that the graphics processing server can adjust its sending strategy.

16. The method according to claim 15, wherein, Generate a congestion notification message, including: The congestion notification message is generated based on the source Internet Protocol address, destination Internet Protocol address, source port identifier, destination port identifier, and queue pair information corresponding to the data processing request.

17. A data processing system for artificial intelligence scenarios, applied in model training or model inference scenarios, comprising: A graphics processing unit server is used to initiate data processing requests to a parallel file storage system via a remote direct memory access protocol. The graphics processing server is a graphics processing server in a virtual private cloud network. A virtual switch is used to forward the data processing request to a virtual router through multiple physical switches; the data processing request forwarded by the virtual switch carries a value identifying the Remote Direct Memory Access Protocol, so that each physical switch can prioritize the data processing request based on a pre-configured priority policy. The virtual router is used to re-encapsulate the data processing request message and forward it to the intelligent computing gateway. The intelligent computing gateway is used to remove the header of the data processing request, perform address mapping, and forward the mapped data processing request message to the parallel file storage system.

18. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-16.

19. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-16.

20. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-16.

Citation Information

Patent Citations

  • Data forwarding method and device under virtual network and computer program product

    CN115225634A

  • Message processing method, gateway equipment and storage system

    CN116566933A

  • Data direct connection method in AI computing power resource remote calling scene

    CN116582576A

  • Recommendation model estimation system and use method thereof

    CN117311960A

  • Remote direct memory access method and device

    CN118860924A