Data processing method, system, apparatus, device, medium, and product
By introducing the DPU and CXL protocols into the computing power awareness network, the computing power information of hardware accelerators and central processing units is directly transmitted, solving the problem of low bandwidth utilization in the interaction between the CPU and the network card, and realizing more efficient utilization of computing resources.
Patent Information
- Application Number
- PCT/CN2025/083373
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-27
- Filing Date
- 2025-03-19
- Publication Date
- 2025-12-04
AI Technical Summary
In computing power-aware networks, bandwidth utilization decreases during data interaction between the CPU and network interface card, leading to a decline in the efficiency of computing resources.
Using the DPU as an intermediary, the computing power information of the hardware accelerator and the central processing unit is directly transmitted to the DPU through an open interconnection protocol. The CXL protocol is used to optimize the data transmission path, reduce the dependence on the CPU, and improve bandwidth utilization.
It improves CPU bandwidth utilization, reduces computing resource consumption, and enhances data transmission efficiency and device compatibility.
Smart Images

Figure CN2025083373_04122025_PF_FP_ABST
Abstract
Description
A data processing method, system, apparatus, equipment, medium, and product
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. CN202410658894.7, filed on May 27, 2024, entitled “A data processing method, system, apparatus, device, medium and product”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of communication technology, and in particular to a data processing method, a distributed storage system, a centralized storage system, a data processing device, a data processing equipment, a medium, and a computer-readable instruction product. Background Technology
[0004] Computational power awareness is a network's comprehensive perception of the deployment location, real-time status, load information, and business needs of computing resources and services. By connecting distributed computing nodes through ubiquitous network connections, it enables automated service deployment, optimal routing, and load balancing, thereby building a new network infrastructure that can perceive computing power. This ensures that the network can schedule computing resources in different locations on demand and in real time, improving the utilization rate of network and computing resources.
[0005] Some computing power-aware networks use a central processing unit (CPU) and a single graphics processing unit (GPU) or a single field-programmable gate array (FPGA) to communicate. All data and network card interactions are implemented through the CPU. Since the computing resources corresponding to such workloads as GPUs or FPGAs consume a lot of bandwidth in the transmission between the CPU and the network card, the remaining bandwidth utilization is reduced when transmitting data between the CPU and the network card. Summary of the Invention
[0006] In a first aspect, embodiments of this application provide a data processing method applied to a computing node. The computing node includes a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is a processor that supports the open interconnect protocol and supports computing power network processing of computing power information corresponding to the hardware accelerator and the CPU, respectively. The DPU is used to receive computing power information from the CPU and computing power information from the hardware accelerator.
[0007] The data processing method includes the following steps:
[0008] Obtain the first data packet;
[0009] The first data packet is parsed to extract the corresponding first computing power information; and the first computing power information is encapsulated to obtain a second data packet conforming to the data frame format transmitted between any two nodes; and
[0010] The routing node to be transmitted is determined based on the flow type corresponding to the first computing power information, so that the second data packet can be transmitted to the routing node.
[0011] The first data packet is obtained by transmitting the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU based on the open interconnection protocol and then performing computing power network processing.
[0012] In some embodiments, the hardware accelerator is at least one or more of a graphics processor, a field-programmable gate array, and an application-specific integrated circuit.
[0013] In some embodiments, transmitting the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU based on an open interconnection protocol includes the following steps:
[0014] Obtain the protocol transmission unit corresponding to the open interconnection protocol;
[0015] The flow type of computing power information is determined based on the computing power information of the hardware accelerator and / or the computing power information of the central processing unit;
[0016] Set the flow type of computing power information to the data slot of the protocol-level message in the protocol transmission unit;
[0017] The computing power information is set in the data slot of the protocol transmission unit, which is used to represent the data block corresponding to the request-response message; and
[0018] Multiple configured protocol transmission units are transmitted to the DPU.
[0019] In some embodiments, the process of determining the flow type of computing power information includes the following steps:
[0020] Obtain the computing tasks corresponding to the hardware accelerators and / or central processing units corresponding to the computing power information; and
[0021] The flow type of computing power information is determined based on the computing task.
[0022] Among them, the flow types include computing power awareness type, computing power announcement type, test type and scheduling type.
[0023] In some embodiments, setting computing power information in a data slot of the protocol transmission unit, which is used to represent a data block corresponding to a request-response message, includes the following steps:
[0024] The type of computing power service identifier for computing power information is determined based on its source.
[0025] Obtain network resource information corresponding to computing power information; and
[0026] The computing power service identifier type and network resource information of the computing power information are saved to the data slot of the protocol transmission unit, which is used to represent the data block corresponding to the request and response message.
[0027] Among them, network resource information includes at least one or more of the following: CPU utilization, memory utilization, GPU utilization, video memory utilization, disk utilization, network packet loss rate, and network bandwidth utilization.
[0028] In some embodiments, determining the computing service identifier type of the computing power information based on its source includes the following steps:
[0029] The target source direction for obtaining computing power information; and
[0030] The computing power service identifier type is determined based on the target source direction.
[0031] The target source is either a central processing unit or a hardware accelerator.
[0032] In some embodiments, encapsulating the first computing power information to obtain a second data packet corresponding to a data frame format for transmission between any two nodes includes the following steps:
[0033] Obtain the computing service identifier type and corresponding network resource information of the first computing power information;
[0034] The type information of the first computing power information is determined based on the computing power service identifier type;
[0035] Match the corresponding actual network resources based on the network resource information of the first computing power information;
[0036] Using actual network resources, primary computing power information, and type information as computing power perception information; and
[0037] The computing power perception information is encapsulated to obtain a second data packet that conforms to the data frame format transmitted between any two nodes.
[0038] In some embodiments, encapsulating the computing power awareness information to obtain a second data packet corresponding to a data frame format that conforms to the transmission between any two nodes includes the following steps:
[0039] Obtain the data space of the payload data corresponding to the first data frame transmitted between any two nodes;
[0040] Obtain the data space corresponding to the computing power perception information; and
[0041] The data space of the payload data corresponding to the first data frame is compressed according to the data space corresponding to the computing power perception information, so as to encapsulate the data space corresponding to the computing power perception information within the first data frame to obtain the second data packet.
[0042] In some embodiments, encapsulating the computing power awareness information to obtain a second data packet corresponding to a data frame format that conforms to the transmission between any two nodes includes the following steps:
[0043] Obtain the data space of the payload data corresponding to the first data frame transmitted between any two nodes;
[0044] Obtain the data space corresponding to the computing power perception information;
[0045] A corresponding preset data space is reserved based on the data space corresponding to the computing power perception information; wherein, the preset data space is greater than or equal to the data space corresponding to the computing power perception information; and
[0046] The data space of the payload data corresponding to the first data frame is compressed according to the preset data space so that the data space corresponding to the computing power perception information is encapsulated in the first data frame to obtain the second data packet.
[0047] In some embodiments, the structure of the first data frame includes at least a protocol header, an Internet Protocol header, a User Datagram Protocol header, an Infinite Bandwidth Protocol header, a data space, and a Cyclic Redundancy Check (CRC) code.
[0048] The data space includes the data space corresponding to the computing power perception information and the data space of the compressed payload data.
[0049] In some embodiments, the protocol header is an Ethernet protocol header, and the first data frame is a data frame based on Ethernet Remote Direct Data Access technology.
[0050] In a second aspect, embodiments of this application also provide a data processing method applied to a routing node, the routing node including a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU, and both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is a processor that supports open interconnect protocols and supports computing power network processing of computing power information corresponding to the hardware accelerator and the CPU, respectively. The DPU is used to receive computing power information from the CPU and the hardware accelerator. The data processing method includes the following steps:
[0051] Obtain the second data packet transmitted by the computing node;
[0052] The second data packet is parsed to extract the corresponding second computing power information; and
[0053] The second computing power information is used to perform computing power scheduling calculations to generate a routing table for computing power scheduling.
[0054] The computing node is used to acquire the first data packet, parse the first data packet to extract the corresponding first computing power information, and encapsulate the first computing power information to obtain the second data packet corresponding to the data frame format that conforms to the transmission between any two nodes. The first data packet is transmitted from the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU of the computing node based on the open interconnection protocol, and is obtained by computing power network processing.
[0055] In some embodiments, after obtaining the second data packet, the method further includes: determining the source of the second data packet.
[0056] In response to the fact that the source of the second data packet is a node other than its own routing node, the next round of computing power awareness process will be carried out after the computing power scheduling is completed.
[0057] In a third aspect, embodiments of this application also provide a distributed storage system, which includes computing nodes and routing nodes. Each computing node and routing node includes a DPU, a hardware accelerator, a memory, and a central processing unit. The memory is connected to the central processing unit. Both the central processing unit and the hardware accelerator are connected to the DPU through an open interconnect protocol. The DPU is used to receive computing power information from the central processing unit and the hardware accelerator. The DPU is a processor that supports open interconnect protocols and supports computing power network processing of the computing power information corresponding to the hardware accelerator and the central processing unit, respectively.
[0058] A computing node is used to acquire a first data packet, parse the first data packet to extract the corresponding first computing power information, and encapsulate the first computing power information to obtain a second data packet corresponding to the data frame format for transmission between any two nodes. Based on the flow type corresponding to the first computing power information, a routing node is determined to transmit the second data packet to the routing node. The first data packet is obtained by transmitting the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU based on an open interconnection protocol and performing computing power network processing.
[0059] The routing node is used to obtain the second data packet, parse the second data packet to extract the corresponding second computing power information, and perform computing power scheduling calculations on the second computing power information to generate a routing table for computing power scheduling.
[0060] In a fourth aspect, embodiments of this application also provide a data processing method applied to a control device, the data processing method comprising the following steps:
[0061] Acquire computational tasks for data processing;
[0062] The control computing node transmits the computing power information of its hardware accelerator and / or central processing unit (CPU) to the computing node's DPU based on an open interconnect protocol, according to the computing task. This allows the DPU to perform computing power network processing to obtain a first data packet, parse the first data packet to extract the corresponding first computing power information, encapsulate the first computing power information to obtain a second data packet conforming to the data frame format for transmission between any two nodes, and determine the routing node for transmission based on the flow type corresponding to the first computing power information, so that the second data packet can be transmitted to the routing node.
[0063] The control routing node obtains the second data packet, parses and extracts the corresponding second computing power information, and performs computing power scheduling calculations on the second computing power information to generate a routing table for computing power scheduling.
[0064] The computing nodes and routing nodes each include a DPU, a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is used to receive computing power information from the CPU and the hardware accelerator. The DPU is a processor that supports the open interconnect protocol and can perform computing power network processing on the computing power information corresponding to the hardware accelerator and the CPU, respectively.
[0065] In a fifth aspect, embodiments of this application also provide a centralized storage system, which includes a control device, computing nodes, and routing nodes. Each computing node and routing node includes a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is used to receive computing power information from the CPU and the hardware accelerator. The DPU is a processor that supports open interconnect protocols and supports computing power network processing of the computing power information corresponding to the hardware accelerator and the CPU, respectively.
[0066] A centralized storage system includes: control devices, computing nodes, and routing nodes.
[0067] Control equipment is used to acquire and process computational data.
[0068] The computing node is used to transmit the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU of the computing node based on the computing task, so that the DPU of the computing node can perform computing power network processing to obtain a first data packet, parse the first data packet to extract the corresponding first computing power information, encapsulate the first computing power information to obtain a second data packet corresponding to the data frame format for transmission between any two nodes, and determine the routing node to be transmitted according to the flow type corresponding to the first computing power information, so as to transmit the second data packet to the routing node.
[0069] The routing node is used to obtain the second data packet, parse the second data packet to extract the corresponding second computing power information, and perform computing power scheduling calculations on the second computing power information to generate a routing table for computing power scheduling.
[0070] In a sixth aspect, embodiments of this application also provide a data processing apparatus applied to a computing node. The computing node includes a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU, and both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is a processor that supports open interconnect protocols and supports computing power network processing of computing power information corresponding to the hardware accelerator and the CPU, respectively. The DPU is used to receive computing power information from the CPU and the hardware accelerator.
[0071] The data processing device includes the following modules:
[0072] The first acquisition module is used to acquire the first data packet;
[0073] The first parsing and extraction module is used to parse and extract the first data packet to obtain the corresponding first computing power information, and to encapsulate the first computing power information to obtain a second data packet corresponding to the data frame format transmitted between any two nodes; and
[0074] The first determining module is used to determine the routing node to be transmitted based on the flow type corresponding to the first computing power information, so as to transmit the second data packet to the routing node.
[0075] The first data packet is obtained by transmitting the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU based on the open interconnection protocol and then performing computing power network processing.
[0076] In a seventh aspect, embodiments of this application also provide a data processing apparatus applied to a routing node. The routing node includes a DPU, a hardware accelerator, a memory, and a central processing unit. The memory is connected to the central processing unit. Both the central processing unit and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is a processor that supports the open interconnect protocol and supports computing power network processing of computing power information corresponding to the hardware accelerator and the central processing unit, respectively. The DPU is used to receive computing power information from the central processing unit and computing power information from the hardware accelerator.
[0077] The data processing device includes:
[0078] The second acquisition module is used to acquire the second data packet transmitted by the computing node;
[0079] The second parsing and extraction module is used to parse and extract the second data packet to obtain the corresponding second computing power information; and
[0080] The computing power scheduling module is used to perform computing power scheduling calculations on the second computing power information to generate a routing table for computing power scheduling.
[0081] The computing node is used to acquire the first data packet, parse the first data packet to extract the corresponding first computing power information, and encapsulate the first computing power information to obtain the second data packet corresponding to the data frame format that conforms to the transmission between any two nodes. The first data packet is transmitted to the DPU based on the computing power information of the hardware accelerator and / or the computing power information of the central processing unit based on the open interconnection protocol, and is obtained by computing power network processing.
[0082] In an eighth aspect, embodiments of this application also provide a data processing apparatus for use in a control device, the data processing apparatus comprising:
[0083] The third acquisition module is used to acquire computational tasks for data processing.
[0084] The first control module is used to control the computing node to transmit the computing power information of its hardware accelerator and / or central processing unit to its DPU based on an open interconnection protocol, according to the computing task. This allows the DPU to perform computing power network processing to obtain a first data packet, parse and extract the corresponding first computing power information from the first data packet, encapsulate the first computing power information to obtain a second data packet conforming to the data frame format for transmission between any two nodes, and determine the routing node to be transmitted based on the flow type corresponding to the first computing power information, so as to transmit the second data packet to the routing node.
[0085] The second control module is used to control the routing node to obtain the second data packet, parse the second data packet to extract the corresponding second computing power information, and perform computing power scheduling calculation on the second computing power information to generate a routing table for computing power scheduling.
[0086] The computing nodes and routing nodes each include a DPU, a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is used to receive computing power information from the CPU and the hardware accelerator. The DPU is a processor that supports the open interconnect protocol and can perform computing power network processing on the computing power information corresponding to the hardware accelerator and the CPU, respectively.
[0087] In a ninth aspect, embodiments of this application also provide a data processing apparatus, the apparatus comprising: a memory and a processor.
[0088] Memory is used to store computer-readable instructions.
[0089] A processor is used to execute computer-readable instructions to implement the steps of the data processing method described above.
[0090] In a tenth aspect, embodiments of this application also provide a non-volatile computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the steps of the data processing method as described in any of the above embodiments.
[0091] In an eleventh aspect, embodiments of this application also provide a computer-readable instruction product, including computer-readable instructions that, when executed by a processor, implement the steps of the data processing method in any of the above embodiments.
[0092] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description
[0093] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0094] Figure 1 is a flowchart of a data processing method applied to a computing node according to an embodiment of this application;
[0095] Figure 2 is a schematic diagram of a traditional computing power network architecture;
[0096] Figure 3 is a schematic diagram of a computing network architecture provided in an embodiment of this application;
[0097] Figure 4 is a schematic diagram of the structure of a protocol transmission unit provided in an embodiment of this application;
[0098] Figure 5 is a schematic diagram of a standard data frame format provided in an embodiment of this application;
[0099] Figure 6 is a schematic diagram of an improved data frame format provided in an embodiment of this application;
[0100] Figure 7 is a flowchart of a data processing method applied to a routing node according to an embodiment of this application;
[0101] Figure 8 is a schematic diagram of a distributed storage system provided in an embodiment of this application;
[0102] Figure 9 is a flowchart of a data processing method provided in an embodiment of this application;
[0103] Figure 10 is a schematic diagram of a centralized storage system provided in an embodiment of this application;
[0104] Figure 11 is a structural diagram of a data processing device applied to a computing node according to an embodiment of this application;
[0105] Figure 12 is a structural diagram of a data processing device applied to a routing node according to an embodiment of this application;
[0106] Figure 13 is a structural diagram of a data processing device applied to a control device according to an embodiment of this application;
[0107] Figure 14 is a structural diagram of a data processing device provided in an embodiment of this application. Detailed Implementation
[0108] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0109] This application provides a data processing method, system, apparatus, device, medium, and product to solve the problem that in computing networks, interaction is only achieved through the CPU and the network card, resulting in reduced utilization of the remaining bandwidth during transmission between the CPU and the network card.
[0110] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0111] Computing power networks are generally classified into three types of computing power network awareness schemes: centralized, distributed, and hybrid. In the centralized scheme, computing and network resources at the cloud, edge, and endpoint levels are collected and distributed uniformly by a centralized orchestrator. The orchestrator also selects the optimal routing and forwarding paths based on computing service requirements and the perceived overall network computing and network resource status, and distributes these paths to the routing and forwarding nodes of the computing power network for data forwarding. The centralized scheme requires minimal modification to the existing network architecture and protocols, making it easy to implement, but it suffers from poor flexibility and scalability. In the distributed scheme, computing service nodes register their computing resource status information with the nearest computing power network node, which then publishes this information to the network. Network devices announce the computing and network resource status and, based on this status, forward computing tasks to the corresponding computing service nodes. The distributed scheme fully leverages the control capabilities of routing nodes in the bearer network, resulting in higher coordination between computing and network resources, greater flexibility, and higher efficiency, but it requires significant modifications to the existing network architecture and protocols. The hybrid approach involves coordinating centralized and distributed solutions, using distributed solutions in certain areas and centralized solutions for critical nodes.
[0112] When computing power fluctuates frequently, distributed computing power awareness increases the load on individual computing nodes. The CPU needs to constantly communicate with various computing resource components, reducing the timeliness of computing power information updates and decreasing the utilization of the CPU and other computing resources. For example, when the CPU communicates with the load, additional overhead is added. Originally utilizing 90% of the GPU, only 1% is used for information collection, resulting in only 89% of computing resources being utilized. Furthermore, the communication process between the CPU and the network card involves the load transmitting data through this communication, consuming CPU computing resources. Due to limited bandwidth, once some of the load's computing resources are used, the remaining loads, such as storage loads, will have less CPU computing resources available, leading to reduced CPU bandwidth utilization.
[0113] Furthermore, since the perception of computing power information requires a redesign of the network information protocol, it will increase the load on network resources by adding extra probe data packets, and will also pose certain difficulties for the development and maintenance of the new network communication protocol. The data processing method provided in this application can solve the above-mentioned technical problems.
[0114] In a first aspect, Figure 1 is a flowchart of a data processing method applied to a computing node according to an embodiment of this application. The computing node includes a DPU (Data Processing Unit), a hardware accelerator, a memory, and a central processing unit. The memory is connected to the central processing unit, and both the central processing unit and the hardware accelerator are connected to the DPU through an open interconnect protocol. The DPU is a processor that supports the open interconnect protocol and supports computing power network processing of the computing power information corresponding to the hardware accelerator and the central processing unit, respectively. The DPU is used to receive the computing power information of the central processing unit and the computing power information of the hardware accelerator.
[0115] As shown in Figure 1, the data processing method includes the following steps:
[0116] Step S11: Obtain the first data packet;
[0117] Step S12: Parse the first data packet to extract the corresponding first computing power information; and encapsulate the first computing power information to obtain a second data packet corresponding to the data frame format transmitted between any two nodes; and
[0118] Step S13: Determine the routing node to be transmitted based on the flow type corresponding to the first computing power information, so as to transmit the second data packet to the routing node.
[0119] The first data packet is obtained by transmitting the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU based on the open interconnection protocol and then performing computing power network processing.
[0120] Specifically, a compute node is a basic unit based on a distributed computing architecture. It possesses independent computing resources and can execute various computing tasks. A compute node can be a physical server, a virtual machine, or a container, and it collaborates with other nodes to complete tasks. Each compute node has its own processor, memory, and storage devices to achieve efficient computing. A network node is a collection of network nodes, compute nodes, and routers that communicate via protocols. If multiple server nodes exist within a certain address range, routing or filtering methods are used to route between compute nodes and router nodes to complete the computing power scheduling process.
[0121] The first data packet is acquired. This first data packet is transmitted to the DPU based on the computing power information of the hardware accelerator and / or the computing power information of the central processing unit (CPU) using an open interconnect protocol, and then processed by the computing power network. The DPU differs from the network interface card (NIC) in traditional computing power network architectures. In this embodiment, the DPU is a processor that supports open interconnect protocols and can perform computing power network processing on the computing power information corresponding to the hardware accelerator and CPU, respectively. The DPU in this application is a software offloading device that can act as a NIC, continuing the functions of a smart NIC such as CPU overhead release, programmability, task acceleration, and process management, and realizing general-purpose programmable acceleration of both the control plane and data plane. The DPU is used to accelerate network, storage, and security tasks in data centers and cloud computing applications. It typically includes components such as a high-speed network interface, a dedicated processor, and memory, and can perform tasks such as high-speed data transmission, network security, and data processing in data centers.
[0122] Figure 2 is a schematic diagram of the traditional computing power network architecture. As shown in Figure 2, in the traditional architecture (computing power routing node), various software and interface tools need to communicate with various computing power components (storage, hardware accelerators (hardware accelerator 0', hardware accelerator 1', hardware accelerator 2', hardware accelerator 3', ..., hardware accelerator N'), GPUs and network cards, etc., through the CPU to collect various computing power information and encapsulate it into independent data packets for network transmission through the network card. This method will consume a lot of CPU computing resources, and the communication latency between devices cannot meet the current big data computing needs of intelligent computing centers.
[0123] Figure 3 is a schematic diagram of a computing power network architecture provided in an embodiment of this application. As shown in Figure 3, the traditional network architecture is removed, and a DPU is directly used to perform the computing power perception task of the computing power network.
[0124] Specifically, the memory and central processing unit (CPU) are connected, with the CPU still managing the storage resources. The CPU connects to the Data Processing Unit (DPU) and communicates via the Compute Express Link (CXL) protocol. In other words, the storage components communicate with the CPU through the CXL protocol. The components corresponding to the computing resources are stored within hardware accelerators (Hardware Accelerator 0', Hardware Accelerator 1', Hardware Accelerator 2', Hardware Accelerator 3', ..., Hardware Accelerator N'), and communication between the hardware accelerators and the DPU also uses the CXL protocol. The CXL protocol is a high-speed serial protocol that allows for fast and reliable data transfer between different components within a computer system. It aims to address bottlenecks in high-performance computing, including memory capacity, memory bandwidth, and input / output (I / O) latency. CXL also enables memory expansion and sharing, and can communicate with peripherals such as computing accelerators (GPUs and FPGAs), providing faster and more flexible data exchange and processing methods.
[0125] The caching portion (CXL.cache) of the CXL protocol defines the interaction between the host and the device, allowing connected CXL devices to efficiently cache host memory with extremely low latency using request and response methods. In this embodiment, the CXL.cache protocol is used to actively acquire computing power information from the CPU and hardware accelerator, significantly improving bandwidth and latency. It should be noted that the first data packet can be obtained from the computing power information processing of the hardware accelerator, the CPU, or a combination of both; no limitation is made here, and it can be set according to the actual situation. This embodiment separates the computing power information corresponding to the hardware accelerator and the CPU, so that the computing resources of the hardware accelerator do not need to be collected by the CPU, thereby improving CPU bandwidth utilization.
[0126] In some embodiments, the hardware accelerator is at least one or more of a graphics processor, a field-programmable gate array, and an application-specific integrated circuit.
[0127] Specifically, this embodiment does not limit the specific components of the hardware accelerator; they can be set according to actual conditions. Hardware acceleration is the process of transferring some software running on the CPU to idle hardware resources. These resources can be graphics cards, sound cards, GPUs, or special devices (such as AI accelerators) to optimize resource usage and performance. Most browsers also have acceleration features.
[0128] The CPU is the core of all computer systems, designed to manage all tasks. However, managing all tasks doesn't guarantee high efficiency, so tasks like video encoding / decoding and graphics rendering are performed on dedicated devices like GPUs. Hardware acceleration offloads routine tasks from the CPU to specially designed hardware that can perform these tasks more efficiently.
[0129] In step S12, the first data packet is parsed to extract the corresponding first computing power information. In this embodiment, considering that the first data packet is obtained from within the computing node, it is necessary to determine the computing power awareness process and the specific task operation. Therefore, the first data packet needs to be parsed to extract the corresponding first computing power information. The specific computing power task can be determined through the first computing power information. During data transmission, the data transmission between the computing node and the routing node is taken into account. Therefore, the first computing power information needs to be encapsulated into a data frame for transmission. Since the CXL protocol only transmits data within hardware components, and computing nodes and routing nodes typically use remote transmission due to long path distances, a protocol suitable for remote transmission is required. Therefore, in this embodiment, the first computing power information is encapsulated in a remote protocol data frame for transmission, forming a second data packet in the form of a data frame during the encapsulation process.
[0130] For the encapsulation process, data in the format of the remote protocol is simply added. The specific remote protocol is not limited and can be set according to the actual situation. Similarly, the specific address where the encapsulated data is stored is also not limited and can be set according to the actual remote protocol.
[0131] In step S13, the specific transmission routing node is determined based on the flow type corresponding to the first computing power information, thereby determining the transmission path between the current computing node and the routing node. The flow type records the computing task of the first computing power information and which specific routing node it is transmitted to, thus revealing the transmission task of the target routing node. The specific path strategy set here is not limited and can be set according to the current routing path algorithm.
[0132] This application provides a data processing method applied to a computing node. The computing node includes a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is a processor that supports the open interconnect protocol and supports computing power network processing of the computing power information corresponding to the hardware accelerator and the CPU, respectively. The DPU is used to receive the computing power information of the CPU and the computing power information of the hardware accelerator.
[0133] The data processing method includes the following steps: acquiring a first data packet, parsing the first data packet to extract the corresponding first computing power information, encapsulating the first computing power information to obtain a second data packet corresponding to the data frame format for transmission between any two nodes, and determining the routing node to be transmitted based on the flow type corresponding to the first computing power information, so as to transmit the second data packet to the routing node.
[0134] The computing resources (hardware accelerators) of the workload are transmitted directly to the DPU via an open interconnect protocol, eliminating the need for a transmission path solely through the CPU and network interface card (NIC). This separate transmission path for computing and storage resources improves CPU bandwidth utilization, provides more bandwidth for other storage resources, and conserves CPU computing resources. Furthermore, data transmission based on the open interconnect protocol significantly improves the bandwidth and latency for information retrieval within the DPU. Simultaneously, the separate transmission paths for the hardware accelerators and CPU allow components that do not support the open interconnect protocol to still utilize the existing CPU for communication, ensuring sufficient device compatibility.
[0135] Secondly, the process of determining the first data packet, while improving CPU bandwidth utilization, is based on the CXL protocol to improve transmission latency, regardless of whether the first data packet comes from the hardware accelerator or the central processing unit.
[0136] In some embodiments, the first data packet is obtained by transmitting the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU based on an open interconnection protocol and performing computing power network processing.
[0137] As shown in Figure 3, the computing power information of the hardware accelerator and / or the computing power information of the central processing unit are transmitted to the DPU using the CXL protocol. Within the DPU, the computing power information is processed by the computing power network to obtain the first data packet. It should be noted that the computing power network processing procedure in this embodiment can be the same as or different from conventional computing power network processing methods, or a new computing power network processing method may be used, etc., without limitation.
[0138] The process for determining the first data packet provided in this embodiment improves transmission latency based on the CXL protocol, regardless of whether the first data packet comes from a hardware accelerator or a central processing unit, while increasing CPU bandwidth utilization.
[0139] In some embodiments, the computing power information of the hardware accelerator and / or the computing power information of the central processing unit are transmitted to the DPU based on an open interconnection protocol. The steps include:
[0140] Obtain the protocol transmission unit corresponding to the open interconnection protocol;
[0141] The flow type of computing power information is determined based on the computing power information of the hardware accelerator and / or the computing power information of the central processing unit;
[0142] Set the flow type of computing power information to the data slot of the protocol-level message in the protocol transmission unit;
[0143] The computing power information is set in the data slot of the protocol transmission unit, which is used to represent the data block corresponding to the request-response message; and
[0144] Multiple configured protocol transmission units are transmitted to the DPU.
[0145] Specifically, the protocol transmission unit (Flit) is the smallest transmission unit of the CXL protocol. Figure 4 is a schematic diagram of the structure of a protocol transmission unit provided in an embodiment of this application. As shown in Figure 4, the Flit size of the device or interface (CXL.mem) for memory communication in CXL.cache / CXL is a fixed 528 bits, including a 16-bit Cyclic Redundancy Check (CRC) and four 16-byte data slots. The leftmost part of the figure shows an overview of the standard Flit data unit, which mainly consists of the protocol transmission unit header field (Flit Header), the protocol-level message data slot (Header Slot), three generic data slots (Generic Slot 1, Generic Slot 2, and Generic Slot 3), and the CRC. The "Header" slot carries the "header" information of link layer specific information, including the rest of the header and the protocol-level message definitions in the other data slots in the Flit. The "Generic" data slot contains one or more request / response messages or a single 16-byte data block.
[0146] The flow type of computing power information defines the type of protocol transmission unit, and the flow direction of computing power information needs to be determined in the computing power network. Simultaneously, the flow type of computing power information is set in the data slot (Header Slot) of the protocol-level message in Flit. Computing power information is set in the data slot (Generic Slot) of Flit, which represents the data block corresponding to the request-response message. The last data slot of Flit serves as the transmission carrier of computing power information, effectively increasing the information content of the data packet and providing the collection of computing power information without adding additional computing power data probe packets. Then, the currently modified multiple Flits are used as a second data packet for transmission into the DPU. In other words, the design of the Flit data frame format enables internal transmission within the computing node.
[0147] In some embodiments, the process of determining the flow type of computing power information includes the following steps: obtaining the computing task corresponding to the hardware accelerator and / or central processing unit corresponding to the computing power information, and determining the flow type of computing power information based on the computing task.
[0148] Among them, the flow types include computing power awareness type, computing power announcement type, test type and scheduling type.
[0149] Specifically, the flow type corresponding to the computing power information needs to be determined through the issued computing tasks or the computing tasks corresponding to the hardware accelerators and / or CPUs. This includes computing power awareness type, computing power announcement type, test type, and scheduling type. The computing power awareness type collects data from internal nodes; the computing power announcement type indicates which node the data specifically flows to; the test type is only used for test packets and is not used for current transmission; and the scheduling type indicates which node to schedule the data to. This scheduling is not the problem addressed in this embodiment. The scheduling process occurs after the computing power announcement type and after the data has been transmitted to the target node. Additionally, the unit data corresponding to the flow type may also include reserved bits. These reserved bits are set for subsequent actual filling to enrich subsequent computing tasks. The reserved bits occupy 3 bits, and their bit positions can be seen in Figure 4, located in the general data slot three.
[0150] The computing power information provided in this embodiment corresponds to the flow type in order to improve data processing efficiency. The flow type indicates which node the data is transmitted to, which facilitates subsequent data transmission.
[0151] In some embodiments, setting computing power information in a data slot of the protocol transmission unit, which is used to represent a data block corresponding to a request-response message, includes the following steps:
[0152] The type of computing power service identifier for computing power information is determined based on its source.
[0153] Obtain network resource information corresponding to computing power information; and
[0154] The computing power service identifier type and network resource information of the computing power information are saved to the data slot of the protocol transmission unit, which is used to represent the data block corresponding to the request and response message.
[0155] Among them, network resource information includes at least one or more of the following: CPU utilization, memory utilization, GPU utilization, video memory utilization, disk utilization, network packet loss rate, and network bandwidth utilization.
[0156] Specifically, the encapsulation within the data slot, as shown in Figure 4, mainly consists of computing power service identification type, CPU utilization, memory utilization, GPU utilization, video memory utilization, disk utilization, network packet loss rate, and network bandwidth utilization. Since the result of computing power perception is used for subsequent services such as computing power scheduling, this application designs the data precision at this location to be a half-precision floating-point number (FP)16, thus reserving 8 bits as reserved bits.
[0157] In some embodiments, determining the computing service identifier type of computing power information based on the source of the computing power information includes the following steps: obtaining the target source direction of the computing power information, and determining the computing service identifier type of the computing power information based on the target source direction.
[0158] The target source is either a central processing unit or a hardware accelerator.
[0159] Specifically, the computing power service identifier can be defined first as whether the encoding comes from the CPU or the hardware accelerator, so as to facilitate differentiation. In actual use, since different computing power service identifier types are different in different computing systems, appropriate replacements and modifications will be made based on the actual situation.
[0160] The combination of Flit and computing power information provided in this embodiment enables transmission within the node based on the CXL protocol, allowing the computing power information to be transmitted to the DPU for subsequent computing power network processing. At the same time, based on the CXL protocol, the bandwidth and latency for obtaining computing power information are improved.
[0161] In some embodiments, encapsulating the first computing power information to obtain a second data packet corresponding to a data frame format for transmission between any two nodes includes the following steps:
[0162] Obtain the computing service identifier type and corresponding network resource information of the first computing power information;
[0163] The type information of the first computing power information is determined based on the computing power service identifier type;
[0164] Match the corresponding actual network resources based on the network resource information of the first computing power information;
[0165] Using actual network resources, primary computing power information, and type information as computing power perception information; and
[0166] The computing power perception information is encapsulated to obtain a second data packet that conforms to the data frame format transmitted between any two nodes.
[0167] Specifically, after obtaining the first computing power information, node-to-node transmission is required. This is implemented using a different protocol than CXL. Since the computing power information is carried over using the CXL protocol and transmitted to the DPU, it needs to be parsed and re-encapsulated. The CPU's first computing power information is encapsulated into a modified Flit, and the hardware accelerator's first computing power information is also encapsulated into a modified Flit; the two Flits are different. After the DPU obtains the first computing power information from the modified Flit, it disassembles it to obtain the computing power statistics information of the CPU and / or the hardware accelerator. Here, it is necessary to know the computing power service identifier type and network resource information of the first computing power information. Based on the computing power service identifier type, the type information of the first computing power information is determined, that is, whether the computing power information is transmitted by the CPU or the hardware accelerator. The network resource information is matched to obtain the actual network resources. The actual network resources, the first computing power information, and the type information are used as computing power awareness information and encapsulated in a data frame that can be transmitted between any two nodes to form the corresponding second data packet.
[0168] In some embodiments, encapsulating the computing power awareness information to obtain a second data packet corresponding to a data frame format that conforms to the transmission between any two nodes includes the following steps:
[0169] Obtain the data space of the payload data corresponding to the first data frame transmitted between any two nodes;
[0170] Obtain the data space corresponding to the computing power perception information; and
[0171] The data space of the payload data corresponding to the first data frame is compressed according to the data space corresponding to the computing power perception information, so as to encapsulate the data space corresponding to the computing power perception information within the first data frame to obtain the second data packet.
[0172] Figure 5 is a schematic diagram of a standard data frame format provided in an embodiment of this application. As shown in Figure 5, the standard data frame format requires data space compression and the addition of a computing power sensing header and computing power information data. Here, the computing power sensing header is computing power information, and the computing power information data is actual network resources. One method of compressing the original data space is to adjust the compression based on the current data length of the first computing power information, ensuring the flexibility of the data length. This allows the length of the data space that needs to be compressed to change at any time, facilitating the carrying of other data. Specifically, the data space of the payload data corresponding to the first data frame is compressed based on the data space corresponding to the computing power sensing information, so that the data space corresponding to the computing power sensing information is encapsulated within the first data frame to obtain the second data packet.
[0173] In other embodiments, the computing power awareness information is encapsulated to obtain a second data packet that conforms to the data frame format transmitted between any two nodes, including the following steps:
[0174] Obtain the data space of the payload data corresponding to the first data frame transmitted between any two nodes;
[0175] Obtain the data space corresponding to the computing power perception information;
[0176] Based on the data space corresponding to the computing power perception information, a corresponding preset data space is reserved; and
[0177] The data space of the payload data corresponding to the first data frame is compressed according to the preset data space so that the data space corresponding to the computing power perception information is encapsulated in the first data frame to obtain the second data packet.
[0178] Among them, the preset data space is greater than or equal to the data space corresponding to the computing power perception information.
[0179] Specifically, a certain amount of data space is reserved for the data corresponding to computing power perception. An estimated value is used, meaning that in most cases, this data space is greater than or equal to the data space corresponding to the actual computing power perception information. This ensures the stability of the originally compressed data space during compression and also facilitates the fixation of the space carried by subsequent data. That is, the data space of the payload data corresponding to the first data frame is compressed according to the preset data space, so that the data space corresponding to the computing power perception information can be encapsulated within the first data frame to obtain the second data packet.
[0180] Figure 6 is a schematic diagram of an improved data frame format provided in an embodiment of this application. As shown in Figure 6, compared with Figure 5, the space of the payload data has been compressed. Here, only 16 bytes and 3 bits of data length need to be added to carry the computing power data, so that the data space includes the data space corresponding to the computing power perception information and the data space of the compressed payload data, so as to facilitate the notification process of the computing power network.
[0181] In some embodiments, the structure of the first data frame includes at least a protocol header, an Internet Protocol header, a User Datagram Protocol header, an Infinite Bandwidth Protocol header, a data space, and a Cyclic Redundancy Check (CRC) code. The data space includes the data space corresponding to the computing power awareness information and the data space for the compressed payload data.
[0182] As shown in Figure 6, the data space includes the data space corresponding to the computing power perception information and the data space of the compressed payload data. This enables the carrying of computing power data so that other computing power routing nodes or servers can receive the improved data frames, parse them to obtain computing power information, and perform computing power scheduling or further processing.
[0183] In some embodiments, the protocol header is an Ethernet protocol header, and the first data frame is a data frame based on Ethernet Remote Direct Data Access technology.
[0184] Specifically, Ethernet is a networking technology that includes the protocols, ports, cables, and computer chips required to quickly transmit data via coaxial or fiber optic cables when a desktop or laptop computer is plugged into a Local Area Network (LAN). It provides a simple user interface for connecting multiple devices, including switches, routers, and personal computers (PCs). A LAN can be built with just one router and a few Ethernet connections, enabling users to communicate between all connected devices. In this embodiment, the second version of the network protocol allowing Remote Direct Memory Access over Converged Ethernet (RoCEv2) is an Ethernet-based Remote Direct Memory Access (RDMA) technology that allows for high-performance data transmission and communication over Ethernet. RoCEv2 is an improvement and extension of RoCEv1, offering higher performance, lower latency, and better compatibility. RoCEv2 allows applications to perform efficient data transfers directly between host memory without CPU intervention. It supports Remote Direct Memory (RDM) operations, including read, write, and atomic operations.
[0185] RoCEv2 is based on the Ethernet protocol stack and can run on existing Ethernet infrastructure without requiring any additional hardware or network device changes. It uses Ethernet frames for data transmission and routes and forwards data through Ethernet switches.
[0186] RoCEv2 uses User Datagram Protocol (UDP) / Internet Protocol (IP) as its transport layer protocol to provide reliable data transmission and flow control. It uses UDP ports to identify and differentiate different RDMA traffic.
[0187] RoCEv2 requires a network adapter that supports RDMA functionality, typically an Ethernet-based RDMA network card. These network cards have hardware and firmware support to implement the RDMA protocol stack and related functions.
[0188] The advent of RoCEv2 has enabled high-performance RDMA over Ethernet, providing a more flexible and scalable interconnect solution for data centers, cloud computing, and storage systems. It can integrate with existing Ethernet infrastructure and deliver performance and functionality similar to traditional InfiniBand, while reducing cost and complexity.
[0189] This embodiment provides a method for transmitting data between nodes using the Ethernet protocol. At the same time, computing power data is added to the Ethernet protocol to realize the processing of the computing power network and improve the data transmission efficiency in the application scenarios of the computing power network.
[0190] In a second aspect, embodiments of this application provide a data processing method applied to a routing node. Figure 7 is a flowchart of a data processing method applied to a routing node provided by an embodiment of this application. The routing node includes a DPU, a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU through an open interconnect protocol. The DPU is a processor that supports the open interconnect protocol and supports computing power network processing of the computing power information corresponding to the hardware accelerator and the CPU, respectively. The DPU is used to receive the computing power information of the CPU and the computing power information of the hardware accelerator.
[0191] As shown in Figure 7, this data processing method includes the following steps:
[0192] Step S21: Obtain the second data packet transmitted by the computing node;
[0193] Step S22: Parse the second data packet to extract the corresponding second computing power information; and
[0194] Step S23: Perform computing power scheduling calculations on the second computing power information to generate a routing table for computing power scheduling.
[0195] The computing node is used to acquire a first data packet, parse and extract the corresponding first computing power information from the first data packet, and encapsulate the first computing power information to obtain a second data packet that conforms to the data frame format transmitted between any two nodes. The first data packet is obtained by transmitting the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU of the computing node based on an open interconnection protocol and performing computing power network processing.
[0196] The second data packet is parsed to extract the corresponding second computing power information. This parsing and extraction process is the same as, but can be different from, the parsing and extraction process applied to the computing node described above; it is not limited here. The obtained second computing power information is then used for computing power scheduling to generate a routing table. It should be noted that the routing node in this embodiment can be used as a computing node, or it can focus on routing paths and computing power scheduling processes to facilitate subsequent computing power scheduling.
[0197] In some embodiments, the second data packet may also be obtained through computing power network processing based on the computing power information of the hardware accelerator and / or the central processing unit within the routing node itself, and this is not limited here. If it is transmitted internally, it is the same as the data processing method applied to the computing node in the above embodiments.
[0198] As a network device, the DPU can offload the network protocol stack from the CPU and send computing power information data to other computing power nodes through the network for computing power scheduling and notification. Since the encapsulation and parsing of network data are processed in the DPU, the network resource information is directly calculated and obtained by the DPU.
[0199] Similarly, other computing components that support the CXL protocol can communicate directly with the DPU at high speed via the CXL protocol, while components that do not support CXL can still communicate via the CPU to maintain sufficient device compatibility.
[0200] The data processing method provided in this application embodiment is applied to a routing node, which includes a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU, and both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU receives computing power information from the CPU and the hardware accelerator. A second data packet transmitted by the computing node is acquired. The DPU is a processor that supports the open interconnect protocol and can perform computing power network processing on the computing power information corresponding to the hardware accelerator and the CPU, respectively. The second data packet is parsed to extract the corresponding second computing power information, and the second computing power information is used for computing power scheduling calculation to generate a routing table for computing power scheduling. The computing resources (hardware accelerator) of the load are transmitted to the DPU separately via the open interconnect protocol, without needing to rely solely on a transmission path such as the CPU and network card. This allows computing resources and storage resources to each occupy a separate transmission path, improving CPU bandwidth utilization and providing more bandwidth utilization for other storage resource loads, while also saving CPU computing resources. Furthermore, data transmission based on the open interconnect protocol significantly improves the bandwidth and latency for information acquisition within the DPU. Meanwhile, the separate transmission paths of the hardware accelerator and the CPU allow components that do not support open interconnect protocols to still use the existing CPU for communication, thus ensuring sufficient device compatibility. Additionally, the data packets carry computing power information, facilitating subsequent computing power scheduling.
[0201] In some embodiments, after acquiring the second data packet, the method further includes: determining the sending source of the second data packet; in response to the sending source of the second data packet being a node other than the routing node itself, proceeding to the next round of computing power awareness process after computing power scheduling is completed.
[0202] Specifically, different types of server nodes have different DPUs that perform different functions. The second data packet received in the routing node needs to determine the data source of the second data packet, that is, the sending source. In response to the fact that the sending source of the second data packet is other computing power nodes, it is necessary to parse and extract the corresponding computing power information, that is, the second computing power information, and pass it to the computing power scheduling module inside the server to which the routing node belongs for calculation, thereby generating a routing table and completing computing power scheduling.
[0203] Although the DPU performs different functions within the routing node and the compute node, the corresponding hardware devices are exactly the same. After the computing power scheduling is completed, there is no need to inform the original compute node again, and the next round of computing power awareness process can be started directly.
[0204] In this embodiment, when a data packet originates from a data frame outside the node, the data frame is directly parsed and sent to the computing power scheduling module. After computing power scheduling is completed, there is no need to notify the original computing node again; the process will proceed to the next round of computing power awareness, thereby improving data processing efficiency.
[0205] In a third aspect, embodiments of this application also provide a distributed storage system, which includes computing nodes and routing nodes. Each computing node and routing node includes a DPU, a hardware accelerator, a memory, and a central processing unit. The memory is connected to the central processing unit. Both the central processing unit and the hardware accelerator are connected to the DPU through an open interconnect protocol. The DPU is used to receive computing power information from the central processing unit and the hardware accelerator. The DPU is a processor that supports open interconnect protocols and supports computing power network processing of the computing power information corresponding to the hardware accelerator and the central processing unit, respectively.
[0206] A computing node is used to acquire a first data packet, parse the first data packet to extract the corresponding first computing power information, and encapsulate the first computing power information to obtain a second data packet that conforms to the data frame format for transmission between any two nodes. The routing node to be transmitted is determined according to the flow type corresponding to the first computing power information so as to transmit the second data packet to the routing node. The first data packet is obtained by transmitting the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU based on the open interconnection protocol and performing computing power network processing.
[0207] The routing node is used to obtain the second data packet, parse the second data packet to extract the corresponding second computing power information, and perform computing power scheduling calculations on the second computing power information to generate a routing table for computing power scheduling.
[0208] Figure 8 is a schematic diagram of a distributed storage system provided in an embodiment of this application. As shown in Figure 8, only one computing node and one routing node are taken as examples. The hardware devices of the computing node and the routing node are the same, both including a central processing unit, a hardware accelerator, and a DPU. The DPU includes a receiving module, a sending module, and a scheduling module. The difference is that in the computing node, the main function is to parse the CXL data packets inside the node, extract the computing power information, and then encapsulate the computing power information into RoCEv2 data frames and send them out through the sending module. In the routing node, the RoCEv2 data frames are received, and the parsed computing power information is passed to the computing power scheduling module for calculation, thereby generating a routing table and completing the scheduling of computing power.
[0209] In both types of nodes, the DPU has identical hardware, but the data flow differs depending on the specific data header. Considering that the computing device may not support the CXL protocol, the computing power information can be processed by the CPU and encapsulated into an improved CXL Flit for information transmission.
[0210] Because the DPU plays different roles in the compute node and the routing node, data type determination is required before the receiving module starts working: In response to a data packet originating from a CXL Flit within the node, the CXL data packet is parsed to obtain the raw computing power information. This information is then re-encapsulated using RoCEv2 by the sending module and sent to the computing power routing node for computing power announcement. In response to a RoCEv2 data frame originating from outside the node, the data frame is directly parsed and sent to the computing power scheduling module. After computing power scheduling is completed, there is no need to inform the original compute node again; the process will proceed to the next round of computing power awareness.
[0211] Figure 9 is a flowchart of a data processing method provided in an embodiment of this application. As shown in Figure 9, the data processing method includes:
[0212] For hardware accelerators and / or central processing units:
[0213] Step S31: Determine whether open interconnection protocols are supported. If open interconnection protocols are supported, proceed to step S32. If open interconnection protocols are not supported, proceed to step S33.
[0214] Step S32: Using the transmission unit with the improved open interconnect protocol, the computing power information is directly transmitted to the receiving module through the open interconnect protocol.
[0215] Step S33: Write computing power information into system memory.
[0216] Step S34: The central processing unit collects computing power information and transmits the computing power information to the receiving module through the transmission unit of the improved open interconnection protocol.
[0217] Regarding the receiving module:
[0218] Step S35: Determine whether the received information is information within the node. If the received information is information within the node, proceed to step S36; if the received information is not information within the node, proceed to step S37.
[0219] Step S36: The server node receives and parses the computing power information and sends it to the sending module.
[0220] Step S38: The sending module encapsulates the computing power information into a data frame to announce the computing power.
[0221] Step S37: Receive computing power notification information, parse it, and send it to the scheduling module.
[0222] Step S39: The scheduling module performs computing power scheduling on the parsed computing power announcement information, and this round of computing power perception ends.
[0223] Specifically, regarding the parsing process in step S36: the DPU receives the data packet corresponding to the transmission unit of the improved open interconnection protocol, and parses the computing power information in the data packet according to the reverse encapsulation process.
[0224] Regarding the process of sending to the routing node in step S38: the DPU re-encapsulates the computing power information into the improved RoCEv2 data frame for computing power announcement.
[0225] In step S37, the process of receiving and parsing computing power announcements is as follows: The DPU receives the improved RoCEv2 data frame and parses the computing power announcement information according to the reverse process of encapsulating the improved RoCEv2 data frame.
[0226] The computing power scheduling process in step S39: The DPU performs computing power scheduling based on the computing power information.
[0227] For an introduction to the distributed storage system provided in this application, please refer to the above method embodiments. This application will not repeat the details here, as it has the same beneficial effects as the above data processing methods.
[0228] In a fourth aspect, embodiments of this application also provide a data processing method for a control device, specifically including the following steps:
[0229] Acquire computational tasks for data processing;
[0230] The control computing node transmits the computing power information of its hardware accelerator and / or central processing unit (CPU) to its DPU based on an open interconnect protocol, according to the computing task. This allows the DPU to perform computing power network processing to obtain a first data packet, parse and extract the corresponding first computing power information from the first data packet, and encapsulate the first computing power information to obtain a second data packet conforming to the data frame format for transmission between any two nodes. The control computing node then determines the routing node to be transmitted based on the flow type corresponding to the first computing power information, so that the second data packet can be transmitted to the routing node.
[0231] The control routing node obtains the second data packet; the second data packet is parsed to extract the corresponding second computing power information; the second computing power information is used to perform computing power scheduling calculations to generate a routing table for computing power scheduling.
[0232] The computing nodes and routing nodes each include a DPU, a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is used to receive computing power information from the CPU and the hardware accelerator. The DPU is a processor that supports the open interconnect protocol and can perform computing power network processing on the computing power information corresponding to the hardware accelerator and the CPU, respectively.
[0233] It is understood that the control device here operates on a unified server node, sending instructions to computing nodes and routing nodes according to computing tasks, and processing data on the computing nodes according to the computing tasks. This data processing process is the same as the data processing process applied to the computing nodes in the above embodiments. Similarly, the data processing process for controlling the routing nodes is the same as the data processing process applied to the routing nodes in the above embodiments, and the above embodiments can be referred to.
[0234] For a description of the data processing method for control equipment provided in this application, please refer to the above method embodiments. This application will not repeat the description here, as it has the same beneficial effects as the above data processing method.
[0235] In a fifth aspect, embodiments of this application also provide a centralized storage system. Figure 10 is a schematic diagram of a centralized storage system provided in an embodiment of this application. As shown in Figure 10, the system includes a control device, computing nodes, and routing nodes. Each computing node and routing node includes a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU receives computing power information from the CPU and the hardware accelerator. The DPU is a processor that supports open interconnect protocols and supports computing power network processing of the computing power information corresponding to the hardware accelerator and the CPU, respectively. This centralized storage system includes: a control device, computing nodes, and routing nodes.
[0236] Control equipment is used to acquire and process computational data.
[0237] The computing node is used to transmit the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU of the computing node based on the computing task, so that the DPU of the computing node can perform computing power network processing to obtain a first data packet, parse the first data packet to extract the corresponding first computing power information, encapsulate the first computing power information to obtain a second data packet corresponding to the data frame format for transmission between any two nodes, and determine the routing node to be transmitted according to the flow type corresponding to the first computing power information, so as to transmit the second data packet to the routing node.
[0238] The routing node is used to obtain the second data packet, parse the second data packet to extract the corresponding second computing power information, and perform computing power scheduling calculations on the second computing power information to generate a routing table for computing power scheduling.
[0239] Specifically, the system has a control device, multiple computing nodes and multiple routing nodes. In Figure 10, only one control device, one computing node and multiple routing nodes (routing node 1 and routing node 2) are used as an example. After the computing node finishes processing the data, it provides specific result information to the control device. The control device determines the corresponding routing node based on the result information to make computing power announcements.
[0240] For a description of the centralized storage system provided in this application, please refer to the above method embodiments. This application will not repeat the description here, as it has the same beneficial effects as the above data processing method.
[0241] The various embodiments corresponding to the data processing method have been described in detail above. Based on this, this application also discloses a data processing apparatus corresponding to the above method.
[0242] In a sixth aspect, Figure 11 is a structural diagram of a data processing device applied to a computing node according to an embodiment of this application. As shown in Figure 11, the computing node includes a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is a processor that supports the open interconnect protocol and supports computing power network processing of the computing power information corresponding to the hardware accelerator and the CPU, respectively. The DPU is used to receive the computing power information of the CPU and the computing power information of the hardware accelerator. The device includes the following modules:
[0243] The first acquisition module 11 is used to acquire a first data packet, wherein the first data packet is obtained by transmitting the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU based on the open interconnection protocol and performing computing power network processing.
[0244] The first parsing and extraction module 12 is used to parse and extract the first data packet to obtain the corresponding first computing power information; and to encapsulate the first computing power information to obtain a second data packet corresponding to the data frame format transmitted between any two nodes; and
[0245] The first determining module 13 is used to determine the routing node to be transmitted based on the flow type corresponding to the first computing power information, so as to transmit the second data packet to the routing node.
[0246] In a seventh aspect, Figure 12 is a structural diagram of a data processing device applied to a routing node according to an embodiment of this application. As shown in Figure 12, the routing node includes a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is a processor that supports the open interconnect protocol and supports computing power network processing of the computing power information corresponding to the hardware accelerator and the CPU, respectively. The DPU is used to receive the computing power information of the CPU and the computing power information of the hardware accelerator. The device includes the following modules:
[0247] The second acquisition module 14 is used to acquire the second data packet transmitted by the computing node. The computing node is used to acquire the first data packet, parse and extract the corresponding first computing power information from the first data packet, and encapsulate the first computing power information to obtain the second data packet corresponding to the data frame format transmitted between any two nodes. The first data packet is transmitted to the DPU based on the computing power information of the hardware accelerator and / or the computing power information of the central processing unit based on the open interconnection protocol, and is obtained by computing power network processing.
[0248] The second parsing and extraction module 15 is used to parse and extract the second data packet to obtain the corresponding second computing power information; and
[0249] The computing power scheduling module 16 is used to perform computing power scheduling calculations on the second computing power information to generate a routing table for computing power scheduling.
[0250] In an eighth aspect, FIG13 is a structural diagram of a data processing apparatus applied to a control device according to an embodiment of the present application. As shown in FIG13, the apparatus includes the following modules:
[0251] The third acquisition module 17 is used to acquire computational tasks for data processing.
[0252] The first control module 18 is used to control the computing node to transmit the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU of the computing node according to the computing task, so that the DPU of the computing node can perform computing power network processing to obtain a first data packet, parse the first data packet to extract the corresponding first computing power information, encapsulate the first computing power information to obtain a second data packet corresponding to the data frame format for transmission between any two nodes, and determine the routing node to be transmitted according to the flow type corresponding to the first computing power information, so as to transmit the second data packet to the routing node; and
[0253] The second control module 19 is used to control the routing node to obtain the second data packet, parse the second data packet to extract the corresponding second computing power information, and perform computing power scheduling calculation on the second computing power information to generate a routing table for computing power scheduling.
[0254] The computing nodes and routing nodes each include a DPU, a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is used to receive computing power information from the CPU and the hardware accelerator. The DPU is a processor that supports the open interconnect protocol and can perform computing power network processing on the computing power information corresponding to the hardware accelerator and the CPU, respectively.
[0255] Since the embodiments of the device section correspond to the embodiments described above, the embodiments of the device section are described with reference to the embodiments of the method section above, and will not be repeated here. For the description of the data processing apparatus provided in this application, please refer to the method embodiments described above; this application will not repeat the description here, as it has the same beneficial effects as the data processing method described above.
[0256] In a ninth aspect, FIG14 is a structural diagram of a data processing device provided in an embodiment of the present application. As shown in FIG14, the device includes: a memory 21 for storing computer-readable instructions, and a processor 22 for implementing the steps of the data processing method when executing the computer-readable instructions.
[0257] The data processing device provided in this embodiment may include, but is not limited to, a tablet computer, a laptop computer, or a desktop computer.
[0258] The processor 22 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 22 may be implemented using at least one hardware form selected from Digital Signal Processors (DSPs), FPGAs, and Programmable Logic Arrays (PLAs). The processor 22 may also include a main processor and a coprocessor. The main processor, also known as a CPU, is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 22 may integrate a GPU, which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 22 may also include an Artificial Intelligence (AI) processor, which handles computational operations related to machine learning.
[0259] The memory 21 may include one or more non-volatile readable storage media, which may be non-transitory. The memory 21 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory 21 is used to store at least the following computer-readable instructions 211, which, after being loaded and executed by the processor 22, can implement the relevant steps of the data processing method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 21 may also include an operating system 212 and data 213, and the storage method may be temporary storage or permanent storage. The operating system 212 may include Windows, Unix, Linux, etc. The data 213 may include, but is not limited to, the data involved in the data processing method.
[0260] In some embodiments, the data processing device may further include a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27.
[0261] Those skilled in the art will understand that the structure shown in Figure 14 does not constitute a limitation on the data processing device and may include more or fewer components than shown.
[0262] The processor 22 implements the data processing method provided in any of the above embodiments by calling instructions stored in the memory 21.
[0263] For a description of the data processing device provided in this application, please refer to the above method embodiments. This application will not repeat the description here, as it has the same beneficial effects as the above data processing method.
[0264] In a tenth aspect, embodiments of this application also provide a non-volatile computer-readable storage medium storing computer-readable instructions, which, when executed by processor 22, implement the steps of the data processing method described above.
[0265] It is understood that if the methods in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0266] For a description of the non-volatile computer-readable storage medium provided in this application, please refer to the above method embodiments. This application will not repeat the description here, as it has the same beneficial effects as the above data processing method.
[0267] In the eleventh aspect, embodiments of this application also provide a computer-readable instruction product, including computer-readable instructions that, when executed by processor 22, implement the steps of the data processing method described above.
[0268] For an introduction to the computer-readable instruction product provided in this application, please refer to the above method embodiments. This application will not repeat the details here, as it has the same beneficial effects as the above data processing method.
[0269] The foregoing has provided a detailed description of a data processing method, system, apparatus, device, medium, and product provided in this application. The various embodiments in the specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatuses disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.
[0270] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
Claims
1. A data processing method, characterized in that, This is applied to a computing node, which includes a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is a processor that supports the open interconnect protocol and performs network processing on the computing power information corresponding to the hardware accelerator and the CPU, respectively. The DPU is used to receive the computing power information from the CPU and the hardware accelerator. This includes: Acquire a first data packet; wherein the first data packet is obtained by transmitting the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU based on an open interconnection protocol and performing computing power network processing. The first data packet is parsed to extract the corresponding first computing power information; and the first computing power information is encapsulated to obtain a second data packet conforming to the data frame format transmitted between any two nodes; and The routing node to be transmitted is determined based on the flow type corresponding to the first computing power information, so as to transmit the second data packet to the routing node.
2. The data processing method according to claim 1, characterized in that, The hardware accelerator is at least one or more of a graphics processor, a field-programmable gate array (FPGA), and an application-specific integrated circuit (ASIC).
3. The data processing method according to claim 1, characterized in that, The process of transmitting the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU based on an open interconnection protocol includes: Obtain the protocol transmission unit corresponding to the open interconnection protocol; The flow type of computing power information is determined based on the computing power information of the hardware accelerator and / or the computing power information of the central processing unit; Set the flow type of computing power information to the data slot of the protocol-level message of the protocol transmission unit; The computing power information is set in the data slot of the protocol transmission unit, which is used to represent the data block corresponding to the request-response message; and Multiple configured protocol transmission units are transmitted to the DPU.
4. The data processing method according to claim 3, characterized in that, The process of determining the flow type of computing power information includes: Obtain the computing tasks corresponding to the hardware accelerators and / or the central processing unit corresponding to the computing power information; and The flow type of computing power information is determined based on the computing task; wherein, the flow type includes computing power awareness type, computing power notification type, testing type and scheduling type.
5. The data processing method according to claim 3, characterized in that, Setting computing power information within the data slots of the protocol transmission unit that represent the data blocks corresponding to request-response messages includes: The type of computing power service identifier for computing power information is determined based on its source. Obtain network resource information corresponding to computing power information; wherein, the network resource information includes at least one or more of the following: CPU utilization, memory utilization, GPU utilization, video memory utilization, disk utilization, network packet loss rate, and network bandwidth utilization; and The computing power service identifier type of the computing power information and the network resource information are saved to the data slot of the protocol transmission unit, which is used to represent the data block corresponding to the request response message.
6. The data processing method according to claim 5, characterized in that, The computing power service identifier type of the computing power information is determined based on its source, including: The target source direction for acquiring computing power information; wherein, the target source direction is the central processing unit or the hardware accelerator; and The computing power service identifier type of the computing power information is determined based on the target source direction.
7. The data processing method according to any one of claims 1 to 6, characterized in that, The first computing power information is encapsulated to obtain a second data packet that conforms to the data frame format transmitted between any two nodes, including: Obtain the computing power service identifier type and corresponding network resource information of the first computing power information; The type information of the first computing power information is determined based on the computing power service identifier type; Match the corresponding actual network resources based on the network resource information of the first computing power information; The actual network resources, the first computing power information, and the type information are used as computing power perception information; and The computing power perception information is encapsulated to obtain the second data packet, which conforms to the data frame format transmitted between any two nodes.
8. The data processing method according to claim 7, characterized in that, The computing power perception information is encapsulated to obtain a second data packet corresponding to the data frame format transmitted between any two nodes, including: Obtain the data space of the payload data corresponding to the first data frame transmitted between any two nodes; Obtain the data space corresponding to the computing power perception information; and The data space of the payload data corresponding to the first data frame is compressed according to the data space corresponding to the computing power perception information, so as to encapsulate the data space corresponding to the computing power perception information within the first data frame to obtain the second data packet.
9. The data processing method according to claim 7, characterized in that, The computing power perception information is encapsulated to obtain a second data packet corresponding to the data frame format transmitted between any two nodes, including: Obtain the data space of the payload data corresponding to the first data frame transmitted between any two nodes; Obtain the data space corresponding to the computing power perception information; A corresponding preset data space is reserved based on the data space corresponding to the computing power perception information; wherein, the preset data space is greater than or equal to the data space corresponding to the computing power perception information; and The data space of the payload data corresponding to the first data frame is compressed according to the preset data space, so as to encapsulate the data space corresponding to the computing power perception information within the first data frame to obtain the second data packet.
10. The data processing method according to claim 8, characterized in that, The structure of the first data frame includes at least a protocol header, an Internet Protocol header, a User Datagram Protocol header, an Infinite Bandwidth Protocol header, a data space, and a Cyclic Redundancy Check (CRC) code; wherein, the data space includes the data space corresponding to the computing power perception information and the data space of the compressed payload data.
11. The data processing method according to claim 10, characterized in that, The protocol header is an Ethernet protocol header, and the first data frame is a data frame based on Ethernet Remote Direct Data Access technology.
12. A data processing method, characterized in that, This is applied to a routing node, which includes a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is a processor that supports the open interconnect protocol and performs network processing on the computing power information corresponding to the hardware accelerator and the CPU, respectively. The DPU is used to receive the computing power information from the CPU and the hardware accelerator. This includes: The computing node acquires a second data packet transmitted by a computing node; wherein the computing node acquires a first data packet, parses and extracts the corresponding first computing power information from the first data packet, and encapsulates the first computing power information to obtain a second data packet corresponding to the data frame format transmitted between any two nodes, wherein the first data packet is transmitted to the DPU of the computing node based on the computing power information of the hardware accelerator and / or the computing power information of the central processing unit based on an open interconnection protocol, and is obtained by computing power network processing; The second data packet is parsed to extract the corresponding second computing power information; and The second computing power information is used to perform computing power scheduling calculations to generate a routing table for computing power scheduling.
13. The data processing method according to claim 12, characterized in that, After obtaining the second data packet, the process also includes: Determine the source of the second data packet; In response to the fact that the source of the second data packet is a node other than its own routing node, the next round of computing power awareness process will be carried out after the computing power scheduling is completed.
14. A distributed storage system, characterized in that, The distributed storage system includes computing nodes and routing nodes; each computing node and routing node includes a DPU, a hardware accelerator, a memory, and a central processing unit (CPU); the memory is connected to the CPU; both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol; the DPU is used to receive computing power information from the CPU and the hardware accelerator; the DPU is a processor that supports open interconnect protocols and supports computing power network processing of the computing power information corresponding to the hardware accelerator and the CPU respectively. The computing node is used to acquire the first data packet; The first data packet is parsed and extracted to obtain the corresponding first computing power information; The first computing power information is encapsulated to obtain a second data packet corresponding to the data frame format for transmission between any two nodes; the routing node to be transmitted is determined according to the flow type corresponding to the first computing power information, so as to transmit the second data packet to the routing node, wherein the first data packet is obtained by transmitting the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU based on an open interconnection protocol and performing computing power network processing; and The routing node is used to acquire the second data packet; parse the second data packet to extract the corresponding second computing power information; and perform computing power scheduling calculation on the second computing power information to generate a routing table for computing power scheduling.
15. A data processing method, characterized in that, Applied to control equipment, including: Acquire computational tasks for data processing; The control computing node transmits the computing power information of its hardware accelerator and / or central processing unit to its DPU based on an open interconnect protocol according to the computing task. This allows the DPU to perform computing power network processing on the computing power information to obtain a first data packet; parse the first data packet to extract corresponding first computing power information; encapsulate the first computing power information to obtain a second data packet conforming to the data frame format for transmission between any two nodes; and determine the routing node to be transmitted based on the flow type corresponding to the first computing power information, so that the second data packet can be transmitted to the routing node. The routing node is controlled to acquire the second data packet; the second data packet is parsed and extracted to obtain the corresponding second computing power information; the second computing power information is used to perform computing power scheduling calculation to generate a routing table for computing power scheduling; The computing node and the routing node each include a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. The CPU and the hardware accelerator are both connected to the DPU via an open interconnect protocol. The DPU is used to receive computing power information from the CPU and the hardware accelerator. The DPU is a processor that supports open interconnect protocols and can perform computing power network processing on the computing power information corresponding to the hardware accelerator and the CPU, respectively.
16. A centralized storage system, characterized in that, The centralized storage system includes a control device, computing nodes, and routing nodes. Each computing node and routing node includes a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is used to receive computing power information from the CPU and the hardware accelerator. The DPU is a processor that supports open interconnect protocols and performs computing power network processing on the computing power information corresponding to the hardware accelerator and the central processing unit, respectively; including: Control equipment used to acquire computational tasks for data processing; The computing node is configured to transmit the computing power information of its hardware accelerator and / or the computing power information of its central processing unit to its DPU based on an open interconnect protocol, according to the computing task. This allows the DPU to perform computing power network processing on the computing power information to obtain a first data packet; parse and extract the corresponding first computing power information from the first data packet; encapsulate the first computing power information to obtain a second data packet conforming to the data frame format for transmission between any two nodes; and determine the routing node to be transmitted based on the flow type corresponding to the first computing power information, so that the second data packet can be transmitted to the routing node. The routing node is used to acquire the second data packet; parse the second data packet to extract the corresponding second computing power information; and perform computing power scheduling calculation on the second computing power information to generate a routing table for computing power scheduling.
17. A data processing apparatus, characterized in that, This is applied to a computing node, which includes a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is a processor that supports open interconnect protocols and performs network processing on the computing power information corresponding to the hardware accelerator and the CPU, respectively. The DPU receives computing power information from the CPU and the hardware accelerator. This includes: The first acquisition module is used to acquire a first data packet; wherein the first data packet is obtained by transmitting the computing power information of the hardware accelerator and / or the computing power information of the central processing unit to the DPU based on an open interconnection protocol and performing computing power network processing. The first parsing and extraction module is used to parse and extract the first data packet to obtain the corresponding first computing power information; and to encapsulate the first computing power information to obtain a second data packet corresponding to the data frame format transmitted between any two nodes; and The first determining module is used to determine the routing node to be transmitted based on the flow type corresponding to the first computing power information, so as to transmit the second data packet to the routing node.
18. A data processing apparatus, characterized in that, This is applied to a routing node, which includes a Data Processing Unit (DPU), a hardware accelerator, a memory, and a central processing unit (CPU). The memory is connected to the CPU. Both the CPU and the hardware accelerator are connected to the DPU via an open interconnect protocol. The DPU is a processor that supports open interconnect protocols and performs network processing on the computing power information corresponding to the hardware accelerator and the CPU. The DPU receives computing power information from the CPU and the hardware accelerator. This includes: The second acquisition module is used to acquire a second data packet transmitted by a computing node; wherein, the computing node is used to acquire a first data packet, parse and extract the corresponding first computing power information from the first data packet, and encapsulate the first computing power information to obtain a second data packet corresponding to the data frame format transmitted between any two nodes, wherein the first data packet is transmitted to the DPU based on the computing power information of the hardware accelerator and / or the computing power information of the central processing unit based on an open interconnection protocol, and is obtained by computing power network processing; The second parsing and extraction module is used to parse and extract the second data packet to obtain the corresponding second computing power information; and The computing power scheduling module is used to perform computing power scheduling calculations on the second computing power information to generate a routing table for computing power scheduling.
19. A data processing device, characterized in that, include: Memory, used to store computer-readable instructions; and A processor, configured to implement the steps of the data processing method as described in any one of claims 1 to 13 or claim 15 when executing the computer-readable instructions.
20. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores computer-readable instructions that, when executed by a processor, implement the steps of the data processing method as described in any one of claims 1 to 13 or claim 15.
21. A computer-readable instruction product, comprising computer-readable instructions, characterized in that, When executed by a processor, the computer-readable instructions implement the steps of any one of claims 1 to 13 or the data processing method of claim 15.
Citation Information
Patent Citations
Computing power processing network system, service processing method and equipment
CN114095579A
Data processing system, method and controller
CN115576661A
Computing power service providing method and device and storage medium
CN116149833A
Data processing method, system, device, equipment, medium and product
CN118227343A
Method and apparatus for sending route calculation information, device, and storage medium
US20230269164A1