An offload card with an accelerator
By installing an accelerator on the unloading card, processing requests are intelligently allocated according to data type, solving the problem of inconsistent CPU performance in cloud computing and achieving efficient control and performance improvement of CPU resources.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2022-10-17
- Publication Date
- 2026-05-19
AI Technical Summary
In cloud computing scenarios, CPUs from different manufacturers are equipped with different types of accelerators, resulting in inconsistent CPU performance. This makes it difficult to effectively control CPU resources and affects the uniformity and efficiency of cloud-native services.
An uninstallation card with an accelerator installed is provided. The uninstallation card receives data processing requests, determines whether to process them itself or send them to the accelerator based on the data type, and provides feedback on the results, thereby realizing intelligent allocation and processing of data types.
It improved CPU performance, reduced CPU stress, enabled unified control of CPU resources, and enhanced the efficiency and consistency of cloud-native services.
Smart Images

Figure CN115686836B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of data processing technology, and in particular to an unloading card with an accelerator installed. Background Technology
[0002] With the continuous development of computer technology, in order to further improve the performance of the CPU (Central Processing Unit) and reduce CPU stress, some CPU manufacturers choose to install various types of accelerators on the CPU to achieve the purpose of improving CPU performance and reducing CPU stress.
[0003] However, in cloud computing scenarios, CPUs from various manufacturers are used, and these CPUs may have different types of accelerators installed or not, resulting in inconsistent CPU performance and making it difficult to regulate CPU resources. Summary of the Invention
[0004] In view of this, embodiments of this specification provide an unloading card with an accelerator installed. One or more embodiments of this specification also relate to a data processing method, a data processing apparatus, a data processing system, a computer-readable storage medium, and a computer program, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, an uninstallation card equipped with an accelerator is provided, wherein...
[0006] The unloading card is configured to receive a data processing request, wherein the data processing request carries data to be processed.
[0007] If the data type of the data to be processed meets the conditions for unloading the card, the data to be processed is processed, and the obtained data processing result is fed back; or
[0008] If the data type of the data to be processed is determined to meet the accelerator's processing conditions, the data to be processed is sent to the accelerator, and the data processing result obtained by the accelerator is fed back.
[0009] According to a second aspect of the embodiments of this specification, a data processing method is provided, applied to an uninstallation card equipped with an accelerator, the method comprising:
[0010] Receive a data processing request, wherein the data processing request carries data to be processed;
[0011] If the data type of the data to be processed meets the conditions for unloading the card, the data to be processed is processed, and the obtained data processing result is fed back; or
[0012] If the data type of the data to be processed is determined to meet the accelerator's processing conditions, the data to be processed is sent to the accelerator, and the data processing result obtained by the accelerator is fed back.
[0013] According to a third aspect of the embodiments of this specification, a data processing system is provided, the system including a CPU, memory, and an offloading card with an accelerator installed, wherein...
[0014] The unloading card is configured to receive data processing requests, wherein the data processing requests carry data to be processed. If the data type of the data to be processed meets the processing conditions of the unloading card, the data to be processed is processed, and the obtained data processing result is fed back to the memory; or
[0015] If the data type of the data to be processed is determined to meet the accelerator processing conditions, the data to be processed is sent to the accelerator, and the data processing result obtained by the accelerator is fed back to the memory.
[0016] The CPU is configured to retrieve the data processing results from the memory.
[0017] According to a fourth aspect of the embodiments of this specification, a data processing apparatus is provided, applied to an unloading card on which an accelerator is installed, the apparatus comprising:
[0018] The receiving module is configured to receive a data processing request, wherein the data processing request carries data to be processed;
[0019] The first processing module is configured to process the data to be processed and return the obtained data processing result when it is determined that the data type of the data to be processed meets the unloading card processing conditions; or
[0020] The second processing module is configured to send the data to be processed to the accelerator and provide feedback on the data processing results obtained by the accelerator when it is determined that the data type of the data to be processed meets the accelerator processing conditions.
[0021] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of the data processing method applied to the unloading card described above.
[0022] According to a sixth aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the data processing method applied to the unloading card described above.
[0023] This specification provides an embodiment of an unloading card equipped with an accelerator, wherein the unloading card is configured to receive a data processing request, wherein the data processing request carries data to be processed; if it is determined that the data type of the data to be processed meets the processing conditions of the unloading card, the unloading card processes the data to be processed and feeds back the obtained data processing result; or if it is determined that the data type of the data to be processed meets the processing conditions of the accelerator, the unloading card sends the data to be processed to the accelerator and feeds back the data processing result obtained by the accelerator.
[0024] The uninstallation card with an accelerator installed in this manual avoids the problem of inconsistent CPU performance and difficulty in controlling CPU resources caused by different types of accelerators installed on the CPU or no accelerators installed at all. Furthermore, when the uninstallation card determines that the data processing request carries the data type of the data to be processed and meets the accelerator's processing conditions, it sends the data to be processed to the accelerator and feeds back the data processing results obtained by the accelerator, thereby improving CPU performance and reducing CPU load. Attached Figure Description
[0025] Figure 1 This is a structural diagram of a CPU solution with an accelerator provided in one embodiment of this specification;
[0026] Figure 2 This is an application diagram of a CPU solution with an accelerator provided in one embodiment of this specification;
[0027] Figure 3 This is a structural diagram of an embodiment of the unloading card + CPU solution provided in this specification;
[0028] Figure 4 This is an application diagram of an embodiment of the unloading card + CPU solution provided in this specification;
[0029] Figure 5 This is an application diagram of an unloading card with an accelerator installed, provided in one embodiment of this specification;
[0030] Figure 6 This is an application diagram of an uninstallation card provided in one embodiment of this specification;
[0031] Figure 7 This is a schematic diagram illustrating the interaction between an unloading card with an accelerator installed and a CPU, provided in one embodiment of this specification.
[0032] Figure 8 This is an application diagram of an unloading card with an accelerator installed, provided in one embodiment of this specification;
[0033] Figure 9 This is a flowchart illustrating a data processing method provided in one embodiment of this specification;
[0034] Figure 10 This is a schematic diagram of the structure of a data processing apparatus provided in one embodiment of this specification;
[0035] Figure 11 This is a schematic diagram of the structure of a data processing system provided in one embodiment of this specification. Detailed Implementation
[0036] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0037] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0038] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0039] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0040] Offload card: This is a type of chip. It is a dedicated processor designed for cloud data centers, specifically for connecting server hardware and cloud virtualization resources; it can replace the CPU as the control and acceleration center for cloud computing. In other words, the offload card can be a chip, such as a cloud-native chip, or a processor.
[0041] Cloud-native chips: From the perspective of computing chips, cloud computing has brought about entirely new application scenarios, thus placing new demands on CPUs. Cloud-native chips are dedicated chips used in cloud computing scenarios to replace CPUs. These cloud-native chips include offloading cards.
[0042] Accelerators: Hardware accelerators for various applications, used to replace inefficient software execution on the CPU, significantly improving the performance of specific applications while freeing up CPU computing power for other general-purpose workloads. These accelerators include, but are not limited to, AMX, AI accelerators, ML engines, HPC accelerators, security coprocessors, GPUs, and other accelerators.
[0043] AI accelerators are a type of specialized hardware accelerator or computer system designed to accelerate the application of artificial intelligence, especially artificial neural networks, machine vision, and machine learning.
[0044] ML engine: Machine Learning (ML) engine.
[0045] AI: refers to Artificial Intelligence.
[0046] AMX: short for Advanced Matrix Extension, is a matrix operation programming framework designed to accelerate machine learning workloads.
[0047] HPC accelerator: short for High Performance Computing accelerator. It generally refers to a high-energy computing accelerator, which is used for high-speed data processing and performing complex calculations.
[0048] Security coprocessor: An independent hardware module added outside the CPU core to handle tasks such as security key management, key generation, encryption and decryption.
[0049] GPU: refers to the graphics processing unit (GPU).
[0050] CPU die: refers to the core of the CPU, which is the most important component of the CPU.
[0051] Offloading Card: In cloud scenarios, to improve the processing speed of input / output (I / O) services, operators can offload some I / O services from the server to low-cost heterogeneous hardware. This frees up the server's Central Processing Unit (CPU) resources and improves CPU efficiency. These heterogeneous hardware components used to offload I / O data are commonly referred to as offloading cards. An offloading card can be a separate Peripheral Component Interconnect Express (PCIe) card that establishes a PCIe channel with the server. When processing I / O services offloaded to the offloading card, the server transmits data to the offloading card for processing via the PCIe channel. This PCIe channel is primarily used for I / O service communication.
[0052] IDC generally refers to Internet Data Center. An Internet Data Center (IDC) is a facility with complete equipment (including high-speed Internet access bandwidth, high-performance local area network, secure and reliable data center environment, etc.), professional management, and a comprehensive application service platform.
[0053] CPU socket: processor socket.
[0054] RAM (Random Access Memory) is also known as main memory or RAM.
[0055] DMA (Direct Memory Access): A key feature of all modern computers, it allows hardware devices of different speeds to communicate without relying on a heavy interrupt load on the CPU. Otherwise, the CPU would need to copy each piece of data from its source to a temporary register and then write it back to the new location. During this time, the CPU would be unavailable for other tasks.
[0056] IaaS (Infrastructure as a Service) refers to a service model that provides IT infrastructure as a service through a network, and charges users based on their actual usage or occupancy of resources.
[0057] RoCE, short for RDMA over Converged Ethernet, is a network protocol that allows applications to access memory remotely over Ethernet. Currently, there are two versions of RoCE: v1 and v2. RoCE v1 is a link-layer protocol that allows direct access between any two hosts within the same broadcast domain. RoCE v2, on the other hand, is an Internet-layer protocol, enabling routing functionality.
[0058] InfiniBand (abbreviated as IB) is a computer network communication standard for high-performance computing. It features extremely high throughput and extremely low latency, and is used for data interconnection between computers. InfiniBand is also used for direct or switched interconnection between servers and storage systems, as well as interconnection between storage systems.
[0059] As computer technology continues to develop, CPUs (Central Processing Units) are also constantly improving. However, limited by Moore's Law, current CPU development follows two paths: one is to continuously improve single-core performance by integrating hardware accelerators within the CPU die; the other is to increase core density, while single-core performance improvements are slower. In this context, cloud providers often use CPUs from multiple vendors, and these CPUs have varying acceleration capabilities—for example, AI engines might only be available on CPUs from specific vendors and not on other CPU platforms. Therefore, building a CPU resource pool based on CPUs with different acceleration capabilities creates a heterogeneous pool that is not very friendly to cloud-native services, making it difficult to manage CPU resources.
[0060] Furthermore, in cloud-native scenarios, current offloading cards (which can be a chip or a processor) only support general network traffic and storage traffic offloading, as well as general encryption and decryption capabilities. They have limited acceleration capabilities for AI (Artificial Intelligence), HPC (High Performance Computing), and ML (Machine Learning), and still rely on CPU computing power and accelerators on the CPU to complete these tasks.
[0061] To address the above problems, this manual provides four solutions, the first of which is a CPU solution with an accelerator.
[0062] Since current CPUs are actually a product of the development of offline data centers, which are generally for single customers and do not have diverse CPU requirements, they can choose only one CPU to support data processing. Therefore, for vertical services such as AI, ML, and HPC, using CPUs with hardware accelerators is a better option.
[0063] See Figure 1 , Figure 1 This is a schematic diagram of a CPU solution with accelerators according to one embodiment of this specification, wherein accelerator 1 and accelerator 2 are installed on the CPU. The CPU is connected to an offloading card via PCIe. The offloading card includes an offloading card control panel for controlling various operations of the offloading card, and also includes a hardware forwarding module for forwarding data to the CPU. It should be noted that the CPU also includes a core (processor core), L3 cache (Level 3 cache), and IMC (Integrated Memory Controller). The CPU is connected to RAM.
[0064] See Figure 2 , Figure 2 This is an application diagram illustrating a CPU solution with an accelerator provided in one embodiment of this specification. The CPU die is installed in a CPU socket, and an ML engine (an accelerator) is installed on the CPU. The ML engine is connected to memory, and the offloading card is also connected to memory. Based on this, data obtained by the ML engine in server A is stored in the connected memory, and the data in memory is transferred to the offloading card of server B via the offloading card. After receiving the data, the offloading card of server B stores the data in memory so that the ML engine in server B can retrieve the data sent by server A from that memory.
[0065] Based on the above Figure 1 , Figure 2 As can be seen, in this CPU solution that supports hardware accelerators, traffic data is sent and received through the network card, and then the data is directly DMA'd to the system memory. After that, the hardware accelerator (ML engine) in the CPU can directly process the data, freeing up CPU computing power.
[0066] However, the diverse needs of cloud-native customers lead to strong demands for diverse CPUs, including x86, ARM, and RISC-V architectures. These CPUs each have distinct characteristics, significant architectural differences, and varying capabilities, particularly in hardware accelerators. It's possible that CPUs from vendor A have a rich array of accelerators, while CPUs from vendors B and C have virtually none. This is especially true for RISC-V, which is still in its early stages, resulting in substantial differences in acceleration capabilities across CPU platforms. Consequently, the performance of these vertically integrated applications varies greatly across different CPU platforms, making it impossible to provide customers with a unified cloud-native service capability.
[0067] The second solution is a CPU + heterogeneous chip solution. The advantage of this solution is its excellent performance, making it suitable for complex, heavy-load vertical services. However, the disadvantage is its high implementation cost, making it unsuitable for light-load vertical services.
[0068] The third option is: unload the card + CPU solution.
[0069] Therefore, since the offloading card only has ordinary I / O traffic offloading capabilities and general encryption / decryption and decompression functions, and does not have the hardware accelerator required for vertical scenarios, this type of data can only be processed in software on the CPU, resulting in extremely low efficiency and performance.
[0070] The fourth option is a CPU solution that does not support hardware accelerators.
[0071] This solution requires the use of an external PCIe accelerator card such as a GPU. After the network card sends and receives packets, it directly DMAs the data to the system memory. Then, the CPU moves the data from the system memory to the GPU memory via PCIe for processing.
[0072] See Figure 3 , Figure 3 This is a schematic diagram of an embodiment of the offloading card + CPU solution provided in this specification. In this embodiment, the CPU does not contain an accelerator, and the CPU and the offloading card are connected via PCIe. For an explanation of the CPU and the offloading card, please refer to [link to relevant documentation]. Figure 1 The corresponding explanation will not be repeated here.
[0073] See Figure 4 , Figure 4This is an application diagram of an offload card + CPU solution provided in one embodiment of this specification. During data processing, the GPU in server A stores the data in GPU memory. The CPU moves this data from GPU memory to system memory via PCIe and sends the data to server B via the network card (i.e., the offload card). After receiving the data packet, the network card (offload card) of server B directly DMAs the data to system memory, and then the CPU moves the data from system memory to GPU memory via PCIe for processing.
[0074] Based on the shortcomings of the four solutions mentioned above, it is clear that these four solutions cannot completely solve the aforementioned technical problems. Therefore, to avoid the issue of heterogeneous CPU resource pools being unfriendly to cloud-native services and having difficulty in controlling CPU resources, there is an urgent need to provide a general cloud-native infrastructure solution to address the acceleration issues of vertical services.
[0075] This specification provides an unloading card with an accelerator installed. This specification also relates to a data processing method, a data processing device, a data processing system, a computer-readable storage medium, and a computer program, which will be described in detail in the following embodiments.
[0076] See Figure 5 , Figure 5 The illustration shows an application diagram of an unloading card equipped with an accelerator according to an embodiment of this specification. The unloading card is configured to receive a data processing request, wherein the data processing request carries data to be processed; if it is determined that the data type of the data to be processed meets the processing conditions of the unloading card, the unloading card processes the data to be processed and feeds back the obtained data processing result; or if it is determined that the data type of the data to be processed meets the processing conditions of the accelerator, the unloading card sends the data to be processed to the accelerator and feeds back the data processing result obtained by the accelerator.
[0077] It should be noted that the uninstallation card with an accelerator installed provided in this manual can be applied to all computing products in the IaaS category of cloud computing, including but not limited to: ECS (Elastic Compute Service), containers, serverless computing, microservices, etc.
[0078] The data processing request can be understood as a request that requires processing by the offloading card. For example, this data processing request could be an AI calculation request, an image rendering request, a machine learning request, an I / O traffic offloading request, or a general encryption / decryption request, etc. This manual does not specify any particular limitation. It should be noted that the offloading card provided in this manual can be a network interface card (NIC) used for sending and receiving packets; based on this, the offloading card can receive data processing requests.
[0079] Data to be processed can be understood as data that needs to be processed. For example, in the case of a data processing request that is an image rendering request, the data to be processed can be the image to be rendered; in the case of a data processing request that is an I / O traffic offloading request, the data to be processed can be the I / O traffic data that needs to be offloaded.
[0080] The data processing result can be understood as the result obtained after the accelerator or offloading card processes the data to be processed. For example, the data processing result can be a rendered image, that is, an image rendering result.
[0081] A data type can be understood as data that uniquely identifies a type of data to be processed. For example, if the data to be processed is an image to be rendered, then the data type is an image type.
[0082] An accelerator can be understood as a hardware device that reduces the computational load on the CPU and accelerates CPU computation. This includes, but is not limited to, any two types of accelerators such as artificial intelligence accelerators, machine learning accelerators, graphics processing accelerators, data security accelerators, and computational accelerators. Specifically, an artificial intelligence accelerator refers to a dedicated hardware accelerator or computer system designed to accelerate artificial intelligence applications; for example, an AI accelerator. A machine learning accelerator is an accelerator used to accelerate machine learning workloads or improve processing efficiency; examples include ML engines and AMX. A graphics processing accelerator is a microprocessor specifically designed for image and graphics-related computations; for example, a graphics processing unit (GPU). A data security accelerator is a device that handles tasks such as secure key management, key generation, encryption, and decryption; for example, a security coprocessor. A computational accelerator is an accelerator that performs high-speed data processing and complex calculations; for example, an HPC accelerator.
[0083] It should be noted that, in one embodiment provided in this specification, the data types processed by the accelerator include artificial intelligence type, machine learning type, graphics type, data security type, and data computation type. Here, artificial intelligence type can be understood as the data type corresponding to artificial intelligence data that supports artificial intelligence implementation; graphics type can be understood as various graphics and image types, such as JPG, PNG, etc. Machine learning type can be understood as the training dataset type, machine learning model type, etc., in the field of machine learning. Data computation type can be understood as the dataset type that requires large-scale data computation in the field of data computation. Data security type can be understood as data types that need to be decrypted or encrypted.
[0084] The fact that the data type meets the processing conditions of the unloading card can be understood as the data type of the data to be processed matching the data type that the unloading card can process.
[0085] Furthermore, in the embodiments provided in this specification, the offloading card can be configured with a data type determination strategy for the data to be processed. When a data processing request is received, the card can determine the corresponding data type for the data to be processed carried in the data processing request based on this data type determination strategy. Specifically, the offloading card can receive various types of data processing requests; for example, image processing requests and machine learning requests. These requests can all carry image data; in this case, how to process the data to be processed becomes a problem to be solved. Based on this, the offloading card can be pre-configured with an association between data processing requests and data types, and this association can be stored in a table. For example, there is an association between image rendering requests and graphics types. Based on this, when the offloading card receives an image rendering request, it can determine the data type of the image to be rendered carried in the image rendering request as a graphics type based on this association. Subsequently, the data to be rendered is sent to the graphics processor based on this graphics type, instead of sending the image to be rendered to other accelerators for processing. In other words, the data type of the data to be processed carried in the data processing request can be determined according to the data processing request; or it can be understood as determining the data type of the data to be processed carried in the data processing request according to the request type of the data processing request.
[0086] Specifically, the offloading card with an accelerator installed, as provided in this manual, can be configured on a server and can receive data processing requests. These requests can be sent by other servers and carry data to be processed. After receiving the data processing request, the offloading card determines whether to process the data itself or through the accelerator installed on the offloading card. Based on this, when the offloading card determines that the data type meets the offloading card's processing conditions, it will process the data using its configured processor and storage medium hardware modules and return the processed data. In other words, although the offloading card is equipped with an accelerator, it can still perform I / O traffic offloading, general encryption and decryption, and compression / decompression functions. Based on this, after receiving a data processing request carrying data to be processed, the offloading card will determine the data type of the data to be processed. If the data type matches the data type it processes, it will determine that the data to be processed is the data it needs to process. Therefore, it will process the data and feed the processing result back to the CPU, thereby reducing the CPU's processing load and improving CPU performance.
[0087] However, if it is determined that the data type of the data to be processed meets the accelerator's processing conditions, the data will be sent to the accelerator for processing. The accelerator will then process the data, provide its processing result, and return the result. Meeting the accelerator's processing conditions means that the data type of the data to be processed is consistent with the data type processed by the accelerator.
[0088] For example, the offloading card itself has I / O traffic offloading capabilities and general encryption / decryption and compression / decompression functions, which are implemented through hardware modules such as the offloading card's processor and storage media. The offloading card with an accelerator installed, as provided in this manual, allows the accelerator to be installed on the offloading card, enabling it to not only have basic capabilities such as I / O traffic offloading and general encryption / decryption and compression / decompression functions, but also to achieve other capabilities through the accelerator. For example, if the accelerator is an artificial intelligence accelerator, then the offloading card can implement artificial intelligence acceleration functions based on the artificial intelligence accelerator; if the accelerator is a graphics processing unit (GPU), then the offloading card with the GPU installed can perform graphics and image processing based on the GPU.
[0089] Based on this, when the unloading card receives an image rendering request carrying an image to be rendered, the unloading card will determine the image to be rendered based on the data type (i.e. image type) of the image to be rendered and determine that the image to be rendered needs to be processed by the graphics processor installed on the unloading card. Therefore, the unloading card sends the image to be rendered to the graphics processor, obtains the image rendering result obtained by the graphics processor after rendering the image to be rendered, and feeds back the image rendering result.
[0090] Alternatively, when the offloading card receives an IO processing request carrying IO traffic data, it will determine whether to process the IO traffic data itself based on the data type (i.e., IO traffic type) and will process the IO traffic data itself without sending it to the accelerator. The offloading card will then provide feedback on the data processing result.
[0091] In practical applications, the architecture diagram of this offloading card in a cloud computing scenario can be found here. Figure 6 , Figure 6 This is an application diagram of an unloading card provided in one embodiment of this specification, based on... Figure 6 It is known that this offloading card can connect to network interface cards (NICs), storage media, heterogeneous chips, CPUs, and GPUs. By connecting to a NIC (such as an RDAM NIC), it can accelerate network performance. By connecting to storage media (such as SSDs), it can accelerate storage performance. Simultaneously, by connecting to heterogeneous chips, CPUs, and GPUs, it can accelerate computation. In other words, by replacing a CPU-centric architecture with an offloading card, server hardware can be better utilized, more virtualization resources can be acquired, and at the software level, the operating system connected to the offloading card can more efficiently orchestrate and schedule virtualization resources. At the hardware level, the offloading card can quickly manage data center physical devices and accelerate network and storage hardware, avoiding wasted CPU computing power and enhancing network and storage performance.
[0092] It should be noted that the unloading card provided in this manual can be a chip, such as a cloud-native chip; or it can be a processor designed specifically for cloud data centers.
[0093] Based on this, by installing accelerators such as AMX accelerators, AI accelerators, ML engines, HPC accelerators, security coprocessors, or GPUs onto the offload card, after the offload card receives a data packet, it selects the target accelerator from the accelerators installed on the offload card to process the data packet. The data in this process is processed only by the target accelerator installed on the network card (offload card), and the final processing result is returned to the system memory for the CPU to perform final processing.
[0094] It should be noted that the offloading card is equipped with a processor and a memory, which are used to implement I / O traffic offloading capabilities, general encryption and decryption functions, and other capabilities inherent to the offloading card itself. It should also be noted that the processor can send the data to be processed to an accelerator or CPU when the data to be processed requires processing by the accelerator or CPU; the processor can also obtain the data processing results from the accelerator. Alternatively, in one embodiment provided in this specification, the offloading card may be equipped with a control unit, which can determine whether the data to be processed carried in the data processing request requires processing by the offloading card, accelerator, or CPU, and send the data to be processed to the processor, accelerator, or CPU of the offloading card for processing.
[0095] In one embodiment provided in this specification, there are at least two accelerators, and at least two accelerators process the same type of data.
[0096] In other words, the uninstallation card provided in this manual can accommodate at least two accelerators, and these accelerators process the same data types. For example, installing at least two graphics processors on the uninstallation card allows these processors to process image and graphics data, thereby enabling the accelerators to process the data to be processed, reducing the CPU's processing load and improving CPU performance.
[0097] In one embodiment provided in this specification, there are at least two accelerators, and the at least two accelerators process different types of data;
[0098] Accordingly, the unloading card is also configured to send the data to be processed to the target accelerator and to feed back the data processing result obtained by the target accelerator, wherein the target accelerator is one of the at least two accelerators, and the data type processed by the target accelerator is the same as the data type of the data to be processed.
[0099] Among them, at least two types of accelerators include, but are not limited to, any two types of accelerators such as artificial intelligence accelerators, machine learning accelerators, graphics processing accelerators, data security accelerators, and computing accelerators.
[0100] Specifically, the uninstallation card provided in this manual can install at least two accelerators, but the data types processed by these accelerators can be different. For example, the uninstallation card may have an image processor and an artificial intelligence accelerator installed. Therefore, when the uninstallation card requires an accelerator to process the data to be processed, it needs to determine the corresponding target accelerator based on the data type of the data to be processed. The target accelerator is one of at least two accelerators, and the data type processed by the target accelerator is consistent with the data type of the data to be processed. This allows the accelerator to process the data, reducing the processing load on the CPU and improving CPU performance.
[0101] In one embodiment of this specification, the unloading card equipped with the accelerator is further configured to determine a data storage unit corresponding to the target accelerator, wherein the data storage unit stores the data processing result obtained by the target accelerator, and the data processing result is obtained by the target accelerator processing the data to be processed; and
[0102] The data processing result is obtained from the data storage unit and then fed back.
[0103] In this context, the data storage unit can be understood as the unit corresponding to the target accelerator, used to store the data required by the target accelerator during data processing, as well as the data processing results of the target accelerator. For example, if the target accelerator is a GPU, the data storage unit can be understood as GPU memory.
[0104] For example, after sending the image to be rendered, carried in the image rendering request, to the GPU, the offloading card can obtain the rendered image from the GPU's corresponding GPU memory and feed it back to the CPU for further processing. This reduces the CPU's computational load by leveraging the target accelerator installed on the offloading card and avoids scheduling difficulties caused by inconsistent CPU acceleration capabilities.
[0105] In practical applications, the offloading card can store the data to be rendered in the GPU memory and instruct the GPU to retrieve the data from the GPU memory for rendering.
[0106] In one embodiment of this specification, the unloading card equipped with the accelerator communicates with the CPU;
[0107] The unloading card is also configured to feed back the data processing results to the CPU.
[0108] For example, the offloading card processes the data in the data processing process only through the offloading card or accelerator, and returns the final processing result to the CPU for final processing, reducing the CPU's computational burden.
[0109] It should be noted that after the accelerator is installed on the uninstallation card, the accelerator originally installed on the CPU will not be used. The accelerator originally installed on the CPU will be turned off, thereby ensuring the uniformity of CPU performance in cloud computing scenarios and avoiding scheduling difficulties caused by inconsistent CPU acceleration capabilities.
[0110] In one embodiment of this specification, the unloading card with the accelerator installed is further configured to determine the memory corresponding to the CPU and store the data processing result in the memory, so that the CPU can obtain the data processing result from the memory.
[0111] Specifically, the unloading card processes the data in the data processing process only through the unloading card or accelerator, and returns the final processing result to the system memory. The CPU can then retrieve the data processed by the unloading card or accelerator from the system memory and perform final processing on it, thus reducing the CPU's computational burden.
[0112] In one embodiment of this specification, the unloading card equipped with the accelerator, wherein,
[0113] The unloading card is further configured to determine the memory corresponding to the CPU, store the data processing result in the memory, and send the storage information of the data processing result in the memory to the CPU, so that the CPU can obtain the data processing result from the memory based on the storage information; or
[0114] The unloading card is further configured to determine the memory corresponding to the CPU and store the data processing result in a preset storage area in the memory, so that the CPU can obtain the data processing result from the preset storage area in the memory.
[0115] The storage information can be understood as the storage location of the data processing results in memory; the preset storage area can be understood as a pre-defined area in memory, which is dedicated to storing the data processing results provided by the CPU to the unloading card; the CPU can periodically check this area and obtain the newly written data processing results from this area.
[0116] Specifically, during the process of the unloading card feeding back the data processing results to the CPU, it needs to determine the memory corresponding to the CPU, store the data processing results in the memory, and send the storage information of the data processing results in the memory to the CPU. After receiving the storage information, the CPU can obtain the data processing results from the memory based on the storage information and perform subsequent processing.
[0117] Alternatively, when the unloading card is feeding back the data processing results to the CPU, it needs to determine the memory corresponding to the CPU, as well as the preset storage area in the memory that transmits data with the CPU, and store the data processing results in the preset storage area in the memory; the CPU can obtain the data processing results from the preset storage area in the memory and perform subsequent processing.
[0118] Based on this, the offloading card reduces the CPU's processing load by providing the data processing results to the CPU.
[0119] In one embodiment of this specification, the unloading card equipped with the accelerator communicates with the CPU;
[0120] The unloading card is further configured to send the data to be processed to the CPU when it is determined that the data type of the data to be processed meets the CPU processing conditions.
[0121] In practical applications, for a data type to satisfy the CPU processing conditions, it can be understood as the data type being consistent with the data type processed by the CPU. Alternatively, it can mean that the data type is different from the data type processed by the offloading card or the data type processed by the accelerator.
[0122] Specifically, the offloading card and the accelerator installed on the offloading card are used to reduce the processing pressure on the CPU. The image processing, I / O traffic, data encryption and decryption functions of the original CPU are offloaded and implemented by the offloading card and the accelerator installed on the offloading card, so that the CPU can handle more important requests, such as user requests and web (World Wide Web) requests.
[0123] Based on this, when the unloading card determines that the data to be processed needs to be processed by the CPU during the process of receiving data packets, it sends the data to be processed to the CPU, thereby ensuring the smooth operation of the CPU.
[0124] In one embodiment of this specification, the unloading card equipped with the accelerator is further configured to determine the data type of the data to be processed, and the data type processed by the at least two accelerators;
[0125] If the data types processed by the at least two accelerators match the data type of the data to be processed, it is determined that the data type of the data to be processed satisfies the accelerator processing conditions.
[0126] The accelerator that processes the data type of the data to be processed among the at least two accelerators is identified as the target accelerator.
[0127] Specifically, after receiving a data processing request carrying data to be processed, the offloading card determines the data type of the data to be processed and the data type processed by each of at least two accelerators. It then matches these two data types. If the data types processed by at least two accelerators match the data type of the data to be processed, the data type of the data to be processed is deemed to meet the accelerator's processing conditions, and the data to be processed needs to be sent to the accelerator for processing. Based on this, the accelerator that processes the data type of the data to be processed is determined from at least two accelerators and designated as the target accelerator. This ensures accurate allocation of the data to be processed to the corresponding accelerator for processing, improving data processing efficiency.
[0128] The uninstallation card with accelerators provided in this manual avoids the problem of inconsistent CPU performance and difficulty in controlling CPU resources caused by installing different types of accelerators or not installing accelerators on the CPU, by installing at least two accelerators on the uninstallation card. Furthermore, when the uninstallation card determines that the data processing request carries the data type to be processed and meets the accelerator's processing conditions, it sends the data to be processed to the target accelerator and feeds back the data processing result obtained by the target accelerator, thereby improving CPU performance and reducing CPU load.
[0129] See Figure 7 , Figure 8 , Figure 7 This diagram illustrates the interaction between an unloading card with an accelerator installed and a CPU, according to one embodiment of this specification. Figure 8 This diagram illustrates an application of an unloading card equipped with an accelerator according to one embodiment of this specification, wherein... Figure 7 For an explanation of CPU and offload card, please refer to the above explanation. Figure 1 The corresponding or relevant content in the explanation is based on Figure 7 As can be seen, the uninstallation card with accelerators installed provided in this manual can offload accelerators such as AMX, AI accelerators, ML engines, HPC accelerators, security coprocessors, and GPUs to the uninstallation card. See also Figure 7 It is known that server A and server B communicate through offload cards. The offload cards of server A and server B can communicate via RoCE and InfiniBand. Based on this, during packet transmission and reception, the offload cards process the data through at least two accelerators installed on them. This allows data processing to be completed solely on the network interface card (i.e., the offload card), and the processing results are ultimately returned to system memory for final CPU processing. This approach offers the advantage of short data links and supports different CPU platforms.
[0130] Based on the above, this solution is a method for achieving general cloud-native chip acceleration. It involves offloading some hardware accelerators from various CPU architectures to offload cards, such as AMX accelerators, AI accelerators, ML engines, HPC accelerators, security coprocessors, and GPUs (CPU manufacturers will consider embedding micro-GPUs in CPUs in the future). Cloud-native CPU chips then eliminate these hardware accelerators, thus reducing the workload on the host machine's CPU and allowing it to focus on providing high computing power. This truly achieves a data-centric computing architecture, where data is processed where it is, bringing a series of benefits, including: unifying acceleration capabilities across different CPU platforms; and eliminating CPU-centric traffic routing and resource consumption.
[0131] See Figure 9 , Figure 9 A flowchart of a data processing method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0132] Step 902: Receive a data processing request, wherein the data processing request carries data to be processed.
[0133] Step 904: If the data type of the data to be processed meets the unloading card processing conditions, process the data to be processed and return the obtained data processing results.
[0134] Step 906: If the data type of the data to be processed meets the accelerator processing conditions, the data to be processed is sent to the accelerator, and the data processing result obtained by the accelerator is fed back.
[0135] For an explanation of this data processing method, please refer to the corresponding or relevant content in the explanation of an uninstallation card with an accelerator installed above, which will not be elaborated on here.
[0136] The data processing method provided in this manual for use with an accelerator-equipped unloading card avoids the problem of inconsistent CPU performance and difficulty in controlling CPU resources caused by different types of accelerators installed on the CPU or no accelerators installed at all. Furthermore, when the unloading card determines that the data processing request carries the data type to be processed and meets the accelerator's processing conditions, it sends the data to be processed to the accelerator and feeds back the data processing results obtained by the accelerator, thereby improving CPU performance and reducing CPU stress.
[0137] The above is an illustrative scheme of a data processing method according to this embodiment. It should be noted that the technical solution of this data processing method belongs to the same concept as the technical solution of the uninstallation card with accelerator installed described above. For details not described in detail in the technical solution of the data processing method, please refer to the description of the technical solution of the uninstallation card with accelerator installed above.
[0138] Corresponding to the above method embodiments, this specification also provides data processing apparatus embodiments. Figure 10 A schematic diagram of the structure of a data processing apparatus according to one embodiment of this specification is shown. Figure 10 As shown, the device is applied to an unloading card equipped with an accelerator, and the device includes:
[0139] The receiving module 1002 is configured to receive a data processing request, wherein the data processing request carries data to be processed;
[0140] The first processing module 1004 is configured to process the data to be processed and return the obtained data processing result when it is determined that the data type of the data to be processed meets the unloading card processing conditions; or
[0141] The second processing module 1006 is configured to send the data to be processed to the accelerator and provide feedback on the data processing result obtained by the accelerator when it is determined that the data type of the data to be processed meets the accelerator processing conditions.
[0142] For an explanation of the data processing device, please refer to the corresponding or relevant content in the explanation of an unloading card equipped with an accelerator mentioned above, which will not be elaborated upon here.
[0143] The data processing device provided in this manual, applied to an unloading card with an accelerator installed, avoids the problem of inconsistent CPU performance and difficulty in controlling CPU resources caused by installing at least two types of accelerators on the unloading card. Furthermore, when the unloading card determines that the data processing request carries the data type to be processed and meets the accelerator's processing conditions, it sends the data to be processed to the accelerator and feeds back the data processing result obtained by the accelerator, thereby improving CPU performance and reducing CPU load.
[0144] The above is an illustrative scheme of a data processing device according to this embodiment. It should be noted that the technical solution of this data processing device belongs to the same concept as the above-described unloading card with an accelerator. For details not described in detail in the technical solution of the data processing device, please refer to the description of the above-described unloading card with an accelerator.
[0145] Corresponding to the above method embodiments, this specification also provides data processing system embodiments. Figure 11 A schematic diagram of the structure of a data processing system according to one embodiment of this specification is shown. Figure 11 As shown, the system includes a CPU 1102, memory 1104, and an offloading card 1108 with an accelerator 1106 installed.
[0146] The unloading card 1108 is configured to receive data processing requests, wherein the data processing requests carry data to be processed. If the data type of the data to be processed meets the processing conditions of the unloading card, the data to be processed is processed, and the obtained data processing result is fed back to the memory 1104; or
[0147] If the data type of the data to be processed is determined to meet the accelerator processing conditions, the data to be processed is sent to the accelerator 1106, and the data processing result obtained by the accelerator 1106 is fed back to the memory 1104.
[0148] The CPU 1102 is configured to retrieve the data processing result from the memory 1104.
[0149] For an explanation of this data processing system, please refer to the corresponding or relevant content in the explanation of an uninstallation card with an accelerator installed above, which will not be elaborated on here.
[0150] The data processing system provided in this manual avoids the problem of inconsistent CPU performance and difficulty in controlling CPU resources caused by installing accelerators on an unloading card. Furthermore, when the unloading card determines that the data processing request carries the data type to be processed and meets the accelerator's processing conditions, it sends the data to be processed to the accelerator and feeds back the data processing results obtained by the accelerator to the memory. This allows the CPU to obtain the data processing results from the memory, thereby improving CPU performance and reducing CPU load.
[0151] The above is an illustrative scheme of a data processing system according to this embodiment. It should be noted that the technical solution of this data processing system belongs to the same concept as the above-mentioned unloading card with an accelerator installed. For details not described in detail in the technical solution of the data processing system, please refer to the description of the above-mentioned unloading card with an accelerator installed.
[0152] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of a data processing method applied to an unloading card.
[0153] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the data processing method applied to the uninstallation card described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the data processing method applied to the uninstallation card described above.
[0154] An embodiment of this specification also provides a computer program, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the data processing method applied to the unloading card described above.
[0155] The above is an illustrative example of a computer program according to this embodiment. It should be noted that the technical solution of this computer program belongs to the same concept as the technical solution of the data processing method applied to the uninstallation card described above. Details not described in detail in the computer program's technical solution can be found in the description of the technical solution of the data processing method applied to the uninstallation card described above.
[0156] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0157] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0158] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0159] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0160] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. An uninstallation card with an accelerator installed, wherein, The unloading card is configured to receive data processing requests, wherein the data processing requests carry data to be processed. If the data type of the data to be processed meets the offloading card processing conditions, the data to be processed is processed, and the obtained data processing result is fed back. The data type processed by the offloading card includes IO traffic type, encryption / decryption type, and compression / decompression type; or If the data type of the data to be processed is determined to meet the accelerator processing conditions, the data to be processed is sent to the accelerator, and the data processing result obtained by the accelerator is fed back. The data types processed by the accelerator include artificial intelligence type, machine learning type, graphics type, data security type, and data computing type.
2. The unloading card with an accelerator installed according to claim 1, wherein, The accelerators are at least two, and at least two accelerators process the same type of data.
3. The unloading card with an accelerator installed according to claim 1, wherein, The accelerators are at least two, and the at least two accelerators process different types of data; Accordingly, the unloading card is also configured to send the data to be processed to the target accelerator and to feed back the data processing result obtained by the target accelerator, wherein the target accelerator is one of the at least two accelerators, and the data type processed by the target accelerator is the same as the data type of the data to be processed.
4. The unloading card with an accelerator installed according to claim 3, wherein, The unloading card is further configured to determine a data storage unit corresponding to the target accelerator, wherein the data storage unit stores the data processing result obtained by the target accelerator, and the data processing result is obtained by the target accelerator processing the data to be processed; and The data processing result is obtained from the data storage unit and then fed back.
5. The unloading card with an accelerator installed according to claim 1, wherein, The unloading card communicates with the CPU; The unloading card is also configured to feed back the data processing results to the CPU.
6. The unloading card with an accelerator installed according to claim 5, wherein, The unloading card is also configured to determine the memory corresponding to the CPU and store the data processing result in the memory so that the CPU can obtain the data processing result from the memory.
7. The unloading card with an accelerator installed according to claim 6, wherein, The unloading card is further configured to determine the memory corresponding to the CPU, store the data processing result in the memory, and send the storage information of the data processing result in the memory to the CPU, so that the CPU can obtain the data processing result from the memory based on the storage information; or The unloading card is further configured to determine the memory corresponding to the CPU and store the data processing result in a preset storage area in the memory, so that the CPU can obtain the data processing result from the preset storage area in the memory.
8. The unloading card with an accelerator installed according to claim 1, wherein, The unloading card communicates with the CPU; The unloading card is further configured to send the data to be processed to the CPU when it is determined that the data type of the data to be processed meets the CPU processing conditions.
9. The unloading card with an accelerator installed according to claim 3, wherein, The unloading card is also configured to determine the data type of the data to be processed, and the data type processed by the at least two accelerators; If the data types processed by the at least two accelerators match the data type of the data to be processed, it is determined that the data type of the data to be processed satisfies the accelerator processing conditions. The accelerator that processes the data type to be processed among the at least two accelerators is identified as the target accelerator.
10. A data processing method applied to an uninstallation card with an accelerator installed, the method comprising: Receive a data processing request, wherein the data processing request carries data to be processed; If the data type of the data to be processed meets the offloading card processing conditions, the data to be processed is processed, and the obtained data processing result is fed back. The data type processed by the offloading card includes IO traffic type, encryption / decryption type, and compression / decompression type; or If the data type of the data to be processed is determined to meet the accelerator processing conditions, the data to be processed is sent to the accelerator, and the data processing result obtained by the accelerator is fed back. The data types processed by the accelerator include artificial intelligence type, machine learning type, graphics type, data security type, and data computing type.
11. A data processing system, the system comprising a CPU, memory, and an offloading card with an accelerator installed, wherein, The offloading card is configured to receive data processing requests, wherein the data processing requests carry data to be processed. If the data type of the data to be processed meets the processing conditions of the offloading card, the data to be processed is processed, and the obtained data processing result is fed back to the memory. The data types processed by the offloading card include IO traffic types, encryption / decryption types, and compression / decompression types; or... If the data type of the data to be processed is determined to meet the accelerator processing conditions, the data to be processed is sent to the accelerator, and the data processing result obtained by the accelerator is fed back to the memory. The data types processed by the accelerator include artificial intelligence type, machine learning type, graphics type, data security type, and data computing type. The CPU is configured to retrieve the data processing results from the memory.