Method and system for efficiently computing large language model inference on basis of heterogeneous computing cluster

The heterogeneous computing cluster system efficiently processes large-scale language model inference by using GPUs for summarization and FPGAs or specialized accelerators for generation, addressing inefficiencies in existing technologies and achieving reduced processing time and costs.

WO2025116306A1PCT designated stage expired Publication Date: 2025-06-05HYPERACCEL CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/016562
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-10-28
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Large-scale language models face inefficiencies in processing inference operations due to the increasing length of input and output sentences, with existing accelerators not optimized for both long input and long output sentences.

Method used

A heterogeneous computing cluster system is configured, utilizing GPUs for the summarization stage and FPGAs or specialized accelerators for the generation stage, allowing for asynchronous key-value value transmission between devices to optimize processing.

Benefits of technology

This approach significantly reduces inference processing time, maximizes hardware utilization, and lowers the Total Cost of Ownership (TCO) by leveraging the optimal computational characteristics of different devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024016562_05062025_PF_FP_ABST
    Figure KR2024016562_05062025_PF_FP_ABST
Patent Text Reader

Abstract

A method and a system for efficiently computing large language model inference (LLM) on the basis of a heterogeneous computing cluster are provided. The inference system based on a heterogeneous computing cluster, according to one embodiment may comprise: a first device including a graphic processing unit (GPU) for processing a summarization stage of an LLM; and one or more second devices, each including an accelerator for a generation stage in order to process the generation stage of the LLM.
Need to check novelty before this filing date? Find Prior Art

Description

Method and system for efficiently computing large-scale language model inference based on heterogeneous computing clusters

[0001] Embodiments of the present invention relate to a method and system for efficiently computing large-scale language model inference based on a heterogeneous computing cluster.

[0002] Large Language Models (LLMs) are models that use artificial neural networks to calculate the probability distribution of natural language sentences, and are widely used in language-related tasks such as question answering and translation.

[0003] Not only is the size of large-scale language models constantly increasing, but the length of input and output sentences requested by users is also increasing, leading to increasingly longer inference processing times. While various accelerators are being developed, the problem remains that none are optimized for both long input and output sentences.

[0004] The inference operation of large-scale language models can be broadly divided into a summarization stage, which extracts key information by calculating input tokens, and a generation stage, which creates output tokens based on that information. The summarization and generation stages differ in their computational characteristics, requiring different optimal solutions and devices. The summarization stage can process input tokens in parallel, making it well-suited for GPUs (Graphics Processing Units). However, the generation stage requires sequential calculations for each output token, resulting in a high number of consecutive operations, making it ideally suited for FPGAs (Field Programmable Gate Arrays).

[0005] A method and system for efficiently computing large-scale language model inference based on a heterogeneous computing cluster can be provided.

[0006] The technical problems of the present invention are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art from the description below.

[0007] An inference system is provided, comprising: a first device including a GPU (Graphics Processing Unit) for processing a summarization stage of a large language model (LLM); and at least one second device including an accelerator for the generation stage, for processing a generation stage of the large language model.

[0008] According to one aspect, the first device may be characterized in that it processes an input token through the summary stage using the GPU to generate a key-value value, transmits the key-value value to a corresponding second device among the at least one second device, and the corresponding second device generates an output token through the generation stage using an accelerator included in the corresponding second device and the key-value value transmitted from the first device.

[0009] According to another aspect, the generation of the key-value value in the first device and the transmission of the key-value value from the first device to the corresponding second device may be performed asynchronously.

[0010] According to another aspect, the inference system may further include a third device including a CPU that converts a user's input sentence into an input token and transmits it to the first device, and converts an output token from each of the at least one second devices into an output sentence and transmits it to the corresponding user.

[0011] According to another aspect, the third device may be characterized in that it processes scheduling for selecting a second device among the at least one second device to receive the processing result of the summary stage of the first device.

[0012] According to another aspect, the first device and the at least one second device may be configured in a ratio of 1 to n, wherein n is a natural number greater than or equal to 1.

[0013] In an inference method of an inference system, a step is provided, wherein a first device included in the inference system processes an input token through a summarization stage of a Large Language Model (LLM) using a GPU (Graphics Processing Unit) included in the first device to generate a key-value value; a step, by the first device, transmitting the key-value value to a corresponding second device among a plurality of second devices further included in the inference system, wherein each of the at least one second device includes an accelerator for a generation stage of the large language model; and a step, by the corresponding second device, processing the key-value value through the generation stage using the accelerator included in the corresponding second device to generate an output token.

[0014] Specific details of other embodiments are included in the detailed description and drawings.

[0015] A method and system for efficiently computing large-scale language model inference based on a heterogeneous computing cluster can be provided.

[0016] By utilizing the characteristics of large-scale language models and the characteristics of computing devices, a heterogeneous computing cluster system can be constructed and inference operations can be accelerated.

[0017] The effects of the present invention are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description of the claims.

[0018] FIG. 1 is a diagram illustrating an example of an inference system based on a heterogeneous computing cluster according to one embodiment of the present invention.

[0019] FIG. 2 is a diagram illustrating an example of communication between a GPU node and an FPGA node in one embodiment of the present invention.

[0020] FIG. 3 is a diagram illustrating an example of a computing timeline of a heterogeneous computing cluster-based inference system according to one embodiment of the present invention.

[0021] Figure 4 is a flowchart illustrating an example of an inference method according to one embodiment of the present invention.

[0022] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below, but may be implemented in various different forms. These embodiments are provided only to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims. Like reference numerals designate like elements throughout the specification.

[0023] When one component is referred to as being "connected to" or "coupled to" another component, it includes both cases where it is directly connected or coupled to the other component, or cases where there is another component intervening therebetween. Conversely, when one component is referred to as being "directly connected to" or "directly coupled to" another component, it indicates that there is no other component intervening therebetween. "And / or" includes each and any combination of one or more of the mentioned items.

[0024] The terminology used herein is for the purpose of describing embodiments only and is not intended to limit the present invention. In this specification, the singular also includes the plural unless specifically stated otherwise. As used herein, the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components, steps, operations, and / or elements.

[0025] Although terms like "first" and "second" are used to describe various components, these components are not limited by these terms. These terms are merely used to distinguish one component from another. Therefore, it should be understood that a "first" component referred to below may also be a "second" component within the technical scope of the present invention.

[0026] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in their common sense to those of ordinary skill in the art to which the present invention pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.

[0027] In embodiments of the present invention, the input sentence and the output sentence may correspond to a 'general language' that a human can read and write, and the input token and the output token may mean an 'integer array' converted through an embedding vector so that the input sentence and the output sentence can be understood by a computer.

[0028] FIG. 1 is a diagram illustrating an example of an inference system based on a heterogeneous computing cluster according to an embodiment of the present invention. Since the size of a large-scale language model is large, a multi-device can be used as a computational device for the large-scale language model. The large-scale language model can be divided into two parts: a summarization stage and a generation stage. The summarization stage, which processes operations on input tokens, involves many parallel computations, and thus can be well processed by a multi-GPU (Graphics Processing Unit), which is a suitable device for this. Conversely, the generation stage, which processes operations on output tokens, involves sequential computations and requires many memory accesses, which reduces the efficiency of a multi-GPU. Therefore, the operations can be well processed by a multi-FPGA (Field Programmable Gate Array).

[0029] Therefore, by using GPUs and FPGAs together to form a heterogeneous computing cluster-based computational system, large-scale language model inference can be performed more efficiently than before. Here, the word "efficient" can have various meanings. It can mean "the execution time of inference is reduced," "the hardware is kept as active as possible (utilize)," and "the total cost of ownership (TCO) can be lowered." TCO can refer to the total amount of money spent on operation, including purchase and maintenance costs. In the embodiment of Fig. 1, a "node" can mean one or more servers. The number of FPGA servers may be the same as the number of CPU servers, but considering the latency of each device, if the number of FPGA servers is greater than the number of CPU servers and GPU servers, more tasks can be processed more efficiently.

[0030] The following examples illustrate a node or device that includes an FPGA to process the generation stage of a large-scale language model. However, the FPGA can be extended with an accelerator specialized for the generation stage. For example, the accelerator for processing the generation stage can be implemented as an FPGA-based accelerator, an NPU (Neural Processing Unit)-based accelerator, and / or a PIM (Processing-In-Memory)-based accelerator.

[0031] In the embodiment of FIG. 1, the heterogeneous computing cluster-based inference system (100) may largely include a CPU node (110), a GPU node (120), and multiple FPGA nodes (130).

[0032] The CPU node (110) can tokenize a user's input sentence to generate an input token, and can translate an output token of a large-scale language model into an output sentence. In addition, the CPU node (110) can perform the role of managing data transmission between devices and scheduling an FPGA node pool.

[0033] The GPU node (120) can process the summary stage using the input token received from the CPU node (110). In addition, the GPU node (120) can perform the role of reshaping and / or permuting the data so that it can be easily used by the FPGA node.

[0034] Each of the plurality of FPGA nodes (130) can process the generation stage with data received from the GPU node (120) to generate an output token. The generated output token can be converted into an output sentence through the CPU node (110) and transmitted to the user.

[0035] At this time, each node can be implemented as a different device.

[0036] FIG. 2 is a diagram illustrating an example of communication between a GPU node and an FPGA node in one embodiment of the present invention. In order for each of the plurality of FPGA nodes (130) to process the generation stage, key-value values ​​must be transmitted from the summary stage processed by the GPU node (120). A large-scale language model is composed of a plurality of decoders, although the number varies for each model, and key-value values ​​can be generated for each decoder. At this time, since the decoder operation and the communication between devices are not dependent on each other, the transmission of key-value values ​​between the GPU node (120) and each of the plurality of FPGA nodes (130) can be performed asynchronously, thereby minimizing latency loss due to communication between devices. The embodiment of FIG. 2 illustrates an example in which a key-value value generated through the summary stage in an arbitrary GPU node is asynchronously transmitted to the generation stage processed by an arbitrary FPGA node via P2P communication. In other words, operations for the summary stage of the GPU node and communication between the GPU node and the FPGA node can be performed asynchronously.

[0037] FIG. 3 is a diagram illustrating an example of a computing timeline of a heterogeneous computing cluster-based inference system according to an embodiment of the present invention. In this embodiment, the ratio of GPU nodes to FPGA nodes may be configured as 1 to n (n is a natural number greater than or equal to 1). The embodiment of FIG. 3 illustrates an example in which GPUs and FPGAs are configured as 1 to n, and the computational results (e.g., key-value values) of the summary stage for each of n FPGAs are transmitted from one GPU. However, the computational time required for the summary stage of the GPU node is shorter than the computational time required for the generation stage of the FPGA node. Therefore, if only one FPGA node is connected to one GPU node, there may be a time when the GPU node is not activated. Therefore, in order to maximize the activation of all devices within the inference system (100), it may be desirable to set n to a natural number greater than or equal to 2. However, this does not exclude the case in which n is set to 1. Therefore, the multiple FPGA nodes (130) described below are only one embodiment and do not limit the use of one FPGA node.

[0038] Figure 4 is a flowchart illustrating an example of an inference method according to an embodiment of the present invention. The inference method according to the present embodiment can be performed by the devices (CPU node (110), GPU node (120), and multiple FPGA nodes (130)) of the inference system (100) described above.

[0039] In step (410), the CPU node (110) tokenizes a sentence input by a user to generate an input token and transmits the generated input token to the GPU node (120). For example, the CPU node (110) may be implemented as a device including a CPU, and may convert an input sentence into an input token using the CPU and then transmit the converted input token to the GPU node (120).

[0040] In step (420), the GPU node (120) may process the summary stage based on the input token to generate a key-value value and then transfer the key-value value to a corresponding FPGA node among the plurality of FPGA nodes (130). The GPU node (120) may be a device including a GPU for processing the summary stage of a large-scale language model. At this time, the GPU node (120) may use the GPU to process the input token through the summary stage of the large-scale language model to generate a key-value value. As previously described, the generation of the key-value value in the GPU node (120) and the transfer of the key-value value from the GPU node (120) to the corresponding FPGA node may be performed asynchronously.

[0041] According to an embodiment, the CPU node (110) may handle scheduling for selecting an FPGA node among a plurality of FPGA nodes (130) to receive a processing result (e.g., a key-value value) of a summary stage. In other words, the FPGA node among the plurality of FPGA nodes (130) to receive the key-value value in step (420) may be determined by scheduling of the CPU node (110).

[0042] For example, when multiple users transmit multiple input sentences, the input sentences of each user can be converted into input tokens through the CPU node (110) and transmitted to the GPU node (120). The GPU node (120) can process a summary stage based on multiple input tokens for multiple users to generate key-value values. At this time, the FPGA node to which the multiple key-value values ​​for the multiple users are transmitted can be selected from among the multiple FPGA nodes (130) through the CPU node (110).

[0043] In step (430), at least one FPGA node among the plurality of FPGA nodes (130) may process the generation stage based on the key-value value to generate an output token and transmit the generated output token to the CPU node (110). The FPGA node may be a device including an FPGA for processing the generation stage of a large-scale language model. The plurality of FPGA nodes (130) may each be implemented as a plurality of devices including an FPGA. In this case, the FPGA node that has received the key-value value may process the key-value value through the generation stage of the large-scale language model using the FPGA to generate an output token.

[0044] In step (440), the CPU node (110) can translate the output token into an output sentence and transmit the translated output sentence to the user. For example, the CPU node (110) can convert the output token into an output sentence using the CPU and then transmit the converted output sentence to the user.

[0045] Thus, according to embodiments of the present invention, a method and system for efficiently computing large-scale language model inference based on a heterogeneous computing cluster can be provided. Furthermore, utilizing the characteristics of large-scale language models and the characteristics of computing devices, a heterogeneous computing cluster system can be configured and inference operations can be accelerated.

[0046] Although the embodiments of the present invention have been described with reference to the attached drawings, those skilled in the art will appreciate that the present invention can be implemented in other specific forms without altering the technical concept or essential features thereof. Therefore, the embodiments described above should be understood to be illustrative in all respects and not restrictive.

Claims

1. A first device including a GPU (Graphics Processing Unit) for processing a summarization stage of large language models (LLM); At least one second device, each device including an accelerator for the generation stage, for processing the generation stage of the large-scale language model; and A third device including a CPU that converts a user's input sentence into an input token and transmits it to the first device, and converts an output token from each of the at least one second devices into an output sentence and transmits it to the corresponding user. Including, The first device processes an input token through the summary stage using the GPU to generate a key-value value, and reshapes or permutes data including the key-value value for the accelerator and transmits the data to a corresponding second device among the at least one second device. The corresponding second device generates an output token through the generation stage using the data including the accelerator included in the corresponding second device and the rearranged or substituted key-value value transmitted from the first device, The third device processes scheduling for selecting a second device among the at least one second device to receive the key-value value generated by the first device, In order to reduce latency loss due to communication between devices, the generation of the key-value value in the first device and the transmission of the key-value value from the first device to the corresponding second device are performed asynchronously through P2P communication. An inference system characterized by .

2. In paragraph 1, The first device and the at least one second device are configured in a ratio of 1 to n, The above n is a natural number greater than or equal to 1. An inference system characterized by .

3. In paragraph 1, An inference system, wherein said at least one second device implements one server or two or more servers.

4. In the inference method of the inference system, A step in which a first device included in the above inference system processes an input token through a summarization stage of a large language model (LLM) using a GPU (Graphic Processing Unit) included in the first device to generate a key-value value; The step of the first device transmitting the key-value value to a corresponding second device among at least one second device further including the inference system, wherein each of the at least one second device includes an accelerator for a generation stage of the large-scale language model, and the first device reshapes or permutes data including the key-value value for the accelerator and transmits the data to the corresponding second device; and A step in which the corresponding second device generates an output token by processing the key-value value through the generation stage using the accelerator included in the corresponding second device. Including, A step in which a third device further included in the above inference system converts a user's input sentence into the input token and transmits it to the first device; The step of the third device converting the output token received from the corresponding second device into an output sentence and transmitting it to the corresponding user; and A step for the third device to process scheduling for selecting a second device among the at least one second device to receive the key-value value Including more, In order to reduce latency loss due to communication between devices, the generation of the key-value value in the first device and the transmission of the key-value value from the first device to the corresponding second device are performed asynchronously through P2P communication. An inference method characterized by .

Citation Information

Patent Citations

  • An automated optical inspection system based on CPU+GPU+FPGA architecture

    JP2021501330A

  • Preparinkg method for rice syrup

    KR102487807B1

  • Apparatus that crushes and sorts waste insulation containers

    KR102599595B1

  • Code partitioning for the array of devices

    US20170262567A1