Large language model inference system and large language model inference method

By utilizing the collaborative work of multiple memory and computing devices in a large language model system, the problem that single hardware cannot meet the inference efficiency of large language models is solved, and more efficient model inference operations are achieved.

CN122332083APending Publication Date: 2026-07-03GIGA BYTE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GIGA BYTE TECH CO LTD
Filing Date
2025-07-18
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

The video random access memory of a single computing device cannot meet the inference requirements of large language models, resulting in low efficiency.

Method used

The first computing device detects the hardware performance information of multiple dynamic random access memories and video random access memories, allocates model data and performs inference operations, and uses multiple computing devices to perform cluster computing to improve efficiency.

Benefits of technology

It effectively improves the inference efficiency of large language models, makes up for the insufficient performance of single hardware, and supports model inference with higher parameters and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122332083A_ABST
    Figure CN122332083A_ABST
Patent Text Reader

Abstract

This invention proposes a large-scale language model inference system and method. The large-scale language model inference system includes a first computing unit. The first computing unit is used to set a model file to be inferred and select an inference framework for inferring the large-scale language model based on the model file. The first computing unit detects hardware performance information of multiple dynamic random access memories (DRAMs) and multiple video random access memories (VRAMs) and obtains multiple model loading parameters. The first computing unit allocates model data of the large-scale language model to at least one of the DRAMs and VRAMs based on the hardware performance information and the multiple model loading parameters. The first computing unit operates at least one of the DRAMs and VRAMs to perform inference operations on the large-scale language model based on at least one inference output filtering parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a data processing system, and more particularly to a large-scale language model inference system and method. Background Technology

[0002] With the continuous expansion of language model size and the increasing number of model parameters, the video random access memory (VRAM) of a single computing device is no longer sufficient to meet the inference requirements of large language models. How to effectively improve the inference efficiency (running efficiency) of large language models is an important issue in this field. Summary of the Invention

[0003] This invention provides a large-scale language model inference system and method, which can achieve highly efficient large-scale language model inference.

[0004] The large language model inference system of the present invention includes a first computing unit. The first computing unit is used to set a model file to be inferred and to select an inference framework for inferring the large language model based on the model file. The first computing unit detects hardware performance information of a plurality of dynamic random access memories (DRAMs) and a plurality of video random access memories (VRAMs) and obtains a plurality of model loading parameters. The first computing unit allocates model data of the large language model to at least one of the plurality of DRAMs and the plurality of VRAMs based on the hardware performance information and the plurality of model loading parameters. The first computing unit operates at least one of the plurality of DRAMs and the plurality of VRAMs to perform inference operations on the large language model according to at least one inference output filtering parameter.

[0005] The large language model inference method of the present invention includes the following steps: setting a model file to be inferred through a first computing device, and selecting an inference framework for inferring the large language model according to the model file to be inferred; detecting hardware performance information of multiple dynamic random access memories and multiple video random access memories through the first computing device; obtaining multiple model loading parameters through the first computing device; allocating model data of the large language model to at least one of the multiple dynamic random access memories and multiple video random access memories through the first computing device according to the hardware performance information and the multiple model loading parameters; and operating at least one of the multiple dynamic random access memories and multiple video random access memories through the first computing device to perform inference operations on the large language model according to at least one inference output filtering parameter.

[0006] Based on the above, the large language model inference system and method of the present invention can automatically allocate the model data of the large language model to at least one of multiple dynamic random access memory and multiple video random access memory, so as to effectively improve the inference efficiency of the large language model.

[0007] To make the above features and advantages of the present invention more apparent and understandable, specific embodiments are described below in conjunction with the accompanying drawings. Attached Figure Description

[0008] Figure 1 This is a schematic diagram of a large-scale language model inference system according to an embodiment of the present invention.

[0009] Figure 2 This is a method diagram of a large language model inference method according to an embodiment of the present invention.

[0010] Figure 3 This is a method diagram of a large language model inference method according to another embodiment of the present invention.

[0011] Figure 4 This is a schematic diagram of a large language model inference system connected to a terminal device according to an embodiment of the present invention.

[0012] The reference numerals in the attached figures are explained as follows:

[0013] 100: Large-scale language model inference system

[0014] 110: First arithmetic unit

[0015] 120_1~120_N: Second arithmetic unit

[0016] S210~S250, S310~S380: Steps

[0017] 400: Terminal device

[0018] 401: Two-dimensional barcode Detailed Implementation

[0019] Some embodiments of the present invention will now be described in detail with reference to the accompanying drawings. Component symbols used in the following description are considered identical or similar when they appear in different drawings. These embodiments are only a part of the present invention and do not disclose all possible implementations of the invention. More precisely, these embodiments are merely examples of the apparatus and methods within the scope of the present invention's patent application.

[0020] Figure 1 This is a schematic diagram of a large-scale language model inference system according to an embodiment of the present invention. (Reference) Figure 1The large-scale language model inference system 100 includes a first computing device 110 and multiple second computing devices 120_1 to 120_N, where N is a positive integer. In this embodiment, the first computing device 110 may be a master computer device 110, and the second computing devices 120_1 to 120_N may be multiple slave computer devices 120_1 to 120_M. In one embodiment, the large-scale language model inference system 100 may not include the second computing devices. In this embodiment, the first computing device 110 can be connected to the second computing devices 120_1 to 120_N via a Secure Shell (SSH) protocol, a twisted pair (TP) cable, or a Thunderbolt connection. In one embodiment, the first computing device 110 may perform large-scale language model inference using only its own computing resources, or the first computing device 110 may also perform large-scale language model inference using its own computing resources and at least one of the computing resources of the second computing devices 120_1 to 120_N. In another embodiment, the first computing device 110 may also be used in conjunction with the computing resources of at least one of the second computing devices 120_1 to 120_N to perform cluster computing or distributed computing.

[0021] In this embodiment, the first computing device 110 and the second computing devices 120_1 to 120_N may each include a processor and a storage device. The processor may include, for example, a Central Processing Unit (CPU) or other programmable general-purpose or special-purpose microprocessors, digital signal processors (DSPs), programmable controllers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), other similar processing devices, or combinations thereof. The storage device may include, for example, dynamic random access memory (DRAM) and video random access memory (VRAM). In this embodiment, the storage device can be used to store related algorithms, communication protocols, and application programs for implementing the steps and operations described in the embodiments of the present invention. Furthermore, the first computing device 110 and the second computing devices 120_1 to 120_N may also each include a communication interface, an input device, and other peripheral functional elements, wherein the communication interface is a connection used to implement the connections between the devices described in the embodiments of the present invention.

[0022] Figure 2 This is a method diagram of a large-scale language model inference method according to an embodiment of the present invention. (Reference) Figure 1 as well as Figure 2 The large language model inference system 100 can execute the following steps S210 to S250. In step S210, the first computing device 110 can set the desired inference model file and select an inference framework for inferring the large language model based on the desired inference model file. In this embodiment, the user can set the desired inference model file by operating the input device of the first computing device 110, and the first computing device 110 can automatically select an inference framework suitable for inferring the large language model based on the format of the desired inference model file. In one embodiment, the format of the desired inference model file can be Safetensor or GPT-generated Unified Format (GGUF).

[0023] In step S220, the first computing device 110 can detect hardware performance information of multiple dynamic random access memories (DRAMs) and multiple video random access memories (VRAMs). In this embodiment, hardware performance information refers to at least one of the following: memory capacity, memory speed, memory bandwidth, performance metrics (read / write rate / random access time), power consumption, compatibility, and application scenario. In this embodiment, when the first computing device 110 is connected to the second computing devices 120_1 to 120_N, the multiple DRAMs and multiple VRAMs are disposed within the first computing device 110 and the second computing devices 120_1 to 120_N. In one embodiment, when the first computing device 110 is not connected to any second computing device, the multiple DRAMs and multiple VRAMs are disposed within the first computing device 110.

[0024] In step S230, the first computing device 110 can obtain multiple model loading parameters. In this embodiment, the first computing device 110 can determine the multiple model loading parameters based on hardware performance information. Alternatively, in one embodiment, the first computing device 110 can also receive multiple model loading parameters via an input device. In this embodiment, the multiple model loading parameters may include GPU offload parameters and CPU thread parameters. In this embodiment, when the first computing device 110 is connected to the second computing devices 120_1 to 120_N, the first computing device 110 can determine the model loading parameters for each computing device based on the individual hardware performance information of each computing device. The first computing device 110 can determine individual model loading parameters based on the individual hardware performance information of each computing device. In one embodiment, when the first computing device 110 is not connected to any second computing device, the first computing device 110 can determine only the model loading parameters for the first computing device 110.

[0025] In step S240, the first computing device 110 may allocate the model data of a large language model to at least one of a plurality of dynamic random access memories (DRAMs) and a plurality of video random access memories (VRAMs) based on hardware performance information and multiple model loading parameters. In one embodiment, the first computing device 110 may allocate the model data of a large language model to at least one DRAM and at least one VRAM to perform model inference using at least one DRAM and at least one VRAM.

[0026] In step S250, the first computing device 110 can operate at least one of multiple dynamic random access memories and multiple video random access memories to perform inference operations on a large language model according to at least one inference output filtering parameter. In this embodiment, the at least one inference output filtering parameter may include at least one of model setting instructions, temperature parameters, and nucleus sampling (or Top-p) parameters, but the present invention is not limited thereto. In this embodiment, the model setting instructions may refer to the instructions given by the user to the large language model regarding inference requests, inference conditions, or related information or content required to guide the model to generate inferences. In this embodiment, the inference output filtering parameter filters the results inferred by the model to select the most suitable response. Furthermore, the inference output filtering parameter can fine-tune the style (i.e., randomness) of the model output. In this embodiment, the temperature parameter and the nucleus sampling parameter can be used to control the diversity and randomness of the text generated by the large language model.

[0027] Therefore, the large language model inference system and method of the present invention can automatically allocate the model data of the large language model to at least one of multiple dynamic random access memory and multiple video random access memory, so as to effectively improve the inference efficiency of the large language model.

[0028] Figure 3 Another embodiment of the large language model inference method of the present invention can be performed as follows: S310 to S380. The first computing device 110 can execute the following steps S310 to S380 to perform large language model inference. In step S310, the first computing device 110 can first detect whether it is connected to the second computing devices 120_1 to 120_N. If yes, in step S320, a connection can be automatically set between the first computing device 110 and the second computing devices 120_1 to 120_N. If no, in step S330, the first computing device 110 can select an inference framework for inferring the large language model. In step S340, the first computing device 110 can obtain output filtering parameters. In step S350, the first computing device 110 can determine whether to manually input model loading parameters. If yes, in step S360, the first computing device 110 can obtain the model loading parameters through an input device. If not, in step S370, the first computing device 110 can determine the model loading parameters based on hardware performance information. In step S380, the first computing device 110 performs model inference. Therefore, the large language model inference system and method of the present invention can also provide users with partial selection functions to design suitable large language model inference settings.

[0029] Figure 4This is a schematic diagram illustrating a large-scale language model inference system connected to a terminal device according to an embodiment of the present invention. (Reference) Figure 1 as well as Figure 4 In this embodiment, when the large language model inference system 100 completes the above... Figure 2 or Figure 3 After the configuration of the embodiment, the first computing device 110 can generate a two-dimensional barcode 401. Furthermore, the terminal device 400 can log in to the first computing device 110 by scanning the two-dimensional barcode 401 to perform inference operations on a large language model. However, in one embodiment, the first computing device 110 may also generate other connection information for the terminal device 400 to make connections.

[0030] In one embodiment, the communication link between the terminal device 400 and the large language model inference system 100 is established in a local network (e.g., a private area network). In other words, the large language model inference system 100 can be built on a local end (e.g., a server within a company), and once the large language model inference system 100 is configured, multiple users can connect via personal terminal devices (e.g., mobile phones, tablets, or laptops) to share the terminal device 400 and simultaneously use the large language model for their respective inferences.

[0031] In summary, the large language model inference system and method of the present invention can distribute the model data of a large language model to at least one of multiple dynamic random access memories and multiple video random access memories of one or more computing devices, thereby effectively improving the inference efficiency of large language models. Furthermore, the large language model inference system and method of the present invention can cascade multiple computing devices to compensate for insufficient hardware performance, and can also load part of the model into dynamic random access memory to achieve the inference of large language models with higher parameters and accuracy.

[0032] Although the present invention has been disclosed above with reference to embodiments, it is not intended to limit the present invention. Anyone skilled in the art can make some modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A large-scale language model inference system, comprising: A first computing device is used to set a desired inference model file and, based on the desired inference model file, select an inference framework for inferring a large language model. The first computing device detects hardware performance information of multiple dynamic random access memories and multiple video random access memories, and obtains multiple model loading parameters. The first computing device allocates the model data of the large language model to at least one of the multiple dynamic random access memories and the multiple video random access memories according to the hardware performance information and the multiple model loading parameters, and the first computing device operates the multiple dynamic random access memories and the multiple video random access memories to perform the inference operation of the large language model according to at least one inference output filtering parameter.

2. The large-scale language model inference system as described in claim 1 further includes: A plurality of second computing devices are coupled to the first computing device, and the plurality of dynamic random access memories and the plurality of video random access memories are disposed in the first computing device and the plurality of second computing devices.

3. The large-scale language model inference system as described in claim 2, wherein when the first computing device determines that the plurality of second computing devices are connected to the first computing device through a secure shell protocol, a twisted pair, or a Thunderbolt connection, the first computing device detects the plurality of dynamic random access memories and the plurality of video random access memories of the plurality of second computing devices to establish the hardware performance information.

4. The large language model inference system of claim 1, wherein the first computing device determines the inference framework of the large language model based on a model format of the large language model and the hardware performance information.

5. The large language model inference system as described in claim 4, wherein the model format of the large language model is a secure tensor storage format or a GPT universal format.

6. The large-scale language model inference system as described in claim 1, wherein the first computing device determines the plurality of model loading parameters based on the hardware performance information.

7. The large-scale language model inference system of claim 1, wherein the first computing device receives the plurality of model loading parameters via an input device.

8. The large-scale language model inference system as described in claim 1, wherein the plurality of model loading parameters include a graphics processor unloading parameter and a central processing unit thread parameter.

9. The large-scale language model inference system of claim 1, wherein at least one inference output filtering parameter includes at least one of a model setting instruction, a temperature parameter, and a kernel sampling parameter.

10. The large language model inference system of claim 1, wherein the first computing device is further configured to generate a two-dimensional barcode, and a terminal device logs into the first computing device by scanning the two-dimensional barcode to perform the inference operation of the large language model.

11. A method for inferring large-scale language models, comprising: A desired inference model file is set up using a first computing device, and an inference framework for inferring a large language model is selected based on the desired inference model file. The first computing device detects hardware performance information of multiple dynamic random access memories and multiple video random access memories; Multiple model loading parameters are obtained through the first computing device; The first computing device allocates the model data of the large language model to at least one of the multiple dynamic random access memories and the multiple video random access memories according to the hardware performance information and the multiple model loading parameters; as well as The first computing device operates at least one of the plurality of dynamic random access memories and the plurality of video random access memories to perform the inference operation of the large language model based on at least one inference output filtering parameter.

12. The large-scale language model inference method as described in claim 11, wherein the plurality of dynamic random access memories and the plurality of video random access memories are disposed in the first computing device and the plurality of second computing devices.

13. The large-scale language model inference method as described in claim 12, further comprising: When the first computing device determines that the plurality of second computing devices are connected to the first computing device through a secure shell protocol, a twisted pair cable or a Thunderbolt cable, the first computing device detects the plurality of dynamic random access memories and the plurality of video random access memories of the plurality of second computing devices to establish the hardware performance information.

14. The large-scale language model inference method as described in claim 11, further comprising: The first computing device determines the inference framework of the large language model based on a model format of the large language model and the hardware performance information.

15. The large language model inference method as described in claim 14, wherein the model format is a secure tensor storage format or a GPT universal format.

16. The large-scale language model inference method as described in claim 11, wherein the step of obtaining the plurality of model loading parameters includes: The first computing device determines the loading parameters of the multiple models based on the hardware performance information.

17. The large-scale language model inference method as described in claim 11, wherein the step of obtaining the plurality of model loading parameters includes: The first computing device receives the multiple model loading parameters via an input device.

18. The large language model inference method as described in claim 11, wherein the plurality of model loading parameters include a graphics processor unloading parameter and a central processing unit thread parameter.

19. The large language model inference method of claim 11, wherein at least one inference output filtering parameter includes at least one of a model setting instruction, a temperature parameter, and a kernel sampling parameter.

20. The large-scale language model inference system as described in claim 1, further comprising: The first computing device generates a two-dimensional barcode. as well as The user logs into the first computing device by scanning the two-dimensional barcode using a terminal device, in order to perform inference operations on the large language model.