Computing system, computing power distribution method, electronic device, and readable storage medium

The computing device is pooled and dynamically adjusted through the resource scheduler, which solves the problems of idle devices and waste of resources in the AI model, and achieves more efficient input information processing.

CN120256133AActive Publication Date: 2025-07-04INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510713566.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-04
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

In the prior art, AI models have problems of idle devices and waste of computing resources during the input information processing process, especially during the pre-filling and decoding stages, which require multiple reads and writes to complete the task.

Method used

Multiple computing devices are pooled through the resource scheduler to form a first computing resource pool and a third computing resource pool, which are used for the pre-filling and decoding stages respectively, and the allocation of computing devices is dynamically adjusted according to the computing power indicators of each pool to realize flexible resource scheduling.

Benefits of technology

It reduces the idle frequency of the device, reduces the waste of computing resources, and improves the efficiency of the AI model to process input information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256133A_ABST
    Figure CN120256133A_ABST
Patent Text Reader

Abstract

The invention discloses a computing system, a computing power distribution method, electronic equipment and a readable storage medium, and relates to the technical field of high-performance computing, and the computing system comprises the steps that a resource scheduler in the computing system pools a plurality of pieces of computing equipment to obtain a first computing resource pool for providing computing power for a pre-filling stage, and the third computing resource pool for providing computing power for the decoding stage, and computing equipment allocated to the first computing resource pool and the third computing resource pool from the second computing resource pool can be adjusted according to the computing power index of the first computing resource pool and / or the third computing resource pool, so that the computing power of the computing resource pools in different stages can be dynamically adjusted, and therefore, the computing power of the computing resource pools in different stages can be dynamically adjusted. The technical problems of equipment idling and large computing resource waste due to the fact that one-time input information processing task can be completed only through multiple times of reading and writing in the prior art can be solved, and the technical effects of reducing the equipment idling frequency and reducing the computing resource waste are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of high-performance computing technology, and particularly relates to a computing system, a computing power allocation method, an electronic device, and a readable storage medium. Background Art

[0002] During the inference process of an AI model (such as a large language model LLM), its prefill stage and decode stage are two separate stages (abbreviated as "PD separation").

[0003] When the input information scale of the AI model provided by the related technology is large, a large amount of time is spent on reading and writing the input information in the prefill stage. After waiting for all the input information reading and writing to be completed, the acceleration card in the decode stage starts to calculate. Even multiple read and write operations are required to complete a processing task of the input information, resulting in device idling and significant waste of computing resources. Summary of the Invention

[0004] This application provides a computing system, a computing power allocation method, an electronic device, and a readable storage medium to at least solve the problem that multiple read and write operations are required to complete a processing task of input information in the related technology, resulting in device idling and significant waste of computing resources.

[0005] This application provides a computing system, including: Multiple computing devices, which are used to provide computing power for a target AI model; A resource scheduler, which is used to pool multiple computing devices to obtain a first computing resource pool, a second computing resource pool, and a third computing resource pool; multiple computing devices in the first computing resource pool are used to provide computing power for the prefill stage of the target AI model; multiple computing devices in the third computing resource pool are used to provide computing power for the decode stage of the target AI model; The resource scheduler is further used to adjust the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool according to the computing power metrics of the first computing resource pool and the third computing resource pool during the process of the target AI model processing input information, so as to adjust the computing power of the first computing resource pool and the third computing resource pool.

[0006] This application also provides a computing power allocation method, which is applied to the resource scheduler of the computing system, including: Pool multiple computing devices of a computing system to obtain a first computing resource pool, a second computing resource pool, and a third computing resource pool; the multiple computing devices are used to provide computing power for a target AI model; the multiple computing devices in the first computing resource pool are used to provide computing power for the pre-filling stage of the target AI model; the multiple computing devices in the third computing resource pool are used to provide computing power for the decoding stage of the target AI model; During the process of the target AI model processing input information, according to the computing power metrics of the first computing resource pool and the third computing resource pool, adjust the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool, so as to adjust the computing power of the first computing resource pool and the third computing resource pool.

[0007] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above computing power allocation methods when executing the computer program.

[0008] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above computing power allocation methods are implemented.

[0009] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above computing power allocation methods are implemented.

[0010] Through this application, since the resource scheduler pools the computing devices of the target AI model, a first computing resource pool for providing computing power for the pre-filling stage of the target AI model, a third computing resource pool for providing computing power for the decoding stage of the target AI model, and a second computing resource pool for allocating its own computing devices to the first computing resource pool and the third computing resource pool are obtained. According to the computing power metrics of the first computing resource pool and / or the third computing resource pool, adjust the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool, and dynamically adjust the computing power of the first computing resource pool and the third computing resource pool. The first computing resource pool and the third computing resource pool can perform parallel computing. Therefore, the technical problem that multiple reads and writes are required to complete the processing task of input information in the related art, resulting in device idling and large computing resource waste, can be solved, and the technical effects of reducing the frequency of device idling and reducing computing resource waste can be achieved.

[0011] In addition, the computing power of the first computing resource pool is higher than that of the second computing resource pool; the computing power of the second computing resource pool is higher than that of the third computing resource pool, which is also set based on the computing power requirements in the pre-filling stage and the decoding stage. On this basis, during the process of processing input information, according to the computing power metrics of the first computing resource pool and / or the third computing resource pool, the computing devices allocated from the second computing resource pool to the first computing resource pool and to the third computing resource pool are adjusted, so as to more deeply achieve the technical effect of adaptively and dynamically allocating computing resources to the first computing resource pool and the third computing resource pool according to the actual computing power requirements, reducing waste of computing resources while improving the efficiency of the target AI model in processing input information. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 A schematic diagram of each computing resource pool obtained after pooling the computing resources of the target AI model provided by the embodiments of the present application; Figure 2 A schematic diagram after allocating the computing devices of the scalable computing resource pool to the pre-filling computing resource pool and the decoding computing resource pool provided by the embodiments of the present application; Figure 3 One of the schematic diagrams of the structure of a computing system provided by the embodiments of the present application; Figure 4 Another schematic diagram of the structure of a computing system provided by the embodiments of the present application; Figure 5 A schematic flowchart of a resource allocation method provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0015] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0016] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following further describes this application in detail with reference to the accompanying drawings and specific embodiments.

[0017] Combined with the specific application environment architecture or specific hardware architecture on which the computing system depends, the specific application environment architecture or specific hardware architecture is described herein.

[0018] The embodiment of this application provides a computing system, which includes: Multiple computing devices, which are used to provide computing power for the target AI model; A resource scheduler, which is used to pool multiple computing devices to obtain a first computing resource pool, a second computing resource pool, and a third computing resource pool; the multiple computing devices in the first computing resource pool are used to provide computing power for the pre-filling stage of the target AI model; the multiple computing devices in the third computing resource pool are used to provide computing power for the decoding stage of the target AI model; The resource scheduler is also used to adjust the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool according to the computing power metrics of the first computing resource pool and the third computing resource pool during the process of the target AI model processing input information, so as to adjust the computing power of the first computing resource pool and the third computing resource pool.

[0019] The computing device in the embodiment of this application refers to the hardware device used to process and calculate AI tasks. The computing device is responsible for performing computing operations in model training, inference (including pre-filling, encoding, decoding, etc.) and other related tasks. Generally speaking, the computing device is used to provide computing power for the target AI model.

[0020] The type of the computing device can be one or more of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Tensor Processing Unit (TPU), a Digital Signal Processor (DSP), and a Field-Programmable Gate Array (FPGA).

[0021] Computational power, also known as computing ability, is usually used to measure the processing ability of a computing device or a computing resource pool, and involves how many computing tasks or operations the system can execute per unit time.

[0022] The computing system of the embodiments of the present application further includes a resource scheduler, which refers to a hardware device capable of allocating or scheduling computing resources. In practical applications, the resource scheduler can be a CXL (Compute ExpressLink) resource scheduler, and the CXL resource scheduler can be deployed in a server.

[0023] The target AI model of the embodiments of the present application can be an AI model with serialized serial output. The target AI model can perform at least one of generation tasks such as image generation, speech generation, text generation, content parsing, and text translation based on the input information of the user, and there is no limitation thereto.

[0024] For example, for the target AI model for text generation, the target AI model can be a Large Language Model (LLM).

[0025] The target AI model of the embodiments of the present application includes a Prefill stage and a Decode stage.

[0026] The Prefill stage refers to the initialization processing stage in a generation task. In this stage, the model will generate feature data of the input information and initial KV cache data based on the given input information (such as text, question, or prompt).

[0027] The feature data includes the feature vectors of each text unit in the input information.

[0028] The parallelism of the Prefill stage can make full use of the computing power of the GPU and belongs to compute-intensive.

[0029] The initial KV cache data includes the KV vectors of each text unit. In related technologies, the query Q (Query) vector and KV vector of the text unit are obtained by linearly projecting the text representation vector of the text unit. The KV vector consists of a K (Key) vector and a V (Value) vector.

[0030] The decoding stage is a process of gradually generating the actual output content based on the feature data obtained in the pre-filling stage and the initial KV cache data. For text generation tasks, the decoding process usually generates the next word or sentence in sequence based on the previous words or sentences until a complete answer, paragraph, or text is generated.

[0031] In the decoding stage, only one token (text unit, such as a character, word, sentence, etc.) is fetched from memory, and its requirement for computing power is not that large. It is mainly limited by the memory bandwidth and belongs to memory-intensive.

[0032] To achieve dynamic allocation of computing resources without affecting the calculations in the pre-filling stage and the decoding stage, the resource scheduler in the embodiments of this application pools the computing resources of the target AI model to obtain a first computing resource pool, a second computing resource pool, and a third computing resource pool. Each of the three computing resource pools is a resource set composed of multiple computing devices.

[0033] Among them, the first computing resource pool is also called the "pre-filling computing resource pool". This first computing resource pool provides computing power for the pre-filling stage of the target AI model and is responsible for the conversion calculation of the input information of the model. Since the model is for multiple users and it is easy to have a scenario where multiple users initiate tasks, there is a large amount of input information. Therefore, this first computing resource pool usually uses acceleration card devices with stronger computing power (such as chips with more computing units), and the quantity can be controlled at a relatively small scale.

[0034] The second computing resource pool is also called the "scalable computing resource pool" with dynamic load allocation. The computing devices in this second computing resource pool can be allocated to the first computing resource pool or the third computing resource pool. That is, when the computing resources or computing power of the first computing resource pool or the third computing resource pool are insufficient, some computing devices are divided to temporarily supplement the computing resources or computing power. This second computing resource pool usually can select computing devices with moderate computing power.

[0035] The third computing resource pool is also called the "decoding computing resource pool" and is responsible for the decoding calculation of the model output. Since this stage often involves multi-batch parallel decoding, the amount of cached data read will increase exponentially. Therefore, this third computing resource pool can use acceleration card devices with weaker computing power but larger access bandwidth (such as devices with HBM high-bandwidth memory).

[0036] Generally speaking, based on the computing power requirements in the pre-filling stage and the decoding stage, the computing power of the first computing resource pool obtained after pooling is higher than that of the second computing resource pool; the computing power of the second computing resource pool is higher than that of the third computing resource pool.

[0037] The computing power of a computing device refers to the ability of a single computing device to execute computing tasks within a unit of time, which can be measured by floating-point operations per second (FLOPS). FLOPS can measure how many floating-point operations a device can execute per second. Of course, other metrics can also be used for measurement, such as the processor frequency (Clock Speed), and there is no limitation on this.

[0038] The computing power of a computing resource pool refers to the total computing power of all the computing devices included in the computing resource pool, which can be measured by the total FLOPS of all the computing devices.

[0039] The foregoing embodiments have illustrated that the target AI model can process the input information of the user, and the input information can be voice input, text input, image input, etc., and there is no limitation on this.

[0040] In the resource scheduler of the embodiments of the present application, when each computing resource pool is performing data processing, the resource scheduler will monitor the indicators of each computing resource pool, and each monitoring indicator includes but is not limited to computing power indicators, temperature data, memory indicators, etc. The computing power indicator can specifically be the computing power utilization rate, which is the ratio of the used computing power to the total computing power. The memory indicator can specifically be the memory utilization rate, which can be the ratio of the used memory to the total memory.

[0041] It should be noted that the foregoing embodiments have illustrated that the computing resources of the second computing resource pool are either allocated to the first computing resource pool or the third computing resource pool. Therefore, monitoring indicators such as the computing power utilization rate of the first computing resource pool or the third computing resource pool take into account the computing devices allocated from the second computing resource pool to the first computing resource pool or the third computing resource pool.

[0042] It can be understood that since the first computing resource pool and the third computing resource pool are usually distributed on different devices, their computing power requirements and computing power utilization rates are also different, and the computing power requirements and computing power utilization rates are also constantly changing. For example, in the initial stage of the pre-filling stage, there is a large amount of input information. In this case, there are high computing power requirements and computing power utilization rates, and the decoding stage has not started yet. In this case, the computing devices in the second computing resource pool can be all allocated to the first computing resource pool.

[0043] Next, when the first computing resource pool processes the input information during the pre-filling stage, it continuously generates the feature data and KV cache data of the input data. As the data volume decreases, the computing power demand decreases, and the computing power utilization rate also continuously decreases. After the feature data and KV cache data of the first input information are generated, the decoding stage begins, and the third computing resource pool needs to process the feature data and KV cache data. In this case, the computing power demand and computing power utilization rate of the third computing resource pool gradually increase, and it may be necessary for the second computing resource pool to allocate computing power devices to it to provide the required computing power.

[0044] Based on the above, it can be found that the target scheduler in the embodiment of the present application dynamically adjusts the computing power allocated from the second computing resource pool to the first computing resource pool or the third computing resource pool according to the computing power metrics of the first computing resource pool and / or the third computing resource pool, so as to meet the computing power required in the pre-filling stage and the computing power required in the decoding stage of the target AI model.

[0045] See Figure 1 , the embodiment of the present application provides a schematic diagram of each computing resource pool obtained by pooling the computing resources of the target AI model. Each computing resource pool includes a pre-filling computing resource pool, a scaling computing resource pool, and a decoding computing resource pool. The computing power of the pre-filling computing resource pool is higher than that of the scaling computing resource pool, and the computing power of the scaling computing resource pool is the same as that of the decoding computing resource pool.

[0046] Each type of computing resource pool includes a number of computing devices. For example, the pre-filling computing resource pool includes 12 pre-filling computing devices, the scaling computing resource pool includes 18 scaling computing devices, and the decoding computing resource pool includes 30 decoding computing devices. The computing power of the pre-filling computing device is higher than that of the scaling computing device, and the computing power of the scaling computing device is higher than that of the decoding computing resource pool.

[0047] See Figure 2 , the embodiment of the present application provides a schematic diagram after the computing devices of the scaling computing resource pool are allocated to the pre-filling computing resource pool and the decoding computing resource pool. Continuing Figure 1 the embodiment of, the pre-filling computing resource pool includes 12 pre-filling computing devices it has and 7 scaling computing devices allocated from the scaling computing resource pool, and the decoding computing resource pool includes 30 decoding computing devices it has and 11 scaling computing devices allocated from the scaling computing resource pool.

[0048] Through the computing system provided by the embodiments of the present application, since the resource scheduler in the computing system pools the computing resources of the target AI model, a first computing resource pool for providing computing power for the pre-filling stage of the target AI model, a third computing resource pool for providing computing power for the decoding stage of the target AI model, and a second computing resource pool for allocating its own resources to the first computing resource pool and the third computing resource pool are obtained. According to the computing power metrics of the first computing resource pool and / or the third computing resource pool, the computing resources allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool are adjusted, realizing dynamic adjustment of the computing power of the first computing resource pool and the third computing resource pool. The first computing resource pool and the third computing resource pool can perform parallel computing (for example, at the same time, the first computing resource pool processes the 5th to 6th input information, and the third computing resource pool processes the feature data and KV cache data of the 3rd to 4th input information). Therefore, the technical problem that multiple reads and writes are required to complete the processing task of one input information in the related art, resulting in device idling and significant computing resource waste, can be solved, and the technical effect of reducing the frequency of device idling and reducing computing resource waste can be achieved.

[0049] In some embodiments, the computing power of the first computing resource pool is higher than that of the second computing resource pool; the computing power of the second computing resource pool is higher than that of the third computing resource pool; The computing power of the computing devices in the first computing resource pool is higher than that of the computing devices in the second computing resource pool; the computing power of the computing devices in the second computing resource pool is higher than that of the computing devices in the third computing resource pool.

[0050] Since the computing volume of the first computing resource pool in the pre-filling stage is larger than that of the third computing resource pool in the decoding stage, and the second computing resource pool is responsible for dynamic allocation, that is, when the computing resources of the first computing resource pool or the third computing resource pool are insufficient, some computing resources are divided for temporary support. Therefore, the computing power of the first computing resource pool obtained after pooling is higher than that of the second computing resource pool; the computing power of the second computing resource pool is higher than that of the third computing resource pool, which meets the computing power requirements of the pre-filling stage and the decoding stage.

[0051] On this basis, during the process of processing input information, according to the computing power metrics of the first computing resource pool and / or the third computing resource pool, the computing resources allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool are adjusted, further achieving the technical effect of adaptively and dynamically allocating computing resources for the first computing resource pool and the third computing resource pool according to the actual computing power requirements. While reducing computing resource waste, the efficiency of the target AI model in processing input information is improved.

[0052] In some embodiments, the resource scheduler includes a scheduling manager; the scheduling manager is further configured to allocate each computing device in the second computing resource pool to the first computing resource pool during the initialization phase of the pre-filling phase.

[0053] Considering the promotion of long text and multimodal applications, the pre-filling phase will receive a large amount of input information at the beginning of inference. Since the decoding phase has not obtained processable feature data and KV cache data and does not require additional computing resources, all computing devices in the second computing resource pool (scalable computing resource pool) will be initialized as pre-filling computing resources during the initialization phase, so as to provide sufficient computing resources for the pre-filling phase and avoid waste of computing resources in the scalable computing resource pool.

[0054] In some embodiments, the computing power metric is the computing power utilization rate; the resource scheduler includes a scheduling manager and a topology manager; The scheduling manager is configured to determine the change amount of the computing devices allocated from the second computing resource pool to the first computing resource pool according to the actual total computing power, the original theoretical total computing power, and the computing power loss of the first computing resource pool when the computing power utilization rate of the first computing resource pool is less than the first utilization threshold but the computing power utilization rate of the third computing resource pool is greater than the second utilization threshold, or when the computing power utilization rate of the third computing resource pool is less than the second utilization threshold; The topology manager is further configured to adjust the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool according to the change amount.

[0055] After all the computing devices in the second computing resource pool are initialized and allocated to the first computing resource pool, the resource scheduler starts dynamic scheduling, that is, adjusts the computing resources allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool according to the computing power metrics of the first computing resource pool and / or the third computing resource pool. The dynamic scheduling process can be executed by the scheduling manager in the resource scheduler.

[0056] The foregoing embodiments have shown that the computing power metric can be the computing power utilization rate. This computing power utilization rate is the ratio between the used computing power and the total computing power. The higher the computing power utilization rate, the less waste of computing power resources; on the contrary, the lower the computing power utilization rate, the more waste of computing power resources.

[0057] The computing power requirements of the first computing resource pool and the third computing resource pool are constantly changing, and the computing power utilization rate is also constantly changing. Although the computing power requirement of the first computing resource pool is usually higher than that of the third computing resource pool, when the computing power utilization rate of the first computing resource pool is low (for example, the computing power utilization rate of the first computing resource pool is less than the first utilization threshold, and the first utilization threshold is a preset threshold, such as 30%), there is a large amount of computing resource waste in the first computing resource pool. And the third computing resource pool may require more computing resources due to the increase in the obtained feature data and KV cache data, that is, the computing power requirement of the third computing resource pool is large (for example, the computing power utilization rate of the third computing resource pool is greater than the second utilization threshold, and the second utilization threshold is also a preset threshold, such as 35%). In this case, it will trigger a change in the computing resources allocated by the second computing resource pool to the first computing resource pool and the third computing resource pool, that is: reducing the number of computing resources allocated to the first computing resource pool and increasing the number of computing resources allocated to the third computing resource pool.

[0058] In addition, as the third computing resource pool processes the feature data and KV cache data, its computing power requirement is also constantly decreasing, and the computing power utilization rate is also constantly decreasing (for example, the computing power utilization rate of the third computing resource pool is less than the second utilization threshold). In this case, to enable the first computing resource pool to have sufficient computing power to process subsequent input information, there will also be a change in the computing resources allocated by the second computing resource pool to the first computing resource pool and the third computing resource pool, that is: reducing the number of computing resources allocated to the third computing resource pool and increasing the number of computing resources allocated to the first computing resource pool.

[0059] Specifically, the scheduling manager can determine the change amount of the computing devices allocated from the second computing resource pool to the first computing resource pool according to the actual total computing power, the original theoretical total computing power, and the computing power loss of the first computing resource pool.

[0060] The actual total computing power of the first computing resource pool includes the sum of the actual computing power of the original computing devices in the first computing resource pool and the actual computing power of the computing devices allocated from the second computing resource pool to the first computing resource pool.

[0061] The original theoretical total computing power of the first computing resource pool is the sum of the peak computing power of the original computing devices in the first computing resource pool.

[0062] The computing power loss refers to the loss of computing power caused by communication or other interferences, which can be determined according to the computing power utilization rate, and is generally set to 50% of the peak computing power of the computing devices in the elastic computing resource pool.

[0063] The computing power difference between the actual total computing power and the theoretical total computing power of the first computing resource pool can be determined; the computing power difference is adjusted according to the computing power loss to obtain the computing power adjustment amount of the first computing resource pool; according to the ratio of the computing power adjustment amount of the first computing resource pool to the peak computing power of the second computing resource pool, the change amount of the computing devices allocated from the second computing resource pool to the first computing resource pool is determined.

[0064] The change amount can be an increase amount or a decrease amount. The topology manager can adjust the computing resources allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool according to the change amount of the computing devices allocated to the first computing resource pool.

[0065] Specifically, when the computing power utilization rate of the first computing resource pool is less than the first utilization threshold but the computing power utilization rate of the third computing resource pool is greater than the second utilization threshold, the change amount is a decrease amount. Therefore, it is necessary to reduce the computing devices allocated to the first computing resource pool and increase the computing devices allocated to the third computing resource pool, so as to reduce the idling of computing devices and the waste of computing resources.

[0066] When the computing power utilization rate of the third computing resource pool is less than the second utilization threshold, the change amount is an increase amount. Therefore, it is necessary to increase the computing devices allocated to the first computing resource pool and reduce the computing devices allocated to the third computing resource pool, so as to reduce the idling of computing devices and the waste of computing resources.

[0067] In some embodiments, the scheduling manager is further configured to determine the computing power difference between the actual total computing power of the first computing resource pool and the original theoretical total computing power of the first computing resource pool; adjust the computing power difference according to the computing power loss to obtain the computing power adjustment amount of the first computing resource pool; The scheduling manager is further configured to determine the change amount of the computing devices allocated from the second computing resource pool to the first computing resource pool according to the ratio of the computing power adjustment amount of the first computing resource pool to the peak computing power of the second computing resource pool.

[0068] Suppose any one of the computing devices in the first computing resource pool and the computing devices allocated from the second computing resource pool to the first computing resource pool is represented by i p characterizes, and the actual computing power of any one computing device is f p , then the actual total computing power of the first computing resource pool can be characterized as .

[0069] The total number of the original computing devices in the first computing resource pool in the embodiment of the present application is N p , assuming that each computing device in the first computing resource pool has the same peak computing power, and the peak computing power is F p , then the original theoretical total computing power of the first computing resource pool can be characterized as 。

[0070] Then, the computing power difference can be characterized as 。

[0071] Next, the scheduling manager can adjust the computing power difference according to the computing power loss to obtain the computing power adjustment amount of the first computing resource pool. Specifically, the sum value (algebraic sum, weighted sum, mean square sum, etc.) between the computing power loss and the computing power difference can be used as the computing power adjustment amount of the first computing resource pool.

[0072] Then, the scheduling manager can determine the ratio of the computing power adjustment amount of the first computing resource pool to the peak computing power of the second computing resource pool, and can round up or round down this ratio to obtain the change amount of the computing devices allocated from the second computing resource pool to the first computing resource pool.

[0073] See the following formula (1): (1), where n represents the change amount of the computing devices allocated from the second computing resource pool to the first computing resource pool, " " represents rounding up, C represents the computing power loss, F S represents the peak computing power of the second computing resource pool, i p represents any one of the computing devices in the first computing resource pool and the computing devices allocated from the second computing resource pool to the first computing resource pool, f p represents the actual computing power of any one of the computing devices in the first computing resource pool and the computing devices allocated from the second computing resource pool to the first computing resource pool is f p , N p represents the total number of the original computing devices in the first computing resource pool, F p represents the peak computing power of the original computing devices.

[0074] In the embodiment of the present application, according to the actual total computing power of the first computing resource pool, the original theoretical total computing power and the computing power loss of the first computing resource pool, the computing power adjustment amount of the first computing resource pool is obtained. Furthermore, according to the ratio of the computing power adjustment amount of the first computing resource pool to the peak computing power of the second computing resource pool, the change amount of the computing devices allocated from the second computing resource pool to the first computing resource pool is determined, which can realize dynamically adjusting the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool according to the computing power requirements of the first computing resource pool and the third computing resource pool, can reduce the waste of computing resources, and can also improve the computing efficiency.

[0075] In some embodiments, the resource scheduler includes a topology manager; the topology manager adjusts the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool by creating or deleting a first high-speed connection between the first computing resource and the computing devices allocated from the second computing resource pool to the first computing resource, and creating or deleting a second high-speed connection between the third computing resource and the computing devices allocated from the second computing resource pool to the third computing resource.

[0076] The topology manager (specifically, a CXL topology manager) in the embodiments of the present application does not process large amounts of characteristic data and cache data itself. As a hardware controller for managing the cluster topology, the topology manager is used to analyze and process the topology relationships after the dynamic scheduling of each resource pool. When the resource pool is dynamically scheduled, the high-speed connection relationships between nodes are updated in real time.

[0077] The high-speed connection can specifically be a CXL connection, and the CXL connection is created based on the CXL protocol. CXL is a high-speed serial protocol that allows for fast and reliable data transfer between different components within a computer system.

[0078] The high-speed connection between the first computing resource and the computing devices allocated from the second computing resource pool to the first computing resource is the first high-speed connection, and the high-speed connection between the third computing resource and the computing devices allocated from the second computing resource pool to the third computing resource is the second high-speed connection. Efficient allocation of computing resources can be achieved by creating and deleting the first high-speed connection or the second high-speed connection.

[0079] In some embodiments, the topology manager is specifically configured to, when the computing power utilization rate of the first computing resource pool is less than the first utilization threshold but the computing power utilization rate of the third computing resource pool is greater than the second utilization threshold, determine, from the second computing resource pool, each first candidate computing device having a first high-speed connection with the first computing resource pool; select a first target computing device with a change amount from the first candidate computing devices; disconnect the first high-speed connection between the first computing resource pool and the first target computing device with the change amount; and create a second high-speed connection between the first target computing device with the change amount and the third computing resource pool.

[0080] As described in the foregoing embodiments, when the computing power utilization rate of the first computing resource pool is less than the first utilization threshold but the computing power utilization rate of the third computing resource pool is greater than the second utilization threshold, the change amount of the computing devices allocated to the first computing resource pool is a reduction amount, that is, the computing devices with the change amount are recovered from the computing devices allocated from the second computing resource pool to the first computing resource pool, and the recovered computing devices are allocated to the third computing resource pool. When the computing power utilization rate of the third computing resource pool is less than the second utilization threshold, the computing devices with the change amount are recovered from the computing devices allocated from the second computing resource pool to the third computing resource pool, and the recovered computing devices are allocated to the first computing resource pool. Thus, dynamic allocation of the computing resources in the second computing resource pool can be achieved, reducing the consumption of computing resources.

[0081] In an embodiment of the present application, when recovering the computing devices with the corresponding change amount from the computing devices allocated from the second computing resource pool to the first computing resource pool, that is, the topology manager determines each first candidate computing device having a first high-speed connection with the first computing resource pool from the second computing resource pool; selects the first target computing device with the corresponding change amount from the first candidate computing devices; and disconnects the first high-speed connection between the first computing resource pool and the first target computing device with the corresponding change amount.

[0082] On the contrary, when allocating the recovered computing devices to the third computing resource pool, that is, the topology manager creates a second high-speed connection between the computing devices with the change amount and the third computing resource pool.

[0083] In some embodiments, the topology manager is specifically configured to, when the computing power utilization rate of the third computing resource pool is less than the second utilization threshold, determine each second candidate computing device having a second high-speed connection with the third computing resource pool from the second computing resource pool; select the second target computing device with the corresponding change amount from the second candidate computing devices; disconnect the second high-speed connection between the third computing resource pool and the second target computing device with the change amount; and create a first high-speed connection between the second target computing device with the change amount and the first computing resource pool.

[0084] Similarly, when the topology manager recovers the computing devices with the change amount from the computing devices allocated from the second computing resource pool to the third computing resource pool, that is, it determines each second candidate computing device having a second high-speed connection with the third computing resource pool from the second computing resource pool; selects the second target computing device with the change amount from the second candidate computing devices; and disconnects the second high-speed connection between the third computing resource pool and the second target computing device with the change amount.

[0085] In addition, the topology manager also needs to allocate the recovered computing devices to the first computing resource pool, that is, create a second high-speed connection between the computing devices with the corresponding change amount and the first computing resource pool.

[0086] In the embodiments of the present application, by disconnecting or creating high-speed connections between computing devices in a computing resource pool, it is possible to dynamically increase or decrease the computing resources in the computing resource pool, achieve adaptive adjustment of computing resources, and have high flexibility.

[0087] In some embodiments, the computing power metric is the computing power utilization rate; the scheduling manager is further configured to maintain the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool when the computing power utilization rate of the first computing resource pool is greater than the first utilization threshold, or when the computing power utilization rate of the third computing resource pool is greater than the second utilization threshold.

[0088] It should be noted that when the computing power utilization rate of the first computing resource pool is greater than the first utilization threshold, it indicates that the computing volume of the first computing resource pool is relatively large. In this case, the computing resources allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool are maintained.

[0089] In addition, when the computing power utilization rate of the third computing resource pool is greater than the second utilization threshold, it indicates that the computing volume of the third computing resource pool is relatively large. In this case, the computing resources allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool are also maintained, thereby reducing waste of computing resources and ensuring system stability.

[0090] In some embodiments, the computing system further includes a memory expansion resource pool; the memory expansion resource pool is obtained after the resource manager pools and expands the memory resources of the target AI model; The computing devices in the first computing resource pool, the second computing resource pool, and the third computing resource pool share the memory expansion resource pool.

[0091] In the embodiments of the present application, the resource scheduler expands and pools the memory resources of the target AI model based on the target protocol to obtain a memory expansion resource pool. The target protocol can be the CXL protocol. The memory expansion resource pool is responsible for the unified access of cached data such as KV cache in the follow-up. The memory expansion resource pool can be composed of CXL type 3 devices.

[0092] In the embodiments of the present application, expanding and pooling the memory resources can meet the greater data processing requirements of the target AI model, especially when processing large-scale data, and avoid performance bottlenecks caused by insufficient memory.

[0093] In addition, in the embodiments of the present application, a shared memory expansion resource pool is set up for each of the first computing resource pool, the second computing resource pool, and the third computing resource pool. The memory resources can be flexibly allocated to different computing resource pools to meet different task requirements, enabling more efficient use of memory resources and avoiding memory surplus or shortage in a certain computing resource pool. Dynamically adjusting memory allocation helps avoid resource waste and improve the overall memory utilization rate of the system.

[0094] In some embodiments, the input information includes multiple text units; the resource scheduler further includes a data exchange module; The first computing resource pool is used to process the input information during the pre-filling stage of each input information, and obtain and output the feature data and initial KV cache data of each input information to the data exchange module; The data exchange module is used to store the feature data and initial KV cache data of each input information into the memory expansion resource pool; The third computing resource pool is used to obtain and process the feature data and the KV cache data of the previous round from the memory expansion resource pool at least once during the decoding stage of each input information, with the goal of updating the KV cache data in the memory expansion resource pool to the target KV cache data; the target KV cache data is used to generate the prediction result of the target AI model; the KV cache data of the first round is the initial KV cache data.

[0095] The feature data includes the feature vectors of each text unit; the initial KV cache data includes the KV vectors of each text unit; The third computing resource pool is used for each round of processing in the decoding stage to process the text vector of a corresponding text unit in the corresponding round and the KV cache data of the previous stage read from the memory expansion resource pool, and obtain and output the incremental KV vector of the corresponding round to the data exchange module; The data exchange module is used to store the incremental KV vector of the corresponding round into the memory expansion resource pool to update the KV cache data in the memory expansion resource pool to obtain the KV cache data of the corresponding round; the KV cache data of the last round is the target KV cache data.

[0096] See Figure 3 , one of the structural schematic diagrams of a computing system provided by the embodiments of the present application includes a first computing resource pool, a second computing resource pool, a third computing resource pool, a memory expansion resource pool, and a resource scheduler.

[0097] Among them, the first computing resource pool includes multiple pre-filling computing devices, for example, 12, and the computing power of the first computing resource pool is relatively large.

[0098] The second computing resource pool includes multiple scalable computing devices, such as 18, and the computing power of the second computing resource pool is moderate.

[0099] The third computing resource pool includes multiple decoding computing devices, such as 30. The computing power of the third computing resource pool is poor, but the access bandwidth is large.

[0100] The first computing resource pool, the second computing resource pool, and the third computing resource pool can interact with the memory expansion resource pool through a resource scheduler.

[0101] Continue Figure 3 , Figure 3 The numbers (1)-(5) of each step are marked, and the interaction process includes the following steps: (1): In the pre-filling stage, the computing devices in the first computing resource pool (including the scalable computing devices allocated from the second computing resource pool to the first computing resource pool) will process the input information to obtain the feature data of the input information and the initial KV cache data.

[0102] (2): The resource scheduler will store the feature data and the initial KV cache data in the memory expansion resource pool through the CXL.mem protocol.

[0103] (3): The resource scheduler will also determine the computing devices allocated from the second computing resource pool to the third computing resource pool according to the computing power metrics of the first computing resource pool and / or the computing power metrics of the third computing resource pool. The computing devices in the third computing resource pool and the pre-filling can obtain the feature data and the KV cache data through the CXL.cache protocol.

[0104] (4): The computing devices allocated from the second computing resource pool to the third computing resource pool, and the computing devices in the third computing resource pool obtain the feature data and the KV cache data through the CXL.cache protocol.

[0105] (5): In each round of processing in the decoding stage, the resource scheduler receives the text vector of a corresponding text unit in this sub-round from the computing devices in the third computing resource pool and the incremental KV vector of the corresponding round obtained after processing the KV cache data of the previous stage read from the memory expansion resource pool, and stores the incremental KV vector of the corresponding round in the memory expansion resource pool to update the KV cache data in the memory expansion resource pool to obtain the KV cache data of the corresponding round; the KV cache data of the first round is the initial KV cache data Repeat the above (4) and (5) until the decoding stage ends (the decoding calculation process outputs an end symbol).

[0106] The KV cache data obtained in the last round can be used as the target KV cache data. The target KV cache data can be output, and after the output of the target KV cache data, the target KV cache data is deleted from the memory expansion resource pool.

[0107] When the first computing resource in the embodiment of the present application processes the input information, it can obtain temporarily added computing resources from the second computing resource. When the third computing resource pool processes the feature data and KV cache data, it can also obtain temporarily added computing resources from the second computing resource pool, so that there are sufficient computing resources in both the pre-filling stage and the decoding stage.

[0108] In addition, the first computing resource pool, the second computing resource pool, and the third computing resource pool in the embodiment of the present application share the feature data and KV cache data in the memory expansion resource pool. The resource pools communicate through the CXL protocol, with low latency, improved communication efficiency, and the memory expansion resource pool can store more data and support higher-concurrency computing tasks.

[0109] In some embodiments, the computing device includes device memory; In the decoding stage, when the computing device allocated from the second computing resource pool to the third computing resource pool remains unchanged, for any round of processing in the decoding stage, the feature vector of the text unit in the corresponding round is obtained by the computing device of the third computing resource pool from its own device memory; the feature data in the device memory is obtained from and cached in the memory expansion resource pool; When the computing device allocated from the second computing resource pool to the third computing resource pool changes, for any round of processing in the decoding stage, the feature vector of the text unit in the corresponding round is read by the computing device of the third computing resource pool from the memory expansion resource pool.

[0110] In the embodiment of the present application, for any input information, after the third computing resource pool decodes the incremental KV vector of any text unit in the input information, the feature data still remains in the device memory of the third computing resource pool. During this period, if some computing devices in the second computing resource pool are scheduled to the first computing resource pool, the feature data is also stored in the memory expansion resource pool, and then the feature vector of the next text unit is continuously obtained from the memory expansion resource pool; on the contrary, when the computing device allocated from the second computing resource pool to the third computing resource pool remains unchanged, for any round of processing in the decoding stage, the feature vector of the text unit in the corresponding round is obtained by the computing device of the third computing resource pool from its own device memory. Thus, the number of accesses to the memory expansion resource pool can be reduced, and the data processing efficiency can be improved.

[0111] In some embodiments, the computing device has an accelerator card memory; the accelerator card memory is used to store at least part of the weight parameters, bias parameters, hyperparameters, and system operation parameters of the target AI model.

[0112] In the embodiments of the present application, the accelerator card memory of the computing devices in each computing resource pool is used to store the weight parameters, bias parameters, hyperparameters, and system operation parameters of the model. This design is to place more model parameters in the accelerator card memory, reduce the cluster scale under the same model, and thus reduce the synchronization communication overhead of the large-scale cluster.

[0113] In some embodiments, the resource scheduler is further configured to: during the process of the target AI model processing the input information, obtain the temperature data and memory occupancy rate of each of the first computing resource pool, the second computing resource pool, and the third computing resource pool; In the case where it is determined that the temperature data of any one of the computing resource pools is greater than the temperature threshold and / or the memory occupancy rate is greater than the memory occupancy rate threshold, perform a downclocking process on the operating system where the computing resource pool is located.

[0114] In the embodiments of the present application, in addition to monitoring the computing power utilization rate, the resource scheduler also monitors the temperature data and the memory occupancy rate. The temperature data can be obtained from a temperature sensor, and the memory occupancy rate refers to the proportion of the system memory occupied by the computing resource pool at a certain moment.

[0115] The temperature data and the memory occupancy rate will affect the overall operation stability of the computing cluster. Therefore, these two metrics will be used for independent stability control scheduling. When the temperature data and the memory occupancy rate exceed the set thresholds (in the case where the temperature data of the computing resource pool is greater than the temperature threshold and / or the memory occupancy rate is greater than the memory occupancy rate threshold), a downclocking operation will be performed on the entire system. The downclocking operation can specifically be to reduce the working frequency of the processor (CPU frequency, GPU frequency, etc.). Thus, heat generation can be reduced and the stability of the system can be improved.

[0116] In some embodiments, the data exchange module is further configured to perform integrity verification on the feature data, the initial KV cache data, and the incremental KV data, and after the verification passes, store the feature data, the initial KV cache data, and the incremental KV data in the memory expansion resource pool.

[0117] In the embodiment of the present application, the data exchange module (specifically, it can be a CXL exchange module), under the guidance of the CXL topology manager, processes the forwarding and integrity verification of feature data and cache data (initial KV cache data and incremental KV data) (the verification can be performed according to the encoding format). After the verification passes, the feature data, initial KV cache data, and incremental KV data are stored in the memory expansion resource pool. In addition, the data exchange module can also transfer the data in the memory expansion resource pool to the specified computing resource pool, which can provide the accuracy and stability of data transmission.

[0118] See Figure 4 , the embodiment of the present application provides a second structural schematic diagram of a computing system, and the computing system includes a resource scheduler, a first computing resource pool, a second computing resource pool, and a third computing resource pool.

[0119] The resource scheduler is a CXL scheduler, and the CXL scheduler includes a CXL topology manager, a CXL exchange module, and a CXL scheduling manager. The CXL exchange module can obtain the monitoring metrics generated by any computing device in any computing resource pool. Each computing device can include multiple acceleration cards, and the acceleration card memory is used to store at least part of the weight parameters, bias parameters, hyperparameters, and system operation parameters of the target AI model. The multiple monitoring metrics include feature data, temperature data (data collected by the acceleration card temperature sensor), computing power metrics, memory occupancy rate, and cache data (KV cache data), etc.

[0120] These monitoring metrics are transmitted to the CXL scheduler through the CXL protocol, and the CXL scheduler can store these monitoring metrics in the memory expansion card of the memory expansion resource pool.

[0121] The CXL topology manager is responsible for creating or deleting high-speed connections between resource pools.

[0122] The CXL scheduling manager is used to adjust the computing resources allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool.

[0123] For the detailed content of each part, please refer to the foregoing embodiments, and no further elaboration will be provided here.

[0124] As an additional component, the CXL scheduler in the embodiment of the present application can be flexibly deployed, reducing the deployment difficulty of the cluster.

[0125] In some embodiments, the resource scheduler is further used for: Obtaining the computing power of each computing device corresponding to the computing resources of the target AI model; Pooling the computing devices with computing power within the first computing power range to obtain the first computing resource pool; Pool the computing devices with computing power in the second computing power range into a second computing resource pool; Pool the computing devices with computing power in the third computing power range into a third computing resource pool; The minimum value of the first computing power range is greater than the maximum value of the second computing power range; the minimum value of the second computing power range is greater than the maximum value of the third computing power range.

[0126] When the resource scheduler in the embodiment of the present application pools the computing resources of the target AI model, it can pool based on the computing power of each computing device. Specifically, each computing device can be classified based on its computing power, such as being classified into computing devices with computing power in the first computing power range, computing devices with computing power in the second computing power range, and computing devices with computing power in the third computing power range.

[0127] Thus, the computing devices with computing power in the first computing power range can be pooled into a first computing resource pool; the computing devices with computing power in the second computing power range can be pooled into a second computing resource pool; and the computing devices with computing power in the third computing power range can be pooled into a third computing resource pool.

[0128] In the embodiment of the present application, each computing device is pooled into a first computing resource pool, a second computing resource pool, and a third computing resource pool with different computing powers based on the computing power of the computing device. The first computing resource pool, the second computing resource pool, and the third computing resource pool can process different model tasks, which helps to improve computing efficiency, optimize resource utilization, enhance system elasticity and scalability, while simplifying management, reducing costs, and supporting diverse computing tasks.

[0129] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.

[0130] See Figure 5 , the embodiment of the present application also provides a computing power allocation method, which is applied to the resource scheduler of a computing system. The method includes steps 510 to 520: Step 510, pool multiple computing devices of the computing system to obtain a first computing resource pool, a second computing resource pool, and a third computing resource pool; the multiple computing devices are used to provide computing power for the target AI model; the multiple computing devices in the first computing resource pool are used to provide computing power for the pre-filling stage of the target AI model; the multiple computing devices in the third computing resource pool are used to provide computing power for the decoding stage of the target AI model; Step 520, during the process of the target AI model processing the input information, according to the computing power metrics of the first computing resource pool and the third computing resource pool, adjust the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool, so as to adjust the computing power of the first computing resource pool and the third computing resource pool.

[0131] In some embodiments, the computing power of the first computing resource pool is higher than that of the second computing resource pool; the computing power of the second computing resource pool is higher than that of the third computing resource pool; The computing power of the computing devices in the first computing resource pool is higher than that of the computing devices in the second computing resource pool; the computing power of the computing devices in the second computing resource pool is higher than that of the computing devices in the third computing resource pool.

[0132] In some embodiments, the computing power metric is the computing power utilization rate; the resource scheduler includes a scheduling manager and a topology manager; according to the computing power metrics of the first computing resource pool and the third computing resource pool, adjusting the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool includes: When the computing power utilization rate of the first computing resource pool is less than the first utilization threshold but the computing power utilization rate of the third computing resource pool is greater than the second utilization threshold, or when the computing power utilization rate of the third computing resource pool is less than the second utilization threshold, the scheduling manager determines the change amount of the computing devices allocated from the second computing resource pool to the first computing resource pool according to the actual total computing power, the original theoretical total computing power, and the computing power loss of the first computing resource pool; The topology manager adjusts the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool according to the change amount.

[0133] In some embodiments, determining the change amount of the computing devices allocated from the second computing resource pool to the first computing resource pool according to the actual total computing power, the original theoretical total computing power, and the computing power loss of the first computing resource pool includes: The scheduling manager determines the computing power difference between the actual total computing power of the first computing resource pool and the original theoretical total computing power of the first computing resource pool; Adjust the computing power difference according to the computing power loss to obtain the computing power adjustment amount of the first computing resource pool; The scheduling manager determines the change amount of the computing devices allocated from the second computing resource pool to the first computing resource pool according to the ratio of the computing power adjustment amount of the first computing resource pool to the peak computing power of the second computing resource pool.

[0134] In some embodiments, the resource scheduler includes a topology manager; the method further includes: adjusting the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool by creating or deleting a first high-speed connection between the first computing resource and the computing devices allocated from the second computing resource pool to the first computing resource, and creating or deleting a second high-speed connection between the third computing resource and the computing devices allocated from the second computing resource pool to the third computing resource, according to the topology manager.

[0135] In some embodiments, when the computing power utilization rate of the first computing resource pool is less than the first utilization threshold but the computing power utilization rate of the third computing resource pool is greater than the second utilization threshold, adjusting the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool includes: Determining, by the topology manager, each first candidate computing device in the second computing resource pool that has a first high-speed connection with the first computing resource pool; Selecting a first target computing device with a change amount from the first candidate computing devices; Disconnecting the first high-speed connection between the first computing resource pool and the first target computing device with a change amount; Creating a second high-speed connection between the first target computing device with a change amount and the third computing resource pool.

[0136] In some embodiments, when the computing power utilization rate of the third computing resource pool is less than the second utilization threshold, adjusting the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool includes: Determining, by the topology manager, each second candidate computing device in the second computing resource pool that has a second high-speed connection with the third computing resource pool; Selecting a second target computing device with a change amount from the second candidate computing devices; Disconnecting the second high-speed connection between the third computing resource pool and the second target computing device with a change amount; Creating a first high-speed connection between the second target computing device with a change amount and the first computing resource pool.

[0137] In some embodiments, the computing power metric is the computing power utilization rate; during the process of the target AI model processing the input information, the method further includes: When the computing power utilization rate of the first computing resource pool is greater than the first utilization threshold, or the computing power utilization rate of the third computing resource pool is greater than the second utilization threshold, maintaining, by the scheduling manager, the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool.

[0138] In some embodiments, before adjusting the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool according to the computing power metrics of the first computing resource pool and the third computing resource pool, the method further includes: In the initialization stage of the pre-filling phase, the scheduling manager allocates each computing device in the second computing resource pool to the first computing resource pool.

[0139] In some embodiments, the computing system further includes a memory expansion resource pool; the memory expansion resource pool is obtained after pooling and expanding the memory resources of the target AI model by the resource manager; The computing devices in the first computing resource pool, the second computing resource pool, and the third computing resource pool share the memory expansion resource pool.

[0140] In some embodiments, the input information includes multiple text units; the resource scheduler further includes a data exchange module; The first computing resource pool is used to process the input information in the pre-filling phase of each input information, and obtain and output the feature data and initial KV cache data of each input information to the data exchange module; the method further includes: Storing the feature data and initial KV cache data of each input information into the memory expansion resource pool through the data exchange module; The third computing resource pool is used to obtain and process the feature data and the KV cache data of the previous round from the memory expansion resource pool at least once in the decoding phase of each input information, so as to update the KV cache data in the memory expansion resource pool to the target KV cache data; the target KV cache data is used to generate the prediction result of the target AI model; the KV cache data of the first round is the initial KV cache data.

[0141] In some embodiments, the feature data includes the feature vector of each text unit; the initial KV cache data includes the KV vector of each text unit; The third computing resource pool is used for each round of processing in the decoding phase, to process the text vector of a corresponding text unit in the corresponding round and the KV cache data of the previous stage read from the memory expansion resource pool, and obtain and output the incremental KV vector of the corresponding round to the data exchange module; The method further includes Storing the incremental KV vector of the corresponding round into the memory expansion resource pool through the data exchange module, so as to update the KV cache data in the memory expansion resource pool to obtain the KV cache data of the corresponding round; the KV cache data of the last round is the target KV cache data.

[0142] In some embodiments, the computing device includes device memory; In the decoding stage, when the computing device allocated from the second computing resource pool to the third computing resource pool remains unchanged, for any round of processing in the decoding stage, the feature vector of the text unit in the corresponding round is obtained by the computing device in the third computing resource pool from its own device memory; the feature data in the device memory is obtained from and cached in the memory expansion resource pool; When the computing device allocated from the second computing resource pool to the third computing resource pool changes, for any round of processing in the decoding stage, the feature vector of the text unit in the corresponding round is read by the computing device in the third computing resource pool from the memory expansion resource pool.

[0143] In some embodiments, the method further includes: Performing integrity verification on the feature data, initial KV cache data, and incremental KV data through the data exchange module, and after the verification passes, storing the feature data, initial KV cache data, and incremental KV data in the memory expansion resource pool.

[0144] In some embodiments, the computing device has an accelerator card memory; the accelerator card memory is used to store at least part of the weight parameters, bias parameters, hyperparameters, and system operation parameters of the target AI model.

[0145] In some embodiments, the method further includes: During the process of the target AI model processing the input information, obtaining the temperature data and memory occupancy rates of each computing resource pool in the first computing resource pool, the second computing resource pool, and the third computing resource pool; When it is determined that the temperature data of any one computing resource pool is greater than the temperature threshold and / or the memory occupancy rate is greater than the memory occupancy rate threshold, performing downclocking on the operating system where the computing resource pool is located.

[0146] In some embodiments, pooling multiple computing devices of the computing system includes Obtaining the computing power of each computing resource corresponding computing device of the target AI model; Pooling the computing devices with computing power in the first computing power range into the first computing resource pool; Pooling the computing devices with computing power in the second computing power range into the second computing resource pool; Pooling the computing devices with computing power in the third computing power range into the third computing resource pool; The minimum value of the first computing power range is greater than the maximum value of the second computing power range; the minimum value of the second computing power range is greater than the maximum value of the third computing power range.

[0147] For the description of the features in the embodiments corresponding to the computing power allocation method, reference can be made to the relevant description of the embodiments corresponding to the computing system, which will not be elaborated here one by one.

[0148] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the resource allocation method.

[0149] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the resource allocation method when running.

[0150] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk, or optical disc, etc., various media that can store computer programs.

[0151] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the resource allocation method.

[0152] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the resource allocation method.

[0153] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0154] The above has introduced in detail a computing system, a computing power allocation method, an electronic device, a computer-readable storage medium, and a computer program product provided by the present application. Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only applicable to helping understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can still be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A computing system, characterized in that, Including: Multiple computing devices for providing computing power for a target AI model; A resource scheduler for pooling the multiple computing devices to obtain a first computing resource pool, a second computing resource pool, and a third computing resource pool; the multiple computing devices in the first computing resource pool are used to provide computing power for the pre-filling stage of the target AI model; The multiple computing devices in the third computing resource pool are used to provide computing power for the decoding stage of the target AI model; The resource scheduler is further configured to adjust the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool according to the computing power metrics of the first computing resource pool and the third computing resource pool during the process of the target AI model processing input information, so as to adjust the computing power of the first computing resource pool and the third computing resource pool.

2. The computing system according to claim 1, wherein The computing power of the first computing resource pool is higher than that of the second computing resource pool; the computing power of the second computing resource pool is higher than that of the third computing resource pool; The computing power of the computing devices in the first computing resource pool is higher than that of the computing devices in the second computing resource pool; The computing power of the computing devices in the second computing resource pool is higher than that of the computing devices in the third computing resource pool.

3. The computing system according to claim 2, wherein The computing power metric is the computing power utilization rate; the resource scheduler includes a scheduling manager and a topology manager; The scheduling manager is configured to determine the change amount of the computing devices allocated from the second computing resource pool to the first computing resource pool according to the actual total computing power, the original theoretical total computing power, and the computing power loss of the first computing resource pool when the computing power utilization rate of the first computing resource pool is less than a first utilization threshold but the computing power utilization rate of the third computing resource pool is greater than a second utilization threshold, or when the computing power utilization rate of the third computing resource pool is less than the second utilization threshold; The topology manager is further configured to adjust the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool according to the change amount.

4. The computing system according to claim 3, wherein The scheduling manager is further configured to determine the computing power difference between the actual total computing power of the first computing resource pool and the original theoretical total computing power of the first computing resource pool; adjust the computing power difference according to the computing power loss to obtain the computing power adjustment amount of the first computing resource pool; The scheduling manager is further configured to determine the change amount of the computing devices allocated from the second computing resource pool to the first computing resource pool according to the ratio of the computing power adjustment amount of the first computing resource pool to the peak computing power of the second computing resource pool.

5. The computing system according to claim 3, wherein The resource scheduler includes a topology manager; the topology manager adjusts the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool by creating or deleting a first high-speed connection between the first computing resource and the computing devices allocated from the second computing resource pool to the first computing resource, and creating or deleting a second high-speed connection between the third computing resource and the computing devices allocated from the second computing resource pool to the third computing resource.

6. The computing system according to claim 5, wherein the topology manager is specifically configured to, when the computing power utilization rate of the first computing resource pool is less than a first utilization rate threshold but the computing power utilization rate of the third computing resource pool is greater than a second utilization rate threshold, determine, from the second computing resource pool, each first candidate computing device having a first high-speed connection with the first computing resource pool; select a first target computing device of the change amount from the first candidate computing devices; disconnect the first high-speed connection between the first computing resource pool and the first target computing device of the change amount; create a second high-speed connection between the first target computing device of the change amount and the third computing resource pool.

7. The computing system according to claim 5, wherein the topology manager is specifically configured to, when the computing power utilization rate of the third computing resource pool is less than the second utilization rate threshold, determine, from the second computing resource pool, each second candidate computing device having a second high-speed connection with the third computing resource pool; select a second target computing device of the change amount from the second candidate computing devices; disconnect the second high-speed connection between the third computing resource pool and the second target computing device of the change amount; create a first high-speed connection between the second target computing device of the change amount and the first computing resource pool.

8. The computing system according to claim 3, wherein The computing power metric is the computing power utilization rate; the scheduling manager is further configured to maintain the computing devices allocated from the second computing resource pool to the first computing resource pool and the third computing resource pool when the computing power utilization rate of the first computing resource pool is greater than the first utilization rate threshold, or the computing power utilization rate of the third computing resource pool is greater than the second utilization rate threshold.

9. The computing system according to any one of claims 3-8, characterized in that, The scheduling manager is further configured to allocate each computing device in the second computing resource pool to the first computing resource pool during the initialization stage of the pre-filling stage.

10. The computing system according to claim 1, wherein The computing system further includes a memory expansion resource pool; the memory expansion resource pool is obtained after the resource manager pools and expands the memory resources of the target AI model. The computing devices in the first computing resource pool, the second computing resource pool, and the third computing resource pool share the memory expansion resource pool.

11. The computing system according to claim 10, wherein The input information includes a plurality of text units; the resource scheduler further includes a data exchange module. The first computing resource pool is used to process the input information during the pre-filling stage of each piece of the input information, and obtain and output to the data exchange module the feature data and initial KV cache data of each piece of the input information; The data exchange module is used to store the feature data and initial KV cache data of each piece of the input information into the memory expansion resource pool; The third computing resource pool is used to, during the decoding stage of each piece of the input information, obtain and process the feature data and the KV cache data of the previous round from the memory expansion resource pool at least once, with the goal of updating the KV cache data in the memory expansion resource pool to target KV cache data; the target KV cache data is used to generate the prediction result of the target AI model; the KV cache data of the first round is the initial KV cache data.

12. The computing system according to claim 11, wherein The feature data includes the feature vector of each text unit; the initial KV cache data includes the KV vector of each text unit; The third computing resource pool is used for each round of processing in the decoding stage to process the text vector of a corresponding text unit in the corresponding round and the KV cache data of the previous stage read from the memory expansion resource pool, and obtain and output to the data exchange module the incremental KV vector of the corresponding round; The data exchange module is used to store the incremental KV vector of the corresponding round into the memory expansion resource pool to update the KV cache data in the memory expansion resource pool to obtain the KV cache data of the corresponding round; the KV cache data of the last round is the target KV cache data.

13. The computing system according to claim 12, wherein The computing device includes device memory; During the decoding stage, when the computing device allocated by the second computing resource pool to the third computing resource pool remains unchanged, for any round of processing in the decoding stage, the feature vector of the text unit in the corresponding round is obtained by the computing device of the third computing resource pool from its own device memory; the feature data in the device memory is obtained and cached from the memory expansion resource pool; When the computing device allocated by the second computing resource pool to the third computing resource pool changes, for any round of processing in the decoding stage, the feature vector of the text unit in the corresponding round is read by the computing device of the third computing resource pool from the memory expansion resource pool.

14. The computing system according to claim 12, wherein The data exchange module is also used to perform integrity verification on the feature data, the initial KV cache data, and the incremental KV data, and after the verification passes, store the feature data, the initial KV cache data, and the incremental KV data into the memory expansion resource pool.

15. The computing system according to claim 1, wherein The computing device has an accelerator card memory; the accelerator card memory is used to store at least part of the weight parameters, bias parameters, hyperparameters, and system operation parameters of the target AI model.

16. The computing system according to claim 1, wherein The resource scheduler is further configured to: during the process of the target AI model processing input information, obtain the temperature data and memory occupancy rates of each of the first computing resource pool, the second computing resource pool, and the third computing resource pool; In the case where it is determined that the temperature data of any one of the computing resource pools is greater than the temperature threshold and / or the memory occupancy rate is greater than the memory occupancy rate threshold, perform a frequency reduction process on the operating system where the computing resource pool is located.

17. The computing system according to claim 1, wherein The resource scheduler is further configured to: Obtain the computing power of the corresponding computing devices of each computing resource of the target AI model; Pool the computing devices with computing power within the first computing power range to obtain a first computing resource pool; Pool the computing devices with computing power within the second computing power range to obtain a second computing resource pool; Pool the computing devices with computing power within the third computing power range to obtain a third computing resource pool; The minimum value of the first computing power range is greater than the maximum value of the second computing power range; the minimum value of the second computing power range is greater than the maximum value of the third computing power range.

18. A computing power allocation method, characterized in that, A resource scheduler applied to a computing system, comprising: Pooling multiple computing devices of the computing system to obtain a first computing resource pool, a second computing resource pool, and a third computing resource pool; the multiple computing devices are used to provide computing power for a target AI model; the multiple computing devices in the first computing resource pool are used to provide computing power for the pre-filling stage of the target AI model; the multiple computing devices in the third computing resource pool are used to provide computing power for the decoding stage of the target AI model; During the process of the target AI model processing input information, according to the computing power metrics of the first computing resource pool and the third computing resource pool, adjust the computing devices allocated from the second computing resource pool to the first computing resource pool and allocated to the third computing resource pool, so as to adjust the computing power of the first computing resource pool and the third computing resource pool.

19. An electronic device, characterized in that, Comprising: A memory for storing a computer program; A processor for implementing the steps of the computing power allocation method as described in claim 18 when executing the computer program.

20. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the computing power allocation method as described in claim 18.

21. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the computing power allocation method as described in claim 18.

Citation Information

Patent Citations

  • Server system, resource scheduling method of server system, chip and chip grain

    CN118210634A

  • Model reasoning scheduling method and device and server cluster

    CN118897736A

  • Resource management method and device for computing power resource pool, electronic equipment and storage medium

    CN119668853A

  • Intelligent computing power scheduling method and device for intelligent common computing power computing center

    CN119938340A

  • Resource scheduling method, job processing method, scheduler, system, and related device

    WO2024221991A1