Request Allocation Method, System, Device, and Medium Based on Heterogeneous Computing System

By obtaining inference task information and heterogeneous computing power equipment performance parameters in heterogeneous computing system, determining the key-value cache reading time and presetting the number of allocation requests, the problem of poor allocation balance of inference requests is solved, and the saving of computing power resources and cost reduction is achieved.

CN119690687BActive Publication Date: 2025-06-13LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510221607.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In heterogeneous computing systems, the allocation balance of inference requests is poor, resulting in a lot of wasted computing resources, especially when calling separate memory or local memory.

Method used

By obtaining inference task information and performance parameters of heterogeneous computing power equipment, the reading time required for key-value cache is determined, and the appropriate heterogeneous computing power equipment is selected based on the reading time, the number of allocation requests is preset based on the target reading time and performance parameters, and the request allocation result is finally determined.

Benefits of technology

It improves the allocation balance of inference requests, saves computing resources, reduces costs, and optimizes the feedback speed of model inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119690687B_ABST
    Figure CN119690687B_ABST
Patent Text Reader

Abstract

The present application discloses a request allocation method, system, device and medium based on a heterogeneous computing system, relating to the field of computer technology. When giving priority to the use of the key-value cache mechanism, the reading time required for multiple heterogeneous computing power devices to access memory using the key-value cache is determined. Considering the performance differences in the computing power information of heterogeneous computing power devices, the memory information corresponding to memory expansion, and the characteristics of inference task information, requests are reasonably allocated. Further, according to the comparison relationship between the preset number of allocated requests and the number of concurrent requests, and different strategies for whether the allocation conditions are met, the rationality of request allocation is improved. Therefore, the technical problem of poor allocation balance for inference requests when calling separate memory or local memory, resulting in a large waste of computing power resources, can be solved, achieving the technical effect of reasonably allocating inference task information to heterogeneous computing power devices to improve the allocation balance while saving computing power resources and reducing costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a request allocation method, system, device, and medium based on a heterogeneous computing system. Background Art

[0002] In a disaggregated memory system, when heterogeneous computing power devices execute inference tasks, they will access local memory or disaggregated memory. When the inference requests corresponding to the inference tasks are allocated to the heterogeneous computing power devices, if the disaggregated memory is called, problems such as the latency and bandwidth of the corresponding memory during the access process will reduce the computing efficiency of the inference tasks, affect the feedback speed of model inference, and reduce the user experience. If the local memory is called for model inference, it will consume more computing power resources and overall reduce the allocation balance of inference requests.

[0003] Therefore, how to improve the allocation balance of inference requests to save computing power resources is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0004] This application provides a request allocation method, system, device, and medium based on a heterogeneous computing system to at least solve the problem in the related art that when calling disaggregated memory or local memory, the allocation balance of inference requests is poor, resulting in a large waste of computing power resources.

[0005] This application provides a request allocation method based on a heterogeneous computing system, including:

[0006] Obtain inference task information, multiple heterogeneous computing power devices, and corresponding performance parameters;

[0007] Determine the required reading time of the key-value cache based on the inference task information and multiple performance parameters; and determine the current heterogeneous computing power device according to the multiple reading times;

[0008] Determine the preset allocation request quantity of the inference task according to the target reading time and target performance parameters of the current heterogeneous computing power device;

[0009] Determine the request allocation result of the current heterogeneous computing power device according to the preset allocation request quantity, the concurrent request quantity of the inference task information, and the allocation conditions.

[0010] This application also provides a heterogeneous computing system, including a host, multiple heterogeneous computing power devices, and multiple storage devices;

[0011] The host is respectively connected to multiple heterogeneous computing power devices; the multiple heterogeneous computing power devices call multiple storage devices;

[0012] The host is used to execute the steps of the above-mentioned request allocation method based on a heterogeneous computing system to complete the request allocation of inference tasks.

[0013] The present application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any one of the above request allocation methods based on a heterogeneous computing system when executing the computer program.

[0014] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any one of the above request allocation methods based on a heterogeneous computing system.

[0015] The present application also provides a computer program product including a computer program, which, when executed by a processor, implements the steps of any one of the above request allocation methods based on a heterogeneous computing system.

[0016] Through the present application, since the inference task information and multiple performance parameters determine the reading time required for the key-value cache, considering that the key-value cache mechanism speeds up the inference speed in the inference task, and giving priority to the use of the key-value cache mechanism, the reading time required for the key-value cache to be used when multiple heterogeneous computing power devices access the memory is determined, so as to fully consider the reading time feedback by the key-value cache during the process of presetting the allocation request quantity. At the same time, the performance parameters of multiple heterogeneous computing power devices are added, which mainly include computing power information and memory information of the disaggregated memory, such as latency, bandwidth and other information, to determine the preset allocation request quantity. Considering the performance differences of the computing power information of heterogeneous computing power devices, the memory information corresponding to memory expansion and the characteristics of inference task information, requests are reasonably allocated. In addition, considering the key-value cache mechanism, the request allocation result of the current heterogeneous computing power device is determined according to the relationship between the preset allocation request quantity, the concurrent request quantity and the allocation conditions, and further considering the comparison relationship between the preset allocation request quantity and the concurrent request quantity, as well as different strategies for whether the allocation conditions are met, the rationality of request allocation and the user experience are improved. Therefore, the technical problem of poor allocation balance for inference requests when calling disaggregated memory or local memory, resulting in a large waste of computing power resources, can be solved, and the inference task information can be reasonably allocated to heterogeneous computing power devices, so as to improve the allocation balance while saving computing power resources and reducing costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0018] Figure 1 It is a structural diagram of a heterogeneous computing system provided by an embodiment of the present application;

[0019] Figure 2 A flowchart of a request allocation method based on a heterogeneous computing system provided by an embodiment of the present application;

[0020] Figure 3 A schematic diagram of a request allocation architecture based on a heterogeneous computing system provided by an embodiment of the present application;

[0021] Figure 4 A structural diagram of a request allocation device based on a heterogeneous computing system provided by an embodiment of the present application. Detailed implementation manners

[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0023] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0024] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0025] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the request allocation method based on the heterogeneous computing system depends, the specific application environment architecture or specific hardware architecture is described herein.

[0026] In a heterogeneous computing system, Figure 1 A structural diagram of a heterogeneous computing system provided by an embodiment of the present application, as Figure 1 shown, for each heterogeneous computing power device (in Figure 1Three heterogeneous computing power devices are given in [reference], namely heterogeneous computing power device 1', heterogeneous computing power device 2' and heterogeneous computing power device 3'), all of which can access local memory and extended memory in a separate memory network, that is, non-local memory devices at the same time; the two memory devices are collectively referred to as storage devices. When a new inference task of a pre-trained language model arrives, two processing methods are usually adopted. One is to execute in local memory. Since the memory space of local memory is limited and the key-value cache (KVCache) occupies a large amount of memory, heterogeneous computing power devices are directly used for model inference in the case of local memory calls, and KVCache is not used. The second is to store KVCache in a separate memory, and use the storage space conversion performance for model inference when calling the separate memory.

[0027] The first processing method mentioned above has a lower latency for feedback inference results, but it consumes a large amount of computing power resources, and a single heterogeneous computing power device can only process a small number of inference requests. Although the second processing method can process more inference requests, it is affected by the latency, bandwidth, etc. of the separate memory, resulting in a larger latency for returning inference results. Therefore, while reasonably selecting the processing method, heterogeneous computing power resources need to be saved.

[0028] In the above heterogeneous computing system, training tasks are assigned to multiple heterogeneous computing power devices. The host is connected to multiple heterogeneous computing power devices respectively, and multiple heterogeneous computing power devices call multiple storage devices. Heterogeneous computing power mainly refers to the computing method of a system composed of computing units using different types of instruction sets and architectures. Common categories of computing units include Central Processing Unit (CPU), Graphics Processing Unit (GPU), Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), etc. Heterogeneous computing power devices are different types of processors, and their performance varies when processing different types of applications.

[0029] The embodiments of the present application provide a request allocation method based on a heterogeneous computing system. Combining the execution process of the request allocation method based on the heterogeneous computing system, the method is described in detail.

[0030] Pre-training is a strategy for training deep learning models. Its core lies in using a large-scale dataset to preliminarily train the model so that the model can learn general feature representations. This process is similar to the basic learning stage of humans before learning new knowledge, where they accumulate experience through extensive reading and observation.

[0031] Pre-trained language models generally refer to designing language model training tasks based on a large-scale corpus (including language training materials such as sentences and paragraphs), training a large-scale neural network algorithm structure to learn and implement. The final large-scale neural network algorithm structure and parameters are the pre-trained language model. For subsequent other tasks, feature extraction or task fine-tuning can be performed on the basis of this model to achieve specific task purposes. The idea of pre-training is to first train a task to obtain a set of model parameters, then use this set of model parameters to initialize the network model parameters, and then use the initialized network model to train other tasks to obtain models adapted to other tasks. By pre-training on a large-scale corpus, neural language representation models can learn powerful language representation capabilities and can extract rich syntactic and semantic information from text. Pre-trained language models can provide word elements (tokens) containing rich semantic information and sentence-level features for downstream tasks, or directly perform fine-tuning for downstream tasks on the pre-trained model to conveniently and quickly obtain downstream-specific models.

[0032] The neural network algorithm structure for pre-training language model training can be a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory (LSTM), etc., or it can be a model constructed by an attention network, such as a transformer structure, a bidirectional encoder representations from transformers (BERT), a generative pre-trained transformer (GPT), a contrastive language-image pre-training (CLIP), etc. This application does not make any limitations here. An attention network refers to a network model that uses an attention mechanism for training. This model assigns different weights to each part of the input sequence, thereby extracting more important feature information from the input sequence and enabling the model to finally obtain a more accurate output.

[0033] Fine-tuning refers to further training on a task-specific dataset based on the use of a pre-trained model to adjust the model parameters to better suit the target task. During the fine-tuning process, most of the layers of the pre-trained model are usually frozen, and only the newly added layers are trained or a small number of key layers are adjusted. This can not only retain the useful features learned by the pre-trained model, but also quickly adapt to the specific needs of the new task. In addition, choosing the right learning rate and training rounds is also the key to successful fine-tuning. Fine-tuning refers to small-scale training for specific task objectives (downstream tasks) and task data (downstream data) based on the pre-trained model, to achieve minor adjustments to the parameters of the pre-trained model, and finally obtain a model adapted to specific tasks and data.

[0034] Figure 2 A flowchart of a request allocation method based on a heterogeneous computing system provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the method includes:

[0035] S11: Obtain inference task information, multiple heterogeneous computing devices, and corresponding performance parameters.

[0036] It is understandable that reasoning task information usually refers to the data, parameters, configuration information, etc. involved in the reasoning-related tasks. Since the pre-trained language model is mainly formed by transformer stacking, it mainly includes model information and task information. Model information is the model parameters of the pre-trained model, mainly including:

[0037] How many transformer structures does the model contain?

[0038] How many attention heads does a transformer have?

[0039] The dimension of the key vector for each attention head;

[0040] The dimension of the value vector of each attention head;

[0041] The maximum generated sequence length supported by the task;

[0042] The precision of the model parameters, such as half precision, single precision, etc.

[0043] The above model information can be manually input or obtained based on tools provided by the deep learning framework, such as a profiler. There is no limitation here and it can be set according to actual conditions.

[0044] Regarding task information, it mainly includes reasoning request information, such as:

[0045] User-specified heterogeneous computing device types, such as H100 and MLU370 single computing devices;

[0046] How many concurrent requests need to be processed at most per second?

[0047] The requirement with the longest response time.

[0048] The above information depends on user needs and can be manually input.

[0049] The multiple heterogeneous computing devices may be all the heterogeneous computing devices corresponding to the entire heterogeneous computing system, or multiple heterogeneous computing devices selected from all the heterogeneous computing devices. The selection requirements here may be specified, or may be selected by one or more factors such as computing power, latency, bandwidth, etc. to select heterogeneous computing devices that meet the requirements, so as to save computing resources when requesting allocation.

[0050] The performance parameters of heterogeneous computing devices include heterogeneous computing performance parameters and memory communication performance parameters. The heterogeneous computing performance parameters cover the performance indicators of different computing units, such as peak computing information, floating-point calculation accuracy, etc. They can be obtained through the performance manual of heterogeneous computing or related public information, and entered into the module in advance for storage and retrieval when needed. They can also be obtained by recording historical heterogeneous computing execution information, that is, using the profiler that comes with the deep learning framework to obtain it during execution.

[0051] Memory communication performance parameters are used to describe the access efficiency and transmission capacity of memory resources, such as the latency and bandwidth information of separate memory. The latency and bandwidth information of each heterogeneous computing power reading the data in its corresponding separate memory (since separate memory can be used to expand the heterogeneous computing power memory in large quantities using double data rate (DDR), it is assumed that the capacity of the remotely expanded separate memory is unlimited). Separate memory may be implemented using Compute Express Link (CXL), Non-Uniform Memory Access (NUMA) and other methods. Therefore, for the extended memory of each heterogeneous computing power, its latency and bandwidth information are different and need to be collected separately. The corresponding latency and bandwidth information can be obtained using public testing tools such as the system benchmark tool (sysbench).

[0052] S12: Determine the read time required for the key-value cache based on the inference task information and multiple performance parameters; and determine the current heterogeneous computing power device based on the multiple read times.

[0053] The reading time required for the key-value cache is the reading time when accessing the disaggregated memory using the key-value cache mechanism for different heterogeneous computing power devices, that is, the feedback time of the key-value cache mechanism. The inference task information mainly considers the maximum key-value cache data volume corresponding to the inference task and the bandwidth and latency parameters corresponding to the disaggregated memory in the performance parameters. The maximum key-value cache data volume is obtained by the maximum floating-point count, corresponding to the floating-point operation ability of the heterogeneous computing power device, and reflects the maximum number of floating-point operations that the device can complete per unit time.

[0054] Here, when different heterogeneous computing power devices access the disaggregated memory using the key-value cache, the reading times may correspond to different values for the latency and bandwidth of different heterogeneous computing power devices accessing the disaggregated memory. It should be noted that there are multiple storage devices in the disaggregated memory. The multiple latencies and multiple bandwidths obtained by different heterogeneous computing power devices respectively calling each storage device can form a matching relationship by pre-selecting the disaggregated memory with the least latency and bandwidth corresponding to different heterogeneous computing power devices based on the two factors of latency and bandwidth, so as to facilitate the determination of the subsequent reading time, that is, each reading time corresponds to one heterogeneous computing power device. It is also possible to calculate the reading time for accessing all disaggregated memories based on all storage devices for the latency and bandwidth generated by each heterogeneous computing power device accessing each storage device, and then screen based on the reading time required for the key-value cache corresponding to each heterogeneous computing power device, and select the minimum reading time as the reading time corresponding to each heterogeneous computing power device, that is, form a matching relationship where there is only one disaggregated memory accessed by different heterogeneous computing power devices. The first way of determining the reading time corresponds to less processing time, and the second way of determining the reading time is more accurate. The determination processes of the above two reading times can be set according to the actual situation and are not limited here.

[0055] After forming the above matching relationship, determine the current heterogeneous computing power device according to multiple reading times. It should be noted that the number of current heterogeneous computing power devices is one, that is, check and iterate through the heterogeneous computing power device corresponding to the minimum reading time among multiple reading times as the current heterogeneous computing power device. This is to facilitate the subsequent allocation of inference requests on the current heterogeneous computing power device to improve the subsequent inference speed.

[0056] It should be noted that the current heterogeneous computing power device is updated based on the real-time matching relationship. For example, there are 10 heterogeneous computing power devices, and their corresponding reading times are each one. After selecting the current heterogeneous computing power device and performing subsequent steps S13 and S14, the matching relationship will be updated, and there will be 9 remaining heterogeneous computing power devices, and then continue to select the next current heterogeneous computing power device.

[0057] S13: Determine the preset allocation request quantity of the inference task according to the target reading time and target performance parameters of the current heterogeneous computing power device.

[0058] Specifically, the preset number of allocation requests is based on the number of requests that the current heterogeneous computing power device can allocate when using KVCache. It is determined based on the reading time and performance parameters corresponding to the current heterogeneous computing power, mainly considering the computing power aspect, and adopting the peak computing power information representing the computing power information and the corresponding maximum floating-point calculation amount under the KVCache mechanism. The preset number of allocation requests for the inference task is determined according to the two parameters, the reading time, and the response time.

[0059] S14: Determine the request allocation result of the current heterogeneous computing power device according to the preset number of allocation requests, the concurrent request number of the inference task information, and the allocation conditions.

[0060] Specifically, the preset number of allocation requests is the number of requests expected to be allocated to the current heterogeneous computing power device. To make the request allocation result accurate, in this embodiment, it is judged whether the preset number of allocation requests meets the allocation conditions. If it meets, the final request allocation result is determined based on the relationship between the preset number of allocation requests and the concurrent request number of the inference task information. At this time, if the actual concurrent request number is less than the preset number of allocation requests, the actual concurrent request number is selected as the request allocation number. This scenario mainly considers the remaining unallocated concurrent request number and the total concurrent request number. If it does not meet, it means that the KVCache mechanism cannot meet the latency requirements, and then the request allocation result of the current heterogeneous computing power device needs to be determined according to the performance parameters, that is, the computing power information.

[0061] It can be understood that in this embodiment, corresponding to the reduction of costs, by allocating concurrent requests to heterogeneous computing power devices, compared with randomly placing rare requests in heterogeneous computing power devices, resulting in more heterogeneous computing power devices and thus more purchase costs, in this embodiment, a reasonable allocation method is adopted to deploy requests corresponding to more inference tasks and allocate them to heterogeneous computing power devices, so as to deploy more inference tasks with fewer heterogeneous computing power devices, thereby reducing the purchase costs. Based on the existing heterogeneous computing power devices, through the deployment and allocation situation of this embodiment, the power consumption of the heterogeneous computing power devices is saved, and the operating costs are saved.

[0062] Through this application, since the inference task information and multiple performance parameters determine the reading time required for the key-value cache, considering that the key-value cache mechanism speeds up the inference in the inference task, when giving priority to the use of the key-value cache mechanism, the reading time required for the key-value cache to be used when multiple heterogeneous computing power devices access the memory is determined, so as to fully consider the reading time feedback by the key-value cache during the process of presetting the allocation request quantity. At the same time, the performance parameters of multiple heterogeneous computing power devices are added, which mainly include computing power information and memory information of the discrete memory, such as latency, bandwidth and other information, to determine the preset allocation request quantity. Considering the performance differences of the computing power information of heterogeneous computing power devices, the memory information corresponding to memory expansion and the characteristics of inference task information, the requests are reasonably allocated. In addition, considering the key-value cache mechanism, the request allocation result of the current heterogeneous computing power device is determined according to the relationship between the preset allocation request quantity, the concurrent request quantity and the allocation conditions. Further considering the comparison relationship between the preset allocation request quantity and the concurrent request quantity, and different strategies for whether the allocation conditions are met, the rationality of request allocation and the user experience are improved. Therefore, the technical problem of poor allocation balance for inference requests when calling discrete memory or local memory, resulting in a large waste of computing power resources, can be solved, and the inference task information can be reasonably allocated to heterogeneous computing power devices to improve the allocation balance while saving computing power resources and reducing costs.

[0063] In some embodiments, the process of selecting multiple heterogeneous computing power devices includes:

[0064] Obtain multiple initial heterogeneous computing power devices and corresponding peak computing power information under the heterogeneous computing system;

[0065] According to the relationship between multiple peak computing power information and preset computing power information, screen out the target initial heterogeneous computing power devices corresponding to the preset computing power information from the multiple peak computing power information;

[0066] Use the target initial heterogeneous computing power devices as multiple heterogeneous computing power devices.

[0067] Specifically, there are multiple heterogeneous computing power devices in the heterogeneous computing system, and they are used as initial heterogeneous computing power devices. Here, based on the determination of the heterogeneous computing power specified by the user, screening needs to be carried out according to the peak computing power information and the preset computing power information, so as to screen out the peak computing power information corresponding to the preset computing power information from the multiple initial heterogeneous computing power devices, and use the initial heterogeneous computing power devices corresponding to the peak computing power information exceeding the preset computing power information as the target initial heterogeneous computing power devices, that is, the multiple heterogeneous computing power devices in step S11.

[0068] The determination process provided in this embodiment for screening and determining the specified heterogeneous computing power device based on the peak computing power information and the preset computing power information improves the computing power performance.

[0069] In some embodiments, determining the read time required for the key-value cache based on the inference task information and multiple performance parameters includes:

[0070] Obtaining the critical key-value cache data volume required by the inference task information, the latency information and the bandwidth information of the performance parameters;

[0071] Determining the read time required for the corresponding key-value cache when multiple heterogeneous computing power devices access the disaggregated memory according to the critical key-value cache data volume, multiple latency information, and multiple bandwidth information.

[0072] Specifically, the critical key-value cache data volume required by the inference task information is the maximum KVCache data volume required by the inference requirement. The performance parameters correspond to the latency and bandwidth of the disaggregated memory.

[0073] Determining the read time required for the corresponding key-value cache when each heterogeneous computing power device accesses the disaggregated memory through the critical key-value cache data volume and the memory characteristic information of the extended memory. The read time determined by taking into account the computing power information and the memory characteristic information of the extended memory enables reasonable allocation during the process of allocating the number of requests, improving the user experience.

[0074] In some embodiments, determining the read time required for the corresponding key-value cache when multiple heterogeneous computing power devices access the disaggregated memory according to the critical key-value cache data volume, multiple latency information, and multiple bandwidth information includes:

[0075] Determining multiple critical transmission latencies according to the critical key-value cache data volume and multiple bandwidth information;

[0076] Determining multiple read times according to the multiple critical transmission latencies and the corresponding multiple latency information.

[0077] The process of determining the read time is as follows:

[0078] ;

[0079] wherein, is the read time required for each heterogeneous computing power device to access the corresponding disaggregated memory key-value cache, is the latency information for accessing the disaggregated memory, is the bandwidth information for accessing the disaggregated memory, is the critical key-value cache data volume, is the th heterogeneous computing power device.

[0080] The process of determining the read time required for the corresponding key-value cache when multiple heterogeneous computing power devices access the disaggregated memory provided in this embodiment lays the foundation for the subsequent process of determining the number of requests allocated using KVCache and ensures reasonable allocation.

[0081] In some embodiments, the process of determining the critical key-value cache data volume required by the inference task information includes:

[0082] Obtain the model information of the pre-trained model corresponding to the inference task information; wherein, the model information at least includes the number of transformer structures, the number of attention heads of the transformer structure, the key vector dimension of the attention head, the value vector dimension, the critical sequence length supported by the inference task, and the number of precisions;

[0083] Determine the critical key-value cache data volume according to the number of transformer structures, the number of attention heads, the key vector dimension, the value vector dimension, the critical sequence length, and the number of precisions.

[0084] Specifically, Table 1 is a table of various parameters involved in the request allocation method. As shown in Table 1, the performance parameters include the peak computing power information of heterogeneous computing power, the latency information and bandwidth information of the extended memory. The inference task information includes model information and task information.

[0085] Table 1 Table of various parameters involved in the request allocation method

[0086]

[0087] Among them, the process of determining the critical key-value cache data volume is specifically as follows:

[0088] ;

[0089] Among them, is the number of transformer structures, is the number of attention heads, is the key vector dimension, is the value vector dimension, is the critical sequence length, is the number of precisions.

[0090] The process of determining the critical key-value cache data volume required by the inference task information provided in this embodiment facilitates the subsequent calculation of the maximum read time required for the key-value cache, improves the rationality of request allocation, and makes the request allocation reasonable.

[0091] In some embodiments, determining the current heterogeneous computing power device according to multiple read times includes:

[0092] Sort the multiple read times from largest to smallest to obtain the sorted read time relationship;

[0093] Select the smallest read time in the read time relationship, and use the heterogeneous computing power device to which the smallest read time belongs as the current heterogeneous computing power device.

[0094] Specifically, one read time in this embodiment corresponds to the read time required for a heterogeneous computing power device to access a disaggregated memory key-value cache. Sorting multiple read times from largest to smallest gives the sorted read time relationship, which can be sorted in a list, a set, etc., without limitation here. Select the smallest read time, which is to select the heterogeneous computing power device to which the smallest read time belongs as the current heterogeneous computing power device. That is , where the represents heterogeneous computing power devices.

[0095] The process of selecting the current heterogeneous computing power device provided in this embodiment iterates to find the heterogeneous computing power device corresponding to the smallest key-value cache data read time, so as to improve computing power when allocating the number of requests and ensure sufficient remaining time for accessing the disaggregated memory.

[0096] In some embodiments, determining the preset allocation request quantity for the inference task according to the target read time and target performance parameter of the current heterogeneous computing power device includes:

[0097] Obtain the first critical computing power information of the key-value cache, the peak computing power information of the target performance parameter, the critical response time of the inference task information, and the target read time;

[0098] Determine the remaining access time according to the target read time and the critical response time;

[0099] Determine the number of requests processed per unit time according to the first critical computing power information and the peak computing power information;

[0100] Determine the access time corresponding to multiple requests according to the number of requests and the remaining access time;

[0101] Perform normalization processing on the access time corresponding to multiple requests and the target read time to obtain the preset allocation request quantity.

[0102] Specifically, the first critical computing power information is the maximum floating-point calculation amount in the presence of KVCache, and the critical response time is the longest response time requirement in Table 1. The formula for the preset allocation request quantity is as follows:

[0103] ;

[0104] Among them, is the target read time, is the critical response time, is the remaining access time, is the first critical computing power information, is the peak computing power information, is the number of requests processed per unit time, is the access time corresponding to multiple requests, is the preset allocation request quantity, is the current heterogeneous computing power device.

[0105] The determination of the above preset allocation request quantity is based on the preset allocation request quantity per unit time, that is, 1 second. The target read time, critical response time, remaining access time, and access time corresponding to multiple requests are all set in seconds as the calculation unit.

[0106] Regarding the determination of the first critical computing power information, its formula is as follows:

[0107] ;

[0108] Among them, is the number of transformer structures, is the number of attention heads, is the key vector dimension, is the value vector dimension, is the critical sequence length.

[0109] The determination process of the preset allocation request quantity for the inference task provided in this embodiment takes into account the number of requests that can be allocated per second when using key-value caching, determines the remaining access time through the difference between the key-value caching time and the critical response time, and performs subsequent normalization processing, making the entire determination process relatively accurate, taking into account the computing power information, memory expansion information, and inference task characteristic information of the heterogeneous computing power device, so as to ensure the reasonable allocation of the inference task while meeting the maximum response delay performance requirements.

[0110] In some embodiments, determining the request allocation result of the current heterogeneous computing power device according to the preset allocation request quantity, the concurrent request quantity of the inference task information, and the allocation condition includes:

[0111] When the preset allocation request quantity meets the allocation condition, determining the final allocation request quantity of the current heterogeneous computing power device according to the relationship between the preset allocation request quantity and the concurrent request quantity;

[0112] When the preset allocation request quantity does not meet the allocation condition, determining a new preset allocation request quantity for the inference task according to the target performance parameter.

[0113] Specifically, when the preset allocation request quantity meets the allocation condition, the final allocation request quantity is determined according to the relationship between the preset allocation request quantity and the concurrent request quantity. The allocation condition in this application is that the request quantity is 1. If the preset allocation request quantity is greater than or equal to 1, then the preset allocation request quantity and the concurrent request quantity are further compared to determine the final allocation request quantity. If the preset allocation request quantity does not meet the allocation condition, that is, the preset allocation request quantity is less than 1, it indicates that using the key-value cache cannot meet the latency requirement. Then, the current allocation does not use the key-value cache, and a new preset allocation request quantity is determined based on the computing power information.

[0114] The comparison between the preset allocation request quantity provided in this embodiment and the allocation condition is used to ensure whether the key-value cache can truly be used to meet the latency requirement, ensure that the current heterogeneous computing power device can meet the latency requirement while using the key-value cache, and determine the request quantity that can be allocated when not using the key-value cache, thereby improving the rationality of the allocation.

[0115] In some embodiments, determining the final allocation request quantity of the current heterogeneous computing power device according to the relationship between the preset allocation request quantity and the concurrent request quantity includes:

[0116] When the preset allocation request quantity is greater than or equal to the concurrent request quantity, determining the final allocation request quantity of the current heterogeneous computing power device according to the concurrent request quantity;

[0117] When the preset allocation request quantity is less than the concurrent request quantity, determining the final allocation request quantity of the current heterogeneous computing power device according to the preset allocation request quantity.

[0118] Specifically, considering the situation where the relationship between the preset allocation request quantity and the concurrent request quantity is not equal, when the preset allocation request quantity is greater than or equal to the concurrent request quantity, it means that using one heterogeneous computing power device can allocate all the concurrent request quantities. At this time, the concurrent request quantity is used as the final allocation request quantity of the current heterogeneous computing power device. When the preset allocation request quantity is less than the concurrent request quantity, the preset allocation request quantity is determined as the final allocation request quantity.

[0119] The final allocation request quantity of the current heterogeneous computing power device determined in this embodiment when there is an unequal relationship between the preset allocation request quantity and the concurrent request quantity ensures the correctness of the allocation result.

[0120] In some embodiments, determining the new preset allocation request quantity of the inference task according to the target performance parameter includes:

[0121] Obtaining the second critical computing power information;

[0122] Determine the new preset allocation request quantity corresponding to the current heterogeneous computing power device according to the second critical computing power information and the peak computing power information;

[0123] Determine the final allocation request quantity of the current heterogeneous computing power device according to the relationship between the new preset allocation request quantity and the concurrent request quantity.

[0124] Specifically, considering that the preset allocation request quantity is determined based on the use of the key-value cache, when the current latency requirement cannot be met and the key-value cache is not used, recalculate the new preset allocation request quantity, and its formula is as follows:

[0125] ;

[0126] Wherein, is the second critical computing power information, that is, the maximum floating-point calculation amount without KVCache; is the peak computing power information, is the new preset allocation request quantity, is the current heterogeneous computing power device.

[0127] The determination process of the second critical computing power information, the formula is as follows:

[0128] ;

[0129] Wherein, is the number of transformer structures, is the number of attention heads, is the key vector dimension, is the value vector dimension, is the critical sequence length.

[0130] After determining the new preset allocation request quantity, compare it with the concurrent request quantity to determine the final allocation request quantity of the current heterogeneous computing power device. The same as the above embodiments, it will not be elaborated here.

[0131] The new preset allocation request quantity of the inference task determined according to the target performance parameters provided in this embodiment realizes the allocation request quantity without using the key-value cache. While improving the allocation rationality, when there is an unequal relationship between the preset allocation request quantity and the concurrent request quantity, the final allocation request quantity of the current heterogeneous computing power device is determined to ensure the correctness of the allocation result.

[0132] In some embodiments, determining the final allocation request quantity of the current heterogeneous computing power device according to the preset allocation request quantity includes:

[0133] If the preset allocation request quantity has a decimal, round down the preset allocation request quantity to obtain the rounded-down preset allocation request quantity;

[0134] Use the rounded preset allocation request quantity as the final allocation request quantity.

[0135] Specifically, if there is a decimal in the preset allocation request quantity, floor rounding is required. Floating-point operations are usually more time-consuming than integer operations. Floor rounding can convert floating-point operations into integer operations, thereby improving calculation efficiency. Integers usually occupy less storage space than floating-point numbers. Floor rounding can reduce storage requirements and improve the overall performance of the system.

[0136] In some embodiments, after determining the request allocation result of the current heterogeneous computing power device, it further includes:

[0137] Obtain the final allocation request quantity of the current heterogeneous computing power device;

[0138] Determine the remaining request quantity according to the concurrent request quantity and the final allocation request quantity;

[0139] Update the read times corresponding to multiple heterogeneous computing power devices according to the read time corresponding to the current heterogeneous computing power device;

[0140] Determine a new current heterogeneous computing power device according to the updated read times of multiple heterogeneous computing power devices, and return to the step of determining the preset allocation request quantity of the inference task based on the target read time and target performance parameters of the current heterogeneous computing power device until the concurrent request quantity is completely allocated among multiple heterogeneous computing power devices.

[0141] Specifically, after determining the request allocation result of the current heterogeneous computing power device, if it is necessary to perform the allocation of the next heterogeneous computing power device, it is necessary to determine the current remaining request quantity according to the concurrent request quantity and the final allocation request quantity of the current heterogeneous computing power device, continue to update the read time relationship of multiple heterogeneous computing power devices to determine the next new current heterogeneous computing power device, and continue to return to step S13 until the concurrent request quantity is completely allocated.

[0142] The request allocation process provided in this embodiment implements a loop method until the concurrent requests of the inference task are completely allocated among multiple heterogeneous computing power devices. While considering the computing power performance differences, memory expansion, and inference task characteristics of heterogeneous computing power devices, reasonable requests are allocated to heterogeneous computing power devices, so that the above-mentioned considered factors are balanced, saving computing power resources and reducing costs.

[0143] In other embodiments, after determining the request allocation results of multiple heterogeneous computing power devices, it further includes:

[0144] Count the first request quantities corresponding to multiple heterogeneous computing power devices;

[0145] Determine the final remaining request quantity according to the concurrent request quantity and multiple first request quantities;

[0146] Determine the remaining heterogeneous computing power devices in the initial heterogeneous computing power device except the target initial heterogeneous computing power device;

[0147] Sort the reading times of the remaining heterogeneous computing power devices, determine the new current heterogeneous computing power device, and return to the step of determining the preset allocation request quantity of the inference task according to the target reading time and target performance parameters of the current heterogeneous computing power device, so as to complete the request allocation of the final remaining request quantity among the remaining heterogeneous computing power devices.

[0148] Specifically, after all the multiple heterogeneous computing power devices in step S11 are allocated requests based on the concurrent request quantity, if there are still remaining request quantities, it is necessary to count the final remaining request quantity. After using up the multiple heterogeneous computing power devices, continue to screen the remaining heterogeneous computing power devices in the original initial heterogeneous computing power device. Here, the screening can be random use, or the preset computing power information can be reduced for further screening, or it can be screened based on other factors such as other computing power information, which is not limited here. Then, sort the reading times of the determined remaining heterogeneous computing power devices for subsequent allocation, which is the same as the technical solution in step S13 and is not limited here until the final remaining request quantity is allocated.

[0149] This embodiment takes into account that even after the multiple heterogeneous computing power devices specified by the user are allocated, if there are still remaining request quantities, the screening range of the heterogeneous computing power devices is continued to be expanded, which improves the allocation efficiency compared with reallocating the multiple heterogeneous computing power devices specified for use.

[0150] Figure 3 It is a schematic diagram of a request allocation architecture based on a heterogeneous computing system provided by an embodiment of the present application. As Figure 3 shown, it includes 4 modules, namely:

[0151] 1. Heterogeneous computing system information collection module: This module regularly collects the heterogeneous computing power and discrete memory information in the heterogeneous computing system. Once a neural network is issued, this information will be sent to the inference task information collection module.

[0152] 2. Inference task information collection module: When an inference task is issued, the heterogeneous computing system collects the task information of the inference task, and this information will be sent to the inference task information collection module.

[0153] 3. Inference task allocation module: According to the collected heterogeneous computing power, discrete memory information, and the collected inference task information, this module calculates the allocation plan for the inference task and sends it to the inference task issuing module.

[0154] 4. Inference Task Distribution Module: According to the distribution plan of the inference task, deploy the inference task into the heterogeneous computing system for actual task execution.

[0155] For the heterogeneous computing system information collection module, regularly collect the computing power in the heterogeneous computing system and its corresponding disaggregated memory information.

[0156] For the inference task information collection module, when a pre-trained language model inference task is issued, it is necessary to collect the inference task information and send it to the inference task distribution module for calculating the distribution plan for this inference task.

[0157] For the inference task distribution module, based on the collected disaggregated memory information and large model inference task information, this module will execute the large language model inference task distribution algorithm for disaggregated memory. Under the condition of ensuring the longest response time that meets the user's requirements, select a suitable processing method for the arriving inference task, and perform reasonable request allocation while trying to save heterogeneous computing power resources.

[0158] For the inference task distribution module, according to the inference task distribution result of the inference task distribution module, this module will issue the inference request according to the calculated distribution result and perform actual deployment in the disaggregated memory heterogeneous computing system.

[0159] The purpose of the algorithm is to select several user-specified heterogeneous computing powers in the heterogeneous computing system for the arriving inference requests and output their request distribution results in the form of [(heterogeneous computing power identifier, whether to use KVCache, how many requests can be concurrent per second), (heterogeneous computing power identifier, whether to use KVCache, how many requests can be concurrent per second),...]. For example, if the user chooses to use MLU370 to execute the LLAMA2-13B inference task with a maximum of 10 concurrent requests per second, the final obtained distribution result may be [(MLU370_1, use KVCache, can be concurrent 6 requests), (MLU370_3, do not use KVCache, can be concurrent 4 requests)].

[0160] During the distribution process, if there are no available heterogeneous computing power devices and there are still requests to be distributed at this time, it is possible to directly return that the demand cannot be met. If there are available heterogeneous computing power devices, it is necessary to use KVCache to calculate the number of requests that can be distributed per second. The specific calculation process has been given in the above embodiments. If , then take the floor and let (if after this calculation , then let ), and obtain the distribution result of is ( , use KVCache, ). If , directly output the allocation result. If , it means that using KVCache cannot meet the latency requirement. At this time, estimate the number of new requests that can be allocated per second without using KVCache. The specific calculation process has been given in the above embodiments. If is assigned and rounded down, and let (if after this calculation , then let ), and obtain the allocation result as ( , without using KVCache, ). If , directly output the allocation result.

[0161] The sorting algorithm output is [(heterogeneous computing power identifier, whether to use KVCache, how many requests can be concurrent per second), (heterogeneous computing power identifier, whether to use KVCache, how many requests can be concurrent per second),...], and return the final result.

[0162] It should be noted that the determination of the first critical computing power information, that is, the floating-point computing amount without using KVCache.

[0163] 1. Key ( ) and Value ( ) vector cache: The previously calculated and are cached, and only the Key and Value vectors of the new token need to be calculated.

[0164] 2. Attention mechanism calculation : The dot product calculation needs to be performed on all current tokens.

[0165] 3. Attention output calculation: Weighted sum of the value vectors of all tokens.

[0166] The computing amount of a single layer (the th token):

[0167] Key and Value of the new token: ;

[0168] Dot product calculation : ;

[0169] Attention weighted sum : ;

[0170] The total computational amount of a single head is:

[0171] ;

[0172] The computational amount of [[number of heads]] heads is:

[0173] ;

[0174] The total computational amount of [[number of layers]] layers (computational amount of a single token) is:

[0175] .

[0176] The cumulative total computational amount (when generating a sequence of length [[sequence length]]) is: );

[0177] For the generated sequence length [[sequence length]], gradually accumulate from the 1st to the [[number of tokens]]th token: , from the 1st to the [[number of tokens]]th token:

[0178] ;

[0179] Split the terms:

[0180] ;

[0181] Sum of the first term:

[0182] ;

[0183] Sum of the second term:

[0184] ;

[0185] After combining:

[0186] ;

[0187] Further simplify:

[0188] .

[0189] Regarding the determination of the second critical computing power information, that is, when not using the floating-point computational amount of KVCache, for each newly generated token:

[0190] 1. Re-compute the Key ( ) and Value ( ) vectors: Each token requires re-computing the Key and Value vectors;

[0191] 2. Compute the attention mechanism : It is necessary to calculate the dot product for all previously generated tokens;

[0192] 3. Attention output calculation ( ): Weighted sum of the value vectors of all tokens.

[0193] Computational cost of a single layer (the th token):

[0194] Recalculate Key and Value: ;

[0195] Among them, is also the length of the current sequence, and each token needs to recalculate Key and Value.

[0196] Dot product calculation : ;

[0197] Attention weighted sum : ;

[0198] The total computational cost of a single head is:

[0199] ;

[0200] The computational cost of

[0201] heads is:

[0202] The total computational cost of a layer (computational cost of a single token) is:

[0203] .

[0204] Cumulative total computational cost (generating a length of ):

[0205] For the generated sequence length , gradually accumulate from the 1st to the th token:

[0206] ;

[0207] Add from 1 to :

[0208] ;

[0209] Use the formula :

[0210] 。

[0211] Regarding KVCache calculation:

[0212] 1. Content stored in KVCache:

[0213] For a multi-layer Transformer architecture:

[0214] 1). Key ( ) vector: Used to store the key values of all previously generated tokens;

[0215] 2). Value ( ) vector: Used to store the values of all previously generated tokens.

[0216] Storage capacity of a single layer:

[0217] The sizes of the Key and Value vectors cached by each attention head are respectively:

[0218] ;

[0219] Among them, is the length of the generated sequence, , are the vector dimensions of Key and Value.

[0220] Total size of the cache for each attention head:

[0221] ;

[0222] If there are attention heads, the cache size of a single layer is:

[0223] ;

[0224] 2. Total KVCache storage capacity:

[0225] For a -layer Transformer, the total storage capacity of KVCache is:

[0226] 。

[0227] Furthermore, this application also provides a heterogeneous computing system, as shown in Figure 1 including a host, multiple heterogeneous computing power devices, and multiple storage devices;

[0228] The host is respectively connected to multiple heterogeneous computing power devices; multiple heterogeneous computing power devices call multiple storage devices;

[0229] A host, configured to execute the steps of the above-mentioned request allocation method based on a heterogeneous computing system to complete the request allocation of an inference task.

[0230] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0231] An embodiment of the present application further provides a request allocation device based on a heterogeneous computing system. Figure 4 It is a structural diagram of a request allocation device based on a heterogeneous computing system provided by an embodiment of the present application. As Figure 4 shown, the request allocation device based on a heterogeneous computing system includes:

[0232] An acquisition module 11: acquiring inference task information, multiple heterogeneous computing power devices and corresponding performance parameters;

[0233] A first determination module 12: determining the required reading time of the key-value cache based on the inference task information and multiple performance parameters; and determining the current heterogeneous computing power device according to the multiple reading times;

[0234] A second determination module 13: determining the preset allocation request quantity of the inference task according to the target reading time and target performance parameters of the current heterogeneous computing power device;

[0235] A third determination module 14: determining the request allocation result of the current heterogeneous computing power device according to the preset allocation request quantity, the concurrent request quantity of the inference task information and the allocation conditions.

[0236] For the description of the features in the corresponding embodiment of the request allocation device based on a heterogeneous computing system, reference can be made to the relevant description of the corresponding embodiment of the request allocation method based on a heterogeneous computing system, which will not be elaborated here one by one.

[0237] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-mentioned embodiments of the request allocation method based on a heterogeneous computing system.

[0238] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above-mentioned embodiments of the request allocation method based on a heterogeneous computing system when running.

[0239] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0240] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the request allocation method based on a heterogeneous computing system.

[0241] Another embodiment of the present application also provides a computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the request allocation method based on a heterogeneous computing system.

[0242] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0243] The above has introduced in detail a request allocation method, system, device, and medium based on a heterogeneous computing system provided by this application. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A request allocation method based on a heterogeneous computing system, characterized in that: include: Obtain inference task information, multiple heterogeneous computing devices, and corresponding performance parameters; Determine the read time required for the key-value cache based on the inference task information and multiple performance parameters; And determine the current heterogeneous computing power device based on multiple reading times; Determine a preset number of allocation requests for the inference task according to the target read time and target performance parameters of the current heterogeneous computing power device; The request allocation result of the current heterogeneous computing power device is determined according to the preset allocation request number, the concurrent request number of the inference task information and the allocation condition.

2. The request allocation method based on heterogeneous computing system according to claim 1, characterized in that: The read time required for the key-value cache is determined based on the inference task information and multiple performance parameters, including: Obtaining the critical key value cache data volume required for the reasoning task information, the latency information and the bandwidth information of the performance parameters; The read time required for the corresponding key-value cache when multiple heterogeneous computing devices access the separate memory is determined according to the critical key-value cache data volume, multiple delay information and multiple bandwidth information.

3. The request allocation method based on heterogeneous computing system according to claim 2, characterized in that: Determining the read time required for the corresponding key-value cache when the plurality of heterogeneous computing devices access the separated memory according to the critical key-value cache data volume, the plurality of delay information and the plurality of bandwidth information, including: Determine multiple critical transmission delays according to the critical key value cache data amount and the multiple bandwidth information; A plurality of the reading times are determined according to the plurality of the critical transmission delays and the corresponding plurality of the delay information.

4. The request allocation method based on a heterogeneous computing system according to any one of claims 1 to 3, characterized in that: Determine the current heterogeneous computing power device based on multiple read times, including: Sorting the plurality of read times from largest to smallest to obtain a sorted read time relationship; The minimum read time is selected from the read time relationship, and the heterogeneous computing power device to which the minimum read time belongs is used as the current heterogeneous computing power device.

5. The request allocation method based on heterogeneous computing system according to claim 4, characterized in that: Determining the preset allocation request quantity of the inference task according to the target reading time and target performance parameters of the current heterogeneous computing power device includes: ; in, Read the time for the target, is the critical response time of the reasoning task information, is the remaining access time, The first critical computing power information for key-value cache, is the peak computing power information of the target performance parameter, is the number of requests processed per unit time, Allocate the requested quantity for the preset, It is the current heterogeneous computing power device.

6. The request allocation method based on heterogeneous computing system according to claim 5, characterized in that: Determining the request allocation result of the current heterogeneous computing power device according to the preset allocation request number, the concurrent request number of the inference task information and the allocation condition, including: In the case where the preset number of allocation requests meets the allocation condition, determining the final number of allocation requests of the current heterogeneous computing power device according to the relationship between the preset number of allocation requests and the number of concurrent requests; When the preset number of allocation requests does not satisfy the allocation condition, a new preset number of allocation requests for the reasoning task is determined according to the target performance parameter.

7. The request allocation method based on heterogeneous computing system according to claim 6, characterized in that: Determining the final number of allocation requests of the current heterogeneous computing power device according to the relationship between the preset number of allocation requests and the number of concurrent requests includes: When the preset number of allocation requests is greater than or equal to the number of concurrent requests, determining the final number of allocation requests of the current heterogeneous computing power device according to the number of concurrent requests; When the preset number of allocation requests is less than the number of concurrent requests, the final number of allocation requests of the current heterogeneous computing power device is determined according to the preset number of allocation requests.

8. The request allocation method based on heterogeneous computing system according to claim 6, characterized in that: Determining a new preset allocation request quantity of the reasoning task according to the target performance parameter includes: Obtaining the second critical computing power information; Determine a new preset allocation request quantity corresponding to the current heterogeneous computing power device according to the second critical computing power information and the peak computing power information; The final number of allocation requests for the current heterogeneous computing power device is determined based on the relationship between the new preset number of allocation requests and the number of concurrent requests.

9. The request allocation method based on heterogeneous computing system according to claim 7, characterized in that: Determining the final allocation request quantity of the current heterogeneous computing power device according to the preset allocation request quantity includes: If the preset allocation request quantity contains a decimal, the preset allocation request quantity is rounded down to obtain a rounded preset allocation request quantity; The rounded preset allocation request quantity is used as the final allocation request quantity.

10. The request allocation method based on heterogeneous computing system according to claim 1, characterized in that: The process of selecting the plurality of heterogeneous computing devices includes: Obtain multiple initial heterogeneous computing power devices and corresponding peak computing power information in a heterogeneous computing system; According to the relationship between the plurality of peak computing power information and the preset computing power information, selecting the target initial heterogeneous computing power device that exceeds the preset computing power information from the plurality of peak computing power information; The target initial heterogeneous computing power device is used as a plurality of the heterogeneous computing power devices.

11. The request allocation method based on heterogeneous computing system according to claim 1, characterized in that: After determining the request allocation result of the current heterogeneous computing power device, the method further includes: Obtain the final allocation request quantity of the current heterogeneous computing power device; Determine the number of remaining requests according to the number of concurrent requests and the number of final allocated requests; Update the reading times corresponding to the plurality of heterogeneous computing devices according to the reading time corresponding to the current heterogeneous computing device; Determine a new current heterogeneous computing device based on the updated reading time of the multiple heterogeneous computing devices, and return to the step of determining the preset number of allocated requests for the inference task based on the target reading time and target performance parameters of the current heterogeneous computing device until the number of concurrent requests is allocated among the multiple heterogeneous computing devices.

12. A heterogeneous computing system, characterized in that: Includes a host, multiple heterogeneous computing devices, and multiple storage devices; The host is respectively connected to a plurality of heterogeneous computing devices; the plurality of heterogeneous computing devices call a plurality of storage devices; The host is used to execute the steps of the request allocation method based on a heterogeneous computing system as described in any one of claims 1 to 11 above to complete the request allocation of the reasoning task.

13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is used to implement the steps of the request allocation method based on a heterogeneous computing system as described in any one of claims 1 to 11 when executing the computer program.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the request allocation method based on a heterogeneous computing system as claimed in any one of claims 1 to 11.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the request allocation method based on a heterogeneous computing system as claimed in any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Memory scheduling method of heterogeneous computing system, heterogeneous computing system and device

    CN118260053A

  • Inference task processing method, system and equipment, medium and program product

    CN119127482A