Computing card group scheduling method and device based on task perception and generation length prediction

Through the scheduling framework of task perception and generation length prediction, the memory management and task scheduling of the large language model inference service system are optimized, and the problems of low resource utilization and long-tail delay in high concurrency scenarios are solved, significantly improving system performance and concurrency capabilities.

CN120295737APending Publication Date: 2025-07-11INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510500087.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the scenarios of high concurrency and differentiation of requests, the traditional scheduling strategy leads to low utilization of memory resources, long tasks occupy computing resources, short tasks respond to a surge in response delays, system throughput is limited, and the accelerator card lacks pagination memory management, resulting in high memory fragmentation rate, which cannot meet the high concurrency needs.

Method used

The scheduling framework of task perception and generation length prediction is adopted, and the generation length is predicted through the dynamic task metadata perception mechanism and the lightweight BERT model, combined with the acceleration card hardware status, the video memory is allocated and scheduled on demand, and the dynamic block merging algorithm is designed to optimize the video memory utilization rate and task scheduling.

Benefits of technology

The memory utilization rate has been increased to 85%, the number of concurrent tasks has increased by 1.7 times, the response delay of short tasks has been reduced by 70%, and the system throughput has been significantly improved, meeting the needs of high-concurrency scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295737A_ABST
    Figure CN120295737A_ABST
Patent Text Reader

Abstract

The invention provides a computing card group scheduling method and device based on task perception and generation length prediction, and the method comprises the steps: sending a plurality of reasoning requests with user prompt words to a computing card group, packaging each reasoning request as a request object, and enabling the computing card group to have a plurality of intelligent computing cards, the video memories of all the intelligent computing cards form a shared video memory pool; adding the request object into a user request pool; performing generation length prediction on the user prompt word of the request object in the user request pool by adopting a predictor module to obtain a predicted generation length; creating a corresponding reasoning task and metadata for each request object in the user request pool, applying for a video memory block as required from the shared video memory pool according to the predicted generation length, obtaining a video memory block handle with a video memory allocation address and size, and writing the metadata into the video memory block handle; generating a task scheduling sequence according to the metadata of the reasoning task and the running state of each computing card; and reasoning tasks are selected from the task scheduling sequence in sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the inference acceleration technology of artificial intelligence applications, and specifically relates to the request scheduling technical field for large language model inference service systems. The present invention proposes a computing card group scheduling method, device, electronic device, computer-readable storage medium, and computer program product based on task awareness and generated length prediction. Background Art

[0002] Existing large language model inference service systems are not optimized in combination with the intelligent acceleration card hardware characteristics of the MLU Arch v3 architecture, and the acceleration technologies applicable to NVIDIA chips cannot be simply replicated to all acceleration card (computing card) environments. For example, for high-concurrency and request-differentiated scenarios, the widely used first-come-first-served scheduling strategy applied to acceleration cards will result in low utilization rate of video memory resources, and cannot meet the user service-level goals in inference scenarios, causing users to wait busy for a long time or not get a system response, and reducing the system concurrency.

[0003] Currently, the main application scenario of large language models is to provide cloud question-and-answer services. In this scenario, there are often high concurrency and large differences in user requests. For example, long tasks block the execution of short tasks, resulting in long-tail latency and head-of-line blocking problems, and long requests occupy a large amount of resources, leading to a decline in system performance. The first-come-first-served scheduling strategy used by mainstream large language model service frameworks on the market currently (such as vLLM, Orca, etc.) cannot meet the requirements of this scenario, and there may even be system crashes or users not getting a response for a long time in high-concurrency scenarios.

[0004] In addition, currently mainstream large language model inference service systems for intelligent acceleration cards often cannot be well adapted to the specific hardware characteristics of the acceleration cards. For example, for some intelligent acceleration card series that do not support the paged video memory allocation strategy, the traditional static allocation strategy is still used in the framework, resulting in a video memory fragmentation rate as high as 40%, thus reducing the maximum number of concurrent tasks of the acceleration card and resulting in low system throughput. Summary of the Invention

[0005] When the inventors deeply studied the large language model inference service system for intelligent computing card groups, they found the following key problems in the prior art: First, traditional scheduling strategies (such as first-come-first-served) do not consider the characteristics of large differences in the generated lengths of user requests in high-concurrency scenarios, resulting in long tasks occupying video memory and computing resources for a long time, a sharp increase in the response delay of short tasks, and serious limitation of system throughput; Second, intelligent edge acceleration cards lack the support of a dynamic video memory paging mechanism (such as PagedAttention of vLLM), and the existing framework uses a static video memory pre-allocation strategy, resulting in a video memory fragmentation rate as high as 40%, making it difficult to increase the number of concurrent tasks.

[0006] In response to the above problems, the inventor proposed a scheduling framework that coordinates task awareness and generation length prediction, and gradually overcame the technical difficulties. First, the inventor designed a dynamic task metadata awareness mechanism. By constructing multi-dimensional metadata information for each user request (including real-time waiting duration, predicted generation length, and user priority weight), and combining the hardware status monitoring interface of the acceleration card (such as video memory occupancy rate and computing unit load), a dynamic priority calculation model was constructed.

[0007] In terms of optimizing video memory management, due to the low-bandwidth characteristics of the acceleration card, virtual memory paging is currently not supported. Therefore, there are a large number of video memory space fragments in the card group environment. In response to this problem, the inventor proposed a video memory on-demand allocation strategy for generation length prediction. Specifically, based on a lightweight BERT model (with a parameter quantity of less than 100MB), the length of the user prompt word is predicted, and the accuracy of predicting the generation length reaches 98.41%. According to the prediction result, the allocator module requests video memory blocks on demand from the shared video memory pool of the card group. The allocator divides the video memory into multi-granularity blocks. When a request is needed, it searches for a block of the appropriate size in the video memory pool and applies for and allocates it. After the inference task is completed, the video memory block is released back to the video memory pool, and the allocator implements a dynamic block merging algorithm. After the task is completed, adjacent free blocks can be merged to form a larger video memory block to meet the requirements of the task for large video memory blocks. This mechanism increases the video memory utilization rate from 60% to 85%, and the maximum number of concurrent tasks that the system can support increases by 1.7 times.

[0008] Aiming at the deficiencies of the prior art, such as Figure 3 As shown, the present invention proposes a computing card group scheduling method based on task awareness and generation length prediction, which includes:

[0009] Initial step: Send multiple inference requests with user prompt words to the computing card group, encapsulate each such inference request as a request object. The computing card group has multiple intelligent computing cards, and the video memories of all intelligent computing cards form a shared video memory pool;

[0010] Prediction step: Add the request object to the user request pool; use the predictor module to perform generation length prediction on the user prompt words of the request objects in the user request pool to obtain the predicted generation length;

[0011] Allocation step: Create corresponding inference tasks and metadata for each request object in the user request pool. The metadata includes request arrival time, request priority, and predicted generation length, etc.; According to the predicted generation length, apply for video memory blocks on demand from the shared video memory pool, obtain a video memory block handle with a video memory allocation address and size, and write it into the metadata; Generate a task scheduling sequence according to the metadata of the inference task and the running status of each computing card;

[0012] Inference step: Sequentially select inference tasks from the task scheduling sequence, allocate intelligent computing cards corresponding to the video memory block handle for inference, generate inference results, and encapsulate them into response objects.

[0013] The described computing card cluster scheduling method based on task awareness and generation length prediction, where the request object contains information such as user prompt words, request ID, user priority label, request status, and timestamp.

[0014] The described computing card cluster scheduling method based on task awareness and generation length prediction, where when the system load of the computing card cluster is greater than the threshold, the user request pool uses the token bucket algorithm to limit the request inflow rate.

[0015] The described computing card cluster scheduling method based on task awareness and generation length prediction, where the predictor is a lightweight BERT model, fine-tuned based on a large-scale Q&A corpus, and used to predict the generation length of user prompt words.

[0016] As Figure 4 shown, the present invention also proposes a computing card cluster scheduling device based on task awareness and generation length prediction, which includes:

[0017] Initial module: Send multiple inference requests with user prompt words to the computing card cluster, encapsulate each inference request into a request object. The computing card cluster has multiple intelligent computing cards, and the video memories of all intelligent computing cards form a shared video memory pool.

[0018] Prediction module: Add the request object to the user request pool; Use the predictor module to predict the generation length of the user prompt words in the request objects in the user request pool to obtain the predicted generation length.

[0019] Allocation module: Create corresponding inference tasks and metadata for each request object in the user request pool. The metadata includes request arrival time, request priority, and predicted generation length, etc.; According to the predicted generation length, apply for video memory blocks as needed from the shared video memory pool, obtain video memory block handles with video memory allocation addresses and sizes, and write them into the metadata; Generate a task scheduling sequence according to the metadata of the inference tasks and the running status of each computing card.

[0020] Inference module: Sequentially select inference tasks from the task scheduling sequence, allocate intelligent computing cards corresponding to the video memory block handle for inference, generate inference results, and encapsulate them into response objects.

[0021] The described computing card cluster scheduling device based on task awareness and generation length prediction, where the request object contains information such as user prompt words, request ID, user priority label, request status, and timestamp.

[0022] The described computing card group scheduling device based on task awareness and generated length prediction, where when the system load of the computing card group is greater than the threshold, the user request pool uses the token bucket algorithm to limit the request inflow rate; the predictor is a lightweight BERT model, fine-tuned based on a large-scale Q&A corpus, and performs generated length prediction on user prompt words.

[0023] The present invention also proposes an electronic device, which includes the described computing card group scheduling device based on task awareness and generated length prediction. The electronic device is either connected to an information display device, which is used to display the inference result with display parameters, attributes set by the user, or through an artificial intelligence model.

[0024] The present invention also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the computing card group scheduling method based on task awareness and generated length prediction.

[0025] The present invention also proposes a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the computing card group scheduling method based on task awareness and generated length prediction.

[0026] From the above solutions, the advantages of the present invention are as follows:

[0027] The present invention synergistically optimizes the large language model inference service system for intelligent computing card groups through task awareness and generated length prediction, significantly improving the resource utilization efficiency and concurrency ability of the system. Aiming at the long-tail latency and system throughput bottleneck problems caused by traditional scheduling strategies in high-concurrency scenarios, a fine-grained dynamic scheduling mechanism based on task metadata is proposed. Combining real-time hardware status monitoring and a task priority calculation model, intelligent scheduling of tasks is achieved, reducing the response latency of short tasks by 70% and improving the user experience. At the same time, aiming at the limitation of the lack of paged video memory management in intelligent acceleration cards, a video memory on-demand allocation strategy based on generated length prediction is designed. Through a lightweight BERT prediction model, a generated length prediction accuracy of 98.41% is achieved, increasing the video memory utilization rate from 60% to 85%, increasing the number of concurrent tasks by 1.7 times, significantly reducing the video memory fragmentation problem, and improving the system throughput capacity. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a schematic diagram of a fine-grained dynamic scheduling mechanism based on inference task awareness;

[0029] Figure 2 It is a flowchart of an embodiment of the present invention;

[0030] Figure 3 It is a flowchart of the method of the present invention;

[0031] Figure 4 This is the block diagram of the device of the present invention;

[0032] Figure 5 This is the schematic structural diagram of the first electronic device of the present invention;

[0033] Figure 6 This is the schematic structural diagram of the application environment of the first electronic device of the present invention;

[0034] Figure 7 This is the schematic structural diagram of the second electronic device of the present invention.

[0035] Reference numerals:

[0036] A - The first electronic device;

[0037] B - The computing card group scheduling device based on task perception and generated length prediction;

[0038] C - The data acquisition device;

[0039] D - The information display device;

[0040] 1000 - The second electronic device;

[0041] Ⅰ - The computing unit;

[0042] Ⅱ - ROM;

[0043] Ⅲ - RAM;

[0044] Ⅳ - The bus;

[0045] Ⅴ - The interface;

[0046] Ⅵ - The input unit;

[0047] Ⅶ - The output unit;

[0048] Ⅷ - The storage medium;

[0049] Ⅸ - The communication unit. Detailed implementation manners

[0050] It should be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0051] Without further limitation, an element qualified by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0052] The processor according to the present invention is the control center of an electronic device, which may be a single processor or a collective term for multiple processing elements. For example, it may be one or more central processing units (CPUs), or it may be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0053] Optionally, the processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0054] In a specific implementation, as an embodiment, the processor may include one or more CPUs. Each of these processors may be a single-CPU or a multi-CPU. Here, the processor may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The electronic device may include: servers, desktop computers, laptop computers, smart phones, tablet computers, embedded computers, etc., where the embedded computer includes vehicles and robots, etc.

[0055] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation manner may refer to the above method embodiments and will not be elaborated here.

[0056] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto. The actual knowledge structure recognition device may include more or fewer components than shown in the drawings, or combine certain components, or have different component arrangements.

[0057] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more sets of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0058] It should also be understood that the term "and / or" herein is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Herein, A and B may be singular or plural. In addition, the character " / " in this text generally represents an "or" relationship between the associated objects before and after, but may also represent an "and / or" relationship, which can be specifically understood with reference to the context.

[0059] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0060] It should also be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0061] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0062] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0063] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0064] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0065] In view of the above problems, the present invention proposes a scheduling framework for task awareness and generated length prediction in a high-concurrency scenario (multiple inference tasks are input into a computing card cluster), designs an inference length predictor for complex user requests, and designs request metadata information for each user request. The metadata information includes information such as request waiting time, expected generated length of the request, request priority, etc., and designs a perception scheduler for task metadata. This scheduler can perform scheduling by combining the request metadata information, rather than only scheduling based on the system arrival time of the user request, which is a single metadata information. This scheduling strategy can reduce the short-task response delay by 70%.

[0066] In addition, in view of the problem that some edge acceleration cards based on the MLU Arch v3 architecture currently do not support the page-based video memory allocation mechanism, the present invention designs a lightweight generated length predictor with a prediction accuracy exceeding 98.41%. Through this predictor, the length of the user's prompt can be predicted. Combining the prediction result with the video memory management interface of the acceleration card, dynamic on-demand allocation of video memory space is realized, the video memory utilization rate is increased to 85%, and the number of concurrent tasks is increased by 1.7 times. This strategy greatly reduces the waste of video memory fragmentation and improves the system concurrency and throughput.

[0067] The hardware environment targeted by the present invention is an efficient intelligent computing card cluster composed of multiple intelligent computing acceleration cards. The architecture of this card cluster has the ability of dynamic expansion, and can flexibly increase or decrease acceleration cards according to computing requirements to achieve elastic scaling of computing resources. The task-aware scheduler designed in the present invention is implemented in the large language model inference service system. The scheduler is the core design module of the modern inference service system, and this scheduler needs to interact with the system runtime, which requires the design of complex interaction logic.

[0068] Such as Figure 1As shown, this is a fine-grained dynamic scheduling mechanism based on inference task awareness. By integrating task metadata information, this mechanism dynamically optimizes the allocation and execution strategies of inference tasks, thereby significantly improving resource utilization efficiency and service quality. The task-aware scheduler continuously monitors and updates task metadata during the execution of inference tasks. At the same time, combined with the real-time status of the intelligent acceleration card (XPU status) provided by the supervisor, the task profiler determines the currently prioritized tasks to be scheduled, so as to achieve efficient resource scheduling and allocation. In addition, the present invention proposes a length predictor module (predictor), which can accurately predict the generation length of inference requests, with a prediction accuracy of up to 98.41%, providing a key basis for video memory allocation and scheduling decisions. At the same time, the metadata information will be dynamically adjusted according to the system runtime status and the requirements of specific inference tasks, thereby further enhancing the flexibility and adaptability of the system. The functions of each component will be introduced one by one below:

[0069] User Request Pool:

[0070] The user request pool is the interface between the system and external users, responsible for receiving, encapsulating, and managing the inference requests sent by users. Its core functions include request reception, encapsulation, and traffic control. When a user sends an inference request through a remote interface, the system first performs preliminary encapsulation on the request, wrapping the prompt provided by the user into a unified request object (Request Object). The request object contains the following key information: the specific prompt content, the unique identifier of the request (Request ID), the user priority label (such as VIP or ordinary user), the request status (such as waiting, processing, completed, or timed out), and the timestamp (recording the request arrival time, waiting duration, and expected timeout threshold). After encapsulation, the request object is added to the user request pool for further processing. To handle high-concurrency scenarios, the user request pool also implements a traffic shaping and admission control mechanism. When the number of concurrent requests exceeds the system load threshold, the request pool will start a waiting queue based on the token bucket algorithm to limit the request inflow rate and prevent the system from overloading and crashing. For high-priority requests (such as real-time conversations or VIP user requests), the request pool will set up a green channel for them to ensure that they skip the waiting queue and directly enter the task pool, thus meeting the low-latency service requirements.

[0071] Task Pool:

[0072] The task pool is the intermediate layer for the transformation of user requests into schedulable tasks. It is responsible for further abstracting and encapsulating the requests, and initializing the metadata and parameters required for the tasks. Its core functions include task instantiation, metadata initialization, and lifecycle management. The task pool creates corresponding inference tasks for each request object and initializes the metadata information of the tasks, including the predicted generation length (provided by the predictor module), the video memory block handle (recording the video memory allocation address and size), the sampling parameters (such as the temperature coefficient, Top-K / P sampling threshold), and the task priority weight.

[0073] Predictor module (Predictor):

[0074] It is responsible for predicting the generation length of the user prompt and recording the prediction result in the metadata information of the task, providing an important basis for dynamic video memory allocation and task scheduling. The predictor module uses a lightweight BERT model to deeply understand the semantics of the user prompt and predicts the length bucket category of the output sequence based on semantic features, so as to determine the generation length range. To improve the prediction accuracy, the predictor module fine-tunes the pre-trained model parameters based on a large-scale Q&A corpus, enabling it to maintain a prediction accuracy of up to 98.41% in the edge computing scenario.

[0075] Allocator module (Allocator):

[0076] The allocator module is responsible for perceiving and managing the intelligent computing card group system. The video memories of multiple acceleration cards in the card group jointly form a logically shared video memory pool. The allocator module can apply for and release video memory blocks from the video memory pool according to needs, so as to achieve efficient utilization of video memory. The predictor module first predicts the length of the user request and sends the generation length information to the allocator module. The allocator module applies for a video memory block of the corresponding size from the shared video memory pool according to the received length information, completes the allocation of the video memory block, and then interacts with the system runtime module (such as passing information such as the address or handle of the video memory block to the runtime module). When the user request is processed, the runtime module will notify the allocator module. The allocator releases the corresponding video memory block back to the shared video memory pool according to the data provided by the runtime module (when the task is completed, the runtime module will notify the allocator module to inform which video memory blocks can be released) and the generation length information, realizing the reuse of video memory, thereby improving the video memory utilization rate.

[0077] To further optimize video memory management, the allocator module divides the video memory into multi-granularity blocks. When video memory needs to be allocated, the allocator searches the video memory pool for a block of the appropriate size for allocation; when the inference task is completed, the allocator releases the video memory block back to the video memory pool. In addition, the allocator module also implements a dynamic block merging algorithm that can merge adjacent free blocks after the task is completed to form larger video memory blocks to meet the requirements of subsequent tasks for large video memory blocks. This mechanism not only improves the utilization rate of video memory but also enhances the flexibility and scalability of the system.

[0078] Task Profiler:

[0079] The Task Profiler is one of the core modules of the task-aware scheduler and is responsible for in-depth analysis and priority sorting of tasks in the task pool. The Task Profiler dynamically calculates the priority weights of each task by integrating task metadata information (such as predicted generation length, waiting time, user priority, etc.) and the real-time status of the accelerator card provided by the monitor (such as video memory occupancy, computing unit load, etc.). Its core functions include metadata aggregation, dynamic priority calculation, and scheduling sequence generation. The Task Profiler obtains the metadata of all tasks to be scheduled in the task pool in real time, including predicted generation length, waiting time, user priority, video memory pre-allocation status, etc., and at the same time receives the real-time status of the accelerator card reported by the monitor (such as remaining video memory capacity, computing unit utilization, task queue depth).

[0080] Supervisor:

[0081] As the bridge between the hardware state and the system runtime, it is responsible for real-time collecting and reporting the resource usage of the accelerator card. Its core functions include resource status monitoring, anomaly detection and recovery, and performance data statistics. The Supervisor obtains key metrics such as video memory occupancy, computing unit load, and task queue depth in real time through the hardware interface of the accelerator card and updates them to the Task Profiler and the Task Dispatcher at a millisecond frequency. When anomalies such as video memory overflow, computing unit overload, or task timeout are detected, it immediately triggers a preemption mechanism to suspend low-priority tasks and release resources to ensure the smooth execution of high-priority tasks. At the same time, the Supervisor records performance metrics such as throughput, average latency, and video memory utilization during system operation to provide data support for subsequent scheduling strategy optimization.

[0082] Dispatcher:

[0083] The Task Dispatcher is the execution module of the scheduler and is responsible for distributing the scheduling sequence generated by the task analyzer to the system runtime. Its core functions include task group scheduling, as well as task preemption and recovery. The Task Dispatcher dynamically adjusts the parallel scale of the task group according to the candidate task list provided by the task analyzer and the video memory status reported by the monitor. For example, when the remaining capacity of the video memory is low, short tasks are preferentially scheduled to release resources. When the monitor detects that a high-priority task needs to be executed immediately and the load on the current card group has reached the upper threshold, the Task Dispatcher will trigger the preemption mechanism, save the context state of the current task and release resources, and resume the execution of the original task after the high-priority task is completed.

[0084] Specifically, in order to achieve the above technical effects, the present invention proposes the following key technical points:

[0085] Key Point 1, a fine-grained dynamic scheduling technology for task perception in the high-concurrency scenario of intelligent computing card groups. Aiming at the long-tail latency effect in the scheduling strategy of the existing large language model service system, this technology proposes a more efficient scheduling optimization scheme to better support high-concurrency and differentiated business scenarios. First, a dynamic priority calculation model is adopted, combined with task metadata (such as predicted generation length, waiting time, user priority, etc.) to optimize the scheduling strategy, and short tasks are preferentially processed, thereby reducing the average response latency of short tasks by 70%, significantly improving the user experience. Second, by real-time monitoring the status of the acceleration card (such as video memory occupancy rate, computing unit load, etc.), the task scheduling strategy is dynamically adjusted to avoid resource overload or idleness and achieve load balancing.

[0086] Key Point 2, a method for on-demand video memory allocation for intelligent computing card groups. The present invention introduces a video memory allocation strategy based on predicted generation length, significantly improving the video memory utilization rate and optimizing the overall concurrency of the system. First, the predicted generation length of the generated task is predicted by the predictor module, and the allocator realizes the on-demand allocation of video memory resources according to the generation length obtained by the predictor, increasing the video memory utilization rate from 60% to 85% and effectively reducing the video memory fragmentation problem. Second, since the on-demand video memory allocation mechanism improves the video memory utilization rate, the maximum number of concurrent tasks supported by the system is increased by 1.7 times, enhancing the high-concurrency processing ability of the system. In addition, a high-precision generation length predictor is constructed by means of a lightweight BERT model, and the prediction accuracy is as high as 98.41%, providing a reliable decision basis for video memory allocation.

[0087] To make the above features and effects of the present invention more clearly understood, specific embodiments are given below and will be described in detail in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are only for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims. AsFigure 2 As shown in the figure, the present invention includes:

[0088] Step 1: The system enables the listening interface and waits for the user to send a request

[0089] After the system starts, it first opens the listening interface and waits for the user to send an inference request through the remote interface. The listening interface is responsible for receiving the user prompt (prompt) and encapsulating it into a unified request object (Request Object). The request object contains information such as the user prompt (prompt), request ID, user priority label, request status, and timestamp.

[0090] Step 2: Request encapsulation and traffic control

[0091] After the user request arrives at the system, the system initially encapsulates the request (prompt) into a request object and then adds it to the user request pool (Request Pool). The user request pool is an encapsulation of a data structure, for example, an array. One element in the array is a user request, which is stored in the CPU's memory waiting for scheduling. To handle high-concurrency scenarios, the request pool starts a traffic shaping and admission control mechanism. When the number of concurrent requests exceeds the system load threshold, the request pool uses the token bucket algorithm to limit the request inflow rate to avoid system overload. High-priority requests (such as VIP user requests) will skip the queuing queue and enter the task pool first to ensure low-latency service requirements.

[0092] If it is detected that the system load is too high (such as the video memory occupancy rate exceeding 90% or the computing unit load exceeding 95%), then hardware status monitoring and anomaly detection are performed, triggering the resource release module and task preemption mechanism, and then jumping back to Step 2 to execute again.

[0093] Step 3: Task instantiation and metadata initialization

[0094] After the request enters the task pool (TaskPool), the system creates a corresponding inference task for each request object and initializes the metadata information of the task. The metadata includes the predicted generation length (provided by the predictor module), video memory block handle, sampling parameters (such as temperature coefficient, Top-K / P sampling threshold), and task priority weight. The task pool is responsible for managing the entire life cycle of the task, from creation to completion or timeout.

[0095] Step 4: Generation length prediction and video memory block allocation

[0096] The Predictor module predicts the generation length of the user's prompt. The predictor uses a lightweight BERT model, which is fine-tuned based on a large-scale Q&A corpus, predicts the length bucket category of the output sequence, and records the prediction result in the metadata information of the task. The bucket category is a technique that discretizes the generation sequence length, which is a continuous value, into multiple intervals (buckets). This bucketing operation can simplify the learning task of the model, especially suitable for scenarios that require classifying or predicting continuous values. For example, if the generation length range is 1 - 10000, training a classification model for this case would result in the model's output categories being 1 - 10000, which is a large number of categories and makes training more complex. Now, I perform bucket classification on the length. There are a total of 10 buckets. Bucket 1 corresponds to the length range 1 - 1000, and bucket 2 is 1001 - 2000. In this way, the output categories of the trained model are 1 - 10 (each corresponding to a bucket), making training simpler and the prediction accuracy higher.

[0097] After fine-tuning, the prediction accuracy of the Predictor module is as high as 98.41%, which provides an important basis for on-demand video memory allocation and task scheduling. Then, the Allocator module requests a video memory block from the shared video memory pool based on the predicted length information of the prediction module, and hands over information such as the video memory block handle and length to the system runtime, waiting for the system runtime to perform the requested inference and return the inference completion information. For example, if the output of the predictor is bucket 2, it means the generation length range of this request is 1001 - 2000, and the size of the requested video memory block can be set to 2000.

[0098] Jump: If the Predictor module detects an abnormal prompt length (such as exceeding the preset maximum length threshold), it directly jumps to step 7 (Result Generation and Return), returns an error prompt message to the user, and releases the relevant resources.

[0099] Step 5: Task Analysis and Priority Sorting

[0100] The Task Profiler obtains the metadata of all tasks to be scheduled from the task pool, including the predicted generation length, waiting time, user priority, video memory pre-allocation status, etc. At the same time, the Task Profiler receives the real-time status of the accelerator card reported by the Supervisor (such as the remaining video memory capacity, computing unit utilization rate, task queue depth, etc.). Based on this information, the Task Profiler dynamically calculates the priority weight of each task and generates a scheduling sequence.

[0101] Step 6: Inference Task Distribution and Execution

[0102] The task dispatcher adjusts the parallel scale of the task group dynamically according to the scheduling sequence generated by the task analyzer and in combination with the video memory status reported by the monitor. When the task dispatcher distributes tasks to the system, the accelerator card is called for inference during system operation. If tensor parallelism is adopted, multiple cards will participate in the inference, and then the inference results are calculated through collective communication primitives and stored on specific video memory blocks.

[0103] Check whether an end marker is generated. If not, jump to step 5 to iteratively generate the result.

[0104] Step 7: Result generation and return

[0105] After the inference task is completed, the system generates the inference result and encapsulates the result into a response object (ResponseObject). The response object contains information such as the generated text content, request ID, task status (such as success or failure), and timestamp. The system returns the response object to the user to complete the processing flow of the entire inference request.

[0106] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiment.

[0107] As Figure 4 shown, the present invention also proposes a computing card group scheduling device based on task awareness and generated length prediction, which includes:

[0108] An initial module that sends multiple inference requests with user prompt words to the computing card group, encapsulates each such inference request into a request object. The computing card group has multiple intelligent computing cards, and the video memories of all intelligent computing cards form a shared video memory pool;

[0109] A prediction module that adds the request object to the user request pool; uses a predictor module to perform generated length prediction on the user prompt words of the request objects in the user request pool to obtain the predicted generated length;

[0110] An allocation module that creates corresponding inference tasks and metadata for each request object in the user request pool. The metadata includes request arrival time, request priority, and predicted generated length, etc.; according to the predicted generated length, applies for video memory blocks as needed from the shared video memory pool, obtains a video memory block handle with video memory allocation address and size and writes it into the metadata; generates a task scheduling sequence according to the metadata of the inference task and the running status of each computing card;

[0111] The inference module sequentially selects inference tasks from the task scheduling sequence, allocates the intelligent computing card corresponding to the video memory block handle for inference, generates inference results, and encapsulates them into response objects.

[0112] The computing card group scheduling device based on task awareness and generation length prediction, wherein the request object includes information such as user prompt words, request ID, user priority label, request status, and timestamp.

[0113] The computing card group scheduling device based on task awareness and generation length prediction, wherein when the system load of the computing card group is greater than the threshold, the user request pool uses the token bucket algorithm to limit the request inflow rate; the predictor is a lightweight BERT model, which is fine-tuned based on a large-scale Q&A corpus to predict the generation length of user prompt words.

[0114] As Figure 5 shown, in another embodiment of the present invention, a first electronic device A is further proposed, which includes the computing card group scheduling device based on task awareness and generation length prediction.

[0115] As Figure 6 shown, the first electronic device A can also be connected to the data acquisition device C and the information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to acquire inference requests, such as the picture classification request of the present invention. The information display device D is used to display the inference results obtained by the analysis of the present invention, such as the specific results of image classification, such as whether there is a certain object in the picture.

[0116] Among them, the information display device D can process and organize the data output by the first electronic device A based on the information display mechanism to improve the readability of the data output by the first electronic device A. The information display mechanism can be preset manually. For example, the data output by the first electronic device A is visually displayed, which can display according to the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll and play, etc. Present the key information specified by the user to the user, so that the user can understand this information more timely without having to access the secondary page or scroll the page, saving the user's operation. Or the information display mechanism can be an artificial intelligence AI display model, which can learn the key information that the user focuses on according to the user's previous usage habits, such as viewing duration, click times, editing times, etc., and then automatically present rich and necessary key information to the user.

[0117] The present invention also provides a computer program product, which includes a computer program that can be stored on a readable storage medium. When the computer program is executed by a processor, the computer can execute the computing card group scheduling method based on task awareness and length prediction provided by the above-mentioned various methods.

[0118] In another embodiment, the present invention further proposes a storage medium VIII for storing a computer program for executing the computing card group scheduling method based on task awareness and length prediction. It should be understood that the storage medium in the embodiments of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0119] Figure 7FIG. shows a schematic block diagram of a second electronic device 1000 that can be used to implement an embodiment of the present invention. The second electronic device 1000 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present invention described and / or claimed herein. The second electronic device 1000 can be the same as or different from the first electronic device A.

[0120] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory II (ROM) or a computer program loaded from a storage medium VIII into a random access memory (RAM) III. In the RAM III, various programs and data required for the operation of the device 1000 can also be stored. The computing unit I, the ROM II, and the RAM III are connected to each other via a bus IV. An input / output (I / O) interface V is also connected to the bus IV.

[0121] A plurality of components in the second electronic device 1000 are connected to the I / O interface V, including: an input unit VI, such as a keyboard, a mouse, etc.; an output unit VII, such as various types of displays, speakers, etc.; a storage medium VIII, such as a magnetic disk, an optical disk, etc.; and a communication unit IX, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0122] The computing unit I can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit I include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit I executes the various methods and processes described above, such as method steps S1 - S4. For example, in some embodiments, the method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via ROM II and / or communication unit IX. When the computer program is loaded into RAM III and executed by the computing unit I, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit I can be configured to execute the method in any other suitable way (e.g., by means of firmware).

[0123] Although the embodiments of the present invention have been disclosed as above, they are not limited to only the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the illustrated and described examples here.

Claims

1. A computing card group scheduling method based on task perception and generated length prediction, characterized in that Including: Initial step: Send multiple inference requests with user prompt words to the computing card group, encapsulate each such inference request as a request object. The computing card group has multiple intelligent computing cards, and the video memories of all intelligent computing cards form a shared video memory pool; Prediction step: Add the request object to the user request pool; Use a predictor module to perform a generation length prediction on the user prompt words of the request objects in the user request pool to obtain a predicted generation length; Allocation step: Create corresponding inference tasks and metadata for each request object in the user request pool. The metadata includes request arrival time, request priority, and predicted generation length, etc.; According to the predicted generation length, apply for video memory blocks as needed from the shared video memory pool, obtain a video memory block handle with a video memory allocation address and size and write it into the metadata; Generate a task scheduling sequence according to the metadata of the inference tasks and the running states of each computing card; Inference step: Select inference tasks from the task scheduling sequence in order, and allocate the intelligent computing card corresponding to the video memory block handle to perform inference, generate an inference result and encapsulate it as a response object.

2. The computing card group scheduling method based on task awareness and generated length prediction according to claim 1, wherein The request object contains information such as user prompt words, request ID, user priority label, request status, and timestamp.

3. The computing card group scheduling method based on task perception and generated length prediction according to claim 1, characterized in that When the system load of the computing card group is greater than the threshold, the user request pool uses the token bucket algorithm to limit the request inflow rate.

4. The computing card group scheduling method based on task perception and generated length prediction according to claim 1, wherein The predictor is a lightweight BERT model, which is fine-tuned based on a large-scale Q&A corpus to perform a generation length prediction on user prompt words.

5. A computing card group scheduling device based on task perception and generated length prediction, characterized in that Including: Initial module: Send multiple inference requests with user prompt words to the computing card group, encapsulate each such inference request as a request object. The computing card group has multiple intelligent computing cards, and the video memories of all intelligent computing cards form a shared video memory pool; Prediction module: Add the request object to the user request pool; Use a predictor module to perform a generation length prediction on the user prompt words of the request objects in the user request pool to obtain a predicted generation length; Allocation module: Create corresponding inference tasks and metadata for each request object in the user request pool. The metadata includes request arrival time, request priority, and predicted generation length, etc.; According to the predicted generation length, apply for video memory blocks as needed from the shared video memory pool, obtain a video memory block handle with a video memory allocation address and size and write it into the metadata; Generate a task scheduling sequence according to the metadata of the inference tasks and the running states of each computing card; Inference module: Select inference tasks from the task scheduling sequence in order, and allocate the intelligent computing card corresponding to the video memory block handle to perform inference, generate an inference result and encapsulate it as a response object.

6. The computing card group scheduling device based on task perception and generated length prediction according to claim 5, characterized in that, The request object contains information such as user prompt words, request ID, user priority label, request status, and timestamp.

7. The computing card group scheduling device based on task perception and generated length prediction according to claim 5, wherein, When the system load of the computing card group is greater than the threshold, the user request pool uses the token bucket algorithm to limit the request inflow rate; The predictor is a lightweight BERT model, which is fine-tuned based on a large-scale Q&A corpus to perform a generation length prediction on user prompt words.

8. An electronic device, characterized in that, Including a computing card cluster scheduling device according to any one of claims 5-7 based on task awareness and generated length prediction, the electronic device is connected to an information display device, and the information display device is configured to display the inference result with display parameters, attributes set by the user, or through an artificial intelligence model.

9. A computer-readable storage medium having a computer program stored thereon, the computer program, when executed by a processor, implements the steps of the computing card cluster scheduling method according to any one of claims 1-4 based on task awareness and generated length prediction.

10. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the computing card cluster scheduling method according to any one of claims 1-4 based on task awareness and generated length prediction.

Citation Information

Cited By

  • Task allocation method, device and equipment, storage medium and computer program product

    CN120762926A

  • Text reasoning method, product, equipment and storage medium

    CN120849131A

  • GPU video memory isolation and scheduling method between containers

    CN121029324A

  • Multi-model dynamic scheduling method for high-concurrency dialogue requests

    CN122316980A

  • Multi-model dynamic scheduling method for high-concurrency dialogue request

    CN122316980B