Large model service-oriented task parallel processing intelligent scheduling method and system
The generation length is predicted by the LaBSE model and LightGBM regressor, combined with semantic integrity constraints and discrete binary particle swarm algorithm to optimize task scheduling, the problem of slow response speed and unreasonable resource configuration of large language model services is solved, and efficient parallel processing and resource utilization are achieved.
Patent Information
- Application Number
- CN202510740546.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
Large language model services are slow to respond in scenarios with high concurrency or high real-time requirements, and the computing complexity and network delay lead to a decline in user experience, unreasonable resource allocation leads to waste of resources or insufficient performance, and existing optimization algorithms are difficult to achieve efficient optimization.
Semantic features are extracted and embedded vectors are generated, and the dimension is reduced through the compression module, combined with the LightGBM regressor to predict the generation length, the text segmentation algorithm based on semantic integrity constraints divides long text into fragments, and uses multi-constrained task batch algorithm and improved discrete binary particle swarm algorithm to optimize task scheduling and reasonably allocate it to heterogeneous cloud resources.
It significantly improves the inference speed and resource utilization of large language model services, reduces calculation costs and delays, ensures the integrity and accuracy of output results, and provides an efficient and stable user experience.
Smart Images

Figure CN120256068A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of service computing, and particularly relates to an intelligent scheduling method and system for task parallel processing for large model services. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Large language model services are intelligent services based on deep learning and natural language processing technologies, which provide users with various types of services such as accurate text generation, dialogue answering, language translation, and code assistance through pre-trained large models. However, with the continuous expansion of the model scale and the number of users, the computing resources and storage requirements for deploying and executing these models have increased sharply, which has brought huge cost and resource burdens to service providers. To address these issues, service providers deploy large language models to the cloud and provide inference services to users through application programming interfaces (APIs). This service model is called Model as a Service (MaaS), which not only reduces the upfront investment and operation and maintenance costs of providers, but also promotes the efficient utilization and large-scale application of models. In addition, this model also supports on-demand invocation of resources, elastic expansion, and continuous update of models, ensuring that users can always use the optimal version of the model without having to maintain and update it themselves.
[0004] The model-as-a-service model has achieved cost optimization for large language model services, but there are certain problems and challenges in practical applications. First, the problem of slow service response speed is prominent. Especially in scenarios with high concurrency or high real-time requirements, the computational complexity and network latency of the model may lead to a decline in user experience. Second, due to the scale and complexity of large models, their model behaviors and outputs are difficult to predict, resulting in the difficulty of existing service system optimization algorithms to achieve efficient optimization of large model services in a short time, increasing the risk of uncertainty. Finally, the resource allocation of the cloud platform may not accurately match the task requirements, resulting in unreasonable resource configuration, which in turn causes problems such as resource waste or insufficient performance.
[0005] The prior art mainly includes methods carried out around three levels: data level, model level, and system level. Data-level optimization mainly reduces the computational burden through input compression and improves the inference parallelism by using output organization. Although it can effectively reduce the computational overhead, its applicability to different tasks has certain limitations. Model-level optimization mainly involves methods such as model quantization, pruning, structure optimization, and dynamic inference. For example, quantization technology reduces storage and computational costs through low-bit precision, but it may lead to a decrease in accuracy and requires additional compensation strategies; pruning can reduce redundant parameters and improve the inference speed, but a high pruning rate may affect the model performance. System-level optimization improves the hardware utilization rate through the optimization of the inference engine and inference system. For example, methods based on speculative decoding and FlashAttention reduce the computational overhead, but they are limited by specific architectures and are not universal; distributed inference and batch processing optimization can improve the throughput, but their scheduling mechanisms and resource management methods are too complex. Generally speaking, the prior art of large language model services has problems such as slow inference speed, high computational cost and storage overhead, low throughput, and high inference latency. Summary of the Invention
[0006] To overcome the deficiencies of the above prior art, the present invention proposes a task parallel processing intelligent scheduling method and system for large model services, comprehensively optimizing the service efficiency of large language models, improving the inference speed and system throughput, and fully utilizing user data characteristics and system resources to achieve double-level optimization acceleration at the data level and system level. First, to solve the problem that the generation length of requests is difficult to predict, the present invention proposes a large language model task generation length prediction model based on LaBSE. This model uses the LaBSE model to extract semantic features and generate embedding vectors, reduces the vector dimension through a compression module, and inputs the low-dimensional vectors into a LightGBM regressor for generation length prediction to improve the accuracy and efficiency of generation length prediction. Second, based on the generation length prediction results, a text segmentation algorithm based on semantic integrity constraints is proposed. Under the semantic integrity constraints, this algorithm divides long texts into multiple segments according to parallel requirements for parallel processing. On this basis, a multi-constraint task batching algorithm is formulated to allocate text segments with similar lengths to the same batch, reducing the waiting time of tasks in the same batch. Finally, to improve the system throughput, a batch task allocation algorithm for heterogeneous cloud resources is proposed. This algorithm uses an improved discrete binary particle swarm optimization algorithm, combines an adaptive strategy and local search optimization, and reasonably allocates batch tasks to a heterogeneous virtual machine cluster to minimize the completion time of batch tasks under memory constraints. Through the above methods, the efficient parallel processing of large models is realized, significantly improving the inference speed, resource utilization rate, and the intelligent level of task scheduling, while reducing the computational cost and latency.
[0007] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions: In the first aspect, an intelligent scheduling method for task parallel processing for large model services is disclosed, including: Receiving the instruction and text of the user task, and judging whether the text can be split into subtasks for parallel execution according to the instruction to obtain parallel subtasks; Using the LaBSE model to extract semantic features from the instructions and text of the parallel subtasks and generate embedding vectors, and reducing the vector dimension of the embedding vectors through a compression module to obtain compressed embedding vectors, and inputting the compressed embedding vectors into a LightGBM regressor to predict the generated length; Based on the generated length, constructing a segmentation algorithm based on semantic integrity constraints to segment the text into a list of text segments and tags; Using a multi-constraint task batching algorithm, according to the length similarity constraint and the same instruction constraint, dividing the segmented text segments in the text segment list into tasks in different batches to obtain a set of batch tasks; Using a task allocation algorithm for heterogeneous resources to optimize the set of batch tasks and determine an optimized subtask scheduling plan; According to the tags of the parallel subtasks, splicing the subtask scheduling plans of the same task in sequence to obtain a task parallel processing scheduling plan.
[0008] In the second aspect, an intelligent scheduling system for task parallel processing for large model services is disclosed, including: A task receiving module, which is configured to: receive the instruction and text of the user task, and judge whether the text can be split into subtasks for parallel execution according to the instruction to obtain parallel subtasks; A generated length prediction module, which is configured to: use the LaBSE model to extract semantic features from the instructions and text of the parallel subtasks and generate embedding vectors, and reduce the vector dimension of the embedding vectors through a compression module to obtain compressed embedding vectors, and input the compressed embedding vectors into a LightGBM regressor to predict the generated length; A text segmentation module, which is configured to: based on the generated length, construct a segmentation algorithm based on semantic integrity constraints to segment the text into a list of text segments and tags; A task batching module, which is configured to: use a multi-constraint task batching algorithm, according to the length similarity constraint and the same instruction constraint, divide the segmented text segments in the text segment list into tasks in different batches to obtain a set of batch tasks; A scheduling optimization module, which is configured to: use a task allocation algorithm for heterogeneous resources to optimize the set of batch tasks and determine an optimized subtask scheduling plan; A scheduling generation module, which is configured to: splice the sub-task scheduling schemes of the same task in sequence according to the tags of parallel sub-tasks to obtain a task parallel processing scheduling scheme.
[0009] In a third aspect, an electronic device is disclosed, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps of the above-mentioned task parallel processing intelligent scheduling method for large model services are completed.
[0010] In a fourth aspect, a computer-readable storage medium is disclosed, which is used to store computer instructions. When the computer instructions are executed by a processor, the steps of the above-mentioned task parallel processing intelligent scheduling method for large model services are completed.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. A large language model task generation length prediction model based on LaBSE: By extracting semantic features through the LaBSE model and combining with the LightGBM regressor to accurately predict the text generation length, the accuracy and efficiency of the prediction are significantly improved, providing a reliable basis for subsequent text segmentation and batch processing, and reducing the waste of computing resources.
[0012] 2. A text segmentation algorithm based on semantic integrity constraints: Based on the predicted generation length, a text segmentation algorithm is used to cut long texts into multiple semantically complete segments, ensuring that the segmented text segments are suitable for parallel processing, effectively reducing the computational complexity, and improving the inference efficiency.
[0013] 3. A multi-constraint task batching algorithm: By the multi-constraint task batching algorithm, text segments with similar lengths and the same instructions are assigned to the same batch, reducing the use of padding tokens and the waiting time of tasks in the same batch, optimizing the service quality, and at the same time reducing the memory overhead and computational cost.
[0014] 4. A text segmentation algorithm based on semantic integrity constraints: An improved discrete binary particle swarm optimization (PSO) algorithm is used for task scheduling, combined with an adaptive strategy and local search optimization, to reasonably allocate batch tasks to virtual machine resources, minimizing the maximum completion time of tasks, and significantly improving the intelligent level of task scheduling and resource utilization.
[0015] 5. By optimizing the task processing flow and resource scheduling strategy, the present invention significantly improves the inference efficiency of large language model services, reduces the computational cost and latency, and at the same time ensures the integrity and accuracy of the output results, providing users with an efficient, stable, and low-cost large language model service experience.
[0016] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0018] Figure 1 Schematic application diagram of the task parallel processing intelligent scheduling method for large model services described in Embodiment 1 of the present invention.
[0019] Figure 2 Overall flowchart of the task parallel processing intelligent scheduling method for large model services described in Embodiment 1 of the present invention.
[0020] Figure 3 Flowchart of the large language model task generation length prediction model based on LaBSE described in Embodiment 1 of the present invention.
[0021] Figure 4 Flowchart of the text segmentation algorithm based on semantic integrity constraints described in Embodiment 1 of the present invention.
[0022] Figure 5 Flowchart of the multi-constraint task batching algorithm described in Embodiment 1 of the present invention.
[0023] Figure 6 Flowchart of the task allocation algorithm for heterogeneous resources described in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0025] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.
[0026] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0027] Embodiment 1 In one or more embodiments, an intelligent scheduling method for task parallel processing for large model services is disclosed, aiming to improve the inference efficiency of large language models and the system throughput in the cloud environment. First, an indivisible task and a divisible task are distinguished by a task type judgment module to avoid ineffective processing. Secondly, a generated length prediction module is used to accurately predict the length of the task text generation, providing a basis for subsequent processing. Then, a text segmentation module is used to segment the long text into segments with similar lengths for parallel processing. Finally, in combination with a text batching module and an intelligent scheduling module, tasks are reasonably allocated to heterogeneous virtual machine resources to minimize the task completion time. By optimizing the task processing flow and resource scheduling strategy, the present invention significantly improves the inference efficiency while ensuring the integrity and accuracy of the output results.
[0028] It should be noted that the embodiments of the present invention use a cloud server cluster as the background implementation platform, which supports the deployment of multiple large models. This cloud platform has high-performance computing resources and an automated management function, and can dynamically adjust computing resources to meet the needs of different tasks. And it supports on-demand elastic expansion, allocating and releasing resources in real time according to the task load, which not only ensures the computing efficiency but also effectively reduces the cost.
[0029] As Figure 1 shown, it is a schematic diagram of a specific application scenario of the present invention. After a user submits a task to the cloud platform, the cloud scheduling system first performs parallel processing on the task, splits the task into multiple subtasks, and performs batching processing according to the task type and resource requirements. Then, the system dynamically allocates the batched tasks to the cloud resources (such as GPU / TPU clusters) where the large models are deployed for execution, making full use of the high-performance computing capabilities of the cloud environment to ensure the efficient completion of tasks. Through parallel processing and intelligent scheduling, the present invention significantly improves the task processing efficiency while reducing the computing cost.
[0030] The specific implementation process of the intelligent scheduling method for task parallel processing for large model services proposed by the present invention, as Figure 2 shown, includes the following steps: Step S1: Receive the instruction and text of the user task, and judge whether the text can be split into subtasks for parallel execution through the large model according to the instruction, to obtain parallel subtasks.
[0031] In this embodiment, the user task refers to a large language model inference task, which is a service request. The service request can be a request sent by the user's large language model service, and can include natural language processing tasks such as text translation, entity recognition, language polishing, text continuation, summary generation, and sentiment analysis.
[0032] Specifically, tasks with weak context dependence such as text translation, entity recognition, and language polishing belong to parallelizable tasks, while tasks with strong context dependence such as text continuation, summary generation, and sentiment analysis belong to non-parallelizable tasks. If the judgment result is negative, direct scheduling processing is performed. If the judgment result is positive, parallel subtasks are obtained and subsequent parallel optimization operations are carried out. Through this type of judgment, large model tasks that can be optimized in parallel can be screened out.
[0033] Step S2, as Figure 3 shown, use the LaBSE model to extract semantic features from the instructions and text of the parallel subtasks and generate embedding vectors, and reduce the vector dimension of the embedding vectors through a compression module, and input the compressed embedding vectors into the LightGBM regressor to predict the generated length.
[0034] Step S2-1: Based on the instructions and text of the parallel subtasks obtained in Step S1, use the LaBSE model to capture the semantic information of the text and generate embedding vectors; Specifically, in this embodiment, the LaBSE model is used to extract semantic features and generate 768-dimensional embedding vectors, and the expression is: (1) (2) Where is the text embedding vector, is the user text, is the instruction embedding vector, is the instruction, , is the vector dimension, and = 768. The high-dimensional embedding vectors generated by the LaBSE model can provide a basis for subsequent compression and prediction.
[0035] Step S2-2: Based on the embedding vectors obtained in Step S2-1, compress the embedding vectors through a compression module to form compressed embedding vectors.
[0036] Specifically, compressing the embedding vectors through a compression module to form compressed embedding vectors is as follows: Divide the vector into groups on average, and divide the vector into groups on average. The number of elements in each group is , . Sum and normalize each group to obtain the compressed embedding vectors and . The calculation formula for the elements in each group is: (3) (4) Among them, is the compressed text embedding vector, is the compressed instruction embedding vector, is the i-th element within a group, is the number of text vector groups, optional = 16, is the number of instruction vector groups, optional = 4 (the optimal result after multiple experiments), and , , is the index of the element within the group.
[0037] Reduce the dimension of the embedding vector through a compression module, reduce the computational complexity, and avoid overfitting problems at the same time.
[0038] Step S2-3: Concatenate the compressed embedding vectors and with the length of the text input by the user to form the first input vector . Subsequently, input the first input vector into the trained LightGBM regressor to predict the generated length. Predict the generated length of the text with the help of this regressor, so as to ensure that the prediction result is both accurate and efficient.
[0039] Step S3: Based on the generated length, construct a segmentation algorithm based on semantic integrity constraints to segment the text into a list of text segments and tags, as Figure 4 shown, for subsequent parallel processing to improve the inference efficiency.
[0040] Step S3-1: Receive the text within a unit time and its corresponding predicted generated length to form an input text list ( is the th input text, is the total number of texts, where ) and its corresponding predicted generated length list ( is the predicted length of ). The service provider sets a parallel threshold according to the resource configuration to control the granularity of text segmentation and ensure that the segmented segments are suitable for parallel processing; Step S3-2: Initialize the text segment list , providing a container for storing the segmented text segments; Step S3-3: Traverse the input text list , for the A text , initialize the text sub - list , calculate the number of text segments to be cut: (5) Wherein, is the predicted generated length of the text, is a preset generated length threshold. Texts longer than this length will be cut to ensure that the lengths of the segmented segments are similar for subsequent parallel acceleration processing. Then The expected value of the size of each text segment after being cut is: (6) Wherein, is the expected value of the size of each text segment, is the text 's actual length; Step S3 - 4, cut the first sub - segment. Starting from the starting index of the input text (here 's starting index , calculate the end position of the first sub - segment : (7) Wherein, is the end position of the sub - segment, is the starting index of the input text.
[0041] To ensure semantic integrity, if the character at is not a full - stop (i.e., ), then search backward for the nearest full - stop and update to ensure that the text segment is not damaged.
[0042] Subsequently, extract the segment as the cut text (here ), store it in the text sub - list . Update the starting index of , increment . Continue to split the subsequent segments until the current text is processed. Set to 1, add the text sub - list to the text segment list . Increment , process the next text until all texts are processed. Increment
[0043] Step S3 - 5, return the text segment list , the text sub - list ( For the text of the nth text segment, where is the number of cut segments of the nth text. The index of this list is denoted as a marker), providing input for subsequent batch processing.
[0044] The above algorithm can be expressed as: a text segmentation algorithm based on semantic integrity constraints, as shown in Table 1.
[0045]
[0046] Step S4: As Figure 5 shown, adopt a multi-constraint task batching algorithm. According to the length similarity constraint and the same instruction constraint, divide the segmented text segments in the text segment list into different batch tasks, which can reduce the use of padding tokens and the differences in batch tasks to ensure that tasks at the same time and subtasks of the same task are completed at similar times, shortening the task execution time.
[0047] Step S4-1: Based on the text segment list segmented in Step S3 , set the relative error threshold and batch capacity according to the resource configuration, controlling the accuracy and scale of batching. This setting can continuously optimize the parameters according to the returned results.
[0048] Step S4-2: Initialize the batch queue and create an empty batch to provide a container for storing the text segments after batching.
[0049] Step S4-3: Traverse each sub-list of the text segment list . At the beginning of each loop, set the initial value of the relative error to , where is the relative error threshold. For each sub-list , calculate the length relative error of the first batch for the text segments in this sub-list: (8) where, is the batch length, that is, the length of the longest text in the batch. Consider the predicted lengths of all segments after cutting the same text as .
[0050] Step S4-4: After traversing each batch in the batch queue , determine whether there is a batch that meets the error threshold and batch capacity constrained batches. If the judgment result is yes, insert the text fragments in the sub-list into the batch with the closest generated length. If the judgment result is no, create a new batch and insert the text fragments in the sub-list. By making conditional judgments, the batch insertion of text fragments is realized, which not only ensures the constraints of the relative error threshold and batch capacity, but also dynamically expands the batch queue. This mechanism can balance precision and flexibility during the batching process, thereby improving the overall efficiency and adaptability of the algorithm.
[0051] The above algorithm can be expressed as: a multi-constrained task batching algorithm, as shown in Table 2.
[0052]
[0053] Step S5: As Figure 6 shown, optimize the batch task set using a task allocation algorithm for heterogeneous resources to determine an optimized sub-task scheduling scheme; Specifically, optimize the allocation of batch tasks using a task allocation algorithm for heterogeneous resources. With the aim of minimizing the task completion time, use the improved binary PSO algorithm to reasonably allocate batch tasks to heterogeneous virtual machine resources, improving the overall inference efficiency; Step S5-1: Randomly form an initial scheduling scheme based on the batch task set and the generated length. The initial scheduling scheme includes a batch task memory requirement vector and a virtual machine memory capacity vector. Based on the batch task set obtained in step 4 , there are a total of batch tasks, and the virtual machine set in the cloud service , there are a total of virtual machines. The estimated batch task execution time matrix , where the element in the matrix represents the estimated running time of the th task on the th virtual machine, indicating the execution time of batch task on virtual machine (i.e., the time of the task with the longest execution time within the batch). The batch task memory requirement vector obtained according to the predicted text generation length is expressed as , where is the element in the memory requirement vector , representing the memory required for the execution of batch task , is the memory requirement vector element index, . The virtual machine memory capacity vector , where is the memory capacity vector Element in the middle, representing the virtual machine The remaining memory capacity after loading the large model weights Is the index of the memory capacity vector element . Since the virtual machines in the cloud environment are heterogeneous, the memory capacities of all virtual machines are inconsistent.
[0054] Step S5-2, Initialize the particle swarm, corresponding to the initial scheduling scheme in the task generated in step S5-1, set the number of particles And the maximum number of iterations , set the learning factor , (where Is the individual learning factor, Is the global learning factor). Each particle Represents a task allocation scheme, encoded using a binary matrix . If Then it means that task Is assigned to virtual machine To execute; if Then it means that the batch task Is not assigned to virtual machine To execute. Initialize the velocity matrix As a zero matrix. The optimization objective (fitness function) is: (9) Among them, Is the total task completion time of virtual machine , Is the fitness value; (10) Among them, Is the batch task queue of the th virtual machine, Is the rd task on the th virtual machine's estimated running time. Each virtual machine VM maintains a FIFO batch task queue, and tasks enter the queue in the order of arrival.
[0055] The total memory requirement of all batch tasks assigned to the virtual machine cannot exceed the remaining memory capacity of the virtual machine. Therefore, the constraint condition is: (11) Among them, Is the allocation scheme, When it means that the th task is assigned to the th virtual machine; When it means that the th task is not assigned to the One virtual machine; ; (12) Wherein, ; Initialize the individual optimal solution and the global optimal solution ; Calculate the initial fitness value .
[0056] Step S5-3, loop until the maximum number of iterations is reached : Calculate the inertia weight for this round; (13) Wherein, is the inertia weight for this round, is the set maximum weight, is the set minimum weight, is the current round number.
[0057] Update the particle velocity for this round: (14) Wherein, is the particle velocity for this round, is the particle velocity of the previous round, , are random numbers uniformly distributed on [0, 1], is the individual optimal solution of the particle.
[0058] Then update the particle position: = , when < ; = , when ; Wherein, is the updated allocation scheme, is the Sigmoid activation function, and ; Map the velocity to the interval [0, 1], is a random number uniformly distributed on [0, 1], is the particle velocity. Check the task allocation constraints, otherwise reallocate. Calculate the fitness value for this round Update the individual optimal solution of each particle , update the global optimal solution of the current round ; Step S5-4: Perform local search optimization every five rounds. If the new solution is better than the global optimal solution , then update the global optimal solution to the new solution; otherwise, accept the new solution with a probability of , where is the fitness difference between the new solution and the current solution, and is the temperature parameter of the current iteration; Step S5-5: Based on the returned task scheduling scheme , the system allocates batch tasks to the most suitable resources for execution according to the task scheduling scheme; Step S6: Return the inference result and continuously optimize: According to the tags of the parallel subtasks, splice the subtask scheduling schemes of the same task in sequence to obtain a task parallel processing scheduling scheme, and return the result to the user.
[0059] In addition, the system continuously optimizes the parameters of the above step models and algorithms according to the task completion situation to achieve the best inference efficiency.
[0060] Embodiment 2 In one or more embodiments, an intelligent task parallel processing scheduling system for large model services is disclosed, which specifically includes: A task receiving module, which is configured to: receive the instructions and text of the user task, and judge whether the text can be split into subtasks for parallel execution according to the instructions to obtain parallel subtasks; A generated length prediction module, which is configured to: extract semantic features from the instructions and text of the parallel subtasks by using the LaBSE model and generate embedding vectors, reduce the vector dimension of the embedding vectors through a compression module to obtain compressed embedding vectors, and input the compressed embedding vectors into a LightGBM regressor to predict the generated length; A text segmentation module, which is configured to: based on the generated length, construct a segmentation algorithm based on semantic integrity constraints, and segment the text to form a text segment list and tags; A task batching module, which is configured to: adopt a multi-constraint task batching algorithm, and divide the segmented text segments in the text segment list into tasks of different batches according to the length similarity constraint and the same instruction constraint to obtain a batch task set; A scheduling optimization module, which is configured to: optimize the batch task set by using a task allocation algorithm for heterogeneous resources to determine an optimized subtask scheduling scheme; A scheduling generation module, which is configured to: splice the subtask scheduling schemes of the same task in sequence according to the tags of the parallel subtasks to obtain a task parallel processing scheduling scheme.
[0061] Embodiment 3 This embodiment provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps of the above-mentioned intelligent scheduling method for task parallel processing for large model services are completed.
[0062] Embodiment 4 This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the above-mentioned intelligent scheduling method for task parallel processing for large model services are completed.
[0063] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0064] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0065] These computer program instructions can also be loaded onto a computer or other programmable data processing device to perform a series of operation steps on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0066] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0067] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An intelligent scheduling method for task parallel processing oriented to large model services, characterized in that Including: Receiving the instructions and text of the user task, judging whether the text can be split into subtasks for parallel execution according to the instructions, and obtaining parallel subtasks; Using the LaBSE model to extract semantic features from the instructions and text of the parallel subtasks and generate embedding vectors, reducing the vector dimension of the embedding vectors through a compression module to obtain compressed embedding vectors, and inputting the compressed embedding vectors into a LightGBM regressor to predict and generate lengths; Based on the generated length, constructing a segmentation algorithm based on semantic integrity constraints to segment the text into a list of text segments and tags; Adopting a multi-constraint task batching algorithm, and dividing the segmented text segments in the text segment list into tasks of different batches according to the length similarity constraint and the same instruction constraint to obtain a set of batch tasks; Optimizing the set of batch tasks by using a task allocation algorithm for heterogeneous resources to determine an optimized subtask scheduling scheme; According to the tags of the parallel subtasks, splicing the subtask scheduling schemes of the same task in sequence to obtain a task parallel processing scheduling scheme.
2. The intelligent scheduling method for task parallel processing for large model services according to claim 1, characterized in that, The specifically reducing the vector dimension of the embedding vectors through the compression module is: Embed the text into a vector Divide evenly into groups, and embed the instruction into a vector Divide evenly into groups, and the number of elements in each group is respectively 、 , sum and normalize each group to obtain the compressed embedding vector. The calculation formula for the elements in the group is: Among them, is the compressed text embedding vector, is the compressed instruction embedding vector, is the i-th in-group element, is the vector dimension, is the number of text vector groups, is the number of instruction vector groups, is the in-group element index.
3. The intelligent scheduling method for task parallel processing for large model services according to claim 1, characterized in that, Based on the generated length, constructing a segmentation algorithm based on semantic integrity constraints to segment the text into a list of text segments and tags, specifically: According to the generated length, forming an input text list and a generated length list; Initializing the text segment list; Traversing the input text list and its corresponding generated length list, and under the parallel threshold constraint, calculating the number of text cut segments of each text and its corresponding segment generated length to ensure that the generated lengths of the text segments after splitting the same text are similar; Cutting each text in the text list according to the calculated number of cut segments and storing it in the corresponding text sub-list; Integrating the segmented sub-lists to obtain a text segment list.
4. The intelligent scheduling method for task parallel processing for large model services according to claim 1, characterized in that, The dividing the segmented text segments in the text segment list into tasks of different batches according to the length similarity constraint and the same instruction constraint to obtain a set of batch tasks includes: Receiving the segmented text segments in the text segment list; Initializing the batch queue and creating a batch; Sequentially accessing each sub-list in the text segment list, and for each sub-list, calculating its relative error value with the longest generated length of each batch task in the batch queue, and locating the batch with the most matching generated length by comparing the error values; Judging whether there is a batch in the batch queue whose relative error meets the error threshold and the generated length of the sub-list meets the batch capacity constraint; if the judgment result is yes, inserting the text in the sub-list into the batch with the closest generated length; if the judgment result is no, creating a new batch and inserting the text in the sub-list to obtain a set of batch tasks.
5. The intelligent scheduling method for task parallel processing for large model services according to claim 4, characterized in that, The relative error value of the generated length is: Among them, is the relative length error, is the batch length, is the predicted generated length, and n is the th text segment.
6. The intelligent scheduling method for task parallel processing for large model services according to claim 1, wherein The optimizing the set of batch tasks by using a task allocation algorithm for heterogeneous resources to determine an optimized subtask scheduling scheme includes: Randomly forming an initial scheduling scheme based on the set of batch tasks and the generated length; Based on the initial iteration scheme, initializing a particle swarm, setting the number of particles, the maximum number of iterations, the learning factor, calculating the initial local and global optimal solutions, and setting the optimization objective; Calculate the weights cyclically and update the velocity and position of each particle to calculate the optimal solution; Based on the number of cycles, perform local search optimization every five rounds to determine the optimized sub-task scheduling scheme.
7. The intelligent scheduling method for task parallel processing for large model services according to claim 6, characterized in that, The optimization objective is: Among them, is the total task completion time of the virtual machine , Among them, is the th virtual machine batch task queue, is the th task's estimated running time on the th virtual machine; The constraint conditions are: wherein, is an element in the memory requirement vector, ; Among them, is the allocation plan, when it indicates that the th task is assigned to the th virtual machine; when it indicates that the th task is not assigned to the th virtual machine, ; The cyclical calculation of weights and updating the velocity and position of each particle to calculate the optimal solution includes: Calculate the inertial weight for this round; Among them, is the inertia weight of this round, is the set maximum weight, is the set minimum weight, is the current round; Update the particle velocity: Among them, is the particle velocity of this round, is the particle velocity of the previous round, , is a random number uniformly distributed on [0, 1], is the individual optimal solution of the particle, is the global optimal solution, is the allocation scheme, is the individual learning factor, is the global learning factor; Then update the particle position: = , when < ; = , when ; Among them, The updated allocation scheme, is the Sigmoid activation function, is a random number uniformly distributed on [0, 1], is the particle velocity.
8. An intelligent scheduling system for task parallel processing oriented to large model services, characterized in that, including: A task receiving module, configured to: receive the instructions and text of the user task, determine whether the text can be split into sub-tasks for parallel execution according to the instructions, and obtain parallel sub-tasks; A length prediction generation module, configured to: use the LaBSE model to extract semantic features from the instructions and text of the parallel sub-tasks and generate embedding vectors, reduce the vector dimension of the embedding vectors through a compression module to obtain compressed embedding vectors, and input the compressed embedding vectors into a LightGBM regressor to predict the generated length; A text segmentation module, configured to: based on the generated length, construct a segmentation algorithm based on semantic integrity constraints, and segment the text to form a list of text segments and tags; A task batching module, configured to: use a multi-constraint task batching algorithm to divide the segmented text segments in the text segment list into tasks of different batches according to the length similarity constraint and the same instruction constraint, and obtain a set of batch tasks; A scheduling optimization module, configured to: use a task allocation algorithm for heterogeneous resources to optimize the set of batch tasks and determine the optimized sub-task scheduling scheme; A scheduling generation module, configured to: splice the sub-task scheduling schemes of the same task in sequence according to the tags of the parallel sub-tasks to obtain a task parallel processing scheduling scheme.
9. An electronic device, characterized in that, It includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the task parallel processing intelligent scheduling method for large model services described in any one of claims 1-7 is completed.
10. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the task parallel processing intelligent scheduling method for large model services described in any one of claims 1-7 is completed.
Citation Information
Patent Citations
Semantic representation model pre-training method and device, electronic equipment and storage medium
CN112560499A
Financial risk prediction method and device based on text pre-training and multi-task learning
CN113743111A
Text mining-based refined fitting transformer fault identification method and device
CN114912460A
Optimization method and system for dynamic reasoning memory allocation based on predictor
CN118227336A
Large language model service request scheduling method and system
CN119960939A
Cited By
Inference request scheduling method and system, electronic equipment and storage medium
CN120849134A
An inference request scheduling method and system, an electronic device, and a storage medium
CN120849134B
Steel production equipment start-stop scheduling optimization method and device based on particle swarm optimization
CN121903324A
Steel production equipment start-stop scheduling optimization method and device based on particle swarm algorithm
CN121903324B