Model fine tuning method based on intelligent calculation platform and intelligent calculation platform

Through the scheduling center of the intelligent computing platform, combined with the capability information of the computing nodes, the optimal scheduling combination of large-model fine-tuning tasks is determined, which solves the problems of low resource utilization and inflexible pricing solutions, and achieves efficient resource utilization and market supply and demand adaptation.

CN120069000AActive Publication Date: 2025-05-30中国联合网络通信有限公司广东省分公司
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510139801.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The large-model fine-tuning task did not fully consider the characteristics of computing and memory capabilities during the scheduling process, resulting in low resource utilization and the existing pricing plan failed to adapt to changes in market supply and demand in a timely manner.

Method used

Through the scheduling center of the intelligent computing platform, the model fine-tuning tasks sent by the user are received, combined with the node capability information of each computing node in the computing cluster, and with the goal of maximizing the total profit, the optimal scheduling combination is determined, the target computing node that meets the computing needs is selected, and the task is judged based on the optimal scheduling combination, and a flexible pricing plan is designed.

Benefits of technology

It improves the resource utilization rate of computing nodes, adapts to changes in market supply and demand relationships, and maximizes the total income of users and service providers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069000A_ABST
    Figure CN120069000A_ABST
Patent Text Reader

Abstract

The invention discloses a model fine tuning method based on an intelligent computing platform and the intelligent computing platform, the intelligent computing platform comprises a user side, a scheduling center and a computing cluster, the method is applied to a scheduling center side, and the method comprises the following steps: receiving a model fine tuning task sent by the user side, the model fine tuning task at least comprising a bid of a user, a fine tuning data set and a computing demand; based on the bid, the fine tuning data set and the calculation demand quantity, combining node capability information of each calculation node in the calculation cluster, and taking total income maximization as a target, determining an optimal scheduling combination of the model fine tuning task, the optimal scheduling combination comprising a target calculation node; judging whether a model fine tuning task is accepted or not according to the optimal scheduling combination; and if the model fine tuning task is accepted, controlling the target computing node to execute the model fine tuning task based on the fine tuning data set. According to the method, the resource utilization rate of the computing nodes can be improved, and a flexible pricing scheme is provided to adapt to continuously changing market supply and demand relationships.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large model fine-tuning, and particularly to a model fine-tuning method and an intelligent computing platform based on an intelligent computing platform. Background Art

[0002] An artificial intelligence large model refers to a deep learning model with a large number of parameters trained using large-scale data and powerful computing capabilities. These models usually have high generality and generalization capabilities and can be applied to intelligent scenarios such as natural language processing, image recognition, and speech recognition. However, the use effect of large models in some specific fields or special scenarios is not good. Therefore, it is necessary to fine-tune general large models on a large amount of specialized data to make them suitable for certain special scenarios.

[0003] In related technologies, the large model fine-tuning task is essentially a deep learning training task. In the scheduling work of deep learning training tasks, the key characteristics of the large model fine-tuning task itself, such as computing and video memory capabilities, are not fully considered, resulting in low resource utilization. And the existing pricing schemes for large model fine-tuning tasks fail to adapt to the constantly changing supply and demand relationship in the market in a timely manner. Summary of the Invention

[0004] In view of this, the present invention provides a model fine-tuning method and an intelligent computing platform based on an intelligent computing platform to improve the resource utilization rate of computing nodes and provide a flexible pricing scheme to adapt to the constantly changing market supply and demand relationship.

[0005] The first aspect of the present invention provides a model fine-tuning method based on an intelligent computing platform. The intelligent computing platform includes a user side, a scheduling center, and a computing cluster. The method is applied to the scheduling center side and includes:

[0006] Receiving a model fine-tuning task sent by the user side, where the model fine-tuning task at least includes the user's bid, a fine-tuning data set, and a computing demand;

[0007] Based on the bid, the fine-tuning data set, and the computing demand, and combining the node capability information of each computing node in the computing cluster, with the goal of maximizing the total revenue, determining an optimal scheduling combination for the model fine-tuning task, where the optimal scheduling combination includes target computing nodes selected from the computing cluster that meet the computing demand; the node capability information includes computing power, video memory capacity, and video memory bandwidth capacity; the total revenue includes the user's revenue and the revenue of the intelligent computing service provider;

[0008] Judging whether to accept the model fine-tuning task according to the optimal scheduling combination;

[0009] If the model fine-tuning task is accepted, control the target computing node to execute the model fine-tuning task based on the fine-tuning data set.

[0010] The second aspect of the present invention provides an intelligent computing platform, which includes a user terminal, a scheduling center, and a computing cluster. The scheduling center side includes the following modules:

[0011] A task receiving module, configured to receive the model fine-tuning task sent by the user terminal, where the model fine-tuning task includes the user's bid, the fine-tuning data set, and the computing demand;

[0012] An optimal scheduling combination determination module, configured to determine the optimal scheduling combination of the model fine-tuning task based on the bid, the fine-tuning data set, and the computing demand, in combination with the node capability information of each computing node in the computing cluster, with the goal of maximizing the total revenue. Among them, the optimal scheduling combination includes the target computing nodes selected from the computing cluster that meet the computing demand; the node capability information includes computing power, video memory capacity, and video memory bandwidth capacity; the total revenue includes the user's revenue and the revenue of the intelligent computing service provider;

[0013] A task acceptance judgment module, configured to judge whether to accept the model fine-tuning task according to the optimal scheduling combination;

[0014] A task execution module, configured to, if the model fine-tuning task is accepted, control the target computing node to execute the model fine-tuning task based on the fine-tuning data set.

[0015] The third aspect of the present invention provides an electronic device, which includes:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the model fine-tuning method based on the intelligent computing platform as described in the first aspect above.

[0019] The fourth aspect of the present invention provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the model fine-tuning method based on the intelligent computing platform as described in the first aspect above.

[0020] A fifth aspect of the present invention provides a computer program product, which includes a computer program that, when executed by a processor, implements the model fine-tuning method based on an intelligent computing platform as described in the first aspect above.

[0021] In this embodiment, the intelligent computing platform uniformly schedules and manages the model fine-tuning tasks from different user terminals, which is beneficial to the flexible and precise control of the tasks. The model fine-tuning tasks include fine-tuning data sets and computing requirements. When determining the target computing nodes from the computing cluster, considering the node capability information such as the computing power, video memory capacity, and video memory bandwidth capacity of each computing node, the selected target computing nodes need to meet the computing requirements, which is beneficial to the smooth execution of the tasks and also beneficial to selecting computing nodes with matching capabilities, thereby improving the resource utilization rate of the computing nodes.

[0022] Further, the model fine-tuning tasks in this embodiment include the bids of users. When determining whether to accept the tasks, the bids of users also need to be considered. With the goal of maximizing the total revenue, the optimal scheduling combination of the model fine-tuning tasks is determined, and the actual payment cost is calculated according to the optimal scheduling combination. By this auction mechanism, a pricing scheme is designed for users and intelligent computing service providers, which can maximize the total revenue of users and service providers and can also flexibly adapt to the changing market supply and demand relationship.

[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0025] Figure 1 is a schematic framework diagram of an intelligent computing platform provided in Embodiment 1 of the present invention.

[0026] Figure 2 is a flowchart of a model fine-tuning method based on an intelligent computing platform provided in Embodiment 1 of the present invention.

[0027] Figure 3 is a schematic diagram of multi-LoRA task training provided in Embodiment 1 of the present invention;

[0028] Figure 4It is a flowchart of a model fine-tuning method based on an intelligent computing platform provided in the second embodiment of the present invention;

[0029] Figure 5 It is a flowchart of a model fine-tuning method based on an intelligent computing platform provided in the third embodiment of the present invention;

[0030] Figure 6 It is a schematic structural diagram of a scheduling center provided in the fourth embodiment of the present invention.

[0031] Figure 7 It is a schematic structural diagram of an electronic device provided in the fifth embodiment of the present invention. Detailed implementation manners

[0032] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can cover sequences other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0034] Embodiment 1

[0035] A model fine-tuning method based on an intelligent computing platform provided in the first embodiment of the present invention can be executed by the intelligent computing platform, and the intelligent computing platform can be implemented in the form of hardware and / or software. Refer to Figure 1 , which shows a framework schematic diagram of the intelligent computing platform. The intelligent computing platform is a cloud service platform that provides services for intelligent computing service providers, including a user side, a scheduling center, and a computing cluster. Among them,

[0036] The client is the client or terminal for user interaction, which is used to provide an interactive interface for users so that they can submit model fine-tuning tasks. In implementation, lightweight fine-tuning techniques represented by LoRA are mostly used for large model fine-tuning. LoRA can keep the pre-trained model parameters unchanged and only fine-tune the parameters of the LoRA adapter. Therefore, the model fine-tuning task in this embodiment can also be called the LoRA fine-tuning task.

[0037] The client is connected to the scheduling center, which is used to schedule and process model fine-tuning tasks.

[0038] Connected to the scheduling center is also a computing cluster, which can be a GPU cluster. This cluster consists of a group of [K] = {1, 2,..., K} GPU computing nodes and is used to execute the LoRA fine-tuning tasks submitted by users.

[0039] In a further embodiment, as Figure 1 shown, connected to the scheduling center is also a data provider in a third-party data market. The data provider is used to perform data preprocessing on the LoRA fine-tuning tasks, such as tagging, cleaning, etc. Specifically, the data provider is a preprocessing service provider that preprocesses the tasks before the computing nodes process the tasks. When users submit each model fine-tuning task, they need to submit a dataset for fine-tuning. However, current large model fine-tuning has strict requirements for the format of the dataset. For ordinary users, it is still a great challenge to directly submit a dataset format that meets the requirements of the intelligent computing platform. Therefore, a third-party data market can be used to assist in data preprocessing. There can be multiple data providers, and in this embodiment, they can be represented as [N] = {1, 2,..., N}.

[0040] Based on Figure 1 the intelligent computing platform, referring to Figure 2 shows a flowchart of a model fine-tuning method based on an intelligent computing platform provided in Embodiment 1 of the present invention. This embodiment can be executed by the scheduling center. As Figure 2 shown, this embodiment can include the following steps:

[0041] Step 101, receive the model fine-tuning task sent by the client. The model fine-tuning task at least includes the user's bid, the fine-tuning dataset, and the computing demand.

[0042] In implementation, the scheduling center can form a fine-tuning task set [I] = {1, 2,..., I} with the received model fine-tuning tasks. It should be noted that this embodiment can be applied to the scenario of fine-tuning a single large model. In this case, the fine-tuning task set targets the same large model; it can also be applied to the scenario of fine-tuning multiple large models. In this case, different large models can correspond to different fine-tuning task sets.

[0043] Each model fine-tuning task i (which can be abbreviated as task i hereinafter) may include the following information: task arrival time, task deadline, training dataset for fine-tuning (i.e., fine-tuning dataset), video memory requirement (i.e., video memory requirement for GPU), computing requirement (i.e., total computing amount required), whether data preprocessing is needed, the user's bid, etc.

[0044] As an example, the model fine-tuning task i can be expressed as where a i represents the task arrival time; d i represents the task deadline, that is, task i needs to be completed before d i ; otherwise, the intelligent computing service provider cannot obtain the reward for executing this task. represents the training dataset for fine-tuning (i.e., fine-tuning dataset); r i represents the video memory requirement of task i; M i represents the computing requirement of task i, that is, the number of total data samples to be processed; f i represents whether task i needs data preprocessing, and b i represents the user's bid, that is, if the intelligent computing service provider executes task i before the deadline d i specified by the user, the fee that the user is willing to pay.

[0045] In this embodiment, the user can submit the model fine-tuning task through the user terminal in a bidding manner, which is convenient for the user to put forward personalized fine-tuning requirements. The fine-tuning dataset and computing requirement are included in the model fine-tuning task, which is conducive to selecting a suitable computing node to process the task when performing computing node matching subsequently, and thus is conducive to improving the resource utilization rate of the computing node. The user's bid is included in the model fine-tuning task, which is conducive to the user to better express the bidding intention without being restricted by a fixed price, and improves the flexibility of pricing.

[0046] Step 102: Based on the bid, fine-tuning dataset, and computing requirement, combined with the node capability information of each computing node in the computing cluster, with the goal of maximizing the total revenue, determine the optimal scheduling combination of the model fine-tuning task, where the optimal scheduling combination includes the target computing nodes selected from the computing cluster that meet the computing requirement.

[0047] In this step, when the model fine-tuning task arrives at the intelligent computing platform, the scheduling center determines the optimal scheduling combination of the model fine-tuning task based on the demand information in the model fine-tuning task and the node capability information of each computing node in the computing cluster, with the goal of maximizing the total revenue.

[0048] Maximizing the total revenue can be understood as maximizing social welfare. In this embodiment, the total revenue includes the revenue of users and the revenue of the intelligent computing service provider. In one implementation, the revenue of users can be the cost saved by users with the pricing scheme in this embodiment. The revenue of the intelligent computing service provider can be the profit remaining after subtracting various expenses from the fees received by the intelligent computing service provider from users.

[0049] When determining the target computing node in this embodiment, the node capability information of each computing node in the computing cluster will be considered. Exemplarily, the node capability information may include computing power, video memory capacity, and video memory bandwidth capacity. Computing power is an indicator that measures the processing ability and performance of a computing node, representing the maximum number of data samples that the computing node can process in a single time period; video memory capacity refers to the remaining video memory. Video memory is a storage device in a graphics card used to store image data. Its capacity determines the number of images and the resolution that the graphics card can process. The larger the video memory, the more pictures can be stored, and the unit is GB; video memory bandwidth refers to the data transfer rate between the display chip and the video memory, with the unit of GB / s. The larger the video memory bandwidth, the faster the GPU processes data, and the video memory bandwidth is jointly determined by the video memory frequency and the video memory bit width.

[0050] A scheduling combination refers to a combination of various execution elements that meet the various requirement information of the model fine-tuning task. The execution elements may include, but are not limited to: task acceptance information (i.e., whether the current task i is accepted), computing node information, data provider information, etc. For example, the scheduling combination can be expressed as l = {u i ,{x ikt} k,t ,{z in} n}, where l represents the scheduling combination, u i represents whether the current task i is accepted; {x ikt} k,t represents the computing node information, that is, whether to execute task i on computing node k in time period t; {z in} n represents the data provider information, that is, whether to select and use the service of data provider n for data preprocessing of task i.

[0051] Step 103, determine whether to accept the model fine-tuning task according to the optimal scheduling combination.

[0052] In this step, it can be determined whether to accept the model fine-tuning task according to the value of u i in the optimal scheduling combination. For example, if u i = 0, it is determined not to accept the model fine-tuning task; if u i > 0, it is determined to accept the model fine-tuning task.

[0053] Step 104: If the model fine-tuning task is accepted, control the target computing node to perform the model fine-tuning task based on the fine-tuning dataset.

[0054] In this step, when task i is accepted, task i can be sent to the target computing node. The target computing node can then use the LoRA fine-tuning technique to complete the fine-tuning of the large model before the task deadline according to the training dataset in task i. The platform can be set to run in discrete time periods [T] = {1, 2, …, T}.

[0055] In one implementation, when the target computing node performs the model fine-tuning task, it can train the large model based on the shared pre-trained parameters. For example, as Figure 3 shown in the training schematic diagram of multiple LoRA tasks, LoRA can keep the pre-trained model parameters unchanged and only fine-tune the parameters of the LoRA adapter. Among them, (x 1 , x 2 , x 3 ) are the training data points of three fine-tuning tasks respectively. In the forward propagation process, the three tasks calculate the forward output results respectively based on the pre-trained parameters and the corresponding LoRA adapters, further calculate the loss function, and backpropagate the gradients to the corresponding LoRA adapters for parameter update. Sharing the pre-trained parameters among different LoRA fine-tuning tasks can further reduce the resources required for the fine-tuning process.

[0056] In one embodiment, if task i selects to use the data provider service for data preprocessing of task i, before the scheduling center sends task i to the target computing node, it first selects the target data provider from the third-party data market and sends the training dataset of task i to the target data provider for data preprocessing. After the scheduling center obtains the preprocessed training dataset from the target data provider, it replaces the training dataset in task i with the preprocessed training dataset and sends task i to the target computing node.

[0057] When selecting the target data provider from the third-party data market, the intelligent computing platform comprehensively considers the price q in charged by each data provider n and the latency h in required for each data provider n to process the training dataset of task i, and then selects the service of one of the data providers for task i.

[0058] In this embodiment, the intelligent computing platform is used to uniformly schedule and manage the model fine-tuning tasks from different client terminals, which is beneficial to the flexible and precise control of the tasks. The model fine-tuning task includes a fine-tuning data set and a computing requirement. When determining the target computing node from the computing cluster, considering the node capability information such as the computing power, video memory capacity, and video memory bandwidth capacity of each computing node, the selected target computing node needs to meet the computing requirement, which is beneficial to the smooth execution of the task and also beneficial to selecting a computing node with a matching capability, thereby improving the resource utilization rate of the computing node.

[0059] Embodiment 2

[0060] Refer to Figure 4 FIG. shows a flowchart of a model fine-tuning method based on an intelligent computing platform provided in Embodiment 2 of the present invention. On the basis of Embodiment 1, the pricing process of the task is described. As Figure 4 shown, this embodiment may include the following steps:

[0061] Step 201, receive a model fine-tuning task sent by a client terminal, where the model fine-tuning task includes at least the user's bid, a fine-tuning data set, and a computing requirement.

[0062] Step 202, based on the bid, the fine-tuning data set, and the computing requirement, and combining the node capability information of each computing node in the computing cluster, with the goal of maximizing the total revenue, determine the optimal scheduling combination of the model fine-tuning task, where the optimal scheduling combination includes the target computing nodes selected from the computing cluster that meet the computing requirement.

[0063] Exemplarily, the node capability information includes computing power, video memory capacity, and video memory bandwidth capacity; the total revenue includes the user's revenue and the revenue of the intelligent computing service provider.

[0064] Step 203, determine whether to accept the model fine-tuning task according to the optimal scheduling combination.

[0065] Step 204, if the model fine-tuning task is accepted, control the target computing node to execute the model fine-tuning task based on the fine-tuning data set.

[0066] Step 205, calculate the actual payment fee according to the optimal scheduling combination.

[0067] In this embodiment, the actual payment fee may not be the same as the user's bid. Generally, the actual payment fee is lower than the user's bid. If the actual payment fee is higher than the user's bid, the task is rejected and the user does not need to pay.

[0068] In implementation, the actual payment cost can be the overhead incurred by the platform for task execution, which includes the cost paid to third-party data providers, the computing overhead of computing nodes, the video memory overhead, the video memory bandwidth overhead, etc.

[0069] Step 206: Return the actual payment cost to the user side so that the user can make a payment according to the actual payment cost.

[0070] In this embodiment, the model fine-tuning task includes the user's bid. When determining whether to accept the task, the user's bid also needs to be considered. With the goal of maximizing the total revenue, the optimal scheduling combination of the model fine-tuning task is determined, and the actual payment cost is calculated according to the optimal scheduling combination. By this auction mechanism, a pricing scheme can be designed for users and intelligent computing service providers, which can maximize the total revenue of users and service providers and can flexibly adapt to the changing market supply and demand relationship.

[0071] Embodiment III

[0072] See Figure 5 shows a flowchart of a model fine-tuning method based on an intelligent computing platform provided in Embodiment III of the present invention. On the basis of Embodiment I or Embodiment II, the determination process of the optimal scheduling combination is described. As Figure 5 shown, this embodiment may include the following steps:

[0073] Step 301: Receive a model fine-tuning task sent by the user side. The model fine-tuning task includes at least the user's bid, the fine-tuning data set, and the computing demand.

[0074] Step 302: Based on the bid, the fine-tuning data set, and the computing demand, combined with the node capability information of each computing node in the computing cluster, with the goal of maximizing the total revenue, model the original optimization problem based on decision variables.

[0075] Among them, the data provider is a preprocessing service provider that preprocesses the task before the computing node processes the task.

[0076] The original optimization problem includes an optimization objective and a first set of constraint conditions. The first set of constraint conditions includes multiple decision variables. Exemplarily, the decision variables include variables related to data providers, computing nodes, and task acceptance.

[0077] In this embodiment, the optimization objective is to maximize the total revenue. In one implementation, the total revenue is expressed as follows:

[0078]

[0079] Among them,

[0080]

[0081] Among them, U represents the total revenue, and U r represents the revenue obtained by all users, and U c represents the revenue obtained by the intelligent computing service provider, and U i represents the revenue obtained by the user who initiates model fine-tuning task i, and b i represents the bid of the user who initiates model fine-tuning task i, and p i represents the actual payment fee of the user after model fine-tuning task i is accepted, and u i represents the task acceptance information on whether model fine-tuning task i is accepted, and p i u i represents the actual payment fee received by the platform for model fine-tuning task i, and z in represents the data provider information on whether model fine-tuning task i requires data preprocessing by data provider n, and q in represents the fee charged by the data provider, and x ikt represents whether model fine-tuning task i is executed on computing node k during time period t, and e ikt represents the system overhead caused by training model fine-tuning task i on computing node k during time period t. n represents the data provider, k represents the computing node, and t represents the execution time period of the computing node.

[0082] ∑ i ∑ n q in z in represents the fee paid by the platform to the data provider; ∑ i ∑ k ∑ t e ikt x ikt represents the overhead of the platform for executing model fine-tuning task i.

[0083] Based on the above expression of the total revenue, taking the maximization of the total revenue as the optimization objective, the original optimization problem based on decision variables can be modeled as follows:

[0084]

[0085]

[0086]

[0087] Among them, P represents the optimization objective of maximizing the total revenue, and the constraint conditions (a1)-(a9) constitute the first set of constraint conditions;

[0088] (a1) is used to ensure that for each model fine-tuning task i, if the task is accepted and requires data preprocessing, then exactly one data provider is selected at most; fi Indicates whether the model fine-tuning task i requires data preprocessing;

[0089] (a2) Used to ensure that each model fine-tuning task i runs on at most one computing node k in each time period t;

[0090] (a3) Used to ensure that each model fine-tuning task i is executed only after arriving at the system and completing data preprocessing; a i Indicates the arrival time of the model fine-tuning task i; h in Indicates the delay required for the data provider to process the training dataset of the model fine-tuning task i;

[0091] (a4) Used to ensure that each model fine-tuning task i is executed before its deadline; d i Indicates the deadline of the model fine-tuning task i;

[0092] (a5) Used to ensure that each model fine-tuning task i has completed sufficient computational volume; M i Indicates the computational demand of the model fine-tuning task i; s ik Indicates the computational volume that can be completed in each time period when the model fine-tuning task i is executed on the computing node k;

[0093] (a6) Used to constrain the computational capacity of each time period and each computing node; C kp Indicates the computing power of the computing node k ∈ [K];

[0094] (a7) Used to constrain the video memory capacity of each time period and each computing node; C km Indicates the video memory capacity of the computing node k;

[0095] (a8) Used to constrain the video memory bandwidth capacity of each time period and each computing node; C kg Indicates the video memory bandwidth capacity of the computing node k;

[0096] (a9) Specifies the value range of each decision variable. u i 、x ikt 、z in Are all decision variables.

[0097] Step 303, simplify the original optimization problem into an equivalent problem based on scheduling combination.

[0098] Since the original optimization problem contains complex constraint conditions, this embodiment introduces the concept of scheduling combination to simplify the original optimization problem. The scheduling combination is a combination of execution elements of the decision variable that satisfy the constraints of the task itself in the first constraint condition set, and the execution elements at least include computing node information, data provider information, and task acceptance information. In one implementation, the original optimization problem can be simplified to an equivalent problem based on the scheduling combination by means of Comp-Exp.

[0099] When implementing, for any task i, its scheduling combination l is defined as the assignment of a set of specific values of the decision variable l = {u i , {x ikt} k,t , {z in} n} that satisfies the constraints (a1)-(a5).

[0100] In one embodiment, the equivalent problem based on the scheduling combination can be expressed as follows:

[0101]

[0102]

[0103]

[0104] Among them, P 1 represents the maximization of the total revenue after equivalence, and the constraint conditions (b1)-(b5) form the second constraint condition set; ζ i refers to the set of scheduling combinations of the model fine-tuning task i that satisfy all necessary constraint conditions; the necessary constraint conditions include the constraints (a1)-(a5); b il represents the increment of the objective function value of the problem P when using the scheduling combination l to execute the model fine-tuning task i; x il represents whether the model fine-tuning task i is executed according to the scheduling combination l; s kt (il), r kt (il), σ kt (il) respectively represent the computing consumption, video memory consumption, and video memory bandwidth consumption when using the scheduling combination l to execute the model fine-tuning task i on the computing node k at the time period t. r b is the GPU video memory space occupied by the pre-trained large model.

[0105] Furthermore, b il can be determined in the following manner:

[0106]

[0107] Among them, the decision variables z in and x iktThe value of is determined according to the scheduling combination l.

[0108] s kt (il), r kt (il), σ kt (il) are determined respectively in the following ways:

[0109] s kt (il) = s ik x ikt , r kt (il) = r i x ikt , σ kt (il) = σ i x ikt

[0110] Among them, x ikt ∈l. For the convenience of expression, t ∈ l is used to represent that the time period t is one of the specified time periods in the scheduling combination l, that is

[0111]

[0112] Step 304: Based on the primal-dual algorithm and the equivalent problem, generate the Lagrangian dual problem corresponding to the primal optimization problem.

[0113] In this embodiment, an online scheduling algorithm is designed based on the idea of online primal-dual. First, the integer constraint of the decision variable x il is relaxed to x il ∈[0,1]. Then, write out the Lagrangian dual problem of the primal problem P.

[0114] In one embodiment, the Lagrangian dual problem is expressed as:

[0115]

[0116] Among them, μ i , λ kt , γ kt respectively represent the dual variables associated with the constraints (b1), (b2), (b3), (b4).

[0117] Step 305: Determine the optimal scheduling combination of the model fine-tuning task based on the Lagrangian dual problem.

[0118] In this step, when the Lagrangian dual problem is determined, the optimal scheduling combination of the model fine-tuning task can be further determined according to this Lagrangian dual problem.

[0119] In one embodiment, the Lagrangian dual problem includes a plurality of dual variables, and the dual variables include variables related to the computing overhead, video memory overhead, and video memory bandwidth overhead of computing nodes when executing tasks; then step 305 further includes the following steps:

[0120] Step 305-1: Determine the values of the dual variables corresponding to each scheduling combination according to the Lagrangian dual problem.

[0121] In one implementation, the values of the dual variables λ kt 、 γ kt can be determined in the following manner:

[0122]

[0123] wherein, can be understood as the increment of social welfare that can be brought by consuming a unit of resources in a single time period. and respectively represent the values of the dual variables λ kt , and γ kt after the online algorithm processes task i.

[0124] In one implementation, can be calculated in the following manner:

[0125]

[0126] Step 305-2: Determine the function representation of each scheduling combination according to the values of the dual variables corresponding to each scheduling combination.

[0127] In one implementation, step 305-2 can be implemented using the following formula:

[0128]

[0129] wherein, F(il) represents the function representation of scheduling combination l; (k,t) ∈ l means the (k,t) pair of computing node k and time period t where x ikt = 1 in scheduling combination l. There is one or more (k,t) pairs in a scheduling combination l because a task may be executed in multiple time periods.

[0130] In this case, the value of the dual variable μ i is:

[0131] μ i = max{0,F(il)}

[0132] The value of the dual variable μ i determines the decision variable ui value of, which means that if the returned scheduling combination l i results in a negative value of F(il), the system will reject task i and set the value of μ i to zero; on the contrary, if μ i > 0, the system accepts task i and executes it according to the computing node and time period specified by the scheduling combination l i specified.

[0133] Step 305-3, according to the function representations of each scheduling combination, use the maximum independent variable point set function to determine the optimal scheduling combination.

[0134] In one implementation, Step 305-3 can be implemented using the following formula:

[0135]

[0136] where l i represents the optimal scheduling combination.

[0137] Step 306, determine whether to accept the model fine-tuning task according to the optimal scheduling combination.

[0138] Step 307, if the model fine-tuning task is accepted, control the target computing node to execute the model fine-tuning task based on the fine-tuning data set.

[0139] After the target computing node executes the model fine-tuning task, return the result to the client before the deadline.

[0140] Step 308, calculate the actual payment cost according to the value of the dual variable corresponding to the optimal scheduling combination and the preset pricing function.

[0141] In one implementation, the actual payment cost is calculated using the following formula:

[0142]

[0143] Or expressed as:

[0144]

[0145] where the values of x ikt and z in are taken from the scheduling combination l; (k, t) ∈ l means the corresponding k and t when x ikt = 1 in the scheduling combination l. can be regarded as the marginal prices of computing, video memory, and video memory bandwidth after processing task i-1. The user's bid price affects whether they can win the auction. If they win, the remuneration the user needs to pay depends on the amount of resources consumed by the task execution.

[0146] Step 309: Return the actual payment amount to the client so that the user can make the payment according to the actual payment amount.

[0147] This embodiment has the following beneficial effects:

[0148] (1) When this embodiment involves the large model fine-tuning task scheduling algorithm, it considers more dimensions of resource constraints, including computing resources, video memory resources, and video memory bandwidth resources. Combining the computing requirements of the task itself and the fine-tuning dataset, the target computing node is determined, fully considering the key characteristics of the large model fine-tuning task itself and the characteristics of the computing node, and improving the resource utilization rate of the computing node.

[0149] (2) In order to adapt to the changing market supply and demand relationship, this embodiment designs a pricing scheme for users and intelligent computing service providers based on the auction mechanism. The intelligent computing service provider is the auctioneer, and the user is the bidder, improving the flexibility of pricing.

[0150] Embodiment Four

[0151] Refer to Figure 6 , which shows a schematic structural diagram of a scheduling center provided by Embodiment Four of the present invention. It may include the following modules:

[0152] Task receiving module 401, configured to receive the model fine-tuning task sent by the client, where the model fine-tuning task includes the user's bid, the fine-tuning dataset, and the computing requirements;

[0153] Optimal scheduling combination determination module 402, configured to determine the optimal scheduling combination of the model fine-tuning task with the goal of maximizing the total revenue based on the bid, the fine-tuning dataset, and the computing requirements, in combination with the node capability information of each computing node in the computing cluster, where the optimal scheduling combination includes the target computing nodes selected from the computing cluster that meet the computing requirements; the node capability information includes computing power, video memory capability, and video memory bandwidth capability; the total revenue includes the user's revenue and the revenue of the intelligent computing service provider;

[0154] Task acceptance judgment module 403, configured to judge whether to accept the model fine-tuning task according to the optimal scheduling combination;

[0155] Task execution module 404, configured to, if accepting the model fine-tuning task, control the target computing node to execute the model fine-tuning task based on the fine-tuning dataset.

[0156] In an embodiment of the present invention, the scheduling center further includes the following module:

[0157] Actual cost determination module, configured to calculate the actual payment amount according to the optimal scheduling combination.

[0158] An actual cost return module, configured to return the actual payment cost to the client, so that the user can make a payment according to the actual payment cost.

[0159] In an embodiment of the present invention, the optimal scheduling combination determination module 402 further includes the following modules:

[0160] A raw optimization problem modeling module, based on the bid, the fine-tuning data set, and the calculated demand, and in combination with the node capability information of each computing node in the computing cluster, models a raw optimization problem based on decision variables with the goal of maximizing the total revenue. The decision variables include variables related to data providers, computing nodes, and task acceptance. The data provider is a preprocessing service provider that preprocesses tasks before the computing nodes process the tasks; the raw optimization problem includes a first set of constraint conditions;

[0161] An equivalent problem modeling module, configured to simplify the raw optimization problem into an equivalent problem based on a scheduling combination, where the scheduling combination is a combination of the execution elements of the decision variables that satisfy the constraints for the task itself in the first set of constraint conditions. The execution elements at least include computing node information, data provider information, and task acceptance information;

[0162] A Lagrangian dual problem generation module, configured to generate a Lagrangian dual problem corresponding to the raw optimization problem based on the primal-dual algorithm and the equivalent problem;

[0163] An optimal combination determination module, configured to determine an optimal scheduling combination for the model fine-tuning task based on the Lagrangian dual problem.

[0164] In an embodiment of the present invention, the Lagrangian dual problem includes a plurality of dual variables, and the dual variables include variables related to the computing overhead, video memory overhead, and video memory bandwidth overhead of the computing node when executing a task; then the optimal combination determination module is specifically configured to:

[0165] Determine the values of the dual variables corresponding to each scheduling combination according to the Lagrangian dual problem;

[0166] Determine the function representation of each scheduling combination according to the values of the dual variables corresponding to each scheduling combination;

[0167] Determine the optimal scheduling combination by using a maximum independent variable point set function according to the function representations of each scheduling combination.

[0168] In an embodiment of the present invention, the actual cost determination module is specifically configured to:

[0169] Calculate the actual payment cost according to the value of the dual variable corresponding to the optimal scheduling combination and the preset pricing function.

[0170] In one embodiment of the present invention, the total revenue is expressed as follows:

[0171]

[0172] Wherein,

[0173]

[0174] Wherein, U represents the total revenue, U r represents the revenue obtained by all users, U c represents the revenue obtained by the intelligent computing service provider, U i represents the revenue obtained by the user who initiates the model fine-tuning task i, b i represents the bid of the user who initiates the model fine-tuning task i, p i represents the actual payment cost of the user after the model fine-tuning task i is accepted, u i represents the task acceptance information on whether the model fine-tuning task i is accepted, z in represents the data provider information on whether the model fine-tuning task i requires data preprocessing by the data provider n, q in represents the fee charged by the data provider, x ikt represents whether to execute the model fine-tuning task i on the computing node k in the time period t, e ikt represents the system overhead caused by the training of the model fine-tuning task i on the computing node k within the time period t.

[0175] In one embodiment of the present invention, the original optimization problem is expressed as:

[0176]

[0177]

[0178] Wherein, P represents maximizing the total revenue, and the constraint conditions (a1)-(a9) constitute the first constraint condition set;

[0179] (a1) is used to ensure that for each model fine-tuning task i, if the task is accepted and requires data preprocessing, then exactly one data provider is selected at most; f i represents whether the model fine-tuning task i requires data preprocessing;

[0180] (a2) is used to ensure that each model fine-tuning task i runs on at most one computing node k in each time period t;

[0181] (a3) To ensure that each model fine-tuning task i is executed only after it arrives at the system and the data preprocessing is completed; a i Represents the arrival time of model fine-tuning task i; h in Represents the delay required for the data provider to process the training dataset of model fine-tuning task i;

[0182] (a4) To ensure that each model fine-tuning task i is executed before its deadline; d i Represents the deadline of model fine-tuning task i;

[0183] (a5) To ensure that each model fine-tuning task i has completed a sufficient amount of computation; M i Represents the computational requirement of model fine-tuning task i; s ik Represents the amount of computation that can be completed per time period when model fine-tuning task i is executed on computing node k;

[0184] (a6) To constrain the computational capacity of each time period and each computing node; C kp Represents the computational power of computing node k ∈ [K];

[0185] (a7) To constrain the video memory capacity of each time period and each computing node; C km Represents the video memory capacity of computing node k;

[0186] (a8) To constrain the video memory bandwidth capacity of each time period and each computing node; C kg Represents the video memory bandwidth capacity of computing node k;

[0187] (a9) Specifies the value range of each decision variable.

[0188] In one embodiment of the present invention, the equivalent problem is expressed as:

[0189]

[0190]

[0191]

[0192] Among them, P 1 Represents maximizing the total revenue after equivalence, and the constraint conditions (b1)-(b5) form the second constraint condition set; ζ i Refers to the set of all scheduling combinations of model fine-tuning task i that satisfy the necessary constraint conditions; the necessary constraint conditions include constraints (a1)-(a5); b il Represents the increment of the objective function value of problem P when using scheduling combination l to execute model fine-tuning task i; x ilIndicates whether the model fine-tuning task i is executed according to the scheduling combination l; s kt (il), r kt (il), σ kt (il) respectively represent the computing consumption, video memory consumption, and video memory bandwidth consumption when executing the model fine-tuning task i on the computing node k using the scheduling combination l during the time period t.

[0193] In an embodiment of the present invention, the Lagrangian dual problem is expressed as:

[0194]

[0195] Among them, μ i , λ kt , γ kt respectively represent the dual variables associated with the constraints (b1), (b2), (b3), and (b4);

[0196] The function representation is expressed as:

[0197]

[0198] Among them, F(il) represents the function representation of the scheduling combination l; (k, t) ∈ l refers to the (k, t) pairs of the computing node k and the time period t where x ikt = 1 in the scheduling combination l, and there is one or more (k, t) pairs in a scheduling combination l;

[0199] The optimal scheduling combination is expressed as:

[0200]

[0201] Among them, l i represents the optimal scheduling combination;

[0202] Judging whether to accept the model fine-tuning task according to the optimal scheduling combination includes:

[0203] μ i = max{0, F(il)}

[0204] If μ i = 0, it is determined not to accept the model fine-tuning task;

[0205] If μ i = F(il)>0, it is determined to accept the model fine-tuning task;

[0206] The actual payment cost is calculated using the following formula:

[0207]

[0208] The scheduling center provided by the embodiments of the present invention can execute the model fine-tuning method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the model fine-tuning method.

[0209] Embodiment 5

[0210] Refer to Figure 7 , which shows a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, blade servers, mainframe computers, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0211] As Figure 7 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0212] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0213] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the model fine-tuning method based on the intelligent computing platform.

[0214] In some embodiments, the model fine-tuning method based on the intelligent computing platform can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the model fine-tuning method based on the intelligent computing platform described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the model fine-tuning method by any other suitable means (e.g., by means of firmware).

[0215] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0216] The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a dedicated computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0217] In the context of the present invention, a computer-readable storage medium may be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0218] Embodiment Six

[0219] The embodiment of the present invention further provides a computer program product, which includes a computer program that, when executed by a processor, implements the model fine-tuning method based on an intelligent computing platform provided in any embodiment of the present invention.

[0220] It should be understood that various forms of the processes shown above may be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0221] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A model fine-tuning method based on an intelligent computing platform, characterized in that: The intelligent computing platform includes a user terminal, a dispatching center, and a computing cluster. The method is applied to the dispatching center side. The method includes: Receiving a model fine-tuning task sent by the user terminal, wherein the model fine-tuning task at least includes a user's bid, a fine-tuning data set, and a computing demand; Based on the bid, the fine-tuning data set and the computing demand, combined with the node capacity information of each computing node in the computing cluster, with the goal of maximizing the total benefit, determine the optimal scheduling combination of the model fine-tuning task, wherein the optimal scheduling combination includes a target computing node selected from the computing cluster that meets the computing demand; the node capacity information includes computing capacity, video memory capacity and video memory bandwidth capacity; the total benefit includes the user's benefit and the benefit of the intelligent computing service provider; Determine whether to accept the model fine-tuning task according to the optimal scheduling combination; If the model fine-tuning task is accepted, the target computing node is controlled to execute the model fine-tuning task based on the fine-tuning data set.

2. The method according to claim 1, characterized in that After determining the optimal scheduling combination of the model fine-tuning tasks, the method further includes: Calculating the actual payment fee according to the optimal scheduling combination; The actual payment fee is returned to the user end so that the user can pay according to the actual payment fee.

3. The method according to claim 1 or 2, characterized in that: The method of determining the optimal scheduling combination of the model fine-tuning task based on the bid, the fine-tuning data set, and the computing demand, combined with the node capacity information of each computing node in the computing cluster, with the goal of maximizing the total benefit, includes: Based on the bid, the fine-tuning data set and the computing demand, combined with the node capacity information of each computing node in the computing cluster, with the goal of maximizing the total benefit, modeling an original optimization problem based on decision variables, wherein the decision variables include variables related to the data supplier, the computing node and the task acceptance, and the data supplier is a pre-processing service provider that pre-processes the task before the computing node processes the task; the original optimization problem includes a first set of constraints; Simplifying the original optimization problem into an equivalent problem based on scheduling combination, wherein the scheduling combination is a combination of execution elements of the decision variables for the task constraints themselves in satisfying the first constraint condition set, and the execution elements at least include computing node information, data supplier information, and task acceptance information; Based on the primal-dual algorithm and the equivalent problem, a Lagrangian dual problem corresponding to the primal optimization problem is generated; An optimal scheduling combination of the model fine-tuning tasks is determined based on the Lagrangian dual problem.

4. The method according to claim 3, characterized in that The Lagrangian dual problem includes a plurality of dual variables, wherein the dual variables include variables related to the computing overhead, video memory overhead, and video memory bandwidth overhead of the computing node when executing the task; The determining the optimal scheduling combination of the model fine-tuning tasks based on the Lagrangian dual problem includes: According to the Lagrangian dual problem, determining the value of the dual variable corresponding to each scheduling combination; Determine the function representation of each scheduling combination according to the value of the dual variable corresponding to each scheduling combination; According to the function representation of each scheduling combination, the optimal scheduling combination is determined by using the maximum independent variable point set function.

5. The method according to claim 4, characterized in that The calculating the actual payment fee according to the optimal scheduling combination includes: The actual payment fee is calculated according to the value of the dual variable corresponding to the optimal scheduling combination and a preset pricing function.

6. The method according to claim 4 or 5, characterized in that: The total revenue is expressed as follows: in, Among them, U represents the total revenue, U r represents the benefits obtained by all users, U c represents the revenue obtained by the intelligent computing service provider, U i represents the benefit obtained by the user who initiated the model fine-tuning task i, b i represents the bid of the user who initiated the model fine-tuning task i, p i represents the actual payment of the user after the model fine-tuning task i is accepted, u i Indicates whether the model fine-tuning task i is accepted, z in The data supplier information indicating whether the model fine-tuning task i requires data supplier n to perform data preprocessing, q in represents the fee charged by the data provider, x ikt Indicates whether to execute model fine-tuning tasks i, e on computing node k in time period t ikt represents the system overhead caused by the training of computing node k for model fine-tuning task i in time period t.

7. The method according to claim 6, characterized in that The original optimization problem based on decision variables modeled based on the bid with the goal of maximizing total revenue is expressed as: Among them, P represents the maximization of total benefits, and the constraints a1-a9 constitute the first constraint set; a1 is used to ensure that for each model fine-tuning task i, if the task is accepted and requires data preprocessing, there is at most one data supplier selected; f i Indicates whether model fine-tuning task i requires data preprocessing; a2 is used to ensure that each model fine-tuning task i runs on at most one computing node k in each time period t; a3 is used to ensure that each model fine-tuning task i is executed only after it reaches the system and completes data preprocessing; a i represents the arrival time of model fine-tuning task i; h in represents the latency required by the data provider to process the training dataset for model fine-tuning task i; a4 is used to ensure that each model fine-tuning task i is executed before its deadline; d i represents the deadline for model fine-tuning task i; a5 is used to ensure that each model fine-tuning task i completes sufficient computation; M i represents the computational requirements of model fine-tuning task i; s ik represents the amount of computation that can be completed in each time period when executing model fine-tuning task i on computing node k; a6 is used to constrain the computing capacity of each computing node in each time period; C kp represents the computing power of computing node k∈[K]; a7 is used to constrain the video memory capacity in each time period and each computing node; C km represents the memory capacity of computing node k; a8 is used to constrain the memory bandwidth capacity in each time period and each computing node; C kg represents the memory bandwidth capability of computing node k; a9 specifies the value range of each decision variable.

8. The method according to claim 7, characterized in that The equivalent problem is expressed as: Among them, P1 represents the maximization of the total benefit after equivalence, and the constraints b1-b5 constitute the second constraint set; ζ i It refers to the set of all scheduling combinations of model fine-tuning task i that meet the necessary constraints; the necessary constraints include constraints a1-a5; b i x represents the increment of the objective function value of problem P when using the scheduling combination l to execute the model fine-tuning task i; x il Indicates whether the model fine-tuning task i is executed according to the scheduling combination l; s kt (il), r kt (il),σ kt (il) represent the computational consumption, memory consumption, and memory bandwidth consumption when using scheduling combination l to execute model fine-tuning task i on computing node k in time period t.

9. The method according to claim 8, characterized in that The Lagrangian dual problem is expressed as: Among them, μ i , kt , γ kt denote the dual variables associated with constraints b1-b4 respectively; The function representation of each scheduling combination determined according to the value of the dual variable corresponding to each scheduling combination is expressed as: Among them, F(il) represents the function representation of the scheduling combination l; (k,t)∈l refers to the x in the scheduling combination l. ikt = 1 for a pair of computing nodes k and time period t. There are one or more (k, t) pairs in a scheduling combination l. According to the function representation of each scheduling combination, the optimal scheduling combination is determined by using the maximum independent variable point set function, which is expressed as: Among them, l i represents the optimal scheduling combination; The determining whether to accept the model fine-tuning task according to the optimal scheduling combination includes: μ i =max{0,F(il)} If μ i =0, it is determined that the model fine-tuning task is not accepted; If μ i =F(il)>0, it is determined that the model fine-tuning task is accepted; The actual payment fee is calculated according to the optimal scheduling combination, using the following formula:

10. An intelligent computing platform, characterized in that: The intelligent computing platform includes a user terminal, a dispatch center, and a computing cluster. The dispatch center side includes the following modules: A task receiving module, used to receive a model fine-tuning task sent by the user terminal, wherein the model fine-tuning task includes a user's bid, a fine-tuning data set, and a computing demand; An optimal scheduling combination determination module is used to determine the optimal scheduling combination of the model fine-tuning task based on the bid, fine-tuning data set and computing demand, combined with the node capacity information of each computing node in the computing cluster, with the goal of maximizing the total benefit, wherein the optimal scheduling combination includes a target computing node selected from the computing cluster that meets the computing demand; the node capacity information includes computing capacity, video memory capacity and video memory bandwidth capacity; the total benefit includes the user's benefit and the benefit of the intelligent computing service provider; A task acceptance judgment module is used to judge whether to accept the model fine-tuning task according to the optimal scheduling combination; The task execution module is used to control the target computing node to execute the model fine-tuning task based on the fine-tuning data set if the model fine-tuning task is accepted.

Citation Information

Patent Citations

  • Dynamic pricing and deploying method for machine learning task in edge cloud network

    CN114139730A

  • Multi-strategy intelligent scheduling method and device oriented to heterogeneous computing power

    CN115237581A

  • Block chain-based trusted distributed computing unloading method

    CN115981807A

  • Dynamic heterogeneous resource management method based on benefit optimization and load awareness in collaborative edge intelligence

    CN117573341A

  • Resource allocation method for maximum profit in computing power network scene

    CN117793104A