Method, device and storage medium for generating computing power leasing solution

By accurately dividing the task computing power levels and dynamic matching cluster configurations, a computing power leasing solution that adapts to business scenarios is generated, which solves the problem of high computing power leasing costs and achieves accurate matching of computing power supply and demand and cost optimization.

CN120234156BActive Publication Date: 2025-08-15SHENZHEN JIEYI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510713611.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-15
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

In the prior art, computing power leasing plans for different business scenarios lack flexibility, resulting in mismatch between execution overhead and expected execution overhead, which increases computing power leasing costs.

Method used

By determining the computing power levels corresponding to each task in the target business scenario, and building a computing power cluster based on the cluster type and cluster configuration information of the computing power level, scheduling task execution, and generating an adaptive computing power leasing plan to achieve accurate matching of computing power supply and demand.

Benefits of technology

It reduces the execution cost of a single task, generates a computing power leasing solution that is adapted to business scenarios, and reduces the total execution cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234156B_ABST
    Figure CN120234156B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device and storage medium for generating a computing power leasing plan, which relates to the field of data processing technology. The disclosed method for generating a computing power leasing plan includes: determining the computing power level corresponding to each task of the target business scenario, and determining the cluster type and cluster configuration information corresponding to each computing power level; building a computing power cluster corresponding to each computing power level according to the cluster type and cluster configuration information corresponding to each computing power level; scheduling each task to the corresponding computing power cluster for execution to obtain a total execution overhead; if the total execution overhead meets the expected execution overhead, generating a computing power leasing plan adapted to the target business scenario according to the cluster type and cluster configuration information corresponding to each computing power level, thereby solving the problem of high computing power leasing costs in the current computing power leasing plans generated based on specific business scenarios and reducing computing power leasing costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, device, and storage medium for generating a computing power leasing solution. Background Art

[0002] With the development of cloud computing and edge computing, efficient utilization of computing resources has become a core requirement for enterprises to reduce costs and increase efficiency. Business scenarios encompass a wide variety of tasks, encompassing real-time inference, model training, big data analysis, and other types, each with significantly different computing resource requirements. For example, real-time tasks such as financial trading and autonomous driving are extremely latency-sensitive and require low-latency network support, but do not demand high single-node computing power. Compute-intensive tasks such as AI (Artificial Intelligence) model training require massively parallel computing power and rely on high-performance GPUs (Graphics Processing Units) or dedicated accelerators. Therefore, adapting appropriate computing power leasing methods to different business scenarios is crucial. Related technologies typically adapt fixed computing power leasing solutions to different business scenarios. If a fixed computing power leasing solution is used for each business scenario, the execution overhead of the recommended computing power leasing solution may differ significantly from the expected execution overhead, increasing computing power leasing costs. Summary of the Invention

[0003] The main purpose of this application is to provide a method, device and storage medium for generating a computing power leasing solution, aiming to solve the technical problem of high computing power leasing costs in current computing power leasing solutions generated based on specific business scenarios.

[0004] To achieve the above objectives, this application proposes a method for generating a computing power leasing solution, including:

[0005] Determine the computing power level corresponding to each task in the target business scenario, and determine the cluster type and cluster configuration information corresponding to each computing power level;

[0006] Build computing clusters corresponding to each computing power level based on the cluster type and cluster configuration information corresponding to each computing power level;

[0007] Schedule each task to the corresponding computing cluster for execution to obtain the total execution cost;

[0008] If the total execution overhead meets the expected execution overhead, a computing power leasing plan that is adapted to the target business scenario is generated based on the cluster type and cluster configuration information corresponding to each computing power level.

[0009] In one embodiment, determining the computing power level corresponding to each task in the target business scenario, and determining the cluster type and cluster configuration information corresponding to each computing power level includes:

[0010] Constructing first prompt text information according to the task type of each task in the target business scenario and the first requirement description text;

[0011] Input the first prompt text information into the first pre-trained large language model to obtain the computing power level corresponding to each task of the target business scenario;

[0012] Constructing second prompt text information based on the computing power level corresponding to each task of the target business scenario, the task information corresponding to each task, and the second requirement description text;

[0013] The second prompt text information is input into the second pre-trained large language model to obtain the cluster type and cluster configuration information corresponding to each computing power level.

[0014] In one embodiment, the method for generating a computing power leasing plan further includes:

[0015] If the total execution overhead does not meet the expected execution overhead, fine-tune the first pre-trained large language model, and use the fine-tuned first pre-trained large language model to obtain the computing power level corresponding to each task of the target business scenario; and / or,

[0016] If the total execution overhead does not meet the expected execution overhead, the second pre-trained large language model is fine-tuned, and the fine-tuned second pre-trained large language model is used to obtain the cluster type and cluster configuration information corresponding to each computing power level.

[0017] In one embodiment, based on the cluster type and cluster configuration information corresponding to each computing power level, building a computing power cluster corresponding to each computing power level includes:

[0018] Constructing third prompt text information based on the cluster type, cluster configuration information, third requirement description text, and sample examples corresponding to each computing power level;

[0019] Inputting the third prompt text information into the third pre-trained large language model to guide the third pre-trained large language model to determine the computing power cluster corresponding to each computing power level based on the cluster type and cluster configuration information corresponding to each computing power level based on the third requirement description text;

[0020] And, based on the sample examples, guide the third pre-trained large language model to determine the output format of the computing power cluster corresponding to each computing power level;

[0021] Output the corresponding computing power cluster based on the output format of each computing power cluster.

[0022] In one embodiment, a computing cluster is constructed from multiple cluster nodes, and each cluster node is of a different node type. Each task is dispatched to the corresponding computing cluster for execution. The total execution overhead includes:

[0023] For a single task, split it into multiple subtasks according to the task type;

[0024] Assign each subtask obtained by splitting to the corresponding cluster node for execution, and obtain the execution overhead corresponding to a single task;

[0025] According to the execution overhead corresponding to each task, the total execution overhead is obtained.

[0026] In one embodiment, each subtask obtained by splitting is assigned to a corresponding cluster node for execution, and the execution overhead corresponding to a single task includes:

[0027] Get the node type corresponding to each cluster node;

[0028] Based on the node type corresponding to each cluster node, each subtask is assigned to the corresponding cluster node for execution, and the execution overhead corresponding to each subtask is obtained;

[0029] Summarize the execution overhead corresponding to each subtask to obtain the execution overhead corresponding to a single task.

[0030] In one embodiment, each subtask obtained by splitting is assigned to a corresponding cluster node for execution, and the execution overhead corresponding to a single task includes:

[0031] Construct a cost matrix based on the node type of each cluster node and each subtask, where the elements in the cost matrix represent the cost of assigning the corresponding subtask to the corresponding cluster node for execution;

[0032] Subtracting the minimum value in each row of the cost matrix from the minimum value in the row, and subtracting the minimum value in each column of the cost matrix from the minimum value in the column, so that each row has at least one zero element and each column has at least one zero element;

[0033] Use the fewest straight lines to cover all zero elements in the cost matrix;

[0034] If the number of straight lines is equal to the order of the cost matrix, the task allocation scheme is obtained based on the zero elements that are not in the same row and column. The number of straight lines equal to the order of the cost matrix indicates that the number of subtasks is equal to the number of cluster nodes.

[0035] According to the task allocation plan, count the cluster nodes to which each subtask is assigned;

[0036] For each subtask, the execution overhead of the subtask on the corresponding cluster node is determined according to the node type of the cluster node to which the subtask is assigned.

[0037] In one embodiment, the total execution overhead obtained based on the execution overhead corresponding to each task includes:

[0038] Determine the weight coefficient corresponding to each task according to the task type of each task;

[0039] A weighted calculation is performed based on the weight coefficient corresponding to each task and the execution cost corresponding to each task to obtain the total execution cost.

[0040] In addition, to achieve the above-mentioned purpose, the present application also proposes a device for generating a computing power leasing plan, including: a memory, a processor, and a computer program stored in the memory and runnable on the processor, the computer program being configured to implement the steps of the method for generating a computing power leasing plan as described above.

[0041] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the method for generating a computing power leasing solution as described above are implemented.

[0042] This application determines the computing power level corresponding to each task of the target business scenario, and determines the cluster information and cluster configuration information corresponding to each computing power level; then, according to the cluster type and cluster configuration information corresponding to each computing power level, builds a computing power cluster corresponding to each computing power level; then, schedules each task to the corresponding computing power cluster for execution, and obtains the total execution overhead corresponding to all tasks; finally, if it is detected that the total execution overhead meets the expected execution overhead, a computing power leasing plan adapted to the target business scenario is generated according to the cluster type and cluster configuration information corresponding to each computing power level. Compared with related technologies, this application achieves precise matching of computing power supply and demand by accurately dividing task computing power levels, dynamically matching cluster configurations, intelligently scheduling task execution, and optimizing computing power leasing plans based on execution overhead feedback, thereby reducing the execution cost of a single task and thus reducing the total execution cost. Finally, a computing power leasing plan adapted to the business scenario is generated, which reduces the computing power leasing cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0045] Figure 1 A flowchart illustrating an embodiment of a method for generating a computing power leasing solution for this application;

[0046] Figure 2 This is a detailed flowchart of step S10 of the method for generating a computing power leasing solution for this application;

[0047] Figure 3 This is a detailed flowchart of step S20 of the method for generating a computing power leasing solution for this application;

[0048] Figure 4 This is a detailed flowchart of step S30 of the method for generating a computing power leasing solution for this application.

[0049] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0050] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0051] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0052] Amid the booming development of cloud computing and edge computing, efficient utilization of computing resources has become a core requirement for enterprises to reduce costs and increase efficiency. The computing resource requirements for tasks in different business scenarios vary significantly. For example, real-time tasks such as financial transactions and autonomous driving are extremely sensitive to latency and require low-latency network support, but do not require high single-node computing power. Compute-intensive tasks such as AI model training require massively parallel computing power and rely on high-performance GPUs or dedicated accelerators.

[0053] However, the computing power leasing solutions recommended for business scenarios in existing technologies are often fixed and lack flexibility. These fixed solutions are unable to adjust computing power resource allocation based on dynamic changes in task characteristics, such as fluctuations in task volume or changes in task type. This results in a mismatch between the execution overhead of the recommended computing power leasing solution and the expected execution overhead. For example, computing power resources configured to handle peak task volumes may sit idle during low task volumes, resulting in wasted resources. Alternatively, when task volumes exceed expectations, computing power resources may be insufficient to support efficient task execution, leading to task delays or failures and increased execution overhead.

[0054] This mismatch increases computing power leasing costs and reduces the cost-effectiveness of enterprises. Therefore, it is crucial to adapt the appropriate computing power leasing method to different business scenarios.

[0055] In response to the above problems, this application proposes a method for generating a computing power leasing plan. The main technical solutions include: determining the computing power level corresponding to each task of the target business scenario, and determining the cluster information and cluster configuration information corresponding to each computing power level; then, building a computing power cluster corresponding to each computing power level according to the cluster type and cluster configuration information corresponding to each computing power level; then, scheduling each task to the corresponding computing power cluster for execution, and obtaining the total execution overhead corresponding to all tasks; finally, if it is detected that the total execution overhead meets the expected execution overhead, then, based on the cluster type and cluster configuration information corresponding to each computing power level, a computing power leasing plan adapted to the target business scenario is generated. Compared with related technologies, this application achieves precise matching of computing power supply and demand by accurately dividing task computing power levels, dynamically matching cluster configurations, intelligently scheduling task execution, and optimizing computing power leasing plans based on execution overhead feedback. This reduces the execution cost of a single task, thereby reducing the total execution cost, and ultimately generates a computing power leasing plan adapted to the business scenario, thereby reducing the computing power leasing cost.

[0056] It should be noted that the execution entity of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or a device that generates a computing power leasing solution that can implement the above functions. The following uses a device that generates a computing power leasing solution as an example to illustrate this embodiment and the following embodiments.

[0057] It should be noted that the method for generating a computing power leasing plan in the embodiment of the present application can be applied to a computing power provider. When the computing power leasing plan generating device of the computing power provider receives a computing power leasing request from a computing power leasing requester, it determines the computing power level corresponding to each task of the target business scenario according to the computing power leasing request, and determines the cluster type and cluster configuration information corresponding to each computing power level. According to the cluster type and cluster configuration information corresponding to each computing power level, it builds a computing power cluster corresponding to each computing power level; schedules each task to the corresponding computing power cluster for execution, and obtains the total execution overhead corresponding to all tasks; if the total execution overhead meets the expected execution overhead, it generates a computing power leasing plan adapted to the target business scenario according to the cluster type and cluster configuration information corresponding to each computing power level. Finally, the computing power provider recommends the generated computing power leasing plan to the computing power leasing requester, so that the computing power leasing requester evaluates whether to lease computing power based on the computing power leasing plan to reduce the computing power leasing cost.

[0058] Based on this, the embodiment of the present application provides a method for generating a computing power leasing solution, referring to Figure 1 , Figure 1 This is a flowchart of an embodiment of a method for generating a computing power leasing solution for this application.

[0059] In this embodiment, the method for generating a computing power leasing plan includes steps S10 to S40:

[0060] Step S10: Determine the computing power level corresponding to each task in the target business scenario, and determine the cluster type and cluster configuration information corresponding to each computing power level.

[0061] Among them, the target business scenario refers to the specific business field or application scenario that requires computing power support, such as financial transactions, autonomous driving, AI model training, etc.

[0062] A task refers to the specific computational or data processing activities that need to be performed within a target business scenario. Different target business scenarios correspond to different task types and quantities. For example, in a financial trading scenario, tasks might include order processing, risk control, and sensor data processing. In an autonomous driving scenario, tasks might include sensor data processing, environmental perception, decision-making and planning, and control execution.

[0063] Among them, computing power level refers to different computing power levels divided according to the computing requirements of the task, such as low-latency level, medium computing level, high computing level, etc.

[0064] The cluster type refers to the computing resource organization form selected based on the computing power level and task characteristics, such as low-latency network clusters, GPU-intensive computing clusters, and storage-intensive computing clusters. Each computing power level can correspond to one or more cluster types. For example, the cluster types corresponding to the low-latency level can include GPU-intensive computing clusters and storage-intensive computing clusters.

[0065] Cluster configuration information refers to detailed information describing the hardware configuration, software configuration, and network configuration of nodes in the cluster.

[0066] For example, in a financial transaction scenario, different tasks have different computing power requirements. Take the financial transaction scenario that includes order processing tasks, risk control tasks, and data analysis tasks as an example:

[0067] Order processing tasks are characterized by high real-time requirements, low computational complexity, and medium data volumes. Therefore, the chosen computing power tier can be a low-latency tier to ensure real-time and low-latency order processing. The cluster type can be a low-latency network cluster; the node type can be equipped with general-purpose servers with low-latency network interfaces; the number of nodes can be dynamically adjusted based on transaction volume to ensure redundancy; the hardware configuration can utilize multi-core CPUs (Central Processing Units), high-bandwidth network cards, and low-latency storage devices; and the software configuration can utilize a low-latency transaction processing framework and an in-memory database.

[0068] Risk control tasks require high real-time performance, moderate computational complexity, and large data volumes. Therefore, the chosen computing power tier is likely to be medium-capacity and low-latency, providing moderate computing power while ensuring low latency. Cluster types can be a hybrid of computing and low-latency clusters. Node types can be configured with servers featuring medium-performance CPUs and low-latency network interfaces. The number of nodes can be adjusted based on risk monitoring needs, supporting distributed computing. Hardware configurations can include multi-core CPUs, high-bandwidth network cards, and medium-capacity memory. Software configurations can utilize real-time data stream processing frameworks and risk model computation libraries.

[0069] Data analysis tasks are characterized by low real-time requirements, high computational complexity, and large data volumes. Therefore, the chosen computing power level can be a high-level computing power level, providing large-scale parallel computing capabilities. The cluster type can be a high-performance computing cluster; the node type can be equipped with servers equipped with high-performance CPUs or GPUs; the number of nodes can be adjusted based on the data analysis task volume, supporting multi-node parallel processing; the hardware configuration can utilize multi-core CPUs / GPUs, large memory capacities, and high-speed storage devices; and the software configuration can utilize distributed computing frameworks and machine learning libraries.

[0070] For example, in autonomous driving scenarios, different tasks have significantly different requirements for computing power levels. Take sensor data processing tasks and environmental perception tasks as examples:

[0071] Sensor data processing tasks require extremely high real-time performance, low computational complexity, and large data volumes. Therefore, the chosen computing power level can be a low-latency level to ensure real-time processing and transmission of sensor data. The cluster type can be an edge computing cluster; the node type can be an embedded device or edge server equipped with a low-latency network interface; the number of nodes can be deployed based on the number of vehicle sensors, supporting localized computing; the hardware configuration can utilize multi-core CPUs, FPGA (Field Programmable Gate Array) or ASIC (Application-Specific Integrated Circuit) accelerators, and high-bandwidth network modules; and the software configuration can utilize a real-time operating system and a sensor data preprocessing framework.

[0072] Environmental perception tasks require high real-time performance, moderate computational complexity, and large data volumes. Therefore, the chosen computing power level can be medium while ensuring low latency. The cluster type can be an edge-cloud collaborative computing cluster; the node types can be edge servers and cloud servers equipped with GPUs or NPUs (Neural Network Processing Units); the number of nodes can be adjusted based on environmental perception needs, supporting distributed computing; the hardware configuration can include GPUs / NPUs, high-bandwidth network cards, and medium-capacity memory; and the software configuration can utilize deep learning frameworks such as TensorRT (primarily used to optimize and accelerate the inference process of deep learning models) and real-time perception algorithms.

[0073] In one feasible implementation, determining the computing power tier corresponding to each task in the target business scenario involves: analyzing each task's characteristics, including real-time requirements, computational complexity, and data volume; and classifying tasks into different computing power tiers based on these characteristics. For example, tasks with high real-time requirements and low computational complexity, such as order processing, are classified into the low-latency tier; tasks with high computational complexity and low real-time requirements, such as AI model training, are classified into the high-computing tier.

[0074] In another feasible implementation, determining the computing power tier corresponding to each task in the target business scenario may include: collecting historical task execution data for the target business scenario; using machine learning algorithms to establish a mapping model between task characteristics and computing power tiers; and predicting the computing power tier of new tasks based on this mapping model. For example, a decision tree algorithm may be used to predict the computing power tier of a task based on characteristics such as its real-time nature, computational complexity, and data volume.

[0075] In one feasible implementation, determining the cluster type and cluster configuration information corresponding to each computing power tier includes: defining cluster types and configuration rules corresponding to different computing power tiers; and selecting the appropriate cluster type and cluster configuration information for each computing power tier based on the configuration rules. For example, for a low-latency tier, a low-latency network cluster is selected, configured with multi-core CPUs, high-bandwidth network cards, and low-latency storage devices; for a high-computing tier, a GPU-intensive computing cluster is selected, configured with high-performance GPUs and large-capacity memory.

[0076] In another possible implementation, determining the cluster type and cluster configuration information corresponding to each computing power tier includes: performing performance benchmark tests on different cluster types and configurations; and selecting the cluster type and cluster configuration information with the best performance for each computing power tier based on the performance benchmark test results. For example, performing latency tests on low-latency network clusters and general-purpose computing clusters, and selecting the cluster type with the lowest latency for the low-latency tier.

[0077] Step S20: Build a computing power cluster corresponding to each computing power level based on the cluster type and cluster configuration information corresponding to each computing power level.

[0078] A computing power cluster refers to a collection of actual computing resources built based on the cluster type and cluster configuration information. Each computing power level has a corresponding computing power cluster.

[0079] In one feasible implementation, a cloud computing platform that supports multiple cluster types can be selected; based on the cluster type and configuration information, corresponding computing clusters can be created on the cloud computing platform. For example, to create a low-latency network cluster on AWS (Amazon's cloud computing service platform, which provides a wide range of powerful cloud computing services), EC2 instances equipped with low-latency network interfaces can be selected. (EC2 is Amazon AWS's Elastic Compute Cloud service, allowing users to rent virtual computers and run applications in the cloud). To create a GPU-intensive computing cluster on Alibaba Cloud, ECS instances equipped with high-performance GPUs can be selected. (ECS is Alibaba Cloud's cloud computing-based virtual server.)

[0080] Another feasible implementation involves using containerization technologies such as Docker and Kubernetes to encapsulate cluster configuration information. Based on the cluster type and configuration, computing clusters can be rapidly deployed and scaled. For example, Kubernetes can be used to deploy a low-latency network cluster, defining node types, hardware configurations, and network configurations using YAML files (a highly readable format for expressing serialized data). Alternatively, Docker containers can be used to encapsulate GPU-intensive computing environments, enabling rapid deployment and resource sharing.

[0081] Step S30: Schedule each task to the corresponding computing cluster for execution to obtain the total execution overhead.

[0082] Among them, the total execution overhead refers to the overall resource consumption required to execute all tasks on the corresponding computing power cluster, including computing time, storage space, network bandwidth, etc.

[0083] It's important to note that different tasks have corresponding computing power levels, and each computing power level has a corresponding cluster type and cluster configuration information. Therefore, each computing power level has a corresponding computing power cluster, and it's also known that different tasks are scheduled to the corresponding computing power cluster. Each task can be scheduled to a single computing power cluster, or multiple computing power clusters can jointly execute the same task.

[0084] In one feasible implementation, scheduling tasks to corresponding computing clusters involves defining task scheduling rules, matching task characteristics with computing power levels, and scheduling tasks to the corresponding computing clusters. For example, low-latency tasks are scheduled to low-latency network clusters, while high-computing tasks are scheduled to GPU-intensive computing clusters.

[0085] In another feasible implementation, scheduling tasks to corresponding computing clusters involves: collecting cluster resource usage and task execution requirements; optimizing task scheduling using intelligent algorithms such as genetic algorithms and reinforcement learning; and dispatching tasks to the optimal computing cluster for execution. For example, reinforcement learning algorithms can be used to dynamically adjust task scheduling strategies based on cluster load and task characteristics.

[0086] In one feasible implementation, obtaining the total execution cost corresponding to all tasks includes: collecting resource usage data from each computing cluster during task execution; calculating the execution cost of each task based on the monitoring data; and summing up the execution costs of all tasks to obtain the total execution cost. For example, using the Prometheus monitoring tool to collect resource data such as CPU utilization, memory usage, and network bandwidth; and calculating the computation time, storage space, and network bandwidth costs of each task based on the resource data and task execution time.

[0087] In another feasible implementation, obtaining the total execution cost corresponding to all tasks includes: establishing a task execution cost prediction model; predicting the execution cost of each task based on task characteristics and cluster configuration information; and summarizing the predicted costs of all tasks to obtain the total execution cost. For example, a regression analysis algorithm is used to predict the execution time, storage space, network bandwidth, and other costs of the task based on its computational complexity, data volume, and cluster configuration information.

[0088] In step S40, if the total execution overhead meets the expected execution overhead, a computing power leasing plan adapted to the target business scenario is generated based on the cluster type and cluster configuration information corresponding to each computing power level.

[0089] Among them, expected execution overhead refers to the user or system's expected target for task execution overhead, which is used to measure the rationality of the computing power leasing plan.

[0090] Among them, the computing power leasing plan refers to a computing resource leasing plan formulated based on the target business scenario and computing power requirements, including cluster type, configuration information, leasing duration and cost, etc.

[0091] In one feasible implementation, data on leasing costs for different cluster types and configurations can be collected. Based on the total execution cost and expected execution cost, the optimal cluster type and configuration can be selected. This can then generate a computing power leasing plan tailored to the target business scenario. For example, for low-latency tasks, a low-latency network cluster with lower leasing costs can be selected; for high-computing tasks, a GPU-intensive computing cluster with higher leasing costs but better performance can be selected.

[0092] Another feasible implementation involves establishing a performance-cost balance model. Based on the total execution overhead and expected execution overhead, a cluster type and configuration that balances performance and cost can be selected. This can then generate a computing power leasing solution tailored to the target business scenario. For example, a multi-objective optimization algorithm can be used to select a low-cost cluster type and configuration while meeting performance requirements.

[0093] In another feasible implementation, if the total execution cost is less than or equal to the expected execution cost, a computing power leasing plan adapted to the target business scenario is generated based on the cluster type and cluster configuration information corresponding to each computing power tier, thereby recommending the optimal computing power leasing plan. The computing power leasing plan includes the cluster type and cluster configuration information corresponding to each computing power tier, and may also include the configuration cost of each computing power tier and the total configuration cost of all computing power tiers, so that users can evaluate whether to lease computing power based on the computing power leasing plan.

[0094] In this embodiment, by determining the computing power level corresponding to each task of the target business scenario, and determining the cluster information and cluster configuration information corresponding to each computing power level; then, according to the cluster type and cluster configuration information corresponding to each computing power level, build a computing power cluster corresponding to each computing power level; then schedule each task to the corresponding computing power cluster for execution, and obtain the total execution overhead corresponding to all tasks; finally, if it is detected that the total execution overhead meets the expected execution overhead, then according to the cluster type and cluster configuration information corresponding to each computing power level, generate a computing power leasing plan adapted to the target business scenario. Compared with related technologies, this application achieves precise matching of computing power supply and demand by accurately dividing task computing power levels, dynamically matching cluster configurations, intelligently scheduling task execution, and optimizing computing power leasing plans based on execution overhead feedback, thereby reducing the execution cost of a single task and thus reducing the total execution cost, and finally generating a computing power leasing plan adapted to the business scenario, thereby reducing the computing power leasing cost.

[0095] Reference Figure 2 , determine the computing power level corresponding to each task in the target business scenario, and determine the cluster type and cluster configuration information corresponding to each computing power level, including:

[0096] Step S11 : constructing first prompt text information according to the task type of each task in the target business scenario and the first requirement description text.

[0097] The task type is the attribute information of the task. For example, in a financial transaction scenario, the corresponding task types may be order processing, risk control, sensor data processing, etc. In an autonomous driving scenario, the corresponding task types may be sensor data processing, environmental perception, decision planning, and control execution.

[0098] The first requirement description is the original requirements document describing the target business scenario. It provides task background information and serves as the input source for constructing the first prompt text. Different task types correspond to different requirement descriptions. For example, for the order processing task type, the corresponding first requirement description might be: "Please help me determine the corresponding computing power tier for this order processing task type." For the sensor data processing task type, the corresponding first requirement description might be: "Please help me determine the corresponding computing power tier for this sensor data processing task type."

[0099] Among them, the first prompt text information is a model input instruction constructed based on the requirement text and task type, which is used to guide the first pre-trained large language model to output computing power level prediction.

[0100] In a feasible implementation, the first requirement description text corresponding to each task type can be obtained according to the task type of each task, and the single task type and its corresponding first requirement description text can be combined to obtain the prompt text corresponding to each task; finally, the prompt texts corresponding to all tasks are combined to obtain the first prompt text information containing all tasks.

[0101] In another feasible implementation, a mapping table between task types and target business scenarios can be established in advance to convert technical terms into natural language descriptions; a rule template is defined according to the task type, and core parameters are extracted from the original requirement text through regular expressions or keyword matching; the extracted core parameters are filled into the predefined natural language template to form the first prompt text information.

[0102] Step S12: input the first prompt text information into the first pre-trained large language model to obtain the computing power level corresponding to each task of the target business scenario.

[0103] Among them, the first pre-trained large language model is a basic model used to analyze demand and predict computing power levels, and is used to convert natural language demand into structured computing power demand.

[0104] In one feasible implementation, the first prompt text information is input into a first pre-trained large language model, and the first prompt text information is parsed using the first pre-trained large language model to obtain a first text feature vector of the first prompt text information; similarity is calculated between the first text feature vector and a preset text feature vector in a first preset database to obtain a first similar text feature vector; and the preset computing power level associated with the first similar text feature vector is determined as the matching computing power level. The above-mentioned similarity calculation can be a cosine similarity measurement, or can be similarity obtained by calculating the Euclidean distance, or can be other similarity calculation methods.

[0105] Step S13: Construct second prompt text information according to the computing power level corresponding to each task of the target business scenario, the task information corresponding to each task, and the second requirement description text.

[0106] Among them, task information is a structured description containing metadata such as task type, input and output specifications, etc., which is used to establish a mapping relationship between tasks and computing resources.

[0107] The second requirement description text is a detailed requirement that supplements the description of cluster configuration constraints and is used to introduce non-functional requirements into configuration decisions.

[0108] Among them, the second prompt text information is a model input instruction that integrates computing power level and task constraints, which is used to drive the model to generate a configuration plan that meets business constraints.

[0109] In a feasible implementation, the second requirement description text corresponding to each task can be obtained based on the computing power level and task information corresponding to each task, and the single task and its corresponding second requirement description text can be combined to obtain the prompt text corresponding to each task; finally, the prompt texts corresponding to all tasks are combined to obtain the second prompt text information containing all tasks.

[0110] In another feasible implementation, a mapping table of computing power levels, task information, and target business scenarios corresponding to each task can be pre-established to convert technical terms into natural language descriptions; rule templates are defined based on the computing power levels and task information corresponding to each task, and core parameters are extracted from the original requirement text through regular expressions or keyword matching; the extracted core parameters are filled into the predefined natural language template to form the second prompt text information.

[0111] Step S14: Input the second prompt text information into the second pre-trained large language model to obtain the cluster type and cluster configuration information corresponding to each computing power level.

[0112] Among them, the second pre-trained large language model is an enhanced large model specifically used for computing resource optimization, which is used to convert computing power requirements into feasible cluster solutions.

[0113] In one feasible embodiment, the second prompt text information is input into a second pre-trained large language model, and the second prompt text information is parsed and processed using the second pre-trained large language model to obtain a second text feature vector of the second prompt text information; similarity is calculated between the second text feature vector and a preset text feature vector in a second preset database to obtain a second similar text feature vector; and the preset cluster type and preset cluster configuration information associated with the second similar text feature vector are determined as the matching cluster type and cluster configuration information. The similarity calculation described above may be a cosine similarity measurement, a Euclidean distance calculation to obtain similarity, or other similarity calculation methods.

[0114] In this embodiment, through the processing of the pre-trained large language model, the computing power level corresponding to each task of the target business scenario can be automatically determined, and the cluster type and cluster configuration information corresponding to each computing power level can be obtained, thereby improving the efficiency of determining the computing power level, and the cluster type and cluster configuration information corresponding to each computing power level; in addition, the computing power level corresponding to each task is determined by the first pre-trained large language model, and the cluster type and cluster configuration information corresponding to each computing power level is determined by the second pre-trained large language model, thereby improving the computing power level corresponding to each task, and the accuracy of the cluster type and cluster configuration information corresponding to each computing power level.

[0115] Furthermore, the method for generating a computing power leasing plan also includes:

[0116] In step S110 , if the total execution overhead does not meet the expected execution overhead, the first pre-trained large language model is fine-tuned, and the fine-tuned first pre-trained large language model is used to obtain the computing power level corresponding to the target business scenario.

[0117] And / or, in step S120, if the total execution overhead does not meet the expected execution overhead, fine-tune the second pre-trained large language model, and use the fine-tuned second pre-trained large language model to obtain the cluster type and cluster configuration information corresponding to each computing power level.

[0118] The above fine-tuning of the first pre-trained large language model and the second pre-trained large language model can be performed using a model based on the Transformer architecture (Transformer is a deep learning model architecture that uses the self-attention mechanism to capture global dependencies in the input sequence).

[0119] In one feasible implementation, an Adapter structure can be designed and embedded within the Transformer structure. During training, the parameters of the original pre-trained model are fixed, and only the newly added Adapter structure is fine-tuned. This method can reduce the introduction of additional parameters while ensuring efficient training.

[0120] Specifically, the Adapter structure usually contains two main parts: a downsampling layer and an upsampling layer, which may contain a nonlinear activation function in the middle. These layers are usually implemented by fully connected layers. Among them, the downsampling layer maps the input feature vector to a lower-dimensional space to reduce the amount of calculation and avoid overfitting; the nonlinear activation function increases the nonlinear expression ability of the model; the upsampling layer maps the downsampled feature vector back to the original dimension to match the input / output dimensions of other layers in the Transformer model. The Adapter structure is embedded after each Transformer layer of the Transformer model, which usually means inserting the Adapter after each multi-head attention and feedforward network; the Adapter structure is added after each relevant layer of the Transformer model to ensure that the input / output dimensions of the Adapter structure match the input / output dimensions of the Transformer layer. During the training process, all parameters of the pre-trained Transformer model are fixed. This can be achieved by not updating the gradients of these parameters during training; only training the parameters of the Adapter structure while keeping the parameters of the pre-trained model unchanged, which can be achieved by only updating the gradients of the Adapter structure during training; usually using small random values to initialize the parameters of the Adapter structure; setting an optimizer for the Adapter structure and specifying hyperparameters such as the learning rate; during training, only calculating and updating the gradients of the Adapter structure, not the gradients of the pre-trained model. The model is trained using the training dataset and evaluated using the validation dataset; during training, changes in loss functions such as cross-entropy loss and the performance of the model on the validation dataset, such as accuracy and F1 score, can be monitored; if the validation performance does not improve over multiple consecutive training sessions, training is stopped to avoid overfitting; learning rate adjustment: adjust the learning rate according to changes in validation performance, such as using learning rate decay or a learning rate scheduler; finally, after training is complete, the fine-tuned Adapter structure is deployed together with the pre-trained model into the actual application. During inference, only the fine-tuned Adapter structure and the fixed pre-trained model are used for prediction.

[0121] Another feasible implementation indirectly trains some dense layers in the neural network by optimizing the rank decomposition matrix of the dense layers during the adaptation process, while keeping the pre-trained weights unchanged. A bypass is added to the original pre-trained language model, performing a dimensionality reduction and then dimensionality increase operation to simulate the intrinsic rank. During training, the intrinsic rank parameters are fixed, and only the dimensionality reduction matrix A and the dimensionality increase matrix B are trained. Compared with other fine-tuning methods, increasing the number of parameters does not lead to a performance degradation, and the performance is equal to or even better than that of full-parameter fine-tuning. Based on the inherent low-rank characteristics of large language models, this method adds a bypass matrix to simulate full-parameter fine-tuning, which can transform various existing large models into specialized models for different fields through lightweight fine-tuning.

[0122] By applying the above-mentioned implementation method to the present application to fine-tune the first pre-trained large language model and / or the second pre-trained large language model, the computing power level corresponding to each task of the target business scenario is made more accurate through the fine-tuned first pre-trained large language model, and the cluster type and cluster configuration information corresponding to each computing power level are made more accurate through the fine-tuned second pre-trained large language model.

[0123] In this embodiment, when it is detected that the total execution overhead does not meet the expected execution overhead, the first pre-trained large language model and / or the second pre-trained large language model are fine-tuned. The computing power level corresponding to each task of the target business scenario is made more accurate through the fine-tuned first pre-trained large language model, and the cluster type and cluster configuration information corresponding to each computing power level are made more accurate through the fine-tuned second pre-trained large language model.

[0124] Based on the above implementation, refer to Figure 3 According to the cluster type and cluster configuration information corresponding to each computing power level, the computing power cluster corresponding to each computing power level is built, including:

[0125] Step S21: Construct a third prompt text message based on the cluster type, cluster configuration information, third requirement description text, and sample examples corresponding to each computing power level.

[0126] Among them, the sample examples are used to help the third pre-trained large language model understand the third requirement description text and the output format expected by the computing power cluster corresponding to each computing power level.

[0127] Step S22: Input the third prompt text information into the third pre-trained large language model to guide the third pre-trained large language model to determine the computing power cluster corresponding to each computing power level based on the cluster type and cluster configuration information corresponding to each computing power level based on the third requirement description text.

[0128] In one feasible embodiment, the third prompt text information is input into a third pre-trained large language model, and the third prompt text information is parsed and processed using the third pre-trained large language model to obtain a third text feature vector of the third prompt text information; similarity is calculated between the third text feature vector and a preset text feature vector in a third preset database to obtain a third similar text feature vector; and the preset computing power cluster associated with the third similar text feature vector is determined as the matching computing power cluster. The above-mentioned similarity calculation can be a cosine similarity measurement, or can be similarity obtained by calculating the Euclidean distance, or can be other similarity calculation methods.

[0129] And, step S23, based on the sample examples, guide the third pre-trained large language model to determine the output format of the computing power cluster corresponding to each computing power level.

[0130] Among them, the output formats of different computing power clusters are different.

[0131] For example, to create a low-latency network cluster on AWS, choose an EC2 instance equipped with a low-latency network interface. To create a GPU-intensive computing cluster on Alibaba Cloud, choose an ECS instance equipped with a high-performance GPU. Use Kubernetes to deploy a low-latency network cluster, defining the node type, hardware configuration, and network configuration using YAML files. Alternatively, use Docker containers to encapsulate a GPU-intensive computing environment for rapid deployment and resource sharing.

[0132] Step S24: Output the corresponding computing power cluster based on the output format of each computing power cluster.

[0133] In this embodiment, the third prompt text information is processed by the third pre-trained large language model to obtain the cluster type and cluster configuration information corresponding to each computing power level, as well as the output format of the computing power cluster corresponding to each computing power level. This not only improves the accuracy of the cluster type and cluster configuration information corresponding to each computing power level, but also standardizes the output format of the computing power cluster corresponding to each computing power level, which is convenient for subsequent processing.

[0134] Based on the above embodiment, refer to Figure 4 The computing power cluster is built by multiple cluster nodes, and each cluster node has a different node type. Each task is dispatched to the corresponding computing power cluster for execution. The total execution overhead corresponding to all tasks includes:

[0135] Step S31 : for a single task, split the single task into multiple subtasks according to the task type.

[0136] In step S32 , each of the split subtasks is assigned to a corresponding cluster node for execution, thereby obtaining the execution overhead corresponding to a single task.

[0137] The following example uses a financial transaction scenario, which includes three tasks: order processing, risk control, and sensor data processing. Taking the order processing task as an example, it can be broken down into the following subtasks: order reception, trade matching, and settlement confirmation. For the order reception subtask, the corresponding cluster node type can be a GPU instance; for the trade matching subtask, the corresponding cluster node type can be a memory-optimized instance; and for the settlement confirmation subtask, the corresponding cluster node type can be a low-latency instance.

[0138] In one feasible implementation, each of the split subtasks is assigned to a corresponding cluster node for execution. Obtaining the execution overhead corresponding to a single task includes: after the subtasks are assigned to the corresponding cluster node for execution, calculating the execution overhead corresponding to the single task. The execution overhead of a single task includes three components: computing overhead, storage overhead, and network overhead. After calculating the computing overhead, storage overhead, and network overhead of the single task, the computing overhead, storage overhead, and network overhead are aggregated to obtain the execution overhead corresponding to the single task.

[0139] Computing overhead = number of CPU cores * usage duration (seconds) * unit cost (RMB / core-hour) + GPU memory (GB) * usage duration (seconds) * GPU price (RMB / GB-hour). For hot data (SSD): Storage overhead = data volume (GB) * SSD price (RMB / GB-month) * storage duration (month). For cold data (HDD): Storage overhead = data volume (TB) * HDD price (RMB / TB-month) * storage duration (month). Network overhead = data volume (GB) × bandwidth price (RMB / GB) * transmission distance factor (intra-city = 1, inter-city = 1.5, cross-border = 3). SSD (Solid State Drive), HDD (Hard Disk Drive).

[0140] In another feasible implementation, each of the split subtasks is assigned to a corresponding cluster node for execution, and the execution overhead corresponding to the individual task is obtained, including: obtaining the node type corresponding to each cluster node; based on the node type corresponding to each cluster node, assigning each of the split subtasks to a corresponding cluster node for execution to obtain the execution overhead corresponding to each subtask; and summing up the execution overhead corresponding to each subtask to obtain the execution overhead corresponding to the individual task. A mapping relationship between node type and subtask task type can be pre-established, and based on this mapping relationship, each of the split subtasks is assigned to a corresponding cluster node for execution, thereby improving the accuracy of the execution overhead corresponding to each subtask.

[0141] Furthermore, the above-mentioned allocating each subtask obtained by splitting to the corresponding cluster node for execution to obtain the execution overhead corresponding to the single task includes: using the Hungarian algorithm to allocate each subtask obtained by splitting to the corresponding cluster node for execution to obtain the execution overhead corresponding to the single task.

[0142] Specifically, each subtask obtained by splitting is assigned to the corresponding cluster node for execution, and the execution overhead corresponding to a single task is obtained, including: constructing a cost matrix according to the node type of each cluster node and each subtask, wherein the elements in the cost matrix represent the overhead of assigning the corresponding subtask to the corresponding cluster node for execution; subtracting the minimum value in each row of the cost matrix, and subtracting the minimum value in each column of the cost matrix, so that each row has at least one zero element and each column has at least one zero element; using the least straight lines to cover all zero elements in the cost matrix; if the number of straight lines is equal to the order of the cost matrix, a task allocation scheme is obtained according to the zero elements that are not in the same row and in the same column, wherein the number of straight lines equal to the order of the cost matrix indicates that the number of subtasks is equal to the number of cluster nodes; according to the task allocation scheme, the cluster nodes to which each subtask is assigned are counted; for each subtask, the execution overhead of the subtask on the corresponding cluster node is determined according to the node type of the cluster node to which the subtask is assigned.

[0143] For example, the specific implementation process of the Hungarian algorithm is described in detail below:

[0144] First, define the cluster node types. Assume that there are multiple types of nodes in the cluster, such as CPU-intensive nodes, GPU-intensive nodes, memory-intensive nodes, etc. Each node type has different computing capabilities and resource characteristics, and is suitable for executing different types of subtasks. A single task can be split into multiple subtasks and classified according to the computing requirements of the subtasks, such as CPU, GPU, memory, etc. Determine the node type that is suitable for executing each subtask. Create a cost matrix in which rows represent subtasks and columns represent cluster nodes. The element Cij in the cost matrix represents the cost of assigning the i-th subtask to the j-th node. The cost can be quantified based on the computing power of the node and the requirements of the subtask, for example, considering factors such as the node's computing speed, resource utilization, and task completion time.

[0145] Next, for each row of the cost matrix, subtract the minimum value in that row, ensuring that each row has at least one zero element. For each column of the cost matrix, subtract the minimum value in that column, ensuring that each column has at least one zero element. Use a minimum number of lines (horizontal or vertical) to cover all zero elements in the cost matrix. If the number of lines equals the order of the matrix, meaning the number of subtasks equals the number of cluster nodes, then the optimal match is found and the algorithm ends. If the number of lines is less than the order of the matrix, find the minimum value among the elements not covered by a line; subtract this minimum value from the rows not covered by a line; add this minimum value to the columns covered by a line; remove the line and continue searching for the optimal match. When the optimal match is found, select zero elements that are not in the same row and column. These zero elements represent the final task assignment solution: that is, if Cij = 0 and the i-th row and j-th column are not occupied by other zero elements, then assign the i-th subtask to the j-th node.

[0146] Next, based on the optimal allocation solution found by the Hungarian algorithm, the cluster nodes to which each subtask is assigned are counted. For each subtask, the execution overhead on the assigned cluster node is calculated based on the node type and resource characteristics. This overhead can be quantified based on factors such as the cluster node's computing speed, resource utilization, and task completion time.

[0147] Finally, the execution overhead of all subtasks is summarized to obtain the execution overhead of a single task on the cluster node.

[0148] For example, suppose a task is split into three subtasks, and there are three cluster nodes in the cluster, with the node types being CPU-intensive, GPU-intensive, and memory-intensive respectively. The cost matrix is constructed as follows:

[0149]

[0150] Next, apply the Hungarian algorithm to perform row and column subtraction operations:

[0151] Row subtraction: Subtract the minimum value of each row from the row to get the matrix:

[0152]

[0153] Column subtraction: Subtract the minimum value of each column to get the matrix:

[0154]

[0155] Next, cover all zero elements with the least number of straight lines. You can find three zero elements (subtask 1 - CPU-intensive node, subtask 1 - memory-intensive node, subtask 3 - CPU-intensive node, subtask 3 - memory-intensive node), but you need to choose zero elements that are not in the same row and column.

[0156] Select subtask 1 - CPU-intensive nodes, subtask 2 - GPU-intensive nodes, and subtask 3 - memory-intensive nodes to obtain the optimal match, i.e., the task allocation solution.

[0157] Finally, calculate the execution cost:

[0158] The execution cost of subtask 1 on a CPU-intensive node is 5.

[0159] The execution cost of subtask 2 on a GPU-intensive node is 4.

[0160] The execution cost of subtask 3 on a memory-intensive node is 3.

[0161] After determining the execution overhead of the subtasks on the corresponding cluster nodes, the execution overhead of a single task is 5+4+3=12.

[0162] By using the Hungarian algorithm to determine the execution overhead of each subtask on the corresponding cluster node, and then determining the execution overhead corresponding to a single task, the execution overhead between tasks is decoupled and determined, thereby improving the accuracy of the execution overhead of each task.

[0163] Step S33: Obtain the total execution overhead according to the execution overhead corresponding to each task.

[0164] In a feasible implementation manner, obtaining the total execution overhead according to the execution overhead corresponding to each task includes: summing up the execution overhead corresponding to each task to obtain the total execution overhead.

[0165] In another feasible embodiment, obtaining the total execution cost based on the execution costs corresponding to each task includes: determining a weight coefficient corresponding to each task based on the task type; and performing a weighted calculation based on the weight coefficient corresponding to each task and the execution cost corresponding to each task to obtain the total execution cost. The weight coefficients corresponding to different tasks are different.

[0166] In another feasible embodiment, the total execution overhead is obtained based on the execution overhead corresponding to each task, including: for a single task, the weight coefficient corresponding to each subtask can be determined according to the task type of each subtask; weighted calculation is performed based on the weight coefficient corresponding to each subtask and the execution overhead corresponding to each subtask to obtain the execution overhead of the single task. Then, the weight coefficient corresponding to each task is determined according to the task type of each task; weighted calculation is performed based on the weight coefficient corresponding to each task and the execution overhead corresponding to each task to obtain the total execution overhead. Among them, different subtasks have different corresponding weight coefficients due to different task types. Among them, the execution overhead corresponding to each subtask can include three parts: computing overhead, storage overhead and network overhead. Corresponding weight coefficients can also be set for each part of the execution overhead to improve the accuracy of the execution overhead of each subtask, thereby improving the accuracy of the total execution overhead.

[0167] In this embodiment, a single task is split into multiple subtasks and assigned to corresponding cluster nodes for execution, resulting in the execution overhead corresponding to each task. Based on the execution overhead of each task, the total execution overhead is calculated. By adapting the cluster nodes to handle each subtask, resources can be rationally planned and utilized, reducing the execution overhead of each task.

[0168] It should be noted that the above examples are only used to understand this application and do not constitute a limitation on the method of generating the computing power leasing solution of this application. More simple transformations based on this technical concept are all within the scope of protection of this application.

[0169] Based on the same inventive concept, the present application provides a device for generating a computing power leasing plan, comprising: at least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for generating a computing power leasing plan in the above embodiment.

[0170] The computing power leasing solution generation device provided in this application adopts the computing power leasing solution generation method in the above-mentioned embodiment, which can solve the technical problem of high computing power leasing costs in computing power leasing solutions currently generated based on specific business scenarios. Compared with the existing technology, the beneficial effects of the computing power leasing solution generation device provided in this application are the same as the beneficial effects of the computing power leasing solution generation method provided in the above-mentioned embodiment, and the other technical features of the computing power leasing solution generation device are the same as the features disclosed in the previous embodiment method, and are not further described here.

[0171] Based on the same inventive concept, the present application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the method for generating the computing power leasing solution in the above-mentioned embodiment.

[0172] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory, read-only memory, erasable programmable read-only memory (EPROM), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.

[0173] The above-mentioned computer-readable storage medium may be included in the device for generating the computing power leasing solution; or it may exist independently without being assembled into the device for generating the computing power leasing solution.

[0174] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the device generating the computing power leasing plan, the device generating the computing power leasing plan can accurately divide the task computing power levels, dynamically match the cluster configuration, intelligently schedule task execution, and optimize the computing power leasing plan based on execution overhead feedback to achieve accurate matching of computing power supply and demand. By reducing the execution cost of a single task, the total execution cost is reduced, and finally a computing power leasing plan adapted to the business scenario is generated, thereby reducing the computing power leasing cost.

[0175] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0176] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0177] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0178] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned method for generating a computing power leasing solution. This computer-readable storage medium can address the technical issue of high computing power leasing costs associated with current computing power leasing solutions generated based on specific business scenarios. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the method for generating a computing power leasing solution provided in the aforementioned embodiments, and are not further elaborated here.

[0179] The above are only some embodiments of the present application and are not intended to limit the patent scope of the present application. All equivalent structural transformations made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for generating a computing power leasing solution, characterized in that: The method for generating the computing power leasing solution includes: Constructing first prompt text information according to the task type of each task in the target business scenario and the first requirement description text; Inputting the first prompt text information into a first pre-trained large language model to obtain a computing power level corresponding to each of the tasks in the target business scenario; Constructing second prompt text information according to the computing power level corresponding to each task of the target business scenario, the task information corresponding to each task, and the second requirement description text; Inputting the second prompt text information into a second pre-trained large language model to obtain cluster types and cluster configuration information corresponding to each computing power level; Building a computing power cluster corresponding to each computing power level according to the cluster type and cluster configuration information corresponding to each computing power level; Dispatching each of the tasks to the corresponding computing cluster for execution to obtain the total execution cost; If the total execution overhead meets the expected execution overhead, a computing power leasing plan adapted to the target business scenario is generated based on the cluster type corresponding to each computing power level and the cluster configuration information.

2. The method for generating a computing power leasing solution according to claim 1, wherein: The method for generating the computing power leasing plan further includes: If the total execution overhead does not meet the expected execution overhead, fine-tune the first pre-trained large language model, and use the fine-tuned first pre-trained large language model to obtain the computing power level corresponding to each task of the target business scenario; and / or, If the total execution overhead does not meet the expected execution overhead, the second pre-trained large language model is fine-tuned, and the fine-tuned second pre-trained large language model is used to obtain the cluster type and cluster configuration information corresponding to each computing power level.

3. The method for generating a computing power leasing solution according to claim 1, wherein: The step of building a computing power cluster corresponding to each computing power level according to the cluster type and cluster configuration information corresponding to each computing power level includes: Constructing third prompt text information according to the cluster type corresponding to each computing power level, the cluster configuration information, the third requirement description text, and the sample example; Inputting the third prompt text information into a third pre-trained large language model to guide the third pre-trained large language model to determine the computing power cluster corresponding to each computing power level based on the cluster type and cluster configuration information corresponding to each computing power level based on the third requirement description text; and, guiding the third pre-trained large language model based on the sample examples, determining an output format of the computing power cluster corresponding to each computing power level; Output the corresponding computing power cluster based on the output format of each computing power cluster.

4. The method for generating a computing power leasing solution according to any one of claims 1 to 3, wherein: The computing power cluster is constructed by multiple cluster nodes, and each cluster node is of a different node type. The scheduling of each task to the corresponding computing power cluster for execution, and the total execution overhead obtained include: For a single task, split the single task into multiple subtasks according to the task type; Allocate each of the split subtasks to a corresponding cluster node for execution, and obtain the execution overhead corresponding to the single task; According to the execution overhead corresponding to each task, the total execution overhead is obtained.

5. The method for generating a computing power leasing solution according to claim 4, wherein: The execution overhead corresponding to each task obtained by splitting is obtained by assigning each of the subtasks to the corresponding cluster node for execution, and includes: Obtaining the node type corresponding to each of the cluster nodes; Based on the node type corresponding to each cluster node, each of the subtasks obtained by splitting is assigned to the corresponding cluster node for execution, thereby obtaining the execution overhead corresponding to each subtask; The execution overhead corresponding to each of the subtasks is aggregated to obtain the execution overhead corresponding to the single task.

6. The method for generating a computing power leasing solution according to claim 4, wherein: The execution overhead corresponding to each task obtained by splitting is obtained by assigning each of the subtasks to the corresponding cluster node for execution, and includes: Constructing a cost matrix according to the node type of each cluster node and each subtask, wherein the elements in the cost matrix represent the cost of allocating the corresponding subtask to the corresponding cluster node for execution; Subtracting a minimum value in each row of the cost matrix and subtracting a minimum value in each column of the cost matrix so that each row has at least one zero element and each column has at least one zero element; Using a minimum number of straight lines to cover all zero elements in the cost matrix; If the number of the straight lines is equal to the order of the cost matrix, a task allocation scheme is obtained according to the zero elements that are not in the same row and in the same column, wherein the number of the straight lines being equal to the order of the cost matrix indicates that the number of subtasks is equal to the number of cluster nodes; According to the task allocation plan, count the cluster nodes to which each subtask is assigned; For each of the subtasks, the execution overhead of the subtask on the corresponding cluster node is determined according to the node type of the cluster node to which the subtask is assigned.

7. The method for generating a computing power leasing solution according to claim 4, wherein: The total execution overhead obtained according to the execution overhead corresponding to each task includes: Determine the weight coefficient corresponding to each task according to the task type of each task; The total execution overhead is obtained by performing weighted calculation according to the weight coefficient corresponding to each task and the execution overhead corresponding to each task.

8. A device for generating a computing power leasing solution, characterized in that: The device for generating the computing power leasing solution includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for generating the computing power leasing solution according to any one of claims 1 to 7.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the method for generating a computing power leasing solution according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Network computing power arrangement method, device and equipment and computer storage medium

    CN116962528A

  • Mobile phone cloud computing power lease scheduling method for virtual computer

    CN119537035A