Resource recommendation method and related device
By leveraging the cloud management platform to provide cost-effective resource recommendations based on AI models and computing unit cluster resources, the problem of low resource utilization among tenants in the same computing unit cluster is solved, thereby improving resource utilization and tenant experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2026-03-05
AI Technical Summary
In existing technologies, the cloud resources purchased by tenants are concentrated in the same computing unit cluster, resulting in high communication bandwidth and low latency, but also high prices and low utilization of computing unit cluster resources, making it difficult to effectively recommend cost-effective resource solutions.
The cloud management platform provides a resource recommendation scheme that offers better cost-effectiveness than other candidate schemes based on the tenant's AI model information and the available resources of the computing unit cluster. This includes the number of computing unit clusters and the number of computing units in each cluster, supports parallel strategy selection, and improves resource utilization.
It provides tenants with cost-effective resource recommendations, improving their user experience and increasing the resource utilization of the computing unit cluster.
Smart Images

Figure CN2025079288_05032026_PF_FP_ABST
Abstract
Description
A resource recommendation method and related equipment
[0001] This application claims priority to Chinese Patent Application No. 202411214716.1, filed with the China National Intellectual Property Administration on August 30, 2024, entitled “A Resource Recommendation Method and Related Equipment”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of cloud computing, and more specifically, to a resource recommendation method, a computing device, a cluster of computing devices, a computer program product, and a computer-readable storage medium. Background Technology
[0003] With the continuous development of artificial intelligence (AI) technology, tenants can utilize cloud resources provided by cloud vendors to train and / or infer AI models. For tenants, when purchasing cloud resources to run AI models, if the purchased cloud resources are located in the same computing unit cluster, the communication bandwidth between computing units is larger, the communication latency is lower, and the efficiency of model training and / or inference is higher; however, such cloud resources are also more expensive. For cloud vendors, the cloud resources purchased by cloud tenants also affect the resource utilization rate within the computing unit cluster. For example, if the cloud resources purchased by a tenant are concentrated in one or more computing unit clusters, there may be a large number of unpurchased fragmented resources within these clusters. These fragmented resources are difficult to sell to other cloud tenants, resulting in low resource utilization within the computing unit cluster.
[0004] Therefore, how to provide tenants with suitable resource recommendation solutions, improve their user experience, and enhance the resource utilization of computing unit clusters has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides a resource recommendation method, a computing device, a computing device cluster, a computer program product, and a computer-readable storage medium, which can provide tenants with at least one cost-effective resource recommendation scheme, improve the tenant's user experience, and improve the resource utilization of the computing unit cluster.
[0006] Firstly, a resource recommendation method is provided. This method is applied to a cloud management platform for managing infrastructure providing cloud services. This infrastructure includes at least one computing unit cluster, and each computing unit cluster includes at least one computing unit. The method includes: receiving first indication information from a tenant, the first indication information indicating model information of a tenant's first AI model; determining at least one first resource recommendation scheme based on available resources in the at least one computing unit cluster and the model information of the first AI model, wherein the at least one first resource recommendation scheme belongs to a first set of candidate resource recommendation schemes, the cost-effectiveness of the at least one first resource recommendation scheme is better than the cost-effectiveness of candidate resource recommendation schemes other than the at least one first resource recommendation scheme in the first set of candidate resource recommendation schemes, the candidate resource recommendation schemes in the first set of candidate resource recommendation schemes meet the operational requirements of the first AI model, and the resources included in the candidate resource recommendation schemes belong to available resources in the at least one computing unit cluster, the candidate resource recommendation scheme includes the number of computing unit clusters used to run the first AI model and the number of computing units in each computing unit cluster; and providing the at least one first resource recommendation scheme to the tenant.
[0007] In this embodiment, the cloud management platform can provide at least one cost-effective resource recommendation scheme for the tenant based on the model information of the AI model to be run by the tenant and the remaining available resources in the computing unit cluster, so that the tenant can purchase suitable resources and improve the resource utilization rate in the computing unit cluster.
[0008] In conjunction with the first aspect, in some implementations, a second instruction information is received from the tenant, which is used to indicate a target resource recommendation scheme, which belongs to at least one first resource recommendation scheme; based on the second instruction information, resources included in the target resource recommendation scheme are provided to the tenant, which are used to run a first AI model.
[0009] In this embodiment of the application, after the cloud management platform provides at least one first resource recommendation scheme to the tenant, it can also determine the target resource recommendation scheme that the tenant wants to use based on the tenant's selection, thereby providing the tenant with the resources included in the target resource recommendation scheme, that is, allocating the computing units and computing unit clusters included in the target resource recommendation scheme to the tenant to run the first AI model.
[0010] In conjunction with the first aspect, in some implementation methods, a first candidate resource recommendation scheme set is determined based on the available resources in at least one computing unit cluster and the model information of the first AI model. The first candidate resource recommendation scheme set includes at least one candidate resource recommendation scheme. The performance and price of each candidate resource recommendation scheme in the first candidate resource recommendation scheme set are determined. Based on the performance and price of each candidate resource recommendation scheme in the first candidate resource recommendation scheme set, at least one first resource recommendation scheme is determined.
[0011] In this embodiment, the cloud management platform determines at least one candidate resource recommendation scheme that meets the operation requirements of the first AI model and includes resources that are available resources, based on the available resources in the computing unit cluster and the model information of the first AI model. This facilitates providing tenants with at least one resource recommendation scheme that has a higher cost-performance ratio among the at least one candidate resource recommendation scheme.
[0012] In conjunction with the first aspect, in some implementations, at least two candidate resource recommendation schemes in the first candidate resource recommendation scheme set include different total numbers of computing units; or, each candidate resource recommendation scheme in the first candidate resource recommendation scheme set includes the same total number of computing units.
[0013] In this embodiment, the first candidate resource recommendation scheme set includes all candidate resource recommendation schemes determined by the cloud management platform, or the first candidate resource recommendation scheme set includes at least one candidate resource recommendation scheme with the same total number of computing units as all candidate resource recommendation schemes determined by the cloud management platform. In other words, the cloud management platform can provide tenants with at least one resource recommendation scheme with a high cost-performance ratio among all candidate resource recommendation schemes, or it can provide tenants with at least one resource recommendation scheme with a high cost-performance ratio for each total number of computing units according to the total number of computing units included in the candidate resource recommendation schemes.
[0014] In conjunction with the first aspect, in some implementations, when each candidate resource recommendation scheme in the first candidate resource recommendation scheme set includes the same total number of computing units, at least one second resource recommendation scheme and at least one first resource recommendation scheme are determined based on the available resources in at least one computing unit cluster and the model information of the first AI model; and at least one first resource recommendation scheme and at least one second resource recommendation scheme are provided to the tenant.
[0015] At least one second resource recommendation scheme belongs to the second candidate resource recommendation scheme set. The cost-effectiveness of at least one second resource recommendation scheme is better than that of the other candidate resource recommendation schemes in the second candidate resource recommendation scheme set. Each candidate resource recommendation scheme in the second candidate resource recommendation scheme set meets the operational requirements of the first AI model, and the resources included in each candidate resource recommendation scheme are available resources. Each candidate resource recommendation scheme in the second candidate resource recommendation scheme set includes the number of computing unit clusters used to run the first AI model and the number of computing units in each computing unit cluster. The total number of computing units included in each candidate resource recommendation scheme in the second candidate resource recommendation scheme set is the same. The total number of computing units included in the candidate resource recommendation schemes in the second candidate resource recommendation scheme set differs from that in the candidate resource recommendation schemes in the first candidate resource recommendation scheme set.
[0016] In this embodiment of the application, the cloud management platform can provide tenants with at least one resource recommendation scheme with a high cost-performance ratio under different total number of computing units, according to the total number of computing units included in the candidate resource recommendation scheme.
[0017] In conjunction with the first aspect, in some implementation methods, the parallel strategy scheme adopted by the first AI model is obtained, which includes the parallel parameters of at least one parallel strategy; based on the parallel strategy scheme, the available resources in at least one computing unit cluster, and the model information of the first AI model, at least one first resource recommendation scheme is determined.
[0018] In conjunction with the first aspect, in some implementations, obtaining the parallel strategy scheme adopted by the first AI model includes: receiving third instruction information from the tenant, which is used to indicate the parallel strategy scheme; or, determining at least one parallel strategy scheme based on the model information of the first AI model.
[0019] In conjunction with the first aspect, in some implementations, at least one parallel strategy includes at least one of the following: pipeline parallelism (PP), tensor parallelism (TP), data parallelism (DP), expert parallelism (EP), and sequence parallelism (SP).
[0020] In this embodiment of the application, the tenant may specify the parallel strategy scheme adopted for the operation of the first AI model, or the cloud management platform may determine the parallel strategy scheme adopted for the operation of the first AI model, and then determine at least one candidate resource recommendation scheme corresponding to each parallel strategy scheme according to the parallel strategy scheme.
[0021] In conjunction with the first aspect, in some implementations, the parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme in the first candidate resource recommendation scheme set are the same; or, the total number of computing units included in each candidate resource recommendation scheme in the first candidate resource recommendation scheme set is the same, and the parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme are the same. Wherein, the first parallel strategy belongs to at least one of the above-mentioned parallel strategies.
[0022] In conjunction with the first aspect, in some implementations, the first parallel strategy includes at least one of the following: data parallelism (DP) strategy and sequence parallelism (SP) strategy.
[0023] In this embodiment, the cloud management platform can provide tenants with at least one cost-effective resource recommendation scheme under each parallel parameter of the first parallel strategy, or can provide tenants with at least one cost-effective resource recommendation scheme under each parallel parameter of the first parallel strategy when the total number of computing units included is the same, thereby making it easier for tenants to choose a more satisfactory resource recommendation scheme and improve the tenant experience.
[0024] In conjunction with the first aspect, in some implementations, the model information of the first AI model includes the parameter scale and / or model structure of the first AI model.
[0025] In conjunction with the first aspect, in some implementations, the computing unit includes at least one of the following: a central processing unit (CPU), a graphics processing unit (GPU), or a neural processing unit (NPU).
[0026] In a second aspect, a computing device is provided. This device includes modules for implementing the first aspect or any possible implementation thereof.
[0027] Thirdly, a computing device cluster is provided, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method of the first aspect or any implementation thereof described above.
[0028] Fourthly, a computer program product containing instructions is provided, which, when executed by a cluster of computer devices, causes the cluster of computer devices to perform the method described in the first aspect or any one of the implementations of the first aspect.
[0029] Fifthly, a computer-readable storage medium is provided, including computer program instructions, which, when executed by a cluster of computing devices, enable the cluster of computing devices to perform the method described in the first aspect or any one of the implementations of the first aspect. Attached Figure Description
[0030] Figure 1 is a schematic structural block diagram of a resource recommendation system according to an embodiment of this application.
[0031] Figure 2 is a schematic flowchart of a resource recommendation method according to an embodiment of this application.
[0032] Figure 3 is a schematic diagram of an available resource and candidate resource recommendation scheme according to an embodiment of this application.
[0033] Figure 4 is a schematic diagram of a second graphical interface according to an embodiment of the present application.
[0034] Figure 5 is a schematic flowchart of a resource recommendation method according to another embodiment of this application.
[0035] Figure 6 is a schematic structural block diagram of a computing device according to an embodiment of the present application.
[0036] Figure 7 is a schematic structural diagram of a computing device according to an embodiment of the present application.
[0037] Figure 8 is a schematic structural diagram of a computing device cluster according to an embodiment of the present application.
[0038] Figure 9 is a schematic diagram of a network connection between computing devices 700A and 700B according to an embodiment of this application. Detailed Implementation
[0039] The technical solutions in this application will now be described with reference to the accompanying drawings.
[0040] This application will present various aspects, embodiments, or features relating to a system comprising multiple devices, components, modules, etc. It should be understood and appreciated that individual systems may include additional devices, components, modules, etc., and / or may not include all the devices, components, modules, etc. discussed in conjunction with the accompanying drawings. Furthermore, combinations of these approaches are also possible.
[0041] Furthermore, in the embodiments of this application, the words "exemplary," "for example," etc., are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" in the embodiments of this application should not be construed as being better or more advantageous than other embodiments or design schemes. Specifically, the use of the term "exemplary" is intended to present the concept in a concrete manner.
[0042] The business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0043] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0044] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0045] The method described in this application embodiment can be applied to various cloud management platforms. This cloud management platform manages the infrastructure providing cloud instances, including at least one computing unit cluster. Each computing unit cluster includes at least one computing unit, and each computing unit includes, for example, at least one of the following: a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), etc. The communication bandwidth between computing units within each computing unit cluster is high, while the communication bandwidth between different computing unit clusters is low. In other words, the communication bandwidth between computing units located within the same computing unit cluster is higher than the communication bandwidth between computing units located in different computing unit clusters. That is, when two computing units are connected via a high communication bandwidth, the two computing units belong to the same computing unit cluster. When two computing units are connected via a low communication bandwidth, the two computing units belong to different computing unit clusters.
[0046] Figure 1 is a schematic structural diagram of a resource recommendation system 100 provided in an embodiment of this application. The resource recommendation system 100 in Figure 1 includes a cloud management platform 110. The cloud management platform 110 is used to manage the infrastructure providing cloud services, which includes at least one data center (e.g., data center 120). Tenants can apply to use resources in data center 120 through the cloud management platform 110. The tenant is a public cloud tenant who has registered a public cloud account and purchased public cloud resources. Data center 120 includes at least one computing unit cluster (e.g., computing unit cluster 130 and / or computing unit cluster 140), and each computing unit cluster includes at least one computing unit. For example, computing unit cluster 130 includes computing units 131 and 132, and computing unit cluster 140 includes computing units 141 and 142.
[0047] For example, the communication bandwidth between computing units within each computing unit cluster is high, while the communication bandwidth between different computing unit clusters is low. In other words, the communication bandwidth between computing units located within the same computing unit cluster is higher than the communication bandwidth between computing units located in different computing unit clusters. That is, when two computing units are connected via a high communication bandwidth, the two computing units belong to the same computing unit cluster. When two computing units are connected via a low communication bandwidth, the two computing units belong to different computing unit clusters. For example, the communication bandwidth between computing unit 131 and computing unit 132 is higher than the communication bandwidth between computing unit 131 and computing unit 141 (or computing unit 142), and the communication bandwidth between computing unit 141 and computing unit 142 is higher than the communication bandwidth between computing unit 141 and computing unit 131 (or computing unit 132).
[0048] For example, in practical applications, computing units can come from physical servers, virtual machines (VMs), elastic cloud servers (ECS), or containers. Accordingly, the computing unit cluster shown in Figure 1 can be a cluster of computing devices consisting of at least one physical server, at least one VM, at least one ECS, or at least one container. For example, each computing unit cluster includes at least one server, and each of these at least one server includes at least one computing unit. At least two computing units in the same computing unit cluster are deployed on the same or different servers. For example, computing unit 131 and computing unit 132 are deployed on the same server or on different servers.
[0049] For example, the computing unit includes at least one of the following: CPU, GPU, NPU.
[0050] The cloud management platform 110 receives a first instruction from a tenant, which indicates the model information of the tenant's first AI model. This embodiment of the application does not limit the specific type or function of the first AI model.
[0051] For example, the model information of the first AI model includes the parameter scale and / or model structure of the first AI model. The parameter scale includes at least one of the following: the number of parameters used in the first AI model, the precision of the parameters, the amount of data of the parameters, etc. The model structure includes at least one of the following: the number of layers of the neural network included in the first AI model, the type of layers, the connection relationship or hierarchical relationship between multiple layers, etc.
[0052] The cloud management platform 110 determines at least one first resource recommendation scheme based on the first instruction information and the available resources in at least one computing unit cluster, and provides the at least one first resource recommendation scheme to the tenant. The at least one first resource recommendation scheme belongs to a first candidate resource recommendation scheme set, which includes at least one candidate resource recommendation scheme. Each candidate resource recommendation scheme in the first candidate resource recommendation scheme set meets the operational requirements of the first AI model, and the resources included in each candidate resource recommendation scheme are available resources. Each candidate resource recommendation scheme includes the number of computing unit clusters used to run the first AI model and the number of computing units in each computing unit cluster. The cost-effectiveness of the at least one first resource recommendation scheme is better than the cost-effectiveness of the other candidate resource recommendation schemes in the first candidate resource recommendation scheme set.
[0053] In some embodiments, the available resources in at least one computing unit cluster include: the available resources in each computing unit cluster of at least one computing unit cluster.
[0054] In some embodiments, running the first AI model includes: training the first AI model and / or using the first AI model for inference.
[0055] In some embodiments, the cost-effectiveness of each candidate resource recommendation is determined based on data metrics used to measure the performance of each candidate resource recommendation and the price of each candidate resource recommendation. A better cost-effectiveness for a candidate resource recommendation means that, at the same price, the candidate resource recommendation performs better. Alternatively, a better cost-effectiveness for a candidate resource recommendation means that, at the same performance, the candidate resource recommendation is cheaper.
[0056] Optionally, the cloud management platform 110 determines a first set of candidate resource recommendation schemes based on the available resources in at least one computing unit cluster and the model information of the first AI model. The cloud management platform determines the performance and price of each candidate resource recommendation scheme in the first set of candidate resource recommendation schemes. Based on the performance and price of each candidate resource recommendation scheme in the first set of candidate resource recommendation schemes, the cloud management platform determines at least one first resource recommendation scheme.
[0057] In some embodiments, at least two candidate resource recommendation schemes in the first candidate resource recommendation scheme set include different total numbers of computing units. Alternatively, each candidate resource recommendation scheme in the first candidate resource recommendation scheme set includes the same total number of computing units. In other words, the first candidate resource recommendation scheme set includes all candidate resource recommendation schemes determined by the cloud management platform 110, or the first candidate resource recommendation scheme set includes at least one candidate resource recommendation scheme whose total number of computing units is the same among all candidate resource recommendation schemes determined by the cloud management platform 110.
[0058] In some embodiments, when each candidate resource recommendation scheme in the first candidate resource recommendation scheme set includes the same total number of computing units, the cloud management platform 110 determines at least one first resource recommendation scheme and at least one second resource recommendation scheme based on the available resources in at least one computing unit cluster and the model information of the first AI model, and provides the at least one first resource recommendation scheme and at least one second resource recommendation scheme to the tenant. The at least one second resource recommendation scheme belongs to a second candidate resource recommendation scheme set, which includes at least one candidate resource recommendation scheme. The cost-effectiveness of the at least one second resource recommendation scheme is better than that of the other candidate resource recommendation schemes in the second candidate resource recommendation scheme set. The candidate resource recommendation schemes included in the second candidate resource recommendation scheme set are similar to those included in the first candidate resource recommendation scheme set; that is, each candidate resource recommendation scheme in the second candidate resource recommendation scheme set meets the operating requirements of the first AI model, and the resources included in each candidate resource recommendation scheme belong to the available resources in at least one computing unit cluster. Each candidate resource recommendation scheme in the second candidate resource recommendation scheme set includes the number of computing unit clusters used to run the first AI model and the number of computing units in each computing unit cluster. Each candidate resource recommendation scheme in the second candidate resource recommendation scheme set includes the same total number of computing units. The total number of computing units included in the candidate resource recommendation schemes in the second candidate resource recommendation scheme set differs from that in the first candidate resource recommendation scheme set.
[0059] Optionally, the cloud management platform 110 obtains the parallel strategy scheme adopted by the first AI model, which includes parallel parameters of at least one parallel strategy. Based on the parallel strategy scheme, the available resources in at least one computing unit cluster, and the model information of the first AI model, the cloud management platform 110 determines at least one first resource recommendation scheme.
[0060] In some embodiments, the cloud management platform receives third indication information from the tenant, which indicates a parallel strategy scheme. Alternatively, the cloud management platform determines at least one parallel strategy scheme based on model information from a first AI model.
[0061] In some embodiments, at least one parallel strategy includes at least one of the following: pipeline parallelism (PP), tensor parallelism (TP), data parallelism (DP), expert parallelism (EP), sequence parallelism (SP), etc.
[0062] For example, the parallel parameters of each parallel strategy in at least one parallel strategy are used to indicate the minimum number of computing units used when running an AI model through that parallel strategy.
[0063] For example, the minimum number of computing unit clusters included in the candidate resource recommendation scheme and / or the minimum number of computing units in each computing unit cluster are determined according to the parallel strategy scheme corresponding to the candidate resource recommendation scheme.
[0064] In some embodiments, the cloud management platform determines the minimum number of computing units required to meet the operational requirements of the first AI model based on the model information of the first AI model. The cloud management platform then determines at least one recommended resource scheme based on the minimum number of computing units required to meet the operational requirements of the first AI model, at least one parallel strategy scheme, and the available resources in at least one computing unit cluster. See Figure 2 for details.
[0065] In some embodiments, the parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme in the first candidate resource recommendation scheme set are the same. Alternatively, the total number of computing units included in each candidate resource recommendation scheme in the first candidate resource recommendation scheme set is the same, and the parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme are the same. Wherein, the first parallel strategy belongs to at least one of the above-mentioned parallel strategies.
[0066] For example, the first parallel strategy includes a DP strategy and / or a SP strategy.
[0067] For example, each resource recommendation scheme further includes: identification information of the computing unit cluster and / or identification information of the computing units in the computing unit cluster.
[0068] Optionally, after providing at least one first resource recommendation scheme to a tenant, the cloud management platform 110 can also receive second instruction information from the tenant. This second instruction information indicates a target resource recommendation scheme, which belongs to the at least one first resource recommendation scheme. The cloud management platform 110 can also provide the tenant with resources included in the target resource recommendation scheme based on the second instruction information. These resources are used to run the first AI model. In other words, the cloud management platform can also determine the target resource recommendation scheme that the tenant wants to use based on the tenant's selection, and then provide the tenant with the resources included in that target resource recommendation scheme, that is, allocate the computing units and computing unit clusters included in the target resource recommendation scheme to the tenant to run the first AI model.
[0069] The resource recommendation system 100 in Figure 1 can provide tenants with cost-effective resource recommendation solutions based on the model information of the tenant's AI model and the remaining available resources in the computing unit cluster. This allows tenants to purchase corresponding resources according to their needs, thereby improving the tenant's user experience and increasing the resource utilization rate in the computing unit cluster.
[0070] Figure 2 is a schematic flowchart of the resource recommendation method provided in an embodiment of this application. The method in Figure 2 can be executed by the cloud management platform in Figure 1. The method in Figure 2 includes the following steps.
[0071] 210, Receive the tenant's first instruction information.
[0072] The cloud management platform receives the tenant's first instruction information, which is used to indicate the model information of the tenant's first AI model.
[0073] In some embodiments, the model information of the first AI model includes the parameter scale and / or model structure of the first AI model. The parameter scale includes at least one of the following: the number of parameters used in the first AI model, the precision of the parameters, the amount of data of the parameters, etc. The model structure includes at least one of the following: the number of layers of the neural network included in the first AI model, the type of layers, the connection relationship or hierarchical relationship between multiple layers, etc.
[0074] Optionally, the cloud management platform provides a first graphical interface for the tenant, allowing the tenant to select, input, or upload model information of the first AI model on this first graphical interface. The specific form of this first graphical interface is not limited in this embodiment.
[0075] 220. Based on the available resources in at least one computing unit cluster and the model information of the first AI model, determine at least one first resource recommendation scheme.
[0076] The at least one computing unit cluster is managed by a cloud management platform. The available resources in the at least one computing unit cluster include the available resources in each of the at least one computing unit clusters. The available resources in each computing unit cluster are resources that are not allocated to other tenants, are available for sale, or are idle. The at least one first resource recommendation scheme belongs to a first candidate resource recommendation scheme set, which includes at least one candidate resource recommendation scheme. Each candidate resource recommendation scheme in the first candidate resource recommendation scheme set meets the operational requirements of the first AI model, and the resources included in each candidate resource recommendation scheme belong to the available resources in at least one computing unit cluster. Each candidate resource recommendation scheme includes the number of computing unit clusters used to run the first AI model and the number of computing units in each computing unit cluster. The cost-effectiveness of the at least one first resource recommendation scheme is better than the cost-effectiveness of the candidate resource recommendation schemes in the first candidate resource recommendation scheme set other than the at least one first resource recommendation scheme.
[0077] Optionally, prior to step 220, the cloud management platform determines the available resources in each computing unit cluster.
[0078] Optionally, the cloud management platform determines the first set of candidate resource recommendations based on the available resources in at least one computing unit cluster and the model information of the first AI model. The cloud management platform determines the performance and price of each candidate resource recommendation in the first set of candidate resource recommendations. Based on the performance and price of each candidate resource recommendation in the first set of candidate resource recommendations, the cloud management platform determines at least one first resource recommendation.
[0079] In some embodiments, at least two candidate resource recommendation schemes in the first candidate resource recommendation scheme set include different total numbers of computing units. Alternatively, each candidate resource recommendation scheme in the first candidate resource recommendation scheme set includes the same total number of computing units. In other words, the first candidate resource recommendation scheme set includes all candidate resource recommendation schemes determined by the cloud management platform, or the first candidate resource recommendation scheme set includes at least one candidate resource recommendation scheme among all candidate resource recommendation schemes determined by the cloud management platform that includes the same total number of computing units.
[0080] In some embodiments, the cloud management platform determines the minimum number of computing units required to meet the operational requirements of the first AI model based on the model information of the first AI model. The cloud management platform then determines at least one candidate resource recommendation scheme based on the minimum number of computing units required to meet the operational requirements of the first AI model and the available resources in at least one computing unit cluster, thereby determining the first candidate resource recommendation scheme set. Each candidate resource recommendation scheme satisfies the operational requirements of the first AI model, meaning that the total number of computing units included in each candidate resource recommendation scheme is greater than or equal to the minimum number of computing units required to meet the operational requirements of the first AI model; for example, the total number of computing units included in each candidate resource recommendation scheme is an integer multiple of the minimum number of computing units required to meet the operational requirements of the first AI model. The resources included in each candidate resource recommendation scheme belong to the available resources in the at least one computing unit cluster, meaning that the computing units included in each candidate resource recommendation scheme are available resources in the computing unit cluster to which the computing unit belongs.
[0081] For example, the cloud management platform determines the minimum number of computing units required to meet the operational needs of the first AI model based on the model information and expert experience. Alternatively, the cloud management platform determines the required storage space capacity for running the first AI model based on its model information, and then determines the minimum number of computing units required to meet the operational needs of the first AI model based on the required storage space capacity and the capacity of the storage unit corresponding to each computing unit. The required storage space capacity for running the first AI model includes a first storage space capacity and a second storage space capacity. The first storage space capacity is determined based on the number and precision of the parameters in the first AI model; for example, the first storage space capacity is the product of the number and precision of the parameters in the first AI model. The second storage space capacity includes the storage capacity required to store information such as gradients, activation parameters, and optimizer states when running the first AI model. The total amount of data related to gradients, activation parameters, and optimizer states when running the first AI model is proportional to the first storage space capacity. For example, the second storage space capacity = a × the first storage space capacity, where a is a positive number. The specific value of 'a' is not limited in this application embodiment; for example, the specific value of 'a' can be determined based on historical experience. The minimum number of computing units required to meet the running requirements of the first AI model is greater than or equal to the ratio of the storage space capacity required to run the first AI model to the capacity of the storage unit corresponding to the computing unit. The storage unit corresponding to the computing unit is used to store the data required or generated by the computing unit during computation. When the computing unit is a CPU, the storage unit corresponding to the computing unit is, for example, memory. When the computing unit is a GPU, the storage unit corresponding to the computing unit is, for example, video memory. The type of storage unit corresponding to the computing unit is not limited in this application embodiment; for example, it includes volatile storage media and / or non-volatile storage media.
[0082] In some embodiments, those skilled in the art can construct an evaluation model based on the computational logic and / or expert experience that determines the minimum number of computing units required to meet the operational requirements of the first AI model according to the model information of the first AI model in the above embodiments. The cloud management platform inputs the model information of the first AI model into the evaluation model, and the model outputs the minimum number of computing units required to meet the operational requirements of the first AI model.
[0083] In some embodiments, running the first AI model includes: training the first AI model and / or using the first AI model for inference.
[0084] In some embodiments, the performance of the candidate resource recommendation scheme is measured according to at least one of the following data metrics: the training time required to train the first AI model using the resources included in the candidate resource recommendation scheme, the throughput required to train the first AI model using the resources included in the candidate resource recommendation scheme, and the inference latency required to perform inference using the first AI model with the resources included in the candidate resource recommendation scheme. The throughput may include, for example, the number of samples trained per second or the amount of data per sample when training the first AI model using the resources included in the candidate resource recommendation scheme. This application embodiment does not limit the units for training time, inference latency, or throughput; for example, the units for training time or inference latency may be seconds, minutes, hours, etc., and the units for throughput may be units, bits, bytes, etc.
[0085] For example, the cloud management platform receives a fourth instruction from the tenant, which instructs that the resources included in the resource recommendation scheme be used to train the first AI model and / or to perform inference using the first AI model.
[0086] For example, when the fourth indication information is used to indicate that the resources included in the resource recommendation scheme are used to train the first AI model, the performance of the candidate resource recommendation scheme is measured according to at least one of the following data metrics: the training time required to train the first AI model using the resources included in the candidate resource recommendation scheme, and the throughput when training the first AI model using the resources included in the candidate resource recommendation scheme. When the fourth indication information is used to indicate that the resources included in the resource recommendation scheme are used for inference using the first AI model, the performance of the candidate resource recommendation scheme is measured according to the following data metric: the inference latency required to perform inference using the first AI model using the resources included in the candidate resource recommendation scheme.
[0087] In some embodiments, the cloud management platform determines the performance of each candidate resource recommendation scheme based on a performance estimation model. Specifically, the cloud management platform inputs the candidate resource recommendation scheme into the performance estimation model to obtain the output of the performance estimation model, i.e., the performance of the candidate resource recommendation scheme. This application embodiment does not limit the specific type of the performance estimation model; any existing model capable of determining the performance of a candidate resource recommendation scheme can be applied to the method provided in this application embodiment.
[0088] For example, the performance estimation model includes a first time estimation model and a second time estimation model. The first time estimation model is used to estimate the computation time required to execute all computational tasks in the first AI model when running the first AI model using the computing units and computing unit clusters included in the candidate resource recommendation scheme. For example, the first time estimation model determines the computation time required to execute all computational tasks in the first AI model based on the formula: computation time = computational load / computing speed. The computational load is determined based on the parameter scale of the first AI model; for example, the computational load is the sum of the number of operations for each parameter of the first AI model. The computing speed is determined based on the total number of computing units included in the candidate resource recommendation scheme and the computing power of each computing unit; for example, the computing speed is the sum of the computing power of each computing unit in the candidate resource recommendation scheme, and the computing power of a computing unit is measured, for example, by floating point operations per second (FLOPS). The second time estimation model is used to estimate the communication time required to transmit the data required and / or the data generated during the operation of the first AI model when running the first AI model using the computing units and computing unit clusters included in the candidate resource recommendation scheme. For example, the second time estimation model determines the communication time required to transmit the data and / or generated data during the operation of the first AI model, based on the formula: Communication Duration = Communication Volume / Communication Bandwidth. The communication volume is, for example, the sum of the data volumes of the datasets that need to be transmitted each time the first AI model runs. The data volume and the number of datasets to be transmitted are determined based on the parallelism strategy and parameter scale of the first AI model. The communication bandwidth is, for example, the communication bandwidth between computing units used to transmit the corresponding datasets.
[0089] For example, the performance estimation model is the Calculon model from the paper "Calculon: a Methodology and Tool for High-Level Codesign of Systems and Large Language Models". The Calculon model estimates the performance of the hardware used to run the first AI model based on the model information of the first AI model input by the user, information about the hardware used to run the first AI model (e.g., the number of computing units, the type of computing units, and the storage capacity of the corresponding storage units), and the parallel strategy of the first AI model.
[0090] In some embodiments, the price of the candidate resource recommendation scheme is, for example, the price required to use the resources in the candidate resource recommendation scheme per unit time. This application embodiment does not limit the length of this unit time, such as 10 minutes, 1 hour, 1 day, etc.
[0091] In some embodiments, the cost-effectiveness of a candidate resource recommendation is determined based on data metrics used to measure the performance of the candidate resource recommendation and its price. A better cost-effectiveness means that, at the same price, the candidate resource recommendation performs better. Alternatively, a better cost-effectiveness means that, at the same performance, the candidate resource recommendation is cheaper.
[0092] For example, the cost-effectiveness of the candidate resource recommendation scheme is measured based on at least one set of the following data metrics: training duration and price per unit time, throughput and price per unit time, and inference latency and price per unit time.
[0093] For example, cost-effectiveness is determined according to at least one of the following formulas: Cost-effectiveness = 1 / (training duration × price per unit time); Cost-effectiveness = training duration × price per unit time; Cost-effectiveness = 1 / (inference latency × price per unit time); Cost-effectiveness = inference latency × price per unit time; Cost-effectiveness = throughput / price per unit time; Cost-effectiveness = price per unit time / throughput, etc.
[0094] For example, when cost-effectiveness = 1 / (training duration × price per unit time), cost-effectiveness = 1 / (inference latency × price per unit time), or cost-effectiveness = throughput / price per unit time, a higher cost-effectiveness value indicates better cost-effectiveness, and a lower cost-effectiveness value indicates worse cost-effectiveness. Similarly, when cost-effectiveness = training duration × price per unit time, cost-effectiveness = inference latency × price per unit time, or cost-effectiveness = price per unit time / throughput, a lower cost-effectiveness value indicates better cost-effectiveness, and a higher cost-effectiveness value indicates worse cost-effectiveness.
[0095] Optionally, the cloud management platform obtains the parallel strategy scheme adopted by the first AI model, which includes parallel parameters of at least one parallel strategy. Based on the parallel strategy scheme, the available resources in at least one computing unit cluster, and the model information of the first AI model, the cloud management platform determines at least one first resource recommendation scheme.
[0096] In some embodiments, the cloud management platform determines a first set of candidate resource recommendation schemes based on the parallel strategy scheme, the available resources in at least one computing unit cluster, and the model information of the first AI model. The cloud management platform then determines the at least one first resource recommendation scheme based on the cost-effectiveness of each candidate resource recommendation scheme in the first set of candidate resource recommendation schemes.
[0097] In some embodiments, before step 220, the cloud management platform receives third indication information from the tenant, which is used to indicate the parallel strategy scheme. Alternatively, before step 220, the cloud management platform determines at least one parallel strategy scheme based on the model information of the first AI model.
[0098] In some embodiments, at least one parallel strategy includes at least one of the following: PP strategy, TP strategy, DP strategy, EP strategy, SP strategy, etc.
[0099] In some embodiments, the parallelism parameter of the parallel strategy is used to indicate the minimum number of computing units used when running an AI model through the parallel strategy.
[0100] In some embodiments, the cloud management platform determines the minimum number of computing units required to meet the operational requirements of the first AI model based on the model information of the first AI model. The cloud management platform then determines at least one candidate resource recommendation scheme based on the minimum number of computing units required to meet the operational requirements of the first AI model, at least one parallel strategy scheme, and available resources in at least one computing unit cluster, thereby determining a first set of candidate resource recommendation schemes. Each candidate resource recommendation scheme satisfies the operational requirements of the first AI model, meaning that the total number of computing units included in each candidate resource recommendation scheme is greater than or equal to the minimum number of computing units required to meet the operational requirements of the first AI model; for example, the total number of computing units included in each candidate resource recommendation scheme is an integer multiple of the minimum number of computing units required to meet the operational requirements of the first AI model. The resources included in each candidate resource recommendation scheme belong to the available resources in the at least one computing unit cluster, meaning that the computing units included in each candidate resource recommendation scheme are available resources in the computing unit cluster to which the computing unit belongs. Each candidate resource recommendation scheme satisfies the parallel requirements of the corresponding parallel strategy scheme, that is, the number of computing unit clusters included in each candidate resource recommendation scheme and / or the number of computing units included in each computing unit cluster are determined according to the parallel strategy scheme.
[0101] In some embodiments, when some or all tasks in the first AI model are run in parallel using at least one of the SP, EP, and TP strategies, the computing units running those tasks belong to the same computing unit cluster. In other words, the parallel parameters of at least one of the SP, EP, and TP strategies are used to determine the minimum number of computing units included in each computing unit cluster in the candidate resource recommendation scheme.
[0102] In some embodiments, when some or all tasks in the first AI model are run in parallel using PP and / or DP strategies, the computing units running those tasks belong to the same computing unit cluster or different computing unit clusters. In other words, the parallel parameters of at least one of the PP and / or DP strategies are used to determine the minimum number of computing unit clusters in the candidate resource recommendation scheme.
[0103] For example, when the parallel strategy schemes corresponding to the two candidate resource recommendation schemes are the same, the number of computing unit clusters included in the two candidate resource recommendation schemes and / or the number of computing units in each computing unit cluster are different.
[0104] For example, the total number of computational units included in the candidate resource recommendation scheme is greater than or equal to the product of each parallel parameter included in the corresponding parallel strategy scheme. Alternatively, the total number of computational units included in the candidate resource recommendation scheme is an integer multiple of the product of each parallel parameter included in the corresponding parallel strategy scheme.
[0105] For example, suppose the parallel strategy scheme 1 corresponding to candidate resource recommendation scheme 1 includes: the parallel parameter of SP is 2, the parallel parameter of EP is 2, and the parallel parameter of TP is 2. Then the total number of computing units included in candidate resource recommendation scheme 1 is greater than or equal to the product of the parallel parameters of SP, EP, and TP (i.e., 2 × 2 × 2 = 8). Alternatively, the total number of computing units included in candidate resource recommendation scheme 1 is, for example, an integer multiple of the product of the parallel parameters of SP, EP, and TP, such as 8 × 2 = 16.
[0106] For example, suppose the available resources in at least one computing unit cluster are as shown in Figure 3. Figure 3 includes six computing unit clusters, each containing 32 computing units. These six computing unit clusters are: Computing Unit Cluster 310, Computing Unit Cluster 320, Computing Unit Cluster 330, Computing Unit Cluster 340, Computing Unit Cluster 350, and Computing Unit Cluster 360. In these six computing unit clusters, black squares represent computing units allocated to tenant A, vertically filled squares represent computing units allocated to tenant B, and white squares represent computing units not allocated to any tenant. As can be seen from Figure 3, in Computing Unit Cluster 310, 16 computing units have been allocated to tenant A, 8 computing units have been allocated to tenant B, and the remaining 8 computing units are available resources. Similarly, in Computing Unit Cluster 320, 16 computing units have been allocated to tenant A, 8 computing units have been allocated to tenant B, and the remaining 8 computing units are available resources. In Computing Unit Cluster 330, 16 computing units have been allocated to tenant A, 8 computing units have been allocated to tenant B, and the remaining 8 computing units are available resources. Eight computing units in computing unit cluster 340 have been allocated to tenant B, leaving 24 computing units available. Computing units in computing unit clusters 350 and 360 have not been allocated and are therefore available.
[0107] Assume that, based on the model information of the first AI model, the minimum number of computing units required to meet the operational requirements of the first AI model is determined to be 32. Assume that at least one parallel strategy scheme includes parallel strategy scheme A, which includes: SP with a parallel parameter of 1, EP with a parallel parameter of 1, TP with a parallel parameter of 8, and PP with a parallel parameter of 4. Based on parallel strategy scheme A, the cloud management platform determines one or more candidate resource recommendation schemes corresponding to parallel strategy scheme A. The total number of computing units included in these one or more candidate resource recommendation schemes is greater than or equal to the product of SP, EP, TP, and PP, i.e., 1 × 1 × 8 × 4 = 32. That is, the total number of computing units included in these one or more candidate resource recommendation schemes is greater than or equal to the minimum number of computing units required to meet the operational requirements of the first AI model. Based on the minimum number of computing units required to meet the operational requirements of the first AI model, parallel strategy scheme A, and the available resources in at least one computing unit cluster, the cloud management platform determines candidate resource recommendation scheme A, candidate resource recommendation scheme B, and candidate resource recommendation scheme C as shown in Figure 3. Candidate resource recommendation scheme A includes four computing unit clusters, with each cluster containing at least eight usable computing units. The four computing unit clusters in candidate resource recommendation scheme A are: Computing Unit Cluster 310, Computing Unit Cluster 320, Computing Unit Cluster 330, and Computing Unit Cluster 340. Candidate resource recommendation scheme B includes two computing unit clusters, with each cluster containing at least 16 usable computing units. The two computing unit clusters in candidate resource recommendation scheme B are: Computing Unit Cluster 340 and Computing Unit Cluster 350. Candidate resource recommendation scheme C includes one computing unit cluster, with each cluster containing at least 32 usable computing units. The computing unit cluster included in candidate resource recommendation scheme C is Computing Unit Cluster 350.
[0108] As shown in Figure 3, candidate resource recommendation scheme A results in fewer remaining fragmented resources in the 6-computing-unit cluster, thus its price is lower, and it can meet the tenant's needs for running the first AI model. Compared to candidate resource recommendation scheme A, although candidate resource recommendation scheme B can also meet the tenant's needs for running the first AI model, and its efficiency is higher than that of running the first AI model using candidate resource recommendation scheme A, candidate resource recommendation scheme B results in more remaining fragmented resources in the 6-computing-unit cluster, therefore its price is higher. Compared to candidate resource recommendation scheme B, although the efficiency of running the first AI model using candidate resource recommendation scheme C is higher than that of running the first AI model using candidate resource recommendation scheme B, candidate resource recommendation scheme C results in more remaining fragmented resources in the 6-computing-unit cluster, therefore its price is higher.
[0109] In some embodiments, the parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme in the first candidate resource recommendation scheme set are the same. Alternatively, the total number of computing units included in each candidate resource recommendation scheme in the first candidate resource recommendation scheme set is the same, and the parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme are the same. Wherein, the first parallel strategy belongs to at least one of the above-mentioned parallel strategies.
[0110] In some embodiments, the first parallel strategy includes a DP strategy and / or a SP strategy.
[0111] Optionally, if each candidate resource recommendation scheme in the first candidate resource recommendation scheme set includes the same total number of computing units, the cloud management platform determines at least one first resource recommendation scheme and at least one second resource recommendation scheme based on the available resources in at least one computing unit cluster and the model information of the first AI model. The at least one second resource recommendation scheme belongs to a second candidate resource recommendation scheme set, which includes at least one candidate resource recommendation scheme. The cost-effectiveness of the at least one second resource recommendation scheme is better than that of the other candidate resource recommendation schemes in the second candidate resource recommendation scheme set. The candidate resource recommendation schemes included in the second candidate resource recommendation scheme set are similar to those included in the first candidate resource recommendation scheme set; that is, each candidate resource recommendation scheme in the second candidate resource recommendation scheme set meets the operational requirements of the first AI model, and the resources included in each candidate resource recommendation scheme belong to the available resources in at least one computing unit cluster. Each candidate resource recommendation scheme in the second candidate resource recommendation scheme set includes the number of computing unit clusters used to run the first AI model and the number of computing units in each computing unit cluster. The total number of computing units included in each candidate resource recommendation scheme in the second candidate resource recommendation scheme set is the same. The number of computing units included in the candidate resource recommendation schemes in the second candidate resource recommendation scheme set differs from the total number of computing units included in the candidate resource recommendation schemes in the first candidate resource recommendation scheme set. In other words, the cloud management platform can determine the k most cost-effective candidate resource recommendation schemes in each of the multiple candidate resource recommendation scheme sets, where k is a positive integer. Each candidate resource recommendation scheme set includes at least one candidate resource recommendation scheme, and the total number of computing units included in the candidate resource recommendation schemes in each candidate resource recommendation scheme set is the same.
[0112] Optionally, when the parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme in the first candidate resource recommendation scheme set are the same, the cloud management platform determines at least one first resource recommendation scheme and at least one third resource recommendation scheme based on the available resources in at least one computing unit cluster and the model information of the first AI model. The at least one third resource recommendation scheme belongs to a third candidate resource recommendation scheme set, which includes at least one candidate resource recommendation scheme. The cost-effectiveness of the at least one third resource recommendation scheme is better than that of the other candidate resource recommendation schemes in the third candidate resource recommendation scheme set. The candidate resource recommendation schemes included in the third candidate resource recommendation scheme set are similar to those included in the first candidate resource recommendation scheme set; that is, each candidate resource recommendation scheme in the third candidate resource recommendation scheme set meets the operational requirements of the first AI model, and the resources included in each candidate resource recommendation scheme belong to the available resources in at least one computing unit cluster. Each candidate resource recommendation scheme in the third candidate resource recommendation scheme set includes the number of computing unit clusters used to run the first AI model and the number of computing units in each computing unit cluster. The parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme in the third candidate resource recommendation scheme set are the same. The parallel parameters of the first parallel strategy corresponding to the candidate resource recommendation schemes in the third candidate resource recommendation scheme set are different from those in the first candidate resource recommendation scheme set. In other words, the cloud management platform can determine the k most cost-effective candidate resource recommendation schemes in each of the multiple candidate resource recommendation scheme sets, where k is a positive integer. Each candidate resource recommendation scheme set includes at least one candidate resource recommendation scheme, and the parallel parameters of the first parallel strategy corresponding to the candidate resource recommendation schemes in each candidate resource recommendation scheme set are the same.
[0113] Optionally, if each candidate resource recommendation scheme in the first candidate resource recommendation scheme set includes the same total number of computing units and the parallel parameters of the first parallel strategy corresponding to each resource recommendation scheme are the same, the cloud management platform determines at least one first resource recommendation scheme and at least one fourth resource recommendation scheme based on the available resources in at least one computing unit cluster and the model information of the first AI model. The at least one fourth resource recommendation scheme belongs to a fourth candidate resource recommendation scheme set, which includes at least one candidate resource recommendation scheme. The cost-effectiveness of the at least one fourth resource recommendation scheme is better than that of the other candidate resource recommendation schemes in the fourth candidate resource recommendation scheme set. The candidate resource recommendation schemes included in the fourth candidate resource recommendation scheme set are similar to those included in the first candidate resource recommendation scheme set; that is, each candidate resource recommendation scheme in the fourth candidate resource recommendation scheme set meets the operational requirements of the first AI model, and the resources included in each candidate resource recommendation scheme belong to the available resources in at least one computing unit cluster. Each candidate resource recommendation scheme in the fourth candidate resource recommendation scheme set includes the number of computing unit clusters used to run the first AI model and the number of computing units in each computing unit cluster. Each candidate resource recommendation scheme in the fourth candidate resource recommendation scheme set includes the same total number of computing units, and the parallel parameters of the first parallel strategy corresponding to each resource recommendation scheme are the same. The candidate resource recommendation schemes in the fourth candidate resource recommendation scheme set differ from those in the first candidate resource recommendation scheme set in terms of the total number of computing units, and / or, the candidate resource recommendation schemes in the fourth candidate resource recommendation scheme set differ from those in the first candidate resource recommendation scheme set in terms of the parallel parameters of the first parallel strategy. In other words, the cloud management platform can determine the k most cost-effective candidate resource recommendation schemes in each of the multiple candidate resource recommendation scheme sets, where k is a positive integer. Each candidate resource recommendation scheme set includes at least one candidate resource recommendation scheme, and the total number of computing units included in the candidate resource recommendation schemes in each candidate resource recommendation scheme set is the same, and the parallel parameters of the first parallel strategy corresponding to each resource recommendation scheme are the same.
[0114] In some embodiments, each candidate resource recommendation scheme further includes: identification information of the computing unit cluster and / or identification information of the computing unit.
[0115] 230, provide tenants with at least one primary resource recommendation option.
[0116] Optionally, the cloud management platform generates a second graphical interface and provides it to the tenant. The cloud management platform provides the tenant with the at least one first resource recommendation scheme within this second graphical interface; that is, the tenant can view the at least one first resource recommendation scheme through the second graphical interface. The specific form of the second graphical interface is not limited in this embodiment.
[0117] For example, the second graphical interface is shown in Figure 4. Figure 4 is a possible schematic diagram of the second graphical interface 400. As shown in Figure 4, the second graphical interface 400 provides tenants with 12 resource recommendation schemes, and provides tenants with the unit time price and inference latency for each resource recommendation scheme. The unit time price is the price required to use the resource recommendation scheme for 1 hour. The inference latency is in seconds. The resource recommendation schemes in the second graphical interface 400 are recommended for an AI model with a parameter scale of 160 billion (B), and the minimum number of GPUs required to meet the running requirements of the AI model is 32. The minimum number of GPUs required to meet the running requirements of the AI model is determined by the cloud management platform based on the parameter scale of the AI model, or is information input by the tenant. For example, the cloud management platform receives a fifth instruction from the tenant, which indicates the minimum number of computing units required to meet the running requirements of the AI model. The resource recommendation scheme 4*8 in Figure 4 indicates that the resource recommendation scheme includes 8 computing unit clusters, and each computing unit cluster includes 4 available computing units. Similarly, a resource recommendation scheme of 8*4 means that the resource recommendation scheme includes 4 computing unit clusters, and each computing unit cluster includes 8 available computing units. And so on.
[0118] For example, the second graphical interface may also include the cost-effectiveness of each resource recommendation scheme. For instance, the cost-effectiveness of the 12 resource recommendation schemes may be presented in the second graphical interface using tables, line graphs, curves, or other similar formats.
[0119] Optionally, the cloud management platform provides the tenant with at least one first resource recommendation scheme and at least one second resource recommendation scheme. The second resource recommendation scheme is described in step 220. Alternatively, the cloud management platform provides the tenant with at least one first resource recommendation scheme and at least one third resource recommendation scheme. The third resource recommendation scheme is described in step 220. Alternatively, the cloud management platform provides the tenant with at least one first resource recommendation scheme and at least one fourth resource recommendation scheme. The fourth resource recommendation scheme is described in step 220.
[0120] For example, the cloud management platform provides the tenant with at least one first resource recommendation scheme and at least one second resource recommendation scheme in the second graphical interface. Alternatively, the cloud management platform provides the tenant with at least one first resource recommendation scheme and at least one third resource recommendation scheme in the second graphical interface. Alternatively, the cloud management platform provides the tenant with at least one first resource recommendation scheme and at least one fourth resource recommendation scheme in the second graphical interface.
[0121] Optionally, when the cloud management platform provides a tenant with at least one first resource recommendation scheme, the cloud management platform also receives second instruction information from the tenant. This second instruction information is used to indicate a target resource recommendation scheme, which belongs to the at least one first resource recommendation scheme.
[0122] For example, after the cloud management platform provides at least one first resource recommendation scheme to the tenant in the second graphical interface, it can also receive the second instruction information based on the tenant's selection or input operation in the second graphical interface. The selection operation is used to select the target resource recommendation scheme from the at least one first resource recommendation scheme. The input operation is used to input the target resource recommendation scheme or the identification information of the target resource recommendation scheme.
[0123] Optionally, when the cloud management platform provides a tenant with at least one first resource recommendation scheme and at least one second resource recommendation scheme, the cloud management platform also receives second instruction information from the tenant. This second instruction information is used to indicate a target resource recommendation scheme, which belongs to at least one first resource recommendation scheme and at least one second resource recommendation scheme. Alternatively, when the cloud management platform provides a tenant with at least one first resource recommendation scheme and at least one third resource recommendation scheme, the cloud management platform also receives second instruction information from the tenant. This second instruction information is used to indicate a target resource recommendation scheme, which belongs to at least one first resource recommendation scheme and at least one third resource recommendation scheme. Alternatively, when the cloud management platform provides a tenant with at least one first resource recommendation scheme and at least one fourth resource recommendation scheme, the cloud management platform also receives second instruction information from the tenant. This second instruction information is used to indicate a target resource recommendation scheme, which belongs to at least one first resource recommendation scheme and at least one fourth resource recommendation scheme.
[0124] For example, after the cloud management platform provides at least one first resource recommendation scheme and at least one second resource recommendation scheme to the tenant in the second graphical interface, it can also receive the second instruction information based on the tenant's selection or input operation in the second graphical interface. The selection operation is used to select the target resource recommendation scheme from the at least one first resource recommendation scheme and at least one second resource recommendation scheme. The input operation is used to input the target resource recommendation scheme or the identification information of the target resource recommendation scheme. Similarly, when the cloud management platform provides at least one first resource recommendation scheme and at least one third resource recommendation scheme to the tenant in the second graphical interface, or when it provides at least one first resource recommendation scheme and at least one fourth resource recommendation scheme to the tenant in the second graphical interface, the way the cloud management platform receives the second instruction information is similar to the above method, and will not be repeated here.
[0125] Optionally, after receiving the second instruction information, the cloud management platform provides the tenant with the resources included in the target resource recommendation scheme, based on the second instruction information. These resources are used to run the first AI model. In other words, after receiving the second instruction information, the cloud management platform allocates the computing nodes and computing node clusters included in the target resource recommendation scheme to the tenant, enabling the tenant to run the first AI model using these computing nodes and computing node clusters.
[0126] The method in Figure 2 can provide tenants with at least one cost-effective first resource recommendation scheme based on the model information of the AI model to be run by the tenant and the remaining available resources in the computing unit cluster, so that the tenant can purchase suitable resources and improve the resource utilization in the computing unit cluster.
[0127] In some embodiments, the method by which the cloud management platform determines at least one first resource recommendation scheme from at least one candidate resource recommendation scheme is shown in Figure 5.
[0128] Figure 5 is a schematic flowchart of the resource recommendation method provided in an embodiment of this application. The method in Figure 5 can be executed by the cloud management platform in Figure 1. The method in Figure 5 includes the following steps.
[0129] 510. Based on the available resources in at least one computing unit cluster and the model information of the first AI model, determine N first sets.
[0130] Optionally, the cloud management platform determines at least one candidate resource recommendation scheme based on the available resources in at least one computing unit cluster and the model information of the first AI model. Alternatively, the cloud management platform determines at least one candidate resource recommendation scheme based on at least one parallel strategy scheme, the available resources in at least one computing unit cluster, and the model information of the first AI model. Each of these at least one candidate resource recommendation schemes satisfies the operational requirements of the first AI model, and the resources included in each candidate resource recommendation scheme belong to the available resources in at least one computing unit cluster. Each candidate resource recommendation scheme includes the number of computing unit clusters used to run the first AI model and the number of computing units in each computing unit cluster. The available resources, the model information of the first AI model, the candidate resource recommendation schemes, and the parallel strategy schemes are described in Figure 2.
[0131] Optionally, after determining at least one candidate resource recommendation scheme, the cloud management platform determines N first sets, where N is a positive integer. Each of the N first sets includes one or more candidate resource recommendation schemes from the at least one candidate resource recommendation scheme.
[0132] In some embodiments, each candidate resource recommendation scheme in the first set includes the same total number of computing units. Alternatively, each candidate resource recommendation scheme in the first set has the same parallel parameters for its corresponding first parallel strategy. Or, each candidate resource recommendation scheme in the first set includes the same total number of computing units, and each candidate resource recommendation scheme has the same parallel parameters for its corresponding first parallel strategy. The first parallel strategy is at least one type of parallel strategy. The at least one parallel strategy and the first parallel strategy are described in step 220.
[0133] In some embodiments, each of the N first sets is sorted according to a preset order. For example, if the total number of computing units included in each candidate resource recommendation scheme in the first set is the same, the N first sets are sorted in ascending order of the total number of computing units included in each candidate resource recommendation scheme in the first set. Alternatively, if the parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme in the first set are the same, the N first sets are sorted in ascending order of the parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme in the first set. Alternatively, if the total number of computing units included in each candidate resource recommendation scheme in the first set is the same, and the parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme are the same, each of the N sets is sorted in ascending order of the total number of computing units included in each candidate resource recommendation scheme in the first set, and if the total number of computing units included in each candidate resource recommendation scheme in the first set is the same, the sets are sorted in ascending order of the parallel parameters of the first parallel strategy corresponding to the candidate resource recommendation scheme in the first set. Alternatively, if each candidate resource recommendation scheme in the first set includes the same total number of computing units and the parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme are the same, then each of the N sets is sorted according to the increasing order of the parallel parameters of the first parallel strategy corresponding to the candidate resource recommendation scheme in each first set. Furthermore, if the parallel parameters of the first parallel strategy corresponding to the candidate resource recommendation scheme in the first set are the same, then they are sorted according to the increasing order of the total number of computing units included in the candidate resource recommendation scheme in the first set.
[0134] 520. Based on the cost-effectiveness of the most cost-effective candidate resource recommendation scheme in each of the N first sets, determine the cost-effectiveness gain rate corresponding to the most cost-effective candidate resource recommendation scheme in each first set.
[0135] Optionally, after the cloud management platform sorts the N first sets according to a preset order, it determines the cost-performance gain rate corresponding to the cost-performance of the most cost-effective candidate resource recommendation scheme in each of the N first sets. The cost-performance gain rate corresponding to the most cost-effective candidate resource recommendation scheme in each first set is used to characterize the cost-performance gain rate between the most cost-effective candidate resource recommendation scheme in that first set and the most cost-effective candidate resource recommendation scheme in the previous first set, where the previous first set is the set preceding the first set when arranged in a preset order. The cost-performance of the candidate resource recommendation schemes is described in step 220.
[0136] In some embodiments, the cost-performance gain rate S of the most cost-effective candidate resource recommendation scheme in the (n+1)th first set of the N first sets is... n+1 The calculation formula is as follows (1): S n+1 =(Y n+1 -Y n ) / Y n Equation (1)
[0137] In equation (1), Y n+1 Y represents the cost-effectiveness of the candidate resource recommendation scheme in the (n+1)th first set among the N first sets. n This represents the cost-effectiveness of the recommended candidate resource in the nth first set of N first sets, where n = 1, ..., N-1. These N first sets are sorted according to a preset order.
[0138] For example, the cost-performance gain rate S1 = 0 for the candidate resource recommendation scheme with the highest cost-performance ratio in the first set of the N first sets.
[0139] 530. Based on the cost-performance gain rate corresponding to the candidate resource recommendation scheme with the highest cost-performance ratio in each first set, determine at least one first resource recommendation scheme.
[0140] The cloud management platform determines whether the cost-effectiveness gain rate has converged based on the cost-effectiveness gain rate corresponding to the most cost-effective candidate resource recommendation scheme in each first set. If the cost-effectiveness gain rate has converged, a fifth resource recommendation scheme is determined. This fifth resource recommendation scheme is the most cost-effective candidate resource recommendation scheme among all candidate resource recommendation schemes determined by the cloud management platform.
[0141] For example, the convergence of the cost-performance gain rate includes: the cost-performance gain rate being less than or equal to a first preset threshold. The specific value of the first preset threshold is not limited in this application embodiment; for example, it may be 0, 0.1, 1, etc.
[0142] For example, if the cost-performance gain rate corresponding to candidate resource recommendation scheme D is greater than or equal to 0, and the cost-performance gain rate corresponding to candidate resource recommendation scheme E is less than 0, then the cloud management platform determines candidate resource recommendation scheme D as the fifth resource recommendation scheme. Candidate resource recommendation scheme D is the most cost-effective candidate resource recommendation scheme in the n1-th first set of the N first sets, where n1 is a positive integer greater than 0 and less than N. Candidate resource recommendation scheme E is the most cost-effective resource recommendation scheme in the (n1+1)-th first set of the N first sets. The N first sets are sorted according to a preset order.
[0143] Optionally, the cloud management platform determines at least one first resource recommendation scheme based on the fifth resource recommendation scheme. The fifth resource recommendation scheme belongs to the at least one first resource recommendation scheme. Specifically, the cloud management platform determines the first set to which the fifth resource recommendation scheme belongs as the first candidate resource recommendation scheme set, and determines the at least one first resource recommendation scheme from the first candidate resource recommendation scheme set. The first candidate resource recommendation scheme set and the first resource recommendation schemes are described in Figure 2. That is, after determining the most cost-effective candidate resource recommendation scheme among the at least one candidate resource recommendation schemes, the cloud management platform can continue to determine one or more other candidate resource recommendation schemes with relatively high cost-effectiveness among the at least one candidate resource recommendation schemes.
[0144] Optionally, the cloud management platform determines at least one first resource recommendation scheme and at least one second resource recommendation scheme based on the fifth resource recommendation scheme. The fifth resource recommendation scheme belongs to the at least one first resource recommendation scheme. Specifically, the implementation method of the cloud management platform determining the at least one first resource recommendation scheme is similar to the above implementation method and will not be repeated here. Based on the fifth resource recommendation scheme, the cloud management platform determines a set of second candidate resource recommendation schemes, which belongs to the N first sets. When the N first sets are sorted according to a preset order, the set of second candidate resource recommendation schemes is either the set preceding or following the set of first candidate resource recommendation schemes. The cloud management platform determines at least one second resource recommendation scheme from the set of second candidate resource recommendation schemes. The set of second candidate resource recommendation schemes and the second resource recommendation schemes are described in Figure 2.
[0145] Optionally, the cloud management platform determines at least one first resource recommendation scheme and at least one third resource recommendation scheme based on the fifth resource recommendation scheme. The fifth resource recommendation scheme belongs to the at least one first resource recommendation scheme. Specifically, the implementation method of the cloud management platform determining the at least one first resource recommendation scheme is similar to the above implementation method and will not be repeated here. Based on the fifth resource recommendation scheme, the cloud management platform determines a set of third candidate resource recommendation schemes, which belongs to the N first sets. When the N first sets are sorted according to a preset order, the set of third candidate resource recommendation schemes is either the set preceding or following the set of first candidate resource recommendation schemes. The cloud management platform determines at least one third resource recommendation scheme from the set of third candidate resource recommendation schemes. The set of third candidate resource recommendation schemes and the third resource recommendation schemes are described in Figure 2.
[0146] Optionally, the cloud management platform determines at least one first resource recommendation scheme and at least one fourth resource recommendation scheme based on the fifth resource recommendation scheme. The fifth resource recommendation scheme belongs to the at least one first resource recommendation scheme. Specifically, the implementation method of the cloud management platform determining the at least one first resource recommendation scheme is similar to the above implementation method and will not be repeated here. Based on the fifth resource recommendation scheme, the cloud management platform determines a set of fourth candidate resource recommendation schemes, which belongs to the N first sets. When the N first sets are sorted according to a preset order, the set of fourth candidate resource recommendation schemes is either the set preceding or following the set of first candidate resource recommendation schemes. The cloud management platform determines at least one fourth resource recommendation scheme from the set of fourth candidate resource recommendation schemes. The set of fourth candidate resource recommendation schemes and the fourth resource recommendation schemes are described in Figure 2.
[0147] In the method shown in Figure 5, the cloud management platform divides the candidate resource recommendation schemes into at least one first set based on the number of computing units included in the candidate resource recommendation schemes and / or the parallel parameters of the first parallel strategy corresponding to the candidate resource recommendation schemes. Then, based on the cost-performance gain rate of the resource recommendation scheme with the highest cost-performance ratio in the first set, it determines the resource recommendation scheme with the highest cost-performance ratio among the at least one resource recommendation schemes, and thus provides tenants with at least one resource recommendation scheme with a high cost-performance ratio.
[0148] Figure 6 is a schematic structural diagram of a computing device provided in an embodiment of this application. The computing device 600 in Figure 6 includes a transceiver module 610 and a processing module 620. The computing device 600 in Figure 6 can be used to execute the methods in Figure 2 or Figure 5. The computing device 600 in Figure 6 can be applied to the cloud management platform in Figure 1.
[0149] When the computing device 600 executes the method in FIG2, the transceiver module 610 is used to receive first instruction information from the tenant and to provide the tenant with at least one first resource recommendation scheme. The transceiver module 610 is used to execute steps 210 and 230 in FIG2. The processing module 620 is used to: determine at least one first resource recommendation scheme based on the available resources in at least one computing unit cluster and the model information of the first AI model. The processing module 620 is used to execute step 220 in FIG2. The first instruction information, the first resource recommendation scheme, the available resources in the computing unit cluster, and the model information of the first AI model are described in FIG2.
[0150] In some embodiments, the transceiver module 610 is further configured to receive second indication information from a tenant. The processing module 620 is further configured to provide the tenant with resources included in the target resource recommendation scheme based on the second indication information. The second indication information and the target resource recommendation scheme are described in FIG2.
[0151] In some embodiments, the transceiver module 610 is further configured to receive third indication information from a tenant, as described in FIG2.
[0152] When the computing device 600 executes the method in FIG5, the processing module 620 is configured to: determine N first sets based on the available resources in at least one computing unit cluster and the model information of the first AI model; determine the cost-performance gain rate corresponding to the cost-performance of the most cost-effective candidate resource recommendation scheme in each of the N first sets based on the cost-performance of the most cost-effective candidate resource recommendation scheme in each first set; and determine at least one first resource recommendation scheme based on the cost-performance gain rate corresponding to the most cost-effective candidate resource recommendation scheme in each first set. The processing module 620 is configured to execute steps 510-530 in FIG5. The first sets, cost-performance, and cost-performance gain rate are described in FIG5.
[0153] Both the transceiver module 610 and the processing module 620 can be implemented in software or in hardware. For example, the implementation of the processing module 620 will be described below. Similarly, the implementation of the transceiver module 610 can be referenced from the implementation of the processing module 620.
[0154] As an example of a software functional unit, processing module 620 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, processing module 620 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed within the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0155] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same VPC or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0156] As an example of a hardware functional unit, the processing module 620 may include at least one computing device, such as a server. Alternatively, the processing module 620 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0157] The processing module 620 includes multiple computing devices that can be distributed within the same region or in different regions. Similarly, the processing module 620 includes multiple computing devices that can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the processing module 620 includes multiple computing devices that can be distributed within the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0158] Therefore, the modules of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0159] It should be noted that the above embodiments of the device, when executing the above methods, are only illustrative examples of the division of functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. For example, the transceiver module 610 can be used to execute any step in the above methods, and the processing module 620 can be used to execute any step in the above methods. The steps implemented by the transceiver module 610 and the processing module 620 can be specified as needed, and the device can achieve all its functions by implementing different steps in the above methods through the transceiver module 610 and the processing module 620 respectively.
[0160] Furthermore, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments above, which will not be repeated here.
[0161] The method provided in this application can be executed by a computing device, which can also be referred to as a computer system. It includes a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as processing units, memory, and memory control units; the functions and structure of this hardware will be described in detail later. The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software. Optionally, the computer system can be a handheld device such as a smartphone, or a terminal device such as a personal computer; this application does not particularly limit this, as long as the method provided in this application can be used. The executing entity of the method provided in this application can be a computing device, or a functional module within the computing device capable of calling and executing programs.
[0162] Figure 7 is a schematic structural block diagram of a computing device 700 provided in an embodiment of this application. The computing device 700 can be a server, a computer, or other device with computing capabilities. The computing device 700 shown in Figure 7 includes at least one processor 710 and a memory 720.
[0163] It should be understood that this application does not limit the number of processors and memories in the computing device 700.
[0164] The processor 710 executes instructions in the memory 720, causing the computing device 700 to implement the method provided in this application. Alternatively, the processor 710 executes instructions in the memory 720, causing the computing device 700 to implement the various functional modules provided in this application, thereby implementing the method provided in this application.
[0165] Optionally, the computing device 700 also includes a communication interface 730. The communication interface 730 uses a transceiver module, such as, but not limited to, a network interface card or a transceiver, to enable communication between the computing device 700 and other devices or communication networks.
[0166] Optionally, the computing device 700 also includes a system bus 740, wherein the processor 710, memory 720, and communication interface 730 are respectively connected to the system bus 740. The processor 710 can access the memory 720 through the system bus 740; for example, the processor 710 can perform data read / write or code execution in the memory 720 through the system bus 740. The system bus 740 is a peripheral component interconnect express (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The system bus 740 is divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in Figure 7, but this does not mean that there is only one bus or one type of bus.
[0167] In one possible implementation, the processor 710 primarily functions to interpret the instructions (or code) of a computer program and process data within the computer software. The instructions of the computer program and the data within the computer software can be stored in the memory 720 or the cache of the processor 710.
[0168] Optionally, the processor 710 may be an integrated circuit chip with signal processing capabilities. By way of example and not limitation, the processor 710 may be a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA, or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Among these, a general-purpose processor is a microprocessor, etc. For example, the processor 710 may be a central processing unit (CPU).
[0169] The memory 720 provides runtime space for processes in the computing device 700. For example, the memory 720 stores the computer program (specifically, the program code) used to generate the process. After the computer program is run by the processor to generate a process, the processor allocates corresponding storage space for the process in the memory 720. Furthermore, the aforementioned storage space further includes text segments, initialized data segments, bit initialized data segments, stack segments, heap segments, etc. The memory 720 stores data generated during the process's execution, such as intermediate data or process data, in the aforementioned process-specific storage space.
[0170] Optionally, the memory, also known as RAM, is used to temporarily store the data processed by the processor 710, as well as data exchanged with external storage devices such as hard disks. As long as the computer is running, the processor 710 will load the data required for processing into RAM for computation, and then transfer the result back out after the computation is complete.
[0171] By way of example and not limitation, memory 720 may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile storage medium may be, for example, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory is random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus DRAM (DRDRAM). It should be noted that the memory 720 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0172] The structure of the computing device 700 listed above is merely illustrative and is not limited thereto. The computing device 700 in this application includes various hardware components in existing computer systems. For example, the computing device 700 also includes other memories besides the memory 720, such as disk storage. Those skilled in the art should understand that the computing device 700 may also include other devices necessary for normal operation. Furthermore, depending on specific needs, those skilled in the art should understand that the computing device 700 may also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that the computing device 700 may only include the devices necessary for implementing the embodiments of this application, and not necessarily all the devices shown in FIG. 7.
[0173] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server. In some embodiments, the computing device may also be a desktop computer, a laptop computer, or a smartphone, or other terminal device.
[0174] As shown in Figure 8, the computing device cluster includes at least one computing device 700. The memory 720 of one or more computing devices 700 in the computing device cluster may store the same instructions for performing the above-described method.
[0175] In some possible implementations, the memory 720 of one or more computing devices 700 in the computing device cluster may also each store a portion of the instructions for executing the above-described method. In other words, a combination of one or more computing devices 700 can jointly execute the instructions of the above-described method.
[0176] It should be noted that the memories 720 in different computing devices 700 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the aforementioned apparatus. That is, the instructions stored in the memories 720 of different computing devices 700 can implement the functions of one or more modules within the aforementioned apparatus.
[0177] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 9 illustrates one possible implementation. As shown in Figure 9, two computing devices, 700A and 800B, are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device.
[0178] It should be understood that the functions of computing device 700A shown in Figure 9 can also be performed by multiple computing devices 700. Similarly, the functions of computing device 700B can also be performed by multiple computing devices 700.
[0179] In this embodiment of the application, a computer program product containing instructions is also provided. The computer program product may be software or a program product containing instructions that can run on a computing device cluster or be stored on any available medium. When run by the computing device cluster, it causes the computing device cluster to perform the methods provided above, or causes the computing device cluster to perform the functions of the apparatus provided above.
[0180] This application embodiment also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., high-density digital video disc (DVD)), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method provided above.
[0181] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0182] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0183] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0184] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0185] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0186] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0187] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A resource recommendation method, characterized in that, The method is applied to a cloud management platform, which manages infrastructure providing cloud services. The infrastructure includes at least one computing unit cluster, and each computing unit cluster includes at least one computing unit. The method includes: Receive first indication information from the tenant, the first indication information being used to indicate the model information of the tenant's first artificial intelligence (AI) model; Based on the available resources in the at least one computing unit cluster and the model information of the first AI model, at least one first resource recommendation scheme is determined. The at least one first resource recommendation scheme belongs to a first candidate resource recommendation scheme set. The cost-effectiveness of the at least one first resource recommendation scheme is better than that of the candidate resource recommendation schemes in the first candidate resource recommendation scheme set other than the at least one first resource recommendation scheme. The candidate resource recommendation schemes in the first candidate resource recommendation scheme set meet the running requirements of the first AI model. The resources included in the candidate resource recommendation schemes belong to the available resources. The candidate resource recommendation schemes include the number of computing unit clusters used to run the first AI model and the number of computing units in each computing unit cluster. Provide the tenant with at least one first resource recommendation scheme.
2. The method according to claim 1, characterized in that, The method further includes: Receive second indication information from the tenant, the second indication information being used to indicate a target resource recommendation scheme, the target resource recommendation scheme belonging to the at least one first resource recommendation scheme; Based on the second instruction information, the tenant is provided with resources included in the target resource recommendation scheme, and the resources included in the target resource recommendation scheme are used to run the first AI model.
3. The method according to claim 1 or 2, characterized in that, The step of determining at least one first resource recommendation scheme based on the available resources in the at least one computing unit cluster and the model information of the first AI model includes: Based on the available resources in the at least one computing unit cluster and the model information of the first AI model, a first candidate resource recommendation scheme set is determined, wherein the first candidate resource recommendation scheme set includes at least one candidate resource recommendation scheme; Determine the performance and price of each candidate resource recommendation scheme in the first candidate resource recommendation scheme set; The at least one first resource recommendation scheme is determined based on the performance and price of each candidate resource recommendation scheme in the first candidate resource recommendation scheme set.
4. The method according to any one of claims 1 to 3, characterized in that, At least two candidate resource recommendation schemes in the first candidate resource recommendation scheme set have a different total number of computing units; or, Each candidate resource recommendation scheme in the first candidate resource recommendation scheme set includes the same total number of computing units.
5. The method according to claim 4, characterized in that, When each candidate resource recommendation scheme in the first candidate resource recommendation scheme set includes the same total number of computing units, determining at least one first resource recommendation scheme based on the available resources in the at least one computing unit cluster and the model information of the first AI model includes: Based on the available resources in the at least one computing unit cluster and the model information of the first AI model, at least one second resource recommendation scheme and the at least one first resource recommendation scheme are determined. The at least one second resource recommendation scheme belongs to a set of second candidate resource recommendation schemes. The cost-effectiveness of the at least one second resource recommendation scheme is better than that of the candidate resource recommendation schemes in the set of second candidate resource recommendation schemes other than the at least one second resource recommendation scheme. The candidate resource recommendation schemes in the set of second candidate resource recommendation schemes meet the running requirements of the first AI model. The resources included in the candidate resource recommendation schemes in the set of second candidate resource recommendation schemes belong to the available resources. The candidate resource recommendation schemes in the set of second candidate resource recommendation schemes include the number of computing unit clusters used to run the first AI model and the number of computing units in each computing unit cluster. The total number of computing units included in each candidate resource recommendation scheme in the set of second candidate resource recommendation schemes is the same. The total number of computing units included in the candidate resource recommendation schemes in the set of second candidate resource recommendation schemes is different from that in the set of first candidate resource recommendation schemes. Providing the tenant with the at least one first resource recommendation scheme includes: Provide the tenant with at least one first resource recommendation scheme and at least one second resource recommendation scheme.
6. The method according to any one of claims 1 to 5, characterized in that, The step of determining at least one first resource recommendation scheme based on the available resources in the at least one computing unit cluster and the model information of the first AI model includes: Obtain the parallel strategy scheme adopted by the first AI model in operation, wherein the parallel strategy scheme includes parallel parameters of at least one parallel strategy; Based on the parallel strategy scheme, the available resources in the at least one computing unit cluster, and the model information of the first AI model, determine the at least one first resource recommendation scheme; The step of obtaining the parallel strategy scheme adopted by the first AI model includes: Receive third indication information from the tenant, the third indication information being used to indicate the parallel strategy scheme; or, Based on the model information of the first AI model, at least one of the parallel strategy schemes is determined.
7. The method according to claim 6, characterized in that, The at least one parallel strategy includes at least one of the following: pipeline parallelism (PP), tensor parallelism (TP), data parallelism (DP), expert parallelism (EP), and sequence parallelism (SP).
8. The method according to claim 6 or 7, characterized in that, The parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme in the first candidate resource recommendation scheme set are the same; or, Each candidate resource recommendation scheme in the first candidate resource recommendation scheme set includes the same total number of computing units, and the parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme are the same. Wherein, the first parallel strategy belongs to the at least one parallel strategy.
9. The method according to claim 8, characterized in that, The first parallel strategy includes at least one of the following: data parallelism (DP) strategy and sequence parallelism (SP) strategy.
10. The method according to any one of claims 1 to 9, characterized in that, The model information of the first AI model includes the parameter scale and / or model structure of the first AI model.
11. The method according to any one of claims 1 to 10, characterized in that, The computing unit includes at least one of the following: a central processing unit (CPU), a graphics processing unit (GPU), and a neural network processing unit (NPU).
12. A computing device, characterized in that, The apparatus is applied to a cloud management platform, which manages infrastructure providing cloud services. The infrastructure includes at least one computing unit cluster, and each computing unit cluster includes at least one computing unit. The apparatus includes: The transceiver module is used to receive first indication information from a tenant, wherein the first indication information is used to indicate model information of the tenant's first artificial intelligence (AI) model; The processing module is configured to determine at least one first resource recommendation scheme based on the available resources in the at least one computing unit cluster and the model information of the first AI model. The at least one first resource recommendation scheme belongs to a first candidate resource recommendation scheme set. The cost-effectiveness of the at least one first resource recommendation scheme is better than that of the candidate resource recommendation schemes in the first candidate resource recommendation scheme set other than the at least one first resource recommendation scheme. The candidate resource recommendation schemes in the first candidate resource recommendation scheme set meet the running requirements of the first AI model. The resources included in the candidate resource recommendation schemes belong to the available resources. The candidate resource recommendation schemes include the number of computing unit clusters used to run the first AI model and the number of computing units in each computing unit cluster. The transceiver module is also used to provide the tenant with the at least one first resource recommendation scheme.
13. The apparatus according to claim 12, characterized in that, The transceiver module is further configured to receive second indication information from the tenant, the second indication information being used to indicate a target resource recommendation scheme, the target resource recommendation scheme belonging to at least one first resource recommendation scheme; The processing module is further configured to provide the tenant with resources included in the target resource recommendation scheme according to the second indication information, wherein the resources included in the target resource recommendation scheme are used to run the first AI model.
14. The apparatus according to claim 12 or 13, characterized in that, The processing module is specifically used for: Based on the available resources in the at least one computing unit cluster and the model information of the first AI model, a first candidate resource recommendation scheme set is determined, wherein the first candidate resource recommendation scheme set includes at least one candidate resource recommendation scheme; Determine the performance and price of each candidate resource recommendation scheme in the first candidate resource recommendation scheme set; The at least one first resource recommendation scheme is determined based on the performance and price of each candidate resource recommendation scheme in the first candidate resource recommendation scheme set.
15. The apparatus according to any one of claims 12 to 14, characterized in that, At least two candidate resource recommendation schemes in the first candidate resource recommendation scheme set have a different total number of computing units; or, Each candidate resource recommendation scheme in the first candidate resource recommendation scheme set includes the same total number of computing units.
16. The apparatus according to claim 15, characterized in that, When each candidate resource recommendation scheme in the first candidate resource recommendation scheme set includes the same total number of computing units, The processing module is specifically configured to determine at least one second resource recommendation scheme and at least one first resource recommendation scheme based on the available resources in the at least one computing unit cluster and the model information of the first AI model. The at least one second resource recommendation scheme belongs to a set of second candidate resource recommendation schemes. The cost-effectiveness of the at least one second resource recommendation scheme is better than that of the candidate resource recommendation schemes in the set of second candidate resource recommendation schemes other than the at least one second resource recommendation scheme. The candidate resource recommendation schemes in the set of second candidate resource recommendation schemes meet the running requirements of the first AI model. The resources included in the candidate resource recommendation schemes in the set of second candidate resource recommendation schemes belong to the available resources. The candidate resource recommendation schemes in the set of second candidate resource recommendation schemes include the number of computing unit clusters used to run the first AI model and the number of computing units in each computing unit cluster. The total number of computing units included in each candidate resource recommendation scheme in the set of second candidate resource recommendation schemes is the same. The total number of computing units included in the candidate resource recommendation schemes in the set of second candidate resource recommendation schemes is different from that in the set of first candidate resource recommendation schemes. The transceiver module is specifically used to provide the tenant with at least one first resource recommendation scheme and at least one second resource recommendation scheme.
17. The apparatus according to any one of claims 12 to 16, characterized in that, The processing module is specifically used for: Obtain the parallel strategy scheme adopted by the first AI model in operation, wherein the parallel strategy scheme includes parallel parameters of at least one parallel strategy; Based on the parallel strategy scheme, the available resources in the at least one computing unit cluster, and the model information of the first AI model, determine the at least one first resource recommendation scheme; Specifically, the transceiver module is used to receive third indication information from the tenant, which is used to indicate the parallel strategy scheme; or, the processing module is used to determine at least one of the parallel strategy schemes based on the model information of the first AI model.
18. The apparatus according to claim 17, characterized in that, The at least one parallel strategy includes at least one of the following: pipeline parallelism (PP), tensor parallelism (TP), data parallelism (DP), expert parallelism (EP), and sequence parallelism (SP).
19. The apparatus according to claim 17 or 18, characterized in that, The parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme in the first candidate resource recommendation scheme set are the same; or, Each candidate resource recommendation scheme in the first candidate resource recommendation scheme set includes the same total number of computing units, and the parallel parameters of the first parallel strategy corresponding to each candidate resource recommendation scheme are the same. Wherein, the first parallel strategy belongs to the at least one parallel strategy.
20. The apparatus according to claim 19, characterized in that, The first parallel strategy includes at least one of the following: data parallelism (DP) strategy and sequence parallelism (SP) strategy.
21. The apparatus according to any one of claims 12 to 20, characterized in that, The model information of the first AI model includes the parameter scale and / or model structure of the first AI model.
22. The apparatus according to any one of claims 12 to 21, characterized in that, The computing unit includes at least one of the following: a central processing unit (CPU), a graphics processing unit (GPU), and a neural network processing unit (NPU).
23. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 11.
24. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1 to 11.
25. A computer-readable storage medium, characterized in that, Includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Cloud system resource set recommendation method and device, and computing device cluster
CN110837417A
Computing resource recommendation method and device, electronic equipment and storage medium
CN111865644A
Cloud resource charging method, cloud management platform, computing device and storage medium
CN117857228A
Evaluation framework for cloud resource optimization
US20230246981A1
Cited By
Resource planning method and device, electronic equipment, storage medium and program product
CN122044892A