Public cloud technology-based cloud resource management method and apparatus
By transferring cloud resources between tenants and efficiently exchanging parameters within a network with the same parameter plane through a cloud management platform, the problem of uneven cloud resource utilization is solved, achieving efficient resource utilization and reducing AI development costs.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-10-09
- Publication Date
- 2026-07-30
AI Technical Summary
Existing cloud management platforms suffer from uneven resource utilization, resulting in higher resource consumption on weekdays and lower resource consumption on non-weekdays. Furthermore, existing technologies such as spot instances are not suitable for artificial intelligence development scenarios, preventing tenants from conducting AI development activities at a lower cost.
The cloud management platform allows cloud resources to be transferred between different tenants by obtaining tenants' successful purchase information and transfer requests, ensuring resource demand matching, and efficiently exchanging model parameters within the same parameter plane network, supporting multiple tenants to participate in model training tasks together.
It improves the utilization rate of cloud resources and the flexibility of tenant usage, reduces the cost of AI development, and enhances the efficiency of model training tasks and resource utilization.
Smart Images

Figure CN2025126588_30072026_PF_FP_ABST
Abstract
Description
Methods and apparatus for managing cloud resources based on public cloud technology
[0001] Cross-reference of related applications
[0002] This application claims priority to Chinese Patent Application No. 202510125728.5, filed on January 26, 2025, entitled "Method and Apparatus for Managing Cloud Resources Based on Public Cloud Technology", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of cloud service technology, and in particular to a method and apparatus for managing cloud resources based on public cloud technology. Background Technology
[0004] Cloud service providers can offer tenants various types of cloud resources to enable them to use these resources to realize their business needs. After purchasing cloud resources from a cloud management platform, tenants can deploy instances within those resources and then deploy applications within those instances to run their business.
[0005] Currently, cloud resources provided by cloud management platforms are subject to uneven usage, with higher resource consumption on weekdays and lower consumption on weekends. Related technologies, such as spot instances, attempt to address this issue. However, this approach requires significant flexibility from tenant businesses and is unsuitable for scenarios involving artificial intelligence (AI) development, preventing tenants from conducting AI development activities at a lower cost. Summary of the Invention
[0006] This application provides a method and apparatus for managing cloud resources based on public cloud technology. This application improves the resource assurance capabilities of the cloud management platform. The technical solution provided by this application is as follows:
[0007] Firstly, this application provides a method for managing cloud resources based on public cloud technology. The method is applied to a cloud management platform, which manages infrastructure providing cloud services. The infrastructure includes multiple servers, each of which has cloud resources. The method includes: the cloud management platform obtaining first successful purchase information for a first cloud resource in the infrastructure from a first tenant; the cloud management platform responding to the first successful purchase information by setting the user of the first cloud resource as the first tenant and setting the usage time as a first duration; the cloud management platform obtaining a transfer request for the first cloud resource from the first tenant, the transfer request including a second duration, the second duration being less than or equal to the first duration; and the cloud management platform responding to the transfer request for the first cloud resource by publishing the first cloud resource. The first resource transfer information is sent to the Internet. The first resource transfer information includes the transfer conditions for the first cloud resource and for other tenants besides the first tenant. The cloud management platform obtains the model training task submitted by the second tenant, confirms the cloud resource requirements of the model training task based on the model training task, and confirms that the first cloud resource meets part or all of the cloud resource requirements. The cloud management platform obtains the second purchase success information of the second tenant for the first resource transfer information, confirms that the second purchase success information meets the transfer conditions, and sets the user of the first cloud resource as the second tenant, sets the usage time as the second duration, and the cloud management platform calls the first cloud resource to participate in the model training task within the second duration.
[0008] Through the above scheme, the cloud management platform can provide cloud resources that meet the cloud resource requirements of the second tenant's model training task. These cloud resources are transferred from the first tenant. Thus, when the first tenant has a low demand for the first cloud resources it purchased, it can transfer the first cloud resources, and the second tenant can use these cloud resources to train the model at a lower cost. This expands the use cases of cloud resources transferred from other tenants, improves the utilization rate of cloud resources, and increases the flexibility of tenants in using cloud resources.
[0009] In one possible implementation of the first aspect, the cloud management platform confirms that the first cloud resources meet some or all of the cloud resource requirements, including: the cloud management platform confirms that each of the first cloud resources is located in the same parameter plane network.
[0010] Through the above scheme, the cloud management platform confirms that the cloud resources in the first cloud resource are located in the same parameter plane network. Therefore, when the cloud management platform calls the first cloud resource to participate in the model training task of the second tenant, it enables each cloud resource to efficiently exchange or synchronize model parameters, thereby improving the efficiency of the second tenant in using cloud resources for model training tasks.
[0011] In one possible implementation of the first aspect, the cloud management platform confirms that the first cloud resource meets part or all of the cloud resource requirements, including: the cloud management platform confirms that the first cloud resource and the second cloud resource in the infrastructure meet part or all of the cloud resource requirements, each of the first cloud resources and each of the second cloud resources are located in the same parametric plane network, and the method further includes: the cloud management platform calling the second cloud resources to participate in the model training task within a second time period.
[0012] Through the above scheme, the cloud management platform confirms that the cloud resources in the first and second cloud resources are located in the same parametric plane network. Therefore, the second tenant can not only use the first cloud resources transferred by the first tenant for model training, but also participate in the model training task together when the first and second cloud resources jointly meet the second tenant's cloud resource needs. This improves the flexibility of tenants in using cloud resources for model training tasks. Furthermore, since the cloud resources in the first and second cloud resources are located in the same parametric plane network, each cloud resource can efficiently exchange or synchronize model parameters for the model training task, improving the efficiency of the second tenant in using cloud resources for model training tasks.
[0013] In one possible implementation of the first aspect, the method further includes: the cloud management platform obtaining third purchase success information of a third tenant for the second cloud resource; the cloud management platform responding to the third purchase success information setting the user of the second cloud resource as the third tenant; the cloud management platform obtaining the transfer request of the third tenant for the second cloud resource; and the cloud management platform responding to the transfer request for the second cloud resource publishing the second resource transfer information of the second cloud resource to the Internet, wherein the second resource transfer information includes transfer conditions for the second cloud resource and for tenants other than the third tenant.
[0014] Through the above scheme, the second cloud resource is the cloud resource transferred from the third tenant. The cloud management platform calls the first cloud resource transferred from the first tenant and the second cloud resource transferred from the third tenant to jointly participate in the second tenant's model training task. When the second tenant uses cloud resources transferred by others, it is not limited to cloud resources transferred by a single tenant. In addition, the cloud management platform will schedule cloud resources transferred from different tenants to match the cloud resource requirements of the tenant's model training task, thereby reducing the cost of the second tenant using cloud resources for model training tasks and improving the efficiency of the second tenant using cloud resources for model training tasks.
[0015] In one possible implementation of the first aspect, the model training task uses a training dataset to train the neural network model. The training dataset includes multiple batches of training data. After the cloud management platform calls the first cloud resource to participate in the model training task within a second time period, the method further includes: in response to a resource interruption request for the first cloud resource, the cloud management platform saves the checkpoint file of the neural network model after the current batch of training data has been trained, and the cloud management platform releases the first cloud resource.
[0016] Through the above solution, in response to a resource interruption request for the first cloud resource, the cloud management platform saves a checkpoint file after the current batch of training data has been trained. This, combined with the cloud management platform's ability to monitor resource interruption requests and control the start and stop of model training tasks, ensures that the second tenant's model training task has completed training on the current batch of data and saved the checkpoint file before the first cloud resource is interrupted or released. This avoids wasting training resources by directly interrupting the current batch of training. When the second tenant resumes model training, it can continue from the next batch based on the saved checkpoint file. This reduces the cost and improves the efficiency of the second tenant using cloud resources for model training.
[0017] In one possible implementation of the first aspect, the resource interruption request is issued based on the end time of the second duration, and the method further includes: the cloud management platform changing the user of the first cloud resource to the first tenant.
[0018] In one possible implementation of the first aspect, the model training task includes a model loading phase and at least one round of model training phase, the cloud management platform bills the second tenant using a time-based billing method, and the method further includes: the cloud management platform updating the billing for the second tenant's purchase of the first cloud resources in response to the end of the model loading phase or the start of at least one round of model training phase.
[0019] By leveraging the cloud management platform's ability to monitor and control the model training phase, the above solution allows for more accurate billing of the second tenant's cloud resource usage for model training. This enables billing to be generated based on the completion of the model loading phase and the start of a specific training phase, improving the accuracy of tenant billing. In one scenario, billing only begins after the model loading phase ends and a training phase starts, reducing the cost for the second tenant using cloud resources for model training.
[0020] In one possible implementation of the first aspect, the resource demand information includes a first bid from a second tenant for the first cloud resource, and a resource interruption request is issued based on a second bid from a fourth tenant for the first cloud resource that is higher than the first bid. The method further includes: the cloud management platform obtaining fourth purchase success information of the fourth tenant for the transfer information of the first resource, confirming that the fourth purchase success information meets the transfer conditions, and setting the user of the first cloud resource as the fourth tenant.
[0021] In one possible implementation of the first aspect, after the cloud management platform responds to a resource interruption request and saves the checkpoint file of the neural network model after training the current batch of training data is completed, the method further includes: the cloud management platform deploying the model training task on the third cloud resource of the second tenant, the cloud management platform loading the checkpoint file, and training the next batch of training data of the current batch on the third cloud resource.
[0022] With the above solution, if the second tenant wants to continue the previous model training task on the third cloud resource, the cloud management platform can restore the model training based on the previously saved checkpoint file. Since the checkpoint file is saved after the training of a certain batch of training data is completed, the next batch of training data can be trained directly on the third cloud resource, which improves the efficiency of the second tenant in using cloud resources for model training.
[0023] In one possible implementation of the first aspect, the model training task submitted by the second tenant includes at least one of the cloud resource quantity and cloud resource specification in the cloud resource requirements. The cloud management platform obtains the model training task submitted by the second tenant and confirms the cloud resource requirements of the model training task based on the model training task, which includes: the cloud management platform obtains the model training task submitted by the second tenant and confirms the cloud resource requirements of the model training task based on at least one of the cloud resource quantity and cloud resource specification.
[0024] Through the above scheme, the model training task submitted by the second tenant includes information on at least one of the cloud resource requirements: the quantity of cloud resources and the cloud resource specifications. The cloud management platform can confirm the cloud resource requirements based on this information.
[0025] In one possible implementation of the first aspect, the model training task submitted by the second tenant includes model template information, which includes the identifier of a pre-set model. The cloud management platform obtains the model training task submitted by the second tenant and confirms the cloud resource requirements of the model training task based on the model training task, which includes: the cloud management platform obtains the model training task submitted by the second tenant and confirms the cloud resource requirements of the model training task based on the identifier of the pre-set model in the model template information.
[0026] Through the above scheme, the model training task submitted by the second tenant includes model template information, which includes the identifier of the pre-set model. The cloud management platform can confirm the cloud resource requirements based on the mapping relationship between the identifier and the cloud resource information.
[0027] In one possible implementation of the first aspect, cloud resources include at least one of virtual machines, containers, bare metal servers, and physical machines.
[0028] In one possible implementation of the first aspect, the transfer conditions include successful registration and payment for the first cloud resource on the cloud management platform.
[0029] Secondly, this application provides a cloud resource management device based on public cloud technology. The device includes: an interaction module for acquiring first purchase success information of a first tenant for a first cloud resource in the infrastructure; a management module for responding to the first purchase success information by setting the user of the first cloud resource as the first tenant and setting the usage time as a first duration; the interaction module is also used to acquire a transfer request from the first tenant for the first cloud resource, the transfer request including a second duration, the second duration being less than or equal to the first duration; the management module is also used to respond to the transfer request for the first cloud resource by publishing first resource transfer information of the first cloud resource to the Internet, the first resource transfer information including transfer conditions for the first cloud resource and for tenants other than the first tenant; the interaction module is also used to acquire a model training task submitted by a second tenant; the management module is also used to confirm the cloud resource requirements of the model training task based on the model training task, and confirm that the first cloud resource meets part or all of the cloud resource requirements; the interaction module is also used to acquire second purchase success information of the second tenant for the transfer information of the first resource; the management module is also used to confirm that the second purchase success information meets the transfer conditions, and set the user of the first cloud resource as the second tenant and set the usage time as a second duration; the management module is also used to call the first cloud resource to participate in the model training task within the second duration.
[0030] The second aspect or any implementation thereof is a device implementation corresponding to the first aspect or any implementation thereof. The description in the first aspect or any implementation thereof applies to the second aspect or any implementation thereof, and will not be repeated here.
[0031] Thirdly, this application provides a computing device including a memory and a processor, wherein the memory stores program instructions and the processor executes the program instructions to implement the methods provided in the first aspect of this application and any of its possible implementations.
[0032] Fourthly, this application provides a computing device cluster, including multiple computing devices, each computing device including multiple processors and multiple memories, the multiple memories storing program instructions, and the multiple processors executing the program instructions, so that the computing device cluster implements the method provided in the first aspect of this application and any possible implementation thereof.
[0033] Fifthly, this application provides a computer-readable storage medium that is a non-volatile computer-readable storage medium, which includes program instructions that, when executed on a computing device cluster, cause the computing device cluster to implement the methods provided in the first aspect of this application and any of its possible implementations.
[0034] Sixthly, this application provides a computer program product containing instructions that, when run on a computer, cause the computer to implement the methods provided in the first aspect of this application and any of its possible implementations. Attached Figure Description
[0035] Figure 1 is a structural diagram of an implementation scenario involving a cloud resource management method based on public cloud technology provided in an embodiment of this application;
[0036] Figure 2 is a structural diagram of an implementation scenario involving a cloud resource management method based on public cloud technology provided in an embodiment of this application;
[0037] Figure 3 is a schematic diagram of the deployment of basic resources in a data center according to an embodiment of this application;
[0038] Figure 4 is a flowchart of a cloud resource management method based on public cloud technology provided in an embodiment of this application;
[0039] Figure 5 is a schematic diagram of a cloud resource management device based on public cloud technology provided in an embodiment of this application;
[0040] Figure 6 is a flowchart of another cloud resource management method based on public cloud technology provided in an embodiment of this application;
[0041] Figure 7 is a flowchart of another cloud resource management method based on public cloud technology provided in an embodiment of this application;
[0042] Figure 8 is a schematic diagram of the structure of a cloud resource management device based on public cloud technology provided in an embodiment of this application;
[0043] Figure 9 is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0044] Figure 10 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;
[0045] Figure 11 is a schematic diagram of another computing device cluster provided in an embodiment of this application. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0047] To facilitate understanding, the technologies and background involved in the embodiments of this application will be introduced below.
[0048] Cloud computing is a type of distributed computing that refers to a network that centrally manages and schedules a large number of computing and storage resources to provide on-demand services to users. These computing and storage resources are provided through clusters of computing devices located in data centers. Furthermore, cloud computing can provide users with various types of services, such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). Infrastructure as a Service provides virtual machines or other resources as a service to tenants. Platform as a Service provides a development platform as a service to tenants. Software as a Service provides applications (Apps) as a service to customers.
[0049] An Internet Data Center (IDC) is a facility and related service system that provides operation and maintenance for equipment that centrally collects, stores, processes, and transmits data, based on the Internet. Conceptually, it can be understood as a public, commercial Internet "server room," and it is also a professional IT service and a crucial infrastructure for the IT industry. IDC is not only a service concept but also a network concept; it constitutes part of the network infrastructure resources, like backbone networks and access networks, providing high-end data delivery and high-speed access services. Generally, a tenant's on-premises IDC can be understood as their physical server room, where the tenant utilizes existing Internet communication lines and bandwidth resources to establish a standardized, telecommunications-grade server room environment to provide comprehensive services such as server hosting, leasing, and related value-added services. A cloud data center is an Internet data center deployed using the infrastructure resources owned by cloud vendors.
[0050] A resource pool is a collection of various hardware and software resources involved in a cloud data center. Typically, resources in a resource pool can be categorized by type, such as computing resources, storage resources, and network resources.
[0051] A physical machine (PM) is the physical resource used to host virtualization technology. It is also called a physical server. Typically, a physical machine is used to deploy virtual instances. A physical machine has multiple physical devices. For example, a physical server has physical devices such as processors and memory. Multiple virtual instances can be deployed on a single physical machine, sharing the machine's physical resources. Depending on the use case, multiple virtual instances deployed on a single physical machine can belong to the same tenant or to different tenants.
[0052] Virtualization is a resource management technology. Virtualization abstracts and transforms various physical resources of a host, such as computing, network, and storage resources, breaking down the indivisible barriers between the host's physical structures. This allows tenants to utilize these resources in a better way than the original configuration. Resources obtained through virtualization are called virtualized resources, and virtualized resources are not limited by the existing physical resource deployment methods, geographical location, or physical configuration.
[0053] Virtualized resources are typically provided to tenants in the form of virtual instances. Virtual instances utilize the host's hardware resources and run on the host's operating system (OS). Applications run within the virtual instance to implement the tenant's business logic. The host's hardware resources can be allocated to one or more tenants at the virtual instance level. Different virtual instances are isolated from each other, allowing tenants to use physical resources conveniently and flexibly while maintaining security and isolation, and significantly improving the utilization of physical resources. Typically, virtual instances can be virtual machines, containers, or independent processes (such as functions). Virtual instances can also be called Elastic Compute Service (ECS) or Elastic Instances (different cloud service providers may use different names).
[0054] A virtual machine (VM) is a complete computer system with full hardware system functionality, simulated using virtualization technology and running in a completely isolated environment. A subset of the instructions in a VM can be processed on the host machine, while other instructions can be executed in a simulated manner. A VM is also called a virtual server. A VM can be viewed as a collection of virtual devices, which possess full hardware system functionality and run in a completely isolated environment. Virtual devices are created by virtualizing physical devices that can share resources. For example, a virtual processor, created by virtualizing a processor, is a virtual device. Similarly, a training card, created by virtualizing a field-programmable gate array (FPGA), is also a virtual device. For instance, the VM in this application can be a kernel-based virtual machine (KVM). Any task that can be performed on a server can also be performed in a VM. When creating a virtual machine on a server, a portion of the physical machine's hard drive and memory capacity is used as the virtual machine's hard drive and memory capacity. Each virtual machine has its own independent hard drive and operating system, and virtual machine tenants can operate the virtual machine as if it were a server. The runtime environments (such as virtual machine applications, operating systems, and virtual hardware) in different virtual machines are completely isolated, and communication between different virtual machines requires the virtual machine manager to forward network packets.
[0055] Containers utilize the namespace and cgroup technologies supported by the Linux kernel to isolate application processes and their dependencies (the runtime environment's bins / libs, specifically all files required to run the application) within an independent runtime environment. Containers provide a lightweight virtual runtime environment. Containers are created by packaging all the code, libraries, and dependencies of a tenant's application into an image. When the image is executed, it runs in a virtual runtime environment. At this point, the container is a runtime instance of the image, similar to a lightweight sandbox, which can be started, stopped, and deleted. The infrastructure for containers can be server hardware or virtual machines in the cloud (i.e., containers can also be deployed within virtual machines). The operating system uses the Linux kernel and supports namespaces and cgroups. Namespaces are used to isolate processes, while cgroups are used to allocate process resources, specifically virtual processors and memory allocated to the process. The container engine, similar to a virtual machine manager, runs within the operating system and is used to manage containers. Compared to virtual machines, which come with their own operating system, containers do not have an operating system. Instead, containers run as processes within the host machine's operating system. As a result, containers start up faster than virtual machines, making them particularly suitable for lightweight applications. Furthermore, a single host machine can run thousands of containers (processes) simultaneously.
[0056] Resource pooling refers to integrating various computing and storage resources into a unified resource pool for unified dynamic allocation and management. Resource pooling enables high resource sharing, improves resource utilization, simplifies resource management, and provides users with flexible on-demand allocation services.
[0057] Parametric Surface Networks (PSNs): In model training, PSNs support parameter synchronization and communication in distributed training. The model training task can be divided into multiple subtasks and distributed to different computing nodes for parallel execution. PSNs are responsible for efficiently exchanging and synchronizing model parameters between these nodes, thereby ensuring the efficiency and consistency of the training process. The technologies used in PSNs can include, for example, InfiniBand (IB), RoCE, etc., or other types of technologies. This application does not limit the technologies used in PSNs.
[0058] A checkpoint file is a file used in machine learning and deep learning to save the state of the training process. A checkpoint file typically contains some or all of the following: Model parameters: including trainable variables such as weights and biases of neural network layers. Optimizer state: such as optimizer-related parameters like momentum and squared gradients. Training progress: such as the current training epoch and batch number. Performance metrics: such as loss value and accuracy. Other metadata: such as learning rate and training start time. Checkpoint files allow users to resume training from a saved point after a training interruption, instead of starting from scratch.
[0059] Spot instances: A billing model in cloud computing that allows users to purchase idle computing resources on cloud servers at a discounted price. Its core feature is that the price changes in real time based on market supply and demand, typically much cheaper than on-demand instances, but it carries the risk of interruption. Users need to set a maximum bid when purchasing a spot instance. The user can acquire the instance if the bid is higher than the current market price and there is sufficient inventory in the resource pool. If the market price exceeds the user's bid, or if the resource pool is insufficient, the system will interrupt and reclaim the instance. The cloud platform usually issues an interruption warning a few minutes in advance. Due to the possibility of interruption, spot instances are more suitable for scenarios with flexible scheduling, flexible instance availability, and cost sensitivity.
[0060] Cloud service providers can offer tenants various types of cloud resources to enable them to use these resources to realize their business needs. For example, cloud service providers can offer tenants virtual machines, instances, bare metal servers, etc., of appropriate specifications, and deploy applications within these cloud resources to enable tenants to realize their business needs by running these applications.
[0061] In some scenarios, tenants need to use cloud resources to perform AI tasks, such as model training or model inference. Cloud service providers can provide cloud resources including AI computing cards or virtual instances containing AI computing cards, for tenants to deploy and run AI-related tasks. For example, users may use cloud resources for model development activities such as model training and model inference. Generally, when running large-scale AI tasks, a single cloud resource is usually allocated to a single tenant. For example, this can be done through bare metal servers, where one cloud resource corresponding to one bare metal server is allocated to one tenant. In one embodiment, a bare metal server can be configured with one or more AI computing cards.
[0062] Currently, the utilization rate of AI computing resources provided by cloud service providers is relatively uneven. For example, on weekdays or during the day, some AI computing resources suitable for work have a higher utilization rate, while on non-working days or in the evening, some AI computing resources suitable for leisure and entertainment have a higher utilization rate, and the utilization rate of AI computing resources is lower during the period from late night to early morning.
[0063] Currently, cloud management platforms offer tenants two ways to sell cloud resources. One method is on-demand selling, where tenants purchase cloud resources using a pay-as-you-go billing model. This allows tenants to purchase cloud resources when needed and release them promptly when no longer required. The other method is periodic selling. For example, tenants can purchase cloud resources on an annual or monthly basis. During the purchased period, the cloud management platform reserves the purchased cloud resources for the tenant, ensuring they can use them at any time.
[0064] However, both of these sales methods have their drawbacks. For example, with on-demand sales, tenants continuously create and delete virtual instances using purchased cloud resources to reduce costs based on actual business needs. After a tenant releases cloud resources, the cloud management platform no longer reserves resources for them. If the cloud management platform's cloud resource inventory is insufficient when the tenant needs to purchase resources, the tenant may be unable to purchase the required cloud resources. On the other hand, with periodic sales, resource waste can occur if there are off-peak business periods during the purchase period. Therefore, the current methods of selling resources through cloud management platforms are relatively inflexible.
[0065] Currently, related technologies use spot instances to allow tenants to obtain the cloud resources they want at a relatively low price at a specific time. However, spot instances in related technologies are more suitable for scenarios with flexible time, flexible instances, and cost sensitivity, and are therefore not suitable for scenarios where tenants need to perform AI development tasks (such as model training and model inference). Tenants cannot smoothly carry out AI development at a low cost.
[0066] In view of this, embodiments of this application provide a method for managing cloud resources based on public cloud technology. This method is executed by a cloud management platform. The cloud management platform is used to manage the infrastructure providing cloud services. The infrastructure includes multiple servers. Each server has cloud resources configured within it. The method includes: a cloud management platform obtaining first purchase success information for a first cloud resource in the infrastructure from a first tenant; the cloud management platform responding to the first purchase success information setting the user of the first cloud resource as the first tenant and setting the usage time as a first duration; the cloud management platform obtaining a transfer request for the first cloud resource from the first tenant, the transfer request including a second duration, the second duration being less than or equal to the first duration; the cloud management platform responding to the transfer request for the first cloud resource publishing first resource transfer information for the first cloud resource to the Internet, the first resource transfer information including transfer conditions for the first cloud resource and for tenants other than the first tenant; the cloud management platform obtaining a model training task submitted by a second tenant, confirming the cloud resource requirements of the model training task based on the model training task, and confirming that the first cloud resource meets some or all of the cloud resource requirements; the cloud management platform obtaining second purchase success information for the transfer information of the first resource from the second tenant; confirming that the second purchase success information meets the transfer conditions, setting the user of the first cloud resource as the second tenant and setting the usage time as a second duration; and the cloud management platform calling the first cloud resource to participate in the model training task within the second duration.
[0067] In other embodiments provided in this application, the cloud management platform can obtain other AI development tasks submitted by the second tenant, such as model inference and model fine-tuning. The embodiments of this application do not limit the types of services that the tenant wants to deploy.
[0068] This article provides a detailed introduction to the technical solution of this application from multiple perspectives, including implementation scenarios, methods and processes, hardware devices, and software devices.
[0069] The following are examples illustrating the implementation scenarios of the embodiments of this application.
[0070] Figure 1 is a structural diagram of an implementation scenario involving a cloud resource management method based on public cloud technology provided in this application embodiment. As shown in Figure 1, the implementation scenario includes: a data center and a first tenant client 10 and a second tenant client 20. The data center and the first tenant client 10 and second tenant client 20 can establish a communication connection through a network. Optionally, this network can be the Internet 30, or other networks; this application embodiment does not limit the specific network used. Tenants can interact with the data center through the first tenant client 10 and the second tenant client 20. For example, tenants can send cloud service requests and other information to the data center through the first tenant client 10 and the second tenant client 20. The data center responds based on the information sent by the first tenant client 10 and the second tenant client 20.
[0071] Data centers house a large amount of infrastructure owned by cloud service providers, such as computing resources, storage resources, and network resources. For example, computing resources can be computing devices (such as servers) capable of providing computing power. As shown in Figure 1, a data center includes a cloud management platform 40 and infrastructure (not shown in Figure 1). The cloud management platform 40 and the infrastructure are connected via an internal data center network. The cloud management platform 40 manages the infrastructure. The infrastructure provides public cloud services. The infrastructure includes multiple servers. Cloud services are optionally deployed on the servers. Cloud services are implemented by running virtual instances, and are therefore also referred to as virtual instances deployed on servers to implement tenant services. Tenants can send cloud service requests and related information to the server through their first tenant client 10 and second tenant client 20. The server can process the cloud service requests and related information and provide cloud services to the tenant based on the processed cloud service requests and related information. For example, the cloud management platform 40 can manage the cloud resources owned by the cloud management platform 40 through the cloud resource management method based on public cloud technology provided in this application embodiment, so as to provide cloud services to tenants.
[0072] The cloud management platform 40 can be logically divided into: tenant console, compute management service, network management service, storage management service, authentication service, and image management service. The tenant console provides a user interface or application programming interface (API) for interaction with tenants. The compute management service manages servers running virtual instances and bare metal servers. The network management service manages network services (such as gateways and firewalls). The storage management service manages storage services (such as data bucket services). The authentication service manages tenant accounts and passwords. The image management service manages virtual instance images.
[0073] In the implementation scenario shown in Figure 1, a data center contains multiple servers. The servers consist of a hardware layer and a software layer. The hardware layer comprises the standard server configuration, including hardware devices such as processors, memory, network interface cards (NICs), disks, and buses. The software layer includes the operating system installed and running on the server. The operating system relative to the virtual machine can be called the host operating system. The host operating system runs a virtual machine manager (also known as a hypervisor). The virtual machine manager's role is to implement compute virtualization, network virtualization, and storage virtualization for the virtual machines, and to manage the virtual machines.
[0074] The virtual machine manager runs a cloud management platform client. This client receives control plane commands from the cloud management platform, creates virtual instances on the server based on these commands, and manages the virtual instances throughout their lifecycle. For example, the client can monitor the hardware resource usage of the server in real time and report it to the cloud management platform. When the cloud management platform confirms that a virtual instance needs to be created on a specific server, it sends a virtual instance creation command to the client on that server. Upon receiving the command, the client creates the virtual instance on that server. In this way, tenants can create, manage, log in to, and operate virtual instances within the data center through the cloud management platform.
[0075] Servers can run virtual machines of different specifications. Virtual machine specifications are categorized as: general-purpose computing, memory-optimized, ultra-large memory, etc., with specific specifications under each type. After a tenant selects a virtual machine specification, the cloud management platform selects a server in the data center that supports that specification and ensures sufficient idle hardware resources on that server. Then, it creates and configures the virtual machine with that specification on that server. Configuring servers through the cloud management platform allows for the analysis and planning of server hardware resources. Based on the server's hardware performance, it plans the corresponding computing products for the physical hardware, such as planning virtual machines of different specifications, to meet the diverse needs of different tenants. Furthermore, differentiated pricing strategies can be implemented based on the performance differences of virtual machines of different specifications. For example, high-performance virtual instances can be sold at a higher price, while ordinary performance virtual instances can be sold at a lower price, allowing tenants to purchase virtual instances as needed.
[0076] In some scenarios, instances can be containers or bare metal servers. In a scenario where the instance is a bare metal server, as shown in Figure 2, the bare metal server is a server exclusively used by a single tenant. The cloud management platform client is located in a smart card. The computing resources of the bare metal server are used by the tenant's operating system. The tenant can remotely log in to this operating system and control the bare metal server as an administrator. The storage and network resources of the bare metal server are provided by the smart card. The tenant can purchase storage and network resources on the cloud management platform, such as network disks, bandwidth, and Elastic IP addresses (EIPs) for external network access. The purchased storage and network resources are then provided to the bare metal server via a pass-through mechanism through the smart card.
[0077] In one implementation, as shown in Figure 3, the location of basic resources in a data center can be described by cloud resource deployment regions (regions) and availability zones (AZs). Tenants can choose to deploy cloud services based on resources in specific regions and AZs. Regions are defined based on geographical location and network latency. Using the same resource pool within the same region can be understood as sharing public services such as elastic computing, block storage, object storage, virtual private cloud (VPC) networks, elastic internet protocol (EIP) addresses, and images. Regions are divided into general regions and dedicated regions. General regions provide general cloud services to public tenants. Dedicated regions are special regions that host the same type of business or provide business services to specific tenants. A region typically includes multiple AZs. Multiple AZs within a region are connected via high-speed fiber optic cables to meet the needs of tenants building high-availability systems across AZs. An AZ is a collection of one or more data centers as shown in Figure 2. Computing, network, and storage resources within an AZ are logically divided into multiple clusters.
[0078] Tenants can send commands to the cloud management platform through their client to create, manage, log in to, and operate virtual instances on the server, and use the cloud services provided by these virtual instances. For example, the cloud management platform can provide an access interface. This interface can be provided either as a user interface or an API. Tenants can remotely access the access interface through their client to register a cloud account and password on the cloud management platform, and then log in using these accounts and passwords. The cloud management platform can also authenticate the cloud account and password. After successful authentication, the tenant can further select and purchase virtual instances of specific specifications (processor, memory, disk) on the cloud management platform. After the tenant successfully purchases a virtual instance, the cloud management platform provides the tenant with a remote login account and password for the purchased virtual instance. The tenant can use this remote login account and password to remotely log in to the virtual instance through their client, install and run their application within the virtual instance, and then use the application to implement their business operations.
[0079] Client devices can be computers, personal computers, laptops, mobile phones, smartphones, tablets, cloud servers, portable mobile terminals, multimedia players, e-book readers, wearable devices, smart home appliances, artificial intelligence devices, smart wearable devices, smart in-vehicle devices, or IoT devices, etc.
[0080] In one implementation, the cloud resource management method based on public cloud technology provided in this application embodiment can be implemented by running an executable program on computing devices in a data center. Optionally, the cloud resource management method based on public cloud technology provided in this application embodiment can be applied to a cloud resource management system based on public cloud technology. This cloud resource management system is deployed on a server managed by a cloud management platform. The cloud resource management system based on public cloud technology can implement the cloud resource management method based on public cloud technology provided in this application embodiment by running the executable program. Furthermore, the executable program implementing the cloud resource management method based on public cloud technology can optionally be presented in the form of an application installation package. After the server installs the application installation package, it can implement the cloud resource management method based on public cloud technology provided in this application embodiment by running the executable program therein.
[0081] It should be understood that the above content is an exemplary description of the implementation scenarios of the cloud resource management method based on public cloud technology provided in the embodiments of this application, and does not constitute a limitation on the implementation scenarios of the cloud resource management method based on public cloud technology. Those skilled in the art will know that as business needs change, the implementation scenarios can be adjusted according to application requirements, and the embodiments of this application do not specifically limit them. Furthermore, when the cloud resource management method based on public cloud technology provided in the embodiments of this application is applied to other scenarios, the executable program of the method can also be presented in the form of an application installation package or in other ways, and the embodiments of this application do not list them all.
[0082] The following describes the implementation process of a cloud resource management method based on public cloud technology provided in this application embodiment, executed by a cloud management platform. Figure 4 is a flowchart of a cloud resource management method based on public cloud technology provided in this application embodiment. As shown in Figure 4, the implementation process includes the following steps S101 to S110, wherein Figure 4 shows steps S101 to S107.
[0083] Step S101: The cloud management platform obtains the first successful purchase information of the first cloud resource in the infrastructure by the first tenant.
[0084] The first tenant can purchase cloud resources from the cloud management platform 40, deploy virtual instances within these resources, and deploy applications within those virtual instances to realize its business operations. When the first tenant needs to purchase cloud resources, it can perform a specified operation on the first tenant client 10. For example, the first tenant logs into the cloud management platform 40 via the internet 30 through the first tenant client 10, selects a first cloud resource 50 that meets its business needs from the cloud resource purchase interface, and sends a purchase request for that first cloud resource 50 to the cloud management platform 40. This first cloud resource can be any one or more of virtual machines, containers, and bare metal servers, or it can be a serverless computing instance, an edge computing instance, a dedicated host instance, etc. This application does not limit the instance type of the first cloud resource and other cloud resources, and will not elaborate further below. The first cloud resource can be one or more cloud resources; this application does not limit the number of first cloud resources.
[0085] In one possible implementation, the cloud management platform 40 may provide a resource purchase interface to the first tenant. The first tenant can trigger a first purchase success message for the first cloud resource 50 based on this resource purchase interface. For example, when the first tenant needs to purchase the first cloud resource 50, it can perform a specified operation on the first tenant client 10 to indicate the conditions that the cloud resource the first tenant wishes to purchase must meet. After the resource purchase interface responds to this specified operation, if the cloud management platform 40 can provide the first tenant with cloud resources that meet the conditions, it will trigger a first purchase success message for the first cloud resource 50. The first purchase success message carries relevant information about the first cloud resource 50. For example, the first purchase success message carries indication information such as the purchase time period, type, and specifications of the first cloud resource 50. After the first tenant triggers the first purchase success message, the cloud management platform 40 can obtain the first purchase success message through the resource purchase interface. For example, the resource purchase interface is implemented through one or more of the following: an application programming interface (API) and a user interface (UI).
[0086] Step S102: In response to the first successful purchase information, the cloud management platform sets the user of the first cloud resource as the first tenant and the usage time as the first duration.
[0087] After obtaining the first successful purchase information, the cloud management platform 40 can determine that the first tenant has purchased the first cloud resource 50 and become a user of the first cloud resource 50. In response, the platform sets the user information of the first cloud resource 50 to the first tenant. For example, the user information (user) of the first cloud resource 50 is marked as the first tenant, and the first tenant can indicate this through their account. Thus, when the first tenant needs to use the first cloud resource 50, the cloud management platform 40 will only access the first cloud resource 50 for the first tenant if the user information of the first cloud resource 50 is the first tenant.
[0088] Optionally, considering that the first tenant may transfer the first cloud resource 50 during a portion of the purchase period, the cloud management platform 40 also needs to maintain the owner information of the first cloud resource. This is so that after the transfer of the first cloud resource 50 is completed, the right to use the first cloud resource can be returned to the first tenant indicated by the owner information based on this owner information. Therefore, in response to the first purchase information, the cloud management platform also sets the owner information of the first cloud resource to the first tenant. In one possible implementation, the first tenant can indicate this through their account. The first tenant can purchase the first cloud resource on a subscription basis or on demand. For example, if the first tenant purchases the first cloud resource on a subscription basis, the cloud management platform can set the usage time of the first cloud resource as a first duration, which is the duration specified by the tenant when purchasing on a subscription basis, such as one week, one month, six months, or one year.
[0089] Step S103: Obtain the transfer request from the first tenant for the first cloud resource. The transfer request includes a second duration, which is less than or equal to the first duration.
[0090] When a first tenant has a low demand for the first cloud resource it purchased, it may choose to transfer some or all of the cloud resources within that resource. When a first tenant needs to transfer the first cloud resource, it can execute a specified operation on its client to trigger a transfer request. This allows the cloud management platform to sell the first cloud resource based on the transfer request, thus realizing the transfer. The first tenant's transfer request for the first cloud resource may include a second duration, which is less than or equal to the first duration. For example, if the first tenant set the first duration to one month when purchasing the first cloud resource (specifically, from January 1, 2025 to February 1, 2025), then the first tenant can set the second duration to January 20, 2025 to February 25, 2025, within which the first tenant wishes to transfer the first cloud resource to other users. For example, when a first tenant purchases first cloud resources, the first duration is set to one month, specifically from January 1, 2025 to February 1, 2025. The first tenant can set the second duration to 0:00-8:00 every day. Therefore, the second duration is from 0:00-8:00 every day during the period from January 1, 2025 to February 1, 2025. During this second duration, the first tenant wishes to transfer the first cloud resources to other users. That is, the second duration can be divided by day, week, month, or hour, or a combination of these methods. This application does not restrict the specific setting method of the second duration.
[0091] In one possible implementation, the cloud management platform can provide a resource transfer interface to the first tenant, who can then trigger a transfer request for a first cloud resource based on this interface. The transfer request carries information about the cloud resource to be transferred. For example, the transfer request may include at least one of the following: the transfer period set by the first tenant for the first cloud resource, the first cloud resource that can be transferred within the transfer period, and the upper limit of the specifications that the first cloud resource can support. The transfer period is within the purchase period of the first tenant's purchase of the first cloud resource, and the upper limit is less than or equal to the specifications of the first cloud resource purchased by the first tenant. After the first tenant triggers the transfer request, the cloud management platform can obtain the transfer request through the resource transfer interface and sell the first cloud resource based on the transfer request. For example, the resource transfer interface can be implemented through one or more of the following: API and UI. For instance, the cloud management platform maintains a resource management interface for each tenant, which displays information about the cloud resources purchased by the tenant. The first tenant can view the resource management interface of the cloud resources they purchased and perform management operations on those resources within the interface. The resource management interface of the first tenant displays all cloud resources purchased by the first tenant and their related information. The information includes the initial purchase time period, quantity, and specifications of the cloud resources. Each cloud resource has multiple management buttons, each corresponding to a different management function. When the first tenant selects a management button, it triggers the corresponding management operation for that cloud resource. These management functions include resource transfer and resource release. When the first tenant selects the management button corresponding to the transfer function of the first cloud resource, it triggers a transfer operation for that resource. In response to the first tenant's selection of the management button for the transfer function of the first cloud resource, the cloud management platform confirms that the first tenant needs to transfer the first cloud resource and displays the transfer configuration interface to the tenant. For example, as shown in Figure 6, the transfer configuration interface has several prompts. These prompts include prompting the first tenant to select the quantity of cloud resources to be transferred, and prompting the first tenant to select the transfer time. After the first tenant completes the settings for the quantity and transfer time of the cloud resources to be transferred, they can click the submit button in the transfer configuration interface to submit their settings to the cloud management platform. The cloud management platform can then receive the transfer request for the first cloud resource configured by the first tenant, obtaining the first cloud resource that the first tenant needs to transfer, the upper limit of the specifications that the first cloud resource can support, and the transfer period. The first cloud resource to be transferred is indicated by its type and quantity. Depending on the purpose of the cloud resource, the type indicates whether it is a computing resource, storage resource, network resource, or other type of cloud resource.The transfer period is the time during which the first tenant can transfer the first cloud resource to other tenants. Typically, the transfer period falls during the first tenant's off-peak hours. The transfer period can be represented by a transfer time window and a time period. The time period indicates the calendar day to which the transfer time window belongs. For example, a transfer time window indicates that the first cloud resource can be transferred between 06:00 and 12:00. A time period indicates that the first cloud resource can be transferred between 06:00 and 12:00 on any given day or a specific day. The specifications that the first cloud resource can support indicate the performance parameters of the virtual resources that can be created based on the first cloud resource. For example, when the cloud resource is a computing resource, its resource specifications indicate the number of cores of the virtual processors that can be created based on that computing resource. Since the cloud management platform sells cloud resources to tenants according to different specifications, the upper limit of the specifications that a tenant can support for a resource with a specified specification is that specified specification. Accordingly, after the first tenant completes the configuration of the first cloud resource to be transferred in the transfer configuration interface, the specifications of the first cloud resource selected by the first tenant are the upper limit of the specifications that the first cloud resource can support. It should be noted that when setting the transfer quantity, the first tenant must comply with the following restrictions: the maximum transfer quantity is the purchase quantity of the first cloud resource. When setting the transfer time period, the first tenant must comply with the following restrictions: the transfer time period must be within the first purchase time period of the first cloud resource.
[0092] Optionally, the first tenant can also set the transfer price of the first cloud resource. Correspondingly, the transfer request also includes the transfer price set by the first tenant for the first cloud resource. Optionally, if the transfer price is lower than a preset multiple of the predetermined price, the first cloud resource is sold to the first tenant at the predetermined price on the cloud management platform. For example, the first cloud resource can be transferred at a 20% premium, in which case the transfer price is 1.2 times the predetermined price. Another example is that the first cloud resource can be transferred at a 20% discount, in which case the transfer price is 0.8 times the predetermined price. The transfer price can be presented as a unit transfer price. The unit transfer price indicates the fee payable for one unit of first cloud resource used for a unit duration, with the usage specifications being the upper limit of the specifications supported by the first cloud resource. The unit duration can be set according to application requirements. For example, when the transfer period is counted in minutes, the unit duration can be minutes. When the transfer period is counted in hours, the unit duration can be hours. In one implementation, the transfer price can also be set through a resource transfer interface. For example, the transfer configuration interface also includes a prompt for the first tenant to select the transfer price of the cloud resources to be transferred. After the first tenant sets the transfer price of the first cloud resource in this prompt and submits, the cloud management platform can obtain that transfer price. In some implementation scenarios, the cloud management platform can restrict the conditions that the transfer price must meet. For example, when the first tenant transfers the first cloud resource at a premium, the cloud management platform can limit the premium range of the first cloud resource.
[0093] It should be noted that since the first cloud resource can only be transferred after the cloud management platform receives a transfer request for it, the operation of setting the owner information of the first cloud resource as the first tenant in step S102 can also be executed after the cloud management platform receives the transfer request. This embodiment does not specifically limit the timing of its execution. As shown in Figure 4, after receiving the transfer request for the first cloud resource, the cloud management platform creates a resource reservation for the first cloud resource, sets the owner information of the first cloud resource as the first tenant, and sets the user information of the first cloud resource as the first tenant.
[0094] The above functions can be achieved through the collaboration of the front-end and back-end of the cloud management platform. The front-end of the cloud management platform can be components that interact with users, such as application programming interfaces (APIs) or user interfaces. The back-end of the cloud management platform can be components used to implement functions such as resource management and resource scheduling. For example, the above functions are achieved through the collaboration of the user interface, management module, and scheduling module. After the first tenant triggers the cloud management platform to display the transfer configuration interface, the transfer configuration interface prompts the first tenant to select the first cloud resource to be transferred from its purchased cloud resources, set the transfer time period, and the transfer price. After the first tenant completes the settings, the transfer configuration interface provides the management module with a transfer request for the first cloud resource, including the first cloud resource, the transfer time period, and the transfer price. The management module saves the transfer request and sends an instruction to the scheduling module, instructing the resource scheduling module to perform relevant scheduling work for the first tenant to transfer the first cloud resource. For example, the scheduling module creates a resource reservation for the first cloud resource and sets the owner information of the first cloud resource to the first tenant. Setting the owner information of the first cloud resource to the first tenant can be seen as the scheduling module reserving the first cloud resource information to prevent the first cloud resource from being purchased by other tenants.
[0095] It should be noted that transfer requests for primary cloud resources can also be obtained through other means. For example, among the cloud resources it manages, the cloud management platform identifies cloud resources that have been sold but are currently idle. Based on the historical usage of these cloud resources, it determines the transfer period for these resources, or instructs the tenant who purchased the cloud resources to set the transfer period. Then, based on the transfer period and the specifications of the cloud resources, it generates a transfer request for the primary cloud resources and sells them based on this request.
[0096] Step S104: In response to the transfer request for the first cloud resource, publish the first resource transfer information of the first cloud resource to the Internet. The first resource transfer information includes the transfer conditions for the first cloud resource and for tenants other than the first tenant.
[0097] After receiving a transfer request, the cloud management platform generates resource transfer information for the first cloud resource and publishes this information on the internet, enabling tenants to purchase the first cloud resource based on this information. Transfer conditions may include the transfer price set by the first tenant for the first cloud resource. Transfer conditions may also include requiring the other tenant being transferred to register on the cloud management platform. Transfer conditions may include successful payment for the first cloud resource, or successful payment for the first resource transfer information for the first cloud resource.
[0098] Step S105: Obtain the model training task submitted by the second tenant, confirm the cloud resource requirements of the model training task based on the model training task, and confirm that the first cloud resource meets part or all of the cloud resource requirements.
[0099] The second tenant is a tenant who has registered an account on the cloud platform. The second tenant uploads its model training data to cloud storage via the second tenant client. Optionally, this cloud storage is a cloud service provided by the aforementioned cloud management platform, and the cloud storage service is connected to the computing instance via a high-speed network. The second tenant writes training scripts for the model training task and uploads these scripts to the cloud management platform via the second tenant client. Optionally, the second tenant can use pre-built base models, pre-built training data, or pre-built scripts provided by the cloud management platform for model training. The usage of the second tenant client is similar to that of the first tenant client and will not be described further here.
[0100] The model training task can be the task of training a neural network model using a training dataset. The training dataset typically consists of a set of input features and corresponding labels. The model learns patterns and relationships within this data to optimize its parameters, enabling it to predict or classify new data. For example, in an image classification task, the training data might contain tens of thousands of labeled images. This training dataset can include multiple batches of training data. A batch refers to a subset of training data used by the model to calculate gradients and update parameters in a single iteration. For example, if the training dataset has 10,000 samples and the batch size is 32, then each batch contains 32 samples. An iteration refers to the process by which the model performs one forward propagation and one backpropagation on a batch of data. In each iteration, the model calculates the loss function based on the data in the current batch and updates the parameters through backpropagation. A round refers to a complete traversal of the entire training dataset. A round contains multiple iterations, the specific number depending on the batch size and the dataset size. If the training dataset has 10,000 samples and the batch size is 32, then one round requires 313 iterations. The model training task submitted by the second tenant can also include commonly used training parameters, such as epoch, batch size, and data source. The model training task submitted by the second tenant can also include an acceptable bid discount rate or the final resource price.
[0101] In one implementation provided in this application, the second tenant selects the type and quantity of cloud resources based on the cloud resource requirements of its model training task. The model training task submitted by the second tenant includes cloud resource quantity information and / or cloud resource specification information. The cloud resources that the second tenant can select include bare metal servers, virtual machines, containers, etc. For example, the second tenant can select different specifications of CPU instances, NPU instances, GPU instances, etc. Preferably, GPU instances or NPU instances can be used, such as p2s.2xlarge.8 (8 cores 64GB, 1x V100 GPU), pi2.2xlarge.8 (8 cores 64GB, 1x A100 GPU), m6.large.8 (2 cores 16GB), m6.xlarge.16 (4 cores 32GB), h3.2xlarge.8 (8 cores 64GB), etc. The second tenant can select the required number of instances of one or more specifications.
[0102] In other implementations provided in this application, cloud resource requirements may also include operating system image requirements, storage space requirements, storage speed requirements, etc., and may also include other types of cloud resource requirements that may arise after this application. This application does not limit the specific types of cloud resource requirements.
[0103] The cloud management platform obtains the model training task submitted by the second tenant, that is, it obtains the training data and training script submitted by the second tenant, or the training data confirmation information and training script training information sent by the second tenant. Furthermore, the cloud management platform obtains the cloud resource requirement information carried in the model training task submitted by the second tenant, which includes cloud resource quantity information and / or cloud resource specification information.
[0104] In one implementation provided in this application, the second tenant submits corresponding model template information according to its model training task. The model template information includes the identifier of the pre-set model. For example, the cloud management platform has pre-set model A, model B and model C in its storage space, or the cloud management platform has established a communication connection with the storage instance that stores model A, model B and model C so that it can download the corresponding model file according to the storage address of each model.
[0105] When the cloud management platform receives the model training task submitted by the second tenant, it receives the identifier of model A contained in the model template information sent by the second tenant. Based on the mapping relationship between the identifier and the model storage address, the cloud management platform downloads the file of model A to the second tenant's cloud storage for the second tenant to perform subsequent model training tasks. In one embodiment provided in this application, there is a mapping relationship between each model and the cloud resource requirements for training that model. For example, training with model A as the base model requires 4 p2s.2xlarge.8 images, or training with model B as the base model requires 8 m6.large.8 images, and so on. The cloud resource requirements here may include cloud resource specifications and / or cloud resource quantity. In addition to the cloud resources listed above, they may also include storage resources (storage type, storage capacity, storage performance), operating system images, etc. This application does not limit the types of cloud resource requirements.
[0106] Step S105': The cloud management platform confirms that the first cloud resource and the second cloud resource 60 meet part or all of the cloud resource requirements.
[0107] A model training task can be run by a single resource node, or it can be completed by a combination of multiple resource nodes. In this case, the cloud management platform needs to obtain suitable multi-node cloud resources from the cloud resource scheduling unit during the allocation of cloud resources for training, such as multiple Ascend 910B servers under the same RoCE network.
[0108] In one implementation provided in this application, the cloud management platform confirms that the first cloud resources meet part or all of the cloud resource requirements of the second tenant. This includes: the cloud management platform confirming that each cloud resource in the first cloud resources is located in the same parameter plane network. For example, if the cloud resources transferred by the first tenant include 4 cloud resources located in parameter plane network A and 8 cloud resources located in parameter plane network B, and the second tenant needs 3 cloud resources, preferably, the cloud management platform selects the 4 cloud resources located in parameter plane network A as the first cloud resources; alternatively, the cloud management platform may also select the 4 cloud resources located in parameter plane network B as the first cloud resources.
[0109] In one implementation provided in this application, the cloud management platform confirms that the first cloud resources meet part or all of the cloud resource requirements of the second tenant. This includes: the cloud management platform confirms that the first and second cloud resources meet part or all of the cloud resource requirements, and each cloud resource in the first and second cloud resources is located in the same parameter plane network. For example, if the cloud resources transferred by the first tenant include 4 cloud resources located in parameter plane network A and 8 cloud resources located in parameter plane network B, and the second tenant needs 16 cloud resources, preferably, the cloud management platform selects 8 cloud resources located in parameter plane network B as the first cloud resources. In addition, the cloud management platform selects another 8 cloud resources located in parameter plane network B as the second cloud resources. The second cloud resources are not the cloud resources transferred by the first tenant, but the second cloud resources and the first cloud resources are located in the same parameter plane network B.
[0110] In this embodiment, in the first scenario, the second cloud resource can be a cloud resource transferred by a third tenant. The cloud management platform obtains the successful purchase information of the third tenant for the second cloud resource. In response to this successful purchase information, the cloud management platform sets the user of the second cloud resource as the third tenant. The cloud management platform then obtains the transfer request of the third tenant for the second cloud resource. In response to this transfer request, the cloud management platform publishes the second resource transfer information of the second cloud resource to the Internet. This second resource transfer information includes transfer conditions for the second cloud resource and is open to tenants other than the third tenant. In this scenario, the process by which the cloud management platform handles the purchase and transfer of the second cloud resource by the third tenant is similar to the process by which the cloud management platform handles the purchase and transfer of the first cloud resource by the first tenant, and will not be described in detail here.
[0111] In this embodiment, in the second case, the second cloud resource can be a cloud resource provided by the cloud management platform that does not belong to other tenants or is transferred from other tenants. The second cloud resource and the first cloud resource are located in the same parameter plane network B.
[0112] Step S106: Obtain the second purchase success information of the second tenant for the transfer information of the first resource; confirm that the second purchase success information meets the transfer conditions, and set the user of the first cloud resource as the second tenant and the usage time as the second duration.
[0113] After the cloud management platform publishes resource transfer information to the internet, when it receives a purchase request from a tenant indicating the need to buy cloud resources, it can display the resource transfer information to the tenant, allowing the tenant to purchase the transferred cloud resources based on this information. For example, the cloud management platform displays the resource transfer information for the first cloud resource on the cloud resource purchase interface. As shown in Figure 8, the cloud resource purchase interface displays resource sources: existing resources and transferred resources. Existing resources are cloud resources managed by the cloud management platform that have not been purchased by tenants. Transferred resources are cloud resources managed by the cloud management platform that have been purchased by tenants but can be temporarily transferred, such as the first cloud resource. After a second tenant selects the resource source for the transferred resources, the cloud management platform can display the resource transfer information for the transferable cloud resources to the second tenant, allowing the second tenant to choose whether to purchase the transferable cloud resources. Optionally, when displaying the resource transfer information for transferable cloud resources, the cloud management platform can also display the deployment location and owner information of the transferable cloud resources. For example, the cloud resources that can be transferred in the cloud management platform are: c6s.8xlarge.2 belonging to tenant 001, c6s.8xlarge.2 belonging to tenant 002, ac6.8xlarge.2 belonging to tenant 003, and c7.8xlarge.2 belonging to tenant 004. The maximum supported specifications are 8U16G, 4U8G, 4U8G, and 4U8G, respectively. The transfer quantities are 50, 40, 30, and 10, respectively. The deployment location for all of them is the availability zone named Beijing 4. The transfer time periods are: 06:00-12:00 daily, 06:00-12:00 on October 13, 2023, 08:00-16:00 daily, and 22:00-02:00 daily. The transfer prices are: 10 yuan per hour (10 / h), 1 yuan per hour (1 / h), 5 yuan per hour (5 / h) and 5 yuan per hour (5 / h).
[0114] After the cloud management platform displays the transfer request of the first cloud resource to the second tenant, the second tenant can perform a specified operation on its client when it needs to use some or all of the cloud resources in the first cloud resource. This triggers a second purchase success message for the resource transfer information, allowing the cloud management platform to transfer the first cloud resource to the second tenant based on this second purchase success message. The second purchase success message indicates the time period during which the second tenant purchased the first cloud resource and the specifications of the first cloud resource. The specifications of the first cloud resource are less than or equal to the upper limit of the specifications supported by the cloud resource transferred by the first tenant. In one possible implementation, the cloud management platform can provide a resource purchase interface to the tenant, allowing the second tenant to trigger the second purchase success message based on this interface. After the second tenant triggers the second purchase success message, the cloud management platform can obtain the second purchase success message through the resource purchase interface. Optionally, the resource purchase interface can include one or more of the following implementation methods: API and UI.
[0115] After obtaining the second successful purchase information set by the second tenant, the cloud management platform can determine that the second tenant purchased the first cloud resource and its specifications within the purchase period of the first cloud resource. In response to this second successful purchase information, the cloud management platform modifies the user information of the first cloud resource to the second tenant within the purchase period of the first cloud resource, thus instructing the second tenant to have the right to use the first cloud resource during that period. This ensures that the second tenant has the right to use the first cloud resource within the purchase period, and that the first tenant has the right to use the first cloud resource during the time outside the purchase period itself. Alternatively, the cloud management platform can choose to modify the user information of the first cloud resource to the second tenant after obtaining the second successful purchase information, and set this user information to be effective within the purchase period of the first cloud resource. Here, the second tenant can instruct via their account.
[0116] When a transfer request for the first cloud resource also specifies the transfer price of the first cloud resource, the cloud management platform needs to calculate the transfer price for the second tenant's purchase of the third cloud resource based on the second successful purchase information and the transfer price of the first cloud resource. Then, it bills the second tenant for the purchase of the third cloud resource based on this transfer price. For example, when the transfer price is a unit price, the charge for the second tenant's purchase of the third cloud resource can be equal to the product of the unit price, the duration of the transfer period, and the quantity of the third cloud resource. In one possible implementation, when calculating the transfer price for the second tenant's purchase of the third cloud resource, if the third cloud resource is all of the first cloud resource and its specifications are equal to those of the first cloud resource, then the transfer price of the third cloud resource is the same as the transfer price of the first cloud resource. If the third cloud resource is a portion of the first cloud resource, and / or its specifications are smaller than those of the first cloud resource, the cloud management platform also needs to discount the transfer price of the first cloud resource to obtain the transfer price of the third cloud resource. The cloud management platform maintains conversion rules for adjusting the transfer price based on the quantity and specifications of cloud resources. The platform can then execute the conversion process according to these rules. In one possible implementation, the conversion rules indicate the weight of the quantity and specifications of resources on the transfer price. After determining the first proportion of the quantity of the third cloud resource within the quantity of the first cloud resource, and the second proportion of the specifications of the third cloud resource within the specifications of the first cloud resource, the converted transfer price can be obtained based on the first proportion and the weight of the quantity's influence on the transfer price, and the second proportion and the weight of the resource's influence on the transfer price. Optionally, the cloud management platform can also convert the transfer price based on the purchase period of the third cloud resource and the transfer period of the first cloud resource; for this implementation, please refer to the implementation method based on quantity and specifications.
[0117] For example, assume the first cloud resource (ac6.8xlarge.2) supports a maximum specification of 4U8G, the transfer quantity is 30, the transfer period is 08:00-16:00 daily, and the transfer price of the first cloud resource is 5 yuan per hour (5 / h). The second successful purchase message indicates that the quantity of the third cloud resource is 30, the purchase period of the third cloud resource is 08:00-16:00 daily, and the specification of the third cloud resource is 2U4G. The conversion rule indicates that the quantity N1 of the first cloud resource, the quantity N2 of the third cloud resource, the specification S1 of the first cloud resource, the specification S2 of the third cloud resource, the transfer price P1, and the converted transfer price satisfy P2: P2 = P1 × (0.3 × N2 / N1 + 0.7 × S2 / S1). Then the transfer price of the third cloud resource P2 can be obtained as 5 yuan per hour × (0.3 × 30 / 30 + 0.7 × 2U4G / 4U8G) = 3.25 yuan per hour.
[0118] The cloud management platform needs to confirm that the second purchase success information meets the transfer conditions. For example, when the transfer conditions are that the tenant of the cloud resource being transferred registers on the cloud management platform and pays successfully, the cloud management platform needs to confirm that the second tenant that triggered the second purchase success information is a tenant registered on the cloud management platform, and that the second tenant has successfully paid for the first cloud resource it wants to purchase. This payment success information can be obtained by the cloud management platform through the payment information confirmation interface provided by this data center or a third-party data center.
[0119] Step S107: During the second time period, call upon the first cloud resource to participate in the model training task.
[0120] Optionally, when the second tenant needs to perform a model training task based on the first cloud resource, it can execute a specified operation on its client to trigger an execution request for the model training task. This allows the cloud management platform to invoke the first cloud resource for the second tenant under the instruction of the execution request, deploy the model training task on the first cloud resource, and execute the model training task. In one possible implementation, the cloud management platform can provide a training task execution interface to the second tenant, which the tenant can use to trigger an execution request. In another possible implementation, since the second tenant has already submitted the model training task, the second tenant can default to the cloud management platform directly invoking the first cloud resource to participate in the model training task. For example, the cloud management platform can configure a suitable operating system and deep learning framework for the first cloud resource, install necessary dependency libraries for the first cloud resource, deploy the model training task on the first cloud resource, and then start model training on the first cloud resource. This application does not limit the specific stage at which the cloud management platform invokes the first cloud resource to participate in the model training task.
[0121] Step S107': During the second time period, call upon the first cloud resources and the second cloud resources to participate in the model training task.
[0122] When the cloud management platform confirms that the first and second cloud resources meet some or all of the cloud resource requirements, the platform calls upon the first and second cloud resources to participate in the second tenant's model training task within a second time period. For example, the second tenant's model training task is broken down into multiple sub-tasks, which are distributed to various instances of the first and second cloud resources for parallel execution. These nodes synchronize model parameters through network communication, ultimately completing the entire training task. This model training task can be data-parallel training or model-parallel training; this application does not restrict the method of distributed model training. Since both the first and second cloud resources are located on the same parameter plane network, nodes performing distributed model training can achieve efficient parameter synchronization and communication, significantly improving the efficiency of distributed training.
[0123] Step S108: In response to the resource interruption request, the cloud management platform saves the checkpoint file of the neural network model after the training of the current batch of training data is completed.
[0124] In some cases, the cloud management platform may receive a resource interruption request for the primary cloud resource.
[0125] For example, the second duration set by the first tenant in the first cloud resource includes a start time and / or an end time. When the end time arrives or is about to arrive (for example, there is a preset time before the end time arrives, such as five minutes), the instance or a component related to the instance (for example, when the instance is a virtual machine, the component can be a virtual machine manager) will send a resource interruption request to the cloud management platform. The resource interruption request can carry the specific interruption time of the first cloud resource.
[0126] For example, the first cloud resource is a auction resource, and its price changes in real time based on market supply and demand. Second tenants can purchase these instances at a discounted price, but the cloud management platform may automatically reclaim these instances based on resource inventory or market price changes. When the market price exceeds the second tenant's bid (e.g., another tenant offers a higher price to purchase these transferred first cloud resources), a resource interruption request will be triggered. Alternatively, a resource interruption request will be triggered when cloud resource inventory is insufficient. Similarly, the resource interruption request can carry the specific interruption time for the first cloud resource.
[0127] As mentioned above, the training dataset used by the second tenant for model training includes multiple batches of training data. In response to a resource interruption request, the cloud management platform can pause the training of the neural network model after the current batch of training data has been trained, and save the checkpoint file of the neural network model. The current batch of training data can be the training data of the batch that was being trained when the cloud management platform received the resource interruption request, or the training data of the batch that was being trained at the time the resource interruption request was issued, or the training data of the batch that was being trained within a certain timeframe before or after the aforementioned times. The training time of the current batch is related to the processing time of the resource interruption request; however, the specific relationship is not limited in this embodiment.
[0128] Since the start and pause of the model training task for the second tenant on the cloud platform can be controlled by the cloud management platform, and the cloud management platform can determine when the cloud resources in the process of model training will be reclaimed or released by receiving resource interruption requests, the cloud management platform can save the checkpoint file of the neural network model after the training data of the current batch is completed, that is, after the forward training and backpropagation of the current batch are completed, or after the gradient calculation of the current batch is completed. This ensures that the gradient calculation and parameter update of the current batch are completed, avoids the waste of resources caused by the need to retrain the batch in the future due to midway stop, avoids the inconsistency of training state caused by midway stop, and even avoids the problems of loss of training state, model corruption or incorrect saving that may be caused by violent interruption.
[0129] In one embodiment, after the model training task is started, during the training of each batch of training data, checkpoint files can be saved periodically. The period can be determined by time or by the number of batches executed. If the cloud resources being used need to be reclaimed or released, the task status will also be saved and the task will be suspended.
[0130] After training the current batch of training data is completed, the cloud management platform pauses model training, synchronizes parameters and optimizer states across all nodes of the first cloud resource (or both the first and second cloud resources), and serializes and saves the model's current weights, biases, and other parameters to a file. Optionally, the optimizer's internal state should also be saved so that the optimization process can continue when training resumes. Information such as the current epoch number, iteration number, loss value, and accuracy is also saved. Finally, the checkpoint file is saved in cloud storage, preferably the cloud storage of the second tenant.
[0131] Step S109: The cloud management platform releases the first cloud resource and modifies the user information of the first cloud resource to the first tenant.
[0132] In one of the aforementioned embodiments, when the cloud management platform also calls upon a second cloud resource located on the same parameter plane network as the first cloud resource to participate in the model training task: First, if the second cloud resource was transferred from a third tenant, the cloud management platform releases the second cloud resource and modifies its user information to that of the third tenant. Second, if the second cloud resource was transferred from a third tenant and the second tenant's usage period for the second cloud resource has not yet ended, the second cloud resource is not released, and the cloud management platform continues to match other cloud resources that, along with the second cloud resource, meet the second tenant's cloud resource requirements. Third, if the second cloud resource is provided by the cloud management platform and does not belong to or was transferred from any other tenant, the cloud management platform releases the second cloud resource. Fourth, the second tenant can continue to use the second cloud resource for a certain period of time, and the cloud management platform continues to match other cloud resources that, along with the second cloud resource, meet the second tenant's cloud resource requirements. The cloud management platform can query whether there are matching cloud resources through the cloud resource scheduling interface. The query conditions can include the resource specifications and type, the bidding price of the resource, the time window for the resource to participate in the bidding sale, the quantity of the resource, and other conditions.
[0133] Step S109': The cloud management platform releases the first cloud resource and modifies the user information of the first cloud resource to the fourth tenant.
[0134] As mentioned earlier, the first cloud resource can be a bid resource, and its price changes in real time according to market supply and demand. The cloud platform may automatically reclaim instances based on market price changes. When the market price exceeds the bid of the second tenant (for example, if another tenant bids a higher price to purchase these transferred first cloud resources), a resource interruption request will be triggered. For instance, if a fourth tenant bids a higher price for the use of the first cloud resource, the first cloud resource will be released under certain conditions (for example, after a certain period of time after the fourth tenant has successfully bid and paid). The cloud management platform will then modify the user information of the first cloud resource to the fourth tenant so that the fourth tenant can use the first cloud resource to deploy its services. This embodiment includes methods related to the handling of bid resources, which will not be elaborated here.
[0135] Step S110: The cloud management platform calls upon the third cloud resources of the second tenant to participate in the model training task.
[0136] As previously mentioned, the second tenant's model training task on the first cloud resource (or the first and second cloud resources) is paused, and the checkpoint file is saved. In one embodiment provided in this application, the cloud management platform continues to match cloud resources that meet the second tenant's cloud resource requirements, and a third cloud resource that meets the second tenant's cloud resource requirements is matched. For example, the third cloud resource can meet the tenant's cloud resource requirements; alternatively, the third cloud resource can meet the tenant's cloud resource requirements together with the aforementioned second cloud resource; or alternatively, the third cloud resource can meet the tenant's cloud resource requirements together with another fourth cloud resource. The conditions for meeting cloud resource requirements have been described above and will not be repeated here. In another embodiment provided in this application, the third cloud resource can be a cloud resource purchased by the second tenant, not transferred from other tenants, and not subject to bidding; that is, the user information of the third cloud resource is the second tenant themselves, and the second tenant purchased the third cloud resource on a per-month or on-demand basis without worrying about the third cloud resource being released. In these embodiments, the cloud management platform modifies the user of the third cloud resource to the second tenant.
[0137] The cloud management platform invokes the second tenant's third cloud resource to participate in the model training task. During this process, the cloud management platform can load checkpoint files previously saved in cloud storage onto this third cloud resource. In previous training sessions, the second tenant's model training task paused after completing the current batch of training data, and the checkpoint file was saved. While the cloud management platform is invoking the second tenant's third cloud resource for model training, it loads the checkpoint file onto the third cloud resource, reads the model state from the checkpoint file (such as model parameters, optimizer state, training progress, and other metadata), and resumes training from the point of interruption—that is, from the next batch of training data in the aforementioned current batch—based on the training progress saved in the checkpoint file. Through these steps, the training state can be safely restored from the checkpoint file, ensuring a smooth recovery of the second tenant's model training task.
[0138] In one embodiment provided in this application, the final model parameter file will be saved after training. Tenants can view or download the trained model parameter file through the tenant client interface. Optionally, during model training or when the model training process is paused, tenants can also view or download the current model training status information through the tenant client interface.
[0139] In one embodiment provided in this application, the model training task includes a model loading phase and at least one round of model training. The cloud management platform bills the second tenant using a time-based billing method. In response to the end of the model loading phase or the start of at least one round of model training, the cloud management platform updates the billing for the second tenant's purchase of the first cloud resources. During model training, training data and model files need to be prepared, and code is written to load the training data or model files into memory using batch processing or a data pipeline. When the model loading phase ends, or when a round of model training (preferably the first round of model training) begins, the cloud management platform updates the billing for the second tenant's purchase of the first cloud resources. For example, in a single model training session, the cloud management platform only bills for the amount of resources used or the duration of resource usage during the iterative training process, while stages such as training program startup and data saving may choose not to trigger billing for the amount of resources used or the duration of resource usage. The training control logic in the cloud management platform can accurately monitor the training process, thereby enabling precise billing for cloud resource usage.
[0140] Please refer to Figure 5, which is a schematic diagram of a cloud resource management device based on public cloud technology provided in an embodiment of this application. The device includes a cloud resource scheduling unit, a bidding training management unit, a storage unit, resource node 1, and resource node 2. Each unit or module included in the device will be described below.
[0141] Reverse Selling Unit: Users holding resources for a specified period can sell these resources through the reverse selling unit provided by the cloud system. This selling is typically based on a time window; for example, after midnight when business is relatively idle, a portion of resource nodes can be allocated from the existing business resource pool for reverse selling, such as for a 10-hour period. These resource nodes are purchased from the cloud service provider, such as an Ascend server, an Nvidia GPU server, a virtual machine, or a container. Users can notify the reverse selling unit to execute the reverse selling through REST API calls or via the cloud service console / console (web). After receiving the reverse selling request, the reverse selling unit can notify the cloud's resource scheduling unit to temporarily release the resources from the tenant. In one embodiment provided in this application, the reverse selling time window does not guarantee immediate economic returns; the profitability of the resources will depend on subsequent bidding and sales.
[0142] Cloud Resource Scheduling Unit: This is a business capability unit currently possessed by all cloud service providers. It schedules basic resources, such as allocating virtual machines, physical machines, and containers to users. The bidding training management unit can query whether there are matching resources through the interface provided by this unit. The query conditions can include resource specifications, bidding price, time window for bidding, quantity, etc. In one embodiment provided in this application, the bidding query and allocation function is provided by the cloud resource scheduling unit. In other embodiments provided in this application, this feature, or features related to reverse reselling, are not provided by the cloud resource scheduling unit, but are supported by the bidding training management unit. This application does not limit which unit provides this feature.
[0143] The bidding training management unit is the main management unit responsible for bidding training. It receives bidding training tasks submitted by users with training needs and supports cooperation with the cloud resource scheduling unit to obtain bidding computing resources. During the resource usage process of bidding training, bills are generated. These bills can be generated by the bidding training management unit, or, in principle, by the cloud resource scheduling unit. The billing information can be notified to the reverse resale unit for processing, handled by the cloud resource scheduling unit, or handled by a dedicated unit responsible for user billing. In one embodiment provided in this application, bills are required to be generated for the resources used during bidding training. The advantage of having the bidding training unit manage the bills is that the training control logic within the bidding training management unit can more accurately understand the training process, thus allowing for more refined billing from the actual business perspective of resource usage (e.g., time spent on program startup and data storage may not be billed).
[0144] Resource nodes: Resource nodes can be physical machines (generally called bare metal servers in public clouds), virtual machines, or containers—systems capable of supporting computing power. Training tasks will run on resource nodes. A training task can be run on a single resource node, or multiple training tasks can be run on a single resource node. In one embodiment provided in this application, isolation technology is also needed to isolate the resources of each training task, allowing users to more accurately pay for computing power. In one embodiment provided in this application, a training task requires a combination of multiple resource nodes to complete. In this case, the bidding training management unit needs to obtain suitable multi-node resources from the cloud resource scheduling unit during the resource allocation training process, such as two Ascend 910B servers under the same ROCE network. Currently, mainstream training platforms have the capability for single-node or multi-node joint training, which will not be elaborated upon here.
[0145] Storage unit: It can store checkpoints during model training, model parameter files, and information about resources sold by the reverse selling unit. The storage unit can be composed of object storage and data. The storage unit here is not a software system, but may be composed of multiple running software systems.
[0146] Please refer to Figure 6, which is a flowchart of a cloud resource management method based on public cloud technology provided in an embodiment of this application.
[0147] The training process will be linked to the effective time window of the auctioned resources. Before the effective time is reached, the resources need to be returned to the original tenant and the previous configuration needs to be maintained.
[0148] During the model training process, each batch of training can be synchronized with the bidding training management unit upon completion. If the resource needs to be reclaimed, the training can be stopped.
[0149] General training process:
[0150] For iteration count (how many times to tune the entire dataset):
[0151] For each batch of data, read from the dataset:
[0152] One parameter tuning
[0153] If certain conditions are met:
[0154] Save the model parameters as a checkpoint (so that the system can be restored in case of system interruption, resource reclamation, etc.).
[0155] Training ended:
[0156] Save the final trained model parameter file
[0157] After the training task starts, parameter tuning will be performed in each batch. Checkpoints (ckpt) will be saved periodically. The period can be determined by time or by the number of batches executed. If resource reclamation occurs, the task status will be saved and the task will be suspended. The resources of that node will be reclaimed. If the bidding training unit matches the appropriate resources, the training task will resume. If the training is completed, the final model parameter file will also be saved.
[0158] Please refer to Figure 7, which is a flowchart of a cloud resource management method based on public cloud technology provided in an embodiment of this application. As shown in Figure 7, the flowchart of a cloud resource management method based on public cloud technology is as follows:
[0159] Resource holders attempt to monetize their idle computing power through a reverse selling unit via the cloud's auction-based training management unit. This can be done through an interface provided by the reverse selling unit. The interface requires parameters, including resource identifiers (e.g., virtual machine or bare metal IDs) and a time window for the sale (e.g., starting the reverse sale now and lasting two days). However, as mentioned above, starting the reverse sale doesn't guarantee immediate economic value; actual revenue depends on the bidding results. The parameters can also include a discount rate or a discount range, such as 30% to 50% off.
[0160] Upon receiving a reverse sale request, the reverse sale unit needs to record the node's previous configuration information, such as the user it belongs to, network VPC configuration, IP configuration, and configurations for different resource types. For example, bare metal services need to retain the previous operating system data source (typically, servers mount elastic block storage hard drives to boot the operating system), and virtual machines also need to record the previously mounted system disks and data disks. This information can be recorded by the reverse sale unit or the cloud resource scheduling unit. The reverse sale unit then notifies the cloud resource scheduling unit that the resource will be included in the auction management.
[0161] Optionally, the bidding training management unit can receive notifications from the reverse selling unit. This is an optional process. The bidding training management unit can also proactively query available bidding training resources from the cloud resource scheduling unit.
[0162] The auction training management unit initializes the resources for auction training, such as preparing the operating system on the physical machine or virtual machine, preparing the software to run, or preparing the container to run the software. This process may not occur after step 3, but may also occur after the training task is generated. Of course, having this process occur earlier can improve the efficiency of subsequent processing. This process may also occur in multiple sub-processes, because preparing the operating system and preparing the software or container image that the training software depends on can be done separately. Since different training tasks have different requirements for training software or containers, some sub-processes can occur after the training task is generated.
[0163] Mirror training requires users to configure a training task. This task needs to include commonly used training parameters, such as epoch, batch size, and data source, as well as an acceptable bid discount rate or the final price of the resources. Of course, there may be other ways to express this, such as using a certain number of training units, with the price per training unit being 50% of the original price. In short, this is a discount on the original resource unit of measurement.
[0164] The bidding training unit first saves the bidding training task information submitted by the user.
[0165] The bidding training unit queries whether there are suitable resources available for training. This query can be performed from its own stored state information or from the cloud resource scheduling unit (depending on the implementation, the bidding training unit and the cloud resource scheduling unit have flexible functional divisions, and they can even be implemented in the same system). If there are no suitable resources at this time, it can query again after a certain period of time, or it can receive a report once suitable resources are available (by subscribing to suitable resources to receive reports).
[0166] Once suitable resources are available, connect to the resource node (which can be one or more resource nodes) to perform the bidding training task.
[0167] Starting a training task requires obtaining information to support the training, including some parameter information submitted by the user (this information can also be carried when the task is executed, instead of being queried from the storage unit), as well as the model and data to be used for training (the acquisition of data does not necessarily have to be a one-time event; the required data can be acquired during training).
[0168] Starting a training task also requires incurring billing fees.
[0169] During training, information needs to be synchronized with the bidding training management unit (optional) to utilize billing information such as the current training status, progress, errors, and faults.
[0170] The training process saves training state information, such as the current checkpoint of the training model, training progress information, and other information.
[0171] A periodic save will also typically occur in the bidding training management unit (optional).
[0172] The bidding training management unit needs to analyze the current resource status, including the bidding time window for resources. For example, if a resource currently in use is about to be released (this "about to" is related to the training process, such as releasing resources after a batch parameter tuning), the corresponding resource node needs to be notified to pause the current training and save the state and release the resources.
[0173] The resource nodes are notified to pause training, save their state, and release resources.
[0174] After a parameter tuning, the bidding training task on the resource node will pause training (an optional implementation is a pause issued by the bidding training management unit, which is a future pause, such as the end of the batch in 5 minutes as the trigger for training pause). The training pause time will be recorded (for subsequent billing), and the training state will be saved to the storage unit. The training state includes the model parameter file and the current training progress information of the model.
[0175] The above-mentioned state is saved to the storage unit.
[0176] The notification is sent to the bidding training management unit, along with training task status information (which may include the current training progress), billing information, etc.
[0177] The bidding training management unit notifies the cloud resource scheduling unit to return resources to the original reverse-selling tenant.
[0178] Resources are reconfigured to the reverse sale tenant, which requires the use of previously saved resource node information for recovery.
[0179] Complete the relevant recovery configuration.
[0180] The bidding training management unit continues to match appropriate training resources for bidding training tasks.
[0181] After configuring new resources, training can continue using the new resources in the previous mode. The resources described earlier can be a single node or a cluster of multiple nodes. For example, a cluster of multiple Ascend 910b nodes.
[0182] Once the training task is completed, record the timing and save the training status.
[0183] Save the training state to the storage unit. The training state includes the model parameter file.
[0184] The bidding training management unit is notified that the training task has ended, and this also includes billing and tracking information.
[0185] Collect overall billing information and mark the resource as idle.
[0186] If idle resources remain unmatched by any task, they need to be returned. Alternatively, this process could involve returning the resources to the cloud resource scheduling unit, which would then match them with a task to trigger their use.
[0187] Users of auction-based training can query the status of their auction-based training tasks. The status can include "Training in Progress," "Completed," "Stopped by Error," "Stopped by Insufficient Funds," and "Waiting in the Queue." They can also query other task information, such as queue time, training time, and estimated completion time. If the task is complete, users can view the training output (which may include model parameter files, test set reports, etc.).
[0188] The training user retrieves the training output from the storage unit.
[0189] The cloud resource management apparatus and method provided in this application embodiment allow users to sell their purchased subscription resources through a reverse resale unit, enabling them to monetize resources during off-peak hours without affecting their business operations. Auction training allows users to complete necessary training tasks at a lower cost, saving resources. The cloud platform can better utilize resources (unsold resources on the cloud platform itself can be used for auction training, and the auction and training platform can be combined to create smooth auction interruptions).
[0190] In other embodiments provided in this application, the cloud resource management device and method provided in this application can also be applied to inference scenarios, especially offline inference scenarios. Other scenarios, such as offline rendering, are also applicable. In these cases, the types of computing power consumption will be different. For example, rendering requires graphics GPU computing power or CPU computing power.
[0191] Furthermore, the order of steps in the cloud resource management method based on public cloud technology provided in this application embodiment can be appropriately adjusted, and steps can also be added or removed as needed. Any variations that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application, and therefore will not be elaborated further.
[0192] The following describes an example of a virtual device in an embodiment of this application.
[0193] The above describes a cloud resource management method based on public cloud technology according to embodiments of this application. Corresponding to the above method, embodiments of this application also provide a cloud resource management device based on public cloud technology. Figure 8 is a schematic diagram of the structure of a cloud resource management device based on public cloud technology provided in an embodiment of this application. Based on the following components shown in Figure 8, the cloud resource management device based on public cloud technology shown in Figure 8 can perform all or part of the operations shown in Figure 3 above. It should be understood that the device may include more additional components than the components shown or omit some of the components shown, and embodiments of this application do not limit this. Optionally, the cloud resource management device based on public cloud technology is deployed on a cloud management platform. The cloud management platform is used to manage the infrastructure that provides cloud services. The infrastructure includes multiple servers. The multiple servers are used to deploy virtual instances that implement tenant services. As shown in Figure 8, the cloud resource management device 120 based on public cloud technology includes:
[0194] The interaction module 1201 is used to obtain the first successful purchase information of the first tenant for the first cloud resource in the infrastructure.
[0195] The management module 1202 is used to set the user of the first cloud resource as the first tenant and the usage time as the first duration in response to the first purchase success information.
[0196] The interaction module 1201 is also used to obtain the transfer request of the first tenant for the first cloud resource, the transfer request including a second duration, the second duration being less than or equal to the first duration.
[0197] The management module 1202 is also configured to respond to the transfer request for the first cloud resource by publishing the first resource transfer information of the first cloud resource to the Internet, wherein the first resource transfer information includes transfer conditions for the first cloud resource and for tenants other than the first tenant.
[0198] The interaction module 1201 is also used to obtain the model training task submitted by the second tenant.
[0199] The management module 1202 is further configured to confirm the cloud resource requirements of the model training task based on the model training task, and confirm that the first cloud resource meets part or all of the cloud resource requirements.
[0200] The interaction module 1201 is also used to obtain the second purchase success information of the second tenant regarding the first resource transfer information.
[0201] The management module 1202 is also used to confirm that the second purchase success information meets the transfer conditions, and to set the user of the first cloud resource as the second tenant and the usage time as the second duration.
[0202] The management module 1202 is also used to call the first cloud resources to participate in the model training task during the second time period.
[0203] In one possible implementation, the transfer request also includes a transfer price set by the first tenant for the first cloud resource, which is lower than a preset multiple of the predetermined price, and the first cloud resource is sold to the first tenant on the cloud management platform at the predetermined price.
[0204] In one possible implementation, the management module 1202 is also used to confirm that each of the first cloud resources in the first cloud resources is located in the same parameter plane network.
[0205] In one possible implementation, the management module 1202 is further configured to confirm that the second cloud resources in the first cloud resources and infrastructure meet part or all of the cloud resource requirements, and that each first cloud resource in the first cloud resources and each second cloud resource in the second cloud resources are located in the same parametric plane network; and to call the second cloud resources to participate in the model training task within a second duration.
[0206] In one possible implementation, the interaction module 1201 is used to: obtain third purchase success information of the third tenant for the second cloud resource; the management module 1202 is used to: respond to the third purchase success information and set the user of the second cloud resource as the third tenant; the interaction module 1201 is used to: obtain the transfer request of the third tenant for the second cloud resource; the management module 1202 is used to: respond to the transfer request for the second cloud resource and publish the second resource transfer information of the second cloud resource to the Internet, the second resource transfer information including the transfer conditions for the second cloud resource and for tenants other than the third tenant.
[0207] In one possible implementation, the model training task uses a training dataset to train the neural network model. The training dataset includes multiple batches of training data. After the first cloud resource is invoked to participate in the model training task within the second duration, the management module 1202 is used to: in response to a resource interruption request for the first cloud resource, after the training of the current batch of training data is completed, save the checkpoint file of the neural network model; and release the first cloud resource.
[0208] In one possible implementation, the resource interruption request is issued based on the end time of the second duration, and the management module 1202 is used to: change the user of the first cloud resource to the first tenant of the cloud management platform.
[0209] In one possible implementation, the model training task includes a model loading phase and at least one round of model training. The cloud management platform bills the second tenant using a time-based billing method. The management module 1202 is used to update the billing for the second tenant's purchase of the first cloud resources in response to the end of the model loading phase or the start of at least one round of model training.
[0210] In one possible implementation, the resource demand information includes the first bid of the second tenant for the first cloud resource, and the resource interruption request is issued based on the second bid of the fourth tenant for the first cloud resource being higher than the first bid. The interaction module 1201 is used to: obtain the fourth purchase success information of the fourth tenant for the transfer information of the first resource; the management module 1202 is used to: confirm that the fourth purchase success information meets the transfer conditions, and set the user of the first cloud resource as the fourth tenant.
[0211] In one possible implementation, after the current batch of training data has been trained and the checkpoint file of the neural network model has been saved in response to a resource interruption request, the management module 1202 is used to: deploy the model training task on the third cloud resource of the second tenant; load the checkpoint file and train the next batch of training data on the third cloud resource.
[0212] In one possible implementation, the model training task submitted by the second tenant includes at least one of the cloud resource quantity and cloud resource specifications in the cloud resource requirements. The interaction module 1201 is used to obtain the model training task submitted by the second tenant, and the management module 1202 is used to confirm the cloud resource requirements of the model training task based on at least one of the cloud resource quantity and cloud resource specifications.
[0213] In one possible implementation, the model training task submitted by the second tenant includes model template information, which includes the identifier of a pre-set model. The interaction module 1201 is used to obtain the model training task submitted by the second tenant, and the management module 1202 is used to confirm the cloud resource requirements of the model training task based on the identifier of the pre-set model in the model template information.
[0214] In one possible implementation, cloud resources include at least one of virtual machines, containers, bare metal servers, and physical machines.
[0215] In one possible implementation, the transfer conditions include registering and successfully paying for the first cloud resource on the cloud management platform, and other tenants including tenants who have registered and successfully paid for the first cloud resource on the cloud management platform.
[0216] Both the interaction module 1201 and the management module 1202 can be implemented in software or in hardware. For example, the implementation of the interaction module 1201 will be described below. Similarly, the implementation of the management module 1202 can refer to the implementation of the interaction module 1201.
[0217] As an example of a software functional unit, the interaction module 1201 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the aforementioned computing instance may be one or more. For example, the interaction module 1201 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one cloud data center or multiple geographically proximate cloud data centers. Typically, a region may include multiple AZs.
[0218] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0219] As an example of a hardware functional unit, the interaction module 1201 may include at least one computing device, such as a server. Alternatively, the interaction module 1201 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0220] The multiple computing devices included in the interaction module 1201 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the interaction module 1201 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the interaction module 1201 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0221] It should be noted that, in other embodiments, either the interaction module 1201 or the management module 1202 can be used to execute any step in the cloud resource management method based on public cloud technology. The steps implemented by the interaction module 1201 and the management module 1202 can be specified as needed. By implementing different steps in the cloud resource management method based on public cloud technology through the interaction module 1201 and the management module 1202 respectively, all functions of the cloud resource management device based on public cloud technology can be realized.
[0222] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of each component described above can be referred to the corresponding content in the foregoing method embodiments, and will not be repeated here.
[0223] The following provides examples illustrating the basic hardware structures involved in the embodiments of this application.
[0224] This application also provides a computing device 1300. As shown in FIG9, the computing device 1300 includes: a bus 1302, a processor 1304, a memory 1306, and a communication interface 1308. The processor 1304, the memory 1306, and the communication interface 1308 communicate with each other via the bus 1302. The computing device 1300 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1300.
[0225] Bus 1302 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 9, but this does not imply that there is only one bus or one type of bus. Bus 1302 can include pathways for transmitting information between various components of computing device 1300 (e.g., memory 1306, processor 1304, communication interface 1308).
[0226] The processor 1304 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0227] The memory 1306 may include volatile memory, such as random access memory (RAM). The processor 1304 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0228] The memory 1306 stores executable program code, and the processor 1304 executes the executable program code to implement the functions of the aforementioned interaction module 1201 and management module 1202, thereby realizing a cloud resource management method based on public cloud technology. That is, the memory 1306 stores instructions for executing the cloud resource management method based on public cloud technology.
[0229] The communication interface 1308 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1300 and other devices or communication networks.
[0230] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0231] As shown in Figure 10, the computing device cluster includes at least one computing device 1300. The memory 1306 of one or more computing devices 1300 in the computing device cluster may store the same instructions for executing cloud resource management methods based on public cloud technology.
[0232] In some possible implementations, the memory 1306 of one or more computing devices 1300 in the computing device cluster may also store partial instructions for executing cloud resource management methods based on public cloud technology. In other words, a combination of one or more computing devices 1300 can jointly execute instructions for executing cloud resource management methods based on public cloud technology.
[0233] It should be noted that the memory 1306 in different computing devices 1300 within the computing device cluster can store different instructions, which are used to execute certain functions of the cloud resource management device based on public cloud technology. That is, the instructions stored in the memory 1306 of different computing devices 1300 can implement the functions of one or more modules in the interaction module 1201 and the management module 1202.
[0234] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 11 illustrates one possible implementation. As shown in Figure 11, two computing devices 1300A and 1300B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 1306 in computing device 1300A stores instructions for executing the functions of the interaction module 1201. Simultaneously, the memory 1306 in computing device 1300B stores instructions for executing the functions of the management module 1202.
[0235] The connection method between the computing device clusters shown in Figure 11 can be considered as follows: taking into account that the cloud resource management method based on public cloud technology provided in this application requires a large amount of data storage, the function implemented by the management module 1202 is to be executed by the computing device 1300B.
[0236] It should be understood that the functions of computing device 1300A shown in Figure 11 can also be performed by multiple computing devices 1300. Similarly, the functions of computing device 1300B can also be performed by multiple computing devices 1300.
[0237] This application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to the connection method of the computing device clusters in Figures 10 and 11. The difference is that the memory 1306 in one or more computing devices 1300 in this computing device cluster can store the same instructions for executing cloud resource management methods based on public cloud technology.
[0238] In some possible implementations, the memory 1306 of one or more computing devices 1300 in the computing device cluster may also store partial instructions for executing cloud resource management methods based on public cloud technology. In other words, a combination of one or more computing devices 1300 can jointly execute instructions for executing cloud resource management methods based on public cloud technology.
[0239] This application also provides a computer program product containing instructions. The computer program product may be software or program products containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product runs on at least one computing device, it causes the at least one computing device to perform a method for managing cloud resources based on public cloud technology.
[0240] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform a cloud resource management method based on public cloud technology, or instruct the computing device to perform a cloud resource management method based on public cloud technology.
[0241] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0242] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the raw data and executable code involved in this application were obtained with full authorization.
[0243] In the embodiments of this application, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. The term "at least one" refers to one or more, and the term "multiple" refers to two or more, unless otherwise expressly defined.
[0244] In this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0245] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the concept and principles of this application should be included within the protection scope of this application.
Claims
1. A cloud resource management method based on public cloud technology, characterized in that, The method is applied to a cloud management platform, which manages infrastructure providing cloud services. The infrastructure includes multiple servers, each of which has cloud resources configured. The method includes: The cloud management platform obtains the first successful purchase information of the first tenant for the first cloud resource in the infrastructure; In response to the first purchase success information, the cloud management platform sets the user of the first cloud resource as the first tenant and the usage time as the first duration. The cloud management platform obtains the first tenant's transfer request for the first cloud resource, and the transfer request includes a second duration, the second duration being less than or equal to the first duration; The cloud management platform responds to the transfer request for the first cloud resource by publishing the first resource transfer information of the first cloud resource to the Internet. The first resource transfer information includes transfer conditions for the first cloud resource and for tenants other than the first tenant. The cloud management platform obtains the model training task submitted by the second tenant, confirms the cloud resource requirements of the model training task based on the model training task, and confirms that the first cloud resources meet part or all of the cloud resource requirements. The cloud management platform obtains the second purchase success information of the second tenant regarding the transfer information of the first resource; confirms that the second purchase success information meets the transfer conditions, and sets the user of the first cloud resource as the second tenant and the usage time as the second duration; The cloud management platform calls upon the first cloud resources to participate in the model training task within the second time period.
2. The method as described in claim 1, characterized in that, The cloud management platform confirms that the first cloud resource meets some or all of the cloud resource requirements, including: The cloud management platform confirms that each of the first cloud resources is located in the same parameter plane network.
3. The method as described in claim 1 or 2, characterized in that, The cloud management platform confirms that the first cloud resource meets some or all of the cloud resource requirements, including: The cloud management platform confirms that the first cloud resource and the second cloud resource in the infrastructure meet part or all of the cloud resource requirements, and each first cloud resource in the first cloud resource and each second cloud resource in the second cloud resource are located in the same parameter plane network. The method further includes: The cloud management platform calls upon the second cloud resources to participate in the model training task within the second time period.
4. The method as described in claim 3, characterized in that, The method further includes: The cloud management platform obtains the third successful purchase information of the third tenant for the second cloud resource; In response to the third purchase success information, the cloud management platform sets the user of the second cloud resource as the third tenant; The cloud management platform obtains the transfer request from the third tenant for the second cloud resource; In response to the transfer request for the second cloud resource, the cloud management platform publishes the second resource transfer information of the second cloud resource to the Internet. The second resource transfer information includes transfer conditions for the second cloud resource and for tenants other than the third tenant.
5. The method according to any one of claims 1 to 4, characterized in that, The model training task uses a training dataset to train the neural network model. The training dataset includes multiple batches of training data. After the cloud management platform calls the first cloud resource to participate in the model training task within the second time period, the method further includes: In response to a resource interruption request for the first cloud resource, the cloud management platform saves the checkpoint file of the neural network model after the training of the current batch of training data is completed. The cloud management platform releases the first cloud resource.
6. The method as described in claim 5, characterized in that, The resource interruption request is issued based on the end time of the second duration, and the method further includes: The cloud management platform changes the user of the first cloud resource to the first tenant.
7. The method according to any one of claims 1 to 6, characterized in that, The model training task includes a model loading phase and at least one round of model training. The cloud management platform bills the second tenant using a time-based billing method. The method further includes: In response to the end of the model loading phase or the start of the at least one round of model training phase, the cloud management platform updates the billing for the second tenant's purchase of the first cloud resources.
8. The method according to any one of claims 1 to 7, characterized in that, The resource demand information includes the second tenant's first bid for the first cloud resource, and the resource interruption request is issued based on the fourth tenant's second bid for the first cloud resource being higher than the first bid. The method further includes: The cloud management platform obtains the fourth purchase success information of the fourth tenant regarding the transfer information of the first resource; confirms that the fourth purchase success information meets the transfer conditions, and sets the user of the first cloud resource as the fourth tenant.
9. The method according to any one of claims 5 to 8, characterized in that, After the cloud management platform, in response to a resource interruption request, saves the checkpoint file of the neural network model after training the current batch of training data is completed, the method further includes: The cloud management platform deploys the model training task on the third cloud resource of the second tenant; The cloud management platform loads the checkpoint file and trains the next batch of training data for the current batch on the third cloud resource.
10. The method according to any one of claims 1 to 9, characterized in that, The model training task submitted by the second tenant includes at least one of the cloud resource quantity and cloud resource specifications in the cloud resource requirements. The cloud management platform obtains the model training task submitted by the second tenant and confirms the cloud resource requirements of the model training task based on the model training task, including: The cloud management platform obtains the model training task submitted by the second tenant and confirms the cloud resource requirements of the model training task based on at least one of the cloud resource quantity and the cloud resource specifications.
11. The method according to any one of claims 1 to 10, characterized in that, The model training task submitted by the second tenant includes model template information, which includes the identifier of a pre-set model. The cloud management platform obtains the model training task submitted by the second tenant and confirms the cloud resource requirements of the model training task based on the model training task, including: The cloud management platform obtains the model training task submitted by the second tenant and confirms the cloud resource requirements of the model training task based on the identifier of the preset model in the model template information.
12. The method according to any one of claims 1 to 11, characterized in that, The cloud resources include at least one of virtual machines, containers, bare metal servers, and physical machines.
13. The method according to any one of claims 1 to 12, characterized in that, The transfer conditions include registering and successfully paying for the first cloud resource on the cloud management platform.
14. A cloud resource management device based on public cloud technology, characterized in that, The device is deployed on a cloud management platform, which manages the infrastructure providing cloud services. The infrastructure includes multiple servers, each containing cloud resources. The device includes: The interaction module is used to obtain the first successful purchase information of the first tenant for the first cloud resource in the infrastructure; The management module is used to respond to the first purchase success information by setting the user of the first cloud resource as the first tenant and the usage time as the first duration; The interaction module is further configured to obtain a transfer request from the first tenant for the first cloud resource, the transfer request including a second duration, the second duration being less than or equal to the first duration; The management module is also configured to respond to the transfer request for the first cloud resource by publishing the first resource transfer information of the first cloud resource to the Internet, wherein the first resource transfer information includes transfer conditions for the first cloud resource and for tenants other than the first tenant. The interaction module is further configured to obtain the model training task submitted by the second tenant, and the management module is further configured to confirm the cloud resource requirements of the model training task based on the model training task, and confirm that the first cloud resource meets part or all of the cloud resource requirements. The interaction module is further configured to obtain the second purchase success information of the second tenant regarding the transfer information of the first resource; the management module is further configured to confirm that the second purchase success information meets the transfer conditions, and set the user of the first cloud resource as the second tenant, and the usage time as the second duration; The management module is also used to call upon the first cloud resources to participate in the model training task within the second time period.
15. The apparatus as claimed in claim 14, characterized in that, The management module is used for: Confirm that each of the first cloud resources in the first cloud resources is located in the same parameter plane network.
16. The apparatus as claimed in claim 14 or 15, characterized in that, The management module is used for: It is confirmed that the first cloud resource and the second cloud resource in the infrastructure meet part or all of the cloud resource requirements, and each first cloud resource in the first cloud resource and each second cloud resource in the second cloud resource are located in the same parameter plane network; During the second duration, the second cloud resource is invoked to participate in the model training task.
17. The apparatus as claimed in claim 16, characterized in that, The interaction module is used to: obtain third successful purchase information of the third tenant for the second cloud resource; The management module is used to: in response to the third purchase success information, set the user of the second cloud resource as the third tenant; The interaction module is used to: obtain the transfer request from the third tenant for the second cloud resource; The management module is used to: respond to the transfer request for the second cloud resource and publish the second resource transfer information of the second cloud resource to the Internet, wherein the second resource transfer information includes transfer conditions for the second cloud resource and for tenants other than the third tenant.
18. The apparatus according to any one of claims 14 to 17, characterized in that, The model training task uses a training dataset to train the neural network model. The training dataset includes multiple batches of training data. After the first cloud resource is invoked to participate in the model training task within the second time period, the management module is used to: in response to a resource interruption request for the first cloud resource, save the checkpoint file of the neural network model after the current batch of training data has been trained. Release the first cloud resource.
19. The apparatus as claimed in claim 18, characterized in that, The resource interruption request is issued based on the end time of the second duration, and the management module is used for: The cloud management platform changes the user of the first cloud resource to the first tenant.
20. The apparatus according to any one of claims 14 to 19, characterized in that, The model training task includes a model loading phase and at least one round of model training. The cloud management platform bills the second tenant using a time-based billing method. The management module is used for: In response to the end of the model loading phase or the start of the at least one round of the model training phase, the billing for the second tenant's purchase of the first cloud resources is updated.
21. The apparatus according to any one of claims 14 to 20, characterized in that, The resource demand information includes the second tenant's first bid for the first cloud resource, and the resource interruption request is issued based on the fourth tenant's second bid for the first cloud resource being higher than the first bid. The interaction module is used to: obtain the fourth purchase success information of the fourth tenant regarding the first resource transfer information; The management module is used to: confirm that the fourth purchase success information meets the transfer conditions, and set the user of the first cloud resource as the fourth tenant.
22. The apparatus according to any one of claims 18 to 21, characterized in that, In response to a resource interruption request, after training the current batch of training data is completed and the checkpoint file of the neural network model is saved, the management module is used to: Deploy the model training task on the third cloud resource of the second tenant; Load the checkpoint file and train the next batch of training data for the current batch on the third cloud resource.
23. The apparatus according to any one of claims 14 to 22, characterized in that, The model training task submitted by the second tenant includes at least one of the cloud resource quantity and cloud resource specification in the cloud resource requirements. The interaction module is used to obtain the model training task submitted by the second tenant. The management module is used to determine the cloud resource requirements of the model training task based on at least one of the cloud resource quantity and the cloud resource specifications.
24. The apparatus according to any one of claims 14 to 23, characterized in that, The model training task submitted by the second tenant includes model template information, which includes the identifier of a pre-set model. The interaction module is used to obtain the model training task submitted by the second tenant. The management module is used to confirm the cloud resource requirements of the model training task based on the identifier of the preset model in the model template information.
25. The apparatus according to any one of claims 14 to 24, characterized in that, The cloud resources include at least one of virtual machines, containers, bare metal servers, and physical machines.
26. The apparatus according to any one of claims 14 to 25, characterized in that, The transfer conditions include registering and successfully paying for the first cloud resource on the cloud management platform. Other tenants include tenants who have registered and successfully paid for the first cloud resource on the cloud management platform.
27. A computing device cluster, characterized in that, The system includes multiple computing devices, each comprising multiple processors and multiple memories, wherein program instructions are stored in the multiple memories, and the multiple processors execute the program instructions, causing the cluster of computing devices to implement the method described in any one of claims 1 to 13.
28. A computer-readable storage medium, characterized in that, Includes program instructions that, when run on a computing device cluster, cause the computing device cluster to perform the method of any one of claims 1 to 13.
29. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1 to 13.