Method and System for Allocating Cloud Server Resources

By using a pre-trained resource allocation model in the Xinchuang Cloud Management Platform, the problem of resource allocation dependence on manual operations in the existing technology is solved, the efficiency and accuracy of resource allocation are improved, and the cost is reduced.

CN119759582BActive Publication Date: 2025-06-24HUAWEI MARINE NETWORKS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510244669.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-24
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

In the existing Xinchuang Cloud Management Platform, resource allocation relies on manual operations, resulting in low efficiency and high cost of resource activation, and high experience requirements for managers, resulting in waste of resources and unreasonable allocation of resources.

Method used

A pre-trained resource allocation model is adopted, combining the available resources of the cloud server and the resource requirements of the application to be deployed, resource allocation parameters are automatically generated, resource allocation and deployment are realized, and resource allocation results are adjusted when needed.

Benefits of technology

It improves the efficiency and accuracy of cloud server resource allocation, reduces the cost of resource allocation and labor demand, and reduces resource waste and error correction costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119759582B_ABST
    Figure CN119759582B_ABST
Patent Text Reader

Abstract

This application relates to the field of computer technology, and particularly to a method and system for allocating cloud server resources. The method includes: obtaining a first resource allocation request, and based on the available resources of the cloud server and the resource requirements of the target application to be deployed, using a pre-trained resource allocation model to obtain first resource allocation parameters corresponding to the first resource allocation request; based on the first resource allocation parameters, allocate first target resources from the available resources of the cloud server for the target application, and deploy the target application in the first target resources; in the case of obtaining a second resource allocation request, allocate second target resources for the target application, and migrate the target application deployed in the first target resources to the second target resources. The method and system for allocating cloud server resources provided by this application can achieve automatic allocation of cloud server resources, and can migrate the deployed target application, improving the accuracy and flexibility of resource allocation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method and system for allocating cloud service resources. Background Art

[0002] The Xinchuang Cloud is a cloud computing service platform built based on Xinchuang products and technologies, which can provide cloud computing resources such as computing, storage, and networking. Using the management platform of the Xinchuang Cloud, resources can be allocated according to the business type and data volume of the entrusted business, so that each entrusted business can run stably in the allocated resources.

[0003] In the current Xinchuang Cloud management platform, resources are usually allocated manually by managers. For example, an independent resource pool or tenant account is created for the entrusted business on the platform, a certain number of servers are allocated to it, and the required applications are deployed on the servers. However, the above manual allocation method not only consumes manpower, affects the resource activation efficiency, but also requires high experience from managers. Managers without rich experience often cannot allocate resources optimally, resulting in resource waste and increased resource costs. Moreover, in the case of unreasonable resource allocation, it is necessary to re-allocate the environment deployment, and the error correction cost is high. Summary of the Invention

[0004] To solve the above problems, this application provides a method and system for allocating cloud server resources, which can allocate target resources for a target application to be deployed from cloud servers, and can improve the efficiency and accuracy of cloud server resource allocation.

[0005] To achieve the above object, in a first aspect, this application provides a method for allocating cloud server resources, including: obtaining a first resource allocation request, where the first resource allocation request includes the resource requirements of the target application to be deployed; based on the available resource amount of the cloud server and the resource requirements of the target application to be deployed, using a pre-trained resource allocation model, obtaining first resource allocation parameters corresponding to the first resource allocation request; based on the first resource allocation parameters, allocating first target resources for the target application from the available resources of the cloud server, and deploying the target application in the first target resources; in the case of obtaining a second resource allocation request including second resource allocation parameters, allocating second target resources for the target application from the available resources of the cloud server, and migrating the target application deployed in the first target resources to the second target resources.

[0006] The resource allocation method for cloud servers provided in this application can utilize a resource allocation model, combine the available resources of cloud services and the resource requirements of the target application to be deployed, generate first resource allocation parameters, and then allocate resources to the target application to be deployed, and automatically deploy the target application to be deployed in the first target resource. At the same time, in the case of obtaining a second resource allocation request, the resource allocation result is adjusted, the second target resource is allocated to the target application according to the second resource allocation parameter, and the target application deployed in the first target resource is migrated to the second target resource, thereby adjusting the resource allocation result, reducing the cost of resource allocation, and improving the efficiency and accuracy of resource allocation.

[0007] In an alternative embodiment, based on the available resources of the cloud server and the resource requirements of the target application to be deployed, a first resource allocation parameter corresponding to the first resource allocation request is obtained by using a pre-trained resource allocation model, including: inputting the available resources of the cloud server and the resource requirements of the target application to be deployed into the resource allocation model; obtaining an allocation policy output by the resource allocation model, where the allocation policy is used to represent allocating a first quantity of resources to be allocated to a second quantity of target servers; generating a first resource allocation parameter according to the allocation policy.

[0008] In this way, it is possible to automatically generate a resource allocation plan using the pre-trained resource allocation model, reduce the cost of resource allocation and the subjectivity of the allocation result, and improve the efficiency of resource allocation.

[0009] In an alternative embodiment, the target application to be deployed includes at least one task to be deployed, the first resource allocation parameter includes at least one piece of allocation information, and the allocation information includes the task name of the task to be deployed, the first deployment address, and the first allocated resource quantity. The first deployment address is used to represent the deployment location of the task to be deployed, and the first allocated resource quantity is used to represent the resource quantity of the resource pool corresponding to the first deployment address.

[0010] In this way, according to the actual situation of the target application to be deployed, the target application can be deployed in multiple resource pools, and resource quantities are allocated to the multiple resource pools according to the first resource allocation parameter.

[0011] In an alternative embodiment, based on the first resource allocation parameter, the first target resource is allocated to the target application from the available resources of the cloud server, and the target application is deployed in the first target resource, including: obtaining the first target resource from the available resources of the cloud server according to the first deployment address and the first allocated resource quantity, where the location of the first target resource is the first deployment address and the resource quantity of the first target resource is the first allocated resource quantity; obtaining the software program corresponding to the task to be deployed; deploying the software program corresponding to the task to be deployed in the first target resource.

[0012] In this way, the first target resource can be obtained from the available resources of the cloud server according to the first allocated resource amount, the first target resource can be set according to the first deployment address, and the target application can be deployed in the first target resource according to the task to be deployed, thereby realizing the allocation of resources for the target application according to the first resource allocation parameter and completing the deployment work of the target application.

[0013] In an alternative embodiment, the second resource allocation parameter includes at least one piece of corrected allocation information, and the corrected allocation information includes the name of the task to be deployed, the second deployment address, and the second allocated resource amount. The second deployment address is used to represent the deployment location of the task to be deployed, and the second allocated resource amount is used to represent the resource amount of the resource pool corresponding to the second deployment address.

[0014] In this way, in the case of unreasonable resource allocation, the second deployment address and the second allocated resource amount corresponding to the name of the task to be deployed can be modified according to the corrected allocation information.

[0015] In an alternative embodiment, allocating a second target resource for the target application from the available resources of the cloud server and migrating the target application deployed in the first target resource to the second target resource includes: obtaining the second target resource from the available resources of the cloud server according to the second deployment address and the second allocated resource amount, where the location of the second target resource is the second deployment address and the resource amount of the second target resource is the second allocated resource amount; obtaining the software program corresponding to the task to be deployed; deploying the software program corresponding to the task to be deployed in the second target resource, and deleting the software program corresponding to the task to be deployed in the first target resource.

[0016] In this way, the deployed target application can be migrated according to the second resource allocation parameter. In the case of unreasonable resource allocation, resources can be reallocated, and the software program corresponding to the task to be deployed in the first target resource can be migrated to the second target resource in one key, reducing the error correction cost of modifying resource allocation.

[0017] In an alternative embodiment, before obtaining the first resource allocation request, it further includes training a resource allocation model. Training the resource allocation model includes: obtaining training data and inputting the training data into the resource allocation model. The training data includes the remaining server resources and the resource requirements to be deployed; according to a policy, generating at least one allocation policy corresponding to the resource requirements to be deployed; calculating the reward score corresponding to the allocation policy, and generating a first list according to the reward score. The reward score is used to evaluate the allocation policy; querying the largest reward score in the first list, and controlling the resource allocation model to output the allocation policy corresponding to the largest reward score.

[0018] In this way, the method of reinforcement learning can be used to train the resource allocation model to generate the best allocation strategy according to the remaining server resources and the resource requirements to be deployed, improving the accuracy of resource allocation.

[0019] In an alternative embodiment, calculating the reward score corresponding to the allocation strategy includes: calculating a first score, where the first score is the difference between a first quantity and a second quantity; calculating a second score according to the occupancy rate score formula, where the occupancy rate score formula is used to calculate the resource occupancy of the target server after deploying the resources to be allocated, and the occupancy rate score formula is:

[0020] ,

[0021] where D represents the second score, n represents the preset occupancy score, and f represents the resource occupancy percentage; calculating the reward score according to the first score and the second score, where the reward score is the sum of the first score and the second score.

[0022] In this way, the allocation strategy can be scored according to the resource occupancy of the target server and the number of target servers, so that the resource allocation model outputs the allocation strategy with the highest score.

[0023] In an alternative embodiment, the method for allocating cloud server resources further includes: in the case of obtaining a second resource allocation request, inputting the second resource allocation parameter into the resource allocation model to update the first list; in the case of not obtaining a second resource allocation request, inputting the first resource allocation parameter into the resource allocation model to update the first list.

[0024] In this way, training data can be continuously provided for the resource allocation model, thereby continuously optimizing the resource allocation model and improving the accuracy of resource allocation.

[0025] In a second aspect, the present application provides a system for allocating cloud server resources, including: an acquisition request module configured to: acquire a first resource allocation request, where the first resource allocation request includes the resource requirements of the target application to be deployed; a resource allocation module configured to: based on the available resources of the cloud server and the resource requirements of the target application to be deployed, use a pre-trained resource allocation model to obtain a first resource allocation parameter corresponding to the first resource allocation request; a resource deployment module configured to: based on the first resource allocation parameter, allocate first target resources for the target application from the available resources of the cloud server and deploy the target application in the first target resources; a resource adjustment module configured to: in the case of obtaining a second resource allocation request including a second resource allocation parameter, allocate second target resources for the target application from the available resources of the cloud server and migrate the target application deployed in the first target resources to the second target resources.

[0026] Understandably, the allocation system of cloud server resources provided in the second aspect is applied to the allocation method of cloud server resources provided above. Therefore, the beneficial effects it can achieve can refer to the beneficial effects in the allocation method of cloud server resources provided above, which will not be elaborated here. Brief Description of the Drawings

[0027] In order to more clearly illustrate the technical solutions of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 Schematic flowchart of an allocation method for cloud server resources provided in this embodiment;

[0029] Figure 2 Schematic flowchart of the training of a resource allocation model provided in this embodiment;

[0030] Figure 3 Schematic flowchart of deploying a target application provided in this embodiment;

[0031] Figure 4 Schematic flowchart of manually confirming an allocation policy provided in this embodiment;

[0032] Figure 5 Schematic structural diagram of an allocation system for cloud server resources provided in this embodiment. Detailed Description of the Embodiment

[0033] The technical solutions in the embodiments of the present application will be clearly described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the protection scope of the present application.

[0034] Hereinafter, terms such as "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.

[0035] In addition, in this application, orientation terms such as "upper", "lower", "inner", and "outer" are defined relative to the orientation of the components shown in the drawings. It should be understood that these directional terms are relative concepts used for description and clarification with respect to each other, and they can change accordingly with the change in the orientation of the components placed in the drawings.

[0036] For the convenience of understanding the solution, the following explains relevant terms:

[0037] Cloud server: A virtual server based on cloud computing technology that can run multiple virtual servers on a single physical server and flexibly allocate resources according to the actual needs of users, thereby improving the utilization efficiency of resources. The cloud server divides resources into different resource pools, including a computing resource pool, a storage resource pool, and a network resource pool. Among them, the computing resource pool includes a central processing unit (CPU) and memory, the storage resource pool includes hard disk resources, and the network resource pool includes bandwidth and an Internet protocol address (IP address). The cloud server allocates resources in the resource pool according to resource allocation requests.

[0038] Reinforcement learning algorithm: An algorithm in machine learning that emphasizes how a model takes a series of actions in an environment to maximize cumulative rewards. The model interacts with the environment, selects actions according to the policy, and learns the optimal behavior policy based on the reward signal feedback from the environment.

[0039] Policy: A method used by a neural network model to balance exploration and exploitation during the reinforcement learning training process. Among them, "exploration" means that the neural network model tries some actions that have not been fully tried before during training to discover better strategies or action values. "Exploitation" When the exploration rate is large, it means that the model is more inclined to explore. When the exploration rate is small, the model tends to be stable and is more inclined to exploit. "Action" is the behavior that the neural network model can take in the environment, and the neural network model interacts with the environment by selecting actions. "Environment" is the external world with which the neural network model interacts.

[0040] Q - value Table: It is used to store the knowledge accumulated by the neural network model during the learning process. The Q - value table records the value information of different actions of the neural network model in different environments. When the Q - value table is trained, this knowledge can be applied in similar scenarios. When the model encounters a state it has learned again, it can make decisions quickly based on the information in the Q - value table. Moreover, when facing some new environments that are partially similar, the knowledge in the Q - value table can be used as an initial reference and adapted to the new environment through further fine - tuning.

[0041] In the field of Xinchuang cloud technology, the management platform of Xinchuang cloud can allocate resources for entrusted business. First, the management staff needs to obtain the business type and data volume of the entrusted business, and then create an independent resource pool for the entrusted business based on the management platform, allocate a certain number of servers for it, and deploy the required applications in the servers, so that the entrusted business can run stably in the allocated resources.

[0042] However, the above - mentioned resource allocation method relies on manual resource allocation by management staff, which not only consumes manpower and affects the resource activation efficiency, but also requires high work experience of management staff. Management staff lacking experience often cannot allocate resources optimally, resulting in resource waste and increased resource costs. For example, when allocating resources, management staff allocate too many resources to users on a certain server, while the resources on other servers are not fully utilized, leading to a decrease in the resource utilization rate of the entire cloud environment and an increase in operation costs. In addition, the resource structure of manual resource allocation is relatively fixed. When the user's business changes and resource re - planning and adjustment are required, the operation difficulty is relatively large, resulting in poor flexibility of resource allocation and high error - correction costs.

[0043] In order to improve the allocation efficiency and accuracy of cloud server resources, the embodiments of this application provide a method for allocating cloud server resources. Refer to Figure 1 The method for allocating cloud server resources provided by this application includes:

[0044] Step S100: Obtain a first resource allocation request.

[0045] In the embodiments of this application, the method for allocating cloud server resources is applied to a cloud server management platform. The cloud server management platform is a software system used for centralized management, monitoring, configuration, and scheduling of cloud server resources. Through logging in to the management account, the administrator can obtain the available resource volume of the cloud server, and then allocate resources for the target application to be deployed from the available resource volume. The available resource volume of the cloud server includes computing resources, storage resources, and network resources.

[0046] Further, the administrator inputs the resource requirements of the target application to be deployed into the cloud server management platform. The cloud server management platform generates a first resource allocation request according to the resource requirements of the target application to be deployed, and obtains the first resource allocation request. The first resource allocation request includes the resource requirements of the target application to be deployed, and is used to apply for resources from the cloud server, so that the target application can be deployed in the resources allocated by the cloud server, and then the target application can run normally.

[0047] Exemplarily, in the first resource allocation request, the resource requirements of the target application to be deployed are as follows: a server with 32GB of memory, the server is configured with an 8-core CPU, the storage capacity is 1TB, and the network bandwidth is 50Mbps.

[0048] Step S200: Based on the available resource amount of the cloud server and the resource requirements of the target application to be deployed, use the pre-trained resource allocation model to obtain the first resource allocation parameters corresponding to the first resource allocation request.

[0049] In the embodiment of the present application, the resource allocation model can allocate resources for the target application to be deployed according to the available resource amount of the cloud server, thereby realizing the automatic allocation of cloud server resources and improving the allocation efficiency.

[0050] First, the cloud server management platform inputs the available resource amount of the cloud server and the resource requirements of the target application to be deployed into the resource allocation model; then obtains the allocation policy output by the resource allocation model, and the allocation policy is used to represent that the first quantity of resources to be allocated is allocated to the second quantity of target servers; finally, generates the first resource allocation parameters according to the allocation policy, and the first resource allocation parameters include the allocation information for allocating target resources to the target application to be deployed.

[0051] In this way, the resource allocation model can calculate the resources required by the target application to be deployed according to the input resource requirements of the target application to be deployed, such as the number of CPU cores, memory size, storage resource capacity, and network resource bandwidth in the computing resources. At the same time, combined with the available resource amount of the current cloud server, such as the performance of the server, the free capacity of the storage device, and the available bandwidth of the network, generate an allocation policy, and generate the first resource allocation parameters according to the allocation policy.

[0052] In order to improve the accuracy and rationality of the first resource allocation parameters, the cloud server management platform needs to perform reinforcement learning training on the resource allocation model using training data before obtaining the first resource allocation request. During the process of training the resource allocation model, first obtain the training data, and the training data includes the server resource surplus and the resource requirements to be deployed, and input the training data into the resource allocation model.

[0053] Further, the cloud server management platform according to The policy generates at least one allocation policy corresponding to the resource requirements to be deployed, calculates the reward score corresponding to the allocation policy, and generates a first list based on the reward score. The reward score is used to evaluate the allocation policy. Finally, the reward score with the largest value is queried from the first list, and the resource allocation model is controlled to output the allocation policy corresponding to the reward score with the largest value.

[0054] During the above training process, the cloud server management platform uses the SARSA algorithm to train the resource allocation model. SARSA is a policy-based temporal difference learning algorithm in reinforcement learning. Its name comes from the sequence used in its update formula: State, Action, Reward, State, Action. The SARSA algorithm is used to train the resource allocation model to output the allocation policy with the highest score.

[0055] During the process of the cloud server management platform using the SARSA algorithm to train the resource allocation model, first, a Q function (usually a table, i.e., the Q-value table) is initialized to store the expected long-term reward for the model to take each action in each environment. Among them, the Q-value table includes Q-values, and the initialization of the Q-values can be random or set according to some prior knowledge.

[0056] Next, according to the policy, an action is selected, that is, the model randomly selects an action according to the exploration rate or selects the action with the highest Q-value with a probability of 1 - After the model executes the selected action, the environment will feedback a reward score, causing the model to transfer to a new state. The reward score is used to evaluate the quality of the action, and the new state is the basis for the model's next decision.

[0057] Finally, the Q-value table is updated using the SARSA update formula, and the update formula is as follows:

[0058] ,

[0059] where, represents the current state, represents the current action, represents the reward obtained after executing the action , represents the next state, represents the next action, represents the learning rate, which is used to control the update step size, represents the discount factor, which is used to measure the importance of future rewards relative to current rewards.

[0060] Use the above update formula to update the Q value. As the number of updates increases, the Q value table will gradually converge. Among them, the learning rate adopts an exponential decay method. As the number of executed actions increases, the learning rate is adjusted using the learning rate decay formula. The learning rate decay formula is as follows:

[0061] ,

[0062] Among them, represents the decay coefficient, and t represents the number of executed actions. For example, the decay coefficient can be 1.

[0063] In the embodiment of this application, the process of the cloud server management platform using the SARSA training resource allocation model is as Figure 2 shown. First, obtain the initialized Q value table and initialize the Q value to 0. Then, set the initial state according to the resource situation of the current cloud server , and according to the policy, select an action , and execute , The policy includes the exploration rate . When executing the action for the first time, set the initial value to 1. As the number of executed actions increases, the exploration rate decays according to the decay formula until it decays to the preset exploration threshold, and then the exploration stops. For example, the preset exploration threshold is 0.05. The decay formula is as follows:

[0064] ,

[0065] Among them, represents the exploration rate, and t represents the number of executed actions.

[0066] It should be understood that the action can be an allocation policy, which is used to represent allocating the first quantity of resources to be allocated to the second quantity of target servers. After executing the action , calculate the reward score corresponding to the action , and then evaluate the action .

[0067] In the process of the cloud server management platform calculating the reward score, first calculate the first score, which is the difference between the first quantity and the second quantity; then calculate the second score according to the occupancy rate score formula. The occupancy rate score formula is used to calculate the resource occupancy situation of the target server after deploying the resources to be allocated. The occupancy rate score formula is:

[0068] ,

[0069] Among them, D represents the second score, n represents the preset occupancy score, and f represents the resource occupancy percentage; finally, according to the first score and the second score, the reward score is calculated, and the reward score is the sum of the first score and the second score.

[0070] Exemplarily, the allocation policy is to allocate 6 resources to be allocated to 2 target servers (target server A and target server B). After deploying the resources to be allocated, the occupancy rate of target server A is 70%, and the occupancy rate of target server B is 80%. The preset occupancy score is 9. Then, using the above method, the first score can be calculated as 4 points, the second score of target server A is 7 points, the second score of target server B is 8 points, and the reward score is 19 points.

[0071] Furthermore, the cloud server management platform performs actions After allocating resources, the cloud server is updated to a new state , and then used again The policy selects the next action according to the new state , and updates the Q value using the update formula, and generates a first list based on the Q value, that is, the Q value table. It should be noted that as the number of executed actions increases, the Q value table will gradually converge. By observing the changes in the Q value after multiple updates, it can be determined whether the resource allocation model has completed training.

[0072] In the embodiments of the present application, if the change in the Q value is less than the preset change threshold after multiple updates, or remains stable after a certain number of iterations, the Q value table is marked as a convergent state. Stop training the resource allocation model, and then enable the resource allocation model to select the best resource allocation policy according to the Q value table in the convergent state.

[0073] If the Q value table does not reach the convergent state, then obtain the state corresponding to the cloud server after performing the action

[0074] , then execute the selected action , and calculate the reward score corresponding to the action , and the state corresponding to the cloud server after performing the action , and according to The policy selects an action , updates the Q value table using the update formula until the Q value table is in a convergent state. Accordingly The policy selects an action , and updates the Q value table using the update formula until the Q value table is in a convergent state.

[0075] In this way, the trained resource allocation model can automatically generate the allocation policy corresponding to the maximum reward score according to the available resources of the cloud server and the resource requirements of the target application to be deployed, thereby making the allocation of cloud server resources more efficient and accurate.

[0076] Step S300: Based on the first resource allocation parameter, allocate the first target resource from the available resources of the cloud server for the target application, and deploy the target application in the first target resource.

[0077] In the embodiments of the present application, the target application to be deployed includes at least one task to be deployed, and the first resource allocation parameter includes at least one piece of allocation information. The allocation information includes the task name (task to be deployed), the first deployment address, and the first allocated resource amount. The first deployment address is used to represent the deployment location of the task to be deployed, and the first allocated resource amount is used to represent the resource amount of the resource pool corresponding to the first deployment address.

[0078] The allocation information and the task to be deployed are in an associated relationship, and the allocation information is used to represent the deployment location and the allocated resource amount of the task to be deployed. Exemplarily, if the target application to be deployed includes Task A, Task B, and Task C, Table 1 is used to represent the first resource allocation parameter.

[0079] Table 1

[0080] 。

[0081] In the embodiments of the present application, as Figure 3 shown, the cloud server management platform obtains the first target resource from the available resources of the cloud server according to the first deployment address and the first allocated resource amount, and obtains the software program corresponding to the task to be deployed, and deploys the software program corresponding to the task to be deployed in the first target resource. Among them, the location of the first target resource is the first deployment address, and the resource amount of the first target resource is the first allocated resource amount. In this way, after resource allocation, the target application to be deployed can be automatically deployed in the corresponding allocated target resource, reducing the difficulty of manually deploying the target application and improving the deployment efficiency.

[0082] Step S400: Determine whether a second resource allocation request is received.

[0083] In the embodiments of the present application, in some cases, the allocation strategy generated by the resource allocation model can meet the user's needs, and the current resource allocation and deployment task is ended.

[0084] However, in other cases, due to incomplete coverage of the training data of the resource allocation model or the limitations of the resource allocation model itself, the output allocation strategy cannot meet the user's needs. At this time, the user can modify the first resource allocation parameter, and then generate a second resource allocation request including the second resource allocation parameter.

[0085] Therefore, the cloud server management platform needs to determine whether a second resource allocation request is received, and then determine whether it is necessary to re-allocate and deploy the cloud server resources.

[0086] Step S4011: When a second resource allocation request containing second resource allocation parameters is obtained, allocate second target resources for the target application from the available resources of the cloud server, and migrate the target application deployed in the first target resources to the second target resources.

[0087] In the embodiments of the present application, when the cloud server management platform obtains a second resource allocation request, it indicates that the user needs to modify the automatic allocation result of the cloud server resources according to the second resource allocation parameters. Among them, the second resource allocation parameters include at least one piece of corrected allocation information, and the corrected allocation information includes the name of the task to be deployed, the second deployment address, and the second allocated resource amount. The second deployment address is used to represent the deployment location of the task to be deployed, and the second allocated resource amount is used to represent the resource amount of the resource pool corresponding to the second deployment address.

[0088] The corrected allocation information has an associated relationship with the task to be deployed, and the corrected allocation information is used to represent the corrected deployment location and allocated resource amount of the task to be deployed. Exemplarily, the user modifies the allocated resource amount of the task to be deployed B in Table 1 to "8-core CPU, storage capacity 256GB, bandwidth 25Mbps", and modifies the deployment address of the task to be deployed C in Table 1 to "Server 1", then Table 2 is used to represent the second resource allocation parameters.

[0089] Table 2

[0090] 。

[0091] In the embodiments of the present application, the cloud server management platform obtains second target resources from the available resources of the cloud server according to the second deployment address and the second allocated resource amount, obtains the software program corresponding to the task to be deployed, deploys the software program corresponding to the task to be deployed in the second target resources, and deletes the software program corresponding to the task to be deployed in the first target resources. Among them, the location of the second target resources is the second deployment address, and the resource amount of the second target resources is the second allocated resource amount.

[0092] In this way, when the cloud server resource allocation scheme generated by the resource allocation model is unreasonable, the cloud server management platform can quickly correct the allocation scheme according to the second resource allocation parameters, and automatically migrate the software program corresponding to the task to be deployed from the first target resources to the second target resources to achieve one-key migration between servers.

[0093] It should be noted that in addition to automatically allocating target resources and deploying target applications based on the first resource allocation parameters, the solution of the present application also provides a manual confirmation method. Refer to Figure 4, after the first resource allocation parameter is generated, the cloud server management platform first obtains the evaluation identifier input by the user. The evaluation identifier can be "satisfied" or "dissatisfied". Among them, "satisfied" is used to indicate that the user determines to allocate the cloud server resources using the allocation policy corresponding to the first resource allocation parameter. "Dissatisfied" is used to indicate that the user needs to modify the first resource allocation parameter.

[0094] Exemplarily, after the user inputs an evaluation identifier including "satisfied" based on the first resource allocation parameter, the resources of the cloud server are allocated according to the first resource allocation parameter, and the target application is deployed.

[0095] Exemplarily, after the user inputs an evaluation identifier including "dissatisfied" based on the first resource allocation parameter, the parameters modified by the user are obtained to generate a second resource allocation request including the second resource allocation parameter, and the one-key migration of the server is implemented according to the second resource allocation parameter, the resource allocation policy of the cloud server is quickly adjusted, and the deployment of the target application is completed according to the second resource allocation parameter.

[0096] Step S4012: Input the second resource allocation parameter into the resource allocation model to update the first list.

[0097] In the embodiment of the present application, in order to improve the accuracy of cloud server resource allocation, the cloud server management platform needs to update the training data, thereby increasing the generalization ability of the resource allocation model, so that the resource allocation model can generate a more accurate allocation policy. By inputting the second resource allocation parameter into the resource allocation model, the resource allocation model sets the action corresponding to the second resource allocation parameter in the first list as the highest reward score, thereby updating the first list and improving the accuracy of cloud server resource allocation.

[0098] Step S4021: In the case where the second resource allocation request is not obtained, input the first resource allocation parameter into the resource allocation model to update the first list.

[0099] In the case where the second resource allocation request is not obtained, it indicates that the first resource allocation parameter is the best solution for the current cloud server resource allocation. The cloud server management platform inputs the first resource allocation parameter into the resource allocation model, and the resource allocation model sets the action corresponding to the first resource allocation parameter in the first list as the highest reward score, thereby updating the first list and improving the accuracy of cloud server resource allocation.

[0100] As can be seen from the above technical solution, the present application provides a method for allocating cloud server resources. The method can utilize a resource allocation model, combine the available resources of the cloud service and the resource requirements of the target application to be deployed, generate first resource allocation parameters, and then allocate resources for the target application to be deployed, and automatically deploy the target application to be deployed in the first target resources. At the same time, in the case of obtaining a second resource allocation request, the resource allocation result is adjusted, the second target resources are allocated for the target application according to the second resource allocation parameters, and the target application deployed in the first target resources is migrated to the second target resources, thereby adjusting the resource allocation result, reducing the cost of resource allocation, and improving the accuracy of resource allocation.

[0101] Based on the method for allocating cloud server resources provided in the above embodiments, some embodiments of the present application further provide a system for allocating cloud server resources. Refer to Figure 5 , the system for allocating cloud server resources is used to execute the method for allocating cloud server resources described in the first embodiment. The system for allocating cloud server resources includes: a cloud server, a communication device, and a computer. Among them, the computer uses the communication device to send a resource allocation request to the cloud server to allocate the available resources of the cloud server. The cloud server management platform software is installed in the computer, and the cloud server management platform includes an acquisition request module 501, a resource allocation module 502, a resource deployment module 503, and a resource adjustment module 504.

[0102] The acquisition request module 501 is configured to: acquire a first resource allocation request, and the first resource allocation request includes the resource requirements of the target application to be deployed. The resource allocation module 502 is configured to: based on the available resources of the cloud server and the resource requirements of the target application to be deployed, utilize a pre-trained resource allocation model to acquire first resource allocation parameters corresponding to the first resource allocation request. The resource deployment module 503 is configured to: based on the first resource allocation parameters, allocate first target resources for the target application from the available resources of the cloud server, and deploy the target application in the first target resources. The resource adjustment module 504 is configured to: in the case of obtaining a second resource allocation request including second resource allocation parameters, allocate second target resources for the target application from the available resources of the cloud server, and migrate the target application deployed in the first target resources to the second target resources.

[0103] The effect of the above system during the execution of the method can be referred to the description in the above method, and will not be elaborated here.

[0104] It should be noted that those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope of the present application is pointed out by the claims.

[0105] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for allocating cloud server resources, characterized in that: include: Acquire a first resource allocation request, where the first resource allocation request includes a resource requirement of a target application to be deployed, and the target application to be deployed includes at least one task to be deployed; Based on the available resources of the cloud server and the resource requirements of the target application to be deployed, using a pre-trained resource allocation model, obtaining a first resource allocation parameter corresponding to the first resource allocation request; Allocating a first target resource to the target application from available resources of the cloud server based on the first resource allocation parameter, and deploying the target application in the first target resource; In case of obtaining a second resource allocation request including a second resource allocation parameter, allocating a second target resource to the target application from available resources of the cloud server, and migrating the target application deployed in the first target resource to the second target resource, so as to achieve application migration between servers, wherein the second resource allocation parameter includes at least one piece of modified allocation information, and the modified allocation information includes a name of a task to be deployed, a second deployment address, and a second allocated resource amount, wherein the second deployment address is used to characterize a deployment location of the task to be deployed, and the second allocated resource amount is used to characterize a resource amount of a resource pool corresponding to the second deployment address; The allocating a second target resource to the target application from the available resources of the cloud server, and migrating the target application deployed in the first target resource to the second target resource includes: According to the second deployment address and the second allocated resource amount, acquiring the second target resource from available resources of the cloud server, where the location of the second target resource is the second deployment address, and the resource amount of the second target resource is the second allocated resource amount; Obtaining the software program corresponding to the task to be deployed; The software program corresponding to the task to be deployed is deployed in the second target resource, and the software program corresponding to the task to be deployed in the first target resource is deleted.

2. The method for allocating cloud server resources according to claim 1, characterized in that: The obtaining, based on the available resources of the cloud server and the resource requirements of the target application to be deployed, a first resource allocation parameter corresponding to the first resource allocation request using a pre-trained resource allocation model includes: Inputting the available resources of the cloud server and the resource requirements of the target application to be deployed into the resource allocation model; Acquire an allocation strategy output by the resource allocation model, where the allocation strategy is used to characterize allocation of a first number of resources to be allocated to a second number of target servers; The first resource allocation parameter is generated according to the allocation strategy.

3. The method for allocating cloud server resources according to claim 2, characterized in that: The first resource allocation parameter includes at least one allocation information, and the allocation information includes a task name (task to be deployed), a first deployment address and a first allocated resource amount, wherein the first deployment address is used to characterize the deployment location of the task to be deployed, and the first allocated resource amount is used to characterize the resource amount of the resource pool corresponding to the first deployment address.

4. The method for allocating cloud server resources according to claim 3, characterized in that: The allocating a first target resource to the target application from available resources of the cloud server based on the first resource allocation parameter, and deploying the target application in the first target resource includes: According to the first deployment address and the first allocated resource amount, acquiring the first target resource from available resources of the cloud server, where the location of the first target resource is the first deployment address, and the resource amount of the first target resource is the first allocated resource amount; Obtaining the software program corresponding to the task to be deployed; Deploy the software program corresponding to the task to be deployed in the first target resource.

5. The method for allocating cloud server resources according to claim 2, characterized in that: Before obtaining the first resource allocation request, the method further includes training the resource allocation model, wherein the training the resource allocation model includes: Acquire training data, and input the training data into the resource allocation model, wherein the training data includes server resource margin and resource requirements to be deployed; according to Strategy generates at least one allocation strategy corresponding to the resource demand to be deployed; Calculating a reward score corresponding to the allocation strategy, and generating a first list according to the reward score, wherein the reward score is used to evaluate the allocation strategy; The reward score with the largest value is searched from the first list, and the resource allocation model is controlled to output the allocation strategy corresponding to the reward score with the largest value.

6. The method for allocating cloud server resources according to claim 5, characterized in that: The calculating the reward score corresponding to the allocation strategy includes: Calculating a first score, the first score being a difference between the first number and the second number; The second score is calculated according to an occupancy rate score formula, wherein the occupancy rate score formula is used to calculate the resource occupancy of the target server after the resources to be allocated are deployed, and the occupancy rate score formula is: , Wherein, D represents the second score, n represents the preset occupancy score, and f represents the resource occupancy percentage; The reward score is calculated according to the first score and the second score, and the reward score is the sum of the first score and the second score.

7. The method for allocating cloud server resources according to claim 5, characterized in that: Also includes: In case of acquiring the second resource allocation request, inputting the second resource allocation parameter into the resource allocation model to update the first list; In a case where the second resource allocation request is not obtained, the first resource allocation parameter is input into the resource allocation model to update the first list.

8. A cloud server resource allocation system, characterized in that: The cloud server resource allocation method applied to any one of claims 1 to 7, wherein the allocation system comprises: The request acquisition module is configured to: acquire a first resource allocation request, wherein the first resource allocation request includes a resource requirement of a target application to be deployed, and the target application to be deployed includes at least one task to be deployed; A resource allocation module is configured to: based on the available resources of the cloud server and the resource requirements of the target application to be deployed, using a pre-trained resource allocation model, obtain a first resource allocation parameter corresponding to the first resource allocation request; a resource deployment module, configured to: allocate a first target resource to the target application from available resources of the cloud server based on the first resource allocation parameter, and deploy the target application in the first target resource; The resource adjustment module is configured to: upon obtaining a second resource allocation request including a second resource allocation parameter, allocate a second target resource to the target application from available resources of the cloud server, and migrate the target application deployed in the first target resource to the second target resource to achieve application migration between servers, wherein the second resource allocation parameter includes at least one piece of modified allocation information, the modified allocation information includes a name of a task to be deployed, a second deployment address, and a second allocated resource amount, wherein the second deployment address is used to characterize a deployment location of the task to be deployed, and the second allocated resource amount is used to characterize a resource amount of a resource pool corresponding to the second deployment address; The resource adjustment module is specifically configured to: obtain the second target resource from the available resources of the cloud server according to the second deployment address and the second allocated resource amount, the location of the second target resource being the second deployment address, and the resource amount of the second target resource being the second allocated resource amount; obtain the software program corresponding to the task to be deployed; deploy the software program corresponding to the task to be deployed in the second target resource, and delete the software program corresponding to the task to be deployed in the first target resource.

Citation Information

Patent Citations

  • Internet of Vehicles system scheduling method and Internet of Vehicles resource scheduling method

    CN117640693A

  • Computing resource dynamic allocation method and device, electronic equipment and storage medium

    CN119396555A