Capacity expansion and contraction method and system, electronic equipment and storage medium
By using preset prediction models in cloud services, automatically predicting resource requirements and performing scaling operations, the problems of manual decision-making dependence and poor prediction accuracy in the prior art are solved, and efficient and automatic resource management and cost reduction are achieved.
Patent Information
- Application Number
- CN202510373017.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, service capacity adjustment mainly relies on manual decision-making, resulting in poor prediction accuracy, and high scheduling complexity and operation and maintenance costs.
By utilizing preset prediction models, the resource demand patterns are output based on the historical usage data of the target cloud service, real-time monitoring data, and usage trend data, the resource demand patterns are determined, the capacity requirements and migration types are determined, and the capacity expansion or reduction operations are automatically performed according to the preset strategy.
It realizes fully automatic prediction of resource demand and fully automatic execution of expansion and expansion capacity, improves the accuracy and efficiency of expansion capacity, reduces operation and maintenance costs, and makes full use of server resources.
Smart Images

Figure CN120223702A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud services, and in particular, to a method, system, electronic device, and storage medium for scaling. Background Art
[0002] As one of the important guarantees for the stable operation of services, the cost pressure generated by service capacity guarantee also needs to be emphasized and concerned. In related technologies, service scaling is usually carried out in the following ways:
[0003] Advance layout: For example, the way for website distributed services to cope with festivals and sudden increases in traffic is to carry out manual layout in advance, that is, to expand the capacity in advance by adding machines. On the day of large traffic, manual duty is carried out, and in combination with service feature monitoring, the number of service requests and the resource usage of machines are judged, the service and machine features are judged. When it is found that it does not meet the expectations, it is judged that the service needs to be scaled, and then according to experience, the manual operation platform is used to complete service scaling.
[0004] Rely on monitoring judgment: After the capacity water level monitoring or the traffic increases, manual decisions are made on whether to expand the capacity and how much to expand. Finally, the execution is also through platform operation. Some companies or systems also need to go through multiple levels of approval because it involves cost increases.
[0005] It can be seen that the current adjustment requirements generated in response to service capacity pressure all rely on manual decisions and operations, which are mainly divided into adjustment requirement perception and specific adjustment decisions, that is, how to make targeted adjustments. Due to the heavy dependence on the capacity experience of operation and maintenance personnel, the prediction accuracy is poor. Summary of the Invention
[0006] In view of this, embodiments of the present invention provide a method, system, electronic device, and storage medium for scaling to improve the accuracy of scaling.
[0007] According to one aspect of the present invention, a method for scaling is provided, and the method includes:
[0008] Using a preset prediction model, based on the historical usage data, real-time monitoring data, and usage trend data of the target cloud service, output the resource demand pattern of the target cloud service, where the resource demand pattern includes the capacity usage trend and user request rate of the target cloud service;
[0009] Based on the resource demand pattern of the target cloud service and the current resource usage data of the target cloud service, determine the capacity demand and migration type of the target cloud service, where the migration type includes scaling down and scaling up;
[0010] When the migration type of the target cloud service is scale-down, based on the resource occupancy of the target cloud service in each server and a preset scale-down policy, determine the scale-down servers, and delete the target cloud service instances of the target cloud service deployed in the scale-down servers, where the preset scale-down policy is to determine the first preset number of servers with the least resource occupancy of the target cloud service as the scale-down servers;
[0011] When the migration type of the target cloud service is scale-up, based on the resource surplus in each idle server and a preset scale-up policy, determine the scale-up servers, and migrate the target cloud service instances to the scale-up servers, where the idle servers are the servers on which the target cloud service is not deployed, and the preset scale-up policy is to determine the second preset number of servers with the most resource surplus as the scale-up servers.
[0012] In a possible embodiment, the preset prediction model is pre-trained through the following steps:
[0013] Obtain a training sample set, where each sample data in the training sample set includes the resource usage data of the service instance, the user request rate, and the service scale data;
[0014] Output each sample data in the training sample set to an initial prediction model, and obtain the capacity usage trend data output by the initial prediction model for the sample data, where the capacity usage trend data includes the capacity usage change rate of the service instance after a preset time interval;
[0015] Based on the difference between the capacity usage trend and the actual usage trend of the corresponding sample data, train the initial prediction model until the difference is less than a preset difference threshold to obtain a preset prediction model; where the actual usage trend of the sample data is obtained based on the resource usage data of the sample data and the resource usage data of the corresponding service instance after the preset time interval.
[0016] In a possible embodiment, the determining the capacity requirement and the migration type of the target cloud service based on the resource demand pattern of the target cloud service and the current resource usage data of the target cloud service includes:
[0017] If the difference between the first product of the capacity usage trend and the current resource usage data of the target cloud service and the second product of the current resource usage data of the target cloud service and a preset resource reservation threshold is greater than 0, it is determined that the target cloud service needs to be scaled down;
[0018] If the difference between the third product of the capacity usage trend, the current usage data of the target cloud service, and the preset buffer base number and the current resource usage amount of the target cloud service is greater than 0, it is determined that the target cloud service needs to be scaled out.
[0019] In a possible embodiment, the method further includes:
[0020] After the scaling out or scaling in is completed, determine whether the resource utilization rates of the servers in the resource pool meet the preset resource utilization rate range;
[0021] If there is a server that does not meet the preset resource utilization rate range, modify the preset scaling out policy or the preset scaling in policy.
[0022] In a possible embodiment, the method further includes:
[0023] For each server in the resource pool, based on the current resource utilization rate of the server, the number of allocated CPU cores in the server, and the preset resource utilization rate, determine the number of machines that can be optimized, where the machine that can be optimized is a server with an allocation rate less than the preset allocation rate threshold;
[0024] Based on the deployment information in each of the servers in the resource pool and the number of machines that can be optimized, determine the servers to be taken offline, where the deployment information includes the number of service deployments;
[0025] Sort the remaining servers in the resource pool in descending order of the remaining resources, and migrate the service instances deployed in the servers to be taken offline to each of the remaining servers in sequence according to the sorting, where the remaining servers are the other servers in the resource pool except the servers to be taken offline.
[0026] In a possible embodiment, the migrating the service instances deployed in the servers to be taken offline to each of the remaining servers in sequence according to the sorting includes
[0027] For each service instance in the servers to be taken offline, obtain the attribute information of the service instance, where the attribute information includes port information, deployment file information, and traffic peak information;
[0028] Based on the attribute information of the service instance and the deployment attribute information in each of the remaining servers, migrate the service instance to the remaining server that does not conflict with the service instance according to the sorting, where not conflicting with the service instance means having a different port from the service instance, a different storage path for the deployment file, and a different traffic peak.
[0029] According to another aspect of the present invention, there is provided a scaling system, the system includes:
[0030] A prediction module, configured to use a preset prediction model to output a resource demand pattern of the target cloud service based on historical usage data, real-time monitoring data, and usage trend data of the target cloud service, where the resource demand pattern includes a capacity usage trend and a user request rate of the target cloud service;
[0031] A determination module, configured to determine a capacity demand and a migration type of the target cloud service based on the resource demand pattern of the target cloud service and current resource usage data of the target cloud service, where the migration type includes capacity reduction and capacity expansion;
[0032] A migration module, configured to, when the migration type of the target cloud service is capacity reduction, determine a capacity reduction server based on the resource occupancy of the target cloud service in each server and a preset capacity reduction policy, and delete a target cloud service instance of the target cloud service deployed in the capacity reduction server, where the preset capacity reduction policy is to determine the first preset number of servers with the least resource occupancy of the target cloud service as the capacity reduction servers;
[0033] When the migration type of the target cloud service is capacity expansion, determine an expansion server based on the resource surplus in each idle server and a preset expansion policy, and migrate the target cloud service instance to the expansion server, where the idle server is a server on which the target cloud service is not deployed, and the preset expansion policy is to determine the second preset number of servers with the most resource surplus as the expansion servers.
[0034] In a possible embodiment, the preset prediction model is pre-trained through the following steps:
[0035] Obtain a training sample set, where each sample data in the training sample set includes resource usage data, user request rate, and service scale data of a service instance;
[0036] Output each sample data in the training sample set to an initial prediction model, and obtain capacity usage trend data output by the initial prediction model for the sample data, where the capacity usage trend data includes a capacity usage change rate of the service instance after a preset time interval;
[0037] Train the initial prediction model based on the difference between the capacity usage trend and the actual usage trend of the corresponding sample data until the difference is less than a preset difference threshold to obtain a preset prediction model; where the actual usage trend of the sample data is obtained based on the resource usage data of the sample data and the resource usage data of the corresponding service instance after the preset time interval;
[0038] Determining the capacity requirement and migration type of the target cloud service based on the resource demand pattern of the target cloud service and the current resource usage data of the target cloud service includes:
[0039] If the difference between the first product of the capacity usage trend and the current resource usage data of the target cloud service and the second product of the current resource usage data of the target cloud service and the preset resource reservation threshold is greater than 0, it is determined that the target cloud service needs to downsize;
[0040] If the difference between the third product of the capacity usage trend, the current usage data of the target cloud service, and the preset buffer base number and the current resource usage amount of the target cloud service is greater than 0, it is determined that the target cloud service needs to upsize;
[0041] The system further includes:
[0042] An adjustment module, configured to determine whether the resource utilization rate in each server in the resource pool meets the preset resource utilization rate range after the upsize or downsize is completed;
[0043] If there is a server that does not meet the preset resource utilization rate range, modify the preset upsize policy or the preset downsize policy;
[0044] A redistribution module, configured to determine the number of optimizable machines for each server in the resource pool based on the current resource utilization rate of the server, the number of allocated CPU cores in the server, and the preset resource utilization rate, where the optimizable machine is a server with an allocation rate less than the preset allocation rate threshold;
[0045] Based on the deployment information in each server in the resource pool and the number of optimizable machines, determine the servers to be taken offline, where the deployment information includes the number of service deployments;
[0046] Sort the remaining servers in the resource pool in descending order of the remaining resources, and migrate the service instances deployed in the servers to be taken offline to each of the remaining servers in sequence according to the sorting, where the remaining servers are other servers in the resource pool except the servers to be taken offline;
[0047] The migrating the service instances deployed in the servers to be taken offline to each of the remaining servers in sequence according to the sorting includes
[0048] For each service instance in the servers to be taken offline, obtain the attribute information of the service instance, where the attribute information includes port information, deployment file information, and traffic peak information;
[0049] Based on the attribute information of the service instance and the deployment attribute information in each of the remaining servers, migrate the service instance to the remaining server that does not conflict with the service instance according to the sorting. "Not conflicting with the service instance" means having a different port from the service instance, a different deployment file storage path, and a different traffic peak.
[0050] According to another aspect of the present invention, there is provided an electronic device, including:
[0051] A processor; and
[0052] A memory storing a program,
[0053] wherein the program includes instructions that, when executed by the processor, cause the processor to execute any one of the above-described scaling methods.
[0054] According to another aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute any one of the above-described scaling methods.
[0055] One or more technical solutions provided in the embodiments of the present invention achieve fully automatic prediction of resource demand by using a prediction model to predict the resource demand pattern of the target cloud service, without manual participation. Moreover, the prediction model is trained with a large amount of data and can output relatively accurate resource demand prediction data compared with manual prediction. Subsequently, automatic scaling of the target cloud service is performed according to the predicted scaling strategy. In the scaling strategy, the scaling servers are determined based on the server resource occupancy, achieving full utilization of the resource data in the servers and reducing the service operation cost to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features, and advantages of the present invention are disclosed. In the drawings:
[0057] Figure 1 is a flowchart of a scaling method provided by an embodiment of the present invention;
[0058] Figure 2 is a flowchart of determining scaling servers in the scaling method provided by an embodiment of the present invention;
[0059] Figure 3 is a flowchart of adjusting instance deployment in the scaling method provided by an embodiment of the present invention;
[0060] Figure 4 is a schematic diagram of a system architecture for implementing the scaling method provided by an embodiment of the present invention;
[0061] Figure 5 This is a schematic structural diagram of the scaling system provided by an embodiment of the present invention;
[0062] Figure 6 The block diagram of the structure of an exemplary electronic device capable of implementing the embodiments of the present invention is shown. Detailed implementation manners
[0063] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.
[0064] It should be understood that the various steps recorded in the method embodiments of the present invention can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.
[0065] As used herein, the term "including" and its variants are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions executed by these devices, modules or units or their interdependent relationships.
[0066] It should be noted that the modifications of "one" and "a plurality" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".
[0067] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0068] The current scaling methods mainly have the following problems:
[0069] Limited prediction accuracy: Currently, there is a heavy reliance on the capacity experience of operation and maintenance personnel, and there is a lack of data evidence, and the scalability is poor
[0070] Intelligent scheduling complexity: The formulation and execution of intelligent scheduling strategies involve multiple dimensions and complex factors, such as resource conflicts, service priorities, cost - effectiveness, etc. In practical applications, it may be difficult for scheduling strategies to fully meet all business requirements, resulting in uneven resource allocation or low scheduling efficiency.
[0071] Difficulty in containerization and microservice management: Although containerization and microservice architectures improve the flexibility and scalability of services, they also increase the complexity of management. For example, it is necessary to maintain a large number of microservice instances, ensure network communication between containers, handle service dependencies, etc. These management tasks require professional skills and tool support, increasing the operation and maintenance costs.
[0072] High system coupling degree: Some systems may have a high coupling degree in design, resulting in the need to adjust the configurations of multiple components or services simultaneously when scaling resources up or down. This coupling not only increases the complexity of the system but also may affect the flexibility and response speed of scaling.
[0073] Insufficient cost consideration: While pursuing resource utilization efficiency and business performance, some systems may neglect cost factors. For example, blindly increasing resource investment during peak resource demands may lead to a sharp rise in costs. Therefore, when designing and implementing a rapid scaling system, it is necessary to comprehensively consider cost - effectiveness and business requirements.
[0074] Based on this, the embodiments of the present invention provide a scaling method, system, electronic device, and storage medium. The scaling method provided by the embodiments of the present invention can be applied to any electronic device with a scaling function. The electronic device may include servers, computers, mobile terminals, etc. In a possible embodiment, the scaling method can be applied to a distributed cluster, which may include multiple electronic devices, specifically, multiple servers. The following describes the solution of the present invention with reference to the accompanying drawings:
[0075] Figure 1 A flowchart of the scaling method provided by the embodiments of the present invention may include the following steps:
[0076] S101: Utilize a preset prediction model to output the resource demand pattern of the target cloud service based on the historical usage data, real - time monitoring data, and usage trend data of the target cloud service. The resource demand pattern includes the capacity usage trend and user request rate of the target cloud service;
[0077] S102: Determine the capacity demand and migration type of the target cloud service based on the resource demand pattern of the target cloud service and the current resource usage data of the target cloud service. The migration type includes scaling down and scaling up;
[0078] S103. When the migration type of the target cloud service is downsizing, based on the resource occupancy of the target cloud service in each server and a preset downsizing policy, determine the servers to be downsized, and delete the target cloud service instances of the target cloud service deployed in the servers to be downsized. The preset downsizing policy is to determine the first preset number of servers with the least resource occupancy of the target cloud service as the servers to be downsized.
[0079] S104. When the migration type of the target cloud service is upsizing, based on the resource surplus in each idle server and a preset upsizing policy, determine the servers to be upsized, and migrate the target cloud service instances to the servers to be upsized. The idle servers are the servers on which the target cloud service is not deployed. The preset upsizing policy is to determine the second preset number of servers with the most resource surplus as the servers to be upsized.
[0080] Applying the embodiments of the present invention, by using a prediction model to predict the resource demand pattern of the target cloud service, the full-automatic prediction of the resource demand is realized, without manual participation. And the prediction model is trained with a large amount of data, and compared with manual work, it can output relatively accurate resource demand prediction data. Then, according to the predicted upsizing or downsizing policy, the target cloud service is automatically downsized or upsized, realizing the full-automatic execution of upsizing and downsizing. And in the upsizing and downsizing policy, the servers to be upsized and the servers to be downsized are determined based on the resource occupancy of the servers, realizing the full utilization of the resource data in the servers, and reducing the service operation cost to a certain extent.
[0081] The following is an exemplary description of the above S101 - S104:
[0082] In a possible embodiment, the upsizing and downsizing of each cloud service deployed in the cluster can be judged according to a preset time period, and this time period can be set according to the actual application scenario, such as 7 days, 10 days, etc. In a possible embodiment, the corresponding cloud service can be upsized or downsized when a capacity warning is received. Exemplarily, a capacity utilization rate range can be preset. When the capacity utilization rate of the cloud service exceeds the upper limit of this capacity utilization rate range, or is lower than the lower limit of this capacity utilization rate range, it can be determined that the current capacity allocation of the cloud service is unreasonable, and a capacity warning needs to be issued to upsize or downsize the cloud service.
[0083] When it is necessary to upsize or downsize the capacity of the cloud service, the historical usage data, real-time monitoring data, and usage trend data of this cloud service can be obtained. For the convenience of description, the cloud service that needs to be upsized or downsized is hereinafter referred to as the target cloud service.
[0084] In a possible embodiment, the resource usage data of each cloud service can be recorded in real time, and the resource usage data can be stored in the historical usage database corresponding to the cloud service identifier. The resource usage data can include resource usage amount, resource usage rate, and corresponding timestamp information. The above resources can include CPU, memory, disk, etc. Correspondingly, the historical usage data of the target cloud service can be obtained from the historical usage database based on the identifier of the target cloud service, that is, the above resource usage data.
[0085] The real-time monitoring data of the target cloud service can be obtained through a cluster monitoring tool, which can be Prometheus, InfluxDB, etc. The real-time monitoring data can include the resource usage amount, resource usage rate, and service deployment data of the target cloud service, etc. The service deployment data can include the server information where each service instance of the target cloud service is located.
[0086] The usage trend data of the target cloud service can be determined based on the number of customers using the target cloud service. The usage trend can include the proportion or quantity of increase or decrease, etc.
[0087] After obtaining the historical usage data, real-time monitoring data, and usage trend data of the target cloud service, they can be input into a pre-prediction model, and the resource demand pattern output by the prediction model can be obtained. The resource demand pattern includes the user request rate and the capacity usage trend of the target cloud service, etc. The user request rate can be QPS (Query Per Second, the number of queries per second), and the capacity usage trend can be the proportion or quantity of the increase or decrease in capacity demand.
[0088] In a possible embodiment, the above prediction model can be pre-trained through the following steps:
[0089] Obtain a training sample set. Each sample data in the training sample set includes the resource usage data of the service instance, the user request rate, and the service scale data; the service scale data is the number of users using the corresponding service. The resource usage data can be the resource usage data corresponding to multiple timestamps, that is, it can be time series data.
[0090] Output each sample data in the training sample set to the initial prediction model, and obtain the capacity usage trend data output by the initial prediction model for the sample data. The capacity usage trend data includes the capacity usage change rate of the service instance after a preset time interval. The structure of this prediction model can be an RNN (Recurrent Neural Network), an LSTM (Long Short-Term Memory), etc., which can be specifically selected according to the actual application scenario, and the present invention does not make specific limitations thereto. The initial prediction model can learn the above time series data to achieve predicting the data of the next moment based on the data of the current moment.
[0091] Train the initial prediction model based on the difference between the capacity usage trend and the actual usage trend of the corresponding sample data until the difference is less than a preset difference threshold to obtain a preset prediction model; wherein, the actual usage trend of the sample data is obtained based on the resource usage data of the sample data and the resource usage data of the service instance corresponding to the sample data after the preset time interval.
[0092] The above difference can be the cosine distance or information entropy loss between the capacity usage trend and the actual usage trend, etc., and the present invention does not make specific limitations thereto. In this step, parameter adjustment methods such as the gradient descent method can be used to train the initial prediction model until the difference converges to obtain the prediction model.
[0093] Through the above technical solution, by collecting multi-source information such as historical usage data, real-time monitoring data, and market trend data of cloud services, the comprehensiveness and accuracy of the prediction model are ensured, and the model can automatically identify and predict the resource demand pattern in the future for a period of time. Moreover, the time series data used in the training process enables the model to enhance the learning of the correspondence between data and time during the prediction process, so that the model can consider various factors such as seasonal changes, periodic fluctuations, and emergencies. In addition, the prediction model has the ability of adaptive learning and optimization, that is, it can dynamically adjust the model parameters according to the latest data and actual operation results to improve the accuracy and timeliness of the prediction.
[0094] After obtaining the resource demand pattern of the target cloud service, the actual capacity demand of the target cloud service and the migration type of the target cloud service can be determined based on the capacity usage trend therein. In a possible embodiment, if the difference between the first product of the capacity usage trend and the current resource usage data of the target cloud service and the second product of the current resource usage data of the target cloud service and the preset resource reservation threshold is greater than 0, it is determined that the target cloud service needs to downsize;
[0095] If the difference between the third product of the capacity usage trend, the current usage data of the target cloud service, and the preset buffer base number and the current resource usage of the target cloud service is greater than 0, it is determined that the target cloud service needs to be scaled out.
[0096] By using a prediction model to predict the capacity change trend of the target cloud service, the average user request rate and capacity demand received by the target cloud service in the next cycle can be obtained through this capacity change trend and the resources currently used by the target cloud service.
[0097] Exemplarily, the capacity demand can be the product of the resource usage decline ratio and the current resource usage of the target cloud service, that is, the above-mentioned first product. If this first product is greater than the minimum resource usage, it means that the capacity of the target cloud service is excessive, and the target cloud service can be scaled in. The minimum resource usage can be the product of the current resource usage of the target cloud service and the minimum resource reservation threshold.
[0098] During the operation of a service instance, it is usually necessary to use a cache to temporarily store user requests. To meet the storage requirements of the target cloud service, a cache buffer can usually be set for each service instance of the target cloud service. Correspondingly, if the difference between the third product of the above-mentioned first product and the preset cache buffer and the current usage of the target cloud service is greater than 0, it means that the current capacity of the target cloud service is insufficient and needs to be scaled out. Specifically, the scaling-out capacity can be obtained by rounding up the ratio of the above-mentioned difference to the resource occupancy of a single instance.
[0099] Exemplarily, if the capacity is expected to increase by 3 times in the next cycle, then the number of instances in each computer room is multiplied by 3 correspondingly, and then multiplied by the reserved buffer. The basic formula is as follows:
[0100] The capacity is excessive and needs to be scaled in: (the resource prediction decline ratio in the next cycle * the current service resource usage - the current service resource usage * the minimum resource reservation threshold) / the resource occupancy of a single instance. If this value is greater than 0, it is adjusted.
[0101] The capacity is insufficient and needs to be scaled out: (the resource prediction decline ratio in the next cycle * the current service resource usage * the resource reservation buffer base number - the current service resource usage) / the resource occupancy of a single instance. If this value is greater than 0, it is adjusted and rounded up.
[0102] After determining the scaling-in capacity or scaling-out capacity of the target cloud service, the scaling-in or scaling-out operation can be executed. In a possible embodiment, a scaling-in strategy and a scaling-out strategy can be preset, and the scaling-in (scaling-out) strategy includes the selection of the scaling-in (scaling-out) server.
[0103] As Figure 2 shown, Figure 2It is a schematic flowchart for determining an expansion server and a contraction server in the scaling method provided by an embodiment of the present invention. Specifically, after receiving a capacity warning, analyze the current capacity requirement, and determine the migration type based on the result of the capacity requirement analysis. If it is determined that the target cloud service needs to be scaled down, sort each server according to the number of service deployments in each server occupied by the target cloud service, so as to determine the top n servers with the fewest service deployments as the contraction servers, and the service instances of the target cloud service deployed in these contraction servers need to be migrated to other servers or deleted.
[0104] If it is determined that the target cloud service needs to be expanded, sort each idle server according to the remaining resources in each idle server in the resource pool, and determine whether the remaining resources in each idle server can meet the capacity requirement of the target cloud service. If not, new servers can be applied for from the standby resource pool, and sort the idle server and the new servers according to the CPU resource utilization rate, and select the top n servers with the lowest CPU usage rate as the expansion servers, and the target cloud service can deploy its service instances in these expansion servers.
[0105] In a possible embodiment, the above method may further include: after the expansion or contraction is completed, determine whether the resource utilization rates in each server in the resource pool meet the preset resource utilization rate range; if there is a server that does not meet the preset resource utilization rate range, modify the preset expansion policy or the preset contraction policy.
[0106] In a possible embodiment, after scaling the target cloud service through the above steps, the service distribution may be uneven, such as distributing multiple service instances in the same server, resulting in a heavy load on this server, while the resources in other servers are not fully utilized. To solve this problem, it is possible to determine the cloud service with an allocation rate lower than the preset threshold, and adjust the deployment of its service instances during the low-traffic peak of this cloud service.
[0107] In a possible embodiment, the deployment of the service instances of the target cloud service can be adjusted through the following steps:
[0108] S105. For each server in the resource pool, determine the number of optimizable machines based on the current resource utilization rate of the server, the number of allocated CPU cores in the server, and the preset resource utilization rate, where the optimizable machine is a server with an allocation rate less than the preset allocation rate threshold.
[0109] In a possible embodiment, the above number of optimizable machines can be obtained through the following formula:
[0110] needCpu = useCpu * current utilization rate / standard utilization rate
[0111] Optimizable machine number = (allCpu * rate - needCpu * buffer) / 32
[0112] Among them, allCpu is the total number of cores in the resource pool, useCpu is the number of cores already allocated in the resource pool, and needCpu is the total actual CPU demand of the resource pool, which can be obtained by summing up the capacity requirements of each service instance. Rate is the utilization rate, and the standard utilization rate can be set according to the actual application scenario. For example, it can be 12%. Exemplarily, if the used capacity in the resource pool is 320c, the CPU utilization rate is 3%, and the standard utilization rate is 12%, then the actual CPU demand is 320 * 3% / 12% = 80c. The total number of cores * rate is the expected allocation rate, needCpu * buffer is the number of cores required by the service plus the reserved buffer. The difference between the two is the idle optimizable number of cores, and dividing by 32 is the number of machines. 32 represents that the machine is a single machine with 32 cores.
[0113] S106. Determine the servers to be taken offline based on the deployment information in each of the servers in the resource pool and the optimizable machine number, where the deployment information includes the number of service deployments;
[0114] S107. Sort the remaining servers in the resource pool in descending order according to the remaining resources, and migrate the service instances deployed in the servers to be taken offline to each of the remaining servers in sequence, where the remaining servers are other servers in the resource pool except the servers to be taken offline.
[0115] In a possible embodiment, the service instances can be sorted in descending order according to the packages used by the service instances in each server with an allocation rate lower than the preset threshold. The above package refers to the amount of resources occupied by the service instance. Then, S107 can be executed for each service instance in sequence based on this sorting.
[0116] In a possible embodiment, the migrating the service instances deployed in the servers to be taken offline to each of the remaining servers in sequence includes
[0117] For each service instance in the servers to be taken offline, obtain the attribute information of the service instance, where the attribute information includes port information, deployment file information, and traffic peak information;
[0118] Based on the attribute information of the service instance and the deployment attribute information in each of the remaining servers, migrate the service instance to the remaining server that does not conflict with the service instance in sequence according to the sorting. The non - conflict with the service instance means different ports, different deployment file storage paths, and different traffic peaks from the service instance.
[0119] The above deployment attribute information may include the identification of each service instance deployed in the remaining servers, the port information used by the service instance, the storage location of the deployment file, the traffic peak and trough information, and so on. If the service instances in the same server use the same port, port conflicts will occur, which can easily cause the service to fail to run properly, so they cannot be deployed in the same server. If the deployment file storage paths of different service instances in the same server are the same, it may cause the service instance to fail to query or store data, thus affecting the service operation. Therefore, they cannot be deployed in the same server either. If the traffic peak times of different service instances in the same server are the same, resource preemption may occur or the server resources may be insufficient, which can easily cause the service operation to be abnormal. Therefore, they cannot be deployed in the same server either.
[0120] By deploying the service instances in servers without conflicts with them, the normal operation of the service instances is ensured, thereby improving the stability of the service operation.
[0121] As Figure 3 shown, Figure 3 FIG. is a schematic flowchart of adjusting the service instance deployment in the scaling method provided by the embodiment of the present invention, which may include the following steps:
[0122] S301. Obtain a list of machines to be taken offline.
[0123] S302. Obtain the instance information on the offline server.
[0124] S303. Sort the instances in all the servers to be taken offline in descending order according to the package.
[0125] S304. Traverse the list of machines to be taken offline, and sort the remaining non-offline servers in descending order according to the remaining CPU.
[0126] S305. For each server to be taken offline, traverse all the non-offline servers according to the above sorting, and determine whether there are instance conflicts between the non-offline server and each service instance in the offline server. If not, adjust the original service instance to the non-offline server. If so, repeat S305.
[0127] S306. Determine whether the traversal is completed. If it is completed, end the process. If not, return to S304.
[0128] Applying the embodiment of the present invention, through a high-precision prediction algorithm, the capacity adjustment requirements are output. The capacity analysis robot produces the current capacity adjustment plan according to each threshold set in the initial capacity and the analysis of the capacity growth trend. The intelligent orchestration robot orchestrates the current service deployment according to the capacity adjustment plan, and balances the cost and efficiency in the optimal deployment manner.
[0129] Furthermore, the multi-dimensional prediction model: In addition to the basic resource utilization rate and performance metrics, the present invention also considers multi-dimensional factors such as user behavior and business trends to construct a more comprehensive prediction model. This multi-dimensional prediction model can more accurately depict the dynamic changes in resource requirements and provide more reliable data support for intelligent scheduling. The present invention adopts an adaptive optimization algorithm that can automatically formulate and execute resource scheduling strategies based on the prediction results and the current resource status. This algorithm can adjust the resource allocation plan in real time to ensure the maximization of resource utilization while meeting business requirements. At the same time, the algorithm also has the ability of self-learning and optimization, and can continuously improve the scheduling efficiency and accuracy as the system runs for a longer time.
[0130] In addition, in the low-traffic scenario, the intelligent orchestration robot can determine whether the current deployment meets the deployment resource requirements, dynamically adjust the layout, re-orchestrate the resources, and optimize the empty spaces to achieve the optimal resource utilization rate and allocation rate.
[0131] In summary, the present invention realizes a rapid scaling mechanism for cloud resources, which can complete the addition or release of resources in a short time. This mechanism can quickly respond to changes in business requirements, ensure that the system can automatically increase resources under high load to avoid performance bottlenecks, and automatically release resources under low load to reduce operating costs.
[0132] As Figure 4 shown, Figure 4 For a system structure diagram for implementing the scaling method provided by the embodiments of the present invention, it may include a data collection layer, a prediction model layer, an intelligent scheduling layer (corresponding to the model & scheduling layer in Figure 4 ), an execution layer, and a monitoring and feedback layer (not shown in the figure). Among them, the data collection layer is responsible for collecting information from various data sources and providing data support for the prediction model. The prediction model layer is used to construct and run a prediction model to accurately predict resource requirements. The intelligent scheduling layer is used to formulate and execute resource scheduling strategies based on the prediction results and the current resource status. The execution layer is responsible for actually executing the scaling operation of resources and adjusting the number and configuration of cloud service instances. The monitoring and feedback layer is used to monitor service performance and resource status in real time and collect feedback information for optimizing the prediction model and scheduling strategies. Users can trigger the scaling task through the cloud-native PaaS platform. Specifically, the cloud-native PaaS can perform resource scheduling through a unified resource scheduling entry, and this cloud-native PaaS platform can also implement functions such as service monitoring, capacity management, model management, fault migration, and configuration management.
[0133] Applying the embodiments of the present invention can achieve the following technical effects:
[0134] 1. Improve resource utilization efficiency: Through dynamic resource prediction technology, analyze business requirements and resource usage in real time, automatically adjust resource allocation to ensure full utilization of resources and avoid waste.
[0135] 2. Enhance system response speed and flexibility: Adopt intelligent scheduling technology, quickly formulate and execute resource scheduling strategies based on prediction results and current resource status to ensure that the system can quickly respond to changes in business requirements, and improve the flexibility and response speed of the system.
[0136] 3. Reduce operation and maintenance costs: Through automated and intelligent resource management, reduce manual intervention, lower operation and maintenance costs, and improve the stability and reliability of the system at the same time.
[0137] 4. Optimize the user experience: Through fast and accurate resource scaling, ensure the stability and continuity of service performance, and enhance the user experience.
[0138] 5. Support complex and changeable business scenarios: Build a flexible and scalable system architecture, support multiple business scenarios and resource types, meet the personalized needs of different users, and enable the system to cope with new business scenarios and challenges that may arise in the future.
[0139] In summary, the application of the embodiments of the present invention aims to solve a series of problems in traditional cloud service management through technological innovation, and promote the development of cloud computing services towards a more efficient, flexible, lower-cost, and better-experience direction.
[0140] Based on the same inventive concept, the embodiments of the present invention also provide a scaling system, as Figure 5 shown. The scaling system 500 may include:
[0141] A prediction module 501, configured to use a preset prediction model to output a resource demand pattern of the target cloud service based on historical usage data, real-time monitoring data, and usage trend data of the target cloud service. The resource demand pattern includes a capacity usage trend and a user request rate of the target cloud service;
[0142] A determination module 502, configured to determine a capacity demand and a migration type of the target cloud service based on the resource demand pattern of the target cloud service and the current resource usage data of the target cloud service. The migration type includes scaling down and scaling up;
[0143] The migration module 503 is used to determine the shrinking servers based on the resource occupancy of the target cloud service in each server and a preset shrinking policy when the migration type of the target cloud service is shrinking, and delete the target cloud service instances of the target cloud service deployed in the shrinking servers. The preset shrinking policy is to determine the first preset number of servers with the least resource occupancy of the target cloud service as the shrinking servers.
[0144] When the migration type of the target cloud service is expansion, determine the expanding servers based on the resource surplus in each idle server and a preset expansion policy, and migrate the target cloud service instances to the expanding servers. The idle servers are the servers on which the target cloud service is not deployed. The preset expansion policy is to determine the second preset number of servers with the most resource surplus as the expanding servers.
[0145] In a possible embodiment, the preset prediction model is pre-trained through the following steps:
[0146] Obtain a training sample set, where each sample data in the training sample set includes the resource usage data of the service instance, the user request rate, and the service scale data.
[0147] Output each sample data in the training sample set to the initial prediction model, and obtain the capacity usage trend data output by the initial prediction model for the sample data. The capacity usage trend data includes the capacity usage change rate of the service instance after a preset time interval.
[0148] Train the initial prediction model based on the difference between the capacity usage trend and the actual usage trend of the corresponding sample data until the difference is less than a preset difference threshold to obtain the preset prediction model. The actual usage trend of the sample data is obtained based on the resource usage data of the sample data and the resource usage data of the corresponding service instance after the preset time interval.
[0149] The determining the capacity requirement and the migration type of the target cloud service based on the resource demand pattern of the target cloud service and the current resource usage data of the target cloud service includes:
[0150] If the difference between the first product of the capacity usage trend and the current resource usage data of the target cloud service and the second product of the current resource usage data of the target cloud service and the preset resource reservation threshold is greater than 0, it is determined that the target cloud service needs to be shrunk.
[0151] If the difference between the third product of the capacity usage trend, the current usage data of the target cloud service, and the preset buffer base number and the current resource usage of the target cloud service is greater than 0, it is determined that the target cloud service needs to be scaled out;
[0152] The system further includes:
[0153] An adjustment module, configured to determine whether the resource utilization rates of the servers in the resource pool meet the preset resource utilization rate range after the scaling out or scaling in is completed;
[0154] If there is a server that does not meet the preset resource utilization rate range, modify the preset scaling out policy or the preset scaling in policy;
[0155] A reallocation module, configured to determine the number of optimizable machines for each server in the resource pool based on the current resource utilization rate of the server, the number of allocated CPU cores in the server, and the preset resource utilization rate, where the optimizable machine is a server with an allocation rate less than the preset allocation rate threshold;
[0156] Based on the deployment information in each server in the resource pool and the number of optimizable machines, determine the servers to be taken offline, where the deployment information includes the number of service deployments;
[0157] Sort the remaining servers in the resource pool in descending order of the remaining resources, and migrate the service instances deployed in the servers to be taken offline to each of the remaining servers in sequence according to the sorting, where the remaining servers are the other servers in the resource pool except the servers to be taken offline;
[0158] The migrating the service instances deployed in the servers to be taken offline to each of the remaining servers in sequence according to the sorting includes
[0159] For each service instance in the servers to be taken offline, obtain the attribute information of the service instance, where the attribute information includes port information, deployment file information, and traffic peak information;
[0160] Based on the attribute information of the service instance and the deployment attribute information in each of the remaining servers, migrate the service instance to the remaining server that does not conflict with the service instance according to the sorting, where not conflicting with the service instance means having a different port from the service instance, a different storage path for the deployment file, and a different traffic peak.
[0161] Among them, in the present invention, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information and other processing all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0162] An exemplary embodiment of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiment of the present invention.
[0163] An exemplary embodiment of the present invention also provides a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present invention.
[0164] An exemplary embodiment of the present invention also provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present invention.
[0165] Reference Figure 6 , the structural block diagram of the electronic device 600 that can be used as a server or a client of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0166] As Figure 6 shown, the electronic device 600 includes a computing unit 601, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 602 or the computer program loaded from the storage unit 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0167] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. The input unit 606 can be any type of device capable of inputting information into the electronic device 600. The input unit 606 can receive input numerical or character information and generate key signal inputs related to user settings and / or function controls of the electronic device. The output unit 607 can be any type of device capable of presenting information and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 608 can include, but is not limited to, magnetic disks and optical discs. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0168] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above. For example, in some embodiments, any of the above-described scaling methods can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. In some embodiments, the computing unit 601 can be configured to execute any of the above-described scaling methods by any other suitable means (e.g., by means of firmware).
[0169] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0170] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0171] As used in the present invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0172] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0173] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0174] A computer system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other.
Claims
1. A method for scaling up and down, characterized in that: The method comprises: Using a preset prediction model, based on historical usage data, real-time monitoring data, and usage trend data of the target cloud service, output a resource demand pattern of the target cloud service, wherein the resource demand pattern includes a capacity usage trend and a user request rate of the target cloud service; Determine the capacity requirement and migration type of the target cloud service based on the resource demand pattern of the target cloud service and the current resource usage data of the target cloud service, where the migration type includes shrinking and expanding. In the case where the migration type of the target cloud service is scaling down, based on the resource usage of the target cloud service in each server and a preset scaling down strategy, a scaling down server is determined, and the target cloud service instance of the target cloud service deployed in the scaling down server is deleted, wherein the preset scaling down strategy is to determine a first preset number of servers with the least resource usage of the target cloud service as the scaling down servers; In the case where the migration type of the target cloud service is capacity expansion, the capacity expansion server is determined based on the resource margin in each idle server and the preset capacity expansion strategy, and the target cloud service instance is migrated to the capacity expansion server, wherein the idle server is a server on which the target cloud service is not deployed, and the preset capacity expansion strategy is to determine a second preset number of servers with the largest resource margin as the capacity expansion server.
2. The method according to claim 1, characterized in that The preset prediction model is pre-trained by the following steps: Acquire a training sample set, wherein each piece of sample data in the training sample set includes resource usage data of a service instance, a user request rate, and service scale data; Outputting each piece of sample data in the training sample set to an initial prediction model, and obtaining capacity usage trend data output by the initial prediction model for the sample data, wherein the capacity usage trend data includes a capacity usage change rate of the service instance after a preset time interval; Based on the difference between the capacity usage trend and the actual usage trend of the corresponding sample data, the initial prediction model is trained until the difference is less than a preset difference threshold, thereby obtaining a preset prediction model; wherein the actual usage trend of the sample data is obtained based on the resource usage data of the sample data and the resource usage data of the service instance corresponding to the sample data after the preset time interval.
3. The method according to claim 1, characterized in that The determining the capacity requirement and migration type of the target cloud service based on the resource requirement pattern of the target cloud service and the current resource usage data of the target cloud service includes: If the difference between the first product of the capacity usage trend and the current resource usage data of the target cloud service and the second product of the current resource usage data of the target cloud service and the preset resource reservation threshold is greater than 0, it is determined that the target cloud service needs to be scaled down; If the difference between the third product of the capacity usage trend, the current usage data of the target cloud service, and the preset buffer base and the current resource usage of the target cloud service is greater than 0, it is determined that the target cloud service needs to be expanded.
4. The method according to claim 1, characterized in that: The method further comprises: After the expansion or reduction is completed, determine whether the resource utilization rate of each server in the resource pool meets the preset resource utilization rate range; If there are servers that do not meet the preset resource usage range, the preset expansion strategy or the preset reduction strategy is modified.
5. The method according to claim 1, characterized in that The method further comprises: For each server in the resource pool, based on the current resource utilization of the server, the number of CPU cores allocated in the server, and the preset resource utilization, determine the number of machines that can be optimized, wherein the machines that can be optimized are servers whose allocation rate is less than the preset allocation rate threshold; Determine the server to be taken offline based on the deployment information of each server in the resource pool and the number of optimizable machines, wherein the deployment information includes the number of service deployments; The remaining servers in the resource pool are sorted in descending order according to the remaining amount of resources, and the service instances deployed in the servers to be taken offline are migrated to each of the remaining servers in sequence according to the sorting, wherein the remaining servers are other servers in the resource pool except the servers to be taken offline.
6. The method according to claim 5, characterized in that The step of migrating the service instances deployed in the server to be taken offline to the remaining servers in sequence according to the sequence includes: For each service instance in the server to be taken offline, obtaining attribute information of the service instance, wherein the attribute information includes port information, deployment file information, and traffic peak information; Based on the attribute information of the service instance and the deployment attribute information in each of the remaining servers, the service instance is migrated to the remaining servers that have no conflict with the service instance in accordance with the sorting, where no conflict with the service instance means that the service instance has a different port, a different deployment file storage path, and a different traffic peak.
7. A capacity expansion and contraction system, characterized in that: The system comprises: A prediction module, configured to output a resource demand pattern of the target cloud service by using a preset prediction model based on historical usage data, real-time monitoring data, and usage trend data of the target cloud service, wherein the resource demand pattern includes a capacity usage trend and a user request rate of the target cloud service; A determination module, configured to determine the capacity requirement and migration type of the target cloud service based on the resource requirement pattern of the target cloud service and the current resource usage data of the target cloud service, wherein the migration type includes shrinking and expanding; a migration module, configured to determine a reduced capacity server based on the resource usage of the target cloud service in each server and a preset reduction strategy when the migration type of the target cloud service is reduction, and delete the target cloud service instance of the target cloud service deployed in the reduced capacity server, wherein the preset reduction strategy is to determine a first preset number of servers with the least resource usage of the target cloud service as reduced capacity servers; In the case where the migration type of the target cloud service is capacity expansion, the capacity expansion server is determined based on the resource margin in each idle server and the preset capacity expansion strategy, and the target cloud service instance is migrated to the capacity expansion server, wherein the idle server is a server on which the target cloud service is not deployed, and the preset capacity expansion strategy is to determine a second preset number of servers with the largest resource margin as the capacity expansion server.
8. The system according to claim 7, characterized in that The preset prediction model is pre-trained by the following steps: Acquire a training sample set, wherein each piece of sample data in the training sample set includes resource usage data of a service instance, a user request rate, and service scale data; Outputting each piece of sample data in the training sample set to an initial prediction model, and obtaining capacity usage trend data output by the initial prediction model for the sample data, wherein the capacity usage trend data includes a capacity usage change rate of the service instance after a preset time interval; Based on the difference between the capacity usage trend and the actual usage trend of the corresponding sample data, the initial prediction model is trained until the difference is less than a preset difference threshold, thereby obtaining a preset prediction model; wherein the actual usage trend of the sample data is obtained based on the resource usage data of the sample data and the resource usage data of the service instance corresponding to the sample data after the preset time interval; The determining the capacity requirement and migration type of the target cloud service based on the resource requirement pattern of the target cloud service and the current resource usage data of the target cloud service includes: Based on the difference between the first product of the capacity usage trend and the current resource usage data of the target cloud service and the second product of the current resource usage data of the target cloud service and a preset resource reservation threshold being greater than 0, determining that the target cloud service needs to be scaled down; If the difference between the third product of the capacity usage trend, the current usage data of the target cloud service, and the preset buffer base and the current resource usage of the target cloud service is greater than 0, it is determined that the target cloud service needs to be expanded; The system further comprises: The adjustment module is used to determine whether the resource utilization rate of each server in the resource pool meets the preset resource utilization rate range after the expansion or reduction is completed; If there is a server that does not meet the preset resource usage rate range, modify the preset expansion strategy or the preset reduction strategy; A reallocation module is used to determine the number of optimizable machines for each server in the resource pool based on the current resource utilization of the server, the number of CPU cores allocated in the server, and the preset resource utilization, wherein the optimizable machine is a server whose allocation rate is less than a preset allocation rate threshold; Determine the server to be taken offline based on the deployment information of each server in the resource pool and the number of optimizable machines, wherein the deployment information includes the number of service deployments; The remaining servers in the resource pool are sorted in descending order according to the remaining amount of resources, and the service instances deployed in the server to be taken offline are sequentially migrated to each of the remaining servers according to the sorting, wherein the remaining servers are other servers in the resource pool except the server to be taken offline; The step of migrating the service instances deployed in the server to be taken offline to the remaining servers in sequence according to the sequence includes: For each service instance in the server to be taken offline, obtaining attribute information of the service instance, wherein the attribute information includes port information, deployment file information, and traffic peak information; Based on the attribute information of the service instance and the deployment attribute information in each of the remaining servers, the service instance is migrated to the remaining servers that have no conflict with the service instance in accordance with the sorting, where no conflict with the service instance means that the service instance has a different port, a different deployment file storage path, and a different traffic peak.
9. An electronic device, comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 6.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1-6.
Citation Information
Cited By
Cloud platform capacity expansion and contraction method, device, equipment and medium
CN120896855A
Resource capacity expansion and contraction method and system
CN120896866A
A method and system for resource scaling
CN120896866B
Cloud resource optimization method and device, computer equipment and storage medium
CN121125528A
Resource processing method and device of resource node, electronic equipment and storage medium
CN121441919A