Method, apparatus, device, and storage medium for evaluating the capacity of a cloud computing system

Through fine-grained resource information evaluation and load balancing modules, the interval is set according to server category and storage space, the problem of low evaluation accuracy in cloud computing systems is solved, and efficient allocation and utilization of resources is achieved.

CN113010576BActive Publication Date: 2025-07-22CHINA CONSTRUCTION BANK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110297121.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-19
Publication Date
2025-07-22
Estimated Expiration
2041-03-19

AI Technical Summary

Technical Problem

The existing cloud computing system capacity evaluation methods lack distinction between different server uses and storage space, resulting in low evaluation accuracy and inability to effectively deal with fluctuations in data volume, affecting resource allocation and utilization.

Method used

By obtaining the historical operation data of the server, setting usage rate and data volume intervals according to the server category and storage space, fine-grained resource information evaluation is carried out, overcapacity or insufficient capacity is identified, and resource expansion or recycling is carried out through the load balancing module to improve the evaluation accuracy.

Benefits of technology

It realizes rapid and accurate capacity evaluation of cloud computing systems, improves resource utilization and system performance, and adapts to resource allocation needs in different business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113010576B_ABST
    Figure CN113010576B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device, and storage medium for evaluating the capacity of a cloud computing system. The method includes obtaining the historical operation data of the servers in the cloud computing system, specifically including the peak CPU usage rate and the peak data volume of the present server within a historical period; for analyzable servers (i.e., servers for which historical operation data has been successfully obtained), reading the usage rate interval set according to the server category to which they belong, and the data volume interval set according to the total storage space; using the historical operation data, usage rate interval, and data volume interval of the analyzable servers to evaluate resource information, and obtaining the evaluation results of the servers; counting the proportion of servers with normal capacity in the evaluation results to obtain the normal rate, and determining the capacity evaluation result of the cloud computing system according to the normal rate. This solution divides different categories based on the different uses of the servers, sets different capacity evaluation criteria for servers of different categories and storage spaces, and improves the accuracy of the capacity evaluation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a method, apparatus, device, and storage medium for evaluating the capacity of a cloud computing system. Background Art

[0002] A cloud computing system is a large computing system composed of multiple servers and can be used to execute large-scale data analysis and processing tasks. During the actual operation of the cloud computing system, the actual amount of data to be processed may fluctuate. For example, the amount of data may be relatively large in some months and then significantly decrease in the following months.

[0003] To cope with the data volume fluctuations, it is necessary to regularly evaluate the capacity of the cloud computing system to check whether the total amount of resources (generally including processor resources and storage resources) of the cloud computing system meets the recent processing requirements, so as to adjust the total amount of resources of the cloud computing system according to the evaluation results.

[0004] In existing capacity evaluation methods, generally a unified resource threshold is set. When the amount of resources recently used by most servers in the cloud computing system exceeds the resource threshold, it is considered that the cloud computing system needs to be expanded.

[0005] Each server in the cloud computing system often has different uses, so the resource usage situations of each server are also different. Therefore, the existing method of evaluating all servers in the cloud computing system according to a unified standard has a low accuracy. Summary of the Invention

[0006] Aiming at the above-mentioned disadvantages of the prior art, the present invention provides a method, apparatus, device, and storage medium for evaluating the capacity of a cloud computing system to improve the accuracy of evaluating the capacity of the cloud computing system.

[0007] The first aspect of this application provides a method for evaluating the capacity of a cloud computing system. The cloud computing system is composed of multiple servers, and the method includes:

[0008] Obtain the historical operation data of each server in the cloud computing system; wherein, the historical operation data of the server includes the peak CPU usage rate and the peak data volume of this server within a preset historical period;

[0009] For each analyzable server, read the usage rate interval set according to the server category to which the analyzable server belongs, and the data volume interval set according to the total storage space of the analyzable server; wherein, the analyzable server refers to the server in the cloud computing system for which historical operation data has been successfully obtained; the server categories include web servers, database servers, online application servers, non-online application servers, and data analysis servers;

[0010] For each of the analyzable servers, using the historical operation data, usage rate range, and data volume range of the analyzable server, conduct a resource information assessment on the analyzable server to obtain the assessment result of the analyzable server; wherein, the assessment result includes overcapacity, normal capacity, and undercapacity.

[0011] Count the proportion of servers with a normal capacity assessment result among all the analyzable servers in the cloud computing system to obtain the normal rate of the cloud computing system.

[0012] Determine the capacity assessment result of the cloud computing system according to the normal rate of the cloud computing system.

[0013] Optionally, the conducting a resource information assessment on the analyzable server using the historical operation data, usage rate range, and data volume range of the analyzable server to obtain the assessment result of the analyzable server includes:

[0014] Judge whether the peak CPU usage rate of the analyzable server is within the usage rate range, and judge whether the peak data volume of the analyzable server is within the data volume range.

[0015] If the peak CPU usage rate of the analyzable server is within the usage rate range and the peak data volume of the analyzable server is within the data volume range, determine that the assessment result of the analyzable server is normal capacity.

[0016] If the peak CPU usage rate of the analyzable server is greater than the upper limit of the usage rate range and the peak data volume of the analyzable server is greater than the upper limit of the data volume range, determine that the assessment result of the analyzable server is undercapacity.

[0017] If the peak CPU usage rate of the analyzable server is less than the lower limit of the usage rate range and the peak data volume of the analyzable server is less than the lower limit of the data volume range, determine that the assessment result of the analyzable server is overcapacity.

[0018] Optionally, after obtaining the historical operation data of each server in the cloud computing system, it further includes:

[0019] For each of the analyzable servers, calculate the load index of the analyzable server using the historical operation data of the analyzable server.

[0020] For each of the analyzable servers, calculate the load weight value and load redundancy value of the analyzable server according to the load index and performance index of the analyzable server; wherein, the load weight value and load redundancy value of the analyzable server are used as the basis for allocating requests in the cloud computing system.

[0021] Optionally, after using the historical operation data, usage rate interval, and data volume interval of the analyzable server to evaluate the resource information of the analyzable server for each of the analyzable servers and obtaining the evaluation result of the analyzable server, it further includes:

[0022] Identify the analyzable servers with an evaluation result of overcapacity and output a recycling prompt message; wherein, the recycling prompt message is used to indicate to perform a resource recycling operation on the analyzable servers with an evaluation result of overcapacity;

[0023] Identify the analyzable servers with an evaluation result of undercapacity and output an expansion prompt message; wherein, the expansion prompt message is used to indicate to perform a resource expansion operation on the analyzable servers with an evaluation result of overcapacity.

[0024] The second aspect of the present application provides a device for evaluating the capacity of a cloud computing system. The cloud computing system is composed of multiple servers. The device includes:

[0025] An acquisition unit for acquiring the historical operation data of each server in the cloud computing system; wherein, the historical operation data of the server includes the CPU usage peak value and data volume peak value of the server itself within a preset historical period;

[0026] A reading unit for reading, for each analyzable server, the usage rate interval set according to the server category to which the analyzable server belongs, and the data volume interval set according to the total storage space of the analyzable server; wherein, the analyzable server refers to the server in the cloud computing system that has successfully obtained historical operation data; the server categories include web servers, database servers, online application servers, non-online application servers, and data analysis servers;

[0027] An evaluation unit for evaluating the resource information of each analyzable server by using the historical operation data, usage rate interval, and data volume interval of the analyzable server to obtain the evaluation result of the analyzable server; wherein, the evaluation result includes overcapacity, normal capacity, and undercapacity;

[0028] A statistics unit for counting the proportion of servers with a normal capacity evaluation result among all analyzable servers in the cloud computing system to obtain the normal rate of the cloud computing system;

[0029] A determination unit, configured to determine a capacity evaluation result of the cloud computing system according to the normal rate of the cloud computing system.

[0030] Optionally, when the evaluation unit performs resource information evaluation on the analyzable server according to the historical operation data, usage rate interval, and data volume interval of the analyzable server to obtain an evaluation result of the analyzable server, it is specifically configured to:

[0031] Judge whether the peak CPU usage rate of the analyzable server is within the usage rate interval, and judge whether the peak data volume of the analyzable server is within the data volume interval;

[0032] If the peak CPU usage rate of the analyzable server is within the usage rate interval and the peak data volume of the analyzable server is within the data volume interval, determine that the evaluation result of the analyzable server is normal in capacity;

[0033] If the peak CPU usage rate of the analyzable server is greater than the upper limit of the usage rate interval and the peak data volume of the analyzable server is greater than the upper limit of the data volume interval, determine that the evaluation result of the analyzable server is insufficient in capacity;

[0034] If the peak CPU usage rate of the analyzable server is less than the lower limit of the usage rate interval and the peak data volume of the analyzable server is less than the lower limit of the data volume interval, determine that the evaluation result of the analyzable server is excessive in capacity.

[0035] Optionally, the device further includes:

[0036] A first calculation unit, configured to calculate a load index of the analyzable server for each of the analyzable servers by using the historical operation data of the analyzable server;

[0037] A second calculation unit, configured to calculate a load weight value and a load redundancy value of the analyzable server for each of the analyzable servers according to the load index and performance index of the analyzable server; wherein, the load weight value and the load redundancy value of the analyzable server are used as a basis for allocating requests in the cloud computing system.

[0038] Optionally, the evaluation unit is further configured to:

[0039] Identify an analyzable server with an evaluation result of excessive capacity and output a recycling prompt message; wherein, the recycling prompt message is used to indicate to perform a resource recycling operation on the analyzable server with an evaluation result of excessive capacity;

[0040] Identify analyzable servers with an evaluation result of insufficient capacity and output an expansion prompt message; wherein, the expansion prompt message is used to indicate to perform a resource expansion operation on the analyzable server with an evaluation result of excessive capacity.

[0041] The third aspect of this application provides a computer storage medium for storing a computer program, which when executed, is specifically used to implement the method for evaluating the capacity of a cloud computing system provided in any item of the first aspect of this application.

[0042] The fourth aspect of this application provides an electronic device, including a memory and a processor;

[0043] Wherein, the memory is used to store a computer program;

[0044] The processor is used to execute the computer program, and is specifically used to implement the method for evaluating the capacity of a cloud computing system provided in any item of the first aspect of this application.

[0045] This application provides a method, device, equipment and storage medium for evaluating the capacity of a cloud computing system, which is applicable to a cloud computing system composed of multiple servers. The method includes obtaining the historical operation data of each server in the cloud computing system; wherein, the historical operation data of the server includes the peak CPU usage rate and the peak data volume of this server within a preset historical period; for each analyzable server, read the usage rate interval set according to the server category to which the analyzable server belongs, and the data volume interval set according to the total storage space of the analyzable server; wherein, the analyzable server refers to the server in the cloud computing system that has successfully obtained the historical operation data; the server categories include web servers, database servers, online application servers, non-online application servers and data analysis servers; for each analyzable server, use the historical operation data, usage rate interval and data volume interval of the analyzable server to evaluate the resource information of the analyzable server, and obtain the evaluation result of the analyzable server; wherein, the evaluation result includes excessive capacity, normal capacity and insufficient capacity; count the proportion of servers with an evaluation result of normal capacity among all analyzable servers in the cloud computing system to obtain the normal rate of the cloud computing system; determine the capacity evaluation result of the cloud computing system according to the normal rate of the cloud computing system. This solution divides different categories based on the different uses of the servers, and then sets different capacity evaluation criteria for servers of different categories and different storage spaces, thereby improving the accuracy of the capacity evaluation results. Description of the Drawings

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided accompanying drawings.

[0047] Figure 1 Schematic diagram of the architecture of a cloud computing system provided by an embodiment of the present application;

[0048] Figure 2 Flowchart of a method for capacity evaluation of a cloud computing system provided by an embodiment of the present application;

[0049] Figure 3 Flowchart of a method for capacity evaluation of a cloud computing system provided by another embodiment of the present application;

[0050] Figure 4 Schematic diagram of the structure of a device for capacity evaluation of a cloud computing system provided by an embodiment of the present application;

[0051] Figure 5 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0053] With the development of computer technology, cloud computing systems (cloud computing platforms) are widely used in various fields of society involving large-scale data processing. For example, the data centers of large banks are configured with cloud computing systems.

[0054] During the operation of the cloud computing system in the data center of a large bank, due to the uneven distribution of computing resources and storage resources in some systems, problems such as high energy consumption and low resource utilization will occur. To solve this problem, it is often necessary to evaluate the capacity of the cloud computing system and perform load balancing on the cloud computing system based on the capacity evaluation results of the cloud computing system.

[0055] Existing methods for capacity evaluation of cloud computing systems generally regard all servers in computing resources as the same deployment unit, thus setting a unified evaluation standard. For example, a capacity threshold is set, and then the resource information of all servers in the cloud computing system is evaluated according to this capacity threshold. Finally, the capacity evaluation result of the entire cloud computing system is obtained by combining the evaluation results of all servers. However, it does not distinguish servers in different scenarios such as online, offline, and data analysis. At the same time, it also does not distinguish file systems with different storage space sizes on different servers, resulting in problems such as poor adaptability and low accuracy, and thus cannot be well applied to multi-system business scenarios.

[0056] In view of the above problems, the present invention proposes a method for capacity evaluation of a cloud computing system. By extracting the peak resource utilization numbers of each server in the system within one month, capacity evaluation, resource expansion or recycling of the cloud computing system are carried out at the server level and the system level by using three modules: resource information collection, resource information evaluation, and resource load balancing, which can provide accurate and rapid capacity analysis solutions for operation and maintenance personnel, enabling them to reasonably allocate resources for servers and systems in different business scenarios.

[0057] Compared with traditional load balancing algorithms based on CPU and disk utilization, the method for capacity evaluation of a cloud computing system proposed in this patent can quickly generate capacity evaluation results, facilitating operation and maintenance personnel to perform resource expansion or recycling.

[0058] The capacity evaluation method provided in this application can be applied to a cloud computing system as Figure 1 shown. As Figure 1 shown, the system is divided into three layers from bottom to top: CVM (Cloud Virtual Machine), System, and Service, and is divided into three modules from left to right: resource information collection, resource information evaluation, and resource load balancing. Among them, the CVM layer and the System layer constitute the main part of the cloud computing system. Each computing task that needs to be processed received by the cloud computing system is processed by the cloud virtual machines in the CVM layer (i.e., Figure 1 C1, C2... Ci shown, equivalent to the servers in the cloud computing system), and multiple cloud virtual machines form a computer cluster in the cloud computing system, i.e., Figure 1 S1, S2... Si shown. The resource information evaluation module specifically includes a server module and a system module.

[0059] In the process of capacity analysis of multiple systems in a cloud computing system, it is divided into four steps. First, the resource information collection module collects the historical operation data of each server in the above-mentioned cloud virtual machine layer. Second, the server module in the resource information evaluation module, according to the category of the server (mainly determined according to the function of the server and the level it belongs to in the network), uses the index threshold method to select the CPU utilization rate and disk utilization rate as capacity indicators, and combines the index threshold range of the server to obtain the capacity evaluation results of each server. Further, the system module in the resource information evaluation module analyzes all the servers under the jurisdiction of the system, summarizes the server capacity evaluation results, and obtains the system capacity evaluation results according to the index threshold range of the system evaluation. Finally, through the resource load balancing module, the computing resources and storage resources of the server are expanded or recycled.

[0060] Through the organic combination of three levels and three major modules, this patent can quickly generate a capacity analysis model for evaluating servers and systems. However, the existing algorithms based on load balancing cannot perform reasonable capacity evaluation on the system, and there are some unreasonable problems in the server resource scheduling, resulting in the ineffective utilization of server resources.

[0061] This patent adopts an improved load balancing algorithm to conduct fine-grained capacity evaluation on servers and systems with different uses, and finally forms a set of operationally convenient and simple capacity analysis model framework. As a back-end algorithm module, this patent only needs to obtain the peak data of CPU and disk utilization rates of the servers under the jurisdiction of the system within one month, divide the servers into different deployment units such as online, non-online, and data analysis according to different applications, determine the corresponding index threshold range, and can display the results in real time on the front-end interface for the reference of operation and maintenance personnel.

[0062] Next, in combination with Figure 1 the system architecture shown, the specific implementation process of the method for capacity evaluation of the cloud computing system provided by this application will be described.

[0063] Please refer to Figure 2 , a method for capacity evaluation of a cloud computing system provided by an embodiment of this application may specifically include the following steps:

[0064] S201. Obtain the historical operation data of each server in the cloud computing system.

[0065] Among them, the historical operation data of the server includes the peak CPU usage rate and the peak data volume of this server within a preset historical period. The preset historical period can be the most recent month. That is to say, step S201 can be to obtain the peak CPU usage rate and the peak data volume of each server in the cloud computing system within the past month.

[0066] Step S201 can be executed by the aforementioned resource information collection module. The resource information collection module is a basic module of the cloud computing platform, which provides the peak data of the server CPU and disk utilization rate and analyzes the data.

[0067] The peak CPU usage rate refers to the highest CPU usage rate within a preset historical period. The data volume refers to the data stored in the server, that is, the disk space already used by the server. The peak data volume is the maximum value of the disk space already used by the server within a preset historical period.

[0068] For example, within the last month, the highest CPU usage rate of server A reached 70%, and the maximum data volume stored reached 3TB. Then the peak CPU usage rate is 70%, and the peak data volume is 3TB.

[0069] The resource information collection module may specifically include three components: Request, Collection, and Analysis.

[0070] Specifically, when executing step S201, the operation and maintenance personnel can send data acquisition requests to each server in the cloud computing system through the request component, so as to obtain the peak CPU usage rate and the peak data volume of the server within the last month by using relevant Shell script commands or through a visual monitoring dashboard, that is, to obtain the historical operation data of the server. Subsequently, the collection component can provide relevant interfaces for data operations, so that the operation and maintenance personnel can view the obtained historical operation data, and at the same time, the data obtained by the request module can be stored in the database.

[0071] The analysis component can then perform preliminary screening and cleaning on the historical operation data of each obtained server to prepare for subsequent resource information evaluation.

[0072] For example, when obtaining historical operation data, due to reasons such as missing records, it may be impossible to obtain the corresponding historical operation data for some servers. The analysis module can divide the servers in the cloud computing system into analyzable servers and non - analyzable servers by analyzing the presence or absence of historical operation data. Among them, an analyzable server refers to a server that has successfully obtained the corresponding historical operation data, while a non - analyzable server refers to a server that has failed to obtain the corresponding historical operation data.

[0073] Furthermore, the analysis module can also identify the obtained historical operation data based on a preset abnormal data recognition model, such as a pre - trained artificial neural network model, so as to discover and eliminate abnormal historical operation data.

[0074] That is to say, the request component can send a request to the server, store the peak data of the server's CPU and disk utilization within one month obtained into the database of the collection component, and then use the analysis component to analyze the data in the database, eliminate the data with missing values, and obtain the data of CPU and disk utilization respectively, forming a complete data collection and analysis process.

[0075] S202. Read the usage rate range set according to the server category to which the analyzable server belongs, and the data volume range set according to the total storage space of the analyzable server.

[0076] It should be noted that both step S202 and the subsequent step S203 are executed for each analyzable server. That is to say, for each server in the cloud computing system, as long as the historical operation data of the server is successfully obtained in step S201, the corresponding range is read and the resource information is evaluated.

[0077] Among them, the analyzable server refers to the server in the cloud computing system for which historical operation data is successfully obtained; the server category includes web server, database server, online application server, non-online application server, and data analysis server.

[0078] The servers in the cloud computing system specifically include three uses: network proxy, application, and database. Therefore, the servers in the cloud computing system can be divided into three types: web (WEB) server, application (Application, AP) server, and database (Database, DB) server. Among them, the application server can be further divided into three types: online application server, non-online application server, and data analysis server. Servers with different uses have different usage conditions for CPU and disk storage space. Therefore, based on the above classification, different data volume ranges and usage rate ranges are set for different types of servers.

[0079] Generally, in terms of CPU usage, the CPU usage rate of the application server is the highest, the web server is slightly lower, and the database server is generally only used for data access and storage, with the lowest CPU usage rate. The usage rate ranges of different types of servers can be set accordingly.

[0080] For example, the usage rate range of the online application server can be set to 85% to 95%, the usage rate range of the non-online application server can be set to 80% to 90%, the usage rate range of the data analysis server can be set to 85% to 95%, the usage rate range of the web server can be set to 65% to 80%, and the usage rate range of the database server can be set to 55% to 70%. Of course, the usage rate ranges of each category can also be set according to specific circumstances, not limited to the above examples.

[0081] The data volume range is set according to the size of the server's file system, that is, according to the size of the server's total storage space. Specifically, the total storage space can be divided into four levels, namely less than 2TB, 2TB to 5TB, 5TB to 10TB, and greater than 10TB. A corresponding data volume range is set for each level, for example:

[0082] For the storage capacity less than 2TB, the data volume range is 1TB to 1.5TB.

[0083] For the 2TB to 5TB range, the data volume range is 3.5TB to 4TB;

[0084] For the 5TB to 10TB range, the data volume range is 8TB to 9TB;

[0085] For data volumes greater than 10TB, the data volume range is 15TB to 16TB.

[0086] On this basis, for each analyzable server, it can be determined to which of the above-mentioned levels the total storage space of the server belongs, and then the data volume interval corresponding to the corresponding level is used as the data volume interval for evaluating the resource information of the server.

[0087] Similarly, the above-mentioned division of gears and the division of data volume intervals of each gear may also be adjusted according to actual conditions.

[0088] S203: Using the historical operation data, usage rate interval and data volume interval of the analyzable server, perform resource information evaluation on the analyzable server to obtain an evaluation result of the analyzable server.

[0089] Among them, the assessment results include excess capacity, normal capacity and insufficient capacity.

[0090] The resource information evaluation module of the present application may set an Accept() function, through which the historical operation data of each server collected by the aforementioned resource information collection module is received, and the resource information evaluation is performed using the historical operation data.

[0091] That is, for each analyzable server, after resource information evaluation, it may be determined that the server has excess capacity, normal capacity, or insufficient capacity.

[0092] Generally, if the peak value in the server's historical operation data is within the corresponding interval, the server capacity is considered to be normal. If the historical operation data is lower than the lower limit of the corresponding interval, the server capacity is considered to be in excess. If the historical operation data is higher than the upper limit of the corresponding interval, the server capacity is considered to be insufficient.

[0093] It can be seen that the above usage rate range and data volume range are equivalent to the evaluation criteria when evaluating the resource information of a server. That is to say, this solution actually sets different evaluation criteria for different types of servers.

[0094] Based on the ranges set in step S202 for servers with different purposes and different storage spaces of the servers, in step S203, if an analyzable server belongs to a database server and the total storage space belongs to the range of 2TB to 5TB, then the usage rate range of 55% to 70% and the data volume range of 3.5TB to 4TB are used to evaluate its resource information. If an analyzable server belongs to an online application server and the total storage space belongs to the range of 5TB to 10TB, then the usage rate range of 85% to 95% and the data volume range of 8TB to 9TB are used to evaluate its resource information.

[0095] S204. Statistically calculate the proportion of servers with normal capacity in all analyzable servers in the cloud computing system to obtain the normal rate of the cloud computing system.

[0096] For example, in step S201, the historical operation data of 100 servers in the cloud computing system is successfully obtained. The number of analyzable servers is 100. After the resource information evaluation in step S203, it is found that the evaluation results of 70 servers are normal capacity, and the proportion in all analyzable servers is 70%. Then the normal rate of the cloud computing system is 70%.

[0097] S205. Determine the capacity evaluation result of the cloud computing system according to the normal rate of the cloud computing system.

[0098] As Figure 1 shown, the resource information evaluation module of this application may specifically include a server module and a system module.

[0099] The capacity evaluation result of the cloud computing system may specifically include excellent, good, medium, poor, and general. It can be stipulated that when the normal rate is less than 60%, the capacity evaluation result of the cloud computing system is poor; when the normal rate is between 60% and 70%, the capacity evaluation result of the cloud computing system is medium; when the normal rate is between 70% and 80%, the capacity evaluation result of the cloud computing system is good; when the normal rate is greater than 80%, the capacity evaluation result of the cloud computing system is excellent.

[0100] The determination criteria for the above various capacity evaluation results can be changed according to the actual situation.

[0101] As Figure 1As shown in the figure, the resource information evaluation module of the present application specifically includes two sub-modules: a server and a system. In the above steps, steps S202 and S203 can be executed by the server sub-module, and steps S204 and S205 can be executed by the system sub-module.

[0102] In the process of resource information evaluation of this solution, servers with different uses are distinguished, and different thresholds are specified, so that the computing resources and storage resources of each server can be quickly evaluated. For servers with normal capacity evaluation, the percentage of them in the total number of instances is used as the normal rate, and a threshold is specified for system-level capacity analysis. The relevant process is simple and easy to operate. Moreover, by setting different intervals in combination with the actual usage of CPU and storage space by servers with different uses, the evaluation results of the resource information of each server can be made more accurate, thereby improving the accuracy of the capacity evaluation results of the cloud computing system.

[0103] In the above embodiment, step S203, that is, resource information evaluation of analyzable servers, is specifically implemented as follows:

[0104] Judge whether the peak CPU usage rate of the analyzable server is within the usage rate interval, and judge whether the peak data volume of the analyzable server is within the data volume interval;

[0105] If the peak CPU usage rate of the analyzable server is within the usage rate interval and the peak data volume of the analyzable server is within the data volume interval, determine that the evaluation result of the analyzable server is normal capacity;

[0106] If the peak CPU usage rate of the analyzable server is greater than the upper limit of the usage rate interval and the peak data volume of the analyzable server is greater than the upper limit of the data volume interval, determine that the evaluation result of the analyzable server is insufficient capacity;

[0107] If the peak CPU usage rate of the analyzable server is less than the lower limit of the usage rate interval and the peak data volume of the analyzable server is less than the lower limit of the data volume interval, determine that the evaluation result of the analyzable server is excessive capacity.

[0108] For example, assume that the usage rate interval of a server is 80% to 90%, and the data volume interval is 8TB to 9TB. In the historical operation data of this server, the peak CPU usage rate is 85% and the peak data volume is 8.4TB, both of which are within the corresponding intervals. Therefore, it is determined that the server has normal capacity. If the peak CPU usage rate is 75% and the peak data volume is 6.4TB, both of which are less than the lower limits of the corresponding intervals, then it is determined that the server has excessive capacity. If the peak CPU usage rate is 96% and the peak data volume is 9.6TB, both of which are greater than the upper limits of the corresponding intervals, then it is determined that the server has insufficient capacity.

[0109] As shown Figure 1 in the figure, the system provided by the present application may further include a resource load balancing module. Based on this module, the embodiments of the present application further provide the following method for evaluating the capacity of a cloud computing system. Please refer to Figure 3 , and this method may include the following steps:

[0110] S301. Obtain the historical operation data of each server in the cloud computing system.

[0111] S302. Read the usage rate interval set according to the server category to which the analyzable server belongs, and the data volume interval set according to the total storage space of the analyzable server.

[0112] S303. Use the historical operation data, usage rate interval, and data volume interval of the analyzable server to evaluate the resource information of the analyzable server, and obtain the evaluation result of the analyzable server.

[0113] S304. Count the proportion of servers with normal capacity in all analyzable servers of the cloud computing system to obtain the normal rate of the cloud computing system.

[0114] S305. Determine the capacity evaluation result of the cloud computing system according to the normal rate of the cloud computing system.

[0115] The process described in steps S301 to S305 is the same as that in steps S201 to S205, and will not be elaborated here.

[0116] S306. Calculate the load index of the analyzable server using the historical operation data of the analyzable server.

[0117] S307. Calculate the load weight value and load redundancy value of the analyzable server according to the load index and performance index of the analyzable server.

[0118] Among them, the load weight value and load redundancy value of the analyzable server serve as the basis for allocating requests in the cloud computing system.

[0119] It should be noted that steps S306 and S307 are also executed for each analyzable server, that is, each analyzable server can calculate the corresponding load weight value and load redundancy value.

[0120] Steps S306 and S307 can be executed by the resource load balancing module in the system of the present application.

[0121] Specifically, the resource load balancing module can receive the peak CPU usage rate and peak data volume of each server obtained by the resource information collection module through the Accept() function, and then use the Update() to execute the calculation processes described in steps S306 and S307 to calculate the load weight and load redundancy value of the analyzable server. Finally, the Judgement() function performs load balancing based on the load weight and load redundancy value of the server.

[0122] Suppose a cloud computing system includes servers C1, C2, …, Cn. For the i-th server Ci among them, P(Ci) represents the performance index of the server, and F(Ci) represents the load index of the server. Among them, P(Ci) can be calculated by the following formula:

[0123] P(Ci) = [α1 × P CPU (Ci) + α2 × P Mem (Ci) + (1 - α1) × P CPU (Ci) 2 + (1 - α2) × P CPU (Ci) 2 ÷ 2

[0124] In the above formula, P CPU (Ci) represents the CPU frequency of the server Ci, P Mem (Ci) represents the disk size of the server Ci, that is, the total storage space of the server Ci. α1 and α2 are preset proportion coefficients, and the square of the sum of these two proportion coefficients is equal to 1.

[0125] The load index F(Ci) of the server can be calculated according to the following formula:

[0126] F(Ci) = [β1 × F CPU (Ci) + β2 × F Mem (Ci) + (1 - β1) × F CPU (Ci) 2 + (1 - β2) × F CPU (Ci) 2 ÷ 2

[0127] In the above formula, F CPU (Ci) represents the peak CPU usage rate of the server Ci obtained in the previous step, F Mem (Ci) represents the peak data volume of the server Ci obtained in the previous step. β1 and β2 are preset proportion coefficients, and the square of the sum of these two proportion coefficients is equal to 1.

[0128] The load index and performance index calculated using the higher-order terms of CPU usage rate and disk storage space can more accurately reflect the server load capacity.

[0129] The load weight is calculated according to the server performance and load capacity. Assuming the load weight of server Ci is W(Ci), its calculation formula is:

[0130] W(Ci) = F(Ci) ÷ P(Ci).

[0131] The load weight W(Ci) of server Ci can represent the load capacity of the server. The larger the value of W(Ci), the greater the load borne by the current server Ci. Therefore, the resource load balancing module always allocates new requests to the server with a smaller W(Ci) value for processing.

[0132] Calculation of load redundancy value:

[0133] For server Ci, use R(Ci) to represent the load redundancy value of the server. By introducing the load redundancy value, use R(Ci) to characterize the ability of server Ci to add new load. Its calculation formula is: To judge the ability of a certain node to add new load at a certain moment, the calculation formula is:

[0134] R(Ci) = P(Ci) ÷ F(Ci).

[0135] In the system, generally a minimum load redundancy value R will be set min , when R(Ci) is greater than R min , it indicates that server Ci has the ability to handle newly arrived requests.

[0136] Optionally, in any embodiment of the present application, after obtaining the evaluation results of each analyzable server in the cloud computing system through resource information evaluation, the following steps can be further executed:

[0137] Identify the analyzable servers with an evaluation result of overcapacity and output a recycling prompt message; wherein, the recycling prompt message is used to indicate to perform a resource recycling operation on the analyzable servers with an evaluation result of overcapacity. Specifically, the recycling prompt message can prompt the operation and maintenance personnel to reduce the total storage space of the overcapacity servers and lower their CPU frequencies.

[0138] Identify the analyzable servers with an evaluation result of undercapacity and output an expansion prompt message; wherein, the expansion prompt message is used to indicate to perform a resource expansion operation on the analyzable servers with an evaluation result of overcapacity. Specifically, the expansion prompt message can prompt the operation and maintenance personnel to increase the total storage space of the undercapacity servers and raise their CPU frequencies.

[0139] This patent adopts three modules: resource information collection, resource information evaluation, and resource load balancing. By obtaining data, analyzing data, and specifying thresholds for the system server cluster, it realizes capacity analysis at the server level and system level, and finally expands or reclaims resources according to the evaluation results of the servers.

[0140] During the process of resource information evaluation, different types of servers are distinguished and different thresholds are specified, which can quickly evaluate computing resources and storage resources. For servers with normal capacity evaluation, the percentage of them in the total number of instances is used as the normal rate, and thresholds are specified for system-level capacity analysis. The relevant process is simple and easy to operate.

[0141] Resource load balancing uses the server load redundancy value as the basis for resource expansion or recovery. It has strong timeliness and takes effect automatically, which can reduce the number of resource changes, avoid unnecessary power-on and power-off operations on instances, and improve the overall performance and resource utilization rate of the system.

[0142] When the cloud computing platform is running, unnecessary resource waste will occur. How to efficiently utilize the resources of the cloud computing platform is one of the key research objects in the current development of cloud computing. In the traditional cloud computing multi-tenant application mode, due to dynamic load changes affecting tenant resource allocation, resource competition or occupation may occur due to excessive load, resulting in a decline in the quality of computing services provided externally. General load balancing algorithms cannot well solve the problems of multi-system server clusters. The existing work results have poor adaptability and need to be improved in terms of timeliness.

[0143] However, the innovative method of this patent is a capacity analysis model based on the cloud computing platform. On the basis of collecting the peak data of CPU and disk utilization rates of different system servers within a month, it conducts capacity analysis on computing resources and storage resources through three modules: resource information collection, resource information evaluation, and resource load balancing, and adjusts resource allocation according to the server load redundancy value to realize the management of server computing resources and storage resources. The test results show that by using this capacity analysis model, it can effectively realize resource adjustment for multi-systems and multi-server clusters, can adapt to the size of the task set, dynamically improve the resource utilization rate of the cloud computing platform and reduce energy consumption, and provides an effective solution for the cloud computing platform to meet the growing business needs.

[0144] The proposal and implementation of this patent can accurately and quickly provide a capacity analysis solution for operation and maintenance personnel, enabling them to reasonably allocate server resources in different business scenarios, which can not only improve the efficiency of operation and maintenance, but also effectively guarantee the service quality of the cloud computing platform for tenants, laying a solid foundation for multi-system capacity analysis.

[0145] Combined with the method for capacity evaluation of a cloud computing system provided in the embodiments of the present application, the embodiments of the present application also provide a device for capacity evaluation of a cloud computing system. Please refer to Figure 4 , and the device may specifically include the following units:

[0146] An acquisition unit 401, configured to acquire historical operation data of each server in the cloud computing system.

[0147] Among them, the historical operation data of the server includes the peak CPU usage rate and the peak data volume of this server within a preset historical period.

[0148] A reading unit 402, configured to, for each analyzable server, read the usage rate interval set according to the server category to which the analyzable server belongs, and the data volume interval set according to the total storage space of the analyzable server.

[0149] Among them, the analyzable server refers to a server in the cloud computing system for which historical operation data has been successfully acquired; the server categories include web servers, database servers, online application servers, non-online application servers, and data analysis servers.

[0150] An evaluation unit 403, configured to, for each analyzable server, use the historical operation data, usage rate interval, and data volume interval of the analyzable server to perform resource information evaluation on the analyzable server, and obtain an evaluation result of the analyzable server.

[0151] Among them, the evaluation result includes overcapacity, normal capacity, and undercapacity.

[0152] A statistics unit 404, configured to count the proportion of servers with a normal capacity evaluation result among all analyzable servers in the cloud computing system, and obtain the normal rate of the cloud computing system.

[0153] A determination unit 405, configured to determine the capacity evaluation result of the cloud computing system according to the normal rate of the cloud computing system.

[0154] Optionally, when the evaluation unit 403 performs resource information evaluation on the analyzable server according to the historical operation data, usage rate interval, and data volume interval of the analyzable server to obtain an evaluation result of the analyzable server, it is specifically configured to:

[0155] Judge whether the peak CPU usage rate of the analyzable server is within the usage rate interval, and judge whether the peak data volume of the analyzable server is within the data volume interval;

[0156] If the peak CPU usage rate of the analyzable server is within the usage rate interval, and the peak data volume of the analyzable server is within the data volume interval, determine that the evaluation result of the analyzable server is normal capacity;

[0157] If the peak CPU usage rate of the analyzable server is greater than the upper limit of the usage rate range, and the peak data volume of the analyzable server is greater than the upper limit of the data volume range, it is determined that the evaluation result of the analyzable server is insufficient capacity;

[0158] If the peak CPU usage rate of the analyzable server is less than the lower limit of the usage rate range, and the peak data volume of the analyzable server is less than the lower limit of the data volume range, it is determined that the evaluation result of the analyzable server is excessive capacity.

[0159] Optionally, the apparatus further includes:

[0160] The first calculation unit 406 is configured to calculate, for each analyzable server, a load metric of the analyzable server by using the historical operation data of the analyzable server;

[0161] The second calculation unit 407 is configured to calculate, for each analyzable server, a load weight value and a load redundancy value of the analyzable server according to the load metric and the performance metric of the analyzable server; wherein, the load weight value and the load redundancy value of the analyzable server are used as the basis for allocation requests in the cloud computing system.

[0162] Optionally, the evaluation unit 403 is further configured to:

[0163] Identify the analyzable servers with an evaluation result of excessive capacity, and output a recycling prompt message; wherein, the recycling prompt message is used to indicate to perform a resource recycling operation on the analyzable servers with an evaluation result of excessive capacity;

[0164] Identify the analyzable servers with an evaluation result of insufficient capacity, and output an expansion prompt message; wherein, the expansion prompt message is used to indicate to perform a resource expansion operation on the analyzable servers with an evaluation result of excessive capacity.

[0165] For the apparatus for cloud computing system capacity evaluation provided in the embodiments of the present application, the specific working principle can refer to the relevant steps of the method for cloud computing system capacity evaluation provided in the embodiments of the present application, which will not be elaborated here.

[0166] The present application provides a device for evaluating the capacity of a cloud computing system, which is applicable to a cloud computing system composed of multiple servers. In this device, an acquisition unit 401 acquires the historical operation data of each server in the cloud computing system; wherein, the historical operation data of the server includes the peak CPU usage rate and the peak data volume of this server within a preset historical period; a reading unit 402 reads, for each analyzable server, the usage rate interval set according to the server category to which the analyzable server belongs, and the data volume interval set according to the total storage space of the analyzable server; wherein, the analyzable server refers to a server in the cloud computing system for which historical operation data has been successfully acquired; the server categories include web servers, database servers, online application servers, non-online application servers, and data analysis servers; an evaluation unit 403 evaluates the resource information of each analyzable server by using the historical operation data, the usage rate interval, and the data volume interval of the analyzable server, and obtains the evaluation result of the analyzable server; wherein, the evaluation result includes capacity surplus, normal capacity, and insufficient capacity; a statistics unit 404 counts the proportion of servers with normal capacity among all analyzable servers in the cloud computing system to obtain the normal rate of the cloud computing system; a determination unit 405 determines the capacity evaluation result of the cloud computing system according to the normal rate of the cloud computing system. This solution divides different categories based on the different uses of the servers, and then sets different capacity evaluation criteria for servers of different categories and different storage spaces, thereby improving the accuracy of the capacity evaluation result.

[0167] An embodiment of the present application further provides a computer storage medium for storing a computer program, which, when executed, is specifically used to implement the method for evaluating the capacity of a cloud computing system provided in any embodiment of the present application.

[0168] An embodiment of the present application further provides an electronic device, as Figure 5 shown, which specifically includes a memory 501 and a processor 502.

[0169] Among them, the memory 501 is used to store a computer program;

[0170] The processor 502 is used to execute the above computer program, and is specifically used to implement the method for evaluating the capacity of a cloud computing system provided in any embodiment of the present application.

[0171] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0172] It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to define the order or interdependence relationship of the functions performed by these devices, modules or units.

[0173] Those of ordinary skill in the art can implement or use the present application. Various modifications to these embodiments will be readily apparent to those of ordinary skill in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for evaluating the capacity of a cloud computing system, the cloud computing system being composed of multiple servers, characterized in that, The method includes: Obtaining the historical operation data of each server in the cloud computing system; wherein, the historical operation data of the server includes the peak CPU usage rate and the peak data volume of this server within a preset historical period; For each analyzable server, reading the usage rate range set according to the server category to which the analyzable server belongs, and the data volume range set according to the total storage space of the analyzable server; wherein, the analyzable server refers to the server in the cloud computing system for which historical operation data has been successfully obtained; the server categories include web servers, database servers, online application servers, non-online application servers, and data analysis servers, and the categories of the servers are obtained by dividing based on the server usage; For each of the analyzable servers, using the historical operation data, usage rate range, and data volume range of the analyzable server to evaluate the resource information of the analyzable server, and obtaining the evaluation result of the analyzable server; wherein, the evaluation result includes overcapacity, normal capacity, and undercapacity; Counting the proportion of the servers with normal capacity in all analyzable servers in the cloud computing system to obtain the normal rate of the cloud computing system; Determining the capacity evaluation result of the cloud computing system according to the normal rate of the cloud computing system.

2. The method according to claim 1, wherein The evaluating the resource information of the analyzable server using the historical operation data, usage rate range, and data volume range of the analyzable server to obtain the evaluation result of the analyzable server includes: Judging whether the peak CPU usage rate of the analyzable server is within the usage rate range, and judging whether the peak data volume of the analyzable server is within the data volume range; If the peak CPU usage rate of the analyzable server is within the usage rate range and the peak data volume of the analyzable server is within the data volume range, determining that the evaluation result of the analyzable server is normal capacity; If the peak CPU usage rate of the analyzable server is greater than the upper limit of the usage rate range and the peak data volume of the analyzable server is greater than the upper limit of the data volume range, determining that the evaluation result of the analyzable server is undercapacity; If the peak CPU usage rate of the analyzable server is less than the lower limit of the usage rate range and the peak data volume of the analyzable server is less than the lower limit of the data volume range, determining that the evaluation result of the analyzable server is overcapacity.

3. The method according to claim 1, wherein After obtaining the historical operation data of each server in the cloud computing system, it further includes: For each of the analyzable servers, calculating the load index of the analyzable server using the historical operation data of the analyzable server; For each of the analyzable servers, calculating the load weight and load redundancy value of the analyzable server according to the load index and performance index of the analyzable server; wherein, the load weight and load redundancy value of the analyzable server are used as the basis for allocating requests in the cloud computing system.

4. The method according to claim 1, characterized in that, After using the historical operation data, usage rate range, and data volume range of each of the analyzable servers to evaluate the resource information of the analyzable servers and obtaining the evaluation results of the analyzable servers, the following steps are further included: Identify the analyzable servers with an evaluation result of overcapacity and output a recycling prompt message; wherein, the recycling prompt message is used to indicate to perform a resource recycling operation on the analyzable servers with an evaluation result of overcapacity; Identify the analyzable servers with an evaluation result of undercapacity and output an expansion prompt message; wherein, the expansion prompt message is used to indicate to perform a resource expansion operation on the analyzable servers with an evaluation result of overcapacity.

5. An apparatus for evaluating the capacity of a cloud computing system, the cloud computing system being composed of multiple servers, characterized in that, The device includes: An acquisition unit, configured to acquire the historical operation data of each server in the cloud computing system; wherein, the historical operation data of the server includes the CPU usage peak value and the data volume peak value of the server itself within a preset historical period; A reading unit, configured to, for each analyzable server, read the usage rate range set according to the server category to which the analyzable server belongs, and the data volume range set according to the total storage space of the analyzable server; wherein, the analyzable server refers to the server in the cloud computing system for which historical operation data has been successfully acquired; the server categories include web servers, database servers, online application servers, non-online application servers, and data analysis servers, and the categories of the servers are obtained by dividing based on the server usage; An evaluation unit, configured to, for each of the analyzable servers, use the historical operation data, usage rate range, and data volume range of the analyzable server to evaluate the resource information of the analyzable server and obtain the evaluation result of the analyzable server; wherein, the evaluation result includes overcapacity, normal capacity, and undercapacity; A statistics unit, configured to count the proportion of the servers with a normal capacity evaluation result among all the analyzable servers in the cloud computing system to obtain the normal rate of the cloud computing system; A determination unit, configured to determine the capacity evaluation result of the cloud computing system according to the normal rate of the cloud computing system.

6. The device according to claim 5, characterized in that When the evaluation unit evaluates the resource information of the analyzable server according to the historical operation data, usage rate range, and data volume range of the analyzable server and obtains the evaluation result of the analyzable server, it specifically is used for: Judge whether the CPU usage peak value of the analyzable server is within the usage rate range and whether the data volume peak value of the analyzable server is within the data volume range; If the CPU usage peak value of the analyzable server is within the usage rate range and the data volume peak value of the analyzable server is within the data volume range, determine that the evaluation result of the analyzable server is normal capacity; If the CPU usage peak value of the analyzable server is greater than the upper limit of the usage rate range and the data volume peak value of the analyzable server is greater than the upper limit of the data volume range, determine that the evaluation result of the analyzable server is undercapacity; If the peak CPU usage rate of the analyzable server is less than the lower limit of the usage rate range, and the peak data volume of the analyzable server is less than the lower limit of the data volume range, determine that the evaluation result of the analyzable server is overcapacity.

7. The device according to claim 5, characterized in that The device further includes: A first calculation unit, configured to calculate, for each of the analyzable servers, a load index of the analyzable server by using historical operation data of the analyzable server; A second calculation unit, configured to calculate, for each of the analyzable servers, a load weight value and a load redundancy value of the analyzable server according to the load index and the performance index of the analyzable server; wherein, the load weight value and the load redundancy value of the analyzable server are used as a basis for allocation requests in the cloud computing system.

8. The device according to claim 5, wherein The evaluation unit is further configured to: Identify an analyzable server whose evaluation result is overcapacity, and output a recycling prompt message; wherein, the recycling prompt message is used to indicate to perform a resource recycling operation on the analyzable server whose evaluation result is overcapacity; Identify an analyzable server whose evaluation result is undercapacity, and output an expansion prompt message; wherein, the expansion prompt message is used to indicate to perform a resource expansion operation on the analyzable server whose evaluation result is overcapacity.

9. A computer storage medium, characterized in that, For storing a computer program, when the computer program is executed, it is specifically configured to implement the method for evaluating the capacity of the cloud computing system according to any one of claims 1 to 4.

10. An electronic device, characterized in that, Comprising a memory and a processor; Wherein, the memory is used to store a computer program; The processor is used to execute the computer program, and is specifically configured to implement the method for evaluating the capacity of the cloud computing system according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • IO performance evaluation method and device of cache server

    CN109062768A

  • Cloud computing efficiency evaluation method, device and equipment

    CN111404974A