Cloud Platform Resource Monitoring Method, Electronic Device, Computer Readable Storage Medium

By obtaining the cloud system hierarchical model on the cloud platform and detecting the indicators of the relevant resource layer, and generating resource monitoring information, the problem of how to effectively monitor the use of application resources on the cloud platform is solved, real-time monitoring and management of resource usage is achieved.

CN115865942BActive Publication Date: 2025-06-24CHINA PING AN LIFE INSURANCE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211460469.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2025-06-24
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

On the cloud platform, how to effectively monitor the resource usage of applications is an urgent problem to be solved.

Method used

A cloud platform resource monitoring method is proposed, by obtaining the cloud system hierarchical model, detecting the indicators of the computing resource layer, resource orchestration layer and application resource layer, and generating resource monitoring information. The specific steps include obtaining the cloud system hierarchical model, detecting computing resource indicators, resource orchestration indicators, and generating resource monitoring information based on these indicators.

Benefits of technology

Real-time monitoring of cloud platform resource usage is realized, and the resource usage of applications can be further monitored based on resource monitoring information, improving the efficiency and accuracy of resource management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115865942B_ABST
    Figure CN115865942B_ABST
Patent Text Reader

Abstract

This application relates to the field of cloud computing technology, and in particular, to a cloud platform resource monitoring method, an electronic device, and a computer-readable storage medium. The cloud platform resource monitoring method of the embodiments of this application is applied to a cloud platform system. First, a cloud system hierarchical model corresponding to the cloud platform system needs to be obtained. The cloud system hierarchical model includes a computing power resource layer, a resource orchestration layer, and an application resource layer. Further, computing power index detection is performed on the computing power resource layer to obtain computing power resource indexes, orchestration index detection is performed on the resource orchestration layer to obtain resource orchestration indexes, and application index detection is performed on the application resource layer to obtain application resource indexes. Still further, based on the computing power resource indexes, resource orchestration indexes, and application resource indexes, resource monitoring information is generated, which can respectively detect and monitor the resource trends of the computing power resource layer, the resource orchestration layer, and the application resource layer, so that the resource usage of application programs on the cloud platform can be further monitored according to the resource monitoring information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing technology, and in particular, to a method for monitoring cloud platform resources, an electronic device, and a computer-readable storage medium. Background Art

[0002] Kubernetes (K8S) is an orchestration and management technology for portable containers born for container services. Based on Docker technology, Kubernetes provides a series of functions such as deployment and operation, resource scheduling, service discovery, and dynamic scaling for containerized applications. Recently, more and more applications have been migrated from the host platform to the cloud platform, and resource abstraction and resource management are realized through the cloud platform. However, the cloud platform architecture is different from the host platform architecture. The host platform architecture mainly uses the processors in the host to utilize local resources to support the operation of applications, while the cloud platform architecture is based on the underlying distributed cluster servers to provide computing power, and further through the cloud operating system, operations such as cluster management and scheduling optimization of the resources required by application programs are carried out. It should be clear that after migrating the application program to the cloud platform, how to monitor the resource usage of the application program on the cloud platform has become an urgent problem to be solved in the industry. Summary of the Invention

[0003] This application aims to at least solve one of the technical problems existing in the prior art. For this reason, this application proposes a method for monitoring cloud platform resources, an electronic device, and a computer-readable storage medium, which can monitor the resource usage of application programs on the cloud platform.

[0004] According to an embodiment of the first aspect of this application, a method for monitoring cloud platform resources is applied to a cloud platform system. The method includes:

[0005] Obtain a cloud system hierarchical model corresponding to the cloud platform system. The cloud system hierarchical model includes a computing power resource layer, a resource orchestration layer, and an application resource layer. Among them, the computing power resource layer is used to provide computing power resources for the cloud platform system, the resource orchestration layer is used to orchestrate the computing power resources provided by the computing power resource layer, and the application resource layer is used to provide allocatable resources for application programs on the cloud platform system;

[0006] Perform computing power index detection on the computing power resource layer to obtain computing power resource indexes;

[0007] Perform orchestration index detection on the resource orchestration layer to obtain resource orchestration indexes;

[0008] Perform application index detection on the application resource layer to obtain application resource indexes;

[0009] Generate resource monitoring information based on the computing power resource indexes, the resource orchestration indexes, and the application resource indexes.

[0010] According to some embodiments of the present application, the computing power resource layer includes a plurality of resource load nodes, the computing power resource metrics include computing power limit metrics, and the detecting the computing power metrics of the computing power resource layer to obtain computing power resource metrics includes:

[0011] Performing cluster partitioning processing on the plurality of resource load nodes to obtain a computing power resource pool;

[0012] Performing a first stress test based on the computing power resource pool to obtain the computing power limit metrics.

[0013] According to some embodiments of the present application, the performing a first stress test based on the computing power resource pool to obtain the computing power limit metrics includes:

[0014] Determining normal simulation nodes and abnormal simulation nodes from the resource load nodes of each computing power resource pool based on preset abnormal simulation information;

[0015] Performing failure processing on the abnormal simulation nodes of each computing power resource pool;

[0016] Performing a first stress test based on the normal simulation nodes and the abnormal simulation nodes after failure processing to obtain the computing power limit metrics.

[0017] According to some embodiments of the present application, the performing a first stress test based on the computing power resource pool to obtain the computing power limit metrics includes:

[0018] Obtaining application label information corresponding to a plurality of the application programs one by one, where the application label information is used to reflect the resource consumption characteristics of the application programs;

[0019] Performing a first stress test on the computing power resource pool based on each piece of application label information to obtain the computing power limit metrics.

[0020] According to some embodiments of the present application, the performing a first stress test on the computing power resource pool based on each piece of application label information to obtain the computing power limit metrics includes:

[0021] Determining the consumption peak periods of the application programs based on each piece of application label information;

[0022] Performing a first stress test on the computing power resource pool based on the application programs whose consumption peak periods are within the same preset interval to obtain the computing power limit metrics.

[0023] According to some embodiments of the present application, the resource orchestration layer includes a resource allocation server. Detecting the orchestration metrics for the resource orchestration layer to obtain resource orchestration metrics includes:

[0024] Performing multiple simulated request operations on the resource allocation server to obtain a plurality of response test information corresponding one-to-one to the multiple simulated request operations;

[0025] Based on the plurality of response test information, determining the response upper limit information of the resource allocation server;

[0026] Based on the response upper limit information, obtaining the orchestration limit metrics;

[0027] According to the orchestration limit metrics, obtaining the resource orchestration metrics.

[0028] According to some embodiments of the present application, the resource orchestration layer further includes a data storage unit. The data storage unit is used to provide a storage function for the resource allocation server. According to the orchestration limit metrics, obtaining the resource orchestration metrics includes:

[0029] Performing storage metrics detection on the data storage unit of the resource orchestration layer to obtain the actual measured orchestration metrics;

[0030] Integrating the orchestration limit metrics and the actual measured orchestration metrics based on a preset monitoring weight to obtain the resource orchestration metrics, and the proportion of the actual measured orchestration metrics in the monitoring weight is higher than the proportion of the orchestration limit metrics.

[0031] According to some embodiments of the present application, performing application metrics detection on the application resource layer to obtain application resource metrics includes:

[0032] Obtaining application log information corresponding to the application resource layer;

[0033] According to the application log information, obtaining resource warning information and scheduling obstacle information;

[0034] According to the resource warning information and the scheduling obstacle information, determining the application resource metrics.

[0035] In a second aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the cloud platform resource monitoring method according to any one of the embodiments in the first aspect of the present application.

[0036] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the cloud platform resource monitoring method according to any one of the embodiments of the first aspect of the present application.

[0037] The cloud platform resource monitoring method, electronic device, and computer-readable storage medium according to the embodiments of the present application at least have the following beneficial effects:

[0038] The cloud platform resource monitoring method of the embodiment of the present application is applied to a cloud platform system. First, a cloud system hierarchical model corresponding to the cloud platform system needs to be obtained. The cloud system hierarchical model includes a computing power resource layer, a resource orchestration layer, and an application resource layer. Among them, the computing power resource layer is used to provide computing power resources for the cloud platform system, the resource orchestration layer is used to orchestrate the computing power resources provided by the computing power resource layer, and the application resource layer is used to provide allocatable resources for application programs on the cloud platform system. Further, computing power index detection is performed on the computing power resource layer to obtain computing power resource indexes, orchestration index detection is performed on the resource orchestration layer to obtain resource orchestration indexes, application index detection is performed on the application resource layer to obtain application resource indexes, and further, based on the computing power resource indexes, resource orchestration indexes, and application resource indexes, resource monitoring information is generated. Based on the computing power resource indexes, resource orchestration indexes, and application resource indexes, it is possible to respectively detect and monitor the resource trends of the computing power resource layer, the resource orchestration layer, and the application resource layer, so that the resource usage of application programs on the cloud platform can be further monitored according to the resource monitoring information.

[0039] The additional aspects and advantages of the present application will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:

[0041] Figure 1 is a schematic flowchart of the cloud platform resource monitoring method provided by an embodiment of the present application;

[0042] Figure 2 is for an embodiment of the present application Figure 1 a schematic flowchart of step S102;

[0043] Figure 3 is for an embodiment of the present application Figure 2 a schematic flowchart of step S202;

[0044] Figure 4 is for an embodiment of the present application Figure 2 another schematic flowchart of step S202;

[0045] Figure 5 For the embodiment of the present application Figure 4 is a schematic flowchart of step S402;

[0046] Figure 6 For the embodiment of the present application Figure 1 is a schematic flowchart of step S103;

[0047] Figure 7 For the embodiment of the present application Figure 6 is a schematic flowchart of step S604;

[0048] Figure 8 For the embodiment of the present application Figure 1 is a schematic flowchart of step S104;

[0049] Figure 9 is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0050] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present application and should not be construed as a limitation of the present application.

[0051] In the description of the present application, the meaning of several is one or more, the meaning of multiple is more than two, greater than, less than, exceeding, etc. are understood as not including the present number, above, below, within, etc. are understood as including the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0052] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In the description of this application, it should be noted that unless otherwise clearly defined, terms such as "setting", "installation", "connection", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in this application in combination with the specific content of the technical solution. In addition, the identification of specific steps hereinafter does not represent a limitation on the step sequence and execution logic. The execution sequence and execution logic between each step should be understood and inferred with reference to the content expressed in the embodiment.

[0053] Kubernetes (K8S) is an orchestration and management technology for portable containers designed for container services. Based on Docker technology, Kubernetes provides a series of functions such as deployment and operation, resource scheduling, service discovery, and dynamic scaling for containerized applications. Recently, more and more applications have been migrated from the host platform to the Kubernetes platform, and resource abstraction and resource management are achieved through the Kubernetes platform. However, the architecture of the Kubernetes platform is different from that of the host platform. The host platform architecture mainly uses the processors in the host to utilize local resources to support the operation of applications, while the Kubernetes platform architecture is based on the underlying distributed cluster servers to provide computing power, and further through the Kubernetes operating system, operations such as cluster management and scheduling optimization of the resources required by application programs are carried out. It should be clear that after migrating the application program to the Kubernetes platform, how to monitor the resource usage of the application program on the Kubernetes platform has become an urgent problem to be solved in the industry.

[0054] This application aims to solve at least one of the technical problems existing in the prior art. For this purpose, this application proposes a cloud platform resource monitoring method, an electronic device, and a computer-readable storage medium, which can monitor the resource usage of application programs on the cloud platform.

[0055] The cloud platform resource monitoring method according to the first aspect embodiment of this application is applied to a cloud platform system. In some more specific embodiments, the cloud platform resource monitoring method is mainly applied to monitor the resource usage of the Kubernetes cloud platform system.

[0056] Figure 1It is an alternative flowchart of the cloud platform resource monitoring method provided by the embodiments of the present application. Figure 1 The method in may include but is not limited to steps S101 to S105.

[0057] Step S101: Obtain a cloud system hierarchical model corresponding to the cloud platform system. The cloud system hierarchical model includes a computing power resource layer, a resource orchestration layer, and an application resource layer. Among them, the computing power resource layer is used to provide computing power resources for the cloud platform system, the resource orchestration layer is used to orchestrate the computing power resources provided by the computing power resource layer, and the application resource layer is used to provide allocatable resources for the application programs on the cloud platform system.

[0058] Step S102: Detect computing power indicators for the computing power resource layer to obtain computing power resource indicators.

[0059] Step S103: Detect orchestration indicators for the resource orchestration layer to obtain resource orchestration indicators.

[0060] Step S104: Detect application indicators for the application resource layer to obtain application resource indicators.

[0061] Step S105: Generate resource monitoring information based on the computing power resource indicators, resource orchestration indicators, and application resource indicators.

[0062] In the embodiments of the present application shown in the above steps S101 to S105, it is necessary to first obtain a cloud system hierarchical model corresponding to the cloud platform system. The cloud system hierarchical model includes a computing power resource layer, a resource orchestration layer, and an application resource layer. Among them, the computing power resource layer is used to provide computing power resources for the cloud platform system, the resource orchestration layer is used to orchestrate the computing power resources provided by the computing power resource layer, and the application resource layer is used to provide allocatable resources for the application programs on the cloud platform system. Further, detect computing power indicators for the computing power resource layer to obtain computing power resource indicators, detect orchestration indicators for the resource orchestration layer to obtain resource orchestration indicators, detect application indicators for the application resource layer to obtain application resource indicators. Furthermore, generate resource monitoring information based on the computing power resource indicators, resource orchestration indicators, and application resource indicators. Based on the computing power resource indicators, resource orchestration indicators, and application resource indicators, it is possible to detect and monitor the resource trends of the computing power resource layer, resource orchestration layer, and application resource layer respectively, so as to further monitor the resource usage of the application programs on the cloud platform according to the resource monitoring information.

[0063] In step S101 of some embodiments of the present application, it is necessary to first obtain a cloud system hierarchical model corresponding to the cloud platform system. The cloud system hierarchical model includes a computing power resource layer, a resource orchestration layer, and an application resource layer. Among them, the computing power resource layer is used to provide computing power resources for the cloud platform system, the resource orchestration layer is used to orchestrate the computing power resources provided by the computing power resource layer, and the application resource layer is used to provide allocable resources for the application programs on the cloud platform system. It should be noted that the computing power resources in the cloud platform system (such as the Kubernetes cloud platform system) are mainly provided by multiple computing power servers (such as the servers corresponding to the Master node and the Node node in the Kubernetes cloud platform system). After the computing power resources provided by many computing nodes are abstracted and summarized, the resource management server (such as the AP IServer of the Kubernetes cloud platform system) is further used to add, delete, modify, view various resource objects and call various resource interfaces, so as to realize the reasonable orchestration of computing power resources and further allocate them to each application program carried by the cloud platform system, thereby supporting the operation of various application programs. Therefore, in order to monitor the usage of computing power resources in the cloud platform system, in some exemplary embodiments of the present application, the cloud platform system is divided according to the architecture characteristics of the cloud platform system resource usage, so as to obtain a cloud system hierarchical model corresponding to the cloud platform system. The cloud system hierarchical model includes a computing power resource layer, a resource orchestration layer, and an application resource layer. Among them, the computing power resource layer includes multiple computing power servers, which are used to provide computing power resources for the cloud platform system; the resource orchestration layer includes a resource management server, which is used to orchestrate the computing power resources provided by the computing power resource layer; the application resource layer may include a resource controller and a resource scheduler, which are used to provide allocable resources for the application programs on the cloud platform system. It should be clear that there are various ways to obtain the cloud system hierarchical model. It can be to call and obtain the cloud system hierarchical model from the database, or to directly divide according to the architecture characteristics of the cloud platform system to obtain the cloud system hierarchical model, or other acquisition methods. It should be understood that on the basis of the cloud system hierarchical model, the embodiments of the present application can measure various indicators that can reflect the resource usage situation according to the characteristics of each layer in the model, so as to realize the monitoring of the resource usage situation of the cloud platform system.

[0064] In step S102 of some embodiments of the present application, the computing power resource layer is subjected to computing power index detection to obtain computing power resource indexes. It should be emphasized that the computing power resource layer includes multiple computing power servers for providing computing power resources for the cloud platform system. In some embodiments of the present application, each computing power server realizes the supply of computing power resources for the cloud platform system based on its input interface, output interface, central processing unit, and memory capacity. In order to monitor the resource usage of the computing power resource layer, it is necessary to perform computing power index detection on the computing power resource layer to obtain computing power resource indexes. It should be clear that the types of computing power resource indexes in the computing power resource layer are diverse and may include, but are not limited to: computing power limit indexes reflecting the computing power supply capacity of the computing power resource layer and computing power measured indexes reflecting the current computing power usage of the computing power resource layer. Among them, the computing power limit indexes can be determined during the process of performing a stress test on the cloud platform system, while the computing power measured indexes can be obtained by detecting various indexes such as the I / O interface data transmission speed, central processing unit usage rate, memory capacity size, and data throughput of each computing power server.

[0065] In some more specific embodiments, the computing power servers in the Kubernetes cloud platform system mainly provide computing power resources for the Master nodes and Node nodes. Among them, the Master node refers to the cluster control node. Each Kubenetes cluster needs to have one Master node to be responsible for the management and control of the entire cluster. Basically, all control commands of Kubenetes will be sent to the Master node, and then the Master node will be responsible for the specific execution process. Due to the importance of the Master node, the Master node usually occupies an independent X86 server (or a virtual machine). The processes running on the Master node can include, but are not limited to: First, the Kubenetes API Server, which is a key service process for providing HTTP Rest interfaces and is the only entry for operations such as addition, deletion, modification, and query of all resources in Kubenetes, and is also the entry process for cluster control; Second, the Kubenetes Controller Manager, which is the automated control center for all resource objects in Kubenetes; Third, the Kubenetes Scheduler, which is a process responsible for resource scheduling. In addition to the Master node, the other computing power servers in the Kubenetes cluster are called Node nodes. A Node node can be a physical host or two virtual machines. A Node node is a workload node in the Kubenetes cluster. Each Node node will be assigned some workloads by the Master node (the workload refers to a container, for example: Docker). When a certain Node node fails, the workloads on it will be automatically transferred to other nodes by the Master node. The processes running on the Node node can include, but are not limited to: Kubelet, which is responsible for tasks such as the creation, start, and stop of the containers corresponding to the Pods, and closely cooperates with the Master node to implement the basic functions of cluster management; Kube-proxy, which is an important component for implementing the communication and load balancing mechanism of Kubenetes; Docker Engine (Docker engine), which is responsible for the creation and management of containers on the local machine.

[0066] Referring to Figure 2 , according to the cloud platform resource monitoring method of some embodiments of the present application, the computing power resource layer includes multiple resource load nodes, the computing power resource indicators include computing power limit indicators, and the computing power indicators of the computing power resource layer are detected to obtain the computing power resource indicators. Step S102 may include, but is not limited to, the following steps S201 to step S202.

[0067] Step S201, perform cluster division processing on multiple resource load nodes to obtain a computing power resource pool;

[0068] Step S202: Perform a first stress test based on the computing power resource pool to obtain computing power limit indicators.

[0069] In step S201 of some embodiments, it is necessary to first perform cluster division processing on multiple resource load nodes to obtain a computing power resource pool. In some exemplary embodiments of the present application, the computing power resource layer includes multiple computing power servers, some of which are used as resource control nodes to provide a computing power foundation for components that orchestrate resources in the resource orchestration layer, while some other computing power servers are used as resource load nodes to mainly bear the workload required by application programs. It should be noted that the computing power resource indicators reflecting the resource occupancy of the computing power resource layer may include measured computing power indicators and computing power limit indicators. Among them, the measured computing power indicators are the actually measured computing power indicators, including but not limited to the measured values of indicators such as central processing unit utilization rate, memory capacity, data throughput, and data transmission rate; and the computing power limit indicators refer to the indicators reflecting the upper limit of the load of the computing power resource layer. It should be noted that by performing a first stress test on the computing power resource layer to simulate the load boundary situation of the computing power resource layer, the computing power limit indicators reflecting the upper limit of the load of the computing power resource layer can be determined.

[0070] It should be understood that the computing power resource pool is a virtual resource pool that integrates computing power resources. As a key element in realizing the converged infrastructure structure, the virtual resource pool is a collection of shared servers, storage, and networks, which can be reconfigured more quickly according to the requirements of application programs, thereby being able to support changes in business requirements more easily and quickly. In a cloud computing environment, resources are no longer scattered hardware, but after integrating physical servers, one or more logically virtual resource pools are formed, sharing computing, storage, and network resources. The resource pool can delegate the control right of host (or cluster) resources, and its advantages are very obvious when using the resource pool to divide all resources within the cluster. Multiple resource pools can be created as direct children of the host or cluster and configured. Then, the control right of the resource pool can be delegated to other individuals or organizations. It should be clear that clustering and dividing multiple resource load nodes to obtain the computing power resource pool has the following advantages: First, by adding, removing, or reorganizing resource pools as needed, or changing resource allocation, it is possible to isolate resource pools from each other and share resources within the resource pool, thus forming a flexible hierarchical organization. After each computing power resource pool forms a hierarchical organization, changes in resource allocation within a certain resource pool will not affect other unrelated resource pools; Second, all virtual machine creation and management operations can be carried out within the scope of resources granted to the resource pool with reference to factors such as current shares, reservations, and limit settings, which is convenient for resource management; Third, the separation of resources from hardware. If a cluster with enabled Dynamic Resource Scheduling (DRS) is used, the resources of all hosts will always be allocated to the cluster, which means that the system can manage resources independently of the actual hosts providing the resources. For example, if three 2GB hosts are replaced with two 3GB hosts, there is no need to change the resource allocation in the system. The characteristic of the separation of resources from hardware can aggregate more computing power without paying attention to the usage status of each host. For the above three reasons, before orchestrating computing power resources, it is often necessary to first divide the computing power resource pool and then abstract and summarize the computing power resources of multiple resource load nodes. Therefore, in some embodiments of this application, in order to better simulate the above situation, before performing the first stress test on the computing power resource layer, it is necessary to first perform clustering and dividing on multiple resource load nodes to obtain the computing power resource pool, and then further perform the first stress test based on the computing power resource pool to simulate the load boundary situation of the computing power resource layer.

[0071] In step S202 of some embodiments, a first stress test is performed based on the computing power resource pool to obtain a computing power limit index. It should be clear that the first stress test is used to simulate the load boundary situation of the computing power resource layer. During the first stress test, various indicators of the computing power resource layer are measured, and thus the computing power limit index can be obtained. It should be noted that the first stress test can be implemented by means of a stress test tool or by using a pre-written program, script, etc. The load boundary situations simulated by the first stress test may include, but are not limited to: specific load situations such as several resource load nodes having abnormalities and the computing power resource pool being at the limit load. For each load boundary situation, an index data group composed of indicators such as central processing unit utilization rate, memory capacity, data throughput, and data transmission rate can be measured, and thus the index data group is determined as the computing power limit index. It should be understood that there are various ways to perform the first stress test based on the computing power resource pool to obtain the computing power limit index, which may include, but are not limited to, the specific embodiments listed above.

[0072] Through the above steps S201 to S202, after performing cluster division processing on multiple resource load nodes to obtain a computing power resource pool, and then performing a first stress test based on the computing power resource pool to obtain a computing power limit index, it is possible to simulate the load boundary situation of the computing power resource layer in the actual usage scenario, thereby obtaining a computing power limit index that is more in line with the actual application scenario.

[0073] Refer to Figure 3 , according to the cloud platform resource monitoring method of some embodiments of the present application, step S202 may include, but is not limited to, the following steps S301 to S303.

[0074] Step S301, based on the preset abnormal simulation information, determine normal simulation nodes and abnormal simulation nodes from the resource load nodes of each computing power resource pool;

[0075] Step S302, perform failure processing on the abnormal simulation nodes of each computing power resource pool;

[0076] Step S303, based on the normal simulation nodes and the abnormal simulation nodes after failure processing, perform a first stress test to obtain a computing power limit index.

[0077] In steps S301 to S302 of some embodiments, first, based on preset abnormal simulation information, normal simulation nodes and abnormal simulation nodes are determined from the resource load nodes of each computing power resource pool, and then the abnormal simulation nodes of each computing power resource pool are subjected to failure processing. It should be noted that the failure processing of the abnormal simulation nodes of each computing power resource pool is used to simulate the load boundary situation where several resource load nodes are abnormal. It should be understood that each computing power resource pool includes multiple resource load nodes, and several load resource nodes in all the resource load nodes of the same computing power resource pool may be abnormal and fail. At this time, the computing power resources that the entire resource pool can provide for the cloud platform system are immediately affected.

[0078] In step S303 of some embodiments, in order to evaluate the impact on the computing power resources when several load resource nodes are abnormal and fail, in some relatively preferred embodiments of the present application, based on the preset abnormal simulation information, normal simulation nodes and abnormal simulation nodes are determined from the resource load nodes of each computing power resource pool. After the abnormal simulation nodes of each computing power resource pool are subjected to failure processing, a first stress test is further performed based on the normal simulation nodes and the abnormal simulation nodes after the failure processing to obtain a computing power limit index. It should be clear that the preset abnormal simulation information refers to the reference information preset for simulating the abnormal conditions of the resource load nodes. For example, based on the preset abnormal simulation information, a group of resource load nodes are randomly selected from each computing power resource pool for failure processing, and then a first stress test is performed based on the normal simulation nodes and the abnormal simulation nodes after the failure processing to simulate the possible fault situations that each computing power resource pool may encounter. According to some relatively specific embodiments of the present application, during the above first stress test, the upper limit of the fault load that the computing power resource pool can bear is determined, and then the safety threshold corresponding to all the resource load nodes in the resource pool when they are normal is determined based on the upper limit of the fault load and used as the computing power limit index, so as to leave enough margin for the possible faults of the computing power resource pool and avoid potential safety hazards.

[0079] Refer to Figure 4 , according to the cloud platform resource monitoring method of some embodiments of the present application, step S202 may further include, but is not limited to, the following steps S401 to S402.

[0080] Step S401, obtain application label information corresponding to multiple application programs one by one, and the application label information is used to reflect the resource consumption characteristics of the application programs;

[0081] Step S402, perform a first stress test on the computing power resource pool based on each application label information to obtain a computing power limit index.

[0082] In steps S401 to S402 in some embodiments, first obtain application label information corresponding one by one to multiple application programs, where the application label information is used to reflect the resource consumption characteristics of the application programs, and then perform a first stress test on the computing power resource pool based on each piece of application label information to obtain a computing power limit index. It should be emphasized that the computing power resource pool is a virtual resource pool integrating computing power resources. As a key element for implementing the converged infrastructure structure, the virtual resource pool is a collection of shared servers, storage, and networks, and can be reconfigured more quickly according to the requirements of application programs, so as to more easily and quickly support changes in business requirements. In some embodiments of the present application, the resource usage situation of the cloud platform system may conform to certain specific resource consumption characteristics, such as the peaks of resource consumption of several application programs often concentrate in the same time period, and several medium-sized application programs generate relatively large resource consumption due to the linkage relationship between them and various other resource consumption characteristics. The application label information is used to reflect the resource consumption characteristics of the application programs. Therefore, by performing a first stress test on the computing power resource pool based on each piece of application label information to obtain a computing power limit index, the cloud platform system with various resource consumption characteristics can be simulated, and further the load boundary situations that each computing power resource pool may encounter can be created, so as to obtain the computing power limit index more accurately.

[0083] According to some more specific embodiments of the present application, during the monitoring of the Kubernetes cloud platform system, the first stress test of the computing power resource pool can be combined with each application label information and the taint mechanism. For example, if the peak resource consumption of several application programs often concentrates in the same time period, then it can be judged according to this rule: if these application programs correspond to different computing power resource pools respectively, then relatively speaking, it saves resource computing power. Similarly, if these application programs correspond to the same computing power resource pool, then relatively speaking, it consumes more resource computing power; based on the above judgment, the taint mechanism can be used to allocate the above-mentioned several application programs to different application programs, measure the minimum value of computing power consumption, and then allocate the above-mentioned several application programs to the same application program, measure the maximum value of computing power consumption. In this way, the reasonable range of computing power consumption can be delimited according to the minimum and maximum values of computing power consumption. When the measured computing power consumption is not within this reasonable range, it means that the computing power consumption is abnormal. It should be noted that the taint of the Kubernetes cloud platform system is applied to the Node node, indicating that there is a taint on this node, and the Pod that cannot tolerate this taint cannot be scheduled / run on this node. And the tolerance is applied to the Pod. The tolerance allows the resource scheduler to schedule the Pod with the corresponding taint, or allows this Pod to continue to run on this node. It can be clear that the cooperation between the taint and the tolerance can be used to prevent the Pod from being allocated / run to an inappropriate node. In addition, one or more taints can be applied to each node, and one or more tolerances can also be applied to each Pod. It should be pointed out that there are various ways to obtain the computing power limit index by performing the first stress test on the computing power resource pool based on each application label information, which can include, but are not limited to, the specific embodiments cited above. It should be emphasized that by performing the first stress test on the computing power resource pool based on each application label information and obtaining the computing power limit index, the cloud platform system with various resource consumption characteristics can be simulated, and further the load boundary situations that each computing power resource pool may encounter can be created, so as to obtain the computing power limit index more accurately.

[0084] Referring to Figure 5 , according to the cloud platform resource monitoring method of some embodiments of the present application, step S402 may include, but are not limited to, the following steps S501 to step S502.

[0085] Step S501, based on each application label information, determine the peak consumption period of each application program;

[0086] Step S502, based on each application program whose peak consumption period is within the same preset interval, perform the first stress test on the computing power resource pool to obtain the computing power limit index.

[0087] In steps S501 to S502 of some embodiments, first, based on each application label information, determine the peak consumption periods of each application program, and then, based on the application programs whose peak consumption periods are within the same preset interval, perform a first stress test on the computing power resource pool to obtain a computing power limit index. It should be emphasized that the application label information is used to reflect the resource consumption characteristics of the application program. Therefore, based on each application label information, the peak consumption periods of each application program can be determined, where the peak consumption period refers to the period when the computing power consumption peak of the application program occurs. If the peaks of resource consumption of several application programs tend to be concentrated in the same time period, then based on the application programs whose peak consumption periods are within the same preset interval, perform a first stress test on the computing power resource pool to obtain a computing power limit index. Specifically, if these application programs correspond to different computing power resource pools respectively, relatively speaking, it saves computing power resources. Similarly, if these application programs correspond to the same computing power resource pool, relatively speaking, it consumes more computing power resources. Therefore, allocate all the application programs whose peak consumption periods are within the same preset interval to different computing power resource pools for the first stress test, measure the minimum value of the computing power resource consumption, and then allocate all the application programs whose peak consumption periods are within the same preset interval to the same computing power resource pool for the first stress test, measure the maximum value of the computing power resource consumption. Then, using the minimum and maximum values of the computing power resource consumption, a reasonable interval of the computing power resource consumption can be delimited, and this reasonable interval is determined as the computing power limit index. Comparing it with the measured computing power index can be used to judge whether there is an abnormality in the computing power consumption. It should be emphasized that there are various ways to perform a first stress test on the computing power resource pool based on each application label information to obtain a computing power limit index, which can include, but is not limited to, the specific embodiments listed above. It should be understood that in actual applications, since there are various situations for the allocation of computing power resource pools for application programs, performing a first stress test on the computing power resource pool based on the application programs whose peak consumption periods are within the same preset interval can more reasonably determine the computing power limit index.

[0088] In step S103 of some embodiments of the present application, perform an orchestration index detection on the resource orchestration layer to obtain a resource orchestration index. It should be emphasized that the resource orchestration layer includes a resource management server for orchestrating the computing power resources provided by the computing power resource layer. It should be noted that the resource orchestration index includes an orchestration limit index reflecting the resource scheduling ability limit and an orchestration measured index reflecting the resource scheduling situation. Among them, the orchestration limit index can be determined by performing a stress test on the resource management server, and the orchestration measured index can be obtained through the information interaction between the resource management server and each module.

[0089] In some more specific embodiments, the resource management server in the Kubernetes cloud platform system refers to the Kubernetes API Server, which is used to provide HTTP Rest interfaces such as creation, deletion, modification, query, and WATCH for various Kubernetes resource objects (such as Pod, RC, Service, etc.), and is the data bus and data center of the entire system. As the core of the cluster, the Kubernetes API Server is responsible for the communication between various functional modules of the cluster. Each functional module in the cluster stores information in etcd through the API Server. When it is necessary to obtain and operate on this data, it is achieved through the REST interfaces (GET / LIST / WATCH methods) provided by the API Server, thereby realizing the information interaction between modules. Specifically, the interactions between the API Server and each module include the following categories: First, the interaction between the Kubelet and the API Server. The Kubelet on each Node node will regularly call the REST interface of the API Server to report its own status. After the API Server receives this information, it updates the node status information to etcd. The Kubelet also listens to Pod information through the WATCH interface of the API Server to manage the PODs on the Node node. Second, the interaction between the Kube-controller-manager and the API Server. The Node Controller module in the Kube-controller-manager monitors the information of the Node node in real time through the WATCH interface provided by the API Server and performs corresponding processing. Through the interface provided by the API Server, the current status of each resource object in the entire cluster can be monitored in real time. When various failures cause changes in the system status, these Node Controllers will try to correct the system from the "existing state" to the "desired state". Third, the interaction between the Kube-Scheduler and the API Server. After the Scheduler listens to the information of the newly created Pod replicas through the WATCH interface of the API Server, it will retrieve the list of all Node nodes that meet the requirements of the Pod and start executing the Pod scheduling logic. After successful scheduling, the Pod will be bound to the target node immediately.

[0090] Referring to Figure 6 , according to some embodiments of the present application, the resource orchestration layer includes a resource allocation server, and step S103 may include, but is not limited to, the following steps S601 to step S604.

[0091] Step S601, perform multiple simulation request operations on the resource allocation server to obtain multiple response test messages corresponding one-to-one to the multiple simulation request operations;

[0092] Step S602, determine the response upper limit information of the resource allocation server based on the multiple response test messages;

[0093] Step S603, obtain the scheduling limit index based on the response upper limit information;

[0094] Step S604, obtain the resource scheduling index according to the scheduling limit index.

[0095] In steps S601 to S604 in some embodiments, first perform multiple simulation request operations on the resource allocation server to obtain multiple response test messages corresponding one-to-one to the multiple simulation request operations, and determine the response upper limit information of the resource allocation server based on the multiple response test messages. Further, obtain the scheduling limit index based on the response upper limit information. Still further, obtain the scheduling limit index based on the response upper limit information. It should be emphasized that the resource scheduling layer is used to schedule the computing power resources provided by the computing power resource layer. It should be understood that if the resource scheduling ability of the cloud platform system is insufficient, even if the computing power resource layer can provide very sufficient resources, it is difficult to successfully achieve reasonable resource invocation in the cloud platform system and reach the ideal goal. Therefore, it is very important to monitor the resource scheduling ability of the cloud platform system. Thus, in some exemplary embodiments of the present application, multiple simulation request operations are performed on the resource allocation server to obtain multiple response test messages corresponding one-to-one to the multiple simulation request operations. The reason is that through multiple simulation request operations and multiple response test messages corresponding one-to-one to the multiple simulation request operations, the response upper limit information of the resource allocation server can be determined based on the multiple response test messages, and the response upper limit information reflects the response upper limit of the resource allocation server, so as to obtain the scheduling limit index reflecting the resource scheduling ability of the cloud platform system based on the response upper limit information in the subsequent steps. It should be clear that performing multiple simulation request operations on the resource allocation server can be achieved by means of a stress test tool or by using a pre-written program, script, etc., and the response test message can be the response time corresponding to the simulation request operation, or the number of valid responses per unit time, or other response parameters corresponding to the simulation request operation.

[0096] According to some more specific embodiments of the present application, in the Kubernetes cloud platform system, through a stress testing tool or a pre-written program / script, four types of simulated request operations of adding, deleting, modifying, and viewing can be repeatedly performed on various Kubernetes API resources (such as Deployment, Service, etc.) to obtain multiple response test information corresponding one by one to the multiple simulated request operations. Thus, based on the multiple response test information, the response upper limit information of the resource allocation server can be determined. Taking the shell script as an example, instructions such as get, delete, and edit of the Kubectl client are used to loop and batch operate to create, delete, and modify various resources such as PODs to simulate daily operation behaviors. The core is to find the response upper limit information that the resource allocation server can bear requests while simulating API operations at a high frequency, so as to further obtain the scheduling limit indicators based on the response upper limit information. It should be noted that the scheduling limit indicators of the resource allocation server may include, but are not limited to, the response time of the Api-server, timeout situations, etc. In addition, in some embodiments of the present application, through a stress testing tool for storage service capabilities, jmeter or writing a script can be used to simulate high-frequency reading and writing of Etcd, and at the same time, a monitoring tool is used to measure the reading and writing time consumption of Etcd, disk IO indicators, etc., which can reflect service capabilities, and determine them as part of the scheduling limit indicators.

[0097] Through the above steps S601 to S604, first perform multiple simulated request operations on the resource allocation server to obtain multiple response test information corresponding one by one to the multiple simulated request operations, then based on the multiple response test information, determine the response upper limit information of the resource allocation server. Further, based on the response upper limit information, obtain the scheduling limit indicators. Still further, according to the scheduling limit indicators, obtain the resource scheduling indicators. It is possible to determine the scheduling limit indicators based on the response upper limit information of the resource allocation server, thereby more reasonably measuring the resource scheduling indicators to monitor the resource scheduling ability of the cloud platform system.

[0098] Refer to Figure 7 , according to the cloud platform resource monitoring method of some embodiments of the present application, the resource scheduling layer further includes a data storage unit. The data storage unit provides a storage function for the resource allocation server. Step S604 may include, but is not limited to, the following steps S701 to S702.

[0099] Step S701, perform storage index detection on the data storage unit of the resource scheduling layer to obtain the measured scheduling index;

[0100] Step S702: Integrate the orchestration limit metrics and the orchestration measured metrics based on a preset monitoring weight to obtain resource orchestration metrics, where the proportion of the orchestration measured metrics in the monitoring weight is higher than the proportion of the orchestration limit metrics.

[0101] In steps S701 to S702 of some embodiments, it is necessary to first detect the storage metrics of the data storage unit in the resource orchestration layer to obtain the orchestration measured metrics, and then integrate the orchestration limit metrics and the orchestration measured metrics based on a preset monitoring weight to obtain resource orchestration metrics, where the proportion of the orchestration measured metrics in the monitoring weight is higher than the proportion of the orchestration limit metrics. It should be emphasized that the resource orchestration metrics include the orchestration limit metrics reflecting the resource scheduling ability limit and the orchestration measured metrics reflecting the resource scheduling situation. It should be clear that the data storage unit in the resource orchestration layer stores the control data, application data, and the cluster status of the cloud platform system for the cloud platform system to call. Since the data storage unit in the resource orchestration layer stores the control data, application data, and the cluster status of the cloud platform system, and the data storage unit provides a storage function for the resource allocation server, the storage service ability of the data storage unit has an important impact on the resource orchestration ability of the resource allocation server. In some embodiments, if the resource orchestration ability of the resource allocation server is restricted, it is necessary to first check whether there is an abnormality in the data storage unit. Therefore, in some more preferred embodiments of the present application, in order to facilitate troubleshooting possible failures of the data storage unit, it is necessary to detect the storage metrics of the data storage unit in the resource orchestration layer to obtain the orchestration measured metrics.

[0102] According to some more specific embodiments of the present application, in the Kubernetes cloud platform system, the data storage unit of the resource orchestration layer can be the core component Etcd of the Kubernetes cloud platform system. It should be noted that Etcd is a highly available key-value storage system in the Kubernetes cloud platform system, mainly used for shared configuration and service discovery in the Kubernetes cluster. It processes log replication through the Raft consensus algorithm to ensure strong consistency and can be regarded as a highly available and strongly consistent service discovery storage repository. Specifically, Etcd needs to centrally manage some configuration information. When the application starts, it actively obtains the configuration information from Etcd once. At the same time, it registers a WATCHer on the Etcd node and waits. Whenever the configuration is updated later, Etcd will notify the subscriber in real time to achieve the purpose of obtaining the latest configuration information. Service discovery is also one of the problems that need to be solved in a distributed system, that is, how can processes or services in the same distributed cluster find each other and establish connections. Essentially, service discovery is to want to know whether there are processes listening on udp or tcp ports in the cluster, and can be found and connected by name. Etcd mainly solves the problem of data consistency in a distributed system. The data in a distributed system is divided into control data and application data. The data type processed by Etcd is control data, and a small amount of application data can also be processed. It should be noted that the Api-server, as a resource allocation server, can be regarded as the front end of Etcd. At the same time, the state of the entire Kubernetes cluster is stored in Etcd. Therefore, the storage service ability of Etcd actually has a crucial impact on the resource orchestration ability of the Api-server. Therefore, in some more preferred embodiments of the present application, it is necessary to measure the indicators such as IOPS, IO bits / S, and Rratf latency of Etcd in real time and determine them as the measured orchestration indicators to reflect the storage service ability of Etcd. After obtaining the indicators such as IOPS, IO bits / S, and Rratf latency of Etcd as the measured orchestration indicators, then further integrate the orchestration limit indicators and the measured orchestration indicators based on the preset monitoring weights to obtain the resource orchestration indicators. The proportion of the measured orchestration indicators in the monitoring weights is higher than the proportion of the orchestration limit indicators. The reason is also that the storage service ability of Etcd actually has a crucial impact on the resource orchestration ability of the Api-server. Therefore, in the case where the resource orchestration ability of the Api-server is limited, the abnormal troubleshooting of Etcd can be carried out first according to the measured orchestration indicators in the monitoring weights.It should be noted that the monitoring weight can be reflected in the resource orchestration metrics to facilitate maintenance work. For example, if the resource orchestration metrics need to be reflected on the monitoring dashboard, a larger screen proportion can be configured for the measured orchestration metrics, and a screen proportion smaller than the measured orchestration metrics can be configured for the restricted orchestration metrics. Or, if the resource allocation server encounters an exception, a pop-up window showing the measured orchestration metrics will pop up first. It should be clear that the monitoring of cloud platform resources involves listing many metrics. If they are not reasonably arranged according to the importance of each metric, the monitoring efficiency will be relatively low. Therefore, in some preferred embodiments of this application, the restricted orchestration metrics and the measured orchestration metrics are integrated based on a preset monitoring weight to obtain resource orchestration metrics, which can feed back the measured orchestration metrics as a relatively important metric to the operation and maintenance department according to the monitoring weight to improve the monitoring efficiency and facilitate subsequent maintenance.

[0103] In step S104 of some embodiments of this application, application metric detection is performed on the application resource layer to obtain application resource metrics. It should be emphasized that the application resource layer can include a resource controller and a resource scheduler, which are used to provide allocatable resources for application programs on the cloud platform system. It should be noted that the application resource metrics include resource alarm metrics reflecting resource alarm information and scheduling obstacle metrics reflecting scheduling obstacle information. Among them, both the resource alarm metrics and the scheduling obstacle metrics can be obtained from the application log information of the cloud platform system.

[0104] According to some more specific embodiments of the present application, the application resource layer of the Kubernetes cloud platform system may include a resource controller Kube Controller and a resource scheduler Kube Scheduler. The functions of the resource controller Kube Controller include: ensuring the number of replicas of expected Pods, ensuring that all Node nodes run the same Pod, planning one-time tasks and scheduled tasks, stateless application deployment, and stateful application deployment. The resource scheduler Kube Scheduler is used to select Node nodes for application deployment according to a preset algorithm. It should be noted that the Deployment controller is a specific type of resource controller in the Kubernetes cloud platform system. Since the Deployment controller does not directly manage pods, but indirectly manages pods by managing replicasets, that is: deployment manages replicasets, and replicasets manage pods. Therefore, the Deployment controller has more powerful functions than replicasets. Thus, in some more preferred embodiments, the Deployment controller is selected as the resource controller to instruct the Kubernetes cloud platform system to create and update instances of the application program, and use the Master node to schedule the application program instances to specific nodes in the nodes. After creating the application program instances, the Deployment controller will continuously monitor these instances. If the Node node running the application program instance shuts down or is deleted, the Deployment controller will recreate a new instance on another Node node with the optimal resources in the cluster, thereby providing a self-healing mechanism to prevent failures or maintenance problems.

[0105] Referring to Figure 8 , according to the cloud platform resource monitoring method of some embodiments of the present application, step S104 may include but is not limited to the following steps S801 to S803.

[0106] Step S801, obtain application log information corresponding to the application resource layer;

[0107] Step S802, obtain resource alarm information and scheduling obstacle information according to the application log information;

[0108] Step S803, determine application resource metrics according to the resource alarm information and the scheduling obstacle information.

[0109] In steps S801 to S803 of some embodiments, for the detection of application metrics in the application resource layer, it is necessary to first obtain the application log information corresponding to the application resource layer, and then, based on the application log information, obtain the resource alarm information and scheduling obstacle information. Furthermore, based on the resource alarm information and scheduling obstacle information, determine the application resource metrics. It should be noted that for each application program, the resource pool seen by the application program is not the entire abstracted cluster resources. In fact, it is multiple already divided small resource pools. Therefore, it is difficult to measure the application resource metrics from the resource pool. Since the application log information in the cloud platform system often includes fields reflecting the application resource status, in some embodiments of this application, the resource alarm information and scheduling obstacle information can be obtained based on the application log information, and then the application resource metrics can be further determined based on the resource alarm information and scheduling obstacle information. The resource alarm information refers to the log information obtained when the resource load node is abnormal. For example, the computing power resources that the resource load node can provide are insufficient, while the scheduling obstacle information refers to the log information obtained when the resource orchestration process is abnormal. For example, the node stagnation time is too long.

[0110] According to some more specific embodiments of this application, the application log information in the Kubernetes cloud platform system includes the evicted keyword (POD_EVICTED) and the pending keyword (POD_PENDING). It should be understood that eviction means to drive away, and the evicted keyword indicates that the resource load node has been driven away. When the resource load node is abnormal, Kubernetes removes the Pod on that node through the corresponding eviction mechanism, which is mostly caused by insufficient resources. And pending means pending. When a Pod is always in the Pending state, it means that the Pod has not been scheduled to a certain node yet, and it is necessary to check the Pod to analyze the problem reason. The reasons why a Pod is always in the Pending state can include, but are not limited to: insufficient node resources, not meeting nodeSelector and affinity, the Node has taints that the Pod cannot tolerate, bugs in the low-version kube-Scheduler, the kube-Scheduler is not running properly, the stateful applications on other available nodes after eviction and the current node are not in the same available zone, etc. Therefore, by traversing the application log information in the Kubernetes cloud platform system, using the evicted keyword as the resource alarm information and the pending keyword as the scheduling obstacle information, the application resource metrics can be further determined based on the resource alarm information and scheduling obstacle information.

[0111] In the embodiments shown in the above steps S801 to S803, the application resource metrics obtained from the resource alarm information and the scheduling obstacle information can more clearly and distinctly reflect the resource allocation situation. Therefore, by monitoring the application resource metrics, it can be determined whether there are abnormalities in the application resource layer of the cloud platform system.

[0112] In step S105 of some embodiments of the present application, resource monitoring information is generated based on the computing power resource metrics, the resource orchestration metrics, and the application resource metrics. It should be noted that since the computing power resource layer is used to provide computing power resources for the cloud platform system, by detecting the computing power metrics of the computing power resource layer, the computing power resource metrics reflecting the resource occupancy of the computing power resource layer can be obtained; since the resource orchestration layer is used to orchestrate the computing power resources provided by the computing power resource layer, by detecting the orchestration metrics of the resource orchestration layer, the resource orchestration metrics reflecting the resource orchestration situation can be obtained; since the application resource layer is used to provide allocatable resources for the application programs on the cloud platform system, by detecting the application metrics of the application resource layer, the application resource metrics reflecting the resource allocation situation can be obtained. Therefore, based on the computing power resource metrics, the resource orchestration metrics, and the application resource metrics, the resource monitoring information for reflecting the resource usage situation of the cloud platform system can be generated.

[0113] Figure 9 An electronic device 900 provided by an embodiment of the present application is shown. The electronic device 900 includes: a processor 901, a memory 902, and a computer program stored on the memory 902 and executable on the processor 901. When the computer program runs, it is used to execute the above-mentioned cloud platform resource monitoring method.

[0114] The processor 901 and the memory 902 can be connected through a bus or other means.

[0115] As a non-transitory computer-readable storage medium, the memory 902 can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the cloud platform resource monitoring method described in the embodiments of the present application. The processor 901 realizes the above-mentioned cloud platform resource monitoring method by running the non-transitory software programs and instructions stored in the memory 902.

[0116] The memory 902 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store the cloud platform resource monitoring method described above. In addition, the memory 902 may include a high-speed random access memory 902, and may also include a non-transitory memory 902, such as at least one storage device storage component, a flash memory component, or other non-transitory solid-state storage components. In some embodiments, the memory 902 may optionally include a memory 902 that is remotely disposed relative to the processor 901, and these remote memories 902 may be connected to the electronic device 900 through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0117] The non-transitory software programs and instructions required to implement the above cloud platform resource monitoring method are stored in the memory 902. When executed by one or more processors 901, the above cloud platform resource monitoring method is executed. For example, execute Figure 1 the method steps S101 to S105 in Figure 2 the method steps S201 to S202 in Figure 3 the method steps S301 to S303 in Figure 4 the method steps S401 to S402 in Figure 5 the method steps S501 to S502 in Figure 6 the method steps S601 to S604 in Figure 7 the method steps S701 to S702 in Figure 8 the method steps S801 to S803 in

[0118] The embodiments of the present application also provide a computer-readable storage medium storing computer-executable instructions for executing the above cloud platform resource monitoring method.

[0119] In one embodiment, the computer-readable storage medium stores computer-executable instructions that are executed by one or more control processors. For example, execute Figure 1 the method steps S101 to S105 in Figure 2 the method steps S201 to S202 in Figure 3 the method steps S301 to S303 in Figure 4 the method steps S401 to S402 in Figure 5 the method steps S501 to S502 in Figure 6 the method steps S601 to S604 in Figure 7 the method steps S701 to S702 in Figure 8The method steps S801 to S803 in

[0120] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0121] Those of ordinary skill in the art can understand that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassette, tape, storage device storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium. It should also be understood that the various embodiments provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects. The above is a specific description of the preferred embodiments of the present application, but the present application is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.

Claims

1. A cloud platform resource monitoring method, characterized in that, Applied to a cloud platform system, the method includes: Obtain a cloud system hierarchical model corresponding to the cloud platform system. The cloud system hierarchical model includes a computing power resource layer, a resource orchestration layer, and an application resource layer. The computing power resource layer includes multiple resource load nodes. Among them, the computing power resource layer is used to provide computing power resources for the cloud platform system, the resource orchestration layer is used to orchestrate the computing power resources provided by the computing power resource layer, and the application resource layer is used to provide allocatable resources for application programs on the cloud platform system; Perform cluster division processing on multiple resource load nodes to obtain a computing power resource pool; Perform a first stress test based on the computing power resource pool to obtain a computing power limit index. Among them, performing a first stress test based on the computing power resource pool specifically includes: Based on preset abnormal simulation information, determine normal simulation nodes and abnormal simulation nodes from the resource load nodes of each computing power resource pool; Perform failure processing on the abnormal simulation nodes of each computing power resource pool; Based on the normal simulation nodes and the abnormal simulation nodes after failure processing, perform a first stress test to obtain a computing power limit index; Perform an orchestration index detection on the resource orchestration layer to obtain a resource orchestration index. Among them, the resource orchestration index includes an orchestration limit index for reflecting the resource scheduling ability limit and an orchestration measured index for reflecting the resource scheduling situation; Perform an application index detection on the application resource layer to obtain an application resource index. Among them, the application resource index includes a resource alarm index for reflecting resource alarm information and a scheduling obstacle index for reflecting scheduling obstacle information; Generate resource monitoring information based on the computing power limit index, the resource orchestration index, and the application resource index.

2. The method according to claim 1, characterized in that, The performing a first stress test based on the computing power resource pool to obtain a computing power limit index further includes: Obtain application label information corresponding to multiple application programs one by one. The application label information is used to reflect the resource consumption characteristics of the application programs; Perform a first stress test on the computing power resource pool based on each application label information to obtain the computing power limit index.

3. The method according to claim 2, characterized in that, The performing a first stress test on the computing power resource pool based on each application label information to obtain the computing power limit index includes: Based on each application label information, determine the consumption peak period of each application program; Perform a first stress test on the computing power resource pool based on the application programs whose consumption peak periods are within the same preset interval to obtain the computing power limit index.

4. The method according to claim 1, characterized in that, The resource orchestration layer includes a resource allocation server. The performing an orchestration index detection on the resource orchestration layer to obtain a resource orchestration index includes: Perform multiple simulated request operations on the resource allocation server to obtain multiple response test information corresponding to the multiple simulated request operations one by one; Based on the multiple response test information, determine the response upper limit information of the resource allocation server; Based on the response upper limit information, obtain an orchestration limit index; Obtain the resource orchestration index according to the orchestration limit index.

5. The method according to claim 4, wherein The resource orchestration layer further includes a data storage unit, and the data storage unit is used to provide a storage function for the resource allocation server. Obtaining the resource orchestration metrics according to the orchestration limit metrics includes: Performing storage metric detection on the data storage unit of the resource orchestration layer to obtain the actual measured orchestration metrics; Integrating the orchestration limit metrics and the actual measured orchestration metrics based on a preset monitoring weight to obtain the resource orchestration metrics, and the proportion of the actual measured orchestration metrics in the monitoring weight is higher than the proportion of the orchestration limit metrics.

6. The method according to claim 1, wherein Performing application metric detection on the application resource layer to obtain application resource metrics, including: Obtaining application log information corresponding to the application resource layer; Obtaining resource alarm information and scheduling obstacle information according to the application log information; Determining the application resource metrics according to the resource alarm information and the scheduling obstacle information.

7. An electronic device, characterized in that, Including: A memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the cloud platform resource monitoring method according to any one of claims 1 to 6 is implemented.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the cloud platform resource monitoring method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • System fault tolerance test method, electronic equipment and storage medium

    CN115221059A