Cloud platform cluster health state detection method and device, electronic equipment and medium
By generating configuration files and sending detection tasks to cloud platform cluster nodes, obtaining detection results and scoring them, the problem of low efficiency in cloud platform cluster health detection is solved, achieving efficient health status detection and reduced operation and maintenance costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUNAN KYLIN XINAN TECH CO LTD
- Filing Date
- 2023-02-27
- Publication Date
- 2026-05-12
AI Technical Summary
Existing cloud platform cluster health monitoring solutions are inefficient and lack unified standards, making it difficult to detect and address the health status of cloud platform clusters in a timely manner.
By generating configuration files, detection tasks are distributed to each node of the cloud platform cluster, the detection tasks are executed, the detection results are obtained, and the final score of the detection items is determined based on the detection results, thus determining the health status of the cloud platform cluster.
It improves the detection efficiency of cloud platform clusters, promptly detects health issues, ensures the normal operation of the cloud platform, and reduces operation and maintenance costs and the frequency of business trips for R&D personnel.
Smart Images

Figure CN116069542B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud platform technology, and in particular to a method, apparatus, electronic device and medium for detecting the health status of a cloud platform cluster. Background Technology
[0002] In recent years, with the continuous development of cloud computing technology, more and more companies are abandoning traditional office models, no longer providing employees with office computers, but instead moving employee office computers and even critical business machines to the cloud. The cloud platform is an infrastructure platform that supports these operations. It deploys office computers and business machines on the cloud platform, meeting office needs while significantly reducing maintenance work and enhancing security. However, at the same time, a large number of virtual machines run on the cloud platform cluster, and the stable operation of the cloud platform cluster must be guaranteed. Therefore, health checks on the cloud platform cluster are particularly important. Regular health checks must be performed on the cloud platform cluster to detect problems early and avoid widespread unavailability of cloud desktops.
[0003] Currently, there is no universal solution for health monitoring of cloud platform clusters. Each cloud platform solution provider has its own solution for monitoring the health status of cloud platform clusters, which generally suffers from low monitoring efficiency. Summary of the Invention
[0004] To address the aforementioned technical problems, embodiments of this application provide a method, apparatus, electronic device, and medium for detecting the health status of a cloud platform cluster.
[0005] In a first aspect, embodiments of this application provide a cloud platform cluster health status detection method, applied to a cloud platform cluster health status detection system, the method comprising:
[0006] Generate configuration files based on the detection items of the cloud platform cluster;
[0007] The detection task is distributed to each node of the cloud platform cluster according to the configuration file. The detection task includes multiple detection items. Each node is controlled to execute the corresponding detection task to obtain the detection result of each detection item.
[0008] The final test score for each test item is determined based on the test results of each test item.
[0009] The total score of the cloud platform cluster is determined based on the final test score of each of the aforementioned test items, and the health status of the cloud platform cluster is determined based on the total score of the cloud platform cluster.
[0010] Secondly, embodiments of this application provide a cloud platform cluster health status detection device, the device comprising:
[0011] The detection task management module is used to generate a configuration file based on the detection items of the cloud platform cluster, and to distribute detection tasks to each node of the cloud platform cluster according to the configuration file. The detection task includes multiple detection items, and controls each node to execute the corresponding detection task to obtain the detection results of each detection item.
[0012] The detection result parsing module is used to determine the final detection score of each detection item based on the detection results of each detection item, determine the total score of the cloud platform cluster based on the final detection scores of each detection item, and determine the health status of the cloud platform cluster based on the total score of the cloud platform cluster.
[0013] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the computer program executes the cloud platform cluster health status detection method provided in the first aspect when the processor is running.
[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which executes the cloud platform cluster health status detection method provided in the first aspect when run on a processor.
[0015] The cloud platform cluster health status detection method, apparatus, electronic device, and medium provided in this application generate a configuration file based on the detection items of the cloud platform cluster; distribute detection tasks to each node of the cloud platform cluster according to the configuration file, wherein each detection task includes multiple detection items, control each node to execute the corresponding detection task, and obtain the detection results of each detection item; determine the final detection score of each detection item based on the detection results of each detection item; determine the total score of the cloud platform cluster based on the final detection scores of each detection item; and determine the health status of the cloud platform cluster based on the total score of the cloud platform cluster. In this way, by detecting the health of key services and hardware resources of the cloud platform cluster, the detection efficiency of the cloud platform cluster is improved, the health status of the cloud platform is detected in a timely manner to ensure the normal operation of the cloud platform. Through the health detection report, R&D personnel can remotely diagnose the health status of the cloud platform cluster and provide handling opinions, reducing the frequency of business trips for R&D personnel and effectively reducing operation and maintenance costs. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be considered as a limitation on the scope of protection of this application. In the various drawings, similar components are numbered similarly.
[0017] Figure 1A schematic diagram of the cloud platform cluster health status detection system provided in an embodiment of this application is shown.
[0018] Figure 2 This paper illustrates a flowchart of a cloud platform cluster health status detection method provided in an embodiment of this application.
[0019] Figure 3 This paper illustrates a schematic diagram showing the relationship between detection items and detection sub-items provided in an embodiment of this application.
[0020] Figure 4 This paper illustrates a schematic diagram showing the relationship between detection tasks and nodes provided in an embodiment of this application.
[0021] Figure 5 A schematic diagram of the cloud platform cluster health status detection device provided in an embodiment of this application is shown.
[0022] Figure 6 This paper presents another structural schematic diagram of the cloud platform cluster health status detection system provided in an embodiment of this application.
[0023] Main icons: 100 - Management node, 200 - Host, 500 - Cloud platform cluster health status detection device, 501 - Detection task management module, 502 - Detection result parsing module, 503 - Detection task execution module. Detailed Implementation
[0024] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0025] The components of the embodiments of this application described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0026] In the following, the terms “comprising,” “having,” and their cognates, which may be used in various embodiments of this application, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as excluding, firstly, the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more features, numbers, steps, operations, elements, components, or combinations thereof.
[0027] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0028] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.
[0029] Example 1
[0030] This application provides a method for detecting the health status of a cloud platform cluster, which is applied to a cloud platform cluster health status detection system.
[0031] Please see Figure 1 The cloud platform cluster health status detection system includes a management node 100 and multiple hosts 200. A host 200 can also be called a node.
[0032] See Figure 2 The cloud platform cluster health status detection method includes steps S201 to S204, which are described below in conjunction with... Figure 1 and Figure 2 Each step is explained.
[0033] Step S201: Generate a configuration file based on the detection items of the cloud platform cluster.
[0034] As an example, multiple detection items can be obtained through the management node 100's page. A corresponding configuration file is generated based on the selected detection items. The configuration file includes a variable configuration file, which contains the variable configurations required to execute the detection tasks. The initialization of the detection result file is also completed. Detection tasks are generated based on the configuration files and sequentially distributed to each host 200 of the cloud platform cluster health status detection system for detection. The variable configuration file controls the execution of the detection tasks; each detection item corresponds to one detection task. The detection tasks are pre-arranged. During initialization, the variable corresponding to the selected detection item is assigned a value of 1. When the detection tasks are distributed to each host 200 for execution, only the detection items with a variable value of 1 are distributed.
[0035] To further clarify, the initialization of the detection result file includes the following process: creating an empty detection result save file for each selected detection item, and a final detection result record file; the final detection result record file generates some initialization data, saved in JSON format, which is divided into two parts. One part contains the `data` keyword, which records the initialization status of the detection result for each selected detection item, such as initializing it to 0; the other part contains the `head` keyword, which records the start and end times of this detection, progress, score, score of each detection item, etc. The initial value of the progress is 0, the initial value of the score of the cloud platform cluster is 100, and the initial value of the score of each detection item is determined by the number of detection items and their weights.
[0036] In one embodiment, obtaining the initial score value for each of the detection items includes:
[0037] The initial score of each detection item is determined based on a preset scoring threshold and the weight of each detection item. The weight of each detection item is determined according to the importance of each detection item in the cloud platform cluster.
[0038] In this embodiment, the initial value of the score for each detection item follows the following rules:
[0039] Formula 1:
[0040] Formula 2: S i =100 / W*W i ;
[0041] Here, the sum of weights for all current detection items is obtained through Formula 1, where w represents the sum of weights, W i Let S represent the weight of the i-th detection item, and n represent the number of selected detection items; calculate the initial score of each detection item using Formula 2. i This represents the initial score of the i-th detection item. The weight is determined based on the impact of the detection item on the cluster health; a larger value indicates a greater impact of the health status of that detection item on the overall health status of the cluster.
[0042] In this embodiment, the detection item includes multiple detection sub-items; please refer to [link to relevant documentation]. Figure 3 Detection item 1 includes detection sub-items 1 to n, detection item 2 includes detection sub-items 1 to n, ..., detection item n includes detection sub-items 1 to n.
[0043] It should be further noted that the detection items include system operation status detection items and hardware resource status detection items. The system operation status detection items include sub-items for detecting the start / stop status of critical services and / or custom detection logic sub-items; the hardware resource status detection items include sub-items for detecting the hardware configuration parameters of each node in the cloud platform cluster.
[0044] For example, system operation status detection consists of several pre-defined detection items. These items can be the start / stop status detection of a critical service or a custom detection logic. Hardware resource status detection items include, but are not limited to, detection of the node's central processing unit (CPU), network card, memory, and disk. Each detection item has a preset detection result; items that do not conform to the preset detection result are considered abnormal. The detection items can also be further subdivided into multiple detection sub-items. For example, the authorization information detection of system operation status can be subdivided into sub-items such as the number of authorizations and the authorization validity period, with each sub-item having a corresponding detection result.
[0045] The example configuration files include both YAML and conf formats. The YAML configuration file includes: the values of variables corresponding to the detection items, used to specify the tasks to be executed; the IP address of the management node, used to specify the management node; and the IP addresses of all nodes in the cloud platform cluster, used to distribute tasks to these nodes. The conf configuration file includes: the weight of each detection item, specific detection sub-items, thresholds, and other information, which allows users to easily calculate the scores of detection items and determine the detection results of detection sub-items.
[0046] To further clarify, in Figure 1 The management node 100 is responsible for displaying, selecting, and distributing detection tasks, as well as displaying detection results, generating detection reports, and sending emails. Detection tasks consist of pre-defined cloud platform cluster detection items, which can be broadly categorized into two types: system operation status items and hardware resource status items.
[0047] Step S202: Distribute detection tasks to each node of the cloud platform cluster according to the configuration file. The detection task includes multiple detection items. Control each node to execute the corresponding detection task and obtain the detection results of each detection item.
[0048] In this embodiment, the management node 100 pre-sorts the detection tasks corresponding to the detection items and sequentially distributes the tasks to each host 200 for execution according to the order of the detection tasks. After each task is completed, the detection results are updated to a file, and the detection results are parsed to update the detection progress, score, and other information on the page. The file storing the detection results is located in the file system of the management node 100.
[0049] It should be noted that after the cloud platform cluster health status detection system is initialized, detection tasks can be sent to each node. To improve task execution efficiency, detection tasks corresponding to the same detection item will be sent to each node simultaneously. The detection tasks corresponding to the selected detection items will be sent out in a preset order and executed one by one. This process is completed by the management node.
[0050] In this embodiment, a detection task can be distributed to multiple nodes, for example, see [link to relevant documentation]. Figure 4 Detection task 1 is for nodes 1 to n, detection task 2 is for nodes 1 to n, ..., detection task n is for nodes 1 to n.
[0051] As an example, host 200 receives a detection task from management node 100, calls the corresponding program to execute the task, and this program is a custom code that implements the detection logic. The host then feeds back the detection results to the management node, where they are recorded in the corresponding detection result file. After obtaining the detection results for each item, host 200 saves the result for each item and sends it to management node 100. For example, the detection process executed by host 200 refers to the execution of the specific detection logic for a particular item; this could be the execution of a system command or a custom logic block. The detection results from management node 100 are saved in file format to shared or distributed storage, accessible to all nodes in the cloud platform cluster. The format can be JSON, XML, or any other easily parsable data format. Each detection item corresponds to a separate detection result file, which stores the detection results for that item from all nodes.
[0052] It should be further explained that the result of a test item itself is determined by the results of its sub-test items. If any sub-test item of a test item has an abnormal result, then the final result of the entire test item will be abnormal. Sub-test item abnormalities are divided into alarm abnormalities and fault abnormalities. An abnormal result for a test item will affect the score of that test item; among them, alarm abnormalities have a smaller impact on the score than fault abnormalities. In addition, if the test result of a certain specific test item is a fault, all scores of its corresponding test item can be forcibly deducted.
[0053] Step S203: Determine the final test score for each test item based on the test results of each test item.
[0054] As an example, the management node 100 receives the detection results of each detection item sent by the host 200, parses the detection result of the current detection item, obtains the detection result of the current detection item, the current detection progress, score and other information, and updates it to the overall detection result record file corresponding to the detection task; displays the detection results of the detection task on the page, and allows downloading the results or automatically sending email notifications to the administrator; the administrator selectively repairs abnormal detection items based on the detection results.
[0055] This allows for health checks of the cloud platform cluster, and also helps administrators and operations personnel quickly locate problems in the cloud platform cluster, enabling remote diagnosis and repair, and saving on operations and maintenance costs.
[0056] In this embodiment, the final detection result of each detection sub-item can be analyzed based on the original detection result data, and then the final result and score of the detection item can be determined based on the final detection result of each detection sub-item.
[0057] For example, consider a management network detection item. This item includes sub-items such as network latency, connectivity, bonding status, and packet loss. The default settings are: network latency exceeding 100ms results in an alarm; network latency exceeding 1000ms results in a fault; connectivity is assumed to be non-connectable, resulting in a fault; bonding status is assumed to be absent, resulting in an alarm; and packet loss is assumed to be 1 to 5 packets lost out of 10 sent, resulting in an alarm, and more than 5 packets lost, resulting in a fault. Based on these settings, if any sub-item of this management network detection item contains an alarm or a fault, the overall detection result for this item is either an alarm or a fault. The score for this management network detection item is determined by considering the number of abnormal detection sub-items and the specific status of each sub-item.
[0058] In this embodiment, each detection item includes multiple detection sub-items, and the acquisition of the detection result of each detection item includes:
[0059] The final test result of each test item is determined based on the final test result of each sub-test item of each test item.
[0060] As an example, for the detection results of each detection sub-item, the final detection result of that sub-item is obtained by comparing the actual detection value with the preset value. If the result does not conform to the preset value, it is considered an abnormal detection sub-item. The detection result of a detection item is determined by the detection results of its corresponding sub-items. If any one sub-item is abnormal, the final detection result of the entire detection item is abnormal.
[0061] In this embodiment, the total score for a healthy cloud platform cluster is 100 points, starting from 100 points initially. Each detection item is assigned an initial score according to preset rules. The final score for each detection item is determined by the detection results of its sub-items. If any sub-item has an abnormal result, the final detection score for that item is obtained by subtracting the score corresponding to the abnormal sub-item from its initial score. The final score of the cloud platform cluster is the sum of the final detection scores for all detection items. The preset rules are determined by the number of detection items and the weight of each detection item.
[0062] In one embodiment, each of the detection items includes multiple detection sub-items, and determining the final detection score of each detection item based on the final detection results of each detection sub-item of each detection item includes:
[0063] The final detection result of a specific detection sub-item of each detection item and / or the number of abnormal detection sub-items are determined based on the final detection result of each detection sub-item of each detection item.
[0064] The final detection score for each detection item is determined based on the final detection result of a specific sub-item of each detection item and / or the number of abnormal detection sub-items.
[0065] As an example, if a specific detection sub-item is included and that specific detection sub-item is an alarm, the final score for the entire detection item is half of the initial value; if the specific detection sub-item is a fault, the score for the entire detection item is 0.
[0066] In one embodiment, the number of anomaly detection sub-items includes the number of alarm sub-items and the number of fault sub-items. Determining the final detection score for each detection item based on the number of anomaly detection sub-items for each detection item includes:
[0067] The final detection score for each detection item is determined based on the number of alarm sub-items and fault sub-items, the initial score of each detection item, the number of detection sub-items for each detection item, and the number of hosts.
[0068] As an example, if no specific test item is included, the score for the test item is calculated using the following Formula 3:
[0069] Formula 3:
[0070] Where S represents the final test score for this test item, S i This represents the initial score assigned to this detection item during initialization. C represents the product of the number of detection sub-items corresponding to this detection item and the number of hosts. The number of hosts can also be understood as the number of other nodes besides the management node. wC represents the number of alarm sub-items in the detection results across all hosts for all detection sub-items. e This indicates the number of faulty sub-items in the detection results across all hosts.
[0071] To further clarify, the specific detection sub-items are some very important preset detection sub-items. Once such detection sub-items trigger alarms or malfunction, they will seriously affect the health status of the cluster. For example, the detection sub-item for whether a critical service of a certain cluster is abnormal can be configured as a specific detection sub-item.
[0072] In one embodiment, the cloud platform cluster health status detection method further includes:
[0073] The testing schedule for each item is determined based on the number of testing items and the estimated testing time.
[0074] The detection progress of the detection task is determined based on the detection progress of each of the aforementioned detection items.
[0075] It should be noted that the total progress of the testing is set to 100%. The progress of each testing item is calculated based on the number of testing items and the estimated testing time. Once the testing of each item is completed at each node, the progress of that item is added to the total progress.
[0076] As an example, the detection progress of the detection task ranges from 0% to 100%. The initial progress is 0%, and the progress increases to 100% after the detection is completed. The detection progress is updated once each detection item is completed. Each detection item is assigned a detection progress value based on the number of detection items and the preset detection time for each detection item. When the corresponding detection item is completed, the total progress is increased by the progress value corresponding to that detection item.
[0077] Step S204: Determine the total score of the cloud platform cluster based on the final detection score of each detection item, and determine the health status of the cloud platform cluster based on the total score of the cloud platform cluster.
[0078] In this embodiment, the management node 100 saves the processed results to a shared or distributed storage that can be accessed by all nodes in the cluster, using a JSON format that is easy for the web to parse. At the same time, it provides corresponding scores, detection progress, and anomaly statistics for each detection item, and analyzes and determines the health status of the cloud platform cluster.
[0079] In this embodiment, a total score of 100 can be set when the cloud platform cluster is healthy. A certain score is assigned to each detection item based on the number of selected detection items and their corresponding weights. The final score is 100 points minus the points deducted from each detection item. The weights of each detection item are used to measure its impact on the final score; the larger the weight, the greater the impact. The final detection score for each detection item is the deducted score, which is calculated based on the number of abnormal items in the corresponding sub-item and the specific sub-item.
[0080] As an example, a maximum score can be set for when the cloud platform cluster is in a healthy state. Different states of the cloud platform cluster can be divided into corresponding score ranges. The health status level of the cloud platform cluster can be determined based on the score range to which the total score of the cloud platform cluster belongs.
[0081] For example, the preset maximum score is 100 points. A final score of 100 indicates that the cloud platform cluster is very healthy; a score between 80 and 100 indicates that the cloud platform cluster is in good health, with only a few abnormal items and no faults in key detection items; a score between 60 and 80 indicates that the cloud platform cluster is in average health, with more abnormal items but no faults in key detection items; a score less than 60 indicates that the cloud platform cluster is in a worrying state and needs to be addressed; the total score is updated after each detection item is completed.
[0082] The cloud platform cluster health status detection method provided in this embodiment generates a configuration file based on the detection items of the cloud platform cluster; it then distributes detection tasks to each node of the cloud platform cluster according to the configuration file, each detection task including multiple detection items, controls each node to execute the corresponding detection task, and obtains the detection results of each detection item; it determines the final detection score of each detection item based on the detection results of each detection item; it determines the total score of the cloud platform cluster based on the final detection scores of each detection item; and it determines the health status of the cloud platform cluster based on the total score. In this way, by conducting health checks on the critical services and hardware resources of the cloud platform cluster, the detection efficiency of the cloud platform cluster is improved, the health status of the cloud platform is detected in a timely manner to ensure the normal operation of the cloud platform. Through the health check report, developers can remotely diagnose the health status of the cloud platform cluster and provide handling opinions, reducing the frequency of business trips for developers and effectively reducing operation and maintenance costs.
[0083] Example 2
[0084] In addition, this application provides a cloud platform cluster health status detection device, which is applied to a cloud platform cluster health status detection system.
[0085] Specifically, such as Figure 5 As shown, the cloud platform cluster health status detection device 500 includes:
[0086] The detection task management module 501 is used to generate a configuration file based on the detection items of the cloud platform cluster, and to distribute detection tasks to each node of the cloud platform cluster according to the configuration file. The detection task includes multiple detection items, and controls each node to execute the corresponding detection task to obtain the detection results of each detection item.
[0087] The detection result parsing module 502 is used to determine the final detection score of each detection item based on the detection results of each detection item, determine the total score of the cloud platform cluster based on the final detection scores of each detection item, and determine the health status of the cloud platform cluster based on the total score of the cloud platform cluster.
[0088] It should be noted that, see [link / reference] Figure 6 The host 200 may include a detection task execution module 503, and the detection task management module 501 may execute the corresponding detection task through the detection task execution module 503.
[0089] In one embodiment, the detection items include system operation status detection items and hardware resource status detection items. The system operation status detection items include start / stop status detection sub-items for critical services and / or custom detection logic sub-items. The hardware resource status detection items include hardware configuration parameter detection sub-items for each node of the cloud platform cluster.
[0090] In one embodiment, each of the detection items includes multiple detection sub-items, and the detection result parsing module 502 is used to determine the final detection result of each of the detection items based on the final detection result of each detection sub-item of each of the detection items.
[0091] In one embodiment, the detection result parsing module 502 is used to determine the final detection result of a specific detection sub-item of each detection item and / or the number of abnormal detection sub-items based on the final detection result of each detection sub-item of each detection item.
[0092] The final detection score for each detection item is determined based on the final detection result of a specific sub-item of each detection item and / or the number of abnormal detection sub-items.
[0093] In one embodiment, the number of anomaly detection sub-items includes the number of alarm sub-items and the number of fault sub-items. The detection result parsing module 502 is used to determine the final detection score of each detection item based on the number of alarm sub-items and the number of fault sub-items, the initial score of each detection item, the number of detection sub-items of each detection item, and the number of hosts.
[0094] In one embodiment, the detection task management module 501 is used to determine the initial score value of each detection item according to a preset score threshold and the weight of each detection item, wherein the weight of each detection item is determined according to the importance of each detection item in the cloud platform cluster.
[0095] In one embodiment, the cloud platform cluster health status detection device 500 further includes:
[0096] The progress processing module is used to determine the detection progress of each detection item based on the number of detection items and the estimated detection time.
[0097] The detection progress of the detection task is determined based on the detection progress of each of the aforementioned detection items.
[0098] The cloud platform cluster health status detection device 500 provided in this embodiment can implement the cloud platform cluster health status detection method provided in Embodiment 1. To avoid repetition, it will not be described again here.
[0099] The cloud platform cluster health status detection device provided in this embodiment generates a configuration file based on the detection items of the cloud platform cluster; it then distributes detection tasks to each node of the cloud platform cluster according to the configuration file. Each detection task includes multiple detection items, and the device controls each node to execute the corresponding detection task, obtaining the detection results for each item. Based on the detection results of each item, it determines the final detection score for each item; based on the final detection scores of each item, it determines the total score of the cloud platform cluster; and based on the total score, it determines the health status of the cloud platform cluster. In this way, by conducting health checks on the critical services and hardware resources of the cloud platform cluster, the device improves the detection efficiency of the cloud platform cluster, promptly detects health issues, and ensures the normal operation of the cloud platform. Through the health detection report, developers can remotely diagnose the health status of the cloud platform cluster and provide handling suggestions, reducing the frequency of business trips for developers and effectively lowering operation and maintenance costs.
[0100] Example 3
[0101] Furthermore, this application provides an electronic device, including a memory and a processor. The memory stores a computer program, which executes the cloud platform cluster health status detection method provided in Embodiment 1 when running on the processor.
[0102] The electronic device provided in this embodiment can implement the cloud platform cluster health status detection method provided in Embodiment 1. To avoid repetition, it will not be described again here.
[0103] Example 4
[0104] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the cloud platform cluster health status detection method provided in Embodiment 1.
[0105] In this embodiment, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0106] The computer-readable storage medium provided in this embodiment can implement the cloud platform cluster health status detection method provided in Embodiment 1. To avoid repetition, it will not be described again here.
[0107] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.
[0108] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0109] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for detecting the health status of a cloud platform cluster, applied to a cloud platform cluster health status detection system, characterized in that, The method includes: The configuration file is generated based on the detection items of the cloud platform cluster; the configuration file includes a variable configuration file, which contains the variable configuration required to execute the detection task; and the initialization of the detection result file is completed. According to the configuration file, the detection task is issued to each node of the cloud platform cluster. The detection task includes multiple detection items. Each node is controlled to execute the corresponding detection task to obtain the detection result of each detection item. The detection items include system operation status detection items and hardware resource status detection items. The system operation status detection items include key service start / stop status detection sub-items and / or custom detection logic sub-items. The final test score for each test item is determined based on the test results of each test item. The total score of the cloud platform cluster is determined based on the final test score of each of the aforementioned test items, and the health status of the cloud platform cluster is determined based on the total score of the cloud platform cluster. Each of the aforementioned test items includes multiple test sub-items, and determining the final test score for each of the aforementioned test items based on the test results includes: The final detection result of a specific detection sub-item and the number of abnormal detection sub-items for each detection item are determined based on the final detection result of each detection sub-item of each detection item. The final detection score for each detection item is determined based on the final detection result of the specific detection sub-item of each detection item and the number of abnormal detection sub-items; The number of anomaly detection sub-items includes the number of alarm sub-items and the number of fault sub-items. Determining the final detection score for each detection item based on the number of anomaly detection sub-items for each detection item includes: The final detection score for each detection item is determined based on the number of alarm sub-items and fault sub-items, the initial score of each detection item, the number of detection sub-items for each detection item, and the number of hosts. The initial score value for each of the aforementioned detection items is obtained, including: The initial score of each detection item is determined according to the preset score threshold and the weight of each detection item. The weight of each detection item is determined according to the importance of each detection item in the cloud platform cluster. The rules for the initial values of the scores for each test item are as follows: Official 1: ; Official 2: ; Formula 1 is used to obtain the sum of the weights of all current detection items, where W represents the sum of weights. Let represent the weight of the i-th detection item, and n represent the number of selected detection items; calculate the initial score for each detection item using Formula 2. This represents the initial value of the score for the i-th detection item; Each of the aforementioned detection items includes multiple detection sub-items, and the acquisition of the detection results of each of the aforementioned detection items includes: The final test result of each test item is determined based on the final test result of each sub-test item of each test item; If a specific detection sub-item is included and that specific detection sub-item is an alarm, the final score of the entire detection item is half of the initial value; if the specific detection sub-item is a fault, the score of the entire detection item is 0. If a specific detection sub-item is not included, the score for the detection item is calculated using the following formula 3: Official 3: ; Where S represents the final test score for this test item. This represents the initial score assigned to this detection item during initialization. This represents the product of the number of detection sub-items corresponding to this detection item and the number of hosts; This indicates the number of alarm sub-items in the detection results across all hosts. This indicates the number of faulty sub-items in the detection results across all hosts.
2. The method according to claim 1, characterized in that, The hardware resource status detection item includes a sub-item for detecting the hardware configuration parameters of each node in the cloud platform cluster.
3. The method according to claim 1, characterized in that, The method further includes: The testing schedule for each item is determined based on the number of testing items and the estimated testing time. The detection progress of the detection task is determined based on the detection progress of each of the aforementioned detection items.
4. A cloud platform cluster health status detection device, characterized in that, The device, used in a cloud platform cluster health status detection system, includes: The detection task management module is used to generate configuration files based on the detection items of the cloud platform cluster. The configuration files include variable configuration files, which contain the variable configurations required to execute the detection tasks. It also initializes the detection result files and distributes detection tasks to each node of the cloud platform cluster according to the configuration files. Each detection task includes multiple detection items, and the module controls each node to execute the corresponding detection task to obtain the detection results for each detection item. The detection items include system operation status detection items and hardware resource status detection items. The system operation status detection items include start / stop status detection sub-items for critical services and / or custom detection logic sub-items. The detection result parsing module is used to determine the final detection score of each detection item based on the detection results of each detection item, determine the total score of the cloud platform cluster based on the final detection scores of each detection item, and determine the health status of the cloud platform cluster based on the total score of the cloud platform cluster. Each of the aforementioned test items includes multiple test sub-items, and determining the final test score for each of the aforementioned test items based on the test results includes: The final detection result of a specific detection sub-item and the number of abnormal detection sub-items for each detection item are determined based on the final detection result of each detection sub-item of each detection item. The final detection score for each detection item is determined based on the final detection result of the specific detection sub-item of each detection item and the number of abnormal detection sub-items; The number of anomaly detection sub-items includes the number of alarm sub-items and the number of fault sub-items. Determining the final detection score for each detection item based on the number of anomaly detection sub-items for each detection item includes: The final detection score for each detection item is determined based on the number of alarm sub-items and fault sub-items, the initial score of each detection item, the number of detection sub-items for each detection item, and the number of hosts. The initial score value for each of the aforementioned detection items is obtained, including: The initial score of each detection item is determined according to the preset score threshold and the weight of each detection item. The weight of each detection item is determined according to the importance of each detection item in the cloud platform cluster. The rules for the initial values of the scores for each test item are as follows: Official 1: ; Official 2: ; Formula 1 is used to obtain the sum of the weights of all current detection items, where W represents the sum of weights. Let represent the weight of the i-th detection item, and n represent the number of selected detection items; calculate the initial score for each detection item using Formula 2. This represents the initial value of the score for the i-th detection item; Each of the aforementioned detection items includes multiple detection sub-items, and the acquisition of the detection results of each of the aforementioned detection items includes: The final test result of each test item is determined based on the final test result of each sub-test item of each test item; If a specific detection sub-item is included and that specific detection sub-item is an alarm, the final score of the entire detection item is half of the initial value; if the specific detection sub-item is a fault, the score of the entire detection item is 0. If a specific detection sub-item is not included, the score for the detection item is calculated using the following formula 3: Official 3: ; Where S represents the final test score for this test item. This represents the initial score assigned to this detection item during initialization. This represents the product of the number of detection sub-items corresponding to this detection item and the number of hosts; This indicates the number of alarm sub-items in the detection results across all hosts. This indicates the number of faulty sub-items in the detection results across all hosts.
5. An electronic device, characterized in that, The system includes a memory and a processor, wherein the memory stores a computer program that executes the cloud platform cluster health status detection method according to any one of claims 1 to 3 when the processor is running.
6. A computer-readable storage medium, characterized in that, It stores a computer program that, when run on a processor, executes the cloud platform cluster health status detection method according to any one of claims 1 to 3.