A method for estimating a health score of a data center and a computing device
By calculating the health score of the data center through an automated inspection system, the problem of inaccurate assessment of data center health in existing technologies is solved, thereby improving operational efficiency and stability.
Patent Information
- Application Number
- CN202410533527.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-29
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-04-29
AI Technical Summary
Existing technologies cannot accurately assess the health of data centers, resulting in low operational efficiency and failing to meet the monitoring needs of hyperscale data centers.
An automated inspection system is adopted to calculate the weight value of each inspected object by receiving inspection results. The health score of the data center is estimated by using the problem level coefficient and decay function to avoid a rapid drop in the health score caused by anomalies in a single inspected object.
It improves the operational efficiency of data centers, ensures their stable operation, and accurately reflects their health.
Smart Images

Figure CN118569831B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data center operation and maintenance, and particularly relates to a data center health score estimation method and a computing device. BACKGROUND
[0002] A data center is a setting containing a large number of computers and related equipment, used for storing, managing and processing a large amount of data to support information technology services and business needs. In the digital era, data centers play a crucial role in supporting the operation and development of various technology applications such as cloud computing, big data sharing, artificial intelligence, etc. In order to ensure the stable operation of the data center, it is necessary to monitor various indicators of the data center. The health of the data center is an important parameter that reflects the overall operation state and performance of the data center. The health score of the data center is an index for quantitative evaluation of the health.
[0003] With the rapid development of science and technology, data centers are developing towards super large scale. As the number, scale and complexity of data centers gradually increase, the manual inspection method of operation and maintenance personnel and the estimation of health score cannot meet the new requirements. Therefore, through an automatic inspection system, the inspection of the data center is realized and the health score is estimated to meet the monitoring needs of the current super large scale data center.
[0004] The health score calculated by the current technology cannot accurately reflect the health of the data center, resulting in low operation efficiency of the data center. SUMMARY
[0005] The data center health score estimation method and the computing device provided by the present application can accurately reflect the health of the data center, thereby improving the operation efficiency of the data center and ensuring the stable operation of the data center.
[0006] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:
[0007] On the first aspect, the present application provides a method for estimating the health score of a data center, including: receiving inspection results for inspection objects and inspection items, wherein the inspection objects are objects to be inspected in the data center; based on the inspection results, calculating the problem weight value corresponding to each inspection object; if the inspection results indicate that the number of inspection items with abnormal status of the inspection object is greater than a preset number, based on the problem boundary coefficient corresponding to the inspection items with abnormal status and a preset attenuation function, calculating the problem weight value corresponding to the inspection object; based on the problem weight value corresponding to each inspection object, estimating the health score of the data center. The present application calculates the problem weight value of each inspection object by combining the problem level coefficient corresponding to the inspection items with abnormal status and the attenuation function, so that the weight of the inspection object in the overall health score continues to decrease, avoiding the problem of the inability to accurately reflect the health of the data center due to a temporary connection failure of a certain inspection object, which leads to a rapid decrease in the health score, thereby improving the operation and maintenance efficiency of the data center and ensuring the stable operation of the data center.
[0008] In a possible implementation, if the inspection item results indicate that the number of inspection items with abnormal status of the inspection object is less than or equal to a preset number, the problem weight value corresponding to the inspection object is calculated based on the problem level coefficient corresponding to the inspection items with abnormal status.
[0009] In a possible implementation, the problem weight value corresponding to the inspection object is the cumulative multiplication result of the problem weight coefficients corresponding to the inspection items with abnormal status of the inspection object; if the inspection item is the first of all the inspection items with abnormal status of the inspection object, y If the inspection item is not the first one among all the inspection items with abnormal status of the inspection object, the problem weight coefficient corresponding to the inspection item is the problem level coefficient corresponding to the inspection item. y Item, then the problem weight coefficient corresponding to the inspection item is the difference between 1 and the attenuation function value corresponding to the inspection item, where the attenuation function value corresponding to the inspection item is the value of the preset attenuation function when the independent variable is the number of items in all inspection items in abnormal status of the inspection item; wherein, y is a preset number, and y Is a positive integer. y When the inspection item is not the previous one, the problem weight coefficient is the problem level coefficient corresponding to the inspection item; when the inspection item is not the previous one, the problem weight coefficient is the problem level coefficient corresponding to the inspection item; y When the item is selected, the problem weight coefficient is the difference between 1 and the corresponding attenuation function value, where the difference between 1 and the corresponding attenuation function value is close to 1, thereby reducing the impact of the inspection item on the overall health score, avoiding the problem of rapid decline in health score due to temporary abnormality of the inspection object, and failing to accurately reflect the health of the data center, thereby improving the operation and maintenance efficiency of the data center and ensuring the stable operation of the data center.
[0010] In a possible implementation, the inspection result of the inspection object and the inspection item is obtained in the following manner: receiving the inspection object and the inspection item input by the user; based on the received inspection object and the inspection item, sending a calling request to an execution service, and the execution service performs the inspection operation on the inspection object and the inspection item in response to the calling request, and returns the inspection result of the inspection object and the inspection item. The inspection is completed in response to the inspection object and the inspection item set by the user, and the inspection result is returned, which can meet the customization of each inspection task.
[0011] In a possible implementation, the execution service includes a business service, and the business service is a service for performing an inspection task; based on the received inspection object and the inspection item, a calling request is sent to a corresponding application program interface, so as to call a corresponding business service through the corresponding business program interface to perform the inspection operation on the inspection object and the inspection item. The inspection of the inspection object and the inspection item is implemented by calling the business service through the application program interface.
[0012] In a possible implementation, the execution service includes an automation job service, and the automation job service is a service for executing an automation script; based on the received inspection object and the inspection item, a calling request carrying an inspection script is sent to the automation job service, so as to execute the inspection script by the automation job service to implement the inspection operation on the inspection object and the inspection item, and the inspection script is an automation script for inspecting the inspection object and the inspection item. The inspection of the inspection object and the inspection item is implemented by executing the automation script through the automation job service.
[0013] In a possible implementation, before receiving the inspection object and the inspection item input by the user, the method further includes: in response to the start of the automation inspection center, obtaining a set of inspection items stored in a database and all inspection items stored in an inspection package; the inspection package is used to store the inspection items and configuration data of the inspection items; based on the set of inspection items and all the inspection items stored in the inspection package, determining the storage condition of the inspection items; based on the storage condition of the inspection items, performing a registration operation corresponding to the storage condition. To ensure that all the inspection items that can be selected by the user are registered to the automation inspection center, thereby realizing the correct receiving of the inspection items and the inspection object selected by the user.
[0014] In a possible implementation, the storage condition of the inspection item includes: a first storage condition, a second storage condition and a third storage condition; the first storage condition is that the inspection item exists in both the inspection package and the inspection item set; the second storage condition is that the inspection item exists in the inspection package but does not exist in the inspection item set; and the third storage condition is that the inspection item exists in the inspection item set but does not exist in the inspection package. When the storage condition of the inspection item is the first storage condition, the configuration data of the inspection item stored in the inspection package is used to update the inspection item stored in the database; when the storage condition of the inspection item is the second storage condition, the configuration data of the inspection item stored in the inspection package is used to add the inspection item to the database; and when the storage condition of the inspection item is the third storage condition, the inspection item stored in the database is deleted. Through different storage conditions, corresponding registration operations are performed to complete correct registration of the inspection item.
[0015] In a possible implementation, the preset attenuation function includes an exponential attenuation function; the exponential attenuation function is expressed as:
[0016] ;
[0017] wherein, A is an initial value of the exponential attenuation function, λ is an attenuation rate, e is a base number of a natural logarithm, x is an independent variable; wherein the initial value of the exponential attenuation function A and the attenuation rate λ are preset. By using the exponential attenuation function as the preset attenuation function, and by setting the initial value of the exponential attenuation function A and the attenuation rate, the exponential attenuation function can gradually approach zero in a suitable range.
[0018] In a second aspect, the present application provides a computing device, including a processor and a memory, the processor being coupled to the memory, and the memory storing computer program instructions, and the processor implementing the data center health score estimation method in the first aspect and possible implementation manners thereof when executing the computer program instructions.
[0019] In a third aspect, the present application provides a computer readable storage medium, the computer readable storage medium storing computer program instructions, and the computer program instructions making the computing device implement the data center health score estimation method in the first aspect and possible implementation manners thereof when running on the computing device.
[0020] In a fourth aspect, the present application provides a computer program product, which comprises computer program instructions, when the computer program instructions are run on a computing device, cause the computing device to implement the data center health score estimation method in the first aspect and possible implementation manners thereof. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 A schematic diagram of a framework structure of an automatic inspection system provided by an embodiment of the present application;
[0022] Figure 2 A schematic diagram of a flow of a data center health score estimation method provided by an embodiment of the present application;
[0023] Figure 3 A schematic diagram of a front-end interface for inputting an inspection item and an inspection object provided by an embodiment of the present application;
[0024] Figure 4 A schematic diagram of a flow of a registration method of an inspection item provided by an embodiment of the present application;
[0025] Figure 5 A schematic diagram of a flow of another data center health score estimation method provided by an embodiment of the present application;
[0026] Figure 6 A schematic diagram of an interface for feeding back a health score provided by an embodiment of the present application;
[0027] Figure 7 A schematic diagram of a structure of a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0028] The terms "first", "second", and "third" and the like in the specification and claims of the present application and the description of the drawings are used to distinguish different objects, and are not used to limit a particular order.
[0029] In the embodiments of the present application, the words "exemplary" or "for example" are used to mean serving as an example, instance, or illustration, and not necessarily to imply any preference or superiority. In fact, an "exemplary" or "for example" embodiment or design scheme should not necessarily be construed as preferred or superior over other embodiments or design schemes.
[0030] For the sake of clear and concise description of each embodiment below, first, a brief introduction of related art is given:
[0031] Health refers to the overall health and operational status of the data center, and is a comprehensive concept that covers various aspects of the data center, including the operational status of equipment, energy utilization efficiency, environmental control, network performance, completeness, etc.
[0032] Health Score is a quantitative evaluation and scoring of the health of the data center. By monitoring, evaluating and analyzing various indicators of the data center, a comprehensive health score can be obtained to reflect the overall health of the data center. Health Score can be used to help operations personnel connect the operational status of the data center, find problems, develop optimization strategies and improvement measures, etc.
[0033] The advantages of the method for estimating the health score of the data center provided by the embodiments of the present application are described below.
[0034] Figure 1 An automatic inspection system is provided for the embodiments of the present application.
[0035] With the new demands of big data, artificial intelligence and other cloud computing applications, such as cloudification and micro-service, the scale of data centers is becoming larger and larger. Data centers mainly include IT hardware devices and software systems. The current common data center generally includes several host rooms, with rack quantities of more than 10,000, IT hardware devices usually more than 100,000, and software systems containing hundreds of systems and tens of thousands of services. For the monitoring of large-scale data centers, the number of monitoring elements that operations personnel need to control increases exponentially, and the configuration of manual monitoring static thresholds is a huge workload, which is prone to phenomena such as alarm storm, false alarm, and failure to alarm. Furthermore, it is difficult for operations personnel to evaluate the overall operation of the data center based on the trend of index fluctuations and alarm conditions, and further comprehensive analysis is needed from multiple dimensions such as infrastructure, middleware, and business applications to assess the health of the data center. Therefore, the manual inspection method of operations personnel and the evaluation of the health of the data center cannot meet the monitoring needs of large-scale data centers.
[0036] In order to meet the monitoring needs of current large-scale data centers, the data center is patrolled and the health score is estimated through an automatic patrol system. The automatic patrol system is a system that patrols and detects data center equipment, environment, safety, etc. using automation technology. Specifically, the automatic patrol system performs rapid patrol on the whole data center by calling an application programming interface (API) or an automation script, and calculates the health score of the data center based on the patrol results. The automatic patrol method replaces the manual patrol method of the operation and maintenance personnel, significantly improves the patrol and diagnosis efficiency of the data center, and can help the operation and maintenance personnel quickly find the faults and hidden dangers of the data center, thereby ensuring the stable operation and reliability of the data center.
[0037] The following will be described in detail Figure 1 The framework structure of the automatic patrol system will be introduced in detail.
[0038] As Figure 1 shown, the automatic patrol system 100 includes an automatic patrol center 110 and an execution layer 120.
[0039] Among them, the automatic patrol center 110 is the core control and supervision unit of the automatic patrol system 100, responsible for overall monitoring, scheduling and management. The operation and maintenance personnel can patrol the strategy in the automatic patrol center, make monitoring rules, view real-time monitoring data, perform fault diagnosis and early warning processing, etc.
[0040] Further, the automatic patrol center 110 provides a visual monitoring interface and reporting function for users, facilitating the operation and maintenance personnel to monitor the patrol task of the data center in real time.
[0041] Specifically, as Figure 1 shown, the framework of the automatic patrol center includes an application layer 111, a storage layer 112 and a logic and calling layer 113.
[0042] Among them, the application layer 111 is used for the operation and maintenance personnel to interact with the automatic patrol system 100. The operation and maintenance personnel can set the patrol item, the patrol object and the patrol task through the application layer 111, and the application layer 111 is also used to generate the patrol report and provide it to the operation and maintenance personnel, facilitating the operation and maintenance personnel to monitor and manage the patrol of the data center. In one possible implementation, the application layer 111 provides a front-end interface of the automatic patrol center 110 to the operation and maintenance personnel, that is, an interface for the operation and maintenance personnel to interact with the automatic patrol system 100. The operation and maintenance personnel can set the patrol item, the patrol object, etc. through the front-end interface, and the automatic patrol center 110 can display the patrol result, the patrol report, etc. to the operation and maintenance personnel through the front-end interface.
[0043] The inspection item refers to a specific index that needs to be inspected. The inspection item can be various performance indexes, running states, configuration information, etc. For example, CPU utilization, memory occupation, network connection state, disk space, etc.
[0044] The inspection object refers to a specific device, system, or other resource that needs to be inspected. The inspection object can be a server, network device, database, application program, or other types of devices or systems. For different inspection objects, the inspection items set may be different.
[0045] The inspection task refers to a series of inspection operations performed on the inspection object. The inspection task mainly includes the inspection item and the inspection object. After the operation and maintenance personnel set the inspection item and the inspection object, the application layer 111 can generate the corresponding inspection task. Further, the inspection task also includes the inspection frequency, the inspection period, the inspection time limit, etc.
[0046] The inspection report is a summary and record of the inspection task result, which usually includes detailed results of the inspection, problems found, solutions, etc. The inspection report provides important reference information for the operation and maintenance personnel, helping them to understand the health status of the data center, find problems, and determine the corresponding measures.
[0047] The storage layer 112 is used to store the data involved in the automatic inspection center. The storage layer 112 usually includes a database (Database, DB) and a file system. The database is used to store structured data, such as device information, inspection results, inspection history data, user information, etc. The database provides efficient data storage, retrieval, and management functions, and can support complex data query and analysis requirements. The file system is used to store unstructured data, such as inspection reports, log files, configuration files, and automation scripts. The file system provides storage, reading, and management functions for files, and can support large-capacity data storage.
[0048] The logic and calling layer 113 is used to implement the execution and monitoring of the inspection task, and is specifically responsible for task scheduling, result analysis, and result storage.
[0049] Task scheduling, the logic and calling layer 113 is responsible for the scheduling and management of tasks. For example, the execution of the inspection task can be scheduled according to a predetermined schedule or specific trigger conditions. For example, the logic and calling layer 113 can call the corresponding business services of the execution layer 120 through the application program interface to implement the inspection task of the data center; or send an automation script to the execution layer 120 to execute the automation script to implement the inspection task of the data center.
[0050] Result parsing: After the execution layer 120 completes the inspection task, it returns the inspection results to the logic and call layer 113 of the automated inspection center 110. The results are then parsed to extract useful information and indicators. Parsing the inspection results may include data cleansing, format conversion, and anomaly detection to ensure that the final results can be correctly stored and analyzed.
[0051] Result storage: The logic and call layer 113 is responsible for storing the parsed inspection results in the storage layer 112.
[0052] The execution layer 120 is the actual execution unit of the automatic inspection system 100 and is responsible for executing specific inspection tasks. The execution layer 120 mainly includes components such as business services and automated operation services.
[0053] Business services are services responsible for executing inspection tasks, including server inspection services and network equipment inspection services. Automated operation services are services responsible for executing automated scripts.
[0054] Through Figure 1 The automated inspection system described above implements automated inspections of the data center and generates automated inspection results. The health score of the data center is estimated based on the inspection results. Two methods for estimating the health score are described below.
[0055] The first method assigns a weight to each inspection item. The health weight of each inspection item is the corresponding weight factor * the number of inspection items with normal status for that inspection item. The total weight of each inspection item is the corresponding weight factor * the total number of inspection items for that inspection item. The health score is estimated as follows: Health Score = Sum of the health weights of all inspection items / Sum of the total weights of all inspection items * 100. In other words, Health Score = (Health Weight of the First Inspection Item + Health Weight of the Second Inspection Item + ... + Health Weight of the Nth Inspection Item) / (Total Weight of the First Inspection Item + Total Weight of the Second Inspection Item + ... + Total Weight of the Nth Inspection Item) * 100, where N is a positive integer.
[0056] For ease of understanding, the formula for estimating the first health score is shown in formula (1).
[0057] (1)
[0058] in, score Score your health. a i For the i The health weight value of each inspection item, sumScoreThe total weight value corresponding to all the inspection items.
[0059] For example, the inspection task includes four inspection items and three inspection objects, that is, the inspection task is to automatically inspect four inspection items of three inspection objects, and the specific information is shown in Table 1:
[0060] Table 1
[0061]
[0062] As shown in Table 1, the three inspection objects are server 1, server 2 and server 3, and the corresponding four inspection items are power status, system health, device location and product name. The weight coefficient corresponding to the first inspection item power status is 40; the weight coefficient corresponding to the second inspection item system health is 50; the weight coefficient corresponding to the third inspection item device location is 30; and the weight coefficient corresponding to the fourth inspection item product name is 30.
[0063] As shown in Table 1, the total weight value corresponding to the four inspection items is sumScore =40*3+50*3+30*3+30*3=450.
[0064] Through the automatic inspection system, the three inspection objects and four inspection items shown in Table 1 are automatically inspected, and the inspection result is that there are 2 normal state inspection objects for the first inspection item (power status), 3 normal state inspection objects for the second inspection item (system health), 2 normal state inspection objects for the third inspection item (device location), and 1 normal state inspection object for the fourth inspection item (product name). Then the health weight values corresponding to the four inspection items are respectively: a 1=40*2=80, a 2=50*3=150, a 3=2*30=60, a 4=30. Then the health score ( score ) = (80+150+60+30) / 450*100=71 points.
[0065] For the first health score estimation method, if an anomaly in a particular inspection item causes a severe problem, but only some of the inspection objects in that inspection item are abnormal, while the remaining inspection objects are normal (i.e., the inspection item has some anomalies), the overall health score remains high, which is inconsistent with expectations. This means that the health score does not accurately reflect the actual operating status and performance of the data center. For example, if an anomaly in the third of the four inspection items shown in Table 1 causes a severe problem, and automated inspections are performed on the four inspection items and their three inspection objects shown in Table 1, the inspection results show that the third inspection item has two normal inspection objects, and one abnormal inspection object. In this case, the anomaly in the third inspection item will result in a severe problem, but the health score will remain at 71 points, and the health score will not decrease. This does not accurately reflect the severe problem in one inspection item for one inspection object in the data center.
[0066] The second method for estimating the health score is to pre-set the problem level for each inspection item. Generally, problems caused by abnormal inspection items are divided into four levels: fatal, severe, general, and prompt. A corresponding problem level coefficient is set for each problem level: 0.79, 0.9, 0.99, and 0.999, respectively. The problem weight value of each inspection item is the mth power of the corresponding problem level coefficient, where m is the number of inspection objects with abnormal status for that inspection item, and m is a non-negative integer. The health score is estimated as follows: Health score = 100 * Problem weight value of the first inspection item * Problem weight value of the second inspection item * … * Problem weight value of the Nth inspection item, where N is a positive integer.
[0067] For ease of understanding, the formula for estimating the second health score is shown in formula (2).
[0068] (2)
[0069] in, score Score your health. b i For the i The problem weight value of each inspection item.
[0070] Exemplarily, the inspection task includes four inspection items and three inspection objects, that is, the inspection task is to perform automated inspections on four inspection items of three inspection objects, as shown in Table 2.
[0071] Table 2
[0072]
[0073] As shown in Table 2, the three inspection objects are server 1, server 2 and server 3 respectively, and the corresponding four inspection items are power state, system health state, device location and product name. The problem level caused by the state abnormality of the first inspection item is fatal, and the corresponding problem level coefficient is 0.79; the problem level caused by the state abnormality of the second inspection item is fatal, and the corresponding problem level coefficient is 0.79; the problem level caused by the state abnormality of the third inspection item is prompt, and the corresponding problem level coefficient is 0.999; the problem level caused by the state abnormality of the fourth inspection item is prompt, and the corresponding problem level coefficient is 0.999.
[0074] Through the automatic inspection of the three inspection objects and the four inspection items shown in Table 2 by the automatic inspection system, the inspection result is that there is 0 inspection object state abnormality in the first inspection item, i.e. m=0, so the problem weight value corresponding to the first inspection item is 1; there is 1 abnormal object in the second inspection item, i.e. m=1, so the problem weight value corresponding to the second inspection item is 0.79; there is 1 inspection object state abnormality in the third inspection item, i.e. m=1, so the problem weight value corresponding to the third inspection item is 0.999; there is 0 inspection object state abnormality in the fourth inspection item, i.e. m=0, so the problem weight value corresponding to the fourth inspection item is 1, and then the health score (H) = 100*1*0.79*0.999*1=78.9. score
[0075] For the second health score estimation method, the problem that the health score cannot reflect the serious abnormality of the data center in the first health score estimation method can be solved, but the second health score estimation method can cause the health score to rapidly decrease due to temporary state abnormality of a certain inspection object, resulting in too low health score and inaccurate reflection of the current health of the data center, that is, the weight of a single inspection object in the entire health score is too large, and the state abnormality of the inspection object that causes the health score to rapidly decrease is generally caused by temporary connection failure (for example, temporary power-off, state locking, excessive load, etc.) of the inspection object. For example, assuming that server 1 temporarily has connection failure, at this time, four inspection items and three inspection objects are automatically inspected, and the inspection result is that the first inspection item, the second inspection item, the third inspection item and the fourth inspection item only have state abnormality of server 1 (that is, state abnormality of 1 inspection object), at this time, the health score = 100*0.79*0.79*0.999*0.999 = 62.3 points, which is caused by temporary connection failure (offline) of server 1, resulting in rapid decrease of the health score of the entire data center, which does not conform to the actual situation. The health score obtained by the second health score estimation method cannot accurately reflect the health of the data center, and the operation and maintenance personnel cannot use the health score to maintain the data center, resulting in low operation and maintenance efficiency of the data center.
[0076] The embodiment of the present application provides a data center health score estimation method, comprising: receiving an inspection result for an inspection object and an inspection item, wherein the inspection object is an object to be inspected in the data center; based on the inspection result, calculating a problem weight value corresponding to each inspection object, if the number of state abnormality inspection items existing in the inspection object is less than or equal to a preset number, calculating the problem weight value corresponding to the inspection object based on the problem level coefficient corresponding to the state abnormality inspection item, if the number of state abnormality inspection items existing in the inspection object is greater than the preset number, calculating the problem weight value corresponding to the inspection object based on the problem level coefficient corresponding to the state abnormality inspection item and a preset attenuation function; based on the problem weight value corresponding to each inspection object, estimating the health score of the data center. Through the problem level coefficient corresponding to the state abnormality inspection item and the exponential attenuation function, the weight of the inspection object in the overall health score is continuously reduced, the problem that the health score cannot accurately reflect the health of the data center caused by rapid decrease of the health score due to temporary connection failure of a certain inspection object is avoided, and the operation and maintenance efficiency of the data center is improved and the stable operation of the data center is ensured.
[0077] Embodiment one:
[0078] The following will be described in combination withFigures 2-6 The application provides a method for estimating the health score of a data center.
[0079] First, the application provides a method for estimating the health score of a data center. Figure 2 The application provides a method for estimating the health score of a data center.
[0080] S201, the user inputs the inspection object and the inspection item.
[0081] In one possible implementation, the user checks the inspection object and the inspection item of the inspection task through the front-end selection page provided by the automatic inspection center. The inspection object can be a server, an operating system, a service, etc., and the inspection item can be a server PSU (Power Supply Unit) health check, a server CPU (Central Processing Unit) health check, etc. Further, after the user checks the inspection object and the inspection item through the front-end selection page, the corresponding inspection task is automatically generated, and the user can also rename the inspection task through the front-end selection page.
[0082] In one possible implementation, the user can also input the corresponding inspection policy, for example, immediate inspection, periodic inspection, etc. If the inspection policy is periodic inspection, the user can also set the inspection period of the periodic inspection, for example, 6 hours, i.e., inspection once every 6 hours.
[0083] For example, the following describes how the user inputs the inspection object and the inspection item in combination with the front-end selection interface shown in FIG. 3. Figure 3 As shown in (a) of FIG. 3, the automatic inspection center displays the inspection selection interface 300 to the user, which includes the inspection object box 301 and the inspection item box 302. The inspection object box 302 includes the inspection objects of multiple data centers, as shown in (a) of FIG. 1.
[0084] Figure 3 As shown in (a) of FIG. 3, the inspection object box 301 includes: server 1, server 2, server 3, switch 4, switch 5, switch 6, etc. The inspection item box 302 includes the inspection items of multiple data centers, as shown in (a) of FIG. 2. Figure 3 Figure 3 As shown in (a) of FIG. 3, the inspection item box 302 includes: power status, CPU utilization, memory usage, network bandwidth utilization, network connection, etc. The user checks the inspection object and the inspection item based on the above. Figure 3 The front-end selection interface 300 shown in (a) is used to select the inspection object and the inspection item of the target inspection by controlling the mouse cursor, for example, the inspection object server 1, server 2, switch 5 and switch 6, and the inspection item power state, CPU utilization, memory usage, and network bandwidth utilization are selected by controlling the mouse cursor, and the confirmation button 303 is clicked by controlling the mouse cursor, and the input of the inspection object and the inspection item is completed.
[0085] Further, after the user controls the mouse cursor to click the confirmation button 303 on the inspection selection interface 300, the automation inspection center displays the inspection task interface 310 to the user, and the inspection task list 311 is displayed on the inspection task interface 310, and the inspection task list 311 includes the inspection task corresponding to the selected inspection object and inspection item, and in general, the first inspection task 1 is arranged in the inspection task list 311, and the user can set the inspection strategy of the inspection task 1 by clicking the inspection task 1 with the mouse cursor, and the specific setting of the inspection strategy is immediate inspection, and the user controls the mouse cursor to click the execution button 312 to execute the inspection task 1. Further, the user controls the mouse cursor to double-click the inspection task 1, and the inspection task 1 can be renamed.
[0086] Further, before the user inputs the inspection object and the inspection item, that is, before the user selects the inspection item and the inspection object through the front-end page provided by the automation inspection center, all the inspection items that can be selected by the user are registered in the automation inspection center through the inspection package when the automation inspection center is started. In order to facilitate understanding, the registration process of the inspection item of the automation inspection center will be introduced below. Figure 4 The registration process of the inspection item of the automation inspection center will be introduced below.
[0087] S401, in response to the starting operation of the user to the automation inspection center, the automation inspection center is started.
[0088] S402, in response to the starting of the automation inspection center, the inspection item set of the database is acquired.
[0089] Specifically, after the automation inspection center is started, all the inspection items stored in the database of the storage layer of the automation inspection center are acquired, and the inspection item set includes all the inspection items stored in the database.
[0090] S403, all the inspection items stored in the inspection package are acquired.
[0091] Specifically, the automation inspection center starts the coroutine to acquire all the inspection items stored in the inspection package.
[0092] In a possible implementation, the automatic inspection center acquires all the inspection items stored under the inspection package corresponding to a specific product scenario according to the product scenario. Each product scenario (platType) corresponds to one or more inspection packages, and the inspection package can exist in the form of a folder. The product scenarios include Human-Computer Interaction (HCI), High Performance Computing (HPC), Artificial Intelligence (AI), and the like.
[0093] S404, determining the storage condition of the inspection item based on the inspection item stored under the inspection package and the set of inspection items stored in the database.
[0094] Specifically, the storage condition of the inspection item includes three kinds, which are referred to as the first storage condition, the second storage condition, and the third storage condition. The first storage condition is that the inspection item exists in both the inspection package and the set of inspection items in the database, that is, the storage condition of the inspection item existing in both the inspection package and the set of inspection items in the database is the first storage condition. The second storage condition is that the inspection item exists in the inspection package but does not exist in the set of inspection items in the database, that is, the storage condition of the inspection item existing in the inspection package but not existing in the set of inspection items in the database is the second storage condition. The third storage condition is that the inspection item exists in the set of inspection items in the database but does not exist in the inspection package, that is, the storage condition of the inspection item existing in the set of inspection items in the database but not existing in the inspection package is the third storage condition.
[0095] When the storage condition of the inspection item is the first storage condition, S405 is performed.
[0096] When the storage condition of the inspection item is the second storage condition, S406 is performed.
[0097] When the storage condition of the inspection item is the second storage condition, S407 is performed.
[0098] S405, updating the inspection item stored in the database based on the configuration data of the inspection item stored under the inspection package.
[0099] The configuration data of the inspection item includes specific items and index data that need to be checked by the inspection item, the inspection rule (for example, inspection execution logic, condition, and the like) corresponding to the inspection item, and the like.
[0100] Specifically, for the inspection item existing in both the inspection package and the set of inspection items in the database, a new inspection item does not need to be re-registered, and the data of the inspection item in the database needs to be updated to the configuration data of the inspection item in the inspection package, that is, the registration of the inspection item can be completed.
[0101] S406, based on the configuration data of the inspection item stored under the inspection package, the inspection item is added to the database.
[0102] Specifically, for the inspection item that exists in the inspection package but does not exist in the inspection item set of the database, based on the configuration data of the inspection item stored under the inspection package, the inspection item is added to the database.
[0103] S407, delete the inspection item stored in the database.
[0104] Specifically, for the inspection item that exists in the inspection item set of the database but does not exist in the inspection package, the inspection item stored in the database is deleted.
[0105] Through the above S401-S407, after being registered to the automatic inspection center through the inspection package, the automatic inspection center can receive the inspection item and the inspection object input by the user.
[0106] S202, the automatic inspection center responds to the received inspection object and inspection item, and sends a calling request to the execution service.
[0107] Among them, the execution service can be understood as a service included in the execution layer 120 as shown in Figure 1 , such as: business service, automatic job service, etc.
[0108] Specifically, the application layer of the automatic inspection center receives the corresponding inspection task of the inspection object and the inspection item, and the logic and calling layer of the automatic inspection center sends a calling request to the execution service, so that the execution service executes the inspection operation on the inspection object and the inspection item of the data center in response to the calling request.
[0109] In one possible implementation, when the execution service is a business service, sending a calling request to the execution service can be: sending a calling request to a pre-set corresponding application program interface, so as to call the corresponding business service through the corresponding application program interface to execute the inspection operation on the inspection object and the inspection item of the data center.
[0110] In another possible implementation, when the execution service is an automatic job service, sending a calling request to the execution service can be: sending a calling request carrying the inspection script corresponding to the inspection task to the automatic job service, so that the automatic job service executes the inspection script, thereby realizing the inspection operation on the inspection object and the inspection item of the data center. Wherein, the inspection script can be obtained by pre-writing the execution logic of the inspection task in the form of a script, or based on the inspection object and the inspection item, the corresponding pre-set inspection script is determined, and the inspection script of the inspection task is obtained based on the corresponding pre-set inspection script. The application does not make specific limitation.
[0111] S203, in response to the invocation request, the execution service performs the inspection operation on the inspection objects and the inspection items of the data center.
[0112] In a possible implementation, in response to the invocation request, the corresponding business service is invoked through the corresponding business service interface to perform the inspection operation on the inspection objects and the inspection items of the data center.
[0113] In another possible implementation, the automation job service performs the inspection script carried in the invocation request in response to the invocation request, thereby implementing the inspection operation on the inspection objects and the inspection items of the data center.
[0114] Further, in a possible implementation, in the execution of the inspection operation on the inspection objects and the inspection items of the data center, the execution service can return the inspection progress to the automation inspection center at a regular time, and the automation inspection center can display the current inspection progress to the user, so that the user can know whether the inspection task is completed.
[0115] S204, when the execution service completes the inspection operation on the inspection objects and the inspection items of the data center, the execution service returns the inspection result to the automation inspection center.
[0116] The inspection result includes the inspection items of each inspection object corresponding to the state exception, and the number of inspection items corresponding to the state exception of each inspection object. In a possible implementation, the inspection result can also include the possible reasons for the state exception of the inspection item and the processing suggestions for solving the state exception of the inspection item.
[0117] Further, when the execution service does not complete the inspection operation on the inspection objects and the inspection items of the data center, but the time length of the execution of the inspection operation reaches a preset time length, the execution service returns the inspection result to the automation inspection center, and at this time, the inspection result includes the inspection result of the partially completed inspection of the inspection objects and the inspection items. In order to avoid the problem that a long time is consumed for waiting for the inspection and the completed inspection result cannot be obtained due to the performance limitation of the execution service or other situations, which seriously affects the efficiency of the automation inspection of the data center and causes poor user experience.
[0118] S205, the automation inspection center stores the received inspection result into a database.
[0119] Specifically, the logic of the automation inspection center parses the received inspection result, and stores the parsed inspection result into the database of the storage layer of the automation inspection center, so as to read the inspection result corresponding to the inspection task from the database in the future.
[0120] S206, the automation inspection center calculates the health score based on the inspection result.
[0121] In the embodiment of the present application, the health score is calculated based on the mechanism of the decay function.
[0122] For the convenience of understanding, the following will introduce in detail how the automatic inspection center calculates the health score based on the inspection results in the embodiment of the present application, in combination with formula (3), formula (4), formula (5), formula (6) and formula (7). First, the preset decay function will be introduced in combination with formula (3).
[0123] The decay function is a function whose value gradually decreases with the increase of the independent variable, and is usually used to describe the law of gradual weakening or reduction of the independent variable. Common decay functions include exponential decay function, power decay function, logarithmic decay function, etc. The preset decay function involved in the embodiment of the present application is an exponential decay function, which is specifically shown in formula (3).
[0124] (3)
[0125] Wherein, A is the initial value of the exponential decay function, λ is the decay rate, e is the base of natural logarithm, x is the independent variable. The characteristic of the exponential decay function is that when the independent variable x continuously increases, the function value will gradually decrease exponentially, but will not decrease to 0, but will gradually approach 0.
[0126] Specifically, the initial value A and the decay rate λ of the exponential decay function are preset, and the decay rate λ of the exponential decay function determines the descending speed of the function value, and the initial value A determines the maximum value of the function value. In order to make the exponential decay function gradually approach zero in a suitable range, the values of the initial value A and the decay rate λ need to be adjusted according to the specific application scenario.
[0127] The above briefly introduces the exponential decay function in combination with formula (3), and the following will introduce in detail the calculation method of the health score in combination with formula (4), formula (5), formula (6) and formula (7).
[0128] First, the problem weight value of each inspection object is calculated by formula (4).
[0129] (4)
[0130] Wherein, O rrepresents the problem weight value of the first r d represents the problem weight value of the first d k p represents the problem weight value of the first d p
[0131] Specifically, as formula (4), when p < y ( y is a positive integer set in advance), k p represents the problem level coefficient corresponding to the first d p p ≤ y The case specifically includes two kinds: the first is d ≤ y , p ≤ y , and p is certainly smaller than y ; the second is p ≤ y < d When p>y, k p is 1- f ( p ), that is, the problem weight coefficient of the first p is not the problem level coefficient corresponding to it, but the difference between 1 and the exponential decay function value of the independent variable x is p . That is, for the inspection item of a certain inspection object, if the inspection item is the first y of all the inspection items of the state abnormality existing in the inspection object, the problem weight coefficient corresponding to the inspection item is the problem level coefficient corresponding to the inspection item; if the inspection item is not the first y of all the inspection items of the state abnormality existing in the inspection object, the problem weight coefficient corresponding to the inspection item is the difference between 1 and the decay function corresponding to the inspection item, wherein the decay function value corresponding to the inspection item is the value of the preset decay function when the independent variable is the number of the inspection item in all the inspection items of the state abnormality.
[0132] For example, assuming y=4, the inspection result is that the first inspection object exists 6 inspection items of state abnormality, and the problem level coefficients corresponding to the 6 inspection items of state abnormality are 0.79, 0.9, 0.99, 0.999, 0.9, 0.79 respectively, and the exponential decay function At this time, the problem weight value of the first inspection object is calculated by the specific formula (5).
[0133] (5)
[0134] Among them, 1≤4, then k 1=0.79; 2≤4, then k 2=0.9; 3≤4, then k 3=0.99; 4≤4, then k 4=0.999; 5>4, then k 5=1- f (5) = 1-0.00005 = 0.99995; 6>4 k 6=1- f (6) = 1-0.00001 = 0.99999. The problem weight value of the first inspection object is 0.703.
[0135] Furthermore, when the inspection object does not have any inspection items with abnormal status, the problem weight value of the inspection object is 1.
[0136] It should be noted that, in one possible implementation, the arrangement order of the inspection items with abnormal status is determined according to a pre-set arrangement order; in another possible implementation, the inspection items with abnormal status are arranged from high to low according to the problem level. For example, taking the problem level coefficient in formula (5) as an example, the arrangement order of the problem level coefficients corresponding to the inspection items with abnormal status is 0.79, 0.79, 0.9, 0.9, 0.99, and 0.999, which is not specifically limited in this application.
[0137] After obtaining the problem weight value of each inspection object, the problem weight value of each inspection object is multiplied to obtain the cumulative result, and the cumulative result is multiplied by 100 to obtain the health score. Specifically, as shown in formula (6):
[0138] (6)
[0139] in, Score Indicates health score, h Indicates the total number of inspection objects in the inspection task. O r Indicates the r The problem weight value of each inspection object.
[0140] For example, the problem levels corresponding to the inspection items are fatal, serious, general and prompt, and the corresponding problem level coefficients of the four problem levels are set to 0.79, 0.9, 0.99 and 0.999 respectively. The exponential decay function is And the preset value y=4 is taken as an example, in the embodiment of the present application, when x is 5, 1- f ( 5 )=1-0.00005=0.99995, that is, 1- f ( 5 )<0.999, and with the increase of the independent variable x , 1- f ( x ) gradually tends to 0, that is, 1- f ( x ) gradually tends to 1, that is, when the number of abnormal state inspection items of a patrol object is greater than the preset value, the weight of the patrol object in the overall patrol is continuously reduced with the increase of the number of items.
[0141] Next, the calculation method of the health degree score in the embodiment of the present application is introduced as a whole in combination with formula (7).
[0142] (7)
[0143] Wherein, Score represents the health degree score, d represents the number of inspection items of the patrol object in this case, and b p represents the d th inspection item in the p state abnormal inspection item corresponding to the problem level coefficient, y represents the preset number (that is, the threshold number of state abnormal inspection items), lastScore represents the health degree score of the last time, lastScore and the initial value of is 100.
[0144] In order to facilitate understanding, the calculation method of the health degree score provided by the embodiment of the present application is introduced by way of example and comparison. For example, the user inputs 6 inspection items and 3 patrol objects for automatic patrol, and obtains the patrol result, wherein the problem level coefficients corresponding to the 6 inspection items are 0.99, 0.79, 0.79, 0.999, 0.79 and 0.9 respectively. The patrol result can be referred to table 3.
[0145] Table 3
[0146]
[0147] Taking the exponential decay function as , and the preset number y =4 as an example, the health degree score is calculated.
[0148] In the embodiment of the present application, for the patrol object 1, there is one state abnormal inspection item, that is,d =1<4, then Score 1(1)=100*0.9=90 (points); for the inspection object 2, there are six state abnormal inspection items, i.e. d =6>4, and lastScore = Score 1(1)=90, then Score 2(6)=90*0.99*0.79*0.79*0.99*(1- f (5))*(1-f(6))=90*0.99*0.79*0.79*0.99*0.99995*0.99999=55.05 (points); for the inspection object 3, there is no state abnormal inspection item, so the final health degree score=55.05 (points).
[0149] From the inspection results shown in Table 3, it can be seen that the state of the inspection object 2 is abnormal (for example: connection failure, temporary power down, etc.), which causes the state of all inspection items of the inspection object 2 to be abnormal. In the embodiment of the present application, not only the problem level coefficient caused by the state abnormality of the inspection item is considered, but also the health degree score is calculated in combination with the exponential decay function. The health degree score calculated in the embodiment of the present application is 50.05 points, while the health degree score calculated in the current technology is 39.49 points. By comparison, it can be seen that the way of calculating in combination with the decay function in the embodiment of the present application reduces the proportion (weight) of the inspection object 2 in the overall health degree score, avoiding the problem that the health degree score rapidly decreases due to the temporary abnormality of the inspection object 2, and cannot accurately reflect the health degree of the data center, thereby improving the operation efficiency of the data center and ensuring the stable operation of the data center.
[0150] Further, the following will be combined Figure 5 Another flowchart of a method for estimating the health degree score of a data center is shown, and the embodiment of the present application will be introduced.
[0151] S501, receiving the inspection results for the inspection objects and the inspection items.
[0152] Among them, the inspection object is an object to be inspected in the data center.
[0153] S502, based on the inspection results, calculating the problem weight value corresponding to each inspection object.
[0154] If the number of state abnormal inspection items existing in the inspection object is less than or equal to the preset number, the problem weight value corresponding to the inspection item is calculated based on the problem level coefficient corresponding to the state abnormal inspection item.
[0155] If the number of the inspection items representing the state abnormality of the inspection object is greater than the preset number, the problem weight value of the inspection item is calculated based on the problem level coefficient of the state abnormality inspection item and the preset attenuation function.
[0156] S503, the health score of the data center is estimated based on the problem weight value corresponding to each inspection object.
[0157] Specifically, the problem weight value corresponding to each inspection object is multiplied, and then multiplied by 100 to obtain the health score of the data center.
[0158] In the embodiments of the present application, the health score of the data center is estimated based on the inspection result through S501-S503. The health score is estimated by combining the problem level coefficient and the attenuation function, which avoids the problem that the health score cannot accurately reflect the health of the data center due to the continuous decline of the weight of the inspection object in the overall health score, and the rapid decline of the health score due to the temporary connection failure of a certain inspection object, thereby improving the operation efficiency of the data center and ensuring the stable operation of the data center.
[0159] S207, the health score is stored by the automatic inspection center.
[0160] Specifically, the health score is stored by the automatic inspection center in the database of the storage layer.
[0161] In a possible implementation, the automatic inspection center can display the health score of the inspection task and the specific inspection result, i.e., the inspection details of the inspection items and the inspection objects of the inspection task, to the user. Further, the automatic inspection center can also display the inspection suggestion corresponding to the inspection task to the user, including the possible reasons causing the state abnormality of the inspection items of the inspection object and the solution opinions, so that the operation and maintenance personnel can quickly maintain the fault of the data center according to the inspection suggestion, improve the operation efficiency of the data center, and ensure the reliability of the data center.
[0162] For ease of understanding, the following will be described in combination with Figure 6 An example will be given to introduce the interface for displaying the health score of the inspection task to the user.
[0163] For example Figure 6The health score interface 600 shown displays the health score of 63 to the user through the arc-shaped index bar. Further, the health score interface 600 can also display the execution time of the inspection task to the user, specifically including the execution start time and the execution end time of the inspection task, and the specific inspection result, for example, the number of normal inspection items, the number of abnormal inspection items, the number of non-involved inspection items, the number of failed inspection items, and the number of unchecked inspection items. The normal inspection item is an inspection item with a normal state, the abnormal inspection item is an inspection item with an abnormal state, the non-involved inspection item is an inspection item not involved in the inspection object, for example, the inspection object in the inspection task is a hardware device not including GPU, and the inspection item in the inspection task includes GPU utilization, at this time, the GPU utilization is a non-involved inspection item for the hardware device not including GPU, the failed inspection item is a failed inspection item, and the unchecked inspection item is an inspection item not completed within the specified time.
[0164] In the embodiment of the present application, after the inspection is received by the automatic inspection center, the problem weight value of each inspection object is calculated based on the inspection result, if the number of abnormal state inspection items existing in the inspection object is less than the preset number y , the problem level coefficients corresponding to the abnormal state inspection items are multiplied; if the number of abnormal state inspection items existing in the inspection object is greater than the preset number y , the problem level coefficients corresponding to the first y inspection items in the abnormal state inspection items are multiplied, then an exponential decay function is calculated according to the number of items in the abnormal state inspection items, and the exponential decay function is multiplied with the multiplication result of the problem level coefficients corresponding to the first y inspection items to obtain the problem weight value of the inspection object; the problem weight value of all inspection objects is multiplied with 100 to obtain the health score. Through the problem level coefficients corresponding to the abnormal state inspection items and the exponential decay function, the way of calculating the problem weight value of each inspection object makes the weight of the inspection object in the overall health score decrease constantly, avoiding the problem that the health score rapidly decreases due to the temporary connection failure of a certain inspection object, which cannot accurately reflect the health of the data center.
[0165] In addition, referring to Figure 7 , the embodiment of the present application also provides a computing device 700, which includes a processor 710 and a memory 720.
[0166] The processor 710 can include one or more processing cores. The processor 710 connects various parts within the computing device 700 by various interfaces and lines, executes the method provided by any one or more embodiments described above by running or executing instructions, programs, code sets or instruction sets stored in the memory 720, and calls data stored in the memory 720. Alternatively, the processor 710 can be implemented in at least one of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 710 can integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. It can be understood that the above-mentioned modem can also not be integrated into the processor 710, but can be implemented by a separate communication chip.
[0167] The memory 720 can include a random access memory (RAM) and can also include a read-only memory (ROM). Alternatively, the memory 720 includes a non-transitory computer-readable storage medium. The memory 720 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 720 can include a program storage area. The program storage area can store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playing function, an image playing function, etc.), and instructions for implementing the method of the embodiments of the present application, etc.
[0168] The processor 710 and the memory 720 are communicatively connected through a bus within the computing device 700, which can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0169] The embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores computer program instructions. When the computer program instructions run on a computing device, the computer program instructions make the computing device execute the data center health score estimation method in the above embodiment.
[0170] The explanation and beneficial effects of the related content in the computer readable storage medium provided above can refer to the corresponding effects of the data center health score estimation method in the above embodiment, and will not be repeated here.
[0171] The embodiment of the present application further provides a computer program product containing instructions, and when the instructions run on a computing device, the instructions make the computing device execute any one of the data center health score estimation methods in the above embodiment.
[0172] Although the present application is described herein in conjunction with various embodiments, those skilled in the art, with the benefit of the drawings, the disclosure, and the appended claims, can understand and implement other variations of the disclosed embodiments in the course of their work. In the claims, the word "comprising" does not exclude other components or steps, and the word "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. Measures described in mutually different dependent claims can be combined and can produce good results.
[0173] Although the present application is described herein in conjunction with specific features and embodiments thereof, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of the application. Accordingly, the specification and drawings are to be regarded simply as illustrative of the present application as defined by the appended claims, and it is intended to cover any and all modifications, variations, combinations or equivalents that fall within the scope of the present application. Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, it is intended to include these modifications and variations.
Claims
1. A method for estimating a health score of a data center, the method comprising: The method is applied to an automatic inspection center, and comprises the following steps: receiving an inspection result of an inspection object and an inspection item; the inspection object is an object to be inspected in a data center; based on the inspection result, calculating a problem weight value corresponding to each inspection object; if the number of inspection items representing state abnormalities of the inspection object is greater than a preset number, calculating the problem weight value corresponding to the inspection object based on a problem level coefficient corresponding to the inspection items representing state abnormalities and a preset attenuation function; if the number of inspection items representing state abnormalities of the inspection object is less than or equal to a preset number, calculating the problem weight value corresponding to the inspection object based on a problem level coefficient corresponding to the inspection items representing state abnormalities; based on the problem weight value corresponding to each inspection object, estimating a health score of the data center.
2. The method of claim 1, wherein, The problem weight value corresponding to the inspection object is a cumulative result of problem weight coefficients corresponding to the inspection items representing state abnormalities of the inspection object; If the inspection item is the first one in all the inspection items of the state abnormality existing in the inspection object, the problem weight coefficient corresponding to the inspection item is the problem level coefficient corresponding to the inspection item. y item, the problem weight coefficient corresponding to the inspection item is the problem level coefficient corresponding to the inspection item. If the inspection item is not the first one among all inspection items with abnormal status of the inspection object y Item, then the problem weight coefficient corresponding to the inspection item is the difference between 1 and the attenuation function value corresponding to the inspection item; the attenuation function value corresponding to the inspection item is the value of the preset attenuation function when the independent variable is the number of the inspection item in all inspection items in the abnormal state; wherein, y is the preset number, and y Is a positive integer.
3. The method of claim 1, wherein, The inspection result of the inspection object and the inspection item is obtained in the following way: receiving an inspection object and an inspection item input by a user; based on the received inspection object and the inspection item, sending a calling request to an execution service, so that the execution service performs an inspection operation on the inspection object and the inspection item in response to the calling request, and returns an inspection result of the inspection object and the inspection item.
4. The method of claim 3, wherein, The execution service comprises a business service; the business service is a service for performing an inspection task; The sending of the calling request to the execution service based on the received inspection object and the inspection item comprises: based on the received inspection object and the inspection item, sending a calling request to a corresponding application program interface, so that a corresponding business service is called through the corresponding application program interface to perform an inspection operation on the inspection object and the inspection item.
5. The method of claim 3, wherein, The execution service comprises an automatic job service; the automatic job service is a service for executing an automatic script; The sending of the calling request to the execution service based on the received inspection object and the inspection item comprises: based on the received inspection object and the inspection item, sending a calling request carrying an inspection script to the automatic job service, so that the automatic job service executes the inspection script to implement an inspection operation on the inspection object and the inspection item; the inspection script is an automatic script for inspecting the inspection object and the inspection item.
6. The method of claim 3, wherein, Before the receiving of the inspection object and the inspection item input by the user, the method further comprises the following steps: in response to the start of the automatic inspection center, obtaining an inspection item set stored in a database and all inspection items stored in an inspection package; the inspection package is used to store inspection items and configuration data of the inspection items; based on the inspection item set and all the inspection items stored in the inspection package, determining a storage condition of the inspection items; based on the storage condition of the inspection items, performing a registration operation corresponding to the storage condition.
7. The method of claim 6, wherein, The storage condition of the inspection item includes a first storage condition, a second storage condition and a third storage condition; the first storage condition is that the inspection item exists in both the inspection package and the inspection item set; the second storage condition is that the inspection item exists in the inspection package but does not exist in the inspection item set; and the third storage condition is that the inspection item exists in the inspection item set but does not exist in the inspection package; The registration operation corresponding to the storage condition of the inspection item is performed based on the storage condition of the inspection item, including: When the storage condition of the inspection item is the first storage condition, the configuration data of the inspection item stored in the inspection package is used to update the inspection item stored in the database; When the storage condition of the inspection item is the second storage condition, the configuration data of the inspection item stored in the inspection package is used to add the inspection item to the database; When the storage condition of the inspection item is the third storage condition, the inspection item stored in the database is deleted.
8. The method according to any one of claims 1 to 7, characterized in that, The preset attenuation function includes an exponential attenuation function; the formula of the exponential attenuation function is: ; wherein, A is an initial value of an exponential decay function, The data center health score estimation method includes: is a decay rate, e is a base of a natural logarithm, x is an argument; wherein the argument is a number of items of the inspection item in all inspection items of the state exception, an initial value of an exponential decay function A and a decay rate A processor and a memory; are preset.
9. A computing device, comprising: The processor and the memory are coupled; The memory is used to store computer program instructions; The processor is used to execute the computer program instructions stored in the memory to implement the data center health score estimation method in any one of claims 1-8.
Citation Information
Patent Citations
Method and system for evaluating health degree of computer equipment based on combined weight
CN115423318A