Method and device for detecting health state of server
By monitoring servers in a standard environment and matching influencing factors using a factor database, the problem of insufficient comprehensive consideration of multi-dimensional factors in existing technologies is solved, enabling dynamic assessment and improved accuracy of server health status, and supporting forward-looking prediction and resource optimization.
Patent Information
- Application Number
- CN202511085457.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies lack comprehensive consideration of multiple factors and cannot effectively integrate multiple related factors such as hardware configuration, software environment, load scenario, network conditions, and maintenance strategies. This results in insufficient accuracy in health status assessment, inability to quickly predict health indicators under different environments, high operation and maintenance costs, and low efficiency, failing to meet the real-time management needs of large-scale clusters.
The target server is monitored in a standard environment to obtain a health baseline value. Environmental impact factors and equipment difference factors are matched through a pre-built factor database. A continuous multiplication model is used to determine the health status index, which is then evaluated in combination with multi-dimensional impact factors.
It enables health status assessment based on multi-dimensional influencing factors, provides dynamic assessment capabilities that adapt to different hardware configurations, software environments, and operating conditions, significantly improves the accuracy and reliability of server health status detection, reduces false alarms and false negatives, and supports forward-looking prediction and resource allocation optimization.
Smart Images

Figure CN120994491A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of server operation and maintenance, and in particular to a method and device for detecting the health status of a server. BACKGROUND
[0002] In the era of cloud computing and big data, the stable operation of server clusters is of great importance. Existing methods for detecting the health status of a server mainly rely on real-time monitoring of single hardware indicators, such as CPU usage, memory occupancy, and other basic parameters, and determine the server status by setting fixed thresholds. However, this monitoring method based on static thresholds has obvious limitations and cannot adapt to dynamically changing operating environments, resulting in a large number of false positives and false negatives, and making it difficult to accurately reflect the true health status of the server.
[0003] Existing technologies lack consideration of the comprehensive influence of multi-dimensional factors and cannot effectively integrate hardware configuration, software environment, load scenario, network conditions, maintenance strategies, and other related factors. Different hardware configurations may exhibit completely different performance characteristics under the same load conditions, and the same server may exhibit different health levels under different network environments or maintenance cycles. Traditional detection methods cannot efficiently integrate these multi-dimensional influencing factors, resulting in insufficient accuracy of health status evaluation and lack of forward-looking prediction capabilities.
[0004] In addition, existing technologies have difficulty in quickly predicting accurate health status indicators when facing different hardware, software, loads, networks, and maintenance conditions, and cannot estimate the relationship between resource consumption and performance fluctuations in advance, making it difficult to quickly locate potential fault risks. The operation and maintenance cost is high and inefficient, and cannot meet the real-time management needs of large-scale clusters. SUMMARY
[0005] The present application provides a method and device for detecting the health status of a server to solve the problems of lack of comprehensive consideration of multi-dimensional factors, insufficient accuracy of health status evaluation, and inability to quickly predict health indicators under different environments in existing technologies.
[0006] In a first aspect, the present application provides a method for detecting the health status of a server, comprising:
[0007] Monitoring the target server under standard environment to obtain the health baseline value corresponding to the target server; the standard environment includes baseline hardware configuration, standard software environment, constant load scenario, standard network condition, and regular maintenance strategy;
[0008] According to the real-time parameters of the target server, matching the corresponding environmental influence factors and device difference factors from the pre-constructed factor database;
[0009] The health benchmark value is multiplied by the environmental impact factor and the equipment difference factor in sequence to determine the health state index corresponding to the target server.
[0010] In a second aspect, the application provides a server health state detection device, comprising:
[0011] The health benchmark value determination module is configured to monitor the target server under a standard environment to obtain a health benchmark value corresponding to the target server; the standard environment comprises a benchmark hardware configuration, a standard software environment, a constant load scenario, standard network conditions and a regular maintenance strategy.
[0012] The factor determination module is configured to match corresponding environmental impact factors and equipment difference factors from a pre-constructed factor database according to real-time parameters of the target server.
[0013] The health state index determination module is configured to multiply the health benchmark value by the environmental impact factor and the equipment difference factor in sequence to determine the health state index corresponding to the target server.
[0014] In a third aspect, the application provides a readable medium comprising execution instructions, when a processor of an electronic device executes the execution instructions, the electronic device executes the method according to any one of the first aspect.
[0015] In a fourth aspect, the application provides an electronic device comprising a processor and a memory storing execution instructions, when the processor executes the execution instructions stored in the memory, the processor executes the method according to any one of the first aspect.
[0016] The application provides a server health state detection method and device, which monitors a target server under a standard environment to obtain a health benchmark value corresponding to the target server; the standard environment comprises a benchmark hardware configuration, a standard software environment, a constant load scenario, standard network conditions and a regular maintenance strategy; matches corresponding environmental impact factors and equipment difference factors from a pre-constructed factor database according to real-time parameters of the target server; and multiplies the health benchmark value by the environmental impact factor and the equipment difference factor in sequence to determine the health state index corresponding to the target server. The health state evaluation based on multi-dimensional impact factors is realized, the limitations of traditional single-index monitoring are effectively solved through the combination of standardized benchmark values and personalized difference factors, dynamic evaluation capabilities suitable for different hardware configurations, software environments and running conditions are provided, and the accuracy and reliability of server health state detection are significantly improved.
[0017] The further effects of the above-mentioned non-conventional preferred modes will be described in the following in combination with the specific embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to make the objectives, technical solutions and advantages of the present application clearer, the accompanying drawings needed in the description of the embodiments or prior art will be briefly introduced. Obviously, the accompanying drawings described below are only some embodiments of the present application, and all other embodiments obtained by a person of ordinary skill in the art without creative work based on these drawings are within the protection scope of the present application.
[0019] Figure 1 A flowchart of a server health state detection method provided by an embodiment of the present application is shown in FIG. 1.
[0020] Figure 2 A flowchart of another server health state detection method provided by an embodiment of the present application is shown in FIG. 2.
[0021] Figure 3 A flowchart of another server health state detection method provided by an embodiment of the present application is shown in FIG. 3.
[0022] Figure 4 A flowchart of another server health state detection method provided by an embodiment of the present application is shown in FIG. 4.
[0023] Figure 5 A structural diagram of a server health state detection device provided by an embodiment of the present application is shown in FIG. 5.
[0024] Figure 6 A structural diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 6. DETAILED DESCRIPTION
[0025] In order to make the objectives, technical solutions and advantages of the present application clearer, the accompanying drawings needed in the description of the embodiments or prior art will be briefly introduced. Obviously, the accompanying drawings described below are only some embodiments of the present application, and all other embodiments obtained by a person of ordinary skill in the art without creative work based on these drawings are within the protection scope of the present application.
[0026] In the era of cloud computing and big data, stable operation of a server cluster is of great importance. Existing server health state detection methods mainly rely on real-time monitoring of single hardware indicators, such as CPU usage, memory occupancy and other basic parameters, and determine the server state by setting fixed thresholds. However, this monitoring method based on static thresholds has obvious limitations, cannot adapt to dynamically changing operating environments, is prone to a large number of false positives and false negatives, and is difficult to accurately reflect the real health state of the server.
[0027] The prior art lacks consideration of the comprehensive influence of multi-dimensional factors, and cannot effectively integrate multiple related factors such as hardware configuration, software environment, load scenario, network condition, and maintenance strategy. Different hardware configurations may exhibit completely different performance characteristics under the same load conditions, and the same server may exhibit different health levels under different network environments or maintenance periods. Traditional detection methods cannot efficiently integrate these multi-dimensional influencing factors, resulting in insufficient accuracy of health status evaluation and lack of forward-looking prediction ability.
[0028] In addition, the prior art has difficulty in quickly predicting accurate health status indicators under different hardware, software, load, network, and maintenance conditions, and cannot estimate the relationship between resource consumption and performance fluctuation in advance, making it difficult to quickly locate potential fault risks. The operation and maintenance cost is high and the efficiency is low, which cannot meet the real-time management needs of large-scale clusters.
[0029] To solve this problem, an embodiment of the present application proposes a server health status detection method, which aims to solve the problems of lack of comprehensive consideration of multi-dimensional factors, insufficient accuracy of health status evaluation, and inability to quickly predict health indicators under different environments in the prior art. In this embodiment, a server health status detection method includes:
[0030] Step 101, monitoring the target server under a standard environment to obtain a health baseline value corresponding to the target server; the standard environment includes a baseline hardware configuration, a standard software environment, a constant load scenario, a standard network condition, and a regular maintenance strategy.
[0031] Monitoring the target server under a standard environment to obtain a health baseline value is the basic link of the entire detection method. The establishment of the standard environment needs to consider the configuration of multiple dimensions, to ensure the formation of a stable test baseline. The baseline hardware configuration usually selects the mainstream server configuration in the market, including specific models of CPU processors, fixed capacity of memory modules, standard specifications of storage devices, and network interface cards and other core hardware components. The standard software environment covers the standardization settings of the operating system version, system patch level, middleware software configuration, and application program version at the software level.
[0032] The setting of the constant load scenario needs to simulate the typical work load in the real production environment, usually selecting 50% CPU utilization as the baseline load level, while coordinating the corresponding memory usage, disk I / O operation frequency, and network transmission volume. The standard network condition includes fixed network bandwidth configuration, such as 1Gbps network connection speed, as well as stable network delay and extremely low packet loss rate. The regular maintenance strategy defines the standard system maintenance period, including the time arrangement of regular restart frequency, patch update plan, log cleaning period, and other operation and maintenance operations.
[0033] Under these standardized conditions, the system needs to continuously monitor the target server for 72 hours or more to comprehensively collect resource consumption data and performance output data of the target server. Resource consumption data includes quantitative indicators such as CPU usage time, memory usage, disk read / write times, network traffic, etc., and performance output data covers business indicators such as response time, processing throughput, transaction completion rate, etc. By calculating the ratio of total resource consumption to total performance output, a dimensionless health benchmark value is obtained, which will be used as a reference standard for all subsequent health status evaluations.
[0034] Step 102, according to the real-time parameters of the target server, match the corresponding environmental impact factors and device difference factors from the pre-constructed factor database.
[0035] The real-time parameters of the target server include: determining the hardware configuration parameters of the target server based on the CPU model, memory capacity and storage type corresponding to the target server; determining the software version parameters of the target server based on the operating system version and application program version corresponding to the target server; determining the environmental dynamic parameters of the target server based on the real-time load ratio, network packet loss rate and maintenance strategy execution period corresponding to the target server; determining the implementation parameters based on the hardware configuration parameters, software version parameters and environmental dynamic parameters.
[0036] The acquisition of real-time parameters of the target server requires multi-dimensional data collection and analysis to fully understand the current running state and configuration characteristics of the server. The determination process of hardware configuration parameters involves deep identification and performance characteristic analysis of server physical components, which is the hardware basis for building an accurate evaluation model. The system accurately identifies the CPU model through various technical means, including reading the CPUID instruction return value of the processor, parsing the system hardware information file, and obtaining detailed specification parameters of the processor through the hardware abstraction layer interface.
[0037] The determination of software version parameters requires comprehensive scanning and version identification of all key software components running on the server. The identification of the operating system version not only obtains the type and major version number of the operating system, but also deeply analyzes the kernel version, patch installation status, system configuration parameters, security policy settings and other detailed information.
[0038] The identification of application program version covers all business-critical application software running on the server, including database management systems, web server software, application server platforms, middleware components, monitoring agent programs and other core software components. The system scans the installation directory, configuration file, process list, service registration information and other information sources of the application program to obtain detailed information such as the version number, configuration parameters, running state, resource occupation of each application program.
[0039] The acquisition of environmental dynamic parameters requires continuous tracking and data collection on the running state of the server through a real-time monitoring system. The calculation of real-time load proportion involves comprehensive evaluation of multiple system resource dimensions, including real-time monitoring of key indicators such as CPU usage, memory occupancy, disk I / O load, and network transmission load.
[0040] The measurement of network packet loss rate adopts a comprehensive monitoring strategy combining active probing and passive monitoring. Active probing sends test packets to the preset target address at regular intervals, and calculates network quality indicators such as transmission success rate, round-trip time, and path tracking information. Passive monitoring identifies network anomalies and performance problems by analyzing network interface statistics, firewall logs, and application network error reports. The system also monitors network performance indicators such as network bandwidth utilization, connection concurrency, TCP connection state distribution, and network protocol error rate to comprehensively assess the impact of the network environment on server performance.
[0041] Tracking the execution cycle of maintenance strategies requires a complete maintenance operation record and scheduling management system. The system records detailed information for each maintenance operation, including operation type, execution time, duration, impact range, and operation results. Maintenance operation types include system restart, software update, configuration change, hardware maintenance, security patch installation, performance tuning, and backup recovery. By analyzing the execution cycle and frequency pattern of maintenance operations, the system can predict the planned time of the next maintenance operation and assess the cumulative impact of the time interval between the current time and the last maintenance operation on system stability and performance.
[0042] The determination process of implementation parameters requires the comprehensive integration and standardization of hardware configuration parameters, software version parameters, and environmental dynamic parameters. The system establishes a multi-dimensional parameter correlation analysis model to identify the mutual influence and coupling effect between different parameters. Hardware configuration parameters provide a basic performance benchmark for system performance evaluation, software version parameters reflect the performance adjustment of the system software environment, and environmental dynamic parameters represent the current actual working state and load level of the system. Through the organic combination of these three dimensions of parameters, the system can construct a complete and accurate real-time state portrait of the server, providing a data foundation for subsequent health status calculation and evaluation.
[0043] By changing any one of the constant load scenario, standard network conditions, or regular maintenance strategy separately while keeping other conditions consistent with the standard environment, monitor the target server state changes to determine the environmental impact factor; by changing any one of the baseline hardware configuration or standard software environment separately while keeping other conditions consistent with the standard environment, monitor the target server state changes to determine the device difference factor; integrate the environmental impact factor and the device difference factor to generate a factor database.
[0044] The construction of the factor database needs to be completed through a large number of single variable control experiments. In the process of determining the environmental impact factor, the system will change one parameter in the load scenario, network condition or maintenance strategy while keeping the benchmark hardware configuration and standard software environment unchanged. For example, the CPU load is adjusted from the standard 50% to 20%, 70% or 90%, and the health status of the server under these different load levels is monitored, and the load impact factor corresponding to different load levels is obtained through comparative analysis. Similarly, by changing the network bandwidth, introducing different degrees of packet loss rate, adjusting the maintenance operation frequency, etc., the system can obtain the values of the network impact factor and the maintenance impact factor. The load impact factor, network impact factor and maintenance impact factor are collectively referred to as environmental impact factors.
[0045] The acquisition process of the device difference factor needs to change the hardware configuration or software environment parameters alone under the standard load, network and maintenance conditions. By using different models of CPU, different capacities of memory, different types of storage devices, etc. Hardware changes, and installing different versions of operating system, different configurations of middleware, etc. Software changes, the system can quantify the specific influence of these differences on the health status of the server. All these environmental impact factors and device difference factors obtained through controlled variable experiments are finally integrated and stored in the factor database, forming a complete knowledge base covering the influence relationship of multi-dimensional parameters.
[0046] In the actual matching process, the system adopts a multi-level matching algorithm to ensure the accuracy and applicability of factor selection. The matching of hardware configuration parameters is first based on the CPU model for accurate matching, and the system will look up the hardware difference factor in the factor database that is completely consistent with the CPU model of the target server. If an accurate matching model cannot be found, the system will use a similarity algorithm to calculate the similarity according to the CPU architecture type, core number, frequency range, cache size and other key parameters, and select the difference factor corresponding to the closest CPU model. The matching of memory capacity is based on capacity ladder for hierarchical matching, and the system pre-sets the impact factors corresponding to different memory capacity ranges, such as 8GB-16GB, 16GB-32GB, 32GB-64GB, etc. The memory configuration of the target server will be classified into the corresponding interval and obtain the corresponding difference factor.
[0047] The matching process of software environment parameters needs to consider the compatibility of operating system versions and the combined effect of software stacks. The system first performs version matching on the operating system, considering not only the major version number but also the differences in patch levels and kernel versions. For application version matching, the system analyzes the performance differences and resource consumption characteristics between different versions and selects the most suitable software difference factor. When encountering software versions not included in the database, the system performs interpolation calculations based on version evolution rules and performance test data to estimate the corresponding impact factor values.
[0048] The matching of environmental dynamic parameters uses a real-time dynamic adjustment mechanism, and the selection of load impact factors is based on weighted calculations of comprehensive indicators such as current CPU utilization, memory usage, and disk I / O load. The system monitors the changes in these load parameters in real time and selects the corresponding load impact factor according to the pre-set load interval mapping rules. The matching of network impact factors not only considers the current network bandwidth utilization but also analyzes network quality indicators such as network delay, packet loss rate, and connection concurrency to determine the most suitable network impact factor through multi-dimensional network state evaluation.
[0049] The matching of maintenance impact factors requires analysis of the execution status of the current maintenance strategy and historical maintenance records. The system evaluates the time interval from the last system restart, patch update, log cleaning, and other maintenance operations at the current time point, combines the frequency and intensity of maintenance operations, and selects the corresponding maintenance impact factor. For servers that are performing maintenance operations, the system applies special maintenance period impact factors to reflect the temporary impact of maintenance operations on server performance and stability.
[0050] Step 103, continuously multiply the health benchmark value with environmental impact factors and device difference factors to determine the health status index corresponding to the target server.
[0051] The calculation of the health status index is achieved by continuously multiplying the health benchmark value with each dimension impact factor. This product model design considers the interaction between each impact factor, which can more realistically reflect the comprehensive impact of multi-dimensional parameters on server health status. When an impact factor is greater than 1, it indicates that the factor has a positive impact on health status, making the server perform better than the benchmark state; when the impact factor is less than 1, it indicates that the factor has a negative impact on health status; when the impact factor is equal to 1, it indicates that the factor remains consistent with the benchmark state. The final calculated health status index can quantitatively represent the health level of the target server under the current running conditions relative to the benchmark state.
[0052] Through this calculation method, the system can not only provide the health status assessment at the current time, but also predict the health status change trend of the server in the future period of time based on the trend analysis of historical data. Meanwhile, the model also supports reverse calculation function, which can calculate the required resource configuration or predict the performance output range under the given resource condition according to the expected health index level, and provides a scientific basis for cluster capacity planning and cost control.
[0053] The health grading threshold corresponding to the target server is determined, and the target server is graded and warned based on the health grading threshold and the health status index.
[0054] Determining the health grading threshold corresponding to the target server needs to consider multiple factors such as server hardware configuration, business application type and historical performance. The system first analyzes the hardware performance level of the target server, and high-performance servers should maintain a higher health status index, while entry-level servers can be relatively relaxed. The hardware differences such as CPU performance level, memory capacity configuration and storage subsystem type directly affect the expected value setting of the health status.
[0055] The business application type has a decisive influence on the threshold setting. Key business applications such as financial transaction systems require high server health status, and strict grading thresholds need to be set. General business applications such as office systems have relatively relaxed requirements. The system establishes a multi-dimensional threshold determination model, which comprehensively considers factors such as technical specifications, business importance, historical performance and statistical benchmarks of similar servers.
[0056] The health grading threshold is usually divided into five levels: excellent, good, normal, warning and danger. For example, the excellent level corresponds to a health status index greater than 1.2, indicating that the performance is significantly better than the benchmark; the good level corresponds to the interval 1.0-1.2, indicating that the running state is good; the normal level corresponds to the interval 0.8-1.0, indicating that it is basically normal but may have slight fluctuations; the warning level corresponds to the interval 0.6-0.8, indicating that there is a significant performance decline; and the danger level is less than 0.6, indicating that there is a serious risk of failure.
[0057] The system compares the real-time calculated health status index with the preset threshold to determine the current health level and trigger corresponding warnings according to the level change. The warning trigger considers multiple dimensions such as absolute level, change trend, duration and change amplitude to avoid false positives caused by short-term fluctuations.
[0058] The system uses time series analysis to identify the health status change trend, and sends a forward-looking warning when a sustained downward trend is detected. The duration assessment of abnormal state distinguishes between transient abnormalities and sustained abnormalities, and the transient abnormalities have a lower warning level, while the sustained abnormalities require immediate human intervention.
[0059] Different early warning levels correspond to different response strategies. For example, a normal level is regularly checked and evaluated; a warning level timely notifies the operation and maintenance personnel through emails, short messages, etc., and provides detailed state analysis and processing suggestions; a dangerous level immediately sends an emergency warning, simultaneously notifies through multiple channels, and automatically starts emergency response processes such as fault isolation, traffic transfer, and the like.
[0060] It can be seen from the above technical solutions that the embodiment has the beneficial effects that:
[0061] The server health state detection method provided by the embodiment monitors a target server under a standard environment to obtain a health benchmark value corresponding to the target server; the standard environment includes a benchmark hardware configuration, a standard software environment, a constant load scenario, standard network conditions, and a routine maintenance strategy; according to real-time parameters of the target server, corresponding environmental impact factors and device difference factors are matched from a pre-constructed factor database; and the health benchmark value is continuously multiplied by the environmental impact factors and the device difference factors to determine a health state index corresponding to the target server. The health state evaluation based on multi-dimensional impact factors is realized, the combination of the standardized benchmark value and the personalized difference factors effectively solves the limitations of traditional single-index monitoring, provides dynamic evaluation capability suitable for different hardware configurations, software environments, and running conditions, and significantly improves the accuracy and reliability of server health state detection.
[0062] Figure 1 The server health state detection method provided by the embodiment monitors a target server under a standard environment to obtain a health benchmark value corresponding to the target server; the standard environment includes a benchmark hardware configuration, a standard software environment, a constant load scenario, standard network conditions, and a routine maintenance strategy; the server health state detection method provided by the embodiment monitors a target server under a standard environment to obtain a health benchmark value corresponding to the target server; the standard environment includes a benchmark hardware configuration, a standard software environment, a constant load scenario, standard network conditions, and a routine maintenance strategy; and the server health state detection method provided by the embodiment monitors a target server under a standard environment to obtain a health benchmark value corresponding to the target server; the standard environment includes a benchmark hardware configuration, a standard software environment, a constant load scenario, standard network conditions, and a routine maintenance strategy.
[0063] As shown in Figure 2 Another specific embodiment of the server health state detection method.
[0064] In the embodiment, the server health state detection method includes the following steps:
[0065] Step 201, monitoring a target server under a standard environment to obtain a health benchmark value corresponding to the target server; the standard environment includes a benchmark hardware configuration, a standard software environment, a constant load scenario, standard network conditions, and a routine maintenance strategy.
[0066] Step 202, obtaining a total amount of resource consumption and a total amount of performance output of the target server under the standard environment.
[0067] Under the control conditions of the standard environment, the system needs to conduct long-term continuous and comprehensive monitoring on the target server. The monitoring period is usually set to 72 to 168 hours to ensure that the running characteristics and performance fluctuations of the server at different time periods can be fully captured. The collection of total resource consumption covers the use of all key resources during server operation, including the cumulative consumption of CPU time slices, memory space occupation statistics, the number and data volume of disk read and write operations, network transmission byte number and connection number, and other multi-dimensional indicators.
[0068] The statistics of CPU resource consumption need to accurately record the working time of the processor in user mode and kernel mode, as well as the distribution of idle time and I / O waiting time. The system will sample the instantaneous usage rate of the CPU at fixed time intervals (usually 1 second or 5 seconds), and then calculate the total CPU time consumption during the entire monitoring period by integration. The statistics of memory resources not only include the use of physical memory, but also need to consider the swapping of virtual memory, the use of cache and buffer, and the degree of memory fragmentation. The monitoring of disk I / O resources needs to separately count the number of read and write operations, data transmission volume, average response time, and queue depth, etc. indicators, fully reflecting the working load of the storage system.
[0069] The collection of network resource consumption needs to monitor the inbound and outbound data traffic, the number of network connection establishment and disconnection, and the processing overhead of the network protocol stack. At the same time, the system also needs to record the power consumption of the server, including the power consumption distribution of CPU, memory, disk, network equipment, and the energy efficiency indicators of the whole machine. These resource consumption data are collected in real time by special monitoring agent programs to ensure the accuracy and integrity of the data.
[0070] The measurement of performance output total quantity focuses on the actual service capacity and business processing effect that the server can provide under standard load conditions. Performance output indicators mainly include system response time, transaction processing throughput, concurrent processing capacity, data processing rate, and other key performance parameters. The statistics of response time need to measure the complete time period from receiving the request to returning the result, including network transmission delay, business logic processing time, database query time, and other time-consuming links. The throughput indicator reflects the number of business operations that the server can complete per unit time, such as the number of HTTP requests processed per second, the number of database transactions, the amount of file transfer, etc.
[0071] The evaluation of concurrent processing capability requires simulating a scenario of multiple users accessing simultaneously, measuring the performance and stability of the server under different levels of concurrency. The system gradually increases the number of concurrent connections, records the response time distribution, success rate, error rate and other indicators at each level of concurrency, until it reaches the performance inflection point of the system. The measurement of data processing rate is aimed at different types of data operations, such as text processing, image processing, numerical calculation, etc., to evaluate the processing efficiency and accuracy of the server.
[0072] Step 203, calculate the ratio of total resource consumption and total performance output, and take the ratio as the health benchmark value.
[0073] Since resource consumption and performance output involve different measurement units and numerical ranges, the system first needs to standardize and non-dimensionalize these raw data. The calculation of total resource consumption uses a weighted comprehensive method, assigning different types of resource consumption such as CPU time consumption, memory usage, disk I / O operation volume, network transmission volume, etc. according to their impact on system performance.
[0074] CPU resource consumption usually has the highest weight proportion, because processor performance directly determines the computing power of the server. The weight of memory resources is second, reflecting the constraint of memory capacity and bandwidth on system performance. The weights of disk I / O and network transmission are adjusted according to the characteristics of specific application scenarios. For I / O intensive applications, the weight of disk performance will be correspondingly increased, while for network service type applications, the weight of network resources is more important. Through this weighted calculation method, the system can obtain a total quantity value that comprehensively reflects the resource consumption level of the server.
[0075] The calculation of total performance output also uses a multi-index comprehensive evaluation method, converting performance indicators such as response time, throughput, concurrent capability, processing rate, etc. into a unified performance score. The response time indicator uses inverse transformation, i.e. the shorter the response time, the higher the performance score. Throughput and processing rate indicators directly reflect the processing capacity of the system, and the larger the value, the better the performance. Concurrent processing capability is quantified by the maximum stable concurrency number, reflecting the load bearing capacity of the system. After normalization, these different performance indicators are weighted according to business importance and the weighted average value is calculated to obtain the total performance output.
[0076] The formula for calculating the health benchmark value is = total resource consumption / total performance output. This ratio reflects the resource utilization efficiency of the server in the standard environment. The smaller the ratio, the higher the performance output the server can achieve with less resource consumption, indicating that the system is running efficiently and in good health. The larger the ratio, the more resources the server needs to consume to achieve the same performance level, indicating that there may be a performance bottleneck or system anomaly. To ensure the stability and representativeness of the benchmark value, the system will statistically analyze the data during the entire monitoring period, eliminate abnormal values and noise data, and use the average value of multiple measurements as the final health benchmark value.
[0077] Step 204, according to the real-time parameters of the target server, match the corresponding environmental impact factors and device difference factors from the pre-constructed factor database.
[0078] Step 205, continuously multiply the health benchmark value with the environmental impact factors and device difference factors to determine the health status index corresponding to the target server.
[0079] Through the above technical solution, the beneficial effects of the present embodiment are: by comprehensively considering multi-dimensional resource consumption and performance, a more accurate and comprehensive health status benchmark is provided. The ratio calculation method of resource consumption and performance output simplifies the complex multi-index evaluation process, and converts multi-dimensional monitoring data into a single numerical index that is easy to understand and apply. This health benchmark value not only reflects the current running state of the server, but also provides an important reference for subsequent performance optimization and capacity planning.
[0080] As shown in Figure 3 , another specific embodiment of a server health state detection method of the present application is provided. The present embodiment is further described on the basis of the foregoing embodiment.
[0081] In the present embodiment, a server health state detection method includes the following steps:
[0082] Step 301, monitoring the target server in a standard environment to obtain the health benchmark value corresponding to the target server; the standard environment includes a benchmark hardware configuration, a standard software environment, a constant load scenario, standard network conditions, and a regular maintenance strategy.
[0083] Step 302, according to the real-time parameters of the target server, match the corresponding environmental impact factors and device difference factors from the pre-constructed factor database.
[0084] Step 303, according to the hardware configuration parameters and software version parameters, call the corresponding device difference factors from the factor database.
[0085] According to the hardware configuration parameters and software version parameters, the corresponding device difference factor is called from the factor database. First, the hardware configuration of the target server needs to be fully identified and parameter extracted, which involves deep scanning and feature analysis of server hardware components. The identification of CPU model not only includes the brand and specific model of the processor, but also needs to obtain detailed technical parameters such as CPU architecture type, core number, thread number, base frequency, maximum acceleration frequency, cache level structure, process technology, etc. The system fully understands the performance characteristics and processing capacity of the CPU by reading the CPUID instruction return value of the CPU, parsing the processor specification document, and running the benchmark test program.
[0086] The determination of memory capacity needs to count the total capacity of all memory modules, and also needs to analyze the type specification, working frequency, timing parameter, channel configuration and other key parameters of the memory. The system will detect the DDR version of the memory, whether it supports ECC error correction, dual-channel or multi-channel configuration state, and the working mode of the memory controller. The identification of storage type covers the complete configuration information of main storage device and auxiliary storage device, including the type of hard disk drive (mechanical hard disk HDD or solid state disk SSD), interface standard (SATA, SAS, NVMe, etc.), capacity size, speed parameter, cache size and other technical indicators.
[0087] After obtaining detailed hardware configuration parameters, the system starts the matching and calling process of device difference factors. The device difference factors in the factor database are stored in layers according to the type of hardware components, and each hardware configuration corresponds to one or more difference factor values. The matching of CPU difference factors first tries to perform accurate matching, and the system will find the record in the database that is exactly the same as the target CPU model. If accurate matching cannot be found, the system will start the intelligent similarity matching algorithm, which calculates the similarity between the CPU model in the database by analyzing the performance benchmark test score, power consumption level, architecture characteristics and other multi-dimensional parameters, and selects the difference factor corresponding to the CPU model with the highest similarity as the reference value.
[0088] The calling of memory difference factors is based on the dual matching mechanism of capacity classification and performance classification. The system first classifies the target server into the corresponding capacity interval according to the total memory capacity, such as below 8GB, 8-16GB, 16-32GB, 32-64GB, 64GB and above, etc. Then further consider the type and performance parameters of the memory, such as the performance difference between DDR4 and DDR5, the influence of different frequency specifications, the stability difference between ECC memory and ordinary memory, etc. Finally, determine the appropriate memory difference factor. The matching of storage difference factors needs to consider multiple dimensions such as the type, capacity, performance level and other factors of the storage device. There is a significant performance difference between SSD and HDD, and different interface standards and capacity specifications will also have different degrees of impact on the overall performance of the system.
[0089] The identification and matching process of software version parameters needs to consider the combined effect and compatibility influence between the operating system and application program version. The determination of the operating system version not only identifies the type and major version number of the operating system, but also needs to obtain detailed information such as kernel version, patch installation status, system configuration parameters, etc. The system scans the version file of the operating system, the kernel module list, the system service configuration, the security update record and other information sources to build a complete operating system environment portrait. The identification of application program version covers all key application software running on the server, including the version information and configuration status of database systems, web servers, application servers, middleware platforms and other core business components.
[0090] The invocation of software difference factors needs to consider the combined effect between the operating system and application program version. Different operating system versions have performance differences in resource management, process scheduling, memory allocation, network processing, etc., and different application program versions also affect the resource consumption mode and performance of the system. The system configures the corresponding difference factor for each software combination by analyzing the compatibility matrix, performance test report, known problem list and other information of the software version. When encountering a software version combination not included in the database, the system will estimate the corresponding software difference factor value based on the version evolution law and performance change trend.
[0091] Step 304, according to the environmental dynamic parameters, call the corresponding environmental impact factors from the factor database.
[0092] Real-time acquisition of environmental dynamic parameters needs to be achieved through various monitoring methods and data collection techniques to ensure that the current running state and environmental changes of the server can be accurately captured. The calculation of real-time load ratio involves comprehensive evaluation of multiple dimensions such as CPU usage, memory occupancy, disk I / O load, network transmission load, etc. The system needs to continuously monitor the instantaneous usage of these resources and calculate the average load level through sliding time window.
[0093] The monitoring of CPU load not only needs to count the overall usage, but also needs to analyze the time allocation of user mode and kernel mode, interrupt processing overhead, context switch frequency and other detailed indicators. The system will use multiple sampling strategies, including timed sampling, event-driven sampling, load threshold triggered sampling, etc., to ensure that load changes and abnormal fluctuations can be discovered in time. The monitoring of memory load needs to focus on key indicators such as physical memory usage, virtual memory swap situation, cache hit rate, memory fragmentation level, etc. These parameters directly affect the response speed and stability of the system.
[0094] The measurement of network packet loss rate requires a combination of active probing and passive monitoring. Active probing sends test packets to a predetermined target address, calculates the packet loss rate and delay jitter of the network link by counting the success rate of packet transmission and round-trip time. Passive monitoring identifies network anomalies and performance problems by analyzing network interface statistics, system log records, and application network error reports. The system also monitors network performance indicators such as network bandwidth utilization, connection concurrency, and TCP connection state distribution to assess the impact of the network environment on server performance.
[0095] Maintaining the tracking of the maintenance policy execution cycle requires a complete maintenance operation record and scheduling management system. The system records detailed information such as the type, execution time, duration, and impact range of each maintenance operation, including system restart, software update, configuration change, hardware maintenance, and security patch installation. By analyzing the execution cycle and frequency pattern of maintenance operations, the system can predict the next maintenance operation time and assess the impact of the time interval between the current time and the last maintenance operation on system stability and performance.
[0096] The calling process of environmental impact factors uses a multi-level matching and interpolation algorithm to ensure the accuracy of factor selection. The matching of load impact factors is based on the load interval mapping table, which defines the impact factor values corresponding to different load levels, such as 0%~20% low load, 20%~50% medium load, 50%~80% high load, and 80%~100% super high load. When the real-time load of the target server is between two preset intervals, the system uses linear interpolation or spline interpolation to calculate the accurate load impact factor. The determination of network impact factors needs to consider multiple network quality indicators such as packet loss rate, delay, and bandwidth utilization. The system combines these indicators into a single network quality score through weighted calculation, and then maps it to the corresponding network impact factor.
[0097] The selection of maintenance impact factors is based on maintenance state analysis and time decay model. The system evaluates the current maintenance state, such as whether it is performing maintenance operations, the time interval from the last maintenance, and the cumulative impact of maintenance operations. For servers that have just completed maintenance operations, the system applies a higher maintenance impact factor to reflect the optimization effect of the system after maintenance.
[0098] Step 305, continuously multiply the health baseline value with the environmental impact factor and the equipment difference factor to determine the health status index corresponding to the target server.
[0099] Through the above technical solutions, the beneficial effects of the present embodiment are: by automatically identifying and analyzing the hardware component characteristics of the server, including processor performance, memory configuration, storage type, and other key parameters, personalized health status evaluation standards are provided for servers with different hardware configurations. The version matching function of the software environment ensures that the differences between the operating system and the application program accurately reflect the evaluation results, avoiding the problem of ignoring the impact of the software environment in traditional methods. Real-time monitoring of environmental dynamic parameters and factor calling mechanisms enable the system to quickly respond to changes in server operating status, providing dynamic health status evaluation results. Through the comprehensive consideration of multi-dimensional parameters, the accuracy and reliability of health status detection are significantly improved, reducing the incidence of false positives and false negatives.
[0100] As shown in Figure 4 , it is another specific embodiment of the server health status detection method of the present application. The present embodiment is further described on the basis of the foregoing embodiment.
[0101] Step 401, monitoring the target server in a standard environment to obtain the health baseline value corresponding to the target server; the standard environment includes baseline hardware configuration, standard software environment, constant load scenario, standard network condition and routine maintenance strategy.
[0102] Step 402, according to the real-time parameters of the target server, matching the corresponding environmental impact factors and device difference factors from the pre-constructed factor database.
[0103] Step 403, continuously multiplying the health baseline value, environmental impact factors and device difference factors to determine the health status index corresponding to the target server.
[0104] Step 404, when the target server runs in the same working condition as the standard environment, collecting the working condition resource consumption and working condition performance data corresponding to the target server.
[0105] When the target server runs in the same working condition as the standard environment, collecting the working condition resource consumption and working condition performance data corresponding to the target server requires establishing a working condition matching mechanism to ensure that the environmental conditions of data collection are highly consistent with the standard environment when establishing the health baseline value. The system will continuously monitor the operating environment parameters of the target server, and when it detects that the key parameters such as hardware configuration, software environment, load scenario, network condition and maintenance strategy match the standard environment specifications, it will automatically start the working condition data collection process.
[0106] The determination of the working condition matching requires accurate comparison of multiple environmental dimensions. The hardware configuration matching requires that the CPU model, memory capacity, storage type, and other core components be completely consistent with the reference environment or within the allowed difference range. The software environment matching requires verification of the consistency of the operating system version, key application program version, system configuration parameters, and other software level factors with the standard environment. The load scenario matching requires that the CPU usage, memory occupancy, disk I / O load, and other indicators of the current server be within a reasonable interval of the standard load level, usually allowing a deviation range of 5% to 10%. The network condition matching requires confirmation that the current network bandwidth, delay, packet loss rate, and other parameters are basically consistent with the standard network environment. The maintenance strategy matching requires that the current maintenance period and operation frequency be consistent with the regular maintenance strategy.
[0107] The collection of working condition resource consumption data covers the overall resource usage during server operation, including CPU time slice consumption, memory space occupancy, disk read / write operation volume, network transmission traffic, and other key resource indicators. The system uses high-precision monitoring tools and sampling techniques to ensure the accuracy and integrity of data collection. The CPU resource consumption statistics need to record the time allocation of user mode and kernel mode, interrupt processing overhead, context switch frequency, and other detailed indicators to obtain comprehensive processor usage data through multiple levels of performance counters. The memory resource consumption monitoring not only includes the usage of physical memory, but also needs to track the running status of memory subsystems such as virtual memory swapping, cache utilization, and memory allocation mode.
[0108] The collection of disk I / O resources needs to separately count the read / write operation times, data transmission volume, average response time, queue depth, and other performance indicators of different storage devices, and distinguish between sequential access and random access modes. The network resource consumption monitoring includes the statistics of inbound and outbound traffic, the overhead of connection establishment and maintenance, the CPU consumption of protocol processing, and other resource usage of the network subsystem. The system also records the overall power consumption of the server, including the energy consumption distribution and efficiency indicators of each hardware component.
[0109] The collection of working condition performance data focuses on the actual business processing capacity and service quality performance of the server under the standard working condition. Performance indicators include system response time, transaction processing throughput, concurrent processing capacity, data processing rate, and other key performance parameters. The response time measurement needs to cover the complete request processing flow, from request reception to result return. The throughput indicator reflects the number of business operations completed by the server per unit time, which needs to be measured according to different business types. The concurrent processing capacity evaluation measures the performance and stability under different concurrent levels by simulating multi-user access scenarios. The data processing rate measurement targets different types of computing tasks to evaluate the processing efficiency and accuracy of the server.
[0110] Step 405, if the deviation of the operating condition resource consumption and / or operating condition performance data exceeds the preset threshold, dynamically update the corresponding environmental impact factor and / or device difference factor in the factor database.
[0111] The setting of the preset threshold is based on historical data statistical analysis, and usually takes the multiple of standard deviation as the judgment standard. For example, when the resource consumption data deviation exceeds 1.5 times the standard deviation of the historical mean, or the performance data deviation exceeds 2 times the standard deviation, the system determines that it is a significant deviation and starts the factor updating process. The threshold setting of different types of data considers the difference in sensitivity. The deviation threshold of CPU and memory resources is relatively strict, and the threshold of disk and network resources is appropriately relaxed.
[0112] The factor updating decision needs to analyze the root cause of the deviation, and distinguish different influencing factors such as environmental changes, device aging, software optimization, etc. The environmental impact factor update mainly aims at the actual impact changes of external environmental factors such as load conditions, network conditions, maintenance strategies, etc. When these factors show different influence degrees under the same conditions, adjust the corresponding values. The device difference factor update is aimed at the performance changes of hardware performance degradation, software version optimization, configuration parameter adjustment and other device-related factors.
[0113] The dynamic updating algorithm uses incremental learning and weighted average method to determine the new factor value through the weighted calculation of the current measurement value and the historical factor value. The weight distribution is based on factors such as data reliability, time freshness, sample size, etc. The system establishes an update verification mechanism to evaluate the prediction accuracy of the updated factors through cross-validation. If the error increases after updating, it will be rolled back to the previous version and the update strategy will be re-evaluated to ensure the scientificity and effectiveness of the factor adjustment.
[0114] Through the above technical scheme, the beneficial effects of the present embodiment are: through the dynamic factor updating mechanism, the self-adaptive optimization of the health state detection system is realized, which can automatically identify the changes of running environment and device performance and adjust the evaluation parameters in time to ensure the long-term detection accuracy. The present embodiment effectively solves the problem of fixed parameters of traditional static model, and makes the system adapt to the influence of dynamic factors such as hardware aging, software upgrading and environmental changes.
[0115] As shown in Figure 5 , it is a specific embodiment of a server health state detection device of the present application. The server health state detection device of the present embodiment is an entity device for executing the server health state detection method of Figures 1 to 4 . Its technical scheme is essentially consistent with the above-mentioned embodiment, and the corresponding description in the above-mentioned embodiment is also applicable to the present embodiment. The server health state detection device of the present embodiment comprises:
[0116] The health benchmark value determination module 501 is configured to monitor the target server in a standard environment to obtain a health benchmark value corresponding to the target server; the standard environment includes a benchmark hardware configuration, a standard software environment, a constant load scenario, standard network conditions and a conventional maintenance strategy;
[0117] The factor determination module 502 is configured to match corresponding environmental impact factors and equipment difference factors from a pre-constructed factor database according to real-time parameters of the target server;
[0118] The health state index determination module 503 is configured to continuously multiply the health benchmark value, the environmental impact factors and the equipment difference factors to determine a health state index corresponding to the target server.
[0119] Figure 6 FIG. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. At the hardware level, the electronic device includes a processor, and optionally further includes an internal bus, a network interface and a memory. The memory can include a memory such as a random-access memory (RAM), and can further include a non-volatile memory such as at least one disk memory. Of course, the electronic device can further include other hardware required by a business.
[0120] The processor, the network interface and the memory can be connected to each other through the internal bus, which can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus and a control bus, etc. For the convenience of representation, Figure 6 Only one bidirectional arrow is used in the figure to represent the bus, but it does not mean that there is only one bus or only one type of bus.
[0121] The memory is used to store execution instructions. Specifically, the execution instructions are computer programs that can be executed. The memory can include a memory and a non-volatile memory, and provides the processor with execution instructions and data.
[0122] In one possible implementation, the processor reads the corresponding execution instructions from non-volatile memory into main memory and then executes them. Alternatively, it may obtain the corresponding execution instructions from other devices to form a server health status detection device at the logical level. The processor executes the execution instructions stored in the memory to implement the server health status detection method provided in any embodiment of this application through the executed instructions.
[0123] The above is as stated in this application. Figure 5 The method executed by the server health status detection device provided in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor.
[0124] The steps of the method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0125] This application also proposes a readable medium storing executable instructions. When these instructions are executed by a processor of an electronic device, the electronic device can perform a server health status detection method provided in any embodiment of this application, specifically for executing, for example... Figure 1 or Figure 2 or Figure 3 or Figure 4 The method shown.
[0126] The electronic device in each of the foregoing embodiments can be a computer.
[0127] Those skilled in the art should understand that the embodiments of the present application can be provided as a method or a computer program product. Therefore, the present application can take a form of a complete hardware embodiment, a complete software embodiment, or a combination of software and hardware.
[0128] Each of the embodiments in the present application is described in a progressive manner, and the same or similar parts between each of the embodiments can be referred to each other. Each of the embodiments mainly explains the difference from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, they are described more simply, and the relevant parts can be referred to the part of the method embodiments.
[0129] It should also be noted that the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article, or apparatus that comprises a list of elements does not include only those elements, but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element defined by the statement "comprising a" does not exclude the existence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0130] The above is only an embodiment of the present application, and is not intended to limit the present application. Those skilled in the art can make various changes and modifications to the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the scope of the claims of the present application.
Claims
1. A method for detecting a health state of a server, the method comprising: The method comprises: monitoring a target server under a standard environment to obtain a health benchmark value corresponding to the target server; the standard environment comprises a benchmark hardware configuration, a standard software environment, a constant load scenario, standard network conditions and a regular maintenance strategy; according to real-time parameters of the target server, matching corresponding environmental impact factors and device difference factors from a pre-constructed factor database; continuously multiplying the health benchmark value, the environmental impact factors and the device difference factors to determine a health status index corresponding to the target server.
2. The method of claim 1, wherein, The factor database further comprises: monitoring the target server state changes to determine the environmental impact factors by individually changing any of the constant load scenario, the standard network conditions or the regular maintenance strategy while keeping other conditions consistent with the standard environment; monitoring the target server state changes to determine the device difference factors by individually changing any of the benchmark hardware configuration or the standard software environment while keeping other conditions consistent with the standard environment; integrating the environmental impact factors and the device difference factors to generate the factor database.
3. The method of claim 2, wherein, The method further comprises: obtaining a total amount of resource consumption and a total amount of performance output of the target server under the standard environment; calculating a ratio of the total amount of resource consumption and the total amount of performance output, and taking the ratio as the health benchmark value.
4. The method of claim 3, wherein, The real-time parameters of the target server comprise: determining hardware configuration parameters corresponding to the target server based on a CPU model, memory capacity and storage type corresponding to the target server; determining software version parameters corresponding to the target server based on an operating system version and an application program version corresponding to the target server; determining environmental dynamic parameters corresponding to the target server based on a real-time load ratio, a network packet loss rate and a maintenance strategy execution period corresponding to the target server; determining the implementation parameters based on the hardware configuration parameters, the software version parameters and the environmental dynamic parameters.
5. The method of claim 4, wherein, The method further comprises: calling corresponding device difference factors from the factor database according to the hardware configuration parameters and the software version parameters; calling corresponding environmental impact factors from the factor database according to the environmental dynamic parameters.
6. The method of claim 5, wherein, The method further comprises: when the target server is running under the same working conditions as the standard environment, collecting working condition resource consumption and working condition performance data corresponding to the target server; if a deviation of the working condition resource consumption and / or the working condition performance data exceeds a preset threshold, dynamically updating corresponding environmental impact factors and / or device difference factors in the factor database.
7. The method of claim 1, wherein, The method further comprises: determining health classification threshold values corresponding to the target server; based on the health classification threshold values and the health status index, performing a classification warning on the target server.
8. A server health state detection apparatus characterized by comprising: The method comprises: a health benchmark value determination module configured to monitor a target server under a standard environment to obtain a health benchmark value corresponding to the target server; The standard environment includes a benchmark hardware configuration, a standard software environment, a constant load scenario, standard network conditions, and a routine maintenance strategy; The factor determination module is configured to match corresponding environmental impact factors and device difference factors from a pre-constructed factor database according to real-time parameters of the target server; The health state index determination module is configured to continuously multiply the health benchmark value, the environmental impact factors, and the device difference factors to determine a health state index corresponding to the target server.
9. A computer readable storage medium, the storage medium having stored thereon a computer program, characterized in that, The computer program is used to execute the server health state detection method in any one of claims 1-7.
10. An electronic device, comprising: The electronic device comprises: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the server health state detection method in any one of claims 1-7.