Operation system health assessment method and system based on multi-dimensional index dynamic weighting
Through the dynamic weighting method of multi-dimensional indicators, combined with hardware and software status information, the weight parameters are dynamically adjusted, and the problem of insufficient evaluation capabilities in traditional monitoring systems is solved, and the panoramic, real-time evaluation and dynamic adaptation of the operating system health status is realized, and real-time warning and automatic notification functions are provided.
Patent Information
- Application Number
- CN202510501209.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-22
AI Technical Summary
There are limitations of a single index threshold alarm mechanism in the existing computer system health monitoring technology, and it is difficult to build a correlation analysis model of multi-dimensional indicators, resulting in insufficient comprehensive evaluation capabilities and the static weight allocation strategy cannot adapt to environmental changes in dynamic operating scenarios.
The operating system health evaluation method based on dynamic weighting based on multi-dimensional indicators is adopted. By collecting the status information of hardware modules and software modules, adjusting weight parameters in real-time load characteristics and business priority, dynamically calculate the hardware and software health scores, and finally determine the health status of the operating system.
It realizes panoramic and real-time assessment of the health status of the operating system, improves the comprehensive assessment capabilities, can dynamically adapt to changes in the actual operating environment, ensures that the evaluation of key performance indicators matches the environment, and provides real-time early warning and automatic notification functions.
Smart Images

Figure CN120353679A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer system status monitoring, and specifically relates to an operating system health assessment method and system based on dynamic weighting of multi-dimensional indicators. Background Art
[0002] Computer system health monitoring is a key means to ensure the reliable operation of computer systems. Existing technical architectures face multiple technical bottlenecks in the field of computer system health monitoring: First, the single indicator threshold alarm mechanism (such as CPU over-limit alarm) commonly used by traditional monitoring systems has significant limitations. Its working mode based on discrete parameter monitoring makes it difficult to build a correlation analysis model for multi-dimensional indicators, resulting in a lack of comprehensive evaluation capabilities for the health status of the system. Secondly, in dynamic operating scenarios (such as drastic fluctuations in server load or frequent switching of business types), static weight allocation strategies will lead to inaccurate importance assessments of key performance indicators (memory usage, disk IO throughput, etc.), and will not be able to dynamically adapt to the changing patterns of the actual operating environment. Summary of the invention
[0003] Technical problem to be solved by the present invention: In view of the above-mentioned problems in the prior art, a method and system for operating system health assessment based on dynamic weighting of multi-dimensional indicators are provided. The present invention aims to solve the systematic defects existing in the existing health monitoring technology system, improve the comprehensive assessment capability of operating system health assessment, and dynamically adapt to the changing laws of the actual operating environment.
[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0005] A method for evaluating the health of an operating system based on dynamic weighting of multi-dimensional indicators comprises the following steps: collecting status information of each hardware module in the operating system, calculating multiple hardware status indicators and weighted summing them to obtain scores of the hardware modules, wherein the hardware modules include part or all of a CPU, a memory, a disk, a network card and a graphics card, the hardware status indicators of the CPU include a CPU usage rate and a CPU temperature, the hardware status indicators of the memory include a memory usage rate and a memory fragmentation rate, the hardware status indicators of the disk include a disk usage rate and a disk damage rate, the hardware status indicators of the network card include a network card bandwidth usage rate and a packet loss rate, the hardware status indicators of the graphics card include a GPU temperature and a GPU utilization rate of the graphics card, adjusting the weight parameters of each hardware module in combination with real-time load characteristics or business priorities, and weighted summing the scores of each hardware module to obtain a hardware health score; collecting software status information in the operating system and respectively calculating the scores of each software module, wherein the software module includes a key component, a key service and a key process, and weighted summing the scores of each software module to obtain a software health score; and weighted summing the hardware health score and the software health score to obtain a health score of the operating system to determine the health status of the operating system.
[0006] Optionally, adjusting the weight parameters of each hardware module in combination with real-time load characteristics includes: calculating the real-time load characteristics of each hardware module respectively, and using the normalized real-time load characteristics of each hardware module as the weight parameters of each hardware module; adjusting the weight parameters of each hardware module in combination with service priorities includes: obtaining the current service scenario in the operating system, querying the pre-reviewed service scenario-hardware module priority mapping table according to the current service scenario to determine the service priorities of each hardware module, and converting the service priorities of each hardware module into the weight parameters of each hardware module.
[0007] Optionally, when calculating multiple hardware status indicators and summing them up with weights to obtain the score of a hardware module, it includes determining whether any of the hardware status indicators has an instantaneous anomaly. If a certain hardware status indicator has an instantaneous anomaly, first reduce the weight parameter of this hardware status indicator and then sum them up with weights to obtain the score of the corresponding hardware module.
[0008] Optionally, the default weight parameters of both the hardware health score and the software health score are 50%. When summing up the hardware health score and the software health score with weights to obtain the health score of the operating system, it also includes obtaining the current disk usage rate of the operating system and the installation quantity or ratio of non-system official software packages. When the disk usage rate exceeds the first preset threshold, increase the weight parameter of the hardware health score and decrease the weight parameter of the software health score on the basis of the default weight parameters; and when the disk usage rate does not exceed the first preset threshold and the installation quantity or ratio of non-system official software packages exceeds the second preset threshold, decrease the weight parameter of the hardware health score and increase the weight parameter of the software health score on the basis of the default weight parameters.
[0009] Optionally, when the operating system is a Linux operating system, when collecting the status information of each hardware module in the operating system, calculating multiple hardware status indicators and summing them up with weights to obtain the score of a hardware module:
[0010] The steps for obtaining the score of the CPU include: opening the file " / proc / stat" to read the first status of the CPU, CPUstat1; delaying for a specified time; opening the file " / proc / stat" to read the second status of the CPU, CPUstat2; calculating the total difference and the idle difference between the first status CPUstat1 and the second status CPUstat2, and dividing the difference between the total difference and the idle difference by the total difference to obtain the CPU usage rate; opening the file " / sys / class / thermal / thermal_zone0 / temp" to read the original CPU temperature, and converting the original CPU temperature from millidegrees Celsius to degrees Celsius to obtain the CPU temperature; summing up the CPU usage rate and the CPU temperature with weights to obtain the score of the CPU;
[0011] The steps to obtain the score of the memory include: opening the file " / proc / meminfo" to read the total physical memory MemTotal, free memory size MemFree, kernel buffer size Buffers, and page cache size Cached of the memory; subtracting the free memory size MemFree, kernel buffer size Buffers, and page cache size Cached from the total physical memory MemTotal in sequence to obtain the used memory size; dividing the used memory size by the total physical memory MemTotal to obtain the memory utilization rate, reading the file " / proc / buddyinfo" and calculating the memory fragmentation rate, and performing a weighted sum of the memory utilization rate and the memory fragmentation rate to obtain the score of the memory; the step of reading the file " / proc / buddyinfo" and calculating the memory fragmentation rate includes: S101, opening the file " / proc / buddyinfo" and initializing variables total_free and small_free to 0; S102, processing each line in the file " / proc / buddyinfo": parsing out the memory area name and the number of free memory blocks count for each order; S103, for each order, calculate the free memory block size block_size according to block_size = 2 order *MIN_UNIT, where order is the index number of the order and MIN_UNIT is the minimum block size; S104, for each order, calculate the free space memory size order_free of this order according to order_free = count * block_size, where count is the number of free memory blocks of this order; S105, add the free space memory size order_free of each order to the variable total_free; S106, accumulate the free space memory size order_free of the first 3 orders to the variable small_free; S107, divide the variable small_free by the variable total_free to obtain the memory fragmentation rate;
[0012] The steps to obtain the score of the disk include: calling the statvfs function to obtain the mount point data of the disk, and calculating the total capacity, free capacity, and used capacity of the disk according to the mount point data of the disk; dividing the used capacity by the total capacity to obtain the disk utilization rate; using popen to call the smartctl command to obtain the SMART data of the disk, extracting the reallocated sector count, number of uncorrectable errors, command timeout count, and total sector count from the SMART data, summing the reallocated sector count, number of uncorrectable errors, and command timeout count and then dividing by the total sector count to obtain the disk damage rate, and performing a weighted sum of the disk utilization rate and the disk damage rate to obtain the score of the disk;
[0013] The steps for obtaining the score of the network card include: opening the file " / proc / net / dev", parsing the initial sampling data to record the received and transmitted bytes and error counts after skipping the first two lines of headers; after a delay, opening the file " / proc / net / dev" again, parsing the re-sampled data to record the received and transmitted bytes and error counts, calculating the traffic during this period based on the data of the two samplings, dividing the traffic during this period by the length of this period and then dividing by the bandwidth of the network card to obtain the bandwidth utilization rate of the network card; calculating the total number of bytes transmitted and the number of bytes of error counts during this period based on the data of the two samplings, dividing the number of bytes of error counts during this period by the total number of bytes transmitted during this period to obtain the packet loss rate of the network card, and performing a weighted sum of the bandwidth utilization rate and the packet loss rate of the network card to obtain the score of the network card;
[0014] The steps for obtaining the score of the graphics card include: obtaining the GPU temperature; initializing the NVML library, obtaining the GPU handle, calling the NVML library to obtain the GPU utilization rate, and closing the NVML library; performing a weighted sum of the GPU temperature and the GPU utilization rate to obtain the score of the graphics card.
[0015] Optionally, when collecting the software status information in the operating system and calculating the scores of each software module respectively, the steps for obtaining the score of the key component include: for the specified key components in the operating system, obtaining the installation status and whether it is an official version of each key component respectively, and calculating the score of the key component according to the following formula:
[0016]
[0017] where S comp is the score of the key component, M1 is the number of key components, R i is the installation status of the key component i, and the installation status value is 1 or 0. A value of 1 indicates installed, and a value of 0 indicates not installed; V i is whether the key component i is an official version, and the value is 1 indicating it is an official version, and the value is 0 indicating it is not an official version;
[0018] The steps for obtaining the score of the key service include: for the specified key services in the operating system, obtaining the start status of each key service respectively, and calculating the score of the key service according to the following formula:
[0019]
[0020] where S serv is the score of the key service, M2 is the number of key services, l j is the start status of the key service j, and the value is 1 indicating it has been started, and the value is 0 indicating it has not been started;
[0021] The steps to obtain the scores of critical processes include: for the specified critical processes in the operating system, obtain the startup status, CPU usage rate, and memory usage rate of each critical process respectively, and calculate the scores of the critical processes according to the following formula:
[0022]
[0023] Among them, S ps is the score of the critical process, N3 is the number of critical processes, P k is the startup status of critical process k, with a value of 1 indicating started and a value of 0 indicating not started; P k,cpu + is the CPU usage rate of critical process k, P k,men is the memory usage rate of critical process k.
[0024] Optionally, the function expression for weighted summation of CPU usage rate and CPU temperature to obtain the score of the CPU is:
[0025] S cpu = W1·S cpuT + W2·S cpuU ,
[0026] Among them, S cpu is the score of the CPU, W1 and W2 are weight parameters, S cpuT is the CPU temperature, S cpuU is the CPU usage rate; the function expression for weighted summation of memory usage rate and memory fragmentation rate to obtain the score of the memory is:
[0027] S memory = W3·S memoryF + W4·S memoryU ,
[0028] Among them, S memory is the score of the memory, W3 and W4 are weight parameters, S cpuT is the memory fragmentation rate, S memoryU is the memory usage rate; the function expression for weighted summation of disk usage rate and disk damage rate to obtain the score of the disk is:
[0029] S disk = W5·S diskBD + W6·S diskU ,
[0030] Among them, S disk is the score of the disk, W5 and W6 are weight parameters, S diskBD is the disk damage rate, S diskU is the disk usage rate; the function expression for weighted summation of the bandwidth usage rate and packet loss rate of the network card to obtain the score of the network card is:
[0031] S network = W7·S networkU + W8·S networkL ,
[0032] where S network is the score of the network card, W7 and W8 are weight parameters, S networkU is the bandwidth utilization rate of the network card, S networkL is the packet loss rate; the function expression for obtaining the score of the graphics card by weighted summation of the GPU temperature and GPU utilization is:
[0033] S gpu = W9·S gpuT + W 10 ·S gpuU ,
[0034] where S gpu is the score of the graphics card, W9 and W 10 are weight parameters, S gpuT is the CPU temperature, S gpuU is the CPU utilization rate; the function expression for obtaining the hardware health score by weighted summation of the scores of each hardware module is:
[0035] S hardware = W 11 ·S cpu + W 12 ·S memory + W 13 ·S disk + W 14 ·S network + W 15 ·S gpu ,
[0036] where S hardware is the hardware health score, W 11 ~ W 15 are weight parameters, S cpu is the score of the CPU, S memory is the score of the memory, S disk is the score of the disk, S network is the score of the network card, S gpu is the score of the graphics card; the function expression for obtaining the software health score by weighted summation of the scores of each software module is:
[0037] S software = W 16 ·S comp + W 17 ·S serv + W 18 ·S ps ,
[0038] Among them, S software is the software health score, W 16 ~W 18 are weight parameters, S comp is the score of key components, S serv is the score of key services, S ps is the score of key processes; the function expression for obtaining the health score of the operating system by weighted summation of the hardware health score and the software health score is:
[0039] S core =W 19 ·S software +W 20 ·S hardware ,
[0040] In the above formula, S core is the health score of the operating system, W 19 and W 20 are weight parameters.
[0041] In addition, the present invention also provides an operating system health assessment system based on dynamic weighting of multi-dimensional indicators, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the operating system health assessment method based on dynamic weighting of multi-dimensional indicators.
[0042] In addition, the present invention also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the operating system health assessment method based on dynamic weighting of multi-dimensional indicators through a processor.
[0043] In addition, the present invention also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the operating system health assessment method based on dynamic weighting of multi-dimensional indicators through a processor.
[0044] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: aiming at the problem that the traditional discrete threshold warning mechanism is difficult to reveal the coupling relationship between multi-dimensional indicators, the present invention realizes a three-dimensional assessment of the system health state by constructing a correlation analysis model integrating multi-source parameters such as CPU load, memory leak rate, and disk IO exception; breaks through the adaptability bottleneck of the static weight allocation strategy in the dynamic business scenario, and establishes a weight self-regulation mechanism based on load fluctuation characteristics and business priorities to ensure that the evaluation weights of core indicators such as memory occupancy rate and network throughput match the changes in the operating environment in real time, and can solve the systematic defects existing in the existing health monitoring technology system, and improve the comprehensive evaluation ability of the operating system health assessment and the dynamic adaptation to the change law of the actual operating environment. Brief Description of the Drawings
[0045] Figure 1 This is a schematic diagram of the basic process of the method according to the embodiments of the present invention.
[0046] Figure 2 This is a schematic diagram of the working process of the system monitoring module in the embodiments of the present invention.
[0047] Figure 3 This is a schematic diagram of the process of CPU data collection in the embodiments of the present invention.
[0048] Figure 4 This is a schematic diagram of the process of memory data collection in the embodiments of the present invention.
[0049] Figure 5 This is a schematic diagram of the process of disk data collection in the embodiments of the present invention.
[0050] Figure 6 This is a schematic diagram of the process of network card data collection in the embodiments of the present invention.
[0051] Figure 7 This is a schematic diagram of the process of graphics card data collection in the embodiments of the present invention.
[0052] Figure 8 This is a schematic diagram of the working process of the software status collection module in the embodiments of the present invention.
[0053] Figure 9 This is a schematic diagram of the process of collecting key component data in the embodiments of the present invention.
[0054] Figure 10 This is a schematic diagram of the process of collecting key service data in the embodiments of the present invention.
[0055] Figure 11 This is a schematic diagram of the process of collecting the startup status data of key processes in the embodiments of the present invention.
[0056] Figure 12 This is a schematic diagram of the process of collecting the CPU usage rate and memory usage rate of key processes in the embodiments of the present invention. Detailed Embodiments
[0057] The core idea of the present invention is to collect hardware and software data in the Linux operating system, analyze it through a dynamic weighted evaluation algorithm based on the index data, and then obtain the health status of the current operating system. Furthermore, the results can be sent to the relevant responsible persons in the form of an email as needed. To enable those skilled in the art to better understand the technical solutions of the present invention, the following will further elaborate on the technical solutions of the present invention in conjunction with the drawings in the embodiments of the present invention.
[0058] As Figure 1As shown in the figure, the operating system health assessment method based on multi-dimensional index dynamic weighting in this embodiment includes the following steps: collecting the status information of each hardware module in the operating system, calculating multiple hardware status indicators and weighted summing them to obtain the score of the hardware module. The hardware modules include some or all of the CPU, memory, disk, network card, and graphics card. The hardware status indicators of the CPU include CPU usage rate and CPU temperature. The hardware status indicators of the memory include memory usage rate and memory fragmentation rate. The hardware status indicators of the disk include disk usage rate and disk damage rate. The hardware status indicators of the network card include network card bandwidth usage rate and packet loss rate. The hardware status indicators of the graphics card include GPU temperature and GPU utilization rate of the graphics card. Combining real-time load characteristics or service priorities to adjust the weight parameters of each hardware module, and weighted summing the scores of each hardware module to obtain the hardware health score; collecting the software status information in the operating system and calculating the scores of each software module respectively. The software modules include key components, key services, and key processes, and weighted summing the scores of each software module to obtain the software health score; weighted summing the hardware health score and the software health score to obtain the health score of the operating system to determine the health status of the operating system.
[0059] A key point of the operating system health assessment method based on multi-dimensional index dynamic weighting in this embodiment lies in the implementation of the dynamic weighting assessment algorithm. It mainly uses the data of the data acquisition module and calculates the scores of each item through the corresponding algorithm. Among them, the weight allocation of each item will be dynamically adjusted as the system runs. After analyzing through the multi-index scoring + weighted calculation algorithm, the current system health score value is given, which includes the following aspects:
[0060] 1) Automatically adjust the hardware index weights based on real-time load characteristics (such as CPU load standard deviation, network traffic mutation rate): Combining real-time load characteristics to adjust the weight parameters of each hardware module includes: calculating the real-time load characteristics of each hardware module respectively, and normalizing the real-time load characteristics of each hardware module as the weight parameters of each hardware module;
[0061] 2) Combine service priority configuration (such as increasing the disk IO weight in the database service scenario) to dynamically reconstruct the weight coefficient in the scoring formula: Combining service priorities to adjust the weight parameters of each hardware module includes: obtaining the current service scenario in the operating system, querying the pre-reviewed service scenario - hardware module priority mapping table according to the current service scenario to determine the service priorities of each hardware module, and converting the service priorities of each hardware module into the weight parameters of each hardware module.
[0062] 3) Abnormal propagation suppression mechanism: When calculating the score of a hardware module by calculating multiple hardware status indicators and performing weighted summation, it includes determining whether each hardware status indicator (such as the network card packet loss rate) has an instantaneous abnormality. If a certain hardware status indicator has an instantaneous abnormality, first reduce the weight parameter of this hardware status indicator and then perform weighted summation to obtain the score of the corresponding hardware module. Among them, the instantaneous abnormality can use the instantaneous value exceeding the preset threshold as the judgment condition. In addition, other more complex outlier judgment algorithms or machine learning classifiers can also be used for outlier judgment methods according to needs. Specifically, it can be selected according to actual needs, which is not within the scope of the improvement of the present invention and will not be elaborated here.
[0063] 4) The default weight parameters of both the hardware health score and the software health score are 50%. When calculating the health score of the operating system by performing weighted summation of the hardware health score and the software health score, it also includes obtaining the current hard disk usage rate of the operating system and the installation quantity or proportion of non-system official software packages. When the hard disk usage rate exceeds the first preset threshold, increase the weight parameter of the hardware health score and decrease the weight parameter of the software health score on the basis of the default weight parameter. And when the hard disk usage rate does not exceed the first preset threshold and the installation quantity or proportion of non-system official software packages exceeds the second preset threshold, decrease the weight parameter of the hardware health score and increase the weight parameter of the software health score on the basis of the default weight parameter.
[0064] Based on the implementation of the above dynamic weighted evaluation algorithm, the operating system health assessment method based on multi-dimensional index dynamic weighting in this embodiment can effectively solve the problem that the existing solution using a fixed weight table (such as CPU: 0.3, memory: 0.2) cannot adapt to the elastic changes in resource requirements in the cloud computing scenario, and can dynamically adapt to the change law of the actual operating environment.
[0065] In this embodiment, a system monitoring module for constructing the hardware status is used to collect the status information of each hardware module for calculating multiple hardware status indicators. Figure 2This is a schematic diagram of the working process of the system monitoring module in this embodiment. After initializing the system monitoring module, the following operations are sequentially performed: collecting CPU data, collecting memory data, collecting hard disk data, collecting network card data, and collecting graphics card data, which specifically include: a. collecting temperature and utilization rate data in the system CPU; b. collecting utilization rate and fragmentation rate data in the system memory; c. collecting damage rate and utilization rate data in the system hard disk; d. collecting network bandwidth utilization rate and packet loss rate data in the system network card; e. collecting utilization rate data in the system graphics card. Then, hardware status indicators are calculated based on the collected data: the hardware status indicators of the CPU include CPU utilization rate and CPU temperature, the hardware status indicators of the memory include memory utilization rate and memory fragmentation rate, the hardware status indicators of the disk include disk utilization rate and disk damage rate, and the hardware status indicators of the network card include network card bandwidth utilization rate and packet loss rate.
[0066] In this embodiment, the operating system is the Linux operating system. When collecting the status information of each hardware module in the operating system, calculating multiple hardware status indicators, and performing weighted summation to obtain the score of the hardware module:
[0067] As Figure 3 shown, the steps to obtain the score of the CPU include: opening the file " / proc / stat" to read the first status of the CPU, CPUstat1; delaying for a specified time; opening the file " / proc / stat" to read the second status of the CPU, CPUstat2; calculating the total difference and idle difference between the first status CPUstat1 and the second status CPUstat2, and dividing the difference between the total difference and the idle difference by the total difference to obtain the CPU utilization rate; opening the file " / sys / class / thermal / thermal_zone0 / temp" to read the original CPU temperature, and converting the original CPU temperature from millidegrees Celsius to degrees Celsius to obtain the CPU temperature; performing weighted summation on the CPU utilization rate and the CPU temperature to obtain the score of the CPU;
[0068] As Figure 4As shown in the figure, the steps to obtain the score of the memory include: opening the file " / proc / meminfo" to read the total physical memory MemTotal, the size of free memory MemFree, the size of the kernel buffer Buffers, and the size of the page cache Cached; subtracting the size of free memory MemFree, the size of the kernel buffer Buffers, and the size of the page cache Cached from the total physical memory MemTotal in sequence to obtain the size of used memory; dividing the size of used memory by the total physical memory MemTotal to obtain the memory utilization rate, reading the file " / proc / buddyinfo" and calculating the memory fragmentation rate, and performing a weighted sum of the memory utilization rate and the memory fragmentation rate to obtain the score of the memory; the step of reading the file " / proc / buddyinfo" and calculating the memory fragmentation rate includes: S101, opening the file " / proc / buddyinfo" and initializing the variables total_free and small_free to 0; S102, processing each line in the file " / proc / buddyinfo": parsing out the memory area name and the number of free memory blocks count for each order; S103, for each order, calculating the size of the free memory block block_size according to block_size = 2 order *MIN_UNIT, where order is the index number of the order and MIN_UNIT is the minimum block size; S104, for each order, calculating the free memory size order_free of this order according to order_free = count * block_size, where count is the number of free memory blocks of this order; S105, adding the free memory size order_free of each order to the variable total_free; S106, accumulating the free memory size order_free of the first 3 orders to the variable small_free; S107, dividing the variable small_free by the variable total_free to obtain the memory fragmentation rate;
[0069] As Figure 5 shown in the figure, the steps to obtain the score of the disk include: calling the statvfs function to obtain the mount point data of the disk, and calculating the total capacity, free capacity, and used capacity of the disk according to the mount point data of the disk; dividing the used capacity by the total capacity to obtain the disk utilization rate; using popen to call the smartctl command to obtain the SMART data of the disk, extracting the reallocated sector count, the number of uncorrectable errors, the command timeout count, and the total number of sectors from the SMART data, summing up the reallocated sector count, the number of uncorrectable errors, and the command timeout count and then dividing by the total number of sectors to obtain the disk damage rate, and performing a weighted sum of the disk utilization rate and the disk damage rate to obtain the score of the disk;
[0070] As Figure 6As shown, the steps to obtain the score of the network card include: opening the file " / proc / net / dev", parsing the initial sampling data after skipping the first two lines of headers to record the received and transmitted bytes and error counts; after a delay, opening the file " / proc / net / dev" again, parsing the re-sampled data after skipping the first two lines of headers to record the received and transmitted bytes and error counts, calculating the traffic during this period based on the data of the two samplings, dividing the traffic during this period by the length of this period and then dividing by the bandwidth of the network card to obtain the bandwidth utilization rate of the network card; calculating the total number of bytes transmitted and the number of bytes of error counts during this period based on the data of the two samplings, dividing the number of bytes of error counts during this period by the total number of bytes transmitted during this period to obtain the packet loss rate of the network card, and performing a weighted sum of the bandwidth utilization rate and the packet loss rate of the network card to obtain the score of the network card;
[0071] As Figure 7 shown, the steps to obtain the score of the graphics card include: obtaining the GPU temperature; initializing the NVML library, obtaining the GPU handle, calling the NVML library to obtain the GPU utilization rate, and closing the NVML library; performing a weighted sum of the GPU temperature and the GPU utilization rate to obtain the score of the graphics card. This process first checks whether the NVML library is available. If available, it initializes NVML, obtains the GPU handle, and calls to obtain the utilization rate, and finally outputs the result; otherwise, it outputs a prompt message indicating that GPU data is not available.
[0072] In this embodiment, a software status collection module is constructed to collect the status information of each software module for calculating the scores of key components. Figure 8 FIG. is the schematic diagram of the working process of the software status collection module in this embodiment, and its working steps are as follows: a. Collecting the installation and manufacturer identification data of the main software packages in the system; b. Collecting the startup status data of the key services in the system; c. Collecting the startup status data of the key processes in the system; d. Collecting the CPU utilization rate and memory occupancy data of the key processes in the system. In this embodiment, when collecting the software status information in the operating system and calculating the scores of each software module respectively, the steps to obtain the scores of key components include: for the specified key components in the operating system, obtaining the installation status and whether it is an official version of each key component respectively, and calculating the score of the key component according to the following formula:
[0073]
[0074] where, S comp is the score of the key component, N1 is the number of key components, R i is the installation status of the key component i, and the installation status takes values of 1 or 0. A value of 1 indicates installed, and a value of 0 indicates not installed; V iWhether the key component i is the official version. The value of 1 indicates the official version, and the value of 0 indicates the non - official version. For example, Figure 9 As shown, obtaining the installation status and whether it is the official version of each key component includes: constructing the command "rpm - qa -- qf '%{NAME} %{VENDOR}'", executing the command and reading the output, parsing the output data, and outputting the software package and manufacturer information. Among them, '%{NAME}' and '%{VENDOR}' are the software package and manufacturer information respectively.
[0075] The steps for obtaining the score of the key service include: for the specified key service in the operating system, obtaining the start status of each key service respectively, and calculating the score of the key service according to the following formula:
[0076]
[0077] Among them, S serv is the score of the key service, N2 is the number of key services, L j is the start status of the key service j. The value of 1 indicates started, and the value of 0 indicates not started. For example, Figure 10 As shown, obtaining the start status of each key service includes: inputting the key service name <service>, construction command "systemctl is-active <service>Execute the command, read and parse the output status, output the service status (active / inactive), and output the standardized service status identifier after parsing the returned text status (active / inactive) to achieve the automated collection and structured feedback of the running status of key services.
[0078] The steps for obtaining the scores of key processes include: for the specified key processes in the operating system, obtain the startup status, CPU usage rate, and memory usage rate of each key process respectively, and calculate the score of the key process according to the following formula:
[0079]
[0080] where S ps is the score of the key process, N3 is the number of key processes, and P k is the startup status of key process k, with a value of 1 indicating started and a value of 0 indicating not started; P k,cpu + is the CPU usage rate of key process k, and P k,men is the memory usage rate of key process k. As Figure 11 shown, obtaining the startup status of each key process includes: inputting the key process name <process>, construct the command "pidof <process>", execute the command. An empty output indicates that the process is not running, and a non-empty output indicates that the process is running, thus outputting the process status. As Figure 12 shown, obtaining the CPU usage rate and memory usage rate of the key process includes: inputting the key process name and constructing the command "ps -C <process>Execute the command "-o%cpu,%mem--no-headers", parse the output data, obtain the CPU usage rate and memory occupancy data, and output the collection results.
[0081] It can be seen that in order to solve the problem that traditional systems use independent hardware monitoring and software log analysis tools (such as Nagios+Logstash), where data sources are fragmented and there is a lack of coupling analysis ability between metrics, the method of this embodiment realizes the correlation capture of hardware anomalies and software failures through a unified collection protocol (such as synchronous detection of soaring memory fragmentation rate and process memory leakage). It utilizes: Hardware-software collaborative collection: For the first time, deeply integrate the collection modules of hardware status (CPU temperature, memory fragmentation rate, hard disk damage rate, etc.) and software status (versions of key components, service start and stop, process resource occupancy) to form a complementary data collection network. Composite metric definition mechanism: Aiming at the defect that traditional monitoring systems only collect basic metrics (such as CPU usage rate), a number of composite metrics are proposed: Memory fragmentation rate (calculate the distribution of free block orders by parsing / proc / buddyinfo); Hard disk damage rate; Software component health (comprehensively verify the installation status and official version). It can solve the problem of fragmented data sources and lack of coupling analysis ability between metrics, and realize comprehensive and comprehensive system health status detection.
[0082] In this embodiment, the function expression for obtaining the score of the CPU by weighted summation of the CPU usage rate and CPU temperature is:
[0083] S cpu =W1·S cpuT +W2·S cpuU ,
[0084] where S cpu is the score of the CPU, W1 and W2 are weight parameters, S cpuT is the CPU temperature, and S cpuU is the CPU usage rate; the sum of the weight parameters W1 and W2 is 1. In this embodiment, the initial weights are 0.1 and 0.9 respectively and can be adjusted as needed. The function expression for obtaining the score of the memory by weighted summation of the memory usage rate and memory fragmentation rate is:
[0085] S memory =W3·S memoryF +W4·S memoryU ,
[0086] where S memory is the score of the memory, W3 and W4 are weight parameters, S cpuT is the memory fragmentation rate, and S memoryU is the memory usage rate; the sum of the weight parameters W3 and W4 is 1. In this embodiment, the initial weights are 0.1 and 0.9 respectively, which can be adjusted as needed. The functional expression for obtaining the disk score by weighted summation of the disk usage rate and the disk damage rate is:
[0087] S disk = W5·S diskBD + W6·S diskU ,
[0088] where S disk is the disk score, W5 and W6 are weight parameters, S diskBD is the disk damage rate, and S diskU is the disk usage rate; the sum of the weight parameters W5 and W6 is 1. In this embodiment, the initial weights are 0.1 and 0.9 respectively, which can be adjusted as needed. The disk damage rate and the disk usage rate can be expressed as:
[0089] S diskBD = S BadBlock / S TotalBlock , S diskU = S UseBlock / S TotalBlock ,
[0090] where S BadBlock is the number of damaged disk blocks, S UseBlock is the number of used disk blocks, and S TotalBlock is the total number of disk blocks. The functional expression for obtaining the network card score by weighted summation of the network card's bandwidth usage rate and the packet loss rate is:
[0091] S network = W7·S networkU + W8·S networkL ,
[0092] where S network is the network card score, W7 and W8 are weight parameters, S networkU is the network card's bandwidth usage rate, and S networkL is the packet loss rate; the functional expression for obtaining the graphics card score by weighted summation of the GPU temperature and the GPU utilization rate is:
[0093] S gpu = W9·S gpuT + W 10 ·S gpuU ,
[0094] where S gpu is the graphics card score, W9 and W 10 are weight parameters, S gpuT is the CPU temperature, and S gpuU is the CPU utilization rate; the weight parameters W9 and W 10 The sum is 1. In this embodiment, the initial weights are 0.1 and 0.9 respectively, which can be adjusted as needed.
[0095] In this embodiment, the function expression for obtaining the hardware health score by weighted summation of the scores of each hardware module is:
[0096] S hardware = W 11 ·S cpu + W 12 ·S memory + W 13 ·S disk + W 14 ·S network + W 15 ·S gpu ,
[0097] where S hardware is the hardware health score, W 11 ~W 15 are the weight parameters, S cpu is the score of the CPU, S memory is the score of the memory, S disk is the score of the disk, S network is the score of the network card, S gpu is the score of the graphics card; the sum of the weight parameters W 11 ~W 15 is 1. In this embodiment, the initial weights are all 0.2, which can be adjusted as needed.
[0098] In this embodiment, the function expression for obtaining the software health score by weighted summation of the scores of each software module is:
[0099] S software = W 16 ·S comp + W 17 ·S serv + W 18 ·S ps ,
[0100] where S software is the software health score, W 16 ~W 18 are the weight parameters, S comp is the score of the key components, S serv is the score of the key services, S ps is the score of the key processes; the sum of the weight parameters W 16 ~W 18 is 1. In this embodiment, the initial weights are 0.25, 0.25, and 0.5 respectively, which can be adjusted as needed.
[0101] To evaluate the system health status, a method of multi-index scoring + weighted calculation is adopted. Scores are given comprehensively considering key components, key service status, key process status, resources, and hardware status, and early warnings are issued for abnormal situations. In this embodiment, the functional expression for obtaining the health score of the operating system by weighted summation of the hardware health score and the software health score is:
[0102] S core =W 19 ·S software +W 20 ·S hardware ,
[0103] In the above formula, S core is the health score of the operating system, and W 19 and W 20 are weight parameters. The sum of the weight parameters W 19 and W 20 is 1. In this embodiment, the initial weights are both 0.5 and can be dynamically adjusted as the system runs.
[0104] In summary, in view of the problems that traditional health monitoring systems rely on single threshold alarms, are difficult to reflect the coupling relationship of multi-dimensional indicators, and the static weight allocation is not adaptable enough in dynamic environments, this embodiment proposes a Linux operating system health assessment system and method based on dynamic weighting of multi-dimensional indicators. Its main effects and characteristics are reflected in the following aspects: 1. Full-dimensional data collection and fusion. Hardware and software data complementation: The system builds a comprehensive monitoring system by simultaneously collecting hardware status data such as CPU, memory, hard disk, network card, graphics card, and software status data such as key software packages, services, and processes. Revealing the relationship between indicators: Using the fusion analysis of multi-dimensional data, the coupling relationship that is difficult to find in the traditional single indicator monitoring mode is effectively captured, which improves the overall accuracy of health assessment. 2. Dynamic weighted evaluation algorithm. Real-time weight adjustment: In view of the problems of business load fluctuations and changes in key performance indicators in dynamic operation scenarios, the system adopts a dynamic weight self-adjustment mechanism based on load fluctuation characteristics and business priorities, so that the influence of various key indicators (such as memory occupancy, network throughput, etc.) always matches the actual operating environment. Multi-index scoring and comprehensive calculation: By scoring key components, services, processes and hardware status separately and performing weighted calculation, a comprehensive health score is formed to effectively reflect the current operating status of the system and help identify anomalies and potential risks in a timely manner. 3. Real-time warning and automatic notification. Rapid response to abnormal situations: During the evaluation process, various indicators are monitored in real time. Once an abnormal situation is detected, the system can trigger the warning mechanism to ensure that fault or risk information can be captured in the first time. Convenient notification mechanism: The health assessment results and warning information are automatically sent to the relevant person in charge through emails, etc., so that maintenance measures can be taken in time to ensure the long-term stable operation of the system. 4. High compatibility and scalability. Modular design: The system adopts a layered and modular structure, which can be smoothly integrated with the existing monitoring architecture and is convenient for subsequent expansion and optimization to meet the needs of various business scenarios. Dynamic adaptation to changing environments: Whether in scenarios where the server load fluctuates violently or the business type switches frequently, the dynamic weighting mechanism can ensure that the evaluation model is highly consistent with the actual operating status, significantly improving the stability and security of the system. In general, this embodiment achieves a panoramic and real-time evaluation of the health status of the operating system by building a comprehensive monitoring system that integrates hardware and software multi-dimensional data collection and dynamic weighted evaluation. This system not only makes up for the shortcomings of the traditional single indicator alarm mechanism, but also accurately reflects the importance of each key indicator in a dynamic environment, providing strong support for system maintenance and security.
[0105] In addition, this embodiment also provides an operating system health assessment system based on dynamic weighting of multi-dimensional indicators, including an interconnected microprocessor and a memory, wherein the microprocessor is programmed or configured to execute the operating system health assessment method based on dynamic weighting of multi-dimensional indicators.
[0106] As an alternative implementation, in order to achieve comprehensive health monitoring of the Linux operating system through multi-dimensional data collection and dynamic weighted evaluation, the system in this embodiment adopts a modular architecture, which includes functional modules such as log management, configuration management, data collection, data analysis, data storage, notification, and display. Each module inherits from the basic module class and cooperates with each other through mechanisms such as the singleton pattern and message queue to achieve the collection, processing, storage, and display of system data. The system in this embodiment supports two operating modes: graphical user interface (GUI) and command line to adapt to different usage scenarios. The following is the detailed implementation plan of the system in this embodiment: (1) Log management module. 1) Function: Record system operation events, errors, and debugging information to provide support for system maintenance and problem troubleshooting. 2) Implementation method: Adopt a C++ log library (such as spdlog), support configuring the log level (debug, info, error, etc.) and output path (file or console). Initialize the log system when the module starts, and other modules record events by calling the log method. 3) Details: Log files are split by date and support multi-threaded safe writing. (2) Configuration management module. 1) Function: Load and manage configuration files, and provide a list of key components (such as RPM packages, services, processes) to be monitored and weight configurations. 2) Implementation method: Use configuration files in INI or JSON format, parse and load the configuration when the module starts, and provide the get_value interface for other modules to read configuration items (such as monitoring list, weight value). 3) Details: Support dynamic update of configuration files, and the configuration can be reloaded during system operation. (3) Data collection module. 1) Function: Collect software and hardware related data from the system to provide basic data for health assessment. 2) Implementation method: Divided into two parts: software collection and hardware collection: 3) Software collection: a. Key components: Through rpm-qa|grep <package>Query the installation status and version of software packages. b. Key services: Use systemctl is-active <service>Check the service running status. c. Key processes: via pidof <process>Check if the process exists, using ps - C <process>- o% cpu, % mem obtains the CPU and memory utilization rates. 4) Hardware collection: a. CPU: Reads / proc / stat to calculate the utilization rate and reads
[0107] / sys / class / thermal / thermal_zone0 / temp to obtain the temperature (unit: millidegrees Celsius). b. Memory: Reads / proc / meminfo to calculate the utilization rate and fragmentation rate. c. Hard disk: Uses statvfs to obtain the disk utilization rate and gets the SMART health status through smartctl. d. Network card: Reads / proc / net / dev to calculate the bandwidth utilization rate and error rate of the network card. e. Graphics card: If the NVML library is supported, uses its API to obtain the GPU utilization rate and temperature. 5) Details: The collection frequency can be adjusted through the configuration file, and the data is passed to the data analysis module through a thread-safe queue. (4) Data analysis module. 1) Function: Analyzes the collected data and calculates the system health score based on the dynamic weighting algorithm. 2) Implementation method: Calculates the health scores of various indicators according to the weights provided by the configuration management module and the collected data. The weights support dynamic adjustment. For example, the weights of the CPU and memory are increased during high load. Example of the health score calculation formula: 3) Details: Supports anomaly detection (such as CPU utilization rate exceeding 90%), and the analysis results are passed to the data storage and notification modules. 4) Initial weight allocation: 50% for hardware indicators and 50% for software indicators. Dynamic adjustment: When the disk utilization rate > 50% is detected, the hardware weight is increased to 60% and the software weight is 40%. When a large number of non-system official software packages are detected in the current system, the software weight is increased to 60% and the hardware weight is 40%. (5) Data storage module. 1) Function: Saves the health assessment results and supports historical data query and trend analysis. 2) Implementation method: Uses a lightweight database (such as SQLite) to store the analysis results. The table structure includes fields such as timestamp, indicator type, and health score. File system storage (such as CSV format) can also be selected. 3) Details: Supports querying data by time range and periodically clears expired records to save space. (6) Notification module. 1) Function: Monitors the health score and sends an alarm when an anomaly is detected. 2) Implementation method: Sets the health score threshold (for example, below 60 points). When the score is below the threshold or a critical anomaly (such as service stop) is detected, an alarm is sent via email or system notification. 3) Details: Supports multiple notification methods (such as SMTP mail, logging), and the notification content includes anomaly details and timestamp. (7) Display module. 1) Function: Provides a visual interface in GUI mode to display the real-time health status and historical trends. 2) Implementation method: Develops the interface based on the Qt framework to display the health score, trend graph (such as line graph), and detailed indicators (such as CPU utilization rate, memory occupancy). 3) Details: The interface supports configuration of the refresh frequency, and users can switch views to view the status of specific modules.
[0108] In the method of the present invention, the system operation process is divided into four stages: startup, initialization, operation, and shutdown: (1) Startup and mode selection. The system selects the operation mode according to the command-line parameters: --no-gui: Enter the command-line mode and output the health check results to the terminal. --help: Display the help information and exit. Default: Enter the GUI mode and start the graphical interface. (2) Module initialization. Start all background modules in parallel: log management, configuration management, data collection, data analysis, data storage, and notification modules. Initialization order: log management → configuration management → other modules to ensure that logs and configurations are available first. (3) Operation mode. GUI mode: Start the Qt application, create the main window, and display the health status interface. Users can view the health score, trend chart, and detailed metrics in real time through the interface. Command-line mode: Enter the main loop, execute the health check logic once per second, and output the results to the terminal (such as "Health score: 85, CPU usage: 45%"). (4) Data flow and processing. 1) Collection: The data collection module continuously collects hardware and software data from the system. 2) Analysis: The data analysis module receives the collected data and calculates the health score. 3) Storage: The data storage module saves the analysis results to the database or file. 4) Notification: The notification module monitors the health score and triggers an alarm when an anomaly occurs. 5) Display: The display module (GUI mode) periodically reads data from the storage module and updates the interface. (5) System shutdown. Receive the shutdown signal (such as Ctrl+C or the GUI close button), and stop all modules in sequence: display → notification → storage → analysis → collection → configuration → log. Ensure that resources (such as file handles, database connections) are correctly released, and log the system shutdown event. (6) Collaboration mechanism. To ensure efficient collaboration between modules, the system adopts the following mechanisms: 1) Singleton pattern. Each module is implemented as a singleton to ensure that there is only one instance in the system, avoiding resource competition and duplicate initialization. Example: The log management module accesses the global instance through LogManager::getInstance(). 2) Thread-safe queue. A thread-safe queue (such as C++'s std::queue combined with a mutex) is used for data transfer between modules to ensure data integrity and concurrent security. Example: The data collection module pushes the collection results into the queue, and the data analysis module retrieves the data from the queue for processing. 3) Event-driven. The notification module triggers an alarm based on the threshold of the health score, adopting an event-driven mechanism to reduce unnecessary polling overhead. Example: When the health score is lower than 60, the notification module immediately sends an alarm email.
[0109] The system of this embodiment adopts a security-performance integrated evaluation model: (1) Software supply chain security verification module: VendorID verification (resolving the vendor identifier through RPM package signature) is introduced in the key component score C, and unofficial components are regarded as high-risk items to directly trigger security warnings. (2) Process behavior baseline comparison: In the key process score P, in addition to resource occupancy detection, a whitelist mechanism is also built-in: 1) Dynamically verify the process execution path hash value; 2) Match the process startup parameters with the security baseline. Traditional health monitoring systems separate security detection (such as vulnerability scanning) from performance monitoring, resulting in the inability to identify the "high-performance but high-risk" state (such as overclocking the CPU to run unsigned drivers). The system of this embodiment realizes the joint evaluation of security attributes and performance indicators through the C and P score items.
[0110] The system of this embodiment adopts a modular real-time collaboration architecture, including: A hybrid master control module of singleton mode + thread-safe queue: 1) The data collection, analysis, and storage modules run in singleton mode to ensure the uniqueness of the global state. 2) The multi-producer-single-consumer model is adopted, and the processing efficiency of key indicators is improved through a priority queue (such as hardware data being processed prior to software data); In addition, a cross-module event trigger chain is also adopted: 1) When the storage module detects that the historical health score continuously decreases, it automatically triggers the weight retraining process of the configuration management module. 2) The hardware overheat event (CPU temperature > 85°C) directly activates the emergency alarm channel of the notification module (such as SMS notification). Existing systems mostly adopt a linear pipeline architecture (collection → analysis → storage), with high coupling degree and poor scalability between modules. The loose coupling architecture of the present invention supports dynamically loading new collection plugins (such as adding GPU monitoring) without reconstructing the core logic.
[0111] The system of this embodiment adopts non-intrusive process resource tracking technology, including a process-level resource portrait generator: 1) By hijacking / proc / <pid>The / stat pseudo file system captures the time series characteristics of the CPU / memory occupancy of processes in real time. 2) The control group (CGroup) technology is used to model the baseline of resource occupancy of key processes and identify deviations from the normal mode (such as the CPU usage rate of the MySQL process suddenly increasing to 95% but without an increase in the query volume). Traditional methods rely on instantaneous sampling of the top or ps commands and cannot capture short-term resource leaks (such as slow memory growth). The continuous tracking mechanism of the present invention can detect progressive anomalies with a period > 30s, and the detection sensitivity is increased by more than 40%.
[0112] In addition, the notification module of the system in this embodiment adopts a multi-modal notification decision tree mechanism: (1) Hierarchical alarm engine: 1) Level 1 (score 70 - 60): Email notification + console log; 2) Level 2 (score 60 - 50): SMS notification + automatically generated fault tree analysis report; 3) Level 3 (score < 50): Trigger an automated repair script (such as restarting the service, isolating the faulty disk); (2) Context-aware notification policy: 1) During working hours, it is preferentially sent to the mobile terminal of the operation and maintenance personnel; 2) During non-working hours, it switches to AI voice alarm. Existing alarm systems usually use a single threshold to trigger fixed actions (such as sending an email when CPU > 90%), lacking a hierarchical response mechanism. The decision tree model in this embodiment combines the health score and the time context, significantly reducing the false alarm rate.
[0113] Through the above implementation solutions, the system of this embodiment can realize multi-dimensional health assessment of the Linux operating system, provide real-time monitoring, historical trend analysis and automated notification functions, and significantly improve the efficiency and security of system maintenance. The system of this embodiment solves the key problems in traditional technologies such as isolated indicators, rigid weights, and the separation of security and performance in the core structures and function combinations such as the hardware-software data fusion acquisition system, the dynamic weight self-adjusting algorithm, and the security-performance integrated evaluation model. Through the combination of the modular architecture and the non-intrusive tracking technology, it realizes the accurate assessment and intelligent response to the health status of the operating system in a complex environment, providing a breakthrough improvement solution for the existing technology system.
[0114] In addition, this embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the operating system health assessment method based on multi-dimensional index dynamic weighting through a processor.
[0115] In addition, this embodiment also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the operating system health assessment method based on multi-dimensional index dynamic weighting through a processor.
[0116] Those skilled in the art should understand that the technical solutions provided by the present invention can be in the form of a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be in the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or a device for realizing the functions specified in multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or a device for realizing the functions specified in multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or a device for realizing the functions specified in multiple blocks.
[0117] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of the present invention.< / pid> < / process> < / process> < / service> < / package> < / process> < / process> < / process> < / service> < / service>
Claims
1. A method for dynamically weighted health assessment of an operating system based on multi-dimensional indicators, characterized in that It includes the following steps: collecting the status information of each hardware module in the operating system, calculating multiple hardware status indicators and weighted summing them to obtain the score of the hardware module, where the hardware module includes some or all of CPU, memory, disk, network card, and graphics card. The hardware status indicators of the CPU include CPU usage rate and CPU temperature, the hardware status indicators of the memory include memory usage rate and memory fragmentation rate, the hardware status indicators of the disk include disk usage rate and disk damage rate, the hardware status indicators of the network card include network card bandwidth usage rate and packet loss rate, the hardware status indicators of the graphics card include GPU temperature and GPU utilization rate of the graphics card, adjusting the weight parameters of each hardware module in combination with real-time load characteristics or business priorities, and weighted summing the scores of each hardware module to obtain the hardware health score; collecting the software status information in the operating system and calculating the scores of each software module respectively, where the software module includes key components, key services, and key processes, and weighted summing the scores of each software module to obtain the software health score; weighted summing the hardware health score and the software health score to obtain the health score of the operating system to determine the health status of the operating system.
2. The method for dynamically weighted operating system health assessment based on multi-dimensional metrics according to claim 1, wherein Adjusting the weight parameters of each hardware module in combination with real-time load characteristics includes: calculating the real-time load characteristics of each hardware module respectively, and normalizing the real-time load characteristics of each hardware module as the weight parameters of each hardware module; adjusting the weight parameters of each hardware module in combination with business priorities includes: obtaining the current business scenario in the operating system, querying the pre-reviewed business scenario-hardware module priority mapping table according to the current business scenario to determine the business priorities of each hardware module, and converting the business priorities of each hardware module into the weight parameters of each hardware module.
3. The operating system health assessment method based on dynamic weighting of multi-dimensional indicators according to claim 1, characterized in that, When calculating multiple hardware status indicators and weighted summing them to obtain the score of the hardware module, it includes determining whether each hardware status indicator has an instantaneous anomaly. If a certain hardware status indicator has an instantaneous anomaly, first reduce the weight parameter of this hardware status indicator and then weighted sum to obtain the score of the corresponding hardware module.
4. The operating system health assessment method based on dynamic weighting of multi-dimensional indicators according to claim 1, characterized in that The default weight parameters of both the hardware health score and the software health score are 50%. When weighted summing the hardware health score and the software health score to obtain the health score of the operating system, it also includes obtaining the current hard disk usage rate of the operating system and the installation quantity or proportion of non-system official software packages. When the hard disk usage rate exceeds the first preset threshold, increase the weight parameter of the hardware health score and decrease the weight parameter of the software health score on the basis of the default weight parameter, and when the hard disk usage rate does not exceed the first preset threshold and the installation quantity or proportion of non-system official software packages exceeds the second preset threshold, decrease the weight parameter of the hardware health score and increase the weight parameter of the software health score on the basis of the default weight parameter.
5. The operating system health assessment method based on dynamic weighting of multi-dimensional indicators according to claim 1, characterized in that The operating system is a Linux operating system. When collecting the status information of each hardware module in the operating system, calculating multiple hardware status indicators and weighted summing them to obtain the score of the hardware module: The steps to obtain the score of the CPU include: opening the file " / proc / stat" to read the first status of the CPU, CPUstat1; delaying for a specified time; opening the file " / proc / stat" to read the second status of the CPU, CPUstat2; calculating the total difference and the idle difference between the first status CPUstat1 and the second status CPUstat2, and dividing the difference between the total difference and the idle difference by the total difference to obtain the CPU usage rate; opening the file " / sys / class / thermal / thermal_zone0 / temp" to read the original CPU temperature, and converting the original CPU temperature from millidegrees Celsius to degrees Celsius to obtain the CPU temperature; and performing a weighted sum of the CPU usage rate and the CPU temperature to obtain the score of the CPU; The steps to obtain the score of the memory include: opening the file " / proc / meminfo" to read the total physical memory MemTotal, the free memory size MemFree, the kernel buffer size Buffers, and the page cache size Cached of the memory; subtracting the free memory size MemFree, the kernel buffer size Buffers, and the page cache size Cached from the total physical memory MemTotal in sequence to obtain the used memory size; dividing the used memory size by the total physical memory MemTotal to obtain the memory usage rate, reading the file " / proc / buddyinfo" and calculating the memory fragmentation rate, and performing a weighted sum of the memory usage rate and the memory fragmentation rate to obtain the score of the memory; the step of reading the file " / proc / buddyinfo" and calculating the memory fragmentation rate includes: S101, opening the file " / proc / buddyinfo" and initialize variables total_free and small_free to 0; S102, process each line in the file " / proc / buddyinfo": parse out the memory region name and the number of free memory blocks count for each order; S103, for each order, calculate the free memory block size block_size according to block_size = 2 order *MIN_UNIT, where order is the index number of the order and MIN_UNIT is the minimum block size; S104, for each order, calculate the free space memory size order_free of this order according to order_free = count * block_size, where count is the number of free memory blocks of this order; S105, add the free space memory size order_free of each order to the variable total_free; S106, accumulate the free space memory size order_free of the first 3 orders to the variable small_free; S107, divide the variable small_free by the variable total_free to obtain the memory fragmentation rate; The steps to obtain the score of the disk include: calling the statvfs function to obtain the mount point data of the disk, and calculating the total capacity, the free capacity, and the used capacity of the disk according to the mount point data of the disk; dividing the used capacity by the total capacity to obtain the disk usage rate; using popen to call the smartctl command to obtain the SMART data of the disk, extracting the reallocated sector count, the number of uncorrectable errors, the command timeout count, and the total number of sectors from the SMART data, summing the reallocated sector count, the number of uncorrectable errors, and the command timeout count and then dividing by the total number of sectors to obtain the disk damage rate, and performing a weighted sum of the disk usage rate and the disk damage rate to obtain the score of the disk; The steps for obtaining the score of the network card include: opening the file " / proc / net / dev", parsing the initial sampling data after skipping the first two lines of headers to record the received and transmitted bytes and error counts; after a delay, opening the file " / proc / net / dev" again, parsing the re-sampled data after skipping the first two lines of headers to record the received and transmitted bytes and error counts, calculating the traffic during this period based on the two sampled data, dividing the traffic during this period by the length of this period and then dividing by the bandwidth of the network card to obtain the bandwidth utilization rate of the network card; calculating the total number of bytes transmitted and the number of bytes of error counts during this period based on the two sampled data, dividing the number of bytes of error counts during this period by the total number of bytes transmitted during this period to obtain the packet loss rate of the network card, and performing a weighted sum of the bandwidth utilization rate and the packet loss rate of the network card to obtain the score of the network card; The steps for obtaining the score of the graphics card include: obtaining the GPU temperature; initializing the NVML library, obtaining the GPU handle, calling the NVML library to obtain the GPU utilization rate, and closing the NVML library; performing a weighted sum of the GPU temperature and the GPU utilization rate to obtain the score of the graphics card.
6. The method for dynamically weighted operating system health assessment based on multi-dimensional indicators according to claim 1, wherein When collecting the software status information in the operating system and calculating the scores of each software module respectively, the steps for obtaining the score of the key component include: for the specified key components in the operating system, obtaining the installation status and whether it is an official version of each key component respectively, and calculating the score of the key component according to the following formula: Among them, S comp is the score of the key component, N1 is the number of key components, R i is the installation status of the key component i, and the installation status value is 1 or 0. A value of 1 indicates installed, and a value of 0 indicates not installed; V i indicates whether the key component i is an official version. A value of 1 indicates an official version, and a value of 0 indicates a non - official version; The steps for obtaining the score of the key service include: for the specified key services in the operating system, obtaining the startup status of each key service respectively, and calculating the score of the key service according to the following formula: Among them, S serv is the score of the key service, N2 is the number of key services, L j is the startup status of the key service j, with a value of 1 indicating started and a value of 0 indicating not started; The steps for obtaining the score of the key process include: for the specified key processes in the operating system, obtaining the startup status, CPU utilization rate and memory utilization rate of each key process respectively, and calculating the score of the key process according to the following formula: Among them, S ps is the score of the critical process, N3 is the number of critical processes, P k is the startup status of critical process k, with a value of 1 indicating started and a value of 0 indicating not started; P k,cpu + is the CPU usage rate of critical process k, P k,men is the memory usage rate of critical process k.
7. The method for operating system health assessment based on dynamic weighting of multi-dimensional indicators according to claim 5, characterized in that, The function expression for performing a weighted sum of the CPU utilization rate and the CPU temperature to obtain the score of the CPU is: S cpu = W1·S cpuT + W2·S cpuU , Among them, S cpu is the score of the CPU, W1 and W2 are weight parameters, and S cpuT is the CPU temperature, and S cpuU is the CPU usage rate; the functional expression for obtaining the score of the memory by weighted summation of the memory usage rate and the memory fragmentation rate is: S memory = W3·S memoryF + W4·S memoryU , Among them, S memory is the score of the memory, W3 and W4 are weight parameters, and S cpuT is the memory fragmentation rate, and S memoryU is the memory utilization rate; the functional expression for obtaining the score of the disk by weighted summation of the disk utilization rate and the disk damage rate is: S disk = W5·S diskBD + W6·S diskU , Among them, S disk is the score of the disk, W5 and W6 are weight parameters, and S diskBD is the disk damage rate, and S diskU is the disk utilization rate; the function expression for obtaining the score of the network card by weighted summation of the bandwidth utilization rate and packet loss rate of the network card is: S network = W7·S networkU + W8·S networkL , Among them, S network is the score of the network card, W7 and W8 are weight parameters, and S networkU is the bandwidth utilization rate of the network card, and S networkL is the packet loss rate; the functional expression for obtaining the score of the graphics card by weighted summation of the GPU temperature and the GPU utilization rate is: S gpu = W9·S gpuT + W 10 ·S gpuU , Among them, S gpu is the score of the graphics card, W9 and W 10 are weight parameters, S gpuT is the CPU temperature, S gpuU is the CPU usage rate; the function expression for obtaining the hardware health score by weighted summation of the scores of each hardware module is: S hardware = W 11 ·S cpu + W 12 ·S memory + W 13 ·S disk + W 14 ·S network + W 15 ·S gpu , Among them, S hardware is the hardware health score, W 11 ~W 15 are weight parameters, S cpu is the score of the CPU, S memory is the score of the memory, S disk is the score of the disk, S network is the score of the network card, S gpu is the score of the graphics card; the function expression for obtaining the software health score by weighted summation of the scores of each software module is: S software = W 16 · S cimp + W 17 · S serv + W 18 · S ps , Among them, S software is the software health score, W 16 to W 18 are weight parameters, S comp is the score of the key component, S serv is the score of the key service, S ps is the score of the key process; the functional expression for obtaining the health score of the operating system by weighted summation of the hardware health score and the software health score is: S core = W 19 · S software + W 20 · S hardware , In the above formula, S core is the health score of the operating system, and W 19 and W 20 are weight parameters.
8. An operating system health assessment system based on dynamic weighting of multi-dimensional indicators, comprising a microprocessor and a memory connected to each other, characterized in that, The microprocessor is programmed or configured to execute the operating system health assessment method based on dynamic weighting of multi-dimensional metrics according to any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program or instructions, characterized in that, The computer program or instruction is programmed or configured to execute the operating system health assessment method based on dynamic weighting of multi-dimensional metrics according to any one of claims 1 to 7 through a processor.
10. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instruction is programmed or configured to execute the operating system health assessment method based on dynamic weighting of multi-dimensional metrics according to any one of claims 1 to 7 through a processor.
Citation Information
Cited By
Guard system and method of multi-process control system
CN120560892A
Distributed database deployment planning method, device, equipment, medium and product
CN120994746A
Hardware health detection system and method for industrial control all-in-one machine
CN121277787A
Comprehensive evaluation method and device for network health degree of industrial control system
CN121486244A
Chip fault repair evaluation method and electronic equipment
CN121979730A