Server hardware failure early warning and recovery method and system based on intelligent optimization algorithm

Through multi-dimensional data collaborative analysis and dynamic strategy optimization based on intelligent optimization algorithm, the problems of low accuracy and long recovery time in server hardware fault warning and recovery are solved, and efficient and accurate fault warning and rapid recovery are achieved.

CN119621442BActive Publication Date: 2025-05-13GUANGZHOU HEDY COMPUTER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510154329.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-13
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

The existing technology has problems such as low accuracy, high false alarm rate, long recovery time and low resource utilization efficiency in server hardware failure warning and recovery, which is difficult to meet the high requirements of modern data centers for accurate warning and rapid recovery.

Method used

Using an intelligent optimization algorithm method, the mapping matrix and optimization mode matrix are constructed through collaborative analysis of multi-dimensional data, combined with the fault-tolerant strategy group to filter suitable fault-tolerant or backup strategies, and through dynamic threshold adjustment and chaotic mapping optimization strategy parameters, the dynamic matching of the optimal decision-making solution is achieved.

Benefits of technology

It improves the accuracy of fault warning, reduces the false alarm rate, shortens the fault recovery time, improves the efficiency of system resource utilization, and ensures the operating stability and response speed of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621442B_ABST
    Figure CN119621442B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of hardware fault early warning, and more specifically, to a server hardware fault early warning and recovery method and system based on an intelligent optimization algorithm, which obtains hardware self-check information; obtains load status information; obtains historical operation data; constructs a preliminary mapping matrix based on the obtained hardware self-check information, load status information and historical operation data; according to the preliminary mapping matrix, extracts key patterns and generates an optimized pattern matrix through a topological space optimization algorithm; based on the optimized pattern matrix, in combination with a fault-tolerant strategy group, screens a fault-tolerant or backup strategy suitable for the current state; according to the fault-tolerant or backup strategy, evaluates the benefit, cost and success probability of the strategy, and generates an optimal decision plan; outputs fault early warning information, including warning time, warning level and triggering strategy; outputs an optimal decision plan, including a fault-tolerant strategy or a backup strategy, which can capture more complex hardware abnormality patterns and greatly improve the accuracy of early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hardware failure early warning, and more specifically, to a server hardware failure early warning and recovery method and system based on an intelligent optimization algorithm. Background Art

[0002] With the rapid development of data centers and high-performance computer clusters, the operational stability and reliability of server hardware have become key factors in ensuring the normal operation of the entire system. However, due to the complex operating environment of server hardware, dynamic changes in workloads, and high-density distribution of hardware components, the prediction and handling of hardware failures face huge challenges. Traditional server hardware fault management methods mostly rely on fixed rules and simple threshold settings, which are difficult to meet the high requirements of modern data centers for accurate early warning and rapid recovery.

[0003] In the prior art, server hardware failure warnings usually use rule-based monitoring systems. This type of method determines whether the hardware status is abnormal by presetting a fixed threshold (for example, the temperature is higher than 75°C to trigger an alarm). However, this method has the following major problems:

[0004] Fixed thresholds cannot adapt to dynamically changing hardware operating environments. For example, under high load conditions, some hardware may operate normally even if it is close to the threshold, while under low load conditions, the same value may mean that the hardware is abnormal. This mechanism lacks dynamic adaptability and is prone to false positives or false negatives.

[0005] Existing methods usually analyze hardware status data or load information separately, ignoring the correlation and synergy between different data, and cannot fully reflect the actual operating status of the system. This single-dimensional analysis leads to low accuracy of fault prediction.

[0006] Traditional methods often adopt a single strategy for fault handling, such as directly shutting down the faulty hardware or enabling redundant devices. This simple handling method may be effective in some cases, but it lacks flexibility and efficiency in complex fault scenarios, which may cause resource waste or further performance degradation.

[0007] In addition, some advanced intelligent optimization methods, such as fault prediction systems based on a single deep learning algorithm, can improve prediction accuracy to a certain extent, but they also have significant limitations. These methods usually rely only on hardware status information and lack comprehensive consideration of load dynamic changes and historical trends, resulting in insufficient applicability in complex fault scenarios. At the same time, since the algorithm is limited to a single prediction function and lacks intelligent design of fault tolerance and recovery strategies, it is difficult to achieve closed-loop management from fault prediction to recovery.

[0008] Based on the above-mentioned deficiencies of the prior art, existing methods often exhibit problems such as insufficient data analysis, delayed policy response, and poor system adaptability when dealing with hardware failures, which seriously affects the operating efficiency and stability of the system. Summary of the invention

[0009] In view of the above-mentioned problems in the prior art, the present invention proposes a server hardware fault warning and recovery method and system based on an intelligent optimization algorithm, aiming to solve the following technical problems: improve the accuracy of fault warning, effectively identify complex hardware failure modes through collaborative analysis of multi-dimensional data; reduce the false alarm rate, combine dynamic threshold adjustment and intelligent optimization algorithm to improve the robustness of fault identification; shorten the fault recovery time, design flexible fault tolerance and backup strategies, and realize dynamic matching of optimal strategies through intelligent optimization methods; improve the efficiency of system resource utilization, optimize the execution of fault tolerance and backup strategies, and avoid resource waste and unnecessary performance loss.

[0010] The present invention provides a server hardware failure early warning and recovery method based on an intelligent optimization algorithm, comprising:

[0011] The acquisition steps include:

[0012] Obtain hardware self-test information, including temperature and voltage information of server sensors;

[0013] Get load status information, including the number of threads and services;

[0014] Obtain historical operation data, including historical fault trend data and hardware aging data;

[0015] Processing steps include:

[0016] Based on the acquired hardware self-test information, load status information and historical operation data, a preliminary mapping matrix is ​​constructed;

[0017] According to the preliminary mapping matrix, the key patterns are extracted and the optimized pattern matrix is ​​generated through the topological space optimization algorithm;

[0018] Based on the optimized mode matrix and the fault-tolerant strategy group, select the fault-tolerant or backup strategy that is suitable for the current state;

[0019] According to the fault tolerance or backup strategy, evaluate the benefits, costs and success probability of the strategy and generate the optimal decision plan;

[0020] Output steps include:

[0021] Output fault warning information, including warning time, warning level, and trigger strategy;

[0022] Output the optimal decision-making plan, including fault-tolerant strategy or backup strategy;

[0023] Based on the optimal decision-making scheme, the strategy parameters are dynamically adjusted using chaotic mapping;

[0024] Based on the dynamic adjustment results, the execution strategy is optimized and the system status is updated.

[0025] Preferably, the obtaining step specifically includes:

[0026] Obtain hardware self-test information of server sensors in real time through the hardware diagnosis module;

[0027] Obtain server load status information in real time through the load monitoring module;

[0028] Read historical fault data and hardware aging data from the history management module, and format and normalize the read data.

[0029] Preferably, the processing steps specifically include:

[0030] Use matrix operations to make preliminary associations between hardware self-check information, load status information, and historical operation data to construct a mapping matrix;

[0031] The weight relationship of the mapping matrix is ​​analyzed through the topological space optimization algorithm to extract the key patterns;

[0032] According to the relationship matrix between key patterns and strategy groups, fault-tolerant strategies are screened through group theory operations.

[0033] Preferably, the fault-tolerant strategy screening process specifically includes:

[0034] Evaluate the benefits of each fault-tolerant strategy and define the benefit function success ,

[0035] in, success is the success probability of the strategy, For the effectiveness of the strategy, The cost of the strategy;

[0036] The strategy corresponding to the maximum value of the benefit function is selected as the optimal fault-tolerant strategy.

[0037] Preferably, the construction of the mapping matrix in the processing step specifically includes:

[0038] Compare the hardware self-check information with the historical operation data through matrix transformation and calculate the similarity;

[0039] The comparison results are used for pattern classification through a recursive neural network, and a pattern weight matrix is ​​generated.

[0040] Preferably, the output step comprises:

[0041] Output the generated optimal decision-making plan to the server management platform;

[0042] Upload the fault-tolerant strategy or backup strategy to the early warning center and generate detailed execution logs for subsequent analysis.

[0043] Preferably, the physical environment information of the server is further obtained, including temperature, humidity, and airflow conditions, and the warning threshold is adjusted in combination with the environmental information.

[0044] Preferably, the generation of the optimized pattern matrix specifically includes:

[0045] The hardware status data and load information are constructed into a three-dimensional matrix through tensor operations;

[0046] Tensor decomposition technology is used to extract synergy patterns, and decisions are optimized based on the extracted synergy patterns.

[0047] Preferably, the execution of the fault-tolerant strategy or the backup strategy includes:

[0048] Dynamically allocate key resources in fault-tolerant strategies;

[0049] If the fault-tolerant strategy fails, it will automatically switch to the backup strategy and record the strategy execution status in real time.

[0050] The server hardware failure early warning and recovery system based on the intelligent optimization algorithm for executing the method comprises:

[0051] Hardware diagnostic module, used to obtain hardware self-test information of server sensors;

[0052] Load monitoring module, used to obtain server load status information;

[0053] History management module, used to store and read historical failure data and hardware aging trend data of the server;

[0054] Data processing module, used to build a mapping matrix based on the acquired information, perform pattern optimization and strategy screening;

[0055] A strategy execution module, used to execute the optimal fault-tolerant strategy or backup strategy;

[0056] The output module is used to output fault warning information and optimal decision-making solutions to the server management platform.

[0057] Specifically, the beneficial effects of the present invention are mainly reflected in the following aspects:

[0058] 1. Collaborative analysis of multi-dimensional data: The present invention realizes comprehensive analysis of hardware status, load information and historical data by constructing a mapping matrix and an optimization pattern matrix. Specifically, hardware status data (such as temperature and voltage) and load dynamic information are fused through tensor operations, and key patterns are extracted using topology optimization technology, thereby improving the accuracy of fault pattern recognition. Compared with the prior art, the method of the present invention can capture more complex hardware abnormality patterns and greatly improve the accuracy of early warning.

[0059] 2. Innovative application of intelligent optimization algorithm: The present invention introduces group theory to design a fault-tolerant strategy screening mechanism, combines the strategy benefit function to evaluate the advantages and disadvantages of different strategies, and ensures the scientific and efficient strategy selection. For example, by dynamically calculating the success probability, benefit and cost of the strategy, the method of the present invention can select the optimal strategy in a quantitative manner, effectively reducing recovery time and resource consumption.

[0060] 3. Dynamic threshold adjustment mechanism: In the prior art, fixed thresholds limit the flexibility of fault warning. The present invention dynamically adjusts the hardware status threshold in combination with physical environment data (such as temperature, humidity, and airflow conditions). For example, in a high temperature and high humidity environment, the hardware temperature warning threshold can be dynamically adjusted from 75°C to 70°C to more accurately reflect the hardware risk status. This dynamic adjustment mechanism effectively reduces the false alarm rate and improves the robustness of system operation.

[0061] 4. Collaborative design of fault tolerance and backup strategies: The present invention realizes the synergy of fault tolerance strategy and backup strategy through intelligent optimization algorithm. When the fault tolerance strategy cannot effectively solve the problem, the system can automatically switch to the backup strategy to achieve a smooth transition of fault handling. This design ensures the continuity of fault recovery and avoids the recovery delay or performance loss caused by the single strategy in traditional methods.

[0062] 5. Comprehensive closed-loop management of the system: From data collection to policy execution, the present invention builds a complete closed-loop management system. The hardware diagnosis module and the load monitoring module provide real-time data input, the data processing module generates the optimal decision through the optimization algorithm, and the policy execution module and the output module realize the dynamic response and feedback closed loop. Through this full-process design, the present invention significantly improves the response speed and reliability of the system.

[0063] By solving the above technical problems, the present invention demonstrates significant innovation and practical value in the field of server hardware fault management, and provides strong technical support for data centers and high-performance computing environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is the overall logic block diagram of the system of the present invention.

[0065] Figure 2It is a logic block diagram of data collection and preliminary processing of the present invention.

[0066] Figure 3 This is the strategy screening and decision logic block diagram of the present invention.

[0067] Figure 4 This is the final decision optimization logic block diagram based on chaos theory of the present invention.

[0068] Figure 5 It is the output and feedback logic block diagram of the present invention. DETAILED DESCRIPTION

[0069] Please refer to Figure 1-5 The present invention provides a server hardware failure early warning and recovery method based on an intelligent optimization algorithm, comprising the following steps:

[0070] First, the hardware self-test information of the server is obtained in real time through the hardware diagnosis module 1, including the temperature information and voltage information of the server sensor group; the load status information of the server is obtained in real time through the load monitoring module 2, including data such as the number of threads and the number of services; further, the historical operation data of the server is read from the history management module 3, including historical fault trend data and hardware aging data, and these data are formatted and normalized for subsequent analysis and processing.

[0071] Use multi-dimensional data (hardware status, load information, historical trends) to build a preliminary mapping relationship of the multi-dimensional fault space and define basic variables.

[0072] Server hardware status data, such as temperature and voltage.

[0073] Load status data, such as the number of threads and task queues.

[0074] It is the historical fault trend data.

[0075] Construct a preliminary mapping matrix:

[0076] ,

[0077] In the processing step, a preliminary mapping matrix is ​​constructed based on the acquired hardware self-test information, load status information, and historical operation data. The initial mapping matrix is ​​defined as follows:

[0078] ,

[0079] in, Hardware self-test information of server sensors, including temperature and voltage information; Load status information, including the number of threads and services; Historical failure trend data and hardware aging data; They are the dimensions of hardware information, load status and historical data respectively.

[0080] By comprehensively considering multiple data sources, we can have a more comprehensive understanding of the server status, thereby improving the accuracy of fault prediction.

[0081] The initial mapping process helps identify the most valuable features for fault early warning, providing more targeted information for the next steps.

[0082] Data usually comes from server log files, sensor readings, and database records. To ensure the quality of the data, preprocessing is required, such as removing noise, filling missing values, and standardizing the data format.

[0083] These data are crucial to the present invention as they are the basis for all subsequent analyses and determine the performance of the fault prediction model.

[0084] Suppose a node in a server cluster has an abnormally high temperature, which may be caused by a fan failure or poor heat dissipation. By analyzing the temperature data of the node 、CPU usage and similar incidents in the past , a mapping matrix can be constructed to evaluate the current situation and predict possible failures.

[0085] According to the above matrix ,Furthermore, the topological space optimization algorithm is used to analyze the weight relationship of the mapping matrix, extract the key patterns and generate the optimized pattern matrix The specific optimization algorithm is as follows:

[0086] 1. Define topological space ,in is the vertex set of the pattern, is the set of associated edges between patterns;

[0087] 2. Assign weight to each mode:

[0088] ,

[0089] 3. Construct optimization objective function:

[0090] ,

[0091] 4. Solve the optimization objective function and obtain the optimized pattern matrix , .

[0092] The final output step includes outputting fault warning information, such as warning time, warning level and triggering strategy, and further outputting the optimal decision plan, including fault tolerance strategy or backup strategy. These decision plans are directly transmitted to the server management platform for execution through the output module 4.

[0093] By classifying failure modes, we can help quickly locate the root cause of the problem and reduce troubleshooting time. Revealing the potential connections between different failure modes helps understand the behavior of complex systems and prevent cascading failures in advance.

[0094] From the preliminary mapping matrix To extract key patterns from the dataset, further clustering analysis or principal component analysis (PCA) may be needed to reduce the dimensionality. When considering the relative importance and frequency of each mode in the entire system.

[0095] In a data center, some servers often have hard disk read and write errors and memory leaks at the same time. By constructing the topological space, we can find the strong correlation between these two failure modes, and then take measures to strengthen storage and memory management to prevent similar situations from happening again in the future.

[0096] Through the above steps, the method of the present invention can collect and comprehensively analyze server hardware status information, load status information and historical operation data in real time, effectively extract key modes based on the topology optimization algorithm, and generate the optimal early warning and recovery plan. Compared with the existing technology, the accuracy of fault warning and the flexibility of recovery strategy are significantly improved.

[0097] The acquisition step specifically includes the following contents:

[0098] The hardware self-test information of the server sensor group is read in real time through the hardware diagnosis module 1, including temperature information and voltage information. Preferably, the monitoring range of the temperature information is -40°C to +85°C, and the monitoring range of the voltage information is 1.0V to 12.0V. Based on empirical analysis, these parameter ranges cover the normal working conditions and abnormal states of most server hardware.

[0099] The load status information of the server, including the number of threads and the number of services, is collected in real time through the load monitoring module 2. Preferably, when the number of threads exceeds 1000 or the number of services reaches 50 or more, the warning level will be triggered. These thresholds are set based on industry experience data and actual operation conditions, and can effectively distinguish between normal and abnormal load states.

[0100] The historical fault data and hardware aging data are read from the history management module 3 and normalized. Preferably, the normalization formula is:

[0101] ,

[0102] in, and are the minimum and maximum values ​​of historical data respectively. is the normalized historical fault data.

[0103] Through the above acquisition steps, the present invention can accurately capture the real-time status of server hardware and load, and perform normalization processing in combination with historical data, thereby providing a high-quality data foundation for subsequent pattern extraction and decision optimization.

[0104] The processing steps specifically include the following:

[0105] First, based on hardware self-test information, load status information and historical operation data, a preliminary mapping matrix is ​​constructed through matrix transformation Preferably, the calculation formula of the mapping matrix is:

[0106] ,

[0107] in, are the weight coefficients of hardware information, load information and historical data, and the preferred values ​​are 0.4, 0.3, and 0.3 respectively to ensure the balance of various types of data. Next, the mapping matrix is ​​optimized through the topological space optimization algorithm. Perform optimization, extract key patterns and generate optimized pattern matrix During the optimization process, the correlation between patterns is defined as:

[0108] ,

[0109] And construct the correlation matrix based on the correlation degree.

[0110] Furthermore, according to the relationship matrix between key patterns and strategy groups, the fault-tolerant strategy suitable for the current state is screened through group theory operations. The specific formula is:

[0111] ,

[0112] in, For the strategy group, For the selected strategies, The generator of the strategy.

[0113] Through the above processing steps, the present invention can build an efficient fault mode recognition and optimization mechanism based on multi-dimensional data, combine group theory operations to accurately screen fault-tolerant strategies, and significantly improve the reliability and decision-making efficiency of the system.

[0114] Multiple strategies can be combined through group operations to produce more effective solutions than a single strategy. Group theory provides a wealth of tools to design and evaluate different fault-tolerance mechanisms to adapt to various possible failure scenarios.

[0115] Strategy Group The design needs to be based on past experience and best practices, and should also take into account the architectural characteristics of the current system. In actual applications, it may be necessary to simulate and test different strategy combinations to determine their feasibility and effectiveness.

[0116] When a database server's disk I / O performance is detected to be degraded, several different fault-tolerance strategies can be tried, such as increasing cache size, migrating hot data to faster storage devices, or starting backup copies. Through group theory analysis, the most appropriate combination can be selected to restore service and minimize the impact on user experience.

[0117] The fault-tolerant strategy screening process specifically includes:

[0118] In an embodiment of the present invention, each fault tolerance strategy is first evaluated for its effectiveness to determine its applicability under the current hardware status and load conditions. Preferably, the present invention defines a benefit function To quantitatively evaluate the fault-tolerant strategy, the specific formula is:

[0119] ,

[0120] in, Fault Tolerance Strategy The success probability is calculated by combining historical data and current operation status prediction; For strategy The benefit preferably represents the quantitative value of the positive improvement on the server operation status after the policy is executed; For strategy The cost usually includes time cost, resource consumption and possible execution risk. Any policy in the policy set.

[0121] Furthermore, the method of the present invention compares the benefit function values ​​of all candidate strategies and selects the strategy corresponding to the maximum benefit function as the optimal fault-tolerant strategy. The specific steps are:

[0122] ,

[0123] In a preferred embodiment of the present invention, the fault tolerance strategy may include redistributing the load, suspending non-critical tasks, and starting the hardware redundancy mechanism. Taking load redistribution as an example, when the number of threads on a server exceeds 1000 and the calculated benefit function value is When , preferably a load redistribution strategy is initiated.

[0124] Through the above-mentioned fault-tolerant strategy screening method, the present invention can select the optimal strategy in a quantitative manner, combining the multi-dimensional factors of success probability, benefit and cost, to ensure the effectiveness and efficiency of strategy execution, and significantly improve the overall reliability and availability of the system.

[0125] Through probability assessment, the risk level of each strategy can be understood before execution, helping decision makers make more informed choices. Not only does it focus on technical feasibility, but it also takes into account economic factors to ensure that the selected solution is both effective and economical. success It can be learned from historical data or estimated through experiments and simulations. and cost Detailed calculations are required based on specific business needs and technical environment.

[0126] If you choose to replace the failed hard drive as a repair measure, you need to evaluate the cost of the new hard drive, the time required for installation, and the additional revenue that can be generated after restoring normal operations. By comparing the benefit functions of different strategies, you can ultimately decide whether to replace the hard drive immediately or take temporary mitigation measures first.

[0127] The construction of the mapping matrix in the processing step specifically includes:

[0128] In an embodiment of the present invention, the hardware self-check information is compared with the historical operation data through matrix transformation to generate a mapping matrix Preferably, the calculation formula of the mapping matrix is:

[0129] ,

[0130] in, For the Hardware self-test information, such as real-time temperature data of a sensor group; For the Load status information, such as the number of threads currently being served; For the Historical fault data, are the weight coefficients of hardware information, load information, and historical data, respectively. The preferred values ​​are 0.4, 0.3, and 0.3 to balance the impact of different data types.

[0131] Furthermore, the method of the present invention uses a recursive neural network to compare the results and classify the patterns, and generates a pattern weight matrix Preferably, the calculation formula of the mode weight is:

[0132] ,

[0133] In practical applications, for example, when a hardware temperature , Number of threads , historical temperature peak When , the weight value in the mapping matrix can be calculated by the above formula , thus judging it as a high-risk mode.

[0134] Through the above method, the present invention can effectively combine current status data with historical data, realize the organic integration of multi-dimensional information through matrix mapping and pattern classification, and provide accurate risk assessment basis for subsequent strategy screening.

[0135] The output step comprises:

[0136] In a preferred embodiment of the present invention, the output content includes fault warning information and optimal decision-making solutions. Preferably, the fault warning information includes warning time, warning level and triggering strategy. For example, when the temperature of a certain hardware Over 75°C and load thread number When it exceeds 1000, the third level warning information is triggered, including:

[0137] Warning time: the timestamp recorded by the system, such as 10:30:00 on January 5, 2025;

[0138] Warning level: preferably divided into level 1 (low), level 2 (medium) and level 3 (high), in this case it is level 3;

[0139] Triggering strategy: The optimal fault-tolerance strategy selected above is used, such as redistributing the load or enabling hardware redundancy.

[0140] Furthermore, the output optimal decision solution is transmitted to the server management platform through the output module 4. Preferably, the output format is a JSON data structure, which is convenient for the system to receive and execute.

[0141] At the same time, the system will generate a detailed execution log to record the entire process of strategy execution, including trigger conditions, execution time, feedback information, etc., to provide data support for subsequent analysis.

[0142] Through the above-mentioned output steps, the present invention can output key warning information and optimal decision-making solutions in a clear and structured manner, and support the traceability of system operation through detailed execution logs, which significantly improves the reliability and intelligence level of the warning and recovery system.

[0143] Further obtain the physical environment information of the server, including temperature, humidity, and airflow conditions, and adjust the warning threshold based on the environmental information.

[0144] In an embodiment of the present invention, the temperature and humidity data and airflow conditions of the server operating environment are collected in real time through the environment sensing module 5. Preferably, the temperature and humidity collection range is -20°C to +50°C, humidity is 10% to 90%, and the airflow speed range is 0.1 to 2.0 m / s. These parameters are selected based on the common environmental conditions of the data center to ensure that most operating scenarios are covered.

[0145] Furthermore, the warning threshold of the hardware is dynamically adjusted in combination with the physical environment information. For example, when the collected ambient temperature is 40°C and the humidity is 80%, the method of the present invention adjusts the third-level warning threshold of the hardware temperature from 75°C to 70°C to adapt to the characteristics of reduced thermal stability of the hardware in a high temperature and high humidity environment.

[0146] By combining the physical environment information of the server, the present invention can dynamically adjust the warning threshold, more accurately reflect the risk level of the hardware status, reduce the possibility of false alarms or missed alarms, and thus improve the accuracy and reliability of the warning system.

[0147] The optimized pattern matrix generation specifically includes:

[0148] In an embodiment of the present invention, hardware status data, load information and historical operation data are used to construct a three-dimensional matrix through tensor operations, thereby realizing multi-dimensional fusion and collaborative analysis of data. Preferably, the three-dimensional matrix is ​​defined as follows:

[0149] ,

[0150] in, It is a three-dimensional tensor matrix, which represents the collaborative relationship between hardware status, load information and historical data; For the Hardware status information, such as temperature and voltage; For the Load information, such as the number of threads and the number of services; For the Historical operation data, are weight coefficients respectively, and the preferred values ​​are 0.5, 0.3, and 0.2 to reflect the importance of the hardware status.

[0151] Furthermore, the method extracts patterns from the three-dimensional matrix through tensor decomposition technology to identify collaborative patterns between data dimensions. Preferably, the tensor decomposition formula is as follows:

[0152] ,

[0153] in, They are the decomposition vectors of hardware status, load information and historical data respectively; is the rank of tensor decomposition, and the preferred value is 3 to ensure the accuracy and efficiency of pattern extraction; It is the outer product operation of tensors.

[0154] Through the above decomposition process, the collaborative mode matrix is ​​generated , and further optimized in combination with the mode weight. Preferably, the calculation formula of the mode weight is:

[0155] ,

[0156] For example, the hardware status temperature data of a server , Number of load threads , historical temperature peak , the corresponding tensor matrix element can be calculated by the above formula: 402.5. Combined with tensor decomposition, the extracted coordination pattern can reflect the potential risks of the server under high temperature and high load conditions.

[0157] Through the above method, the present invention can use tensor decomposition technology to efficiently extract collaborative patterns, and combine pattern weights to achieve optimization, which significantly improves the comprehensive analysis capabilities of hardware status and load information, and provides a more accurate basis for the formulation of subsequent fault tolerance and backup strategies.

[0158] The execution of the fault tolerance strategy or backup strategy includes:

[0159] In a preferred embodiment of the present invention, the execution process of the fault-tolerant strategy first dynamically allocates key resources to reduce the pressure on high-load hardware, specifically including adjusting the distribution of task threads or suspending non-critical services. Preferably, when the number of threads of a server exceeds 1200 and the load evaluation result exceeds 80%, the method of the present invention will give priority to starting the task reallocation strategy. For example, part of the computing tasks will be transferred to other servers with lower loads, and the adjustment effect will be monitored in real time.

[0160] Furthermore, the method supports the dynamic activation of the hardware redundancy mechanism. When the status information of a hardware component triggers a third-level warning (such as a temperature exceeding 80°C) and the fault-tolerant strategy fails to significantly reduce its risk, the hardware redundancy backup mechanism is preferably activated. Specifically, the current task is switched to the backup hardware, and the abnormal hardware is diagnosed and repaired offline.

[0161] At the same time, this method will record the execution of the fault-tolerant strategy or backup strategy and generate a strategy execution log, including the following:

[0162] 1. The specific time when the policy is triggered, such as 10:35:00 on January 5, 2025;

[0163] 2. The type of execution strategy, such as load redistribution or hardware redundancy;

[0164] 3. The success rate of strategy execution, such as 98%;

[0165] 4. Impact on system status, such as the number of threads reduced from 1200 to 800 and the hardware temperature reduced from 80°C to 70°C.

[0166] By dynamically allocating resources and enabling hardware redundancy mechanisms, the method of the present invention can significantly improve the fault tolerance of the system, while reducing the risk of hardware failures, ensuring task continuity and system stability. Combined with detailed execution logs, this method provides a transparent feedback mechanism for system management, facilitating subsequent optimization and improvement.

[0167] In the preferred embodiment of the present invention, in order to further improve the robustness and flexibility of strategy execution, a final decision optimization step based on chaos theory is designed. Specifically, the method of the present invention dynamically adjusts the parameters of the optimal strategy through chaos mapping to ensure the optimality and adaptability of the strategy in a complex and dynamic environment.

[0168] Preferably, the chaotic mapping adopts the standard Logistic equation, which is defined as follows:

[0169]

[0170] in, is the state variable of the current strategy parameters; The strategy parameters adjusted for the next step; is the control parameter of the chaotic system, and the preferred value is 3.5, which is used to introduce nonlinear dynamic characteristics. The final optimization decision matrix is ​​generated:

[0171] ,

[0172] In the specific implementation process, the method of the present invention dynamically adjusts the policy parameters based on the above-mentioned chaotic mapping formula before each policy execution. For example, when the initial value of the thread allocation weight of the fault-tolerant policy is 0.5, the parameter adjusted by the chaotic mapping may become 0.525. This adjustment value further optimizes the task allocation efficiency in combination with the current load conditions.

[0173] By introducing chaotic mapping, the strategy parameters of the present invention can dynamically adapt to the complexity of the environment during execution, avoiding the system rigidity problem that may be caused by fixed parameters. Compared with systems that do not use chaotic optimization, the method of the present invention significantly improves the flexibility and adaptability of strategy execution, especially in complex and variable load environments.

[0174] Chaos theory allows the system to flexibly adjust according to real-time changes, improving the ability to respond to emergencies. By introducing a certain degree of randomness, the system can avoid falling into a local optimal solution, thereby achieving better long-term performance.

[0175] Dynamic adjustment coefficient The setting should be based on the specific characteristics of the system, such as response time and stability requirements.

[0176] Final decision matrix It is the result of multiple iterations of optimization, which integrates information from all previous levels to form a complete fault warning and recovery plan.

[0177] In the face of a sudden surge in network traffic, the server may be under tremendous pressure. At this time, through dynamic adjustments guided by chaos theory, resource allocation strategies can be changed in a timely manner, such as adding temporary computing nodes or restricting access to certain non-critical services, to maintain the overall service quality.

[0178] The present invention also discloses a server hardware fault early warning and recovery system based on an intelligent optimization algorithm for executing the method. The system specifically includes: a hardware diagnosis module 1, a load monitoring module 2, a history management module 3, a data processing module 4, a strategy execution module 5 and an output module 6. Each module operates collaboratively through the server system intranet to achieve accurate early warning and recovery of hardware faults.

[0179] In a preferred embodiment of the present invention, the hardware diagnosis module 1 is used to obtain the hardware self-test information of the server sensor in real time. Preferably, the module integrates a temperature sensor and a voltage sensor, and can detect hardware status data with a temperature range of -40°C to +85°C and a voltage range of 1.0V to 12.0V. The hardware diagnosis module 1 transmits the collected real-time information to the data processing module 4 through the integrated circuit bus.

[0180] The load monitoring module 2 obtains load status information in real time by connecting to the central processing unit system. The load information includes the current number of threads and the number of services. Preferably, when the number of threads exceeds 1000 or the number of services reaches 50, the load monitoring module 2 will mark it as a high load state and transmit the load data to the data processing module 4.

[0181] The history management module 3 is used to record and manage the historical operation data of the server, including historical fault trend data and hardware aging trend data. Furthermore, the history management module 3 formats the data into a standard input form through normalization processing so that the data processing module 4 can perform subsequent comprehensive analysis.

[0182] The data processing module 4 is the core module of the system, which is used to analyze the hardware self-test information, load status information and historical operation data, including constructing a preliminary mapping matrix, optimizing the mode matrix and screening the optimal fault-tolerant strategy. In a preferred embodiment of the present invention, the data processing module 4 combines intelligent optimization algorithms such as tensor operations, topology optimization and group theory operations to ensure the accuracy of data analysis and the efficiency of strategy screening. For example, when the hardware temperature information is 75°C and the number of threads is 1200, the data processing module 4 can generate a key mode matrix through an optimization algorithm and screen the applicable load redistribution strategy.

[0183] The policy execution module 5 is used to execute the optimal decision solution selected by the data processing module 4. Preferably, when the fault-tolerant strategy is triggered, the policy execution module 5 will dynamically allocate resources and transfer high-load tasks to other servers; when the backup strategy is triggered, the policy execution module 5 will automatically switch to the hardware redundant device and record detailed policy execution information.

[0184] The output module 6 is responsible for transmitting the fault warning information and the optimal decision-making plan to the server management platform. Preferably, the output content includes the warning time, warning level, the type of strategy executed and its execution status. For example, when the temperature of a certain hardware exceeds 80°C, the output module 6 will output the third-level warning information and send the optimal backup strategy to the management platform in JSON format.

[0185] The real-time data collected by the hardware diagnosis module 1 and the load monitoring module 2 are transmitted to the data processing module 4 through the server system intranet. Combined with the historical data provided by the history management module 3, the data processing module 4 can accurately analyze the current hardware status and generate the optimal strategy. The strategy execution module 5 and the output module 6 realize the dynamic execution of the strategy and the feedback loop of information based on the decision results of the data processing module 4.

[0186] The system of the present invention realizes a complete closed loop from real-time monitoring to policy execution of hardware failures through the division of labor and cooperation among various modules, significantly improving the system's fault warning accuracy and recovery efficiency. At the same time, the optimized design of each module in the system ensures stability and efficiency during operation. For example, the real-time data acquisition capabilities of the hardware diagnosis module 1 and the load monitoring module 2, combined with the intelligent optimization algorithm of the data processing module 4, effectively reduce the risk of system downtime caused by hardware anomalies.

[0187] In the preferred embodiment of the present invention, each module is seamlessly integrated with a standardized interface, which is suitable for complex computing environments such as data centers and high-performance computer clusters, and provides a highly intelligent solution for hardware management in these scenarios. Through the execution log function of the output module 6, the transparency and traceability of the system are further guaranteed, providing detailed data support and optimization basis for the operation of the system. The present invention aims to improve the warning accuracy, recovery efficiency and system stability of server hardware failures through a server hardware failure warning and recovery method and system based on an intelligent optimization algorithm. In order to verify the superiority of the present invention, a special embodiment and two comparative examples are compared and analyzed, and the performance of each solution is evaluated by a standardized test method.

[0188] The test indicators are as follows:

[0189] 1. Warning accuracy (%): refers to the ratio of the system to accurately issue a warning before an actual hardware failure occurs, which evaluates the system's ability to perceive failure risks.

[0190] 2. Mean recovery time from failure (minutes): The average time from when a failure occurs to when the system resumes normal operation, which evaluates the response efficiency of the system.

[0191] 3. False alarm rate (%): The rate at which the system incorrectly issues an early warning when there is no actual fault, which evaluates the robustness of the algorithm.

[0192] 4. System resource utilization (%): After the fault tolerance or backup strategy is executed, the usage of system hardware resources is used to evaluate the execution efficiency of the strategy.

[0193] The testing standards and methods are as follows:

[0194] Testing standards:

[0195] 1. The accuracy of early warning is calculated as the percentage of successful warning times to the total number of failures, and it must reach more than 90% to be considered excellent.

[0196] 2. The average failure recovery time is measured in minutes, preferably less than 10 minutes.

[0197] 3. The false alarm rate should be less than 5% to be considered good.

[0198] 4. System resource utilization should be over 80% for optimal performance.

[0199] The detection method is as follows:

[0200] The test environment simulates the typical load conditions of a data center and sets random hardware failure events during 48 hours of operation, including sensor overheating, voltage anomalies, and load exceeding the limit, to evaluate the system's performance under different failure scenarios.

[0201] The embodiments and comparative examples are set as follows:

[0202] Embodiment (system of the present invention):

[0203] Based on intelligent optimization algorithms, combined with tensor operations, topology optimization and group theory strategy screening, chaos optimization and other technologies, accurate early warning and recovery of hardware failures can be achieved.

[0204] Comparative Example 1 (Traditional Rule Algorithm):

[0205] Fault warning and strategy selection are based on fixed rules and threshold settings, and lack the ability for dynamic adjustment and intelligent optimization.

[0206] Comparative Example 2 (optimization system based on a single algorithm):

[0207] Hardware status prediction is only based on a single deep learning model (such as CNN), without multi-dimensional data collaborative analysis and strategy optimization capabilities.

[0208] The test results are shown in Table 1:

[0209] Table 1. Comparison of test results of Example 1, Comparative Example 1 and Comparative Example 2

[0210] index Example Comparative Example 1 Comparative Example 2 Early warning accuracy rate (%) 97.5 75.2 85.6 Mean time to recover from a failure (minutes) 5.8 15.4 11.2 False alarm rate (%) 2.6 13.7 7.8 System resource utilization (%) 91.3 78.5 83.1

[0211] According to the test results in Table 1, it can be seen that:

[0212] 1. Early warning accuracy: The early warning accuracy of the embodiment of the present invention reaches 97.5%, which is much higher than 75.2% of Comparative Example 1 and 85.6% of Comparative Example 2. This advantage is due to the fact that the present invention introduces tensor operations and topology optimization algorithms in the data processing module, performs multi-dimensional collaborative analysis of hardware status, load information and historical data, can capture complex hardware abnormality patterns, and significantly improves the accuracy of early warning.

[0213] 2. Average fault recovery time: The average fault recovery time of the embodiment is 5.8 minutes, which is significantly shorter than 15.4 minutes of comparative example 1 and 11.2 minutes of comparative example 2. The present invention ensures the rapidity and efficiency of strategy execution by dynamically adjusting strategy parameters through group theory strategy screening and chaos optimization, especially in load redistribution and redundant switching strategies.

[0214] 3. False alarm rate: The false alarm rate of the present invention is only 2.6%, which is significantly lower than 13.7% of Comparative Example 1 and 7.8% of Comparative Example 2. Through the dynamic threshold adjustment mechanism, the present invention can accurately distinguish between normal fluctuations of hardware and real abnormalities, reduce unnecessary alarm interference, and further improve the stability of the system.

[0215] 4. System resource utilization: The resource utilization of the embodiment reaches 91.3%, which is significantly higher than 78.5% of Comparative Example 1 and 83.1% of Comparative Example 2. This high efficiency is due to the optimized design of the fault tolerance and backup strategy of the present invention, which further improves the resource allocation efficiency and the balance of task execution through the dynamic adjustment of chaos theory.

[0216] The comprehensive test results show that the system of the present invention has significant advantages in terms of early warning accuracy, fault recovery time, false alarm rate and resource utilization, especially in complex hardware failure scenarios, the multi-dimensional collaborative ability of its intelligent optimization algorithm significantly improves system performance. Therefore, the embodiment of the present invention is selected as the best embodiment.

[0217] Through comparative tests of the embodiments and comparative examples, the method and system of the present invention show excellent performance in server hardware fault warning and recovery. Its innovation is reflected in the comprehensive application of multi-dimensional data collaborative analysis, dynamic fault-tolerant strategy screening and intelligent optimization algorithm. These advantages not only improve the accuracy of fault warning and recovery efficiency, but also significantly reduce the false alarm rate and resource waste, providing reliable guarantee for the efficient operation and maintenance of modern data centers.

[0218] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modification, replacement, and improvement made within the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A server hardware failure early warning and recovery method based on an intelligent optimization algorithm, characterized in that: include: The acquisition steps include: Obtain hardware self-test information, including temperature and voltage information of server sensors; Get load status information, including the number of threads and services; Obtain historical operation data, including historical fault trend data and hardware aging data; Processing steps include: Based on the acquired hardware self-test information, load status information and historical operation data, a preliminary mapping matrix is ​​constructed; According to the preliminary mapping matrix, the key patterns are extracted and the optimized pattern matrix is ​​generated through the topological space optimization algorithm; Based on the optimized mode matrix and the fault-tolerant strategy group, select the fault-tolerant or backup strategy that is suitable for the current state; According to the fault tolerance or backup strategy, evaluate the benefits, costs and success probability of the strategy and generate the optimal decision plan; Output steps include: Output fault warning information, including warning time, warning level, and trigger strategy; Output the optimal decision-making plan, including fault-tolerant strategy or backup strategy; Based on the optimal decision-making scheme, the strategy parameters are dynamically adjusted using chaotic mapping; According to the dynamic adjustment results, optimize the execution strategy and update the system status; The processing steps specifically include: Use matrix operations to make preliminary associations between hardware self-check information, load status information, and historical operation data to construct a mapping matrix; The weight relationship of the mapping matrix is ​​analyzed through the topological space optimization algorithm to extract the key patterns; According to the relationship matrix between key patterns and strategy groups, fault-tolerant strategies are screened through group theory operations.

2. The method according to claim 1, characterized in that The acquisition step specifically includes: Obtain hardware self-test information of server sensors in real time through the hardware diagnosis module; Obtain server load status information in real time through the load monitoring module; Read historical fault data and hardware aging data from the history management module, and format and normalize the read data.

3. The method according to claim 1, characterized in that The fault-tolerant strategy screening process specifically includes: Evaluate the benefits of each fault-tolerant strategy and define the benefit function F(g i )=P(success|g i )·B(g i )-C(g i ), Among them, P(success|g i ) is the success probability of the strategy, B(g i ) is the benefit of the strategy, C(g i ) is the cost of the strategy; The strategy corresponding to the maximum value of the benefit function is selected as the optimal fault-tolerant strategy.

4. The method according to claim 1, characterized in that The construction of the mapping matrix in the processing step specifically includes: Compare the hardware self-check information, load status information and historical operation data through matrix transformation and calculate the similarity; The comparison results are used for pattern classification through a recursive neural network, and a pattern weight matrix is ​​generated.

5. The method according to claim 1, characterized in that The output step comprises: Output the generated optimal decision-making plan to the server management platform; Upload the fault-tolerant strategy or backup strategy to the early warning center and generate detailed execution logs for subsequent analysis.

6. The method according to claim 1, characterized in that Further obtain the physical environment information of the server, including temperature, humidity, and airflow conditions, and adjust the warning threshold based on the environmental information.

7. The method according to claim 1, characterized in that The optimized pattern matrix generation specifically includes: The hardware status data, load information and historical operation data are constructed into a three-dimensional matrix through tensor operations; Tensor decomposition technology is used to extract synergy patterns, and decisions are optimized based on the extracted synergy patterns.

8. The method according to claim 1, characterized in that The execution of the fault tolerance strategy or backup strategy includes: Dynamically allocate key resources in fault-tolerant strategies; If the fault-tolerant strategy fails, it will automatically switch to the backup strategy and record the strategy execution status in real time.

9. A server hardware failure early warning and recovery system based on an intelligent optimization algorithm that executes the method according to any one of claims 1 to 8, characterized in that: include: Hardware diagnostic module, used to obtain hardware self-test information of server sensors; Load monitoring module, used to obtain server load status information; History management module, used to store and read historical failure data and hardware aging trend data of the server; Data processing module, used to build a mapping matrix based on the acquired information, perform pattern optimization and strategy screening; A strategy execution module, used to execute the optimal fault-tolerant strategy or backup strategy; The output module is used to output fault warning information and optimal decision-making solutions to the server management platform.

Citation Information

Patent Citations

  • Method and system for self-detection and self-recovery of software fault of energy controller

    CN118152172A

  • Server fault prediction method and device

    CN118394559A