A method and system for monitoring and diagnosing computer motherboard faults
By constructing a virtual rail index model and multi-level evaluation method, real-time monitoring of the blurring and energy absorption of power rails is solved, and the inadequacy of power rail instability monitoring in the existing technology is achieved, high-precision fault diagnosis and early warning are achieved, and the long-term stability and operation and maintenance efficiency of the equipment are improved.
Patent Information
- Application Number
- CN202510558628.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing technology has shortcomings in power rail instability monitoring and fault diagnosis, and it is impossible to identify potential hidden faults in a timely manner. The reliability of relying on backup power supplies affects the long-term stability of the system, and lacks a global power supply health assessment mechanism, resulting in insufficient fault warning capabilities.
By collecting multi-power rail data in real time, pre-processing with BIOS interface and high-precision sensors, building a virtual rail index model, calculating the voltage rail blur index VEI, lightweight behavior risk index LBPI, and power supply health assessment index PHS, conducting multi-level evaluations, identifying the health status of the power system and triggering corresponding response measures.
It improves the accuracy of fault diagnosis and early warning capabilities, reduces equipment downtime and maintenance costs, improves the long-term reliability and stability of the equipment, and reduces unnecessary maintenance intervention.
Smart Images

Figure CN120085735B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of motherboard monitoring, and particularly to a method and system for monitoring and diagnosing computer motherboard failures. Background Art
[0002] With the rapid development of technology, the technology of computer hardware is constantly updated, especially in the design and functions of motherboards. As one of the core components of a computer, the motherboard undertakes important tasks, including connecting various hardware modules, processors, memories, storage devices, etc. The power supply system of the motherboard, especially the stability of the power rails, is crucial for the performance and operation stability of the entire computer. Unstable or faulty power rails often lead to system crashes, performance degradation, or long-term failure shutdowns. In this large field, computer hardware fault monitoring and diagnosis technology is becoming an important means to improve equipment reliability and reduce system failures. In particular, the diagnostic method for motherboard faults has become the focus of industry attention.
[0003] In the Chinese invention application with the application number CN202210905574.8, a baseboard management circuit, method, computer device, and motherboard control system are disclosed, which relate to a baseboard management circuit, method, computer device, and motherboard control system. Through a power management module respectively connected to the motherboard power supply and the backup power supply, when the current power supply mode of the baseboard management circuit is the motherboard power supply and the motherboard power supply is abnormal, it switches to the backup power supply for power supply; a control module respectively connected to the motherboard, the motherboard power supply, and the power management module, when the power management module switches to the backup power supply for power supply due to the abnormal motherboard power supply, generates a fault correction instruction according to the operating state of the motherboard and / or the motherboard power supply, so that the baseboard management circuit is in a preset working state; realizes the normal power supply of the baseboard management circuit when the motherboard power supply is abnormal, thereby ensuring the monitoring of the operating states of the motherboard and the motherboard power supply and the fault troubleshooting, and improving the operating stability and reliability of the motherboard and the motherboard power supply.
[0004] Combined with the existing technology, the above application still has the following deficiencies:
[0005] Although this solution can ensure the continuity of power supply through the backup power supply switching, relying on the performance of the backup power supply remains a potential problem for this solution. The capacity and reliability of the backup power supply directly affect whether the system can operate stably for a long time, and the process of backup power supply switching may lead to short-term system instability or performance degradation, which is consistent with the potential risk of abnormal power rail fluctuation response ability. In addition, the generation of fault correction instructions also depends on the accurate detection and analysis of the power supply state. However, if the dynamic changes of the power rail cannot be accurately captured or misdiagnosed, it may result in ineffective or incorrect fault repair instructions, thus affecting the effect of system recovery. The power supply monitoring in the existing solutions mainly focuses on the switching between the main power supply and the backup power supply, and does not deeply monitor the dynamic behavior of the power rail, such as voltage changes, frequency regulation, and virtual power rail problems. This makes it may not be able to give early warnings in time when facing potential hidden faults such as power rail fluctuations and virtual power rails. In addition, the solution lacks a global power supply health assessment mechanism and cannot accurately judge the health status of the entire power supply system through comprehensive power supply behavior analysis, resulting in difficulty in providing comprehensive fault diagnosis and prevention measures in a complex power supply environment. In summary, although this invention provides an effective solution for power supply fault switching and repair, introducing a more accurate power rail monitoring, dynamic behavior analysis, and comprehensive health assessment mechanism will help improve the stability and fault warning ability of the system and further enhance the long-term reliability of the device. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the present invention provides a method and system for monitoring and diagnosing computer motherboard faults, which solves the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: including the following steps:
[0008] S1. Real-time collect the power rail data of multiple power rails of the motherboard through the motherboard BIOS interface, and transmit the power rail data to the motherboard fault monitoring system, and preprocess the power rail data to obtain a standardized virtual rail data set;
[0009] S2. In the motherboard fault monitoring system, construct a virtual rail index model, extract the standardized virtual rail data set and input it into the virtual rail index model, calculate and output the voltage rail virtualization index VEI, and based on the output result of the voltage rail virtualization index VEI, conduct a preliminary comparison and evaluation to judge the virtual power rail situation of the motherboard;
[0010] When the result of the preliminary comparison and evaluation is determined to be a virtual power rail, trigger a lightweight response mechanism. The lightweight response mechanism collects the performance data of the corresponding motherboard module by starting the BIOS interface, calculates and outputs the lightweight behavior risk index LBPI based on the performance data, and conducts a lightweight evaluation and analysis based on the output result of the lightweight behavior risk index LBPI;
[0011] S4. Trigger the abnormal energy absorption behavior analysis based on the lightweight evaluation analysis results, and perform calculations to output the abnormal energy absorption index EAAI and the module self-healing ability index REI of the main board module;
[0012] S5. Extract the abnormal energy absorption index EAAI and the module self-healing ability index REI, perform summary calculations to output the power supply health assessment index PHS, and conduct comprehensive comparison and evaluation based on the output results of the power supply health assessment index PHS to analyze the power supply stability of the computer main board.
[0013] Preferably, the S1 includes S11 and S12;
[0014] S11. Install sensors on different power rails of the computer main board, and set the acquisition frequency to 100 ms through the computer main board BIOS interface ACPI to collect power rail data in real time;
[0015] The power rails include +12V main voltage, +5V system voltage, +3.3V / IO voltage, VCORE core power supply, VDD auxiliary rail, and VPP auxiliary rail;
[0016] The sensors include a micro voltage sensor and an oscilloscope;
[0017] The power rail data includes voltage data and response time t;
[0018] S12. Install a main board fault monitoring system in the central processing unit of the computer main board, directly transmit the collected power rail data to the main board fault detection system, and preprocess the power rail data in the main board fault monitoring system to obtain a standardized virtual rail data set;
[0019] The preprocessing includes data processing, timestamp alignment, and normalization processing;
[0020] The timestamp alignment is used to unify the power rail data after data processing to a unified time reference based on the starting time point of the same acquisition frequency. For power rail data with different acquisition frequencies, interpolation is used for unification;
[0021] The data processing is performed by conducting steady-state comparison, voltage change analysis, track signal response duration processing, and spectrum analysis on the voltage data of the power rail data after timestamp alignment, and extracting the power rail virtualization data after data processing;
[0022] The power rail virtualization data includes steady-state voltage V0, voltage change rate △V, track signal response duration tr, and voltage jitter frequency F;
[0023] The steady-state voltage V0 is obtained through steady-state comparison;
[0024] The voltage change rate △V is obtained through voltage change analysis;
[0025] The track signal response duration tr is obtained through track signal response duration processing;
[0026] The voltage jitter frequency F is obtained through spectrum analysis, and voltage fluctuation data is used for spectrum analysis;
[0027] The normalization process uses the Z-Score normalization method to convert the power rail virtualization data into a standardized virtual rail data set with a standard normal distribution, eliminating the influence of the dimension of the power rail virtualization data.
[0028] Preferably, S2 includes S21 and S22;
[0029] S21. In the motherboard fault monitoring system, construct a virtual rail index model, extract the preprocessed standardized virtual rail data set, input it into the virtual rail index model, calculate and output the voltage rail virtualization index VEI, and analyze the health status of the power rail;
[0030] The voltage rail virtualization index VEI is calculated and output through the following virtual rail index model;
[0031] ;
[0032] In the formula, VEI i represents the voltage rail virtualization index of the i-th power rail, Vnominal i represents the standard steady-state voltage reference value of the i-th power rail, △V i represents the voltage change rate of the i-th power rail, F i represents the voltage jitter frequency of the i-th power rail, Fref represents the standardized voltage jitter frequency reference value, tri represents the track signal response duration of the i-th power rail, represents the balance factor.
[0033] Preferably, S22. Based on the results output by the voltage rail virtualization index VEI of each power rail, conduct a preliminary comparative evaluation to analyze the virtualization situation of the power rails of the computer motherboard. The specific evaluation content is as follows;
[0034] When the voltage rail virtualization index VEI of the i-th power rail i ≥1, it indicates that there is a virtual power rail in the power rail, and at this time, a lightweight response mechanism is triggered;
[0035] When the voltage rail virtualization index VEI of the i-th power rail i <1, it indicates that the power rail is normal, and the current monitoring is maintained.
[0036] Preferably, S3 includes S31 and S32;
[0037] S31. After triggering the lightweight response mechanism through preliminary comparison and evaluation, start real-time collection of the performance data of the motherboard module with virtual power rails on the BIOS interface, and perform normalization processing on the performance data to eliminate the dimensional influence between parameters in the performance data;
[0038] The performance data includes the abnormal restart rate Rrestart, the performance decay rate ΔPerf under voltage micro-offset, and the frequency modulation behavior frequency Fdvfs;
[0039] Based on the performance data, calculate and output the lightweight behavior risk index LBPI to comprehensively quantify the abnormal behaviors of micro-restart times, performance degradation, virtual rail offset, and frequency adjustment of different modules of the computer motherboard;
[0040] The lightweight behavior risk index LBPI is calculated and output through the following algorithm formula;
[0041] ;
[0042] In the formula, LBPI j represents the lightweight behavior risk index of the j-th module, Rrestart j represents the abnormal restart rate of the j-th module, T represents the time window length, ΔPerf j represents the performance decay rate of the j-th module under voltage micro-offset, Fdvfs j represents the frequency modulation behavior frequency of the j-th module, and a1, a2, and a3 are the preset weight values of the abnormal restart rate Rrestart, the performance decay rate ΔPerf under voltage micro-offset, and the frequency modulation behavior frequency Fdvfs. The specific values are set by the user, and a1 + a2 + a3 = 1;
[0043] S32. Based on the output result of the lightweight behavior risk index LBPI, perform lightweight evaluation and analysis to judge the abnormal behavior of the computer motherboard module in the case of virtual power rails. The specific evaluation content is as follows;
[0044] When the lightweight behavior risk index LBPI of the j-th module j < 0.3, it means that the current computer motherboard module behaves normally, and continue to monitor at this time;
[0045] When 0.3 ≤ the lightweight behavior risk index LBPI of the j-th module j ≤ 0.7, it means that the current computer motherboard module has a sub-healthy trend in behavior. At this time, generate the first prompt warning message to prompt that the current j-th module is sub-healthy, and adjust the collection frequency to 50ms;
[0046] When the lightweight behavior risk index LBPI of the j-th module jWhen it is greater than 0.7, it indicates that the behavior of the current computer motherboard module in response to power rail fluctuations is abnormal, and at this time, the analysis of abnormal energy absorption behavior is triggered.
[0047] Preferably, S4 includes S41 and S42;
[0048] S41. After the power rail fluctuation response ability at the lightweight evaluation is abnormal, start the BIOS interface to collect the energy Pabnormal absorbed during fluctuations, the power consumption Pexpected under normal conditions, and the recovery curve stabilization time Tstable of all modules of the computer motherboard, and perform normalization processing and then execute the abnormal energy absorption behavior analysis. The abnormal energy absorption behavior analysis includes abnormal energy absorption analysis and module self-healing ability analysis;
[0049] The abnormal energy absorption analysis extracts the energy Pabnormal absorbed during fluctuations, the power consumption Pexpected under normal conditions, and the recovery curve stabilization time Tstable for joint calculation to output the abnormal energy absorption index EAAI, and analyzes the abnormal degree of the module in terms of energy absorption;
[0050] The abnormal energy absorption index EAAI is calculated and output through the following algorithm formula;
[0051] ;
[0052] In the formula, EAAI j represents the abnormal energy absorption index of the jth module, Pabnormal j (t) represents the energy absorbed by the jth module during fluctuations at time t, Tstable j (t) represents the recovery curve stabilization time of the jth module at time t, Pexpected j represents the power consumption of the jth module under normal conditions, and Tavg represents the average recovery time of all modules.
[0053] Preferably, S42. After the abnormal energy absorption analysis, perform the module self-healing ability analysis. The module self-healing ability analysis is based on the recovery curve stabilization time Tstable of all computer motherboard modules, and performs an averaging calculation to obtain the module self-healing ability index REI, and analyzes the self-healing repair ability of the module;
[0054] The module self-healing ability index REI is calculated and output through the following algorithm formula;
[0055] ;
[0056] In the formula, REIj represents the module self-healing ability index of the jth module, N represents the number of times of self-healing ability evaluation, Tstable j(k) represents the stabilization time of the recovery curve during the k-th recovery process of the j-th module.
[0057] Preferably, the S5 includes S51 and S52;
[0058] S51. Based on the abnormal energy absorption index EAAI and the module self-healing ability index REI output from the analysis of abnormal energy absorption behavior, perform a summary calculation to output the power supply health assessment index PHS of all computer motherboard modules, and quantify the power supply health of the motherboard;
[0059] The power supply health assessment index PHS is calculated and output through the following algorithm formula;
[0060] ;
[0061] In the formula, M represents the total number of modules, Var(Pabnormal) represents the variance of the energy absorbed during fluctuations, which is used to measure the changes in power supply fluctuations over different time periods, 、 and represent the preset weight values of the abnormal energy absorption index EAAI, the module self-healing ability index REI, and the variance of the energy absorbed during fluctuations. The specific values are set by the user, and + + = 1.
[0062] Preferably, S52. Based on the output result of the power supply health assessment index PHS, perform a comprehensive comparison and evaluation, analyze the power supply health of the overall modules of the computer motherboard, and based on the comprehensive comparison and evaluation result, trigger different response measures. The specific evaluation content is as follows;
[0063] When the power supply health assessment index PHS < 0.5, it means that the power supply is stable in the case of abnormal power rail fluctuations, and observation continues at this time;
[0064] When the power supply health assessment index PHS ≥ 0.5, it means that there is an abnormality in the computer motherboard power supply in the case of abnormal power rail fluctuations. At this time, a prompt warning is generated through the motherboard fault monitoring system, and the power supply group components of the computer motherboard are replaced.
[0065] A monitoring and diagnostic system for computer motherboard faults includes a multi-power rail monitoring module, a virtual power rail identification module, a lightweight module, an abnormal behavior analysis module, and a comprehensive power supply analysis module;
[0066] The multi-power rail monitoring module collects the power rail data of the motherboard multi-power rail in real time through the motherboard BIOS interface, and transmits the power rail data to the motherboard fault monitoring system, and preprocesses the power rail data to obtain a standardized virtual rail data set;
[0067] The virtual power rail identification module constructs a virtual rail index model in the mainboard fault monitoring system, extracts a standardized virtual rail data set and inputs it into the virtual rail index model to calculate and output the voltage rail virtualization index VEI. Based on the output result of the voltage rail virtualization index VEI, a preliminary comparison and evaluation is carried out to judge the situation of the virtual power rail of the mainboard;
[0068] When the lightweight module determines that it is a virtual power rail through the preliminary comparison and evaluation result, it triggers a lightweight response mechanism. The lightweight response mechanism collects the performance data of the corresponding mainboard module by starting the BIOS interface, calculates and outputs the lightweight behavior risk index LBPI based on the performance data, and conducts a lightweight evaluation and analysis based on the output result of the lightweight behavior risk index LBPI;
[0069] The abnormal behavior analysis module triggers the analysis of abnormal energy absorption behavior based on the lightweight evaluation and analysis result, calculates and outputs the abnormal energy absorption index EAAI and the module self-healing ability index REI of the mainboard module;
[0070] The comprehensive power supply analysis module extracts the abnormal energy absorption index EAAI and the module self-healing ability index REI, sums them up and calculates and outputs the power supply health evaluation index PHS, and conducts a comprehensive comparison and evaluation based on the output result of the power supply health evaluation index PHS to analyze the power supply stability of the computer mainboard.
[0071] The present invention provides a method and system for monitoring and diagnosing computer mainboard faults. It has the following beneficial effects:
[0072] (1) This method collects data of multiple power rails in real time, and uses the mainboard BIOS interface and high-precision sensors for data transmission and preprocessing. This process can comprehensively analyze the power behaviors of multiple dimensions such as the steady-state voltage, voltage change rate, voltage response duration, and voltage jitter frequency of the power rail. After calculation by the virtual rail index model, it can accurately identify potential faults such as power rail virtualization, load change, and power instability. After optimization, this method not only improves the accuracy of fault diagnosis, but also can identify hidden faults of the power rail in advance, greatly enhancing the early warning ability of the system. Through real-time monitoring and data analysis, it can effectively avoid equipment failures caused by power rail problems and improve the stability and safety of equipment operation.
[0073] (2) Through the analysis based on the lightweight behavior risk index LBPI and abnormal energy absorption behavior, the fault monitoring and diagnosis method for computer motherboards can deeply analyze the performance of different motherboard modules under power fluctuations, such as the abnormal restart rate, the performance decay rate of voltage micro-offset, and the frequency modulation behavior frequency. By calculating the abnormal energy absorption index EAAI and the module self-healing ability index REI of the module, the health status of each module is further evaluated, and a quantitative score of the overall power health of the device is provided through the power supply health assessment index PHS. After optimization, this method can specifically identify and warn of potential problems in the power supply system, thereby reducing unnecessary maintenance and replacement costs of the device and improving the accuracy of maintenance decisions. An effective fault warning mechanism can avoid unnecessary downtime and repairs, thereby reducing the overall maintenance cost and device downtime and improving the operation and maintenance efficiency of the system.
[0074] (3) Through the multi-level evaluation by introducing the voltage rail virtualization index VEI, the lightweight behavior risk index LBPI, and the power supply health assessment index PHS, this method can comprehensively analyze and optimize the health status of the motherboard power supply system to ensure that the device can maintain stable power support during long-term operation. By monitoring the virtualization, energy absorption, and repair ability of the power rail, power problems can be identified and repaired in a timely manner to prevent system crashes or performance degradation caused by unstable power rails. After optimization, it can evaluate the dynamic behavior of the power rail and each module in real time, predict and avoid the occurrence of potential faults, improve the long-term reliability and stability of the device, and reduce system downtime caused by unexpected faults. In addition, accurate power health assessment can also reduce over-maintenance and unnecessary intervention during operation, making the performance of the device more stable during long-term operation. Brief Description of the Drawings
[0075] Figure 1 Schematic diagram of the steps of a method for monitoring and diagnosing faults in a computer motherboard according to the present invention;
[0076] Figure 2 Schematic diagram of the process of a system for monitoring and diagnosing faults in a computer motherboard according to the present invention;
[0077] Figure 3 Schematic diagram of the evaluation process of a method for monitoring and diagnosing faults in a computer motherboard according to the present invention;
[0078] Figure 4 Curve diagram of the voltage rail virtualization index VEI over time. Detailed Embodiments
[0079] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0080] Embodiment 1
[0081] Please refer to Figure 1 、 Figure 3 and Figure 4 The present invention provides a method for monitoring and diagnosing computer motherboard failures. To achieve the above objectives, the present invention is implemented through the following technical solutions: including the following steps:
[0082] S1. Real-time collect the power rail data of multiple power rails of the motherboard through the motherboard BIOS interface, and transmit the power rail data to the motherboard failure monitoring system, and preprocess the power rail data to obtain a standardized virtual rail data set;
[0083] S2. In the motherboard failure monitoring system, construct a virtual rail index model, extract the standardized virtual rail data set and input it into the virtual rail index model, calculate and output the voltage rail virtualization index VEI, and based on the output result of the voltage rail virtualization index VEI, conduct a preliminary comparison and evaluation to judge the virtual power rail situation of the motherboard;
[0084] S3. When the preliminary comparison and evaluation result is determined to be a virtual power rail, trigger a lightweight response mechanism. The lightweight response mechanism collects the performance data of the corresponding motherboard module by starting the BIOS interface, calculates and outputs the lightweight behavior risk index LBPI based on the performance data, and conducts a lightweight evaluation and analysis based on the output result of the lightweight behavior risk index LBPI;
[0085] S4. Based on the lightweight evaluation and analysis result, trigger an abnormal energy absorption behavior analysis, calculate and output the abnormal energy absorption index EAAI and the module self-healing ability index REI of the motherboard module;
[0086] S5. Extract the abnormal energy absorption index EAAI and the module self-healing ability index REI, conduct a summary calculation and output the power supply health assessment index PHS, and conduct a comprehensive comparison and evaluation based on the output result of the power supply health assessment index PHS to analyze the power supply stability of the computer motherboard.
[0087] In this embodiment, the method realizes the accurate assessment of the health status of the motherboard power supply system through a series of multi-level analyses and calculations, combined with power rail data acquisition, virtual rail index model, lightweight behavior risk assessment, and abnormal energy absorption analysis. The specific implementation method includes collecting data of multiple power rails in real time through the motherboard BIOS interface, and preprocessing and normalizing the collected data to accurately reflect the health status of the power rails in subsequent analyses. By constructing a virtual rail index model VEI, the virtualization of the power rails can be detected, and the virtualization index can be calculated and output to identify potential power rail fault problems. Further, a lightweight response mechanism can calculate the lightweight behavior risk index LBPI through performance data to conduct risk assessments on different modules of the motherboard, so as to accurately identify abnormal behaviors such as performance degradation, virtual rail offset, and frequency adjustment. Based on these assessment results, abnormal energy absorption analysis and module self-healing ability analysis can also be triggered, and the abnormal energy absorption index EAAI and self-healing ability index REI are calculated and output, and then the power supply health assessment index PHS is obtained. Finally, through the comprehensive assessment of the power supply health assessment index PHS, the stability of the power rails and the health status of the power supply system can be clearly identified. The implementation of this method achieves the purpose of accurate diagnosis and real-time warning. Through multi-dimensional index monitoring and analysis, it can not only identify potential faults and performance degradation of the power rails at an early stage, but also predict and prevent equipment crashes or significant performance degradation caused by unstable power supply systems in a timely manner. After optimization, this method greatly improves the accuracy and efficiency of fault detection, identifies power problems in advance, effectively reduces maintenance costs and downtime, and improves the long-term stability and operation reliability of the equipment. In addition, based on the comprehensive comparative assessment of the power supply health assessment index PHS, maintenance personnel can take targeted measures according to specific risk assessment results, avoiding unnecessary maintenance interventions, and further improving the maintainability and overall operation efficiency of the equipment.
[0088] Embodiment 2
[0089] Please refer to Figure 1 , specifically: S1 includes S11 and S12;
[0090] S11. Install sensors on different power rails of the computer motherboard, and set the acquisition frequency to 100 ms through the computer motherboard BIOS interface ACPI to collect power rail data in real time;
[0091] The power rails include +12V main voltage, +5V system voltage, +3.3V / IO voltage, VCORE core power supply, VDD auxiliary rail, and VPP auxiliary rail;
[0092] The sensors include micro voltage sensors and oscilloscopes;
[0093] The power rail data includes voltage data and response time t;
[0094] S12. Install a motherboard fault monitoring system in the central processing unit of the computer motherboard, directly transmit the collected power rail data to the motherboard fault detection system, and preprocess the power rail data in the motherboard fault monitoring system to obtain a standardized virtual rail dataset;
[0095] The preprocessing includes data processing, timestamp alignment, and normalization processing;
[0096] Timestamp alignment is used to unify the power rail data after data processing to a unified time base based on the starting time point of the same acquisition frequency. For power rail data with different acquisition frequencies, interpolation is used for unification;
[0097] Data processing is performed on the voltage data of the power rail data after timestamp alignment through steady-state comparison, voltage change analysis, track signal response duration processing, and spectrum analysis. After data processing, the power rail virtualization data is extracted;
[0098] The power rail virtualization data includes steady-state voltage V0, voltage change rate △V, track signal response duration tr, and voltage jitter frequency F;
[0099] The steady-state voltage V0 is obtained through steady-state comparison. By installing a high-precision voltage sensor on each power rail to monitor its voltage level in real time, the steady-state voltage can be measured by the sensor to obtain the voltage value predetermined during motherboard design, and it is used to compare with the real-time data to check for deviations. For example, for the +12V rail, the theoretical regulated voltage value should be 12V, and for the +5V rail, the steady-state voltage should be 5V, and so on;
[0100] The voltage change rate △V is obtained through voltage change analysis. By quickly sampling to capture the instantaneous change of voltage data during load jump, for example, using a quickly sampled voltage sensor to capture the voltage change of the power rail during load jump and calculate the voltage change rate △V. The magnitude of the voltage change rate △V can reflect the stability of the power rail. Excessive fluctuations may indicate that the power rail cannot provide stable voltage or there are potential faults;
[0101] The track signal response duration tr is obtained through track signal response duration processing. Using a high-frequency response voltage sensor to monitor the voltage change of the power rail in real time. When the load of the voltage rail changes, record the time required for the power rail voltage to start changing until it reaches the steady-state voltage V0. Use an oscilloscope or a fast data acquisition card to accurately measure this response time. Usually, if the voltage fluctuation cannot be stabilized for too long, it means that the response ability of the power rail is poor;
[0102] The voltage jitter frequency F is obtained through spectrum analysis. The voltage fluctuation data is analyzed using spectrum analysis to analyze the frequency components of the voltage fluctuations. Through the Fourier transform FFT, the time-domain signal is converted into a frequency-domain signal to obtain the distribution of the jitter frequency. Generally, the jitter frequency of the power rail should be in a relatively low and stable range. An excessively high jitter frequency may indicate poor power quality or other problems;
[0103] The normalization process uses the Z-Score standardization method to convert the power rail virtualization data into a standardized virtual rail data set with a standard normal distribution, eliminating the influence of the dimension of the power rail virtualization data.
[0104] In this embodiment, the method installs sensors on different power rails of the computer motherboard and collects power rail data in real time through the motherboard BIOS interface ACPI to ensure multi-dimensional monitoring of the motherboard power supply system. Specifically, a micro voltage sensor and an oscilloscope are used to monitor the data of multiple power rails, including voltage values and response times. Through the real-time collected data, the motherboard fault monitoring system performs preprocessing, including data processing, timestamp alignment, and normalization processing, thereby eliminating the differences in data acquisition frequencies and dimensions of different power rails. The processed data is subjected to steady-state comparison, voltage change analysis, track signal response duration, and spectrum analysis to extract the virtualization data of the power rail, helping to accurately diagnose the health status of the power rail. This method provides a high-precision fault diagnosis mechanism through a standardized virtual rail data set, can monitor the stability of the power rail in real time, and quickly identify potential virtual power rails and power instability problems. After optimization, the method improves the health assessment ability of the power supply system. Through detailed voltage fluctuation analysis, response duration monitoring, and spectrum analysis, it can early warn of power supply faults and ensure the long-term stability and reliability of the device. At the same time, the normalization processing of the data ensures the comparability between different power rails, improves the diagnostic accuracy, and avoids the common deviation problems in traditional power supply monitoring methods.
[0105] Embodiment 3
[0106] Please refer to Figure 3 and Figure 4 , specifically: S2 includes S21 and S22;
[0107] S21. In the motherboard fault monitoring system, construct a virtual rail index model, extract the preprocessed standardized virtual rail data set, input it into the virtual rail index model, calculate and output the voltage rail virtualization index VEI, and analyze the health status of the power rail;
[0108] The voltage rail virtualization index VEI is calculated and output through the following virtual rail index model;
[0109] ;
[0110] where, VEI i represents the voltage rail virtualization index of the i-th power rail, Vnominal i represents the standard steady-state voltage reference value of the i-th power rail, △V i represents the voltage change rate of the i-th power rail, F i represents the voltage jitter frequency of the i-th power rail, Fref represents the standardized voltage jitter frequency reference value, and tri represents the track signal response duration of the i-th power rail, represents the balance factor, which is used to adjust the influence of the response delay on the voltage rail virtualization index VEI according to the working environment and technical requirements of the power rail. The introduction of this factor helps to make the calculation more accurate under dynamic voltage changes.
[0111] S22. Based on the results output by the voltage rail virtualization index VEI of each power rail, perform a preliminary comparative evaluation to analyze the virtualization of the power rails on the computer motherboard. The specific evaluation content is as follows;
[0112] When the voltage rail virtualization index VEI of the i-th power rail i ≥1, it indicates that there is a virtual power rail in the power rail, and at this time, a lightweight response mechanism is triggered;
[0113] When the voltage rail virtualization index VEI of the i-th power rail i <1, it indicates that the power rail is normal, and the current monitoring is maintained.
[0114] In this embodiment, the method comprehensively evaluates the health status of the main board power rail by constructing a voltage rail virtualization index (VEI) model and using a standardized virtual rail dataset. During the specific implementation process, the main board fault monitoring system inputs the preprocessed data into the virtual rail index model for calculation, and outputs the virtualization index VEI of each power rail. The calculation of this index comprehensively considers multiple-dimensional factors in the power rail virtualization data, and can comprehensively evaluate the stability and response ability of the power rail. Through this multi-dimensional evaluation, power rails that seemingly appear normal on the surface but actually have potential faults can be identified in a timely manner, avoiding system failures caused by power problems. Further, by preliminarily comparing the evaluation results, if the virtualization index VEI of a certain power rail is greater than or equal to 1, it indicates that there may be a problem with the virtual power rail of this power rail, and the system will trigger a lightweight response mechanism to further analyze by real-time collecting the performance data of the corresponding main board module. This method effectively avoids failures caused by abnormal power rails, and at the same time has a predictive maintenance function, which can give early warnings when latent faults occur in the power system, and is particularly suitable for devices running for a long time, such as unmanned devices and vehicle-mounted devices. After optimization, it can improve the fault detection accuracy of the system, reduce the downtime caused by power problems, and at the same time enhance the reliability and maintainability of the device. At the same time, the overall physical meaning of the voltage rail virtualization index VEI formula is that the voltage rail virtualization index VEI evaluates the stability, response ability and its long-term working condition of the power rail by quantifying the behavior of the power rail in a dynamic environment. Its core goal is to identify in advance those power rails that seemingly appear normal on the surface but actually have potential faults, avoiding system failures due to power problems; the voltage rail virtualization index VEI can comprehensively reflect the health status of the power rail by considering multiple-dimensional factors such as voltage fluctuations, frequency jitters and response delays. This multi-dimensional evaluation method avoids potential problems that may be ignored by a single indicator, such as only using voltage values or voltage changes, improving the accuracy of fault detection; the design of the voltage rail virtualization index VEI not only helps to detect existing power rail problems, but also can predict potential problems that may occur in the future, especially those latent faults that cannot be discovered by conventional monitoring methods, such as structural virtual power rails. This method is particularly suitable for devices running for a long time, and can improve the maintainability and reliability of the device; the voltage rail virtualization index VEI formula can be widely applied to the power systems of various electronic devices, especially in environments such as industrial control, unmanned devices, and vehicle-mounted devices that require long-term stable operation. By real-time monitoring the voltage rail virtualization index VEI, the device can timely discover potential problems with the power rail and take measures such as repair or replacement to avoid major impacts on system operation caused by faults.
[0115] Embodiment 4
[0116] Please refer to Figure 1 and Figure 3 , specifically: S3 includes S31 and S32;
[0117] S31. After triggering the lightweight response mechanism through preliminary comparison and evaluation, start real-time collection of the performance data of the motherboard module with virtual power rails on the BIOS interface, and perform normalization processing on the performance data to eliminate the dimensional influence between parameters in the performance data;
[0118] The performance data includes the abnormal restart rate Rrestart, the performance decay rate ΔPerf under voltage micro-offset, and the frequency modulation behavior frequency Fdvfs;
[0119] Based on the performance data, calculate and output the lightweight behavior risk index LBPI to comprehensively quantify the behavioral anomalies of micro-restart times, performance degradation, virtual rail offset, and frequency adjustment of different modules of the computer motherboard;
[0120] The lightweight behavior risk index LBPI is calculated and output through the following algorithm formula;
[0121] ;
[0122] In the formula, LBPI j represents the lightweight behavior risk index of the j-th module, Rrestart j represents the abnormal restart rate of the j-th module, T represents the time window length, ΔPerf j represents the performance decay rate of the j-th module under voltage micro-offset, Fdvfs j represents the frequency modulation behavior frequency of the j-th module, and a1, a2, and a3 are the preset weight values of the abnormal restart rate Rrestart, the performance decay rate ΔPerf under voltage micro-offset, and the frequency modulation behavior frequency Fdvfs. The specific values are set by the user, and a1 + a2 + a3 = 1;
[0123] represents the abnormal restart rate that occurs within the specified time window length segment of the j-th module. Through this ratio, the stability of the module can be quantified. If the number of micro-restarts is high, it indicates that the power rail, load changes, or other system parameters may cause the module to be unstable;
[0124] represents the relationship between voltage fluctuations and module performance degradation, quantifying the impact of unstable power rail voltage on module performance. The greater the voltage fluctuation of the power rail and the proportional relationship with the performance degradation amplitude indicate that the power system cannot effectively support the operation of the module, which may lead to device performance degradation or failure;
[0125] S32. Based on the output result of the lightweight behavior risk index LBPI, perform lightweight evaluation and analysis to judge the behavioral anomalies of the computer motherboard module in the case of virtual power rails. The specific evaluation content is as follows;
[0126] When the lightweight behavior risk index LBPI of the j-th module j < 0.3, it indicates that the behavior of the current computer motherboard module is normal and has sufficient flexibility. At this time, continue to monitor;
[0127] When 0.3 ≤ the lightweight behavior risk index LBPI of the j-th module j ≤ 0.7, it indicates that there is a sub-healthy trend in the behavior of the current computer motherboard module. At this time, generate the first prompt warning message to prompt that the current j-th module is sub-healthy, and adjust the acquisition frequency to 50 ms;
[0128] When the lightweight behavior risk index LBPI of the j-th module j > 0.7, it indicates that the response ability of the current computer motherboard module to the power rail fluctuation is abnormal. At this time, trigger the abnormal energy absorption behavior analysis, indicating that the response ability of the computer motherboard module to the power rail fluctuation is weakened, and a failure or a significant performance decline may occur. The device should enter the abnormal energy absorption behavior analysis stage to perform more detailed analysis and processing on the power rail and system behavior.
[0129] In this embodiment, the method triggers a lightweight response mechanism after preliminary comparative evaluation, achieving precise monitoring and risk assessment of the motherboard module. First, the motherboard fault monitoring system collects the performance data of the module with virtual power rails in real time through the BIOS interface and performs normalization processing to eliminate the dimensional influence between different parameters. Based on the module performance data, a lightweight behavior risk index LBPI is calculated and output. This index comprehensively evaluates multiple factors such as the number of micro restarts, performance degradation, virtual rail offset, and frequency adjustment of the motherboard module, and can accurately identify potential faults caused by problems such as unstable power rails and virtual power rails. By performing lightweight evaluation and analysis on the output result of LBPI, the system can judge the health status of the module according to different lightweight behavior risk index LBPI values. This process can not only identify power rail problems in a timely manner but also avoid equipment failures caused by unstable power supplies. The implementation of this method achieves the purpose of precise monitoring and early warning, and is particularly suitable for long-term running equipment. It can diagnose in a timely manner when the power rail is unstable, avoiding frequent restarts or performance degradation of the equipment due to power problems. At the same time, the physical meaning of the formula is that the lightweight behavior risk index LBPI formula comprehensively considers various dynamic behaviors such as the number of micro restarts, the degree of performance degradation, the degree of virtual rail offset, and the frequency of frequency adjustment of the module. Through the weighted evaluation of these parameters, a comprehensive risk index is given. This risk index can reflect whether there are potential power or performance problems in the module, especially those hidden faults that may be caused by problems such as unstable power rails and virtual power rails. The lightweight behavior risk index LBPI is a lightweight risk assessment tool that can be applied to various equipment monitoring scenarios, especially in environments with unstable power quality or frequent load fluctuations. For example, in unattended retail equipment, such long-term online running equipment, if the power rail is unstable, it may lead to frequent micro restarts or performance degradation. By monitoring the health status of these equipment through the lightweight behavior risk index LBPI, power problems can be identified in advance to avoid equipment failures, and at the same time, the overall computing load of the equipment is indirectly reduced.
[0130] Embodiment 5
[0131] Please refer to Figure 1 and Figure 3 , specifically: S4 includes S41 and S42;
[0132] S41. After the power rail fluctuation response ability at the lightweight evaluation is abnormal, start the BIOS interface to collect the energy Pabnormal absorbed during the fluctuation of all modules of the computer motherboard, the power consumption Pexpected under normal conditions, and the stable time Tstable of the recovery curve, and perform normalization processing and then execute abnormal energy absorption behavior analysis. The abnormal energy absorption behavior analysis includes abnormal energy absorption analysis and module self-healing ability analysis;
[0133] The abnormal energy absorption analysis calculates and outputs the abnormal energy absorption index EAAI by jointly extracting the energy Pabnormal absorbed during fluctuations, the power consumption Pexpected under normal conditions, and the recovery curve stabilization time Tstable, and analyzes the abnormal degree of the module in terms of energy absorption;
[0134] The abnormal energy absorption index EAAI is calculated and output through the following algorithm formula;
[0135] ;
[0136] In the formula, EAAI j represents the abnormal energy absorption index of the j-th module, and Pabnormal j (t) represents the energy absorbed during fluctuations of the j-th module at time t, and Tstable j (t) represents the recovery curve stabilization time of the j-th module at time t, and Pexpected j represents the power consumption of the j-th module under normal conditions, and Tavg represents the average recovery time of the entire module;
[0137] The physical meaning of the formula is to quantify the ratio between the extra energy consumed by the module under power fluctuations or unstable power supply conditions and the energy under normal conditions. A high abnormal energy absorption index EAAI indicates that the module has poor adaptability to power fluctuations, which may mean that there is instability in the power supply system or power rail, the electrical behavior of the module is unstable, which may lead to failures or performance degradation. This formula helps to identify the impact of power instability on the device and provides a basis for fault prediction by evaluating the relationship between the abnormal energy absorption amount, the recovery time, and the standard power consumption.
[0138] S42. After the abnormal energy absorption analysis, perform the module self-healing ability analysis. The module self-healing ability analysis is based on the recovery curve stabilization time Tstable of all computer motherboard modules, and performs an averaging calculation to obtain the module self-healing ability index REI and analyze the self-healing and repair ability of the module;
[0139] The module self-healing ability index REI is calculated and output through the following algorithm formula;
[0140] ;
[0141] In the formula, REIj represents the module self-healing ability index of the j-th module, N represents the number of self-healing ability evaluations, usually multiple measurements are carried out within a specific time window, and Tstable j (k) represents the recovery curve stabilization time of the j-th module during the k-th recovery process;
[0142] The physical meaning of the formula is that the module self-healing ability index REI quantifies the self-healing ability of the module by calculating the average time required for the module to return to a stable state. A higher module self-healing ability index REI indicates that the module can return to the normal working state more quickly and has a stronger self-healing ability; while a lower module self-healing ability index REI indicates that the module has a longer recovery time, which may be due to reasons such as an unstable power supply system or hardware failure and cannot recover quickly. A decrease in the module self-healing ability index REI may mean that the device is increasingly unable to maintain stability under power or load changes, increasing the risk of failure.
[0143] In this embodiment, when the power rail fluctuation response ability is abnormal, the method starts the BIOS interface to collect the energy Pabnormal absorbed during fluctuations, the power consumption Pexpected under normal conditions, and the recovery curve stable time Tstable of all modules on the motherboard in real time, and performs abnormal energy absorption behavior analysis after normalizing these data. This process quantifies the difference between the additional energy consumed by the module under power fluctuations or instability and the normal power consumption by calculating the abnormal energy absorption index EAAI. Through this analysis, it can be judged whether the electrical behavior of the module is stable and potential problems of the power supply system or power rail instability can be identified, providing a basis for fault prediction in a timely manner. In addition, after the abnormal energy absorption analysis, the system further performs module self-healing ability analysis, and calculates the module self-healing ability index REI by evaluating the recovery time Tstable of the module. This index reflects the recovery speed of the module after power fluctuations or load changes. A higher REI indicates that the module has a stronger self-healing ability and can return to the normal working state faster, while a lower REI indicates a decrease in the module's recovery ability, which may lead to an increase in the risk of failure. Through this implementation method, the impact of power fluctuations on the motherboard module can be comprehensively evaluated, the abnormal energy absorption and repair ability can be quantified, and a precise health assessment of the power supply system can be provided. After optimization, this method effectively improves the diagnostic accuracy of power problems and can detect and analyze potential faults of the power rail and module in a timely manner.
[0144] Embodiment 6
[0145] Please refer to Figure 1 and Figure 3 , specifically: S5 includes S51 and S52;
[0146] S51. Based on the abnormal energy absorption index EAAI and the module self-healing ability index REI output by the abnormal energy absorption behavior analysis, perform a summary calculation to output the power supply health assessment index PHS of all computer motherboard modules, and quantify the power supply health status of the motherboard;
[0147] The power supply health assessment index PHS is calculated and output through the following algorithm formula;
[0148] ;
[0149] In the formula, M represents the total number of modules, and Var(Pabnormal) represents the variance of the energy absorbed during fluctuations, which is used to measure the changes in power supply fluctuations over different time periods. 、 and represent the preset weight values of the abnormal energy absorption index EAAI, the module self-healing ability index REI, and the variance of the energy absorbed during fluctuations. The specific values are set by the user, and + + = 1;
[0150] The significance of the formula is that the power supply health assessment index PHS is a comprehensive assessment index. By integrating multiple factors such as the energy absorption, recovery ability, and power consumption fluctuations of the computer motherboard modules, it provides a quantitative power supply health score. A high value of the power supply health assessment index PHS indicates that there may be serious instability factors in the power supply system, affecting the reliability of the device.
[0151] S52. Conduct a comprehensive comparative assessment based on the output result of the power supply health assessment index PHS, analyze the power supply health status of the overall modules of the computer motherboard, and trigger different response measures based on the comprehensive comparative assessment result. The specific assessment content is as follows;
[0152] When the power supply health assessment index PHS < 0.5, it indicates that the power supply is stable in the case of abnormal power rail fluctuation response. At this time, continue to observe;
[0153] When the power supply health assessment index PHS ≥ 0.5, it indicates that there is an abnormality in the computer motherboard power supply in the case of abnormal power rail fluctuation response. At this time, generate a prompt warning through the motherboard fault monitoring system and replace the power supply group components of the computer motherboard.
[0154] In this embodiment, the method quantifies the impact of power rail fluctuations on the module and the self-healing ability of the module by calculating the abnormal energy absorption index EAAI and the module self-healing ability index REI, and then combines these data into the power supply health assessment index PHS to form a comprehensive power supply health score. The power supply health assessment index PHS provides a quantitative power supply health assessment by considering multiple factors such as the energy absorption, recovery ability, and power consumption fluctuations of each module. The output result of this index can help determine whether the motherboard power supply system is stable and trigger corresponding response measures according to the change of the power supply health assessment index PHS value. The optimized implementation can identify the unstable factors existing in the power supply system in a timely manner through real-time monitoring and comprehensive comparison and evaluation in the case of abnormal power rail fluctuations. When the power supply health assessment index PHS value is less than 0.5, the system indicates that the power supply is stable and continuous observation is sufficient; while when the power supply health assessment index PHS value is greater than or equal to 0.5, the system determines that there is an abnormality in the power rail, triggers an alarm and recommends replacing the power supply components. This method greatly improves the diagnostic accuracy and response speed of power supply system problems through multi-level and dynamic power supply health assessment. After optimization, the system can not only more accurately monitor the operating state of the power supply system, but also predict and handle potential faults caused by power supply instability in advance, reduce the downtime and maintenance costs caused by power supply problems, and improve the reliability, long-term stability and overall operation and maintenance efficiency of the equipment.
[0155] Embodiment 7
[0156] Please refer to Figure 2 , a monitoring and diagnostic system for computer motherboard faults, including a multi-power rail monitoring module, a virtual power rail identification module, a lightweight module, an abnormal behavior analysis module, and a comprehensive power supply analysis module;
[0157] The multi-power rail monitoring module collects the power rail data of the motherboard multi-power rail in real time through the motherboard BIOS interface, and transmits the power rail data to the motherboard fault monitoring system, and preprocesses the power rail data to obtain a standardized virtual rail data set;
[0158] The virtual power rail identification module constructs a virtual rail index model in the motherboard fault monitoring system, extracts the standardized virtual rail data set and inputs it into the virtual rail index model for calculation to output the voltage rail virtualization index VEI, and based on the output result of the voltage rail virtualization index VEI, conducts a preliminary comparison and evaluation to judge the situation of the motherboard virtual power rail;
[0159] When the lightweight module determines that it is a virtual power rail through the preliminary comparison and evaluation result, it triggers a lightweight response mechanism. The lightweight response mechanism starts the BIOS interface to collect the performance data of the corresponding motherboard module, calculates and outputs the lightweight behavior risk index LBPI based on the performance data, and conducts a lightweight evaluation and analysis based on the output result of the lightweight behavior risk index LBPI;
[0160] Based on the results of lightweight evaluation analysis, the abnormal behavior analysis module triggers the analysis of abnormal energy absorption behavior, and calculates and outputs the abnormal energy absorption index EAAI and the module self-healing ability index REI of the main board module;
[0161] The integrated power supply analysis module extracts the abnormal energy absorption index EAAI and the module self-healing ability index REI, conducts summary calculations to output the power supply health assessment index PHS, and conducts comprehensive comparative evaluations based on the output results of the power supply health assessment index PHS to analyze the power supply stability of the computer motherboard.
[0162] Specific example:
[0163] Assumption: +12V power rail: steady-state voltage V0 = 12.0, voltage change rate ΔV = 0.15, rail signal response duration tr = 2.5, voltage jitter frequency F = 50;
[0164] +5V power rail: steady-state voltage V0 = 5.0, voltage change rate ΔV = 0.1, rail signal response duration tr = 2.0
[0165] Voltage jitter frequency F = 60;
[0166] +3.3V power rail: steady-state voltage V0 = 3.3, voltage change rate ΔV = 0.2, rail signal response duration tr = 3.0
[0167] Voltage jitter frequency F = 55;
[0168] = 0.5, Fref = 100;
[0169] +12V power rail: VEI = 0.15 / 12.0 + 50 / 100 + 0.5 * 2.5 = 1.7625;
[0170] +5V power rail: VEI = 0.1 / 5.0 + 60 / 100 + 0.5 * 2.0 = 1.62;
[0171] +3.3V power rail: VEI = 0.2 / 3.3 + 55 / 100 + 0.5 * 3.0 = 2.1106;
[0172] 12V power rail = 1.7625 ≥ 1.0, triggering the lightweight response mechanism;
[0173] +5V power rail = 1.62 ≥ 1.0, triggering the lightweight response mechanism;
[0174] +3.3V power rail = 2.1106 ≥ 1.0, triggering the lightweight response mechanism;
[0175] The system startup BIOS interface collects the following performance data:
[0176] Abnormal restart rate Rrestart: for the +12V power rail: 0.02, for the +5V power rail: 0.03, and for the +3.3V power rail: 0.01;
[0177] Performance decay rate ΔPerf under voltage micro-offset: for the +12V power rail: 0.05, for the +5V power rail: 0.03, and for the +3.3V power rail: 0.07;
[0178] Frequency modulation behavior frequency Fdvfs: for the +12V power rail: 0.1, for the +5V power rail: 0.2, and for the +3.3V power rail: 0.15;
[0179] For the +12V power rail: LBPI = 0.4 * 0.02 + 0.4 * 0.05 + 0.2 * 0.1 = 0.048;
[0180] For the +5V power rail: LBPI = 0.4 * 0.03 + 0.4 * 0.03 + 0.2 * 0.2 = 0.064;
[0181] For the +3.3V power rail: LBPI = 0.4 * 0.01 + 0.4 * 0.07 + 0.2 * 0.15 = 0.062;
[0182] For the +12V power rail LBPI = 0.048: Normal LBPI < 0.3, continue monitoring.
[0183] For the +5V power rail LBPI5V = 0.064: Normal LBPI < 0.3, continue monitoring.
[0184] For the +3.3V power rail LBPI3.3V = 0.062: Normal LBPI < 0.3, continue monitoring.
[0185] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention.
Claims
1. A method for monitoring and diagnosing computer motherboard failures, characterized in that: The following steps are involved: S1. Real-time acquisition of power rail data of multiple power rails of the motherboard is performed through the motherboard BIOS interface, and the power rail data is transmitted to the motherboard fault monitoring system, and a standardized virtual rail data set is obtained by pre-processing the power rail data; S2. In the motherboard fault monitoring system, a virtual rail index model is constructed, and a standardized virtual rail data set is extracted and input into the virtual rail index model to calculate and output a voltage rail virtualization index VEI for analyzing the health status of the power rail, and based on the output result of the voltage rail virtualization index VEI, a preliminary comparative evaluation is performed to determine the status of the motherboard virtual rail; S3. When the preliminary comparison and evaluation result is determined to be a virtual rail, a lightweight response mechanism is triggered. The lightweight response mechanism collects the performance data of the corresponding motherboard module by starting the BIOS interface, and calculates and outputs a lightweight behavior risk index LBPI based on the performance data, which is used to comprehensively quantify the number of micro-restarts, performance degradation, virtual rail offset and frequency adjustment behavior anomalies of different modules of the computer motherboard, and performs lightweight evaluation and analysis based on the output result of the lightweight behavior risk index LBPI; S4. Based on the lightweight assessment analysis results, abnormal energy absorption behavior analysis is triggered to calculate and output the abnormal energy absorption index EAAI and module self-healing ability index REI of the mainboard module, and analyze the abnormal degree of energy absorption of the sub-module and the self-healing repair ability of the module respectively; S5. Extract the abnormal energy absorption index EAAI and the module self-healing ability index REI, summarize and calculate the output power supply health assessment index PHS, and conduct a comprehensive comparative evaluation based on the output results of the power supply health assessment index PHS to analyze the power supply stability of the computer motherboard.
2. A computer motherboard fault monitoring and diagnosis method according to claim 1, characterized in that: Said S1 includes S11 and S12; S11. Install sensors on different power rails of the computer motherboard, and set the acquisition frequency to 100ms through the computer motherboard BIOS interface ACPI to collect power rail data in real time; The power rails include +12V main voltage, +5V system voltage, +3.3V / IO voltage, VCORE core power supply, VDD auxiliary rail and VPP auxiliary rail; The sensor includes a micro voltage sensor and an oscilloscope; The power rail data includes voltage data and response time t; S12, installing a motherboard fault monitoring system in the central processing unit of the computer motherboard, and directly transmitting the collected power rail data to the motherboard fault detection system, and preprocessing the power rail data in the motherboard fault monitoring system to obtain a standardized virtual rail data set; The preprocessing includes data processing, timestamp alignment and normalization processing; The timestamp alignment is used to unify the processed power rail data into a unified time base based on the starting time point of the same acquisition frequency, and to unify the power rail data with different acquisition frequencies using an interpolation method; The data processing performs steady-state comparison, voltage change analysis, rail signal response duration processing and spectrum analysis based on the voltage data of the power rail data after timestamp alignment, and extracts the power rail virtualization data after data processing; The power rail virtualization data includes a steady-state voltage V0, a voltage change rate ΔV, a rail signal response time tr, and a voltage jitter frequency F; The steady-state voltage V0 is obtained by steady-state comparison; The voltage change rate ΔV is obtained by voltage change analysis; The track signal response time tr is obtained by processing the track signal response time; The voltage jitter frequency F is obtained by spectrum analysis, and the voltage fluctuation data is analyzed by spectrum analysis; The normalization process converts the power rail virtual data into a standardized virtual rail data set of a standard normal distribution by using a Z-Score normalization method, thereby eliminating the dimensional influence of the power rail virtual data.
3. A computer motherboard fault monitoring and diagnosis method according to claim 2, characterized in that: The S2 includes S21 and S22; S21. In the motherboard fault monitoring system, a virtual rail index model is constructed, the preprocessed standardized virtual rail data set is extracted, and the data is input into the virtual rail index model to calculate the output voltage rail virtualization index VEI, and analyze the health status of the power rail; The voltage rail virtualization index VEI is calculated and output by the following virtual rail index model; ; Where, VEI i Represents the voltage rail virtualization index of the i-th power rail, Vnominal i Represents the standard steady-state voltage reference value of the i-th power rail, △V i represents the voltage change rate of the ith power rail, F i represents the voltage jitter frequency of the i-th power rail, Fref represents the standardized voltage jitter frequency reference value, tr i Indicates the track signal response time of the i-th power rail, Represents the balance factor.
4. A computer motherboard fault monitoring and diagnosis method according to claim 3, characterized in that: S22. Based on the output results of the voltage rail virtualization index VEI of each power rail, a preliminary comparative evaluation is performed to analyze the virtualization of the power rail of the computer motherboard. The specific evaluation contents are as follows; When the voltage rail virtualization index VEI of the i-th power rail i When ≥1, it indicates that there is a virtual power rail on the power rail, and the lightweight response mechanism is triggered; When the voltage rail virtualization index VEI of the i-th power rail i When <1, it indicates that the power rail is normal and current monitoring is maintained.
5. A computer motherboard fault monitoring and diagnosis method according to claim 4, characterized in that: The S3 includes S31 and S32; S31. After a lightweight response mechanism is triggered through preliminary comparative evaluation, the BIOS interface is started to collect performance data of the motherboard module with a virtual power rail in real time, and the performance data is normalized to eliminate the dimensional influence between parameters in the performance data. The performance data includes an abnormal restart rate Rrestart, a performance degradation rate ΔPerf under a voltage micro-offset, and a frequency modulation behavior frequency Fdvfs; Calculate and output the Lightweight Behavior Risk Index (LBPI) based on performance data; The lightweight behavior risk index LBPI is calculated and output by the following algorithm formula; ; Where LBPI j Represents the lightweight behavior risk index of the jth module, Rrestart j represents the abnormal restart rate of the jth module, T represents the time window length, ΔPerf j Fdvfs represents the performance degradation rate of the jth module under a slight voltage offset. j Indicates the frequency modulation behavior frequency of the jth module, the preset weight values of the abnormal restart rate Rrestart of a1, a2 and a3, the performance degradation rate ΔPerf under voltage micro-offset and the frequency modulation behavior frequency Fdvfs. The specific value is set by the user, and a1+a2+a3=1; S32, based on the output result of the lightweight behavior risk index LBPI, a lightweight assessment analysis is performed to determine the abnormal behavior of the computer motherboard module under the virtual power rail condition, and the specific assessment contents are as follows; When the lightweight behavioral risk index LBPI of the jth module j When <0.3, it means that the current computer motherboard module is behaving normally, and monitoring continues at this time; When 0.3≤LBPI of the jth module j When ≤0.7, it indicates that the current computer motherboard module behavior has a sub-health trend. At this time, the first warning message is generated, indicating that the current j-th module is sub-healthy, and the collection frequency is adjusted to 50ms; When the lightweight behavioral risk index LBPI of the jth module j When it is >0.7, it indicates that the current computer motherboard module behavior is abnormal in its ability to respond to power rail fluctuations, and abnormal energy absorption behavior analysis is triggered.
6. A computer motherboard fault monitoring and diagnosis method according to claim 1, characterized in that: The S4 includes S41 and S42; S41, after the power rail fluctuation response capability is abnormal at the lightweight assessment, the BIOS interface is started to collect the energy Pabnormal absorbed by all modules of the computer motherboard during fluctuation, the power consumption Pexpected under normal conditions and the recovery curve stabilization time Tstable, and after normalization, abnormal energy absorption behavior analysis is performed, wherein the abnormal energy absorption behavior analysis includes abnormal energy absorption analysis and module self-healing ability analysis; The abnormal energy absorption analysis extracts the energy absorbed during fluctuations Pabnormal, the power consumption under normal conditions Pexpected and the recovery curve stabilization time Tstable, and jointly calculates and outputs the abnormal energy absorption index EAAI to analyze the abnormal degree of the module in terms of energy absorption; The abnormal energy absorption index EAAI is calculated and output by the following algorithm formula: ; Where, EAAI j Pabnormal represents the abnormal energy absorption index of the jth module. j (t) represents the energy absorbed by the jth module during the fluctuation at time t, Tstable j (t) represents the recovery curve stabilization time of the jth module at time t, Pexpected j It represents the power consumption of the jth module under normal conditions, and Tavg represents the average recovery time of all modules.
7. A computer motherboard fault monitoring and diagnosis method according to claim 6, characterized in that: S42, after the abnormal energy absorption analysis, performing module self-healing ability analysis, the module self-healing ability analysis is based on the recovery curve stabilization time Tstable of all computer motherboard modules, performing an average calculation, obtaining the module self-healing ability index REI, and analyzing the module's self-healing repair ability; The module self-healing ability index REI is calculated and output by the following algorithm formula; ; In the formula, REI j represents the module self-healing capability index of the jth module, N represents the number of self-healing capability evaluations, and Tstable j (k) represents the stabilization time of the recovery curve of the jth module during the kth recovery process.
8. A computer motherboard fault monitoring and diagnosis method according to claim 6, characterized in that: The S5 includes S51 and S52; S51, based on the abnormal energy absorption index EAAI and the module self-healing ability index REI output by the abnormal energy absorption behavior analysis, a power supply health assessment index PHS of all computer motherboard modules is summarized and calculated to quantify the power supply health of the motherboard; The power supply health assessment index PHS is calculated and output by the following algorithm formula; ; In the formula, M represents the total number of modules, Var (Pabnormal) represents the variance of energy absorbed during fluctuation, which is used to measure the change of power fluctuation in different time periods. , and Indicates the preset weight values of abnormal energy absorption index EAAI, module self-healing ability index REI and energy variance absorbed during fluctuation. The specific value is set by the user, and + + =1.
9. A computer motherboard fault monitoring and diagnosis method according to claim 8, characterized in that: S52, perform a comprehensive comparative evaluation based on the output results of the power supply health evaluation index PHS, analyze the power supply health of the overall module of the computer motherboard, and trigger different response measures based on the comprehensive comparative evaluation results. The specific evaluation contents are as follows; When the power supply health assessment index PHS is less than 0.5, it means that the power supply is stable under the condition of abnormal power rail fluctuation response, and it is necessary to continue to observe; When the power health assessment index PHS ≥ 0.5, it means that the power rail fluctuation response is abnormal and there is an abnormality in the power supply of the computer motherboard. At this time, a prompt warning is generated through the motherboard fault monitoring system, and the power supply group components of the computer motherboard are replaced.
10. A computer motherboard fault monitoring and diagnosis system, applied to a computer motherboard fault monitoring and diagnosis method according to any one of claims 1 to 9, characterized in that: It includes multi-power rail monitoring module, virtual power rail identification module, lightweight module, abnormal behavior analysis module and comprehensive power supply analysis module; The multi-power rail monitoring module collects power rail data of the multi-power rails of the mainboard in real time through the mainboard BIOS interface, transmits the power rail data to the mainboard fault monitoring system, and pre-processes the power rail data to obtain a standardized virtual rail data set; The virtual power rail identification module constructs a virtual rail index model in the mainboard fault monitoring system, extracts a standardized virtual rail data set and inputs it into the virtual rail index model to calculate and output a voltage rail virtualization index VEI, and performs a preliminary comparative evaluation based on the output result of the voltage rail virtualization index VEI to determine the mainboard virtual power rail situation; When the lightweight module determines that the power rail is a virtual rail through preliminary comparison and evaluation results, the lightweight response mechanism is triggered. The lightweight response mechanism collects the performance data of the corresponding motherboard module by starting the BIOS interface, and calculates and outputs the lightweight behavior risk index LBPI based on the performance data, and performs lightweight evaluation analysis based on the output result of the lightweight behavior risk index LBPI; The abnormal behavior analysis module triggers abnormal energy absorption behavior analysis based on the lightweight assessment analysis results, calculates and outputs the abnormal energy absorption index EAAI of the mainboard module and the module self-healing ability index REI; The comprehensive power supply analysis module extracts the abnormal energy absorption index EAAI and the module self-healing ability index REI, summarizes and calculates the output power supply health assessment index PHS, and performs a comprehensive comparative evaluation based on the output results of the power supply health assessment index PHS to analyze the power supply stability of the computer motherboard.
Citation Information
Patent Citations
Substrate management circuit and method, computer equipment and mainboard control system
CN115407860A
Methods for detecting an imminent power failure in time to protect local design state
CN110546590A
Information Handling System And Methods To Detect Power Rail Failures And Test Other Components Of A System Motherboard
US20200233766A1