Fault diagnosis method of HBA card and electronic equipment

By acquiring time-series data from HBA cards, extracting multimodal features, and selecting features with high contribution, fault response information is generated. This solves the problems of high false alarm rate and low processing efficiency in traditional HBA card fault diagnosis methods, and achieves efficient automated fault processing.

CN120973587AActive Publication Date: 2025-11-18LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511500904.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2025-11-18
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

Traditional HBA card fault diagnosis methods cannot accurately distinguish between instantaneous fluctuations and real faults, resulting in a high false alarm rate, low fault handling efficiency, and difficulty in detecting progressive faults such as fiber optic aging.

Method used

By acquiring the time-series data of the HBA card, extracting the time-series feature set, including time-domain statistical features, frequency-domain energy features, and event burst features, filtering out features with high contribution, and generating fault response information based on fault labels and early warning rules to achieve automated processing.

Benefits of technology

It improves the accuracy and efficiency of fault diagnosis, reduces the false alarm rate, and enables timely and automated handling of HBA card faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973587A_ABST
    Figure CN120973587A_ABST
Patent Text Reader

Abstract

The invention discloses a fault diagnosis method for an HBA card and electronic equipment, and relates to the technical field of servers, and the method comprises the steps: extracting a time sequence feature set from time sequence data of the HBA card; the time domain statistical features of the time sequence feature set comprise one or more of I / O delay variance, temperature slope and link retry times; the frequency domain energy characteristics comprise error frequency and / or a dominant frequency energy ratio; the event burst feature includes signal strength fluctuations. And according to the contribution degree of each feature in the time sequence feature set, screening out a plurality of time sequence features for fault diagnosis from the time sequence feature set. And determining a fault label according to the relevance between the plurality of time sequence characteristics and a set fault category. The fault label comprises the fault category and the fault probability of the HBA card. And generating fault response information based on the fault label and a fault early warning rule matched with the plurality of time sequence characteristics. And the fault processing efficiency is improved while the fault diagnosis accuracy is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of server technology, and in particular to a fault diagnosis method and electronic device for an HBA card. Background Technology

[0002] The Host Bus Adapter (HBA) is a core connection component between servers and storage networks, responsible for protocol conversion and high-speed data transmission. Its failure can lead to a series of problems, including data loss, service disruption, and soaring maintenance costs.

[0003] Traditional fault diagnosis methods rely on preset fixed thresholds to diagnose HBA faults, such as triggering an alarm for temperatures above 85°C. This approach cannot distinguish between instantaneous fluctuations and genuine faults, resulting in a high false alarm rate. For progressive faults like fiber optic aging, triggering the threshold is difficult, leading to significant missed detections. Furthermore, this fault diagnosis method only provides alarms and cannot address faults promptly. Engineers must diagnose HBA faults by parsing log files, resulting in an inefficient and unreliable diagnostic process.

[0004] It is evident that how to improve fault handling efficiency while ensuring the accuracy of fault diagnosis is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0005] This application provides a fault diagnosis method and electronic device for HBA cards, which at least solves the problems of poor fault diagnosis accuracy and low fault handling efficiency in related technologies.

[0006] This application provides a fault diagnosis method for HBA cards, including: Obtain timing data from the HBA card; Extract a time-series feature set from the time-series data; the time-series feature set includes time-domain statistical features, frequency-domain energy features, and event burst features; the time-domain statistical features include one or more of I / O delay variance, temperature slope, and link retry count; the frequency-domain energy features include error frequency and / or main frequency energy ratio; the event burst features include signal strength fluctuations; Based on the contribution of each feature in the time series feature set, multiple time series features for fault diagnosis are selected from the time series feature set. Fault labels are determined based on the correlation between multiple time-series characteristics and the set fault categories; Fault response information is generated based on fault labels and fault warning rules matched with multiple time-series features.

[0007] This application also provides a fault diagnosis device for an HBA card, including an acquisition unit, an extraction unit, a filtering unit, a determination unit, and a generation unit; The acquisition unit is used to acquire timing data from the HBA card. The extraction unit is used to extract a time-series feature set from the time-series data. The time-series feature set includes time-domain statistical features, frequency-domain energy features, and event burst features. The time-domain statistical features include one or more of I / O delay variance, temperature slope, and link retry count. The frequency-domain energy features include error frequency and / or main frequency energy ratio. The event burst features include signal strength fluctuations. The filtering unit is used to filter out multiple time-series features for fault diagnosis from the time-series feature set based on the contribution of each feature in the time-series feature set. The determination unit is used to determine the fault label based on the correlation between multiple timing features and the set fault categories; The generation unit is used to generate fault response information based on fault labels and fault warning rules matched by multiple time-series features.

[0008] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described fault diagnosis methods for HBA cards.

[0009] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described fault diagnosis methods for HBA cards.

[0010] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described HBA card fault diagnosis methods.

[0011] This application obtains timing data of an HBA card and extracts a timing feature set from it. The timing feature set includes multimodal features such as time-domain statistical features, frequency-domain energy features, and event burst features. Time-domain statistical features capture the overall distribution and trend of the timing data, frequency-domain energy features capture periodic faults, and event burst features capture millisecond-level sudden faults. The error frequency and I / O delay variance of the HBA card are the most important features for fault detection. The main frequency energy ratio and temperature slope are important reference values ​​for HBA card fault cause analysis. Signal strength fluctuation and link retry count are important parameters for HBA card link quality assessment. Therefore, the time-domain statistical features include one or more of I / O delay variance, temperature slope, and link retry count; the frequency-domain energy features include error frequency and / or main frequency energy ratio; and the event burst features include signal strength fluctuation. By extracting multimodal timing features, the detection sensitivity is improved, and reliable data support is provided for subsequent fault detection. The temporal feature set contains a large number of temporal features, but some of these features contribute very little to fault detection and are not worth analyzing. Therefore, to improve fault detection efficiency, multiple temporal features for fault diagnosis can be selected from the set based on their contribution. Fault labels are determined based on the correlation between multiple temporal features and defined fault categories. These labels include the HBA card's fault category and probability. Different fault categories are associated with different temporal features. To automate fault handling, all possible combinations of temporal features under different fault categories can be summarized to set corresponding fault warning rules. These rules can include corresponding fault handling methods. After selecting the temporal features and determining the fault labels, fault response information can be generated based on the fault labels and the fault warning rules matched by multiple temporal features. This fault response information includes the response actions required to resolve the current HBA card fault. In this technical solution, the detection sensitivity is improved by extracting multimodal temporal features. Based on the fault warning rules, multiple time-series features and fault labels are comprehensively analyzed to avoid misoperation caused by fluctuations in a single feature. Furthermore, based on the fault response information, faults can be automatically processed in a timely manner, improving fault handling efficiency. Attached Figure Description

[0012] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A flowchart illustrating a fault diagnosis method for an HBA card provided in this application embodiment; Figure 2 A flowchart illustrating a method for obtaining timing data of multiple windows based on dynamic window segmentation, provided in an embodiment of this application; Figure 3 This application provides a data cleaning method flow for embodiments; Figure 4 A flowchart illustrating a method for filtering time-series features provided in this application embodiment; Figure 5 This is a schematic diagram of the structure of an HBA card fault diagnosis device provided in an embodiment of this application. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0015] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0016] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0017] Figure 1 A flowchart of a fault diagnosis method for an HBA card provided in this application embodiment is included. The method includes: S101: Obtain timing data from the HBA card.

[0018] In this application, the fault diagnosis of the HBA card is used as an example for description. The fault diagnosis process of the HBA card can be executed by the management software or by the Baseboard Management Controller (BMC), and there is no limitation here.

[0019] A distributed data acquisition method is used to collect raw time-series data from the HBA card; the raw time-series data may include hardware sensor data, performance indicators, and log information.

[0020] In practical applications, time-series data related to HBA performance and health status are collected from three layers: hardware sensors, performance monitoring agents, and log parsing engines. The hardware sensor layer collects real-time hardware sensor data, including information such as HBA card temperature, voltage, and signal strength. The performance metrics obtained from the operating system at the performance monitoring agent layer include HBA card input / output (I / O) throughput, latency, and queue depth. The log parsing engine layer parses the HBA card driver logs in real-time, extracting information such as error counts (CRC), link retries, and timeout events.

[0021] To balance real-time performance and feature integrity, the original time-series data can be sampled using a dynamic windowing method to obtain multiple window time-series data. To ensure the quality of the time-series data, data cleaning can be performed on the multiple window time-series data to obtain the time-series data for the HBA card.

[0022] The dynamic window segmentation method adaptively adjusts the window length based on the fluctuation characteristics of time series data to ensure that short windows respond quickly and long windows capture long-term trends, thereby achieving a balance between real-time performance and feature integrity.

[0023] Data cleaning aims to improve the quality of HBA card time-series data. Data cleaning can include operations such as noise removal, missing value filling, and unit standardization, providing reliable input for subsequent feature extraction and fault identification.

[0024] S102: Extract time series feature sets from time series data.

[0025] To facilitate subsequent fault detection and analysis, a set of time-series features with high information content and low redundancy can be extracted from the time-series data of the HBA card.

[0026] The time series feature set can include time-domain statistical features, frequency-domain energy features, and event burst features.

[0027] For the extraction of time-domain statistical features, the time-domain statistical features can be determined based on the distribution trend of time-series data within each window.

[0028] Time-domain statistical features are used to capture the overall distribution and trend of time-series data. These features can include any one or any combination of mean, variance, skewness, kurtosis, and slope.

[0029] The mean represents the average value of the data within the window, and the calculation formula is as follows: Where n represents the number of sampled data points within the window, μ represents the mean, and x... i This represents the i-th time series data.

[0030] Variance represents the degree of dispersion of data, and the calculation formula is as follows: ;in, Indicates variance.

[0031] Skewness represents the asymmetry of the data distribution, and the calculation formula is as follows: Skewness represents the degree of skewness.

[0032] Kurtosis represents the sharpness of a data distribution, and is calculated using the following formula: Where, Kurtosis represents kurtosis.

[0033] The slope represents the trend of the data calculated using linear regression, and the calculation formula is as follows: Where s represents the slope, t i This represents the i-th time point.

[0034] The purpose of extracting frequency domain energy features is to capture periodic patterns in time-series data, such as link jitter and periodic errors caused by firmware defects. In practical applications, time-series data can be converted into frequency domain signals. ; where X k represents the frequency domain signal; k represents a constant, which is a pre-set fixed value.

[0035] The dominant frequency component is extracted from the frequency domain signal to determine the dominant frequency energy ratio. ;in, E represents the main frequency energy. i The energy of the i-th frequency domain signal is represented by Energy Ratio, which represents the energy ratio of the main frequency.

[0036] Extracting sudden event characteristics is to capture sudden patterns of abnormal events, such as a sudden increase in CRC errors or a surge in the number of link retries.

[0037] A sudden increase in CRC errors refers to an error count growth rate exceeding a threshold within a unit of time, such as an increase of ≥50 CRC errors within 10 seconds.

[0038] Different emergencies have their own associated performance indicators. Therefore, the characteristics of an emergency can be extracted based on the correlation between performance indicators and failure events in time series data.

[0039] In practical applications, the correlation coefficient between emergencies and performance indicators can be calculated using the following formula. Where r represents the correlation coefficient, x i Let y represent the value of the i-th event. i μ represents the performance metric for the i-th sudden event. x μ represents the mean of sudden events. y This represents the mean of the performance metric.

[0040] For example, the correlation coefficient r = 0.82 between CRC error surge events and I / O delay indicates a strong correlation between the two.

[0041] Considering that the CRC error frequency and I / O delay variance of HBA cards are the most important features for fault detection, the main frequency energy ratio and temperature slope are of great reference value for HBA card fault cause analysis, and signal strength fluctuation and link retry count are important parameters for HBA card link quality assessment, the time domain statistical features can include I / O delay variance, temperature slope, and link retry count; the frequency domain energy features can include error frequency and main frequency energy ratio; and the event burst features can include signal strength fluctuation.

[0042] S103: Based on the contribution of each feature in the time series feature set, select multiple time series features from the time series feature set for fault diagnosis.

[0043] The time series feature set contains a large number of time series features, but some time series features contribute very little to fault detection and do not need to be analyzed. Therefore, in order to improve the efficiency of fault detection, multiple time series features for fault diagnosis can be selected from the time series feature set based on the contribution of each feature.

[0044] SHAP (SHapley Additive exPlanations) values ​​represent the degree to which a model's predictions deviate from the baseline. In practical applications, SHAP values ​​can be used as contribution values ​​to select the most influential time-series features from the set of time-series features.

[0045] Table 1 below lists the time series features that contribute significantly to the model's predictions, including CRC error frequency, I / O delay variance, main frequency energy ratio, temperature slope, signal fluctuation intensity, mean queue depth, number of link retries, voltage fluctuation, throughput degradation rate, and error code burst frequency. Therefore, the selected time series features can include these 10 time series features.

[0046] Table 1 Table 1 shows that CRC error frequency and I / O delay variance are the most important timing features for fault detection, contributing a combined 43%. The main frequency energy ratio and temperature slope are of significant reference value for root cause analysis. Signal strength fluctuations and link retry counts are mainly used for link quality assessment.

[0047] S104: Determine the fault label based on the correlation between multiple timing features and the set fault category.

[0048] In this embodiment of the application, multiple time-series features can be input into the fault classification model to obtain the fault probability corresponding to each fault category; wherein, the fault classification model is trained based on historical sample data, and the historical sample data includes a historical time-series feature set and its corresponding fault categories; the fault categories may include hardware faults, software faults and link faults.

[0049] After determining the failure probability corresponding to each failure category, the failure category with the highest failure probability and its probability can be used as the failure label. The failure label can include the failure category and failure probability of the HBA card.

[0050] In practical applications, the fault classification model can employ the Lightweight Gradient Boosting Machine (LightGBM). The probability value output by the LightGBM model ranges from 0 to 1, representing the likelihood of a fault in the HBA card within the current time window. LightGBM can map the model output to the [0, 1] interval using an activation function (Sigmoid). The formula for the activation function is as follows: ; in, This represents the original output of the fault classification model for the input feature x. This represents the final output failure probability.

[0051] A fault threshold can be set to assess whether a fault has occurred. For example, the fault threshold can be set to 0.5. When the probability of a fault exceeds the fault threshold, a fault is considered to have occurred.

[0052] In this embodiment, the fault threshold can be dynamically adjusted according to business needs to balance the false alarm rate and the false negative rate. For highly sensitive scenarios, the fault threshold can be adjusted to 0.3 to reduce the false negative rate. For high-precision scenarios, the fault threshold can be adjusted to 0.7 to reduce the false alarm rate.

[0053] HBA card faults can be categorized into three main types: hardware faults, software faults, and link faults. LightGBM can output the fault probability for each fault category through a multi-classification task.

[0054] For ease of description, P can be used. 硬件 P represents the probability of a hardware failure. 链路 P represents the probability of a link failure. 软件 This represents the probability of a software fault.

[0055] The failure probability can be normalized using the following probability distribution function (Softmax): ;in, .

[0056] Select the fault category with the highest probability of failure and set the fault label: .

[0057] The timing characteristics associated with different fault categories vary. Table 2 below records the list of relationships between fault categories and timing characteristics.

[0058] Table 2 For example, input characteristics include CRC error frequency = 150 times / minute, I / O delay variance = 50ms², and main frequency energy ratio = 40%. Output characteristics include P... 硬件 =0.1, P 链路 =0.85, P 软件 =0.05, then the probability of a fault label indicating a link fault is 0.85.

[0059] S105: Generate fault response information based on fault labels and fault warning rules matched by multiple time-series features.

[0060] Different fault categories are associated with different time-series characteristics, and the same fault category can be further subdivided into different fault problems. The distribution of different time-series characteristics under the same fault category reflects different fault problems, and each fault problem has its corresponding fault handling method.

[0061] In this embodiment, to automate fault handling, all possible combinations of time-series features under different fault categories can be aggregated to set corresponding fault warning rules. These rules can include corresponding fault handling methods. After filtering out time-series features and determining fault labels, fault response information can be generated based on the fault labels and the fault warning rules matching multiple time-series features.

[0062] In practical implementation, target fault warning rules that match fault labels and target time-series features can be filtered from a defined rule base. Each fault warning rule in the rule base has a corresponding threshold range, which includes the threshold range for fault probability and the threshold range for each time-series feature. Based on the threshold ranges corresponding to each fault warning rule in the rule base, fault warning rules that match fault labels and target time-series features can be filtered from the rule base.

[0063] The fault tags and target time-series characteristics are analyzed according to the target fault early warning rules to determine the fault priority. Based on the response action mapping table, the fault response information matching the fault priority is determined; the fault response information includes the automated response method and the notification method.

[0064] The syntax structure of the fault warning rule is as follows: IF <conditional expression>; AND / OR <conditional expression>; THEN <action>; WITH PRIORITY. Taking link failure as an example, the following conditions must be met: IF failure probability ≥ 0.9; AND CRC error frequency ≥ 100 times / minute; AND delay variance ≥ 20ms. 2 THEN triggers "Link Anomaly Warning"; WITH PRIORITY 1# (highest priority).

[0065] Taking heat dissipation failure as an example of hardware failure, IF failure probability ≥ 0.8; AND temperature slope ≥ 0.5℃ / minute; AND throughput drop rate ≥ 20%; THEN triggers "heat dissipation anomaly warning"; WITH PRIORITY 2.

[0066] In this embodiment, response time requirements can be set for different priorities. The higher the priority level, the shorter the corresponding response time, so as to ensure that high-priority faults can be processed more quickly. Table 3 below is a list of priority levels.

[0067] Table 3 Corresponding automated response actions can be set for different fault priorities. In this embodiment, to ensure timely fault handling, a response action mapping table can be pre-built, recording the automated response methods and notification methods corresponding to different priorities. To ensure timely fault handling, an escalation mechanism can also be set in the response action mapping table. Table 4 below shows the response action mapping table.

[0068] Table 4 Taking the sudden failure of the HBA card in the data center as an example, the event sequence is as follows: time=00:00: CRC error frequency exceeds the threshold (120 times / minute); time=00:05: I / O latency variance exceeds the limit (25ms²); time=00:06: Failure probability rises to 0.92.

[0069] Rule triggering: IF Failure probability ≥ 0.9; AND CRC errors ≥ 100 times / minute; AND latency variance ≥ 20ms 2 THEN triggers "Link Anomaly"; WITH PRIORITY 1.

[0070] Response process: time=00:06: Automatically isolate the faulty card and switch to the redundant link; time=00:07: Notify the on-duty engineer via SMS; time=00:10: The engineer confirms the fault and prompts "Check the fiber optic interface"; time=00:15: Replace the small pluggable optical module (SFP) and the fault is resolved.

[0071] As can be seen from the above technical solution, the process involves acquiring timing data from the HBA card and extracting a timing feature set from this data. This timing feature set includes multimodal features such as time-domain statistical features, frequency-domain energy features, and event burst features. Time-domain statistical features can capture the overall distribution and trend of the timing data, frequency-domain energy features can capture periodic faults in the timing data, and event burst features can capture millisecond-level sudden faults. The error frequency and I / O delay variance of the HBA card are the most important features for fault detection. The main frequency energy ratio and temperature slope are important reference values ​​for HBA card fault cause analysis. Signal strength fluctuation and link retry count are important parameters for HBA card link quality assessment. Therefore, the time-domain statistical features include one or more of I / O delay variance, temperature slope, and link retry count; the frequency-domain energy features include error frequency and / or main frequency energy ratio; and the event burst features include signal strength fluctuation. By extracting multimodal timing features, the detection sensitivity is improved, and reliable data support is provided for subsequent fault detection. The temporal feature set contains a large number of temporal features, but some of these features contribute very little to fault detection and are not worth analyzing. Therefore, to improve fault detection efficiency, multiple temporal features for fault diagnosis can be selected from the set based on their contribution. Fault labels are determined based on the correlation between multiple temporal features and defined fault categories. These labels include the HBA card's fault category and probability. Different fault categories are associated with different temporal features. To automate fault handling, all possible combinations of temporal features under different fault categories can be summarized to set corresponding fault warning rules. These rules can include corresponding fault handling methods. After selecting the temporal features and determining the fault labels, fault response information can be generated based on the fault labels and the fault warning rules matched by multiple temporal features. This fault response information includes the response actions required to resolve the current HBA card fault. In this technical solution, the detection sensitivity is improved by extracting multimodal temporal features. Based on the fault warning rules, multiple time-series features and fault labels are comprehensively analyzed to avoid misoperation caused by fluctuations in a single feature. Furthermore, based on the fault response information, faults can be automatically processed in a timely manner, improving fault handling efficiency.

[0072] In this embodiment of the application, in order to better adapt to the changes in the state of the HBA card and ensure the rationality of the threshold range setting, the threshold range can be adjusted.

[0073] Adjustment methods can include baseline adaptive adjustment, which dynamically calculates the threshold based on historical device data. Taking error threshold adjustment as an example, the average and standard deviation of error counts within a set time period can be statistically analyzed; the error threshold is then determined based on the average and standard deviation.

[0074] For example, base = moving average (historical CRC errors, window = 24h); threshold = base + 3 * standard deviation (historical CRC errors); where base represents the average error count over 24 hours; threshold represents the adjusted error threshold.

[0075] Adjustment methods can include environmental factor compensation: adjusting the threshold based on environmental parameters such as data center temperature and load rate. Taking the adjustment of the temperature slope threshold as an example, a compensation coefficient can be determined based on the difference between the current room temperature and the standard room temperature; the quotient of the temperature slope threshold and the compensation coefficient is used as the adjusted temperature slope threshold.

[0076] For example, the compensation coefficient = 1 + 0.1 * (current room temperature - standard room temperature 25℃); the effective slope = measured slope / compensation coefficient. Here, the measured slope represents the current temperature slope threshold, and the effective slope represents the adjusted temperature slope threshold.

[0077] In this embodiment, a dynamic threshold adjustment mechanism is used to adjust the threshold range in the fault warning rule, so that the threshold range can better adapt to changes in equipment performance and environment, ensuring the rationality of the threshold range setting, thereby filtering out fault warning rules that match the current time series characteristics.

[0078] In practical applications, there may be multiple fault warning rules matching fault labels and target time-series characteristics. Different fault warning rules may determine different fault handling methods, potentially leading to diagnostic conflicts. Therefore, a conflict arbitration mechanism can be adopted to eliminate these conflicts. This mechanism can include elements such as priority based on rank and contribution weight.

[0079] Each fault warning rule has its corresponding priority. When multiple fault warning rules are matched, the target fault warning rule can be determined according to its priority level, that is, the fault warning rule with the highest priority is selected as the target fault warning rule. When there are multiple fault warning rules with the highest priority, the target fault warning rule can be determined according to its contribution weight, that is, the fault warning rule with the highest contribution is selected as the target fault warning rule.

[0080] In this embodiment, the priority of each fault warning rule can be preset. In practical applications, the priority of the target fault warning rule can be adjusted according to the number of times the target fault warning rule is triggered within a period. The period can be flexibly set; for example, it can be set to 24 hours (h).

[0081] The more times a target fault warning rule is triggered, the higher its rule priority can be adjusted.

[0082] The rule priority is adjusted as follows: adjusted_priority = base_priority * (1 + 0.2 * log(trigger count + 1)); Here, adjusted_priority represents the adjusted rule priority, and base_priority represents the rule priority before adjustment.

[0083] In this embodiment, a conflict arbitration mechanism is introduced to prioritize the fault warning rule with the highest priority as the target fault warning rule. When there are multiple fault warning rules with the highest priority, the fault warning rule with the highest contribution is selected as the target fault warning rule, effectively eliminating diagnostic conflicts. Furthermore, the rule priority is adjusted based on the triggering frequency of the fault warning rule, making the rule priority of the fault warning rule more closely match actual needs.

[0084] In order to map the early warning faults to specific root causes and achieve accurate fault location, root cause analysis can be performed on the faults.

[0085] Considering that bidirectional long short-term memory (Bi-LSTM) network models can capture multi-scale temporal dependencies, achieve feature enhancement, and support interpretable diagnostics, root cause analysis can be performed using a Bi-LSTM network model in this embodiment.

[0086] In practical implementation, time-series data and its corresponding time-series feature set can be input into a bidirectional long short-term memory network model to obtain output results. These output results can include predicted probabilities of various faults and feature vectors used for fault mode matching. Based on the feature similarity between the output results and the root cause rules in the fault mode library, the root cause analysis results are determined.

[0087] Bi-LSTM can capture multi-scale temporal dependencies. Forward LSTM can be used to analyze future impacts, while backward LSTM can be used to trace historical causes. Table 5 below lists the Bi-LSTM processing methods at different time scales.

[0088] Table 5 Based on Bi-LSTM feature enhancement, a 128-dimensional high-value feature vector can be output from a 50-dimensional input time-series feature, improving the retention rate of key features and the accuracy of pattern matching.

[0089] Bi-LSTM supports interpretable diagnostics.

[0090] Example of attention weights (t=12:30, weight 0.92): Timeline: [12:00, 12:15, 12:30, 12:45]; Attention: [0.02, 0.05, 0.92, 0.01] → Instructs engineers to focus on checking the 12:30 log.

[0091] Feature contribution analysis, quantifying the impact of features: Contribution of main frequency energy ratio (1-5Hz): +37% (driving the "firmware defect" judgment); Voltage fluctuation contribution: -5% (suppresses false positives for "power failure").

[0092] When performing root cause analysis using Bi-LSTM, historical data needs to be extracted. This historical data can include time-series data from the most recent 24 hours and time-series features with a 50-dimensional feature set. This historical data is used as input to the Bi-LSTM, yielding outputs including fault type probability distributions and 128-dimensional feature vectors.

[0093] Here, the fault type probability distribution represents the predicted probability of each type of fault, and its dimension is equal to the number of fault types. The 128-dimensional feature vector (feature_vector) is the final hidden state, representing the feature vector used for fault mode matching.

[0094] The output results are shown below: "fault_probabilities":{ Firmware Defect: 0.91 "Link unstable": 0.07 "Overheating failure": 0.02 }, "feature_vector":[0.24, -1.32, 0.57, ...] / / 128-dimensional compressed features }

[0095] The output of the Bi-LSTM model is matched with rules in the fault mode library using feature similarity, and the rule with the highest similarity is taken as the final root cause analysis result. The feature similarity calculation formula is as follows: ; Among them, F i T represents the eigenvector; i This represents the predefined feature template values ​​in the fault mode library, which are standardized feature representations belonging to known fault types; w i The feature weights can be determined by their contribution. Table 6 below lists the failure mode library.

[0096] Table 6 The typical pattern is as follows: { "Pattern_ID":"F007", "Name":"Firmware Defect" "Key_Features":[ "FFT detected a periodic signal of 0.0033Hz". "Voltage fluctuation <2%" Error code burst interval 300±10 seconds ], "Repair_Guide":[ "Upgrade firmware to v3.2.1", "Rollback to stable version v2.8.4" ], "Confidence": 0.96 }

[0097] By collecting new fault samples, features are extracted and standardized. The standardized features are then trained using a Bi-LSTM model to generate feature templates. These newly generated templates are added to a fault mode library, and their confidence level is initialized to 0.8. Continuous improvement of the fault module library ensures the accuracy of root cause analysis results.

[0098] In this embodiment of the application, time-series data can be obtained using a dynamic window segmentation method. Figure 2 A flowchart illustrating a method for obtaining timing data of multiple windows based on dynamic window segmentation, provided in this application embodiment, is included in the following method: S201: In the initial state, the first window of timing data is extracted from the original timing data according to the set window length.

[0099] Before obtaining time-series data using dynamic window segmentation, the parameters of the dynamic window segmentation algorithm are defined as follows: Input parameters: raw timing data Examples include I / O latency and CRC error count; where X represents the original timing data, x t This represents the value of the time series at the t-th sliding window, i.e., time point t.

[0100] Output parameters: Dynamic window sequence Each window W i It contains a set of continuous data points.

[0101] Key parameter: Initial window length W int The default value can be set to 300 seconds.

[0102] Standard deviation threshold σ threshold You can set it based on the distribution of historical data, such as 10% of the data range, to avoid being too sensitive to noise.

[0103] Overlap rate R overlap This represents the overlap ratio between windows, which helps avoid missing features and balances computational efficiency with feature continuity. It is usually set to a value between 30% and 50%.

[0104] With an initial window length W int Divide the data and generate the first window W1. For example, if the data sampling interval is 1 second, W... int =300 seconds, then W1 contains 300 data points.

[0105] S202: Adjust the window length of the next window based on the fluctuation index of the time series data in the previous window.

[0106] In this embodiment, the volatility index may include standard deviation and slope. Standard deviation σ i Used to measure the degree of dispersion of data; slope s i Data trends, such as the rate of temperature increase, are calculated using linear regression.

[0107] The formula for calculating standard deviation is as follows: ; The formula for calculating the slope is as follows: ; Where, x j This represents the j-th sampled data within the sliding window; μ i t represents the average value of the sampled data within the sliding window; j Represents the timestamp; n represents the number of sampled data points within the sliding window; σ i s represents the standard deviation. i Indicates the slope.

[0108] The principle for adjusting the window length is to reduce the window length to improve sensitivity when the data fluctuates drastically; to extend the window length to capture periodic faults when the data is stable; and to keep the window length unchanged when the data fluctuation is within a reasonable range.

[0109] In practical implementation, if the fluctuation index of the time series data in the previous window meets the data fluctuation condition, the length of the previous window can be reduced according to the set shrinking rule to serve as the window length of the next window. If the fluctuation index of the time series data in the previous window meets the data stationarity condition, the length of the previous window can be increased according to the set extending rule to serve as the window length of the next window. If the fluctuation index of the time series data in the previous window does not meet either the data fluctuation condition or the data stationarity condition, it indicates that the fluctuation of the time series data is within a reasonable range. In this case, the window length can be kept unchanged, that is, the length of the previous window can be used as the window length of the next window.

[0110] Data fluctuation conditions may include determining whether the standard deviation of the time series data in the previous window is greater than a first standard threshold and whether the absolute value of the slope of the time series data in the previous window is greater than a first slope threshold.

[0111] If the standard deviation of the time series data in the previous window is greater than the first standard threshold and the absolute value of the slope of the time series data in the previous window is greater than the first slope threshold, it indicates that the fluctuation index of the time series data in the previous window meets the data fluctuation condition. At this time, the window length of the next window can be determined according to the set minimum window value, the length of the previous window and its corresponding reduction ratio.

[0112] The data stationarity condition may include determining whether the standard deviation of the time series data in the previous window is less than or equal to a second standard threshold and whether the absolute value of the slope of the time series data in the previous window is less than or equal to a second slope threshold.

[0113] If the standard deviation of the time series data in the previous window is less than or equal to the second standard threshold and the absolute value of the slope of the time series data in the previous window is less than or equal to the second slope threshold, it indicates that the fluctuation index of the time series data in the previous window meets the data stationarity condition. At this time, the window length of the next window can be determined according to the set maximum window value, the length of the previous window and its corresponding extension ratio. Among them, the second standard threshold is less than the first standard threshold, and the second slope threshold is less than the first slope threshold.

[0114] For setting the threshold, half of the first standard threshold can be used as the second standard threshold, and half of the first slope threshold can be used as the second slope threshold.

[0115] In practical applications, the window length can be adjusted using the following formula: ; Among them, W min W represents the minimum value of the window. max W represents the maximum window size. i W represents the length of the i-th window. i+1 σ represents the length of the (i+1)th window. threshold s represents the first standard threshold. threshold This represents the first slope threshold, 0.5σ. threshold This indicates the second standard threshold, 0.5s. threshold This represents the second slope threshold.

[0116] S203: Extract the window time series data of the next window from the original time series data according to the window length and overlap rate of the next window, until all the original time series data has been extracted.

[0117] According to the adjusted W i+1 and overlap rate R overlap This can generate the next window of time series data.

[0118] The starting position of the next window = the starting position of the previous window + (1 - R) overlap )*W i .

[0119] For example, if W i =300 seconds, R overlap If the value is 50%, then the starting position of the next window, i.e., the new window, will be 150 seconds after the previous window.

[0120] In this embodiment, by adjusting the window length of the next window based on the fluctuations in the time-series data within the previous window, the problem of traditional fixed windows failing to capture rapidly changing instantaneous faults and missing sudden anomalies is solved. Redundant computation caused by using short windows for stable data is also resolved, avoiding resource waste.

[0121] In this embodiment of the application, in order to improve the quality of time series data, data cleaning can be performed on time series data from multiple windows. Figure 3 A flowchart of a data cleaning method provided in this application embodiment is shown. The method includes: S301: Perform noise filtering on multiple window time series data according to the set noise filtering method to obtain filtered multiple window time series data.

[0122] HBA card sensors may generate instantaneous noise (such as temperature spikes, I / O delay anomalies) due to electromagnetic interference and signal jitter. Therefore, noise filtering is used to filter the noise. Here, median filtering is used to process the noise.

[0123] Median filtering: Replaces the current point with the median value of the data within a sliding window, effectively suppressing impulse noise. The median filtering formula is as follows: ; Among them, y t x represents the data at time point t after median filtering. t This represents the value of the time series at time point t before median filtering, where median represents the function that takes the median, and the window size k is dynamically adjusted according to the duration of the noise.

[0124] For example, for transient noise (duration < 1 second), k=3 is selected, which means 3 sampling points. If multiple abnormal points are detected continuously, such as voltage fluctuations > 10% within 5 windows, the window is expanded to k=5.

[0125] For example: Original temperature data: [61, 63, 62, 119, 61] → After filtering: [61, 63, 62, 61, 61]. The 119℃ peak has been removed.

[0126] S302: Based on linear interpolation, the first type of missing data in the filtered multi-window time series data is supplemented to obtain the first supplemented multi-window time series data.

[0127] The first type of missing data consists of data where the number of consecutively missing sampling points is less than the set missing value.

[0128] In practical applications, data points may be missing due to transmission interruptions or sensor malfunctions, so missing value processing is required. Based on the amount of missing data, it can be divided into the first type of missing data, namely short-term missing data, and the second type of missing data, namely long-term missing data.

[0129] The missing value can be preset. For example, if the missing value can be 3, then if the number of consecutive missing sampling points is less than 3 consecutive points, it means that the missing data belongs to the first type of missing data; if the number of consecutive missing sampling points is greater than or equal to 3 consecutive points, it means that the missing data belongs to the second type of missing data.

[0130] For short-term missing data, linear interpolation is used: imputation is based on the preceding and following data points. The formula for linear interpolation is as follows: ; where x t This represents the value of the time series at time point t.

[0131] For example: Original delayed sequence: [5, NaN, NaN, 8] → after interpolation: [5, 6, 7, 8].

[0132] S303: Based on the influence weight of historical data on missing data, supplement the second type of missing data in the multiple window time series data after the initial supplementation to obtain the final supplemented multiple window time series data.

[0133] The second type of missing data refers to data in which the number of consecutively missing sampling points is greater than or equal to the set missing value.

[0134] For long-term missing data, a prediction method based on historical patterns can be used, employing the ARIMA model to predict missing segments. This method is suitable for periodic indicators, such as I / O load at fixed times each day.

[0135] The formula corresponding to the ARIMA model is as follows: ; Where, x t This represents the value of the time series at time point t, such as the I / O delay or temperature of an HBA card; c represents a constant term, which represents the mean or baseline of the time series. Represents the autoregressive coefficient, used to characterize the weight of the influence of the past p time points on the current value; This represents the moving average coefficient, used to characterize the weight of the influence of random errors from the past q time points on the current value; This represents the random error at time point t, which is white noise. It is usually assumed to follow a mean of 0 and a variance of σ. 2 The value follows a normal distribution; p represents the autoregression order, used to characterize the prediction of the current value using the values ​​of the past p time points; q represents the moving average order, used to characterize the prediction of the current value using the random error of the past q time points.

[0136] S304: Normalize the final supplemented timing data of multiple windows to obtain the timing data of the HBA card.

[0137] The data from multiple sources have significantly different units, such as temperature in °C and latency in ms, which can affect the model's convergence speed and accuracy. Therefore, data normalization is necessary. Here, Min-Max normalization is used for data processing. The normalization formula is as follows: ; in, x represents the normalized time series data at time point t. max x represents the maximum value of the time series data. min This represents the minimum value of the time series data.

[0138] To accommodate baseline drift caused by equipment aging, x can be reset every 24 hours. max and x min .

[0139] In this embodiment, median filtering and linear interpolation are used to supplement short-term missing data, while long-term missing data is supplemented based on the influence weight of historical data on the missing data. This effectively balances denoising effect and information preservation, avoiding the loss of key abnormal patterns due to excessive cleaning.

[0140] Given that the time series feature set contains a large number of time series features, and some time series features are not actually helpful for fault diagnosis, the time series features contained in the time series feature set can be filtered. Figure 4 A flowchart of a method for filtering time-series features provided in this application embodiment is shown. The method includes: S401: Take the remaining features in the time series feature set other than the target feature as the target subset.

[0141] The target feature is any one of the features contained in the time series feature set.

[0142] The time series feature set contains time-domain statistical features, frequency-domain energy features, and event burst features. Using the SHAP value, we can select several key features that contribute the most to the model prediction from these time series feature sets.

[0143] S402: Determine the contribution of the target features based on the time series feature set, the target subset and their respective predicted values.

[0144] In the specific implementation, the time series feature set and the target subset can be used as input features of the fault classification model to determine the predicted values ​​corresponding to the time series feature set and the target subset respectively; the time series feature set and the target subset are processed according to the set weight calculation rules to obtain the weights; the difference between the predicted value corresponding to the time series feature set and the predicted value corresponding to the target subset is multiplied by the weight to obtain the contribution of the target feature.

[0145] The formula for calculating the contribution of time series feature j is: ; in, Let f(j) represent the contribution of time series feature j, F represent the set of all features, i.e., the time series feature set; S represent the target subset that does not contain feature j; f(S) represents the model's prediction of the target variable when only features in the target subset S are used. Indicates the weight.

[0146] When calculating f(S), features not included in S are masked or filled with default values, such as the mean, median, or zero. Weighting ensures fairness in contribution allocation.

[0147] S403: Select a set number of time series features from the time series feature set in descending order of contribution.

[0148] The number can be set to 10. Ten time series features can be selected from the time series feature set according to the order of contribution from high to low. The 10 time series features can be referred to Table 1 for introduction, and will not be repeated here.

[0149] In this embodiment, by integrating multi-dimensional sensor data, a comprehensive solution for HBA card fault diagnosis is provided, improving fault diagnosis accuracy and reducing false alarm rate. Furthermore, during fault analysis, a predetermined number of time-series features with the highest contribution are first selected from the time-series feature set. This significantly reduces computational load and improves fault analysis efficiency while ensuring analysis accuracy.

[0150] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0151] The above descriptions all use HBA card fault diagnosis as an example. The HBA card fault diagnosis method provided in this application can also be widely applied to electrocardiogram (ECG) abnormality detection, such as converting HBA card CRC error surge detection into ST segment elevation recognition of ECG signals; and using an adaptive window adjustment mechanism to address heart rate variability in different patients. It can also be applied in the industrial Internet of Things (IoT) field, such as converting HBA card I / O delay variance into machine tool spindle vibration spectrum analysis in CNC machine tool vibration monitoring; and using frequency domain energy ratio to predict tool wear (abnormal 1-5Hz components indicate wear), etc.

[0152] Figure 5 A schematic diagram of the structure of an HBA card fault diagnosis device provided in this application embodiment includes an acquisition unit 51, an extraction unit 52, a filtering unit 53, a determination unit 54, and a generation unit 55; Acquisition unit 51 is used to acquire timing data from the HBA card; Extraction unit 52 is used to extract a time series feature set from time series data; wherein, the time series feature set includes time domain statistical features, frequency domain energy features, and event burst features; The filtering unit 53 is used to filter out multiple time-series features for fault diagnosis from the time-series feature set based on the contribution of each feature in the time-series feature set. The determination unit 54 is used to determine the fault label based on the correlation between multiple timing features and the set fault category; The generation unit 55 is used to generate fault response information based on fault labels and fault warning rules matched by multiple time-series features.

[0153] In some embodiments, the acquisition unit includes an acquisition subunit, a sampling subunit, and a cleaning subunit; The acquisition subunit is used to acquire raw timing data from the HBA card; the raw timing data includes hardware sensor data, performance indicators, and log information. The sampling subunit is used to sample the original time series data according to the dynamic window segmentation method to obtain multiple window time series data; The cleaning subunit is used to clean the time-series data from multiple windows to obtain the time-series data of the HBA card.

[0154] In some embodiments, the sampling subunit is used to extract the first window of timing data from the original timing data according to a set window length in the initial state; Adjust the window length of the next window based on the fluctuation index of the time series data in the previous window; Based on the window length and overlap rate of the next window, extract the window time series data of the next window from the original time series data, until all the original time series data has been extracted.

[0155] In some embodiments, the sampling subunit is used to reduce the length of the previous window as the window length of the next window according to a set shrinking rule, provided that the fluctuation index of the time series data in the previous window meets the data fluctuation condition. If the fluctuation index of the time series data in the previous window meets the data stability condition, the length of the previous window is increased according to the set extension rules to serve as the window length of the next window. If the fluctuation index of the time series data in the previous window does not meet the data fluctuation condition and the data stationarity condition, the length of the previous window shall be used as the length of the next window.

[0156] In some embodiments, the sampling subunit is used to determine the window length of the next window based on a set minimum window value, the length of the previous window, and its corresponding reduction ratio, when the standard deviation of the time series data in the previous window is greater than a first standard threshold and the absolute value of the slope of the time series data in the previous window is greater than a first slope threshold.

[0157] In some embodiments, the sampling subunit is used to determine the window length of the next window based on a set maximum window value, the length of the previous window and its corresponding extension ratio, when the standard deviation of the time series data in the previous window is less than or equal to a second standard threshold and the absolute value of the slope of the time series data in the previous window is less than or equal to a second slope threshold; wherein, the second standard threshold is less than a first standard threshold and the second slope threshold is less than a first slope threshold.

[0158] In some embodiments, the cleaning subunit is used to perform noise filtering on multiple window time series data according to a set noise filtering method to obtain filtered multiple window time series data. The first type of missing data in the filtered multi-window time series data is supplemented using linear interpolation to obtain the first supplemented multi-window time series data; wherein, the first type of missing data is data in which the number of consecutively missing sampling points is less than a set missing value; Based on the influence weight of historical data on missing data, the second type of missing data in the multiple window time series data after the initial supplementation is supplemented to obtain the final supplemented multiple window time series data; where the second type of missing data is data in which the number of consecutively missing sampling points is greater than or equal to the set missing value; The final supplemented timing data from multiple windows is normalized to obtain the timing data for the HBA card.

[0159] In some embodiments, the extraction unit is used to determine time-domain statistical features based on the distribution trend of time-series data within each window; wherein, the time-domain statistical features include any one or any combination of mean, variance, skewness, kurtosis, and slope; Convert the time-series data into a frequency domain signal; extract the main frequency component from the frequency domain signal to determine the main frequency energy ratio; Based on the correlation between performance indicators and failure events in time-series data, the suddenness characteristics of events are extracted.

[0160] In some embodiments, the screening unit includes a subunit, a contribution determination subunit, and a feature screening subunit; As a sub-unit, it is used to take the remaining features in the time series feature set excluding the target feature as the target subset; wherein, the target feature is any one of the features contained in the time series feature set; The contribution determination subunit is used to determine the contribution of the target feature based on the time series feature set, the target subset and their respective predicted values. The feature filtering subunit is used to filter a set number of time series features from the time series feature set in descending order of contribution.

[0161] In some embodiments, the contribution determination subunit is used to take the time series feature set and the target subset as input features of the fault classification model, respectively, to determine the predicted values ​​corresponding to the time series feature set and the target subset. According to the set weight calculation rules, the time series feature set and target subset are processed to obtain the weights; The difference between the predicted value corresponding to the time series feature set and the predicted value corresponding to the target subset is multiplied by the weight to obtain the contribution of the target feature.

[0162] In some embodiments, the determining unit is used to input multiple time-series features into a fault classification model to obtain the fault probability corresponding to each fault category; wherein, the fault classification model is trained based on historical sample data, and the historical sample data includes a historical time-series feature set and its corresponding fault categories; the fault categories include hardware faults, software faults and link faults; The fault category with the highest failure probability and its probability are used as fault labels.

[0163] In some embodiments, the generation unit includes a matching subunit, an analysis subunit, and a response subunit; The matching subunit is used to filter out target fault warning rules that match the fault label and target time series features from the set rule base; wherein, each fault warning rule in the rule base has its corresponding threshold range. The analysis subunit is used to analyze fault labels and target time-series characteristics according to the target fault early warning rules in order to determine the fault priority; The response subunit is used to determine the fault response information that matches the fault priority based on the response action mapping table; the fault response information includes the automatic response method and the notification method.

[0164] In some embodiments, the threshold range includes an error threshold and a temperature slope threshold; for adjusting the threshold range, the apparatus further includes a statistical unit, a first determining unit, a second determining unit, and an adjusting unit; The statistical unit is used to calculate the average and standard deviation of error counts within a specified time period. The first determining unit is used to determine the error threshold based on the mean and standard deviation; The second determining unit is used to determine the compensation coefficient based on the difference between the current room temperature and the standard room temperature. The adjustment unit is used to take the quotient of the temperature slope threshold and the compensation coefficient as the adjusted temperature slope threshold.

[0165] In some embodiments, the matching subunit is used to filter out fault warning rules that match the fault label and the target time series features from a set rule base according to the threshold range corresponding to each fault warning rule in the rule base; wherein, each fault warning rule has its corresponding rule priority; When multiple fault warning rules are matched, the fault warning rule with the highest rule priority is selected as the target fault warning rule. When there are multiple fault warning rules with the highest priority, the fault warning rule with the highest contribution is selected as the target fault warning rule.

[0166] In some embodiments, a priority adjustment unit is also included; The priority adjustment unit is used to adjust the priority of the target fault warning rule based on the number of times the target fault warning rule is triggered within a period of time.

[0167] In some embodiments, after generating fault response information based on fault labels and fault warning rules matched by multiple time-series features, the system further includes an analysis unit and a root cause determination unit. The analysis unit is used to input time-series data and its corresponding time-series feature set into the bidirectional long short-term memory network model to obtain the output results; the output results include the predicted probabilities of various faults and feature vectors for fault mode matching. The root cause determination unit is used to determine the root cause analysis results based on the feature similarity between the output results and each root cause rule in the fault mode library.

[0168] For a description of the features in the embodiment corresponding to the fault diagnosis device for HBA cards, please refer to the relevant description in the embodiment corresponding to the fault diagnosis method for HBA cards, which will not be repeated here.

[0169] As can be seen from the above technical solution, the process involves acquiring timing data from the HBA card and extracting a timing feature set from this data. This timing feature set includes multimodal features such as time-domain statistical features, frequency-domain energy features, and event burst features. Time-domain statistical features can capture the overall distribution and trend of the timing data, frequency-domain energy features can capture periodic faults in the timing data, and event burst features can capture millisecond-level sudden faults. The error frequency and I / O delay variance of the HBA card are the most important features for fault detection. The main frequency energy ratio and temperature slope are important reference values ​​for HBA card fault cause analysis. Signal strength fluctuation and link retry count are important parameters for HBA card link quality assessment. Therefore, the time-domain statistical features include one or more of I / O delay variance, temperature slope, and link retry count; the frequency-domain energy features include error frequency and / or main frequency energy ratio; and the event burst features include signal strength fluctuation. By extracting multimodal timing features, the detection sensitivity is improved, and reliable data support is provided for subsequent fault detection. The temporal feature set contains a large number of temporal features, but some of these features contribute very little to fault detection and are not worth analyzing. Therefore, to improve fault detection efficiency, multiple temporal features for fault diagnosis can be selected from the set based on their contribution. Fault labels are determined based on the correlation between multiple temporal features and defined fault categories. These labels include the HBA card's fault category and probability. Different fault categories are associated with different temporal features. To automate fault handling, all possible combinations of temporal features under different fault categories can be summarized to set corresponding fault warning rules. These rules can include corresponding fault handling methods. After selecting the temporal features and determining the fault labels, fault response information can be generated based on the fault labels and the fault warning rules matched by multiple temporal features. This fault response information includes the response actions required to resolve the current HBA card fault. In this technical solution, the detection sensitivity is improved by extracting multimodal temporal features. Based on the fault warning rules, multiple time-series features and fault labels are comprehensively analyzed to avoid misoperation caused by fluctuations in a single feature. Furthermore, based on the fault response information, faults can be automatically processed in a timely manner, improving fault handling efficiency.

[0170] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the fault diagnosis method for HBA cards.

[0171] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the fault diagnosis method for HBA cards when running.

[0172] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0173] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the fault diagnosis method for HBA cards.

[0174] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described HBA card fault diagnosis method embodiments.

[0175] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0176] The above provides a detailed description of a fault diagnosis method and electronic device for an HBA card provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A fault diagnosis method for an HBA card, characterized in that, include: Obtain timing data from the HBA card; Extract a time-series feature set from the time-series data; wherein the time-series feature set includes time-domain statistical features, frequency-domain energy features, and event burst features; the time-domain statistical features include one or more of I / O delay variance, temperature slope, and link retry count; the frequency-domain energy features include error frequency and / or main frequency energy ratio; the event burst features include signal strength fluctuations; Based on the contribution of each feature in the time-series feature set, multiple time-series features for fault diagnosis are selected from the time-series feature set. Fault labels are determined based on the correlation between multiple time-series characteristics and the set fault categories; Based on the fault labels and fault warning rules matched by multiple time-series features, fault response information is generated.

2. The fault diagnosis method for an HBA card according to claim 1, characterized in that, Obtain timing data from the HBA card, including: Collect raw timing data from the HBA card; wherein, the raw timing data includes hardware sensor data, performance indicators, and log information; The original time series data is sampled according to a dynamic window segmentation method to obtain multiple window time series data. Data cleaning is performed on the timing data of multiple windows to obtain the timing data of the HBA card.

3. The fault diagnosis method for an HBA card according to claim 2, characterized in that, The original time-series data is sampled according to a dynamic window segmentation method to obtain multiple window time-series data, including: In the initial state, the first window of timing data is extracted from the original timing data according to the set window length; Adjust the window length of the next window based on the fluctuation index of the time series data in the previous window; Based on the window length and overlap rate of the next window, extract the window timing data of the next window from the original timing data until all the original timing data has been extracted.

4. The fault diagnosis method for an HBA card according to claim 3, characterized in that, Based on the fluctuation indicators of the time series data in the previous window, adjust the window length of the next window, including: If the fluctuation index of the time series data in the previous window meets the data fluctuation conditions, the length of the previous window is reduced according to the set shrinking rules to serve as the window length of the next window. If the fluctuation index of the time series data in the previous window meets the data stability condition, the length of the previous window is increased according to the set extension rules to serve as the window length of the next window. If the fluctuation index of the time series data in the previous window does not meet the data fluctuation condition and the data stationarity condition, the length of the previous window shall be used as the length of the next window.

5. The fault diagnosis method for an HBA card according to claim 4, characterized in that, If the fluctuation indicators of the time series data in the previous window meet the data fluctuation conditions, the length of the previous window is reduced according to the set shrinking rules to serve as the window length of the next window, including: If the standard deviation of the time series data in the previous window is greater than the first standard threshold and the absolute value of the slope of the time series data in the previous window is greater than the first slope threshold, the window length of the next window is determined based on the set minimum window value, the length of the previous window and its corresponding reduction ratio.

6. The fault diagnosis method for an HBA card according to claim 5, characterized in that, If the fluctuation index of the time series data in the previous window meets the data stationarity condition, the length of the previous window is increased according to the set extension rules to serve as the window length of the next window, including: If the standard deviation of the time series data in the previous window is less than or equal to the second standard threshold and the absolute value of the slope of the time series data in the previous window is less than or equal to the second slope threshold, the window length of the next window is determined according to the set maximum window value, the length of the previous window and its corresponding extension ratio; wherein, the second standard threshold is less than the first standard threshold and the second slope threshold is less than the first slope threshold.

7. The fault diagnosis method for an HBA card according to claim 2, characterized in that, Data cleaning is performed on multiple window timing data to obtain the timing data of the HBA card, including: The noise filtering is performed on multiple window time series data according to the set noise filtering method to obtain filtered multiple window time series data; The first type of missing data in the filtered multi-window time series data is supplemented using linear interpolation to obtain the first supplemented multi-window time series data; wherein, the first type of missing data is data in which the number of consecutively missing sampling points is less than a set missing value; Based on the influence weight of historical data on missing data, the second type of missing data in the multiple window time series data after the initial supplementation is supplemented to obtain the final supplemented multiple window time series data; where the second type of missing data is data in which the number of consecutively missing sampling points is greater than or equal to the set missing value; The final supplemented timing data from multiple windows is normalized to obtain the timing data of the HBA card.

8. The fault diagnosis method for an HBA card according to claim 1, characterized in that, Based on the contribution of each feature in the time-series feature set, multiple time-series features for fault diagnosis are selected from the time-series feature set, including: The remaining features in the time-series feature set, excluding the target feature, are taken as the target subset; wherein, the target feature is any one of all features contained in the time-series feature set; The contribution of the target feature is determined based on the time-series feature set, the target subset and their respective predicted values. A set number of time-series features are selected from the time-series feature set according to their contribution levels from high to low.

9. The fault diagnosis method for an HBA card according to claim 8, characterized in that, Based on the temporal feature set, the target subset, and their respective predicted values, the contribution of the target feature is determined, including: The time-series feature set and the target subset are used as input features of the fault classification model to determine the predicted values ​​corresponding to the time-series feature set and the target subset, respectively. According to the set weight calculation rules, the time series feature set and the target subset are processed to obtain the weights; The difference between the predicted value corresponding to the time series feature set and the predicted value corresponding to the target subset is multiplied by the weight to obtain the contribution degree corresponding to the target feature.

10. The fault diagnosis method for an HBA card according to any one of claims 1 to 9, characterized in that, Based on the fault labels and fault warning rules matched by multiple time-series features, fault response information is generated, including: Target fault warning rules that match the fault label and target time-series features are selected from the set rule base; wherein, each fault warning rule in the rule base has its corresponding threshold range. The fault labels and target time-series characteristics are analyzed according to the target fault early warning rules to determine the fault priority. Based on the response action mapping table, the fault response information matching the fault priority is determined; wherein, the fault response information includes automated response methods and notification methods.

11. The fault diagnosis method for an HBA card according to claim 10, characterized in that, The threshold range includes the error threshold and the temperature slope threshold; Regarding the adjustment of the threshold range, the method further includes: Calculate the mean and standard deviation of error counts within a specified time period; An error threshold is determined based on the average value and the standard deviation. The compensation coefficient is determined based on the difference between the current room temperature and the standard room temperature. The quotient of the temperature slope threshold and the compensation coefficient is used as the adjusted temperature slope threshold.

12. The fault diagnosis method for an HBA card according to claim 10, characterized in that, Target fault warning rules that match the fault label and the target time-series characteristics are selected from the established rule base, including: Based on the threshold range corresponding to each fault warning rule in the rule base, fault warning rules that match the fault label and the target time-series features are selected from the set rule base; wherein, each fault warning rule has its corresponding rule priority. When multiple fault warning rules are matched, the fault warning rule with the highest rule priority is selected as the target fault warning rule. When there are multiple fault warning rules with the highest priority, the fault warning rule with the highest contribution is selected as the target fault warning rule.

13. The fault diagnosis method for an HBA card according to claim 12, characterized in that, Also includes: The priority of the target fault warning rule is adjusted based on the number of times the target fault warning rule is triggered within a period of time.

14. The fault diagnosis method for an HBA card according to claim 1, characterized in that, After generating fault response information based on the fault label and fault warning rules matched by multiple time-series features, the method further includes: The time series data and its corresponding time series feature set are input into a bidirectional long short-term memory network model to obtain the output results; wherein, the output results include the predicted probabilities of various faults and feature vectors for fault mode matching; The root cause analysis results are determined based on the feature similarity between the output results and each root cause rule in the fault mode library.

15. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the fault diagnosis method for the HBA card as described in any one of claims 1 to 14 when executing the computer program.

Citation Information

Patent Citations

  • Fault diagnosis method and device, computer equipment and storage medium

    CN116204768A

  • Detection method and device of host bus adapter, electronic equipment and storage medium

    CN118227418A

  • Internet of Things intelligent monitoring alarm method and system for water quality detection

    CN119107774A

  • Fault diagnosis method, device, medium and program product

    CN119847809A

  • Fault analysis method and device, storage medium and electronic equipment

    CN120358131A