Storage link fault diagnosis method, device, equipment, medium and program product

By collecting multi-source data of the storage link, building feature vectors and associated feature vectors, dynamically adjusting weights and fusion of evidence, the high false alarm rate problem of storage link fault diagnosis is solved, and more accurate and efficient fault diagnosis is achieved.

CN120353631BActive Publication Date: 2025-08-22INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510828093.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-08-22
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing storage link fault diagnosis methods rely on a single data source, resulting in high false alarm rates and incomplete diagnosis, making it difficult to distinguish faults caused by mutual cooperation between link levels.

Method used

Collect multi-source data on the storage link, build feature vectors and associated feature vectors of electronic devices, dynamically adjust the contribution weight, generate trust and likelihood through multi-source evidence fusion processing, and determine the fault type of the storage link.

Benefits of technology

It improves the accuracy and efficiency of fault diagnosis, reduces the false alarm rate, enhances the adaptability and reliability of the method, and can have a more comprehensive understanding of the operating status of the storage link.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353631B_ABST
    Figure CN120353631B_ABST
Patent Text Reader

Abstract

The present disclosure provides a storage link fault diagnosis method, apparatus, device, medium, and program product, relating to the field of computer technology. The method collects multi-source data from different electronic devices on a storage link; analyzes the multi-source data to construct feature vectors for each electronic device and associated feature vectors between the devices; adjusts the contribution weight corresponding to each feature vector based on the operating status of each electronic device; performs multi-source evidence fusion processing on the contribution weight corresponding to each feature vector and each feature vector to generate a confidence level and likelihood level corresponding to each electronic device; and determines the storage link fault type based on the confidence level and likelihood level corresponding to each electronic device. This method can improve fault diagnosis accuracy and efficiency and reduce false alarm rates.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a storage link fault diagnosis method, apparatus, device, medium, and program product. Background Art

[0002] With the explosive growth of data volumes and the increasing complexity of storage systems, storage links (including RAID (Redundant Array of Independent Disks) controllers, hard drives, backplanes, and so on) are becoming increasingly complex. For example, storage links employ multi-tiered architectures, leading to diverse fault types. Related technologies typically rely on a single data source, such as hard drive SMART (Self-Monitoring, Analysis, and Reporting Technology) logs or RAID status, resulting in a high false alarm rate. Summary of the Invention

[0003] The present disclosure provides a storage link fault diagnosis method, apparatus, device, medium, and program product, the main purpose of which is to solve the problem of high false alarm rate in the storage link fault diagnosis method.

[0004] According to a first aspect of the present disclosure, a storage link fault diagnosis method is provided, comprising:

[0005] Collect and store multi-source data from different electronic devices on the link;

[0006] Analyzing the multi-source data to construct a feature vector of each electronic device and a correlation feature vector between the electronic devices;

[0007] Adjusting the contribution weight corresponding to each eigenvector according to the operating status of each of the electronic devices;

[0008] Performing multi-source evidence fusion processing on the contribution weight corresponding to each eigenvector and each eigenvector to generate a trustworthiness and likelihood corresponding to each electronic device;

[0009] The fault type of the storage link is determined according to the trustworthiness and likelihood corresponding to each electronic component.

[0010] According to a second aspect of the present disclosure, a storage link fault diagnosis device is provided, comprising:

[0011] A data acquisition module is used to collect multi-source data from different electronic devices on the storage link;

[0012] A vector construction module, configured to analyze the multi-source data and construct a feature vector of each electronic device and a correlation feature vector between the electronic devices;

[0013] A weight adjustment module, configured to adjust the contribution weight corresponding to each eigenvector according to the operating status of each of the electronic devices;

[0014] a data fusion module, configured to perform multi-source evidence fusion processing on the contribution weight corresponding to each eigenvector and each eigenvector, and generate a trustworthiness and likelihood corresponding to each electronic device;

[0015] The fault diagnosis module is used to determine the fault type of the storage link according to the trust and likelihood corresponding to each electronic component.

[0016] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0017] at least one processor; and,

[0018] a memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.

[0020] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect.

[0021] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method as described in the first aspect above.

[0022] In the embodiment provided by the present disclosure, multi-source data of different electronic devices on the storage link is collected; the multi-source data is analyzed to construct the feature vectors of each electronic device and the associated feature vectors between each electronic device; the contribution weight corresponding to each feature vector is adjusted according to the operating status of each electronic device; the contribution weight corresponding to each feature vector and each feature vector are subjected to multi-source evidence fusion processing to generate the trust and likelihood corresponding to each electronic device; and the fault type of the storage link is determined according to the trust and likelihood corresponding to each electronic device. In this way, by fusing multi-source data of multiple electronic devices, dynamically adjusting the weights of the multi-source data and performing multi-source evidence fusion processing, the trust and likelihood of each electronic device are obtained, and fault diagnosis in the storage link is realized according to the trust and likelihood of each electronic device, a solution can be provided for fault diagnosis of the storage link, the accuracy and efficiency of fault diagnosis can be improved, and the false alarm rate can be reduced.

[0023] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0025] Figure 1 A flowchart of a storage link fault diagnosis method provided in the related art;

[0026] Figure 2 A flowchart of a storage link fault diagnosis method provided by an embodiment of the present disclosure;

[0027] Figure 3 A flowchart of another storage link fault diagnosis method provided by an embodiment of the present disclosure;

[0028] Figure 4 A flow chart of a dynamic weight allocation mechanism provided in an embodiment of the present disclosure;

[0029] Figure 5 A schematic diagram of a process for dynamically adjusting the conflict amount allocation method through a redistribution strategy provided by an embodiment of the present disclosure;

[0030] Figure 6 A flowchart of a decision-making mechanism provided in an embodiment of the present disclosure;

[0031] Figure 7 A schematic diagram of the structure of a storage link fault diagnosis device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0032] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0033] As can be seen from the background technology, traditional storage link fault diagnosis methods often use single-point diagnosis methods and usually rely on a single data source (such as the hard drive SMART log or RAID status), which leads to high false alarm rates and has limitations. For example, a hard drive response timeout may be caused by unstable backplane power supply, but it is mistakenly diagnosed as a hard drive failure; a RAID reconstruction failure may be caused by backplane signal interference, but it is mistakenly diagnosed as hard drive media damage, etc. Moreover, traditional fault repair relies on manual troubleshooting, which is inefficient. For example, a backplane failure requires checking the power module and signal line node by node; a hard drive failure requires checking the SMART log and media status of each disk.

[0034] For example, Figure 1 As shown, the storage link failure detection method in the related art generally includes: creating a monitoring virtual machine on each host node; mounting volumes corresponding to different storage backends to the corresponding monitoring virtual machine; obtaining health status detection data; and determining the data link status between the monitoring virtual machine and the storage backend based on the health status detection data. Creating a monitoring virtual machine on each host node can be: based on the monitoring server and through the Nova service, creating a monitoring virtual machine on each host node. Mounting volumes corresponding to different storage backends to the corresponding monitoring virtual machine includes: based on the monitoring server and through the Cinder service, mounting volumes corresponding to different storage backends to the corresponding monitoring virtual machine.

[0035] This technical solution relies on a single data source, making it difficult to distinguish between faults within the storage link caused by inter-layer coordination. Examples include hard drive performance degradation due to backplane power fluctuations, RAID degradation caused by physical hard drive damage, and CRC (Cyclic Redundancy Check) errors caused by signal interference on SAS (Serial Attached SCSI) links. Furthermore, while this technical solution describes diagnosing storage link faults, it lacks diagnostics for RAID cards and hard drive backplanes. This means it only diagnoses hard drive faults and lacks comprehensive coverage of the entire storage link.

[0036] The following describes the storage link fault diagnosis method, apparatus, device medium, and program product according to the embodiments of the present disclosure with reference to the accompanying drawings.

[0037] Figure 2 This is a flow chart of a storage link fault diagnosis method provided by an embodiment of the present disclosure. Figure 1 As shown, the method comprises the following steps:

[0038] Step 201: Collect multi-source data from different electronic devices on a storage link.

[0039] In this case, multi-source data from different electronic devices on the storage link can be collected, and there can be at least two electronic devices, that is, data from each electronic device can be collected separately. These data are the basis for subsequent fault diagnosis.

[0040] The electronic device may include at least a hard disk, a controller, and a backplane. The controller may be, for example, a RAID controller. The multi-source data may include data from the hard disk, data from the controller, and data from the backplane.

[0041] The hard drive data may include at least one of differential signal amplitude data, clock jitter data, bit error rate, spindle motor current ripple detection data, head cantilever vibration spectrum analysis data, disk medium temperature gradient monitoring data, voltage values, and hard drive log information; the controller data may include at least one of voltage values ​​and controller log information; and the backplane data may include at least one of clock jitter data, spindle motor current ripple detection data, and voltage values. For example, the hard drive and backplane data may be acquired using sensors.

[0042] As an example, it can capture in real time the differential signal amplitude, clock jitter, bit error rate, spindle motor current ripple detection (bandwidth DC-10MHz), head cantilever vibration spectrum analysis (0-5kHz FFT), and disk medium temperature gradient monitoring (4 PT100 sensors deployed on each side) on the storage link (including RAID cards (RAID controllers), hard drives, hard drive backplanes, etc.); monitor in real time the voltage values ​​of each unit module on the storage link, data on RAID controllers, hard drive sensors, and backplane sensors; and collect in real time log information on RAID cards and hard drive SMART log information.

[0043] Step 202 : Analyze multi-source data to construct feature vectors of each electronic device and associated feature vectors between electronic devices.

[0044] This approach analyzes collected multi-source data from different electronic devices, extracting features relevant to fault diagnosis and constructing a feature vector for each device. Furthermore, it calculates the associated feature vectors between different devices to account for their mutual influence. By constructing feature vectors and associated feature vectors, complex multi-source data can be transformed into quantifiable feature representations, providing structured data support for subsequent weight adjustment and evidence fusion.

[0045] Step 203: Adjust the contribution weight corresponding to each eigenvector according to the operating status of each electronic device.

[0046] The contribution weight of each eigenvector can be dynamically adjusted based on the operating state of each electronic device (such as load, temperature, and task type). Because changes in the operating state of an electronic device can affect the importance of a feature, the contribution weight of each eigenvector needs to be adjusted in real time to reflect the current operating environment. This dynamic weight adjustment allows the fault diagnosis method to better adapt to different operating conditions, improving diagnostic accuracy and adaptability.

[0047] Step 204 : Perform multi-source evidence fusion processing on the contribution weight corresponding to each eigenvector and each eigenvector to generate the trustworthiness and likelihood corresponding to each electronic device.

[0048] Multi-source evidence fusion methods (such as the Dempster-Shafer theory) can be used to combine each eigenvector and its corresponding contribution weight to generate a confidence and likelihood for each electronic component. The confidence and likelihood reflect the likelihood of each electronic component failing. For example, the confidence (Bel) represents the minimum level of support for a particular fault type (e.g., a fault in an electronic component); the likelihood (Pl) represents the maximum possible level of support for a particular fault type (e.g., a fault in an electronic component). Multi-source evidence fusion, by comprehensively considering different eigenvectors and their weights, generates more accurate confidence and likelihood, thereby improving the reliability of fault diagnosis.

[0049] Step 205: Determine the fault type of the storage link according to the trust and likelihood corresponding to each electronic component.

[0050] The storage link fault type can be determined based on the trust and likelihood of each electronic component. For example, if the trust and likelihood of a particular electronic component meet preset fault judgment criteria, the electronic component is determined to have failed. Examples of fault types include hard drive failure, RAID failure, and backplane failure.

[0051] In summary, in the embodiments provided by the present disclosure, multi-source data of different electronic devices on the storage link is collected; the multi-source data is analyzed to construct the feature vectors of each electronic device and the associated feature vectors between each electronic device; the contribution weight corresponding to each feature vector is adjusted according to the operating status of each electronic device; the contribution weight corresponding to each feature vector and each feature vector are subjected to multi-source evidence fusion processing to generate the trust and likelihood corresponding to each electronic device; and the fault type of the storage link is determined based on the trust and likelihood corresponding to each electronic device. In this way, by fusing multi-source data of multiple electronic devices, dynamically adjusting the weights of the multi-source data and performing multi-source evidence fusion processing on the multi-source data, the trust and likelihood of each electronic device are obtained, and fault diagnosis in the storage link is implemented based on the trust and likelihood of each electronic device. This can provide a solution for fault diagnosis of the storage link, improve the accuracy and efficiency of fault diagnosis, and reduce the false alarm rate. At the same time, the adaptability and reliability of the method can also be enhanced.

[0052] It should be noted that the embodiments of the present disclosure may include multiple steps. For the convenience of description, these steps are numbered, but these numbers do not limit the execution time slots or execution order between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not limit this.

[0053] Furthermore, in a possible implementation of this embodiment, analyzing multi-source data to construct feature vectors of each electronic device and associated feature vectors between electronic devices includes:

[0054] Input multi-source data into a preset fault correlation model;

[0055] By using a preset fault correlation model based on multi-source data, the correlation between various electronic components is quantified to obtain the characteristic vector of each electronic component;

[0056] Determining correlation feature vectors between the electronic devices based on the feature vectors of the electronic devices;

[0057] Based on the feature vectors of each electronic device and the associated feature vectors between the electronic devices, a feature vector of the multi-source data is constructed.

[0058] The multi-source data collected from different electronic devices on the storage link can be input into a preset fault correlation model. The preset fault correlation model can be a pre-trained model used to generate feature vectors for each electronic device. The preset fault correlation model can then be used to analyze and process the input multi-source data to quantify the correlations between the various electronic devices. For example, by analyzing correlations and causal relationships between the data, a feature vector can be generated for each electronic device. These feature vectors reflect the operating status and fault characteristics of each electronic device. Subsequently, based on the feature vectors of each electronic device, correlation feature vectors between the various electronic devices can be further calculated. These correlation feature vectors reflect the interdependencies between different electronic devices and help understand the fault propagation path. Finally, the feature vectors of each electronic device and the correlation feature vectors can be combined to construct a final multi-source data feature vector. This feature vector combines the feature vectors of all electronic devices and the correlation feature vectors between them, providing a comprehensive feature representation for subsequent fault diagnosis. In this way, by constructing feature vectors and correlation feature vectors, the correlations between electronic devices are quantified, and a comprehensive feature vector is generated. This not only improves the accuracy of fault diagnosis results, but also enhances the adaptability and reliability of the system, thereby enabling a more comprehensive understanding of the operating status of the storage link and achieving accurate diagnosis of faults.

[0059] As a specific example, data collected from storage links (multi-source data) can be analyzed using a fault correlation model (pre-set fault correlation model). This fault correlation model can establish a cross-level fault causal chain by quantifying the dynamic correlations between the RAID controller, hard drive hardware, and backplane physical states. First, a feature vector can be defined. The data dimensions required for defining the feature vector are shown in the following table:

[0060]

[0061] Among them, 1) the formula for constructing the characteristic vector of the RAID controller is as follows:

[0062]

[0063] Uncorrectable_Errors indicates the number of errors detected but uncorrectable by the RAID controller, Total_IO indicates the total number of I / O requests (including read and write operations), and Reconstruct_Time indicates the reconstruction time, which is the time it takes for the RAID system to recover (rebuild) the data on a hard drive using redundant data after a hard drive in the RAID array fails. Indicates the uncorrectable error rate. If the value is high, the data integrity of the RAID controller is seriously threatened, which may lead to data loss. If the value is low, it means that the RAID controller is operating normally and the data reliability is high.

[0064] 2) The formula for constructing the hard disk sensor's characteristic vector is as follows:

[0065]

[0066] Among them, Seek_Error_Rate refers to the error frequency that occurs when the head seeks a specific data location. It is one of the hard drive SMART (Self-Monitoring, Analysis and Reporting Technology) attributes and is used to evaluate the mechanical performance of the hard drive. Raw_Vibration represents the vibration spectrum, that is, the vibration characteristics of the spindle motor or head suspension. Temperature represents the hard drive platter temperature.

[0067] 3) The formula for constructing the feature vector related to the hard disk backplane is as follows:

[0068]

[0069] Voltage_Ripple represents the power supply voltage fluctuation amplitude; Threshold represents the temperature difference between different areas of the backplane; UI represents the unit interval; and Jitter_RMS represents a key indicator of signal timing jitter, which represents the root mean square deviation of the signal edge relative to the ideal clock edge. The calculation formula is:

[0070]

[0071] Among them, TIE i Indicates the timing deviation of the i-th signal edge; μ TIE represents the mean value of TIE; N represents the total number of signal edges.

[0072] 4) The formula for constructing cross-level correlation feature vectors is as follows:

[0073]

[0074] This formula is used to quantify the correlation between the hard disk layer (HDD) and the backplane layer (BP), and is weighted by the uncertainty of the RAID log to capture the potential pattern of cross-level failures. When the hard disk and backplane characteristics are highly correlated, the Corr value is large, indicating that there may be cross-level failures. When the RAID log uncertainty is high (high entropy), f Cross As the value increases, the sensitivity to cross-level failures increases. Represents the hard disk layer feature f HDD and backplane layer characteristics fBP The correlation coefficient is used to measure the degree of linear correlation between the two, and the calculation formula is:

[0075]

[0076] Among them, Corr(f HDD , f BP ) represents the covariance, which measures the co-variation of the two; and f HDD and f BP When Corr(f HDD , f BP ) is close to 1: the hard drive and backplane characteristics are highly positively correlated (for example, unstable backplane power supply leads to degraded hard drive performance); when Corr(f HDD , f BP ) is close to -1: highly negatively correlated (e.g., an increase in backplane temperature causes a decrease in hard drive speed); when close to Corr(f HDD , f BP )0: No obvious linear relationship.

[0077] Entropy(RAID_Logs) represents the RAID log information entropy, which is used to quantify the uncertainty of the RAID log. The calculation formula is:

[0078]

[0079] Among them, p i represents the probability of occurrence of the i-th type of log event (such as CRC error, reconstruction failure, etc.); N represents the total number of log event types.

[0080] In summary, the feature vector of multi-source data is constructed as follows:

[0081]

[0082] Furthermore, in a possible implementation of this embodiment, a preset fault correlation model is used to quantify the correlation between the electronic components based on multi-source data to obtain a feature vector of each electronic component, including:

[0083] By using a preset fault correlation model based on multi-source data, the correlation between the hard disk, controller, and backplane is quantified to obtain the feature vectors of the hard disk, controller, and backplane.

[0084] Based on the characteristic vectors of the respective electronic devices, determining the associated characteristic vectors between the respective electronic devices, including:

[0085] Based on the feature vector of the hard disk and the feature vector of the backplane, a correlation feature vector between the hard disk and the backplane is constructed;

[0086] Based on the feature vectors of each electronic device and the associated feature vectors between the electronic devices, the feature vectors of multi-source data are constructed, including:

[0087] Based on the feature vectors of the hard disk, the controller, the backplane, and the associated feature vectors between the hard disk and the backplane, a feature vector of the multi-source data is constructed.

[0088] Among them, if the electronic device includes at least a hard disk, a controller and a backplane. Then, the correlation between the hard disk, the controller and the backplane can be quantified based on multi-source data through a preset fault correlation model to obtain the feature vector of the hard disk, the feature vector of the controller and the feature vector of the backplane. Then, based on the feature vector of the hard disk and the feature vector of the backplane, the correlation feature vector between the hard disk and the backplane can be constructed. After that, based on the feature vector of the hard disk, the feature vector of the controller and the feature vector of the backplane, as well as the correlation feature vector between the hard disk and the backplane, the feature vector of the multi-source data can be constructed. The specific implementation of each step in this embodiment is similar to that of the previous embodiment and will not be repeated here.

[0089] Furthermore, in a possible implementation of this embodiment, adjusting the contribution weight corresponding to each eigenvector according to the operating status of each electronic device includes:

[0090] Obtaining the system operating status; wherein the system operating status includes at least the load status, temperature, and task type of each electronic component;

[0091] Determine the feature significance score based on the historical normal baseline, trend sensitivity coefficient and cross-feature correlation weight;

[0092] Adjust the characteristic type sensitivity coefficient according to the system operation status and the preset sensitivity coefficient matrix;

[0093] Input the feature type sensitivity coefficient and feature significance score into the preset weight factor calculation model;

[0094] Through the preset weight factor calculation model, based on the feature type sensitivity coefficient and feature significance score, the normalized weight corresponding to each feature vector is generated;

[0095] By analyzing historical fault data, the contribution of each system's operating status to fault diagnosis is calculated;

[0096] Based on the contribution of each system operating state to fault diagnosis, the normalized weight of each eigenvector is adjusted to obtain the contribution weight corresponding to each eigenvector.

[0097] The operating status of the storage link system can be obtained, including the load status, temperature, task type, etc. of each electronic device. These operating statuses can reflect the current operating environment and working conditions of the system.

[0098] The operating status of each electronic component in the storage link system can be obtained, such as its load status, temperature, and task type. Load status can include, for example, normal load (e.g., IOPS < 5k) or high load (e.g., IOPS > 10k), temperature can include high temperature warnings (e.g., T > 60°C), and task types can include RAID reconstruction. IOPS (Input / Output Operations Per Second) is a key metric for measuring storage system performance, describing the number of input / output operations a storage device (such as a hard drive, solid-state drive, or storage array) can complete per unit time. A historical normal baseline, trend sensitivity coefficient, and cross-feature correlation weight can then be obtained. A feature significance score is then calculated based on these historical normal baselines, trend sensitivity coefficients, and cross-feature correlation weights. The historical normal baseline can be, for example, a rolling average over a certain period of time (e.g., 30 days). The trend sensitivity coefficient can represent the amplification effect of a sudden increase or decrease and can be a constant, such as 0.7. The cross-feature correlation weight can also be a constant, such as 0.4. The feature significance score can be calculated as follows:

[0099]

[0100] in, : Rolling 30-day average, representing the historical normal baseline; : Constant 0.7, trend sensitivity coefficient; : Constant 0.4, cross-feature correlation weight; : Cross-level correlation feature vector.

[0101] Afterwards, the feature type sensitivity coefficient can be adjusted based on the system operating state and a preset sensitivity coefficient matrix. The preset sensitivity coefficient matrix can be a predefined quantified matrix that quantifies the sensitivity of features to different operating states. Based on the currently acquired system operating state (load, temperature, task type, etc.), combined with the preset sensitivity coefficient matrix, the sensitivity coefficient of each feature type can be dynamically adjusted to better reflect the current operating conditions. An example of a preset sensitivity coefficient matrix is ​​shown in the table below. When multiple system operating states are triggered simultaneously, the feature type sensitivity coefficient α for each state can be adjusted using a weighted sum.

[0102]

[0103] The determined feature significance score and the adjusted feature type sensitivity coefficient are then used as input parameters and input into a preset weight factor calculation model. The weight factor calculation model is a preset mathematical model that comprehensively considers the feature type sensitivity coefficient and the feature significance score. The preset weight factor calculation model is used to normalize the feature type sensitivity coefficient and the feature significance score so that they are within a certain range (e.g., 0 to 1) to facilitate subsequent comparison and application. The weight factor calculation model can be, for example, as follows:

[0104]

[0105] in: : The normalized weight of the i-th feature at time t ( ); : Feature type sensitivity coefficient (preset empirical value); : Feature significance score (calculated in real time).

[0106] Afterward, we can obtain historical fault data from the system's operating status during past faults. This historical data can be analyzed to determine the importance of different operating states (such as high load, high temperature, and specific task types) in fault diagnosis and quantify their contributions. Finally, based on the contribution of each system operating state to fault diagnosis, we adjust the normalized weight of each eigenvector to obtain the corresponding contribution weight for each eigenvector. For example, we can further adjust the generated normalized weight based on the statistical contribution, ultimately obtaining the corresponding contribution weight for each eigenvector in fault diagnosis, which better meets the needs of actual fault diagnosis.

[0107] As an example, by analyzing historical data, we can calculate the contribution of each state to fault diagnosis and dynamically adjust the weight as follows:

[0108]

[0109] in: : Contribution of the jth state in historical faults.

[0110] For example: When RAID reconstruction and high temperature alarm are triggered at the same time, and , ,So The values ​​are:

[0111]

[0112] Furthermore, in some possible implementations, based on the contribution of each system operating state to fault diagnosis, the normalized weight of each eigenvector is adjusted to obtain the contribution weight corresponding to each eigenvector, including:

[0113] Perform median filtering on the derivative of each eigenvector;

[0114] The exponentially weighted moving average method is used to smooth the normalized weight of each eigenvector after median filtering to obtain the contribution weight corresponding to each eigenvector.

[0115] In an embodiment of the present application, an outlier filtering mechanism is added to the anti-interference design, that is, median filtering is performed on the derivative terms in Si, as follows:

[0116]

[0117] This eliminates the influence of outliers.

[0118] In addition, to prevent sudden changes in weights, the weights are smoothed using an exponentially weighted moving average, as follows:

[0119]

[0120] Furthermore, in some possible implementations, multi-source evidence fusion processing is performed on the contribution weight corresponding to each eigenvector and each eigenvector to generate the trust and likelihood corresponding to each electronic device, including:

[0121] Based on each eigenvector and the contribution weight corresponding to each eigenvector, the confidence and likelihood of each electronic component failure are generated through an adaptive membership function;

[0122] The redistribution factor is used to fuse the confidence and likelihood of each electronic component failure to generate the corresponding confidence and likelihood of each electronic component.

[0123] Each eigenvector and its corresponding contribution weight can be used as input. A multi-source evidence fusion method comprehensively considers the information from different eigenvectors to generate the confidence and likelihood of each electronic device failure. Based on each eigenvector and its corresponding contribution weight, an adaptive membership function is then used to generate the confidence and likelihood of each electronic device failure. The adaptive membership function is a dynamically adjusted function that calculates the confidence and likelihood of each electronic device failure based on the eigenvector and contribution weight. Belief (hereinafter abbreviated as Bel) represents the minimum level of support for a proposition, while plausibility (hereinafter abbreviated as Pl) represents the maximum possible level of support for a proposition. The adaptive membership function maps eigenvectors and contribution weights to confidence and likelihood, providing a foundation for subsequent fusion processing.

[0124] Among them, the redistribution factor is used to adjust the weights of the trust and likelihood to ensure that the fusion results are more reasonable. Fusion processing: The trust and likelihood of each electronic device are weighted and summed using the redistribution factor or other fusion methods to generate the final trust and likelihood. For example, each eigenvector and its corresponding contribution weight are used as input, and the adaptive membership function is used to calculate the trust and likelihood of each electronic device based on the eigenvector and contribution weight. The trust and likelihood are then weighted and fused using the redistribution factor to generate the final trust and likelihood of each electronic device. Through fusion processing, the contribution of different eigenvectors to fault diagnosis is comprehensively considered to generate the final trust and likelihood of each electronic device, which is used to more accurately assess the failure risk of electronic devices.

[0125] Furthermore, a redistribution factor is used to fuse the confidence and likelihood of each electronic component failure to generate the corresponding confidence and likelihood of each electronic component, including:

[0126] determining a conflict amount based on the confidence level of each electronic component failing;

[0127] Based on the conflict amount, the redistribution factor is calculated using the preset multi-evidence theory algorithm;

[0128] The confidence level and likelihood level corresponding to each electronic component are determined based on the redistribution factor, the first redistribution factor allocation ratio, and the confidence level of each electronic component that a failure has occurred.

[0129] And, based on the distribution ratio of the redistribution factor and the second redistribution factor, determining the confidence level of the global uncertainty when the fault type cannot be determined.

[0130] The amount of conflict can be calculated by comparing the differences between the trust levels generated by different feature vectors. For example, the following formula can be used:

[0131]

[0132] Where m1, m2: basic probability distribution (confidence) of two independent pieces of evidence (different feature vectors); A, B: two mutually exclusive fault types (such as "backplane fault" and "hard drive fault"). For example: Evidence 1 believes that the probability of "backplane fault" is , Evidence 2 suggests that the probability of “hard disk failure” is ; If A and B are mutually exclusive, then the conflict quantity .

[0133] Among them, multi-evidence theory algorithms, such as the Dempster-Shafer theory (DS theory), are used to handle uncertainty and conflicting information. The redistribution factor is a factor calculated by a preset multi-evidence theory algorithm based on the amount of conflict, and is used to adjust the distribution of trust and likelihood. Based on the amount of conflict K, the redistribution factor is calculated using the combination rules in DS theory (such as the Dempster rule) as follows:

[0134]

[0135] in: : represents the joint information entropy of evidence, used to measure uncertainty,

[0136]

[0137] When the value of Entropy is large, it means that there is high uncertainty in the evidence. At this time, the evidence is confusing, and the amount of conflict redistribution will be reduced to reduce the risk of misjudgment; when the value of Entropy is small, it means that the evidence is clear, and the amount of conflict redistribution will be increased to enhance the conflict handling ability.

[0138] The first redistribution factor allocation ratio is a preset ratio used to allocate the impact of the redistribution factor on the confidence and likelihood, for example, 40%. Fusion processing is used to adjust the confidence and likelihood of each electronic component using the redistribution factor and the first redistribution factor allocation ratio to generate the final confidence and likelihood. Furthermore, the second redistribution factor allocation ratio is a preset ratio used to allocate the impact of the redistribution factor on global uncertainty, for example, 60%. Global uncertainty confidence: When the fault type cannot be clearly determined, the confidence level represents the overall uncertainty of the system.

[0139] Furthermore, it also includes:

[0140] If the conflict amount is greater than the first conflict amount threshold, loading a preset weight allocation rule, and adjusting the trust level corresponding to the electronic device based on the preset weight allocation rule;

[0141] If the conflict amount is greater than the second conflict amount threshold, the confidence level of the global uncertainty is increased according to a preset adjustment method.

[0142] The first conflict threshold is a preset threshold used to determine whether the conflict level requires a confidence adjustment. Conditional judgment: When the calculated conflict level K is greater than the first conflict threshold, it indicates that the conflict between different pieces of evidence is severe and requires confidence adjustment. Preset weight assignment rules are a set of predefined rules used to adjust confidence allocation when the conflict level is high. Rule content may include redistributing the contribution weights of different feature vectors or making specific adjustments to confidence. After loading the preset weight assignment rules, the confidence of each electronic component can be adjusted based on the preset weight assignment rules to reduce the impact of the conflict level. As an example, the preset weight assignment rules can be expert rules stored in a database. For example, if the conflict level is greater than the first conflict threshold, such as 0.5, a predefined expert rule library (the rules in the expert library are compiled based on troubleshooting experience) can be activated, matching expert rules loaded from the database, and confidence allocation adjusted according to the rule weights.

[0143] Among them, the second conflict threshold: another preset threshold, used to determine whether the conflict amount has reached a level that requires an increase in the confidence level of global uncertainty, such as 0.8. Conditional judgment: When the calculated conflict amount K is greater than the second conflict threshold, it indicates that the contradiction between different evidences is very serious and the confidence level of global uncertainty needs to be increased. Preset adjustment method: a predefined adjustment method used to increase the confidence level of global uncertainty when the conflict amount is extremely large. If the conflict amount is greater than the second conflict threshold, the confidence level of global uncertainty is increased according to the preset adjustment method to reflect the overall uncertainty of the system. For example, when the conflict amount K>0.8, the system automatically increases the confidence level of the global uncertain proposition, as follows:

[0144]

[0145] By setting two conflict thresholds, each corresponding to a different adjustment strategy, we can effectively handle conflict situations of varying levels. When the conflict is high, the confidence level is adjusted using a pre-set weight distribution rule to reduce the impact of the conflict. When the conflict is extremely high, the confidence level is increased to account for global uncertainty, reflecting the overall uncertainty of the system. This approach can improve the accuracy and reliability of fault diagnosis, especially when faced with highly uncertain and conflicting information.

[0146] As an example, to address the problem of evidence synthesis in high-conflict scenarios and ensure the rationality and stability of the fusion results, a redistribution strategy is used to dynamically adjust the distribution of conflict quantities. For example: 1) Redistribute the conflict quantity K: Distribute K proportionally between the globally uncertain proposition and the high-probability single-fault proposition. 2) Uncertainty compensation: When the conflict quantity K is large, increase the confidence level of the globally uncertain proposition to avoid blindly trusting a single piece of evidence.

[0147] The redistribution formula is as follows:

[0148]

[0149] : Redistribution factor:

[0150]

[0151] In the specific redistribution strategy, the allocation can be carried out according to the following proportions:

[0152] 1) 60% is allocated to global uncertainty propositions; 2) 40% is allocated to high-probability single fault propositions according to weight.

[0153] Furthermore, the fault type of the storage link is determined based on the trust and likelihood corresponding to each electronic component, including:

[0154] For any fault type A, if the trust of fault type A is higher than the preset trust threshold, and the difference between the likelihood and the trust of fault type A is less than the first preset difference threshold, and the ratio of the trust of fault type A to the trust of global uncertainty is greater than the preset ratio threshold, then the fault type of the storage link is determined to be fault type A; wherein, fault type A is a hard disk failure, a controller failure, a backplane failure, or the fault type cannot be determined.

[0155] Among them, the confidence level of fault type A (Bel(A)) is higher than the preset confidence level threshold, the difference between the likelihood of fault type A Pl(A) and the confidence level is less than the first preset difference threshold, and the ratio of the confidence level of fault type A to the confidence level of global uncertainty Bel() is greater than the preset ratio threshold. This is a preset decision-making mechanism. If a proposition A (e.g., a certain fault type A) satisfies this decision-making mechanism, then fault type A is considered to be established. As an example: If a proposition A (e.g., a certain fault type A) meets the following three conditions simultaneously, proposition A is determined to be faulty:

[0156] 1) Bel(A)>threshold (preset trust threshold, initial value is 0.9);

[0157] 2) Pl(A) - Bel(A) < threshold (first preset difference threshold, initial value is 0.1);

[0158] 3) Bel(A) / Bel() > threshold (preset ratio threshold, initial value is 3.0).

[0159] Furthermore, it also includes:

[0160] If the difference between the likelihood and the confidence of fault type A is greater than or equal to a first preset difference threshold, determining whether the storage link is in a high-load scenario based on the system operating status;

[0161] If the storage link is in a high-load scenario, the preset trust threshold is lowered according to the preset adjustment rules;

[0162] If the storage link is in a low-load scenario, the preset trust threshold is increased according to the preset adjustment rules.

[0163] If the difference between the likelihood and confidence of fault type A is greater than or equal to a first preset difference threshold, for example, when Pl(A) - Bel(A) is large, indicating high uncertainty in proposition A, further analysis is required. The decision threshold should be dynamically adjusted based on the system state. In high-load scenarios, the threshold should be lowered to increase sensitivity; in low-load scenarios, the threshold should be raised to reduce false positives. For example, in high-load scenarios, the threshold can be adjusted from 0.9 to 0.8. In low-load scenarios, the threshold can be adjusted from 0.9 to 0.95.

[0164] Furthermore, it also includes:

[0165] If the difference between the likelihood and the confidence of fault type A is greater than a second preset difference threshold, the expert system is called to determine the fault type of the storage link based on historical fault data.

[0166] If the difference between the likelihood and confidence of fault type A exceeds a second preset difference threshold, indicating high uncertainty, an expert system can be invoked for in-depth analysis. Based on historical data (historical fault data) and expert experience, the expert system can be used to identify the most likely fault type. For example, if Pl(A) − Bel(A) > 0.2, due to high uncertainty, the expert system needs to be invoked for further analysis and to provide the most likely diagnosis.

[0167] Furthermore, it also includes:

[0168] The misjudgment rate of the storage link fault diagnosis method is evaluated by Monte Carlo simulation method.

[0169] Among them, the Monte Carlo simulation method can be used to evaluate the compensation effect of the uncertainty compensation mechanism and evaluate the storage link fault diagnosis method. As an example, the evaluation results are as follows:

[0170]

[0171] Furthermore, before analyzing multi-source data and constructing the feature vectors of each electronic device and the associated feature vectors between the electronic devices, the following steps are also included:

[0172] Perform data preprocessing on multi-source data; wherein data preprocessing includes at least one of data cleaning, noise removal, missing value and outlier processing, normalization processing, and data dimensionality reduction processing.

[0173] Data cleaning can remove errors, duplications, or inconsistencies in data to ensure data quality, such as removing duplicate data, correcting erroneous data, and standardizing data formats. Noise removal can reduce random errors or interference in data and improve data purity. For example, methods such as moving average and median filtering can be used to smooth data, wavelet transforms can be used to remove high-frequency noise, and statistical analysis can be used to identify and remove noise points that do not conform to the data distribution pattern. Missing data can be addressed to avoid analytical bias caused by missing values. For example, missing values ​​can be deleted and filled with the mean, median, mode, or model-based predictions. Linear interpolation and spline interpolation are used for missing value handling in time series data. Outlier handling can identify and address abnormal data points that may significantly affect analytical results. For example, statistical methods such as standard deviation and interquartile range (IQR) can be used to detect outliers; clustering algorithms or rule-based methods can be used to detect outliers; and outliers can be deleted, corrected, or labeled for subsequent analysis. Normalization can convert data to a uniform scale to prevent the impact of dimensional differences between different features on analytical results. Data dimensionality reduction can reduce data dimensions, lower computational complexity, and remove redundant information. Through data cleaning, noise removal, handling missing values ​​and outliers, normalization, and data dimensionality reduction, data quality and usability can be significantly improved, providing a solid foundation for subsequent analysis and modeling.

[0174] To make the storage link fault diagnosis method provided by the embodiment of the present disclosure clearer, the following is described with reference to a specific example. The method includes the following processing:

[0175] 1.1.1 Combination Figure 3It can capture the differential signal amplitude, clock jitter, bit error rate, spindle motor current ripple detection (bandwidth DC-10MHz), head cantilever vibration spectrum analysis (0-5kHz FFT), and disk medium temperature gradient monitoring (4 PT100 sensors deployed on each side) on the storage link (including RAID cards, hard drives, hard drive backplanes, etc.) in real time; monitor the voltage values ​​of each unit module on the storage link, data on the RAID controller, hard drive sensors, and backplane sensors in real time; and collect log information on the RAID card and hard drive SMART log information in real time.

[0176] 1.1.2 Analyze the collected data on the storage link using a fault correlation model. This model quantifies the dynamic correlation between the physical states of the RAID controller, hard drive hardware, and backplane to establish a cross-level fault causal chain. First, define a feature vector. The data dimensions required for defining the feature vector are shown in the following table:

[0177]

[0178] 1) The formula for constructing RAID-related feature vectors is as follows:

[0179]

[0180] Uncorrectable_Errors represents the number of errors detected but uncorrectable by the RAID controller, Total_IO represents the total number of I / O requests (including read and write operations), and Reconstruct_Time represents the reconstruction time, which is the time required for the RAID system to recover (rebuild) the data on a hard drive using redundant data after a hard drive in the RAID array fails. Represents the uncorrectable error rate. If the value is high, it means that the data integrity of the RAID system is seriously threatened, which may lead to data loss. If the value is low, it means that the RAID system is operating normally and the data reliability is high.

[0181] 2) The formula for constructing the feature vector extracted by the hard disk sensor is as follows:

[0182]

[0183] The Seek_Error_Rate (seek error rate) indicates the frequency of errors when the drive head seeks to a specific data location. It is one of the hard drive's SMART (Self-Monitoring, Analysis, and Reporting Technology) attributes used to assess the drive's mechanical performance. Raw_Vibration represents the vibration spectrum, or the vibration characteristics of the spindle motor or head suspension. Temperature represents the hard drive platter temperature.

[0184] 3) The formula for constructing the feature vector related to the hard disk backplane is as follows:

[0185]

[0186] Where Voltage_Ripple represents the power supply voltage fluctuation amplitude; Threshold represents the temperature difference between different areas of the backplane; UI stands for Unit Interval; and Jitter_RMS represents a key indicator of signal timing jitter, indicating the root mean square deviation of the signal edge relative to the ideal clock edge. The calculation formula is:

[0187]

[0188] in: : Timing deviation of the i-th signal edge; : the mean value of TIE; N: the total number of signal edges.

[0189] 4) The formula for constructing cross-level correlation feature vectors is as follows:

[0190]

[0191] This formula is used to quantify the correlation between the hard disk layer (HDD) and the backplane layer (BP), and is weighted by the uncertainty of the RAID log to capture the potential pattern of cross-level failures. When the hard disk and backplane features are highly correlated, the Corr value is large, indicating that there may be cross-level failures. When the RAID log uncertainty is high (high entropy), f Cross As the value increases, the sensitivity to cross-level failures increases.

[0192] The correlation coefficient between the hard disk layer characteristics and the backplane layer characteristics is used to measure the degree of linear correlation between the two. The calculation formula is:

[0193]

[0194] Where: Cov(f HDD , f BP ): covariance, measuring the co-variation of the two; and :respectively f HDD and f BP The standard deviation of .

[0195] When Cov(f HDD , f BP ) is close to 1: the hard drive and backplane characteristics are highly positively correlated (for example, unstable backplane power supply leads to degraded hard drive performance); when Cov(f HDD , f BP) is close to -1: highly negatively correlated (e.g., an increase in backplane temperature causes the hard drive to slow down); when Cov(f HDD , f BP ) is close to 0: there is no obvious linear relationship.

[0196] Entropy(RAID_Logs) represents the RAID log information entropy, which is used to quantify the uncertainty of the RAID log. The calculation formula is:

[0197]

[0198] Where: p i : The probability of occurrence of the i-th type of log event (such as CRC error, reconstruction failure, etc.); N: The total number of log event types.

[0199] To sum up, the feature vector is constructed as follows:

[0200]

[0201] 1.1.3 Using a dynamic weight allocation mechanism to perceive the system status (load, temperature, task type, etc.) in real time, dynamically adjust the contribution weight of each monitoring dimension to achieve accurate adaptation of fault diagnosis. The specific implementation method is as follows: Figure 4 As shown:

[0202] 1) Weight factor calculation model

[0203]

[0204] in: : The normalized weight of the i-th feature at time t ( ); : Feature type sensitivity coefficient (preset empirical value); : Feature significance score (calculated in real time).

[0205] 2) Feature significance score calculation formula:

[0206]

[0207] in, : Rolling 30-day average, representing the historical normal baseline; : Constant 0.7, trend sensitivity coefficient (amplification effect of sudden increase / sudden decrease); : Constant 0.4, cross-feature correlation weight; : Cross-level correlation feature vector.

[0208] 3) Dynamic adjustment strategy of sensitivity coefficient

[0209] The preset sensitivity coefficient matrix is ​​shown in the following table:

[0210]

[0211] When the system triggers multiple states at the same time, the coefficients of each state are adjusted by weighted sum. By analyzing historical data, the contribution of each state to fault diagnosis is calculated and the weights are adjusted dynamically:

[0212]

[0213] in: : Contribution of the jth state in historical faults.

[0214] For example: When RAID reconstruction and high temperature alarm are triggered at the same time, and , ,So The values ​​are:

[0215] .

[0216] 4) Anti-interference design: add an outlier filtering mechanism to the anti-interference design, that is: The derivative term in is median filtered as follows to eliminate the influence of outliers.

[0217]

[0218] 5) Weight smoothing

[0219] To prevent weight In case of sudden changes, the weights are smoothed using exponentially weighted moving average. The processing method is as follows:

[0220]

[0221] 1.1.4 A multi-source evidence fusion algorithm can be used to achieve effective cross-level data fusion. This algorithm is based on the improved Dempster-Shafer (DS) evidence theory and solves the failure problem of traditional methods in high-conflict scenarios. Assume that the set of fault types is:

[0222]

[0223] Among them: A1: backplane power failure; A2: SAS link signal integrity failure; A3: hard disk media damage; A4: RAID controller logic error; A5: cross-level coupling failure (such as signal interference caused by vibration).

[0224] Power set space Contains all possible fault combinations (total elements), for example: combination Indicates that both the backplane power supply and the hard disk media are faulty.

[0225] 1.1.5 For each monitoring feature, generate its confidence in each proposition through adaptive membership function:

[0226]

[0227] in: : Feature weights output by the dynamic weight allocation module; : Fault type sensitivity coefficient (trained by historical data); : decision threshold of type j; : Constant value , mainly to prevent division by 0 errors.

[0228] 1.1.6 In complex storage systems, multiple sensors or monitoring modules may provide conflicting evidence (for example, a backplane sensor may indicate normal voltage, but a hard drive may frequently report errors). Traditional DS evidence theory can produce counterintuitive results when dealing with high-conflict scenarios (for example, forcibly assigning high conflict to uncertain propositions). To address this issue, this embodiment introduces a conflict redistribution factor and an uncertainty compensation mechanism. By dynamically adjusting the conflict allocation strategy, it improves the rationality of the fusion results and the robustness of the system.

[0229] 1.1.7 In the conflict redistribution factor, the conflict quantity K represents the degree of conflict between different pieces of evidence and is the core indicator of DS theory:

[0230]

[0231] in, , : Basic probability distribution of two independent pieces of evidence; A, B: Two non-intersecting propositions (such as "backplane failure" and "hard drive failure").

[0232] For example: Evidence 1 states that the probability of “backplane failure” is ; Evidence 2 believes that the probability of "hard disk failure" is ; If A and B are mutually exclusive, then the conflict quantity .

[0233] 1.1.8 Traditional DS (Dempster-Shafer) directly distributes conflicting items proportionally to non-conflicting items. However, this can lead to distorted results in high-conflict scenarios. This embodiment introduces a dynamic redistribution factor to address this distorted result issue in DS. The design of the dynamic redistribution factor is based on the following principles:

[0234] Conflict weight adjustment: Dynamically adjust the redistribution ratio based on the conflict amount K and evidence uncertainty.

[0235] Fusion of historical experience: Optimizing allocation strategies by combining a library of historical conflict scenarios.

[0236] The definition formula is as follows:

[0237]

[0238] : represents the joint information entropy of evidence and is used to measure uncertainty.

[0239]

[0240] When the value of Entropy is large, it means that there is high uncertainty in the evidence. At this time, the evidence is confusing, and the amount of conflict redistribution will be reduced to reduce the risk of misjudgment; when the value of Entropy is small, it means that the evidence is clear, and the amount of conflict redistribution will be increased to enhance the conflict handling ability.

[0241] 1.1.9 To address the evidence synthesis problem in high-conflict scenarios and ensure the rationality and stability of the fusion results, a redistribution strategy can be used to dynamically adjust the distribution of conflicting elements. This can be done as follows:

[0242] 1) Redistribute the conflict quantity K: Distribute the conflict quantity K to the global uncertainty proposition and the high-probability single-fault proposition in a certain proportion.

[0243] 2) Uncertainty compensation: When the conflict amount K is large, the trust in the global uncertain proposition is increased to avoid blindly trusting a single piece of evidence.

[0244] Redistribution formula:

[0245]

[0246] is the redistribution factor

[0247]

[0248] 1.1.10 In the redistribution strategy, follow the example below:

[0249] 1) 60% is allocated to global uncertain propositions;

[0250] 2) 40% is weighted towards high probability single fault propositions.

[0251] To better understand the process of redistribution strategy, let's take an example: Backplane power fluctuations cause hard disk performance to degrade. The calculation process is as follows: Figure 5 As shown:

[0252] a) Input evidence

[0253] Evidence 1 (back panel):

[0254] Evidence 2 (hard drive):

[0255] b) Calculation of the conflict quantity K. In this example, backplane failure and hard drive failure are mutually exclusive. Therefore:

[0256]

[0257] c) The redistribution strategy is as follows:

[0258] Assuming entropy ,but:

[0259] 60% allocated to global uncertain propositions :

[0260] 40% is allocated to high probability single fault propositions according to weights:

[0261]

[0262] d) Calculation of synthesis results

[0263] First calculate

[0264] For backplane failure:

[0265] For hard drive failure:

[0266] for :

[0267] Add the conflict allocation amount K:

[0268]

[0269]

[0270] e) Normalization

[0271] Since the sum of the synthesis results may not be 1, normalization is required:

[0272]

[0273] Then the normalized result is:

[0274]

[0275] The end result is: , , .

[0276] 1.1.11 In storage links, due to uncertainty in the source of faults, conflicting evidence can arise, leading to conflicting judgments on the same event from different monitoring modules. Data noise can also lead to sensor sampling errors or environmental interference (such as electromagnetic interference). Model limitations can also lead to new fault modes not being covered (such as signal attenuation caused by quantum effects). Uncertainty compensation mechanisms are currently being used to address these issues:

[0277] 1) Global uncertainty enhancement: When the conflict value K>0.8, the system automatically increases the trust of the globally uncertain proposition as follows. This has the effect of avoiding blindly trusting a single piece of evidence in the face of high conflict.

[0278]

[0279] 2) Expert rules intervention

[0280] When the conflict value K exceeds 0.5, the predefined expert rule base (rules in the expert base are compiled based on troubleshooting experience) is activated. Matching expert rules are loaded from the database, and trust allocation is adjusted based on rule weights. Example rules are as follows: Rule 1: If multiple hard drives report errors simultaneously and the backplane temperature is > 70°C, it is determined to be a backplane heat dissipation failure (weight 0.3); Rule 2: If RAID reconstruction fails and the power supply ripple increases suddenly, it is determined to be a backplane capacitor aging (weight 0.5).

[0281] 1.1.12 Verification of the compensation effect of the uncertainty compensation mechanism: Monte Carlo simulation is used to evaluate the effectiveness of the compensation mechanism:

[0282] 1.1.13 In the decision-making mechanism, the most likely fault type or state is determined based on the synthesized belief and likelihood. The decision-making model is as follows: Figure 6 As shown:

[0283] Belief (hereinafter referred to as Bel) represents the minimum level of support for a proposition, and is calculated as follows:

[0284]

[0285] Where: m(B): basic probability assignment (BPA) of proposition B; : represents the subset of all those supporting proposition A; if Bel(A) is high: it indicates that there is more evidence supporting proposition A; if Bel(A) is low: it indicates that there is less evidence supporting proposition A.

[0286] 1.1.14 Plausibility (hereinafter abbreviated as Pl) represents the highest possible level of support for a proposition and is calculated as:

[0287] in, : represents all subsets that intersect with proposition A.

[0288] If Pl(A) is high, it indicates that proposition A is likely to be true; if Pl(A) is low, it indicates that proposition A is unlikely to be true.

[0289] 1.1.15 In the decision-making mechanism, the basic rules for determination are as follows:

[0290] When the following three conditions are met at the same time, it is determined that Proposition A is faulty:

[0291] 1) Bel(A)>threshold (initial value is 0.9);

[0292] 2) Pl(A) - Bel(A) < threshold (initial value is 0.1);

[0293] 3) Bel(A) / Bel() > threshold (initial value is 3.0).

[0294] However, when Pl(A) - Bel(A) is large, it indicates that the uncertainty of proposition A is high and requires further analysis. Therefore, the decision-making mechanism needs to be further optimized, including dynamic adjustment of thresholds and intervention of expert systems. The following are the optimization options:

[0295] 1) Dynamically adjust the threshold: Dynamically adjust the judgment threshold according to the system status. When in a high-load scenario, the threshold needs to be lowered to increase sensitivity; when in a low-load scenario, the threshold needs to be increased to reduce false alarms. For example:

[0296] Under high load, the threshold is adjusted from 0.9 to 0.8; under low load, the threshold is adjusted from 0.9 to 0.95.

[0297] 2) Expert system intervention: When uncertainty is high, the expert system is called to conduct in-depth analysis and match the most likely fault type based on historical data and expert experience. For example:

[0298] When Pl(A)−Bel(A)>0.2, due to the high uncertainty, it is necessary to call the expert system for further analysis and give the most likely diagnosis result.

[0299] In this way, by integrating multi-dimensional sensor data, a comprehensive solution can be provided for fault diagnosis of storage links (including RAID controllers, hard drives, backplanes, etc.), improving the accuracy of fault diagnosis and reducing false alarm rates; through a unified management platform, cross-level data collection, analysis and decision-making can be achieved, improving operation and maintenance efficiency; through automated fault diagnosis, dependence on professional operation and maintenance personnel can be reduced; the system supports dynamic expansion to adapt to the storage needs of ultra-large-scale data centers; it also has high versatility and can diagnose faults of hard drives, RAID cards, and hard drive backplanes of different manufacturers and various models.

[0300] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0301] According to an embodiment of the present disclosure, the present disclosure also provides a storage link fault diagnosis device.

[0302] For example, Figure 7 This is a schematic diagram of the structure of a storage link fault diagnosis device provided by an embodiment of the present disclosure. The storage link fault diagnosis device 700 includes:

[0303] Data acquisition module 710, used to collect multi-source data from different electronic devices on the storage link;

[0304] A vector construction module 720 is configured to analyze the multi-source data and construct a feature vector for each electronic device and a correlation feature vector between the electronic devices;

[0305] A weight adjustment module 730 is configured to adjust the contribution weight corresponding to each eigenvector according to the operating status of each electronic device;

[0306] A data fusion module 740 is configured to perform multi-source evidence fusion processing on the contribution weight corresponding to each feature vector and each feature vector to generate a trustworthiness and likelihood corresponding to each electronic device;

[0307] The fault diagnosis module 750 is configured to determine the fault type of the storage link according to the trust and likelihood corresponding to each electronic component.

[0308] Furthermore, the vector construction module 720 is used to:

[0309] Inputting the multi-source data into a preset fault correlation model;

[0310] quantifying the correlation between the electronic components based on the multi-source data using the preset fault correlation model to obtain a feature vector of each electronic component;

[0311] determining, based on the characteristic vectors of the respective electronic devices, correlation characteristic vectors between the respective electronic devices;

[0312] Based on the feature vectors of the electronic devices and the associated feature vectors between the electronic devices, a feature vector of the multi-source data is constructed.

[0313] Furthermore, the electronic device at least includes a hard disk, a controller and a backplane.

[0314] Furthermore, the hard disk data includes at least one of differential signal amplitude data, clock jitter data, bit error rate, spindle motor current ripple detection data, head cantilever vibration spectrum analysis data, disk medium temperature gradient monitoring data, voltage value, and log information of the hard disk;

[0315] The data of the controller includes at least one of a voltage value and log information of the controller;

[0316] The backplane data includes at least one of clock jitter data, spindle motor current ripple detection data and voltage value.

[0317] Furthermore, the vector construction module 720 is used to:

[0318] quantifying the correlation between the hard disk, the controller, and the backplane based on the multi-source data using the preset fault correlation model to obtain a feature vector of the hard disk, a feature vector of the controller, and a feature vector of the backplane;

[0319] Constructing an associated feature vector of the hard disk and the backplane based on the feature vector of the hard disk and the feature vector of the backplane;

[0320] A feature vector of the multi-source data is constructed based on the feature vector of the hard disk, the feature vector of the controller, the feature vector of the backplane, and the associated feature vector of the hard disk and the backplane.

[0321] Furthermore, the weight adjustment module 730 is configured to:

[0322] Acquire the system operation status; wherein the system operation status includes at least the load status, temperature, and task type of each of the electronic components;

[0323] Determine the feature significance score based on the historical normal baseline, trend sensitivity coefficient and cross-feature correlation weight;

[0324] Adjusting the characteristic type sensitivity coefficient according to the system operating state and a preset sensitivity coefficient matrix;

[0325] Inputting the feature type sensitivity coefficient and the feature significance score into a preset weight factor calculation model;

[0326] Generate a normalized weight corresponding to each feature vector based on the feature type sensitivity coefficient and the feature significance score through the preset weight factor calculation model;

[0327] By analyzing historical fault data, the contribution of each system operation status to fault diagnosis is calculated;

[0328] Based on the contribution of each of the system operating states to fault diagnosis, the normalized weight of each eigenvector is adjusted to obtain the contribution weight corresponding to each eigenvector.

[0329] Furthermore, the weight adjustment module 730 is configured to:

[0330] Performing median filtering on the derivative item of each eigenvector;

[0331] The normalized weight of each eigenvector after the median filtering is smoothed by using an exponentially weighted moving average method to obtain a contribution weight corresponding to each eigenvector.

[0332] Furthermore, the data fusion module 740 is configured to:

[0333] Based on each eigenvector and the contribution weight corresponding to each eigenvector, generating the confidence and likelihood of failure of each electronic component through an adaptive membership function;

[0334] The confidence and likelihood of each electronic component failing are fused using a redistribution factor to generate a corresponding confidence and likelihood for each electronic component.

[0335] Furthermore, the data fusion module 740 is further configured to:

[0336] determining a conflict amount based on a confidence level that each of the electronic components has failed;

[0337] Based on the conflict amount, a redistribution factor is calculated using a preset multiple evidence theory algorithm;

[0338] The confidence level and likelihood level corresponding to each electronic component are determined based on the redistribution factor, the first redistribution factor allocation ratio, and the confidence level of each electronic component that a failure has occurred.

[0339] Furthermore, a confirmation module is included, which is used to:

[0340] Based on the distribution ratio of the redistribution factor and the second redistribution factor, a confidence level of global uncertainty when the fault type cannot be determined is determined.

[0341] Furthermore, it also includes a trust adjustment module for:

[0342] If the conflict amount is greater than a first conflict amount threshold, loading a preset weight allocation rule, and adjusting the trust level corresponding to the electronic device based on the preset weight allocation rule;

[0343] If the conflict amount is greater than a second conflict amount threshold, the confidence level of the global uncertainty is increased according to a preset adjustment method.

[0344] Furthermore, the fault diagnosis module 750 is configured to:

[0345] For any fault type A, if the trust level of the fault type A is higher than the preset trust level threshold, and the difference between the likelihood and the trust level of the fault type A is less than the first preset difference threshold, and the ratio of the trust level of the fault type A to the trust level of the global uncertainty is greater than the preset ratio threshold, then the fault type of the storage link is determined to be the fault type A; wherein, the fault type A is a hard disk failure, a controller failure, a backplane failure, or the fault type cannot be determined.

[0346] Furthermore, a threshold adjustment module is included, which is used to:

[0347] If the difference between the likelihood and the confidence of the fault type A is greater than or equal to the first preset difference threshold, determining whether the storage link is in a high-load scenario based on the system operation status;

[0348] If the storage link is in the high-load scenario, lowering the preset trust threshold according to a preset adjustment rule;

[0349] If the storage link is in a low-load scenario, the preset trust threshold is increased according to the preset adjustment rule.

[0350] Furthermore, a fault determination module is included, which is used to:

[0351] If the difference between the likelihood and the confidence of the fault type A is greater than a second preset difference threshold, an expert system is called to determine the fault type of the storage link based on historical fault data.

[0352] Furthermore, an evaluation module is included for:

[0353] The misjudgment rate of the storage link fault diagnosis method is evaluated by using a Monte Carlo simulation method.

[0354] Furthermore, a pre-processing module is included for:

[0355] The multi-source data is subjected to data preprocessing, wherein the data preprocessing includes at least one of data cleaning, noise removal, missing value and outlier processing, normalization processing, and data dimensionality reduction processing.

[0356] It should be noted that, for the description of the features in the embodiment corresponding to the storage link fault diagnosis apparatus, reference can be made to the relevant description of the embodiment corresponding to the storage link fault diagnosis method, which will not be described in detail here.

[0357] An embodiment of the present disclosure further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above storage link fault diagnosis method embodiments.

[0358] An embodiment of the present disclosure further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned storage link fault diagnosis method embodiments when running.

[0359] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0360] An embodiment of the present disclosure further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned storage link fault diagnosis method embodiments are implemented.

[0361] An embodiment of the present disclosure further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned storage link fault diagnosis method embodiments are implemented.

[0362] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0363] The above is a detailed introduction to a target detection method provided by the present disclosure. Specific examples are used herein to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method of the present disclosure and its core ideas. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present disclosure, several improvements and modifications can be made to the present disclosure, and these improvements and modifications also fall within the scope of protection of the claims of the present disclosure.

Claims

1. A storage link fault diagnosis method, characterized in that: include: Collect multi-source data from different electronic devices on a storage link; the electronic devices include at least a hard disk, a controller, and a backplane; Analyzing the multi-source data to construct a feature vector of each electronic device and a correlation feature vector between the electronic devices; Adjusting the contribution weight corresponding to each eigenvector according to the operating status of each of the electronic devices; Performing multi-source evidence fusion processing on the contribution weight corresponding to each eigenvector and each eigenvector to generate a trustworthiness and likelihood corresponding to each electronic device; Determining the fault type of the storage link according to the trust and likelihood corresponding to each of the electronic components; The analyzing the multi-source data to construct the feature vectors of the respective electronic devices and the associated feature vectors between the respective electronic devices includes: quantifying the features of the hard disk, the controller, and the backplane based on the multi-source data using a preset fault correlation model to obtain a feature vector of the hard disk, a feature vector of the controller, and a feature vector of the backplane; Constructing an associated feature vector of the hard disk and the backplane based on the feature vector of the hard disk and the feature vector of the backplane; The determining the fault type of the storage link according to the trust and likelihood corresponding to each electronic component includes: For any fault type A, if the trust level of the fault type A is higher than the preset trust level threshold, and the difference between the likelihood and the trust level of the fault type A is less than the first preset difference threshold, and the ratio of the trust level of the fault type A to the trust level of the global uncertainty is greater than the preset ratio threshold, then the fault type of the storage link is determined to be the fault type A; wherein, the fault type A is a hard disk failure, a controller failure, a backplane failure, or the fault type cannot be determined.

2. The method according to claim 1, characterized in that The analyzing the multi-source data to construct the feature vectors of the respective electronic devices and the associated feature vectors between the respective electronic devices includes: Inputting the multi-source data into a preset fault correlation model; quantifying the characteristics of each of the electronic components based on the multi-source data using the preset fault correlation model to obtain a characteristic vector of each of the electronic components; determining, based on the characteristic vectors of the respective electronic devices, correlation characteristic vectors between the respective electronic devices; Based on the feature vectors of the electronic devices and the associated feature vectors between the electronic devices, a feature vector of the multi-source data is constructed.

3. The method according to claim 1, characterized in that The hard disk data includes at least one of differential signal amplitude data, clock jitter data, bit error rate, spindle motor current ripple detection data, head cantilever vibration spectrum analysis data, disk medium temperature gradient monitoring data, voltage value and log information of the hard disk; The data of the controller includes at least one of a voltage value and log information of the controller; The backplane data includes at least one of clock jitter data, spindle motor current ripple detection data and voltage value.

4. The method according to claim 2, characterized in that The constructing the feature vector of the multi-source data based on the feature vector of each electronic device and the associated feature vectors between the electronic devices includes: A feature vector of the multi-source data is constructed based on the feature vector of the hard disk, the feature vector of the controller, the feature vector of the backplane, and the associated feature vector of the hard disk and the backplane.

5. The method according to claim 1, wherein The step of adjusting the contribution weight corresponding to each eigenvector according to the operating status of each of the electronic devices includes: Acquire the system operation status; wherein the system operation status includes at least the load status, temperature, and task type of each of the electronic components; Determine the feature significance score based on the historical normal baseline, trend sensitivity coefficient and cross-feature correlation weight; Adjusting the characteristic type sensitivity coefficient according to the system operating state and a preset sensitivity coefficient matrix; Inputting the feature type sensitivity coefficient and the feature significance score into a preset weight factor calculation model; Generate a normalized weight corresponding to each feature vector based on the feature type sensitivity coefficient and the feature significance score through the preset weight factor calculation model; By analyzing historical fault data, the contribution of each system operation status to fault diagnosis is calculated; Based on the contribution of each of the system operating states to fault diagnosis, the normalized weight of each eigenvector is adjusted to obtain the contribution weight corresponding to each eigenvector.

6. The method according to claim 5, characterized in that The adjusting the normalized weight of each eigenvector based on the contribution of each system operating state to the fault diagnosis to obtain the contribution weight corresponding to each eigenvector includes: Performing median filtering on the derivative item of each eigenvector; The normalized weight of each eigenvector after the median filtering is smoothed by using an exponentially weighted moving average method to obtain a contribution weight corresponding to each eigenvector.

7. The method according to claim 1, characterized in that The performing multi-source evidence fusion processing on the contribution weight corresponding to each eigenvector and each eigenvector to generate the trust and likelihood corresponding to each electronic device includes: Based on each eigenvector and the contribution weight corresponding to each eigenvector, generating the confidence and likelihood of failure of each electronic component through an adaptive membership function; The confidence and likelihood of each electronic component failing are fused using a redistribution factor to generate a corresponding confidence and likelihood for each electronic component.

8. The method according to claim 7, characterized in that The redistribution factor is used to fuse the confidence and likelihood of each electronic component failing to generate the confidence and likelihood corresponding to each electronic component, including: determining a conflict amount based on a confidence level that each of the electronic components has failed; Based on the conflict amount, a redistribution factor is calculated using a preset multiple evidence theory algorithm; The confidence level and likelihood level corresponding to each electronic component are determined based on the redistribution factor, the first redistribution factor allocation ratio, and the confidence level of each electronic component that a failure has occurred.

9. The method according to claim 8, characterized in that Also includes: Based on the distribution ratio of the redistribution factor and the second redistribution factor, a confidence level of global uncertainty when the fault type cannot be determined is determined.

10. The method according to claim 9, characterized in that Also includes: If the conflict amount is greater than a first conflict amount threshold, loading a preset weight allocation rule, and adjusting the trust level corresponding to the electronic device based on the preset weight allocation rule; If the conflict amount is greater than a second conflict amount threshold, the confidence level of the global uncertainty is increased according to a preset adjustment method.

11. The method according to claim 1, wherein Also includes: If the difference between the likelihood and the confidence of the fault type A is greater than or equal to the first preset difference threshold, determining whether the storage link is in a high-load scenario based on the system operation status; If the storage link is in the high-load scenario, lowering the preset trust threshold according to a preset adjustment rule; If the storage link is in a low-load scenario, the preset trust threshold is increased according to the preset adjustment rule.

12. The method according to claim 11, characterized in that Also includes: If the difference between the likelihood and the confidence of the fault type A is greater than a second preset difference threshold, an expert system is called to determine the fault type of the storage link based on historical fault data.

13. The method according to claim 10, characterized in that Also includes: The misjudgment rate of the storage link fault diagnosis method is evaluated by using a Monte Carlo simulation method.

14. The method according to claim 1, wherein Before analyzing the multi-source data and constructing the feature vectors of the electronic devices and the associated feature vectors between the electronic devices, the method further includes: The multi-source data is subjected to data preprocessing, wherein the data preprocessing includes at least one of data cleaning, noise removal, missing value and outlier processing, normalization processing, and data dimensionality reduction processing.

15. A storage link fault diagnosis device, characterized in that: include: A data acquisition module, configured to acquire multi-source data from different electronic devices on a storage link; the electronic devices at least comprising a hard disk, a controller, and a backplane; A vector construction module, configured to analyze the multi-source data and construct a feature vector of each electronic device and a correlation feature vector between the electronic devices; A weight adjustment module, configured to adjust the contribution weight corresponding to each eigenvector according to the operating status of each of the electronic devices; a data fusion module, configured to perform multi-source evidence fusion processing on the contribution weight corresponding to each eigenvector and each eigenvector, and generate a trustworthiness and likelihood corresponding to each electronic device; a fault diagnosis module, configured to determine a fault type of the storage link based on the trust and likelihood corresponding to each of the electronic components; The vector building module is used to: quantifying the features of the hard disk, the controller, and the backplane based on the multi-source data using a preset fault correlation model to obtain a feature vector of the hard disk, a feature vector of the controller, and a feature vector of the backplane; Constructing an associated feature vector of the hard disk and the backplane based on the feature vector of the hard disk and the feature vector of the backplane; The fault diagnosis module is used to: For any fault type A, if the trust level of the fault type A is higher than the preset trust level threshold, and the difference between the likelihood and the trust level of the fault type A is less than the first preset difference threshold, and the ratio of the trust level of the fault type A to the trust level of the global uncertainty is greater than the preset ratio threshold, then the fault type of the storage link is determined to be the fault type A; wherein, the fault type A is a hard disk failure, a controller failure, a backplane failure, or the fault type cannot be determined.

16. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 14.

17. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 14.

18. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 14 when executed by a processor.

Citation Information

Patent Citations

  • Method and device for automatically detecting malicious codes and computer readable storage medium

    CN116010945A

  • Multi-element load prediction method and device for integrated energy system

    CN117035154A