Storage link fault diagnosis method and device, equipment, medium and program product

By collecting multi-source data to build feature vectors and associated feature vectors, dynamically adjusting the weights and performing multi-source evidence fusion, the high false alarm rate problem in storage link fault diagnosis is solved, and more efficient and accurate fault diagnosis is achieved.

CN120353631AActive Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510828093.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-22
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing storage link fault diagnosis methods rely on a single data source to cause high false alarm rates, making it difficult to distinguish faults caused by mutual cooperation between link levels, and traditional diagnostic methods are inefficient.

Method used

Collect multi-source data on the storage link, build feature vectors and associated feature vectors of electronic devices, dynamically adjust the contribution weight, generate trust and likelihood through multi-source evidence fusion processing, and determine the fault type of the storage link.

Benefits of technology

It improves the accuracy and efficiency of fault diagnosis, reduces the false alarm rate, enhances the adaptability and reliability of the method, and can have a more comprehensive understanding of the operating status of the storage link.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353631A_ABST
    Figure CN120353631A_ABST
Patent Text Reader

Abstract

The invention provides a storage link fault diagnosis method and device, equipment, a medium and a program product, and relates to the technical field of computers. Multi-source data of different electronic devices on a storage link are collected; analyzing the multi-source data, and constructing feature vectors of the electronic devices and associated feature vectors among the electronic devices; according to the operation state of each electronic device, adjusting a contribution weight corresponding to each feature vector; performing multi-source evidence fusion processing on the contribution weight corresponding to each feature vector and each feature vector to generate credibility and likelihood corresponding to each electronic device; and determining the fault type of the storage link according to the credibility and the likelihood corresponding to each electronic device. The fault diagnosis accuracy and diagnosis efficiency can be improved, and the false alarm rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, apparatus, device, medium, and program product for diagnosing storage link failures. Background Art

[0002] With the explosive growth of data scale and the increasing complexity of storage systems, the complexity of storage links (including RAID (Redundant Array of Independent Disks) controllers, hard disks, backplanes, etc.) is also getting higher and higher. For example, the storage link adopts a multi-level architecture, resulting in diverse fault types. In related technologies, storage link fault diagnosis usually relies on a single data source (such as hard disk SMART (Self-Monitoring, Analysis, and Reporting Technology) logs or RAID status), which consequently leads to a high false alarm rate. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, device, medium, and program product for diagnosing storage link failures. Its main purpose is to solve the problem of high false alarm rate in storage link fault diagnosis methods.

[0004] According to a first aspect of the present disclosure, there is provided a method for diagnosing storage link failures, including: Collecting multi-source data of different electronic devices on the storage link; Analyzing the multi-source data to construct feature vectors of each of the electronic devices and correlation feature vectors between each of the electronic devices; Adjusting contribution weights corresponding to each feature vector according to the operating states of each of the electronic devices; Performing multi-source evidence fusion processing on the contribution weights corresponding to each feature vector and each feature vector to generate a trust degree and a likelihood degree corresponding to each of the electronic devices; Determining the fault type of the storage link according to the trust degree and the likelihood degree corresponding to each of the electronic devices.

[0005] According to a second aspect of the present disclosure, there is provided a device for diagnosing storage link failures, including: A data acquisition module for collecting multi-source data of different electronic devices on the storage link; A vector construction module for analyzing the multi-source data to construct feature vectors of each of the electronic devices and correlation feature vectors between each of the electronic devices; A weight adjustment module for adjusting contribution weights corresponding to each feature vector according to the operating states of each of the electronic devices; A data fusion module for performing multi-source evidence fusion processing on the contribution weights corresponding to each of the feature vectors and each of the feature vectors to generate a confidence level and a likelihood level corresponding to each of the electronic devices; A fault diagnosis module for determining a fault type of the storage link according to the confidence level and the likelihood level corresponding to each of the electronic devices.

[0006] According to a third aspect of the present disclosure, there is provided an electronic device, including: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the foregoing first aspect.

[0007] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the foregoing first aspect.

[0008] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the method described in the foregoing first aspect when executed by a processor.

[0009] In the embodiments provided by the present disclosure, by collecting multi-source data of different electronic devices on a storage link; analyzing the multi-source data to construct feature vectors of each of the electronic devices and correlation feature vectors between each of the electronic devices; adjusting contribution weights corresponding to each feature vector according to the operating states of each of the electronic devices; performing multi-source evidence fusion processing on the contribution weights corresponding to each feature vector and each feature vector to generate a confidence level and a likelihood level corresponding to each of the electronic devices; determining a fault type of the storage link according to the confidence level and the likelihood level corresponding to each of the electronic devices. In this way, by fusing multi-source data of multiple electronic devices, performing dynamic weight adjustment and multi-source evidence fusion processing on the multi-source data, obtaining the confidence level and the likelihood level of each electronic device, and implementing fault diagnosis in the storage link according to the confidence level and the likelihood level of each electronic device, a solution for fault diagnosis of the storage link can be provided, the accuracy and efficiency of fault diagnosis can be improved, and the false alarm rate can be reduced.

[0010] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0011] The accompanying drawings are used to better understand the present solution and do not constitute a limitation to the present disclosure. Among them: Figure 1 is a schematic flowchart of a storage link fault diagnosis method provided by the related art; Figure 2 is a schematic flowchart of a storage link fault diagnosis method provided by an embodiment of the present disclosure; Figure 3 is a schematic flowchart of another storage link fault diagnosis method provided by an embodiment of the present disclosure; Figure 4 is a schematic flowchart of a dynamic weight allocation mechanism provided by an embodiment of the present disclosure; Figure 5 is a schematic flowchart of a method for dynamically adjusting the conflict amount allocation method through a reallocation strategy provided by an embodiment of the present disclosure; Figure 6 is a schematic flowchart of a decision-making and determination mechanism provided by an embodiment of the present disclosure; Figure 7 is a schematic structural diagram of a storage link fault diagnosis device provided by an embodiment of the present disclosure. Detailed implementation manners

[0012] The following makes an explanation of exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0013] As can be seen from the background art, traditional storage link fault diagnosis methods often adopt a single-point diagnosis method that usually relies on a single data source (such as hard disk SMART logs or RAID status), which consequently leads to a high false alarm rate and has limitations. For example: a hard disk response timeout may be caused by unstable backplane power supply but is misjudged as a hard disk fault; a RAID reconstruction failure may be caused by backplane signal interference but is misjudged as a hard disk media damage, etc. Moreover, traditional fault repair relies on manual troubleshooting, which is inefficient. For example, for a backplane fault, the power supply module and signal lines need to be detected node by node; for a hard disk fault, the SMART logs and media status need to be checked disk by disk.

[0014] Exemplarily, such as Figure 1As shown in the figure, the storage link fault detection method in the related art generally includes: creating a monitoring virtual machine on each host node; mounting the volumes corresponding to different storage backends to the corresponding monitoring virtual machines; obtaining health status detection data; and determining the data link status between the monitoring virtual machine and the storage backend based on the health status detection data. And creating a monitoring virtual machine on each host node can be: based on the monitoring server and through the Nova service, creating a monitoring virtual machine on each host node. Among them, mounting the volumes corresponding to different storage backends to the corresponding monitoring virtual machines includes: based on the monitoring server and through the Cinder service, mounting the volumes corresponding to different storage backends to the corresponding monitoring virtual machines.

[0015] It can be seen that the technical solution adopted depends on a single data source and it is difficult to distinguish faults caused by the mutual cooperation between link levels in the storage link, such as: the performance degradation of hard disks caused by the power fluctuation of the backplane, the RAID degradation caused by the physical damage of hard disks, and the CRC (Cyclic Redundancy Check) errors caused by the signal interference of the SAS (Serial Attached SCSI) link. Moreover, although the technical solution introduces the diagnosis of storage link faults, there is no diagnosis of the faults of the RAID card and the hard disk backplane. That is to say, this is only the fault diagnosis of hard disks and does not comprehensively cover the entire storage link.

[0016] Next, the storage link fault diagnosis method, device, equipment medium and program product of the embodiments of the present disclosure will be described with reference to the accompanying drawings.

[0017] Figure 2 The flowchart of a storage link fault diagnosis method provided by an embodiment of the present disclosure is shown in Figure 1 As shown in the figure, the method includes the following steps: Step 201, collect multi-source data of different electronic devices on the storage link.

[0018] Among them, multi-source data of different electronic devices on the storage link can be collected. The electronic devices can be at least two, that is, the data of each electronic device is collected respectively. These data are the basis for subsequent fault diagnosis.

[0019] Among them, the electronic devices can at least include hard disks, controllers and backplanes. Among them, the controller can be, for example, the RAID control right. The multi-source data can include the data of hard disks, the data of controllers and the data of backplanes.

[0020] Among them, the data of the hard disk may include at least one of differential signal amplitude data, clock jitter data, bit error rate, spindle motor current ripple detection data, head suspension vibration spectrum analysis data, disk media temperature gradient monitoring data, voltage value, and the log information of the hard disk; the data of the controller may include at least one of voltage value and the log information of the controller; the data of the backplane may include at least one of clock jitter data, spindle motor current ripple detection data, and voltage value. For example, the data of the hard disk and the backplane can be obtained through sensors.

[0021] As an example, the differential signal amplitude, clock jitter, bit error rate, spindle motor current ripple detection (bandwidth DC - 10MHz), head suspension vibration spectrum analysis (0 - 5kHz FFT), and disk media temperature gradient monitoring (deploying 4 PT100 sensors per surface) on the storage link (including RAID cards (RAID controllers), hard disks, hard disk backplanes, etc.) can be captured in real time; the voltage values on each unit module of the storage link, and the data on the RAID controller, hard disk sensors, and backplane sensors can be monitored in real time; the log information on the RAID card, the SMART log information of the hard disk, etc. can be collected in real time.

[0022] Step 202, analyze the multi-source data, and construct the feature vectors of each electronic device and the correlation feature vectors between each electronic device.

[0023] Among them, the multi-source data of different electronic devices collected can be analyzed, features related to fault diagnosis can be extracted, and the feature vectors of each electronic device can be constructed. At the same time, the correlation feature vectors between different electronic devices can also be calculated to consider their mutual influence. In this way, by constructing the feature vectors and correlation feature vectors, the complex multi-source data can be transformed into a quantifiable feature representation, providing structured data support for subsequent weight adjustment and evidence fusion.

[0024] Step 203, according to the operating states of each electronic device, adjust the contribution weights corresponding to each feature vector.

[0025] Among them, according to the operating states of each electronic device (such as load, temperature, task type, etc.), the contribution weights corresponding to each feature vector can be dynamically adjusted. Since the change in the operating state of the electronic device may affect the importance of the features, it is necessary to adjust the contribution weights corresponding to each feature vector in real time to reflect the current operating environment. In this way, by dynamically adjusting the weights, the fault diagnosis method can better adapt to different operating states and improve the accuracy and adaptability of the diagnosis.

[0026] Step 204, perform multi-source evidence fusion processing on the contribution weights corresponding to each feature vector and each feature vector, and generate the trust degree and likelihood degree corresponding to each electronic device.

[0027] Among them, a multi-source evidence fusion method (such as Dempster-Shafer theory) can be adopted to combine each feature vector and its corresponding contribution weight to generate the belief degree and plausibility degree corresponding to each electronic device. The belief degree and plausibility degree can reflect the possibility of each electronic device having a fault. For example, the belief degree (abbreviated as Bel) can represent the lowest degree of support for a certain fault type (such as a certain electronic device fault); the plausibility degree (abbreviated as Pl) can represent the highest possible degree of support for a certain fault type (such as a certain electronic device fault). In this way, through multi-source evidence fusion, different feature vectors and their weights can be comprehensively considered to generate more accurate belief degrees and plausibility degrees, thereby improving the reliability of fault diagnosis.

[0028] Step 205: Determine the fault type of the storage link according to the belief degree and plausibility degree corresponding to each electronic device.

[0029] Among them, the fault type of the storage link can be determined according to the belief degree and plausibility degree of each electronic device. Exemplarily, if the belief degree and plausibility degree of a certain electronic device meet the preset fault judgment conditions, it is determined that the electronic device has a fault. The fault type can include, for example, hard disk fault, RAID fault, backplane fault, etc.

[0030] In summary, in the embodiment provided by the present disclosure, by collecting multi-source data of different electronic devices on the storage link; analyzing the multi-source data to construct feature vectors of each of the electronic devices and correlation feature vectors between each of the electronic devices; adjusting the contribution weight corresponding to each feature vector according to the operating state of each of the electronic devices; performing multi-source evidence fusion processing on the contribution weight corresponding to each feature vector and each feature vector to generate the belief degree and plausibility degree corresponding to each of the electronic devices; determining the fault type of the storage link according to the belief degree and plausibility degree corresponding to each of the electronic devices. In this way, by fusing multi-source data of multiple electronic devices, dynamically adjusting the weights of the multi-source data and performing multi-source evidence fusion processing, the belief degree and plausibility degree of each electronic device are obtained, and fault diagnosis in the storage link is realized according to the belief degree and plausibility degree of each electronic device, which can provide a solution for fault diagnosis of the storage link, improve the accuracy and efficiency of fault diagnosis, and reduce the false alarm rate. At the same time, the adaptability and reliability of the method can also be enhanced.

[0031] It should be noted that the embodiments of the present disclosure may include multiple steps. For the convenience of description, these steps are numbered, but these numbers are not intended to limit the execution time slots and execution orders between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not make any limitations in this regard.

[0032] Further, in a possible implementation manner of this embodiment, analyzing multi-source data and constructing feature vectors of each electronic device and correlation feature vectors between each electronic device includes: Inputting the multi-source data into a preset fault correlation degree model; Based on the multi-source data through the preset fault correlation degree model, quantifying the correlation between each electronic device to obtain the feature vectors of each electronic device; Based on the feature vectors of each electronic device, determining the correlation feature vectors between each electronic device; Based on the feature vectors of each electronic device and the correlation feature vectors between each electronic device, constructing the feature vector of the multi-source data.

[0033] Among them, the multi-source data of different electronic devices on the storage link collected can be input into the preset fault correlation degree model. The preset fault correlation degree model can be a pre-trained model for generating the feature vectors of each electronic device. Then, the preset fault correlation degree model can be used to analyze and process the input multi-source data to quantify the correlation between each electronic device. For example, by analyzing the correlation and causal relationship between data, the feature vector of each electronic device can be generated. These feature vectors reflect the operating state and fault characteristics of each electronic device. After that, based on the feature vectors of each electronic device, the correlation feature vectors between each electronic device can be further calculated. These correlation feature vectors can reflect the interdependent relationship between different electronic devices and help understand the fault propagation path. Finally, the feature vectors of each electronic device and the correlation feature vectors can be comprehensively combined to construct the final feature vector of the multi-source data. This feature vector combines all the feature vectors of the electronic devices and the correlation feature vectors between them, providing a comprehensive feature representation for subsequent fault diagnosis. In this way, by constructing the feature vectors and correlation feature vectors, the correlation between the electronic devices is quantified, and a comprehensive feature vector is generated. This can not only improve the accuracy of the fault diagnosis result but also enhance the adaptability and reliability of the system. Thus, the operating state of the storage link can be more comprehensively understood, and accurate diagnosis of faults can be achieved.

[0034] As a specific example, the data information (multi-source data) on the storage link collected can be analyzed using a fault correlation degree model (preset fault correlation degree model). The fault correlation degree model can establish a cross-level fault causal chain by quantifying the dynamic correlation between the RAID controller, hard disk hardware, and the physical state of the backplane. First, the feature vectors can be defined. The data dimensions required for defining the feature vectors are shown in the following table:

[0035] Among them, 1) The construction formula of the feature vector of the RAID controller is as follows:

[0036] Among them, Uncorrectable_Errors represents the number of errors detected by the RAID controller but unable to be corrected, Total_IO represents the total number of IO requests (including read and write operations); Reconstruct_Time represents the reconstruction time, that is, when a hard disk in the RAID array fails, the time required for the RAID system to recover (reconstruct) the data of this hard disk through redundant data. represents the uncorrectable error rate. If this value is relatively high, it indicates that the data integrity of the RAID controller is seriously threatened and may lead to data loss; if this value is relatively low, it indicates that the RAID controller is operating normally and the data reliability is high.

[0037] 2) The construction formula of the feature vector of the hard disk sensor is as follows:

[0038] Among them, Seek_Error_Rate (seek error rate) refers to the error frequency that occurs when the magnetic head is looking for a specific data position. It is one of the hard disk SMART (Self-Monitoring, Analysis and Reporting Technology) attributes and is used to evaluate the mechanical performance of the hard disk; Raw_Vibration represents the vibration spectrum, that is, the vibration characteristics of the spindle motor or the magnetic head suspension, and Temperature represents the temperature of the hard disk platter.

[0039] 3) The construction formula of the feature vector related to the hard disk backplane is as follows:

[0040] Among them, Voltage_Ripple represents the amplitude of the power supply voltage fluctuation; Threshold represents the temperature difference of each area of the backplane; UI represents Unit Interval (unit interval); Jitter_RMS represents the key index of the signal timing jitter, which represents the root mean square deviation of the signal edge relative to the ideal clock edge. The calculation formula is:

[0041] Among them, TIE i represents the timing deviation of the i-th signal edge; μ TIE represents the mean value of TIE; N represents the total number of signal edges.

[0042] 4) The construction formula of the cross-level correlation feature vector is as follows:

[0043] This formula is used to quantify the correlation between the hard disk drive (HDD) layer and the backplane (BP) layer, and is weighted by combining the uncertainty of the RAID logs to capture potential patterns of cross-level failures; when the characteristics of the hard disk and the backplane are highly correlated, the Corr value is large, indicating that there may be cross-level failures; when the uncertainty of the RAID logs is high (high entropy), the f Cross value increases, enhancing the sensitivity to cross-level failures. represents the hard disk drive layer feature f HDD and the backplane layer feature f BP 's correlation coefficient, which is used to measure the degree of linear correlation between the two. The calculation formula is:

[0044] where Corr(f HDD , f BP ) represents the covariance, measuring the co-variation between the two; and are the standard deviations of f HDD and f BP respectively. When Corr(f HDD , f BP ) is close to 1: the hard disk and backplane characteristics are highly positively correlated (e.g., unstable backplane power supply leads to a decrease in hard disk performance); when Corr(f HDD , f BP ) is close to -1: highly negatively correlated (e.g., an increase in backplane temperature leads to a decrease in hard disk speed); when close to Corr(f HDD , f BP ) 0: there is no obvious linear relationship.

[0045] Entropy(RAID_Logs) represents the RAID log information entropy, which is used to quantify the uncertainty of the RAID logs. The calculation formula is:

[0046] where p i represents the occurrence probability of the i-th type of log event (such as CRC error, reconstruction failure, etc.); N represents the total number of log event types.

[0047] In summary, the feature vectors of multi-source data are constructed as follows:

[0048] Furthermore, in a possible implementation manner of this embodiment, based on multi-source data through a preset fault correlation model, the correlation between each electronic device is quantified, and the feature vectors of each electronic device are obtained, including: Based on multi-source data through a preset fault correlation model, quantify the correlation between the hard disk, the controller, and the backplane, and obtain the feature vectors of the hard disk, the feature vectors of the controller, and the feature vectors of the backplane; Based on the feature vectors of each electronic device, determine the correlation feature vectors between each electronic device, including: Based on the feature vectors of the hard disk and the feature vectors of the backplane, construct the correlation feature vectors between the hard disk and the backplane; Based on the feature vectors of each electronic device and the correlation feature vectors between each electronic device, construct the feature vectors of multi-source data, including: Based on the feature vectors of the hard disk, the feature vectors of the controller, the feature vectors of the backplane, and the correlation feature vectors between the hard disk and the backplane, construct the feature vectors of multi-source data.

[0049] Among them, if the electronic devices at least include a hard disk, a controller, and a backplane. Then, through a preset fault correlation model based on multi-source data, quantify the correlation between the hard disk, the controller, and the backplane, and obtain the feature vectors of the hard disk, the feature vectors of the controller, and the feature vectors of the backplane. Then, based on the feature vectors of the hard disk and the feature vectors of the backplane, construct the correlation feature vectors between the hard disk and the backplane. After that, based on the feature vectors of the hard disk, the feature vectors of the controller, the feature vectors of the backplane, and the correlation feature vectors between the hard disk and the backplane, construct the feature vectors of multi-source data. The specific implementation of each step in this embodiment is similar to that of the previous embodiment and will not be elaborated here.

[0050] Further, in a possible implementation manner of this embodiment, according to the operating states of each electronic device, adjust the contribution weight corresponding to each feature vector, including: Obtain the system operating state; where the system operating state at least includes the load state, temperature, and task type of each electronic device; According to the historical normal baseline, trend sensitivity coefficient, and cross-feature correlation weight, determine the feature significance score; According to the system operating state and the preset sensitivity coefficient matrix, adjust the feature type sensitivity coefficient; Input the feature type sensitivity coefficient and the feature significance score into the preset weight factor calculation model; Through the preset weight factor calculation model, based on the feature type sensitivity coefficient and the feature significance score, generate the normalized weight corresponding to each feature vector; By analyzing historical fault data, statistically analyze the contribution degree of each system operating state to fault diagnosis; Based on the contribution degree of each system operating state to fault diagnosis, adjust the normalized weight of each feature vector to obtain the contribution weight corresponding to each feature vector.

[0051] Among them, the operating status of the storage link system can be obtained, including the load status, temperature, task type, etc. of each electronic device. These operating statuses can reflect the current operating environment and working conditions of the system.

[0052] Among them, the operating status of each electronic device can be obtained from the storage link system. For example, the load status, temperature, and task type of each electronic device. The load status can include, for example, normal load (e.g., IOPS < 5k), high load (e.g., IOPS > 10k). The temperature can include, for example, high temperature warning (e.g., T > 60°C). The task type can include, for example, RAID reconstruction. Among them, IOPS (Input / Output Operations Per Second, that is, the number of input / output operations per second) is an important indicator to measure the performance of the storage system, used to describe the number of input / output operations that a storage device (such as a hard disk, solid-state drive, storage array, etc.) can complete per unit time. Then, the historical normal baseline, trend sensitivity coefficient, and cross-feature correlation weight can be obtained, and then the feature significance score can be calculated based on the historical normal baseline, trend sensitivity coefficient, and cross-feature correlation weight. Among them, the historical normal baseline can be, for example, the mean value in a certain rolling past period (e.g., 30 days); the trend sensitivity coefficient can represent the amplification effect of sudden increase or decrease and can be a constant, for example, 0.7; the cross-feature correlation weight can be a constant, for example, 0.4. The calculation method of the feature significance score can be as follows:

[0053] Among them, : The rolling 30-day mean value represents the historical normal baseline; : The constant 0.7, the trend sensitivity coefficient; : The constant 0.4, the cross-feature correlation weight; : The cross-level associated feature vector.

[0054] After that, the feature type sensitivity coefficient can be adjusted according to the system operating status and the preset sensitivity coefficient matrix. The preset sensitivity coefficient matrix can be a pre-defined quantization matrix of the feature sensitivity degree according to different operating statuses. According to the currently obtained system operating status (load, temperature, task type, etc.), combined with the preset sensitivity coefficient matrix, the sensitivity coefficient of each feature type can be dynamically adjusted to make it more in line with the current actual operating situation. An example of the preset sensitivity coefficient matrix is as follows in the table. When multiple system operating statuses are triggered simultaneously, the feature type sensitivity coefficient α of each status can be adjusted by the weighted sum method.

[0055]

[0056] Then, use the determined feature significance score and the adjusted feature type sensitivity coefficient as input parameters and input them into a preset weight factor calculation model. The weight factor calculation model is a preset mathematical model that comprehensively considers the feature type sensitivity coefficient and the feature significance score. Through the preset weight factor calculation model, normalization processing is performed based on the feature type sensitivity coefficient and the feature significance score to make it within a certain range (such as 0 to 1), which is convenient for subsequent comparison and application. Among them, the weight factor calculation model can be, for example, as follows:

[0057] Where: : The normalized weight of the i-th type of feature at time t ( ); : Feature type sensitivity coefficient (preset empirical value); : Feature significance score (real-time calculated value).

[0058] After that, the operation state records when the system had faults in the past can be obtained as historical fault data, and the historical fault data is analyzed to count the importance of different operation states (such as high load, high temperature, specific task type, etc.) in fault diagnosis and quantify it as the contribution degree. Finally, based on the contribution degree of each system operation state to fault diagnosis, the normalized weight of each feature vector is adjusted to obtain the contribution weight corresponding to each feature vector. For example, the generated normalized weight can be further adjusted according to the statistically obtained contribution degree, and finally the contribution weight corresponding to each feature vector in fault diagnosis is obtained to make it more in line with the actual needs of fault diagnosis.

[0059] As an example, by analyzing historical data, counting the contribution degree of each state to fault diagnosis, and dynamically adjusting the weights, as follows:

[0060] Where: : The contribution degree of the j-th state in historical faults.

[0061] For example: When both RAID reconstruction and high temperature alarm are triggered simultaneously, and , , then the value of is:

[0062] Furthermore, in some possible implementation manners, adjusting the normalized weight of each feature vector based on the contribution degree of each system operation state to fault diagnosis to obtain the contribution weight corresponding to each feature vector includes: Performing median filtering processing on the derivative term of each feature vector; By using the method of exponential weighted moving average, the normalized weights of each feature vector after median filtering are smoothed to obtain the contribution weights corresponding to each feature vector.

[0063] In the embodiment of the present application, an outlier filtering mechanism is added to the anti-interference design, that is: median filtering is performed on the derivative term in Si, as follows:

[0064] This can eliminate the influence of outliers.

[0065] In addition, to prevent the weight from mutating, the weight is smoothed by using the method of exponential weighted moving average, and the processing method is as follows:

[0066] Furthermore, in some possible implementation manners, multi-source evidence fusion processing is performed on the contribution weights corresponding to each feature vector and each feature vector to generate the trust degree and likelihood degree corresponding to each electronic device, including: Based on each feature vector and the contribution weight corresponding to each feature vector, the trust degree and likelihood degree of each electronic device having a fault are generated through an adaptive membership function; A redistribution factor is used to perform fusion processing on the trust degree and likelihood degree of each electronic device having a fault to generate the trust degree and likelihood degree corresponding to each electronic device.

[0067] Among them, each feature vector and its corresponding contribution weight can be used as inputs, and through the multi-source evidence fusion method, the information of different feature vectors is comprehensively considered to generate the fault trust degree and likelihood degree of each electronic device. Then, based on each feature vector and the contribution weight corresponding to each feature vector, the trust degree and likelihood degree of each electronic device having a fault are generated through an adaptive membership function. Among them, the adaptive membership function is a dynamically adjusted function, which calculates the trust degree and likelihood degree of each electronic device having a fault according to the feature vector and the contribution weight; the trust degree (Belief, hereinafter abbreviated as Bel) represents the lowest support degree for a certain proposition, and the likelihood degree (Plausibility, hereinafter abbreviated as Pl) represents the highest possible support degree for a certain proposition. Through the adaptive membership function, the feature vector and the contribution weight are mapped into the trust degree and the likelihood degree, providing a basis for subsequent fusion processing.

[0068] Among them, the redistribution factor is used to adjust the weights of the confidence and likelihood, ensuring that the fused result is more reasonable. Fusion processing: The confidence and likelihood of each electronic device are weighted and summed through the redistribution factor or other fusion methods to generate the final confidence and likelihood. For example, taking each feature vector and its corresponding contribution weight as inputs, through an adaptive membership function, based on the feature vector and the contribution weight, calculate the confidence and likelihood of each electronic device, and use the redistribution factor to perform weighted fusion on the confidence and likelihood to generate the final confidence and likelihood of each electronic device. Through the fusion processing, comprehensively consider the contributions of different feature vectors to the fault diagnosis, generate the final confidence and likelihood of each electronic device, and use them to more accurately evaluate the fault risk of the electronic device.

[0069] Further, the redistribution factor is used to perform fusion processing on the confidence and likelihood of each electronic device having a fault, generating the corresponding confidence and likelihood of each electronic device, including: Based on the confidence of each electronic device having a fault, determine the conflict amount; Based on the conflict amount, use a preset multi-evidence theory algorithm to calculate the redistribution factor; Based on the redistribution factor, the first redistribution factor allocation ratio, and the confidence of each electronic device having a fault, determine the corresponding confidence and likelihood of each electronic device.

[0070] And, based on the redistribution factor and the allocation ratio of the second redistribution factor, determine the confidence of the global uncertainty when the fault type cannot be judged.

[0071] Among them, the conflict amount can be calculated by comparing the differences between the confidences generated by different feature vectors. For example, the following formula can be used:

[0072] Among them, m1 and m2 are the basic probability assignments (confidences) of two independent evidences (different feature vectors); A and B are two mutually exclusive fault types (such as "backplane fault" and "hard disk fault"). For example: Evidence 1 believes that the probability of "backplane fault" is , and evidence 2 believes that the probability of "hard disk fault" is ; if A and B are mutually exclusive, then the conflict amount .

[0073] Among them, the multi-evidence theory algorithm: such as the Dempster-Shafer theory (D-S theory), etc., is used to process uncertain and conflicting information; the redistribution factor: a factor calculated by a preset multi-evidence theory algorithm according to the amount of conflict, which is used to adjust the distribution of belief and plausibility. According to the conflict amount K, the redistribution factor is calculated using the combination rule (such as the Dempster rule) in the D-S theory as follows:

[0074] Wherein: represents the joint information entropy of evidence, which is used to measure uncertainty,

[0075] When the value of Entropy is large, it indicates that there is high uncertainty in the evidence. At this time, the evidence is chaotic and the conflict redistribution amount will decrease to reduce the risk of misjudgment; when the value of Entropy is small, it indicates that the evidence is clear and the conflict redistribution amount will increase to enhance the conflict handling ability.

[0076] Among them, the distribution ratio of the first redistribution factor: a preset ratio, which is used to distribute the influence of the redistribution factor on belief and plausibility, such as 40%. Fusion processing: Adjust the belief and plausibility of each electronic device through the redistribution factor and the distribution ratio of the first redistribution factor to generate the final belief and plausibility. And, the distribution ratio of the second redistribution factor: a preset ratio, which is used to distribute the influence of the redistribution factor on global uncertainty, such as 60%; the global uncertainty belief: when the fault type cannot be clearly judged, it represents the belief of the overall uncertainty of the system.

[0077] Furthermore, it also includes: If the conflict amount is greater than the first conflict amount threshold, then load the preset weight distribution rule and adjust the belief corresponding to the electronic device based on the preset weight distribution rule; If the conflict amount is greater than the second conflict amount threshold, then increase the belief of global uncertainty according to the preset adjustment method.

[0078] Among them, the first conflict amount threshold: a preset threshold used to determine whether the conflict amount reaches the level where the confidence needs to be adjusted; condition judgment: when the calculated conflict amount K is greater than the first conflict amount threshold, it indicates that the contradiction between different evidences is relatively serious and the confidence needs to be adjusted. Preset weight assignment rules: a predefined set of rules used to adjust the confidence assignment when the conflict amount is large; rule content: may include reassigning the contribution weights of different feature vectors or making specific adjustments to the confidence, etc. After loading the preset weight assignment rules, the confidence of each electronic device can be adjusted according to the preset weight assignment rules to reduce the impact of the conflict amount. As an example, the preset weight assignment rules can be expert rules stored in a database. For example, if the conflict amount is greater than the first conflict amount threshold, such as 0.5, the predefined expert rule library (the rules in the expert library are sorted based on fault handling experience) can be activated, the matching expert rules can be loaded from the database, and the trust assignment can be adjusted according to the rule weights.

[0079] Among them, the second conflict amount threshold: another preset threshold used to determine whether the conflict amount reaches the level where the global uncertainty confidence needs to be increased, such as 0.8. Condition judgment: when the calculated conflict amount K is greater than the second conflict amount threshold, it indicates that the contradiction between different evidences is extremely serious and the confidence of global uncertainty needs to be increased. Preset adjustment method: a predefined adjustment method used to increase the confidence of global uncertainty when the conflict amount is extremely large. If the conflict amount is greater than the second conflict amount threshold, according to the preset adjustment method, the confidence of global uncertainty is increased to reflect the overall uncertainty of the system. For example, when the conflict amount K > 0.8, the system automatically increases the confidence of the global uncertain proposition, specifically as follows:

[0080] In this way, by setting two conflict amount thresholds corresponding to different adjustment strategies, different levels of conflict situations can be effectively handled. When the conflict amount is large, the confidence is adjusted through the preset weight assignment rules to reduce the impact of the conflict; when the conflict amount is extremely large, the confidence of global uncertainty is increased to reflect the overall uncertainty of the system. This method can improve the accuracy and reliability of fault diagnosis, especially when facing highly uncertain and conflicting information.

[0081] As an example, to solve the problem of evidence synthesis in high-conflict scenarios and ensure the rationality and stability of the fusion result, a redistribution strategy is used to dynamically adjust the distribution method of the conflict amount to solve the problem in high-conflict scenarios. For example: 1) Reassign the conflict amount K: The conflict amount K is distributed to the global uncertain proposition and the high-probability single-fault proposition in a certain proportion. 2) Uncertainty compensation: When the conflict amount K is large, increase the confidence of the global uncertain proposition to avoid blindly trusting a single piece of evidence.

[0082] The redistribution formula is as follows:

[0083] : Redistribution factor:

[0084] In a specific redistribution strategy, the distribution can be carried out according to the following ratio: 1) 60% is allocated to global uncertain propositions; 2) 40% is allocated to high-probability single-fault propositions according to the weight.

[0085] Furthermore, according to the trust degree and likelihood degree corresponding to each electronic device, the fault type of the storage link is determined, including: For any fault type A, if the trust degree of fault type A is higher than the preset trust degree threshold, and the difference between the likelihood degree and the trust degree of fault type A is less than the first preset difference threshold, and the ratio of the trust degree of fault type A to the trust degree of global uncertainty is greater than the preset ratio threshold, then the fault type of the storage link is determined to be fault type A; where, fault type A is a hard disk fault, a controller fault, a backplane fault, or the fault type cannot be determined.

[0086] Among them, the trust degree of fault type A (Bel(A)) is higher than the preset trust degree threshold, and the difference between the likelihood degree Pl(A) of fault type A and the trust degree is less than the first preset difference threshold, and the ratio of the trust degree of fault type A to the trust degree of global uncertainty Bel() is greater than the preset ratio threshold, which is a preset decision-making mechanism. If a certain proposition A (such as a certain fault type A) satisfies this decision-making mechanism, then it is considered that fault type A holds. As an example: If a certain proposition A (such as a certain fault type A) satisfies the following 3 conditions at the same time, it is determined that proposition A has a fault: 1) Bel(A) > threshold (preset trust degree threshold, initial value is 0.9); 2) Pl(A) - Bel(A) < threshold (first preset difference threshold, initial value is 0.1); 3) Bel(A) / Bel() > threshold (preset ratio threshold, initial value is 3.0).

[0087] Furthermore, it also includes: If the difference between the likelihood degree and the trust degree of fault type A is greater than or equal to the first preset difference threshold, then according to the system operation state, it is determined whether the storage link is in a high-load scenario; If the storage link is in a high-load scenario, the preset trust degree threshold is reduced according to the preset adjustment rule; If the storage link is in a low-load scenario, the preset trust threshold is increased according to the preset adjustment rule.

[0088] Among them, if the difference between the likelihood and the trust of fault type A is greater than or equal to the first preset difference threshold, for example, when Pl(A) - Bel(A) is large, it means that the uncertainty of proposition A is high, and further analysis is required. The decision threshold is dynamically adjusted according to the system state. When in a high-load scenario, the threshold needs to be reduced to improve sensitivity; when in a low-load scenario, the threshold needs to be increased to reduce false alarms. For example, when in a high-load situation, the threshold is adjusted from 0.9 to 0.8. When in a low-load situation, the threshold is adjusted from 0.9 to 0.95.

[0089] Furthermore, it also includes: If the difference between the likelihood and the trust of fault type A is greater than the second preset difference threshold, the expert system is called to determine the fault type of the storage link based on historical fault data.

[0090] Among them, if the difference between the likelihood and the trust of fault type A is greater than the second preset difference threshold, that is, when the uncertainty is high, the expert system can be called for in-depth analysis. According to historical data (historical fault data) and expert experience, the most likely fault type is matched. For example, when Pl(A) - Bel(A) > 0.2, due to the high uncertainty, the expert system needs to be called for further analysis and the most likely diagnostic result is given.

[0091] Furthermore, it also includes: The misjudgment rate of the storage link fault diagnosis method is evaluated by the Monte Carlo simulation method.

[0092] Among them, the compensation effect verification of the uncertainty compensation mechanism can be evaluated by the Monte Carlo simulation method to evaluate the storage link fault diagnosis method. As an example, the evaluation results are as follows:

[0093] Furthermore, before analyzing multi-source data and constructing the feature vectors of each electronic device and the correlation feature vectors between each electronic device, it also includes: Perform data preprocessing on the multi-source data; among them, the data preprocessing includes at least one of data cleaning, noise removal, missing value and outlier processing, normalization processing, and data dimensionality reduction processing.

[0094] Among them, data cleaning can remove errors, duplicates, or inconsistent information in the data to ensure data quality. For example, removing duplicate data, correcting incorrect data, and unifying data formats. Removing noise can reduce random errors or interference in the data and improve data purity. For example, using methods such as moving average and median filtering to smooth the data, using wavelet transform to remove high-frequency noise, and identifying and removing noise points that do not conform to the data distribution law through statistical analysis. Handling missing parts in the data can avoid analysis biases caused by missing values. For example, deleting missing values, filling missing values with mean, median, mode, or model-based predicted values; using linear interpolation and spline interpolation for handling missing values in time series data. Outlier handling can identify and process outliers in the data, which may have a greater impact on the analysis results. For example, using statistical methods such as standard deviation and interquartile range (IQR) to detect outliers; using clustering algorithms or rule-based methods to detect outliers; deleting outliers, correcting outliers, or marking outliers for subsequent analysis. Normalization processing can transform the data to a unified scale to avoid the impact of dimensional differences between different features on the analysis results. Data dimensionality reduction processing can reduce the dimensionality of the data, lower the computational complexity, and remove redundant information at the same time. Through data cleaning, noise removal, handling of missing values and outliers, normalization processing, and data dimensionality reduction processing, the quality and usability of the data can be significantly improved, providing a solid foundation for subsequent analysis and modeling.

[0095] To make the storage link fault diagnosis method provided by the embodiments of the present disclosure clearer, the following will be described with specific examples. The method includes the following processes: 1.1.1 Combine Figure 3 , and it is possible to capture in real time the differential signal amplitude, clock jitter, bit error rate, spindle motor current ripple detection (bandwidth DC - 10MHz), head suspension vibration spectrum analysis (0 - 5kHz FFT), and disc medium temperature gradient monitoring (deploying 4 PT100 sensors per surface) on the storage link (including RAID cards, hard disks, hard disk backplanes, etc.); monitor in real time the voltage values on each unit module of the storage link, and the data on the RAID controller, hard disk sensors, and backplane sensors; collect in real time the log information on the RAID card, the SMART log information of the hard disk, etc.

[0096] 1.1.2 Analyze the data information on the storage link collected by using a fault correlation model. The fault correlation model quantifies the dynamic correlation between the physical states of the RAID controller, hard disk hardware, and backplane to establish a cross-level fault causal chain. First, it is necessary to define feature vectors, and the data dimensions required for defining feature vectors are shown in the following table:

[0097] 1) The construction formula for the feature vector related to RAID is as follows:

[0098] Among them, Uncorrectable_Errors represents the number of errors detected by the RAID controller but unable to be corrected, Total_IO represents the total number of IO requests (including read and write operations); Reconstruct_Time represents the reconstruction time, that is, when a hard disk in the RAID array fails, the time required for the RAID system to recover (reconstruct) the data of this hard disk through redundant data; represents the uncorrectable error rate. If this value is relatively high, it indicates that the data integrity of the RAID system is seriously threatened and may lead to data loss; if this value is relatively low, it indicates that the RAID system is running normally and the data reliability is high.

[0099] 2) The construction formula for the feature vector extracted by the hard disk sensor is as follows:

[0100] Among them, Seek_Error_Rate (seek error rate) refers to the error frequency that occurs when the magnetic head is looking for a specific data position. It is one of the hard disk SMART (Self-Monitoring, Analysis and Reporting Technology) attributes and is used to evaluate the mechanical performance of the hard disk; Raw_Vibration represents the vibration spectrum, that is, the vibration characteristics of the spindle motor or the magnetic head cantilever, and Temperature represents the temperature of the hard disk platter.

[0101] 3) The construction formula for the feature vector related to the hard disk backplane is as follows:

[0102] Among them, Voltage_Ripple represents the amplitude of the power supply voltage fluctuation; Threshold represents the temperature difference in each area of the backplane; UI: Unit Interval, unit interval; Jitter_RMS represents a key indicator of signal timing jitter, indicating the root mean square deviation of the signal edge relative to the ideal clock edge. The calculation formula is:

[0103] Among them: : The timing deviation of the i-th signal edge; : The mean value of TIE; N: The total number of signal edges.

[0104] 4) The construction formula for the cross-level correlation feature vector is as follows:

[0105] This formula is used to quantify the correlation between the hard disk drive (HDD) layer and the backplane (BP) layer, and is weighted in combination with the uncertainty of the RAID log to capture potential patterns of cross - layer failures; when the characteristics of the hard disk and the backplane are highly correlated, the Corr value is large, indicating that there may be cross - layer failures; when the uncertainty of the RAID log is high (high entropy), the f Cross value increases, enhancing the sensitivity to cross - layer failures.

[0106] represents the correlation coefficient between the characteristics of the hard disk drive layer and the backplane layer, which is used to measure the degree of linear correlation between the two. The calculation formula is:

[0107] where: Cov(f HDD , f BP ): covariance, which measures the co - variation between the two; and : are the standard deviations of f HDD and f BP respectively.

[0108] When Cov(f HDD , f BP ) is close to 1: the characteristics of the hard disk and the backplane are highly positively correlated (for example, unstable power supply on the backplane leads to a decrease in hard disk performance); when Cov(f HDD , f BP ) is close to - 1: highly negatively correlated (for example, an increase in the backplane temperature leads to a decrease in the hard disk speed); when Cov(f HDD , f BP ) is close to 0: there is no obvious linear relationship.

[0109] Entropy(RAID_Logs) represents the information entropy of the RAID log, which is used to quantify the uncertainty of the RAID log. The calculation formula is:

[0110] where: p i : the occurrence probability of the i - th type of log event (such as CRC error, reconstruction failure, etc.); N: the total number of log event types.

[0111] To sum up, the feature vector is constructed as follows:

[0112] 1.1.3 Adopt a dynamic weight allocation mechanism to dynamically adjust the contribution weights of each monitoring dimension by real - time sensing the system state (load, temperature, task type, etc.), so as to achieve precise adaptation for fault diagnosis. The specific implementation method is as Figure 4 shown: 1) Weight factor calculation model

[0113] Wherein: : The normalized weight of the i-th type of feature at time t ( ); : Feature type sensitivity coefficient (preset empirical value); : Feature significance score (real-time calculated value).

[0114] 2) Feature significance score calculation formula:

[0115] Wherein, : The rolling 30-day average, representing the historical normal baseline; : Constant 0.7, trend sensitivity coefficient (amplification effect of sudden increase / sudden decrease); : Constant 0.4, cross-feature correlation weight; : Cross-level associated feature vector.

[0116] 3) Sensitive coefficient dynamic adjustment strategy The preset sensitive coefficient matrix is shown in the following table:

[0117] When the system triggers multiple states simultaneously, the coefficients of each state are adjusted by the weighted sum method. By analyzing historical data, the contribution degree of each state to fault diagnosis is statistically analyzed, and the weight is dynamically adjusted:

[0118] Wherein: : The contribution degree of the j-th state in historical faults.

[0119] For example: When RAID reconstruction and high-temperature alarm are triggered simultaneously, and , , then The value of is: .

[0120] 4) Anti-interference design. An outlier filtering mechanism is added in the anti-interference design, that is: for The derivative term in is median filtered as follows, which can eliminate the influence of outliers.

[0121]

[0122] 5) Weight smoothing processing To prevent the weight from mutating, the exponential weighted moving average method is used to smooth the weight, and the processing method is as follows:

[0123] 1.1.4 The multi-source evidence fusion algorithm can be used to achieve the effective fusion of cross-level data. This algorithm is based on the improved Dempster-Shafer (D-S) evidence theory to solve the failure problem of traditional methods in high-conflict scenarios. Assume that the set of fault types is as follows:

[0124] where: A1: Backplane power supply failure; A2: SAS link signal integrity failure; A3: Hard disk media damage; A4: RAID controller logic error; A5: Cross-level coupling failure (such as signal interference caused by vibration).

[0125] Power set space contains all possible fault combinations (a total of elements), for example: the combination represents the simultaneous failure of the backplane power supply and the hard disk media.

[0126] 1.1.5 For each monitoring feature, generate its belief degree for each proposition through an adaptive membership function:

[0127] where: : The feature weight output by the dynamic weight allocation module; : The fault type sensitivity coefficient (trained through historical data); : The decision threshold of type j; : Constant value , mainly to prevent the error of dividing by zero.

[0128] 1.1.6 In a complex storage system, multiple sensors or monitoring modules may provide conflicting evidence (for example, the backplane sensor shows normal voltage, but the hard disk reports errors frequently). Traditional D-S evidence theory may produce counter-intuitive results when dealing with high-conflict scenarios (for example, forcibly allocating high conflicts to uncertain propositions). To solve this problem, this embodiment introduces a conflict redistribution factor and an uncertainty compensation mechanism, and improves the rationality of the fusion result and the system robustness by dynamically adjusting the conflict amount allocation strategy.

[0129] 1.1.7 In the conflict redistribution factor, the conflict amount K represents the degree of conflict between different evidences and is the core index of D-S theory:

[0130] where, , : Basic probability assignment of two independent evidences; A, B: Two mutually exclusive propositions (such as "backplane failure" and "hard disk failure").

[0131] For example: Evidence 1 believes that the probability of "backplane failure" is ; Evidence 2 believes that the probability of "hard disk failure" is ; If A and B are mutually exclusive, then the conflict amount .

[0132] 1.1.8 Traditional D-S (Dempster-Shafer) directly distributes the conflict amount proportionally to non-conflicting propositions, but in high-conflict scenarios () it will lead to distorted results. In this embodiment, a dynamic redistribution factor is introduced to solve the problem of distorted results in D-S. The design of the dynamic redistribution factor is based on the following principles: Conflict weight adjustment: Dynamically adjust the redistribution ratio according to the conflict amount K and evidence uncertainty.

[0133] Historical experience integration: Optimize the distribution strategy by combining the historical conflict scenario library.

[0134] The definition formula is as follows:

[0135] : Represents the joint information entropy of evidence, used to measure uncertainty.

[0136]

[0137] When the value of Entropy is large, it indicates that there is high uncertainty in the evidence. At this time, the evidence is chaotic and the conflict redistribution amount will be reduced to reduce the risk of misjudgment; when the value of Entropy is small, it indicates that the evidence is clear and the conflict redistribution amount will be increased to enhance the conflict handling ability.

[0138] 1.1.9 To solve the problem of evidence combination in high-conflict scenarios and ensure the rationality and stability of the fusion result, a redistribution strategy can be used to dynamically adjust the distribution method of the conflict amount and solve the problem in high-conflict scenarios as follows: 1) Redistribute the conflict amount K: Distribute the conflict amount K to the global uncertain proposition and the high-probability single-fault proposition according to a certain ratio.

[0139] 2) Uncertainty compensation: When the conflict amount K is large, increase the confidence of the global uncertain proposition to avoid blindly trusting a single evidence.

[0140] Redistribution formula:

[0141] is the redistribution factor

[0142] 1.1.10 In the redistribution strategy, as shown in the following example: 1) 60% is allocated to the global uncertain proposition; 2) 40% is allocated to the high-probability single-fault proposition according to the weight.

[0143] To better understand the process of the redistribution strategy, now take an example: the backplane power supply fluctuation causes the hard disk performance to decline, and the calculation process is as follows Figure 5 as shown: a) Input evidence Evidence 1 (backplane):

[0144] Evidence 2 (hard disk):

[0145] b) Calculate the K of the conflict volume. In this example, the backplane fault and the hard disk fault are mutually exclusive propositions. Therefore:

[0146] c) The redistribution strategy is as follows: Assume the entropy , then:

[0147] 60% is allocated to the global uncertain proposition :

[0148] 40% is allocated to the high-probability single-fault proposition according to the weight:

[0149] d) Calculate the synthesis result First calculate

[0150] For the backplane fault:

[0151] For the hard disk fault:

[0152] For :

[0153] Plus the conflict allocation volume K:

[0154]

[0155] e) Normalization processing Since the sum of the synthesis results may not be 1, normalization is required:

[0156] Then the normalization result:

[0157] The final result is: , , .

[0158] 1.1.11 In the storage link, due to the uncertainty of the fault source, there is evidence conflict, and thus there are contradictions in the judgments of the same event by different monitoring modules; due to data noise, sensor sampling errors or environmental interference (such as electromagnetic interference) occur; problems such as model limitations leading to the failure to cover new fault modes (such as signal attenuation caused by quantum effects) exist. Now, an uncertainty compensation mechanism is adopted to solve these problems: 1) Global uncertainty enhancement: When the conflict amount K > 0.8, the system automatically increases the trust degree of the global uncertain proposition as follows. The effect of this is to avoid blindly trusting a single piece of evidence in the case of high conflict.

[0159]

[0160] 2) Expert rule intervention When the conflict amount K > 0.5, the predefined expert rule library is activated (the rules in the expert library are sorted out based on fault handling experience), the matching expert rules are loaded from the database, and the trust distribution is adjusted according to the rule weights. Example rules are as follows: Rule 1, if multiple hard disks report errors simultaneously and the backplane temperature > 70 °C, then it is determined as a backplane heat dissipation fault (weight 0.3); Rule 2, if the RAID reconstruction fails and the power supply ripple suddenly increases, then it is determined as backplane capacitor aging (weight 0.5).

[0161] 1.1.12 Verification of the compensation effect of the uncertainty compensation mechanism. The effectiveness of the compensation mechanism is evaluated through Monte Carlo simulation: 1.1.13 In the decision-making judgment mechanism, according to the synthesized belief and plausibility, the most likely fault type or state is determined. The decision-making judgment model is as follows Figure 6 shown: Among them, the belief (Belief, abbreviated as Bel later) represents the lowest degree of support for a certain proposition, and the calculation formula is:

[0162] Among them: m(B): the basic probability assignment (BPA) of proposition B; : represents all subsets that support proposition A; if Bel(A) is high: it indicates that there is a lot of evidence to support proposition A; if Bel(A) is low: it indicates that there is little evidence to support proposition A.

[0163] 1.1.14 Plausibility (subsequently abbreviated as Pl) represents the highest possible degree of support for a certain proposition, and the calculation formula is:

[0164] where, : represents all subsets that have an intersection with proposition A.

[0165] If Pl(A) is high: it indicates that proposition A may be true; if Pl(A) is low: it indicates that proposition A is unlikely to be true.

[0166] 1.1.15 In the decision-making judgment mechanism, the basic judgment rules are as follows: When the following 3 conditions are met simultaneously, it is determined that proposition A has a fault: 1) Bel(A) > threshold (the initial value is 0.9); 2) Pl(A) - Bel(A) < threshold (the initial value is 0.1); 3) Bel(A) / Bel() > threshold (the initial value is 3.0).

[0167] However, when Pl(A) - Bel(A) is large, it indicates that the uncertainty of proposition A is high and further analysis is required. Therefore, it is necessary to further optimize the decision-making judgment mechanism, and the optimization items include: dynamically adjusting the threshold and the intervention of the expert system. As follows: 1) Dynamically adjusting the threshold: Dynamically adjust the judgment threshold according to the system state. When in a high-load scenario, the threshold needs to be reduced to improve sensitivity; when in a low-load scenario, the threshold needs to be increased to reduce false alarms. For example: In a high-load situation, the threshold is adjusted from 0.9 to 0.8; in a low-load situation, the threshold is adjusted from 0.9 to 0.95.

[0168] 2) Intervention of the expert system: When the uncertainty is high, call the expert system for in-depth analysis, and match the most likely fault type according to historical data and expert experience. For example: When Pl(A) - Bel(A) > 0.2, due to the high uncertainty, it is necessary to call the expert system for further analysis and give the most likely diagnostic result.

[0169] In this way, by integrating multi-dimensional sensor data, a comprehensive solution can be provided for the fault diagnosis of the storage link (including RAID controllers, hard disks, backplanes, etc.), improving the accuracy of fault diagnosis and reducing the false alarm rate; through a unified management platform, cross-layer data collection, analysis, and decision-making are realized, enhancing the operation and maintenance efficiency; through automated fault diagnosis, the dependence on professional operation and maintenance personnel is reduced; the system supports dynamic expansion to adapt to the storage requirements of ultra-large-scale data centers; it also has high versatility and can perform fault diagnosis on hard disks, RAID cards, and hard disk backplanes of different manufacturers and various models.

[0170] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0171] According to an embodiment of the present disclosure, the present disclosure also provides a storage link fault diagnosis device.

[0172] Exemplarily, Figure 7 FIG. is a schematic structural diagram of a storage link fault diagnosis device provided by an embodiment of the present disclosure. The storage link fault diagnosis device 700 includes: A data acquisition module 710, configured to acquire multi-source data of different electronic devices on the storage link; A vector construction module 720, configured to analyze the multi-source data and construct feature vectors of each of the electronic devices and correlation feature vectors between each of the electronic devices; A weight adjustment module 730, configured to adjust the contribution weight corresponding to each feature vector according to the operating state of each of the electronic devices; A data fusion module 740, configured to perform multi-source evidence fusion processing on the contribution weight corresponding to each feature vector and each feature vector to generate a trust degree and a likelihood degree corresponding to each of the electronic devices; A fault diagnosis module 750, configured to determine the fault type of the storage link according to the trust degree and the likelihood degree corresponding to each of the electronic devices.

[0173] Further, the vector construction module 720 is configured to: Input the multi-source data into a preset fault correlation model; Quantify the correlation between each of the electronic devices based on the multi-source data through the preset fault correlation model to obtain feature vectors of each of the electronic devices; Determine the correlation feature vectors between each of the electronic devices based on the feature vectors of each of the electronic devices; Construct the feature vectors of the multi-source data based on the feature vectors of each of the electronic devices and the correlation feature vectors between each of the electronic devices.

[0174] Further, the electronic devices at least include a hard disk, a controller, and a backplane.

[0175] Further, the data of the hard disk includes at least one of differential signal amplitude data, clock jitter data, bit error rate, spindle motor current ripple detection data, head suspension vibration spectrum analysis data, disc media temperature gradient monitoring data, voltage value, and the log information of the hard disk; The data of the controller includes at least one of a voltage value and the log information of the controller; The data of the backplane includes at least one of clock jitter data, spindle motor current ripple detection data, and voltage value.

[0176] Further, the vector construction module 720 is configured to: Quantify the correlation between the hard disk, the controller, and the backplane based on the multi-source data through the preset fault correlation degree model, and obtain the feature vector of the hard disk, the feature vector of the controller, and the feature vector of the backplane; Construct the correlation feature vector between the hard disk and the backplane based on the feature vector of the hard disk and the feature vector of the backplane; Construct the feature vectors of the multi-source data based on the feature vector of the hard disk, the feature vector of the controller, the feature vector of the backplane, and the correlation feature vector between the hard disk and the backplane.

[0177] Further, the weight adjustment module 730 is configured to: Obtain the system operation state; wherein, the system operation state at least includes the load state, temperature, and task type of each of the electronic devices; Determine the feature significance score according to the historical normal baseline, trend sensitivity coefficient, and cross-feature correlation weight; Adjust the feature type sensitivity coefficient according to the system operation state and the preset sensitivity coefficient matrix; Input the feature type sensitivity coefficient and the feature significance score into the preset weight factor calculation model; Generate the normalized weight corresponding to each feature vector based on the feature type sensitivity coefficient and the feature significance score through the preset weight factor calculation model; Statistically analyze the contribution degree of each of the system operation states to fault diagnosis by analyzing historical fault data; Adjust the normalization weight of each eigenvector based on the contribution degree of each system operating state to fault diagnosis to obtain the contribution weight corresponding to each eigenvector.

[0178] Further, the weight adjustment module 730 is configured to: Perform median filtering on the derivative term of each eigenvector; Smooth the normalization weight of each eigenvector after median filtering by using the exponentially weighted moving average method to obtain the contribution weight corresponding to each eigenvector.

[0179] Further, the data fusion module 740 is configured to: Generate the trust degree and likelihood degree of each electronic device having a fault based on each eigenvector and the contribution weight corresponding to each eigenvector through an adaptive membership function; Use a redistribution factor to perform fusion processing on the trust degree and likelihood degree of each electronic device having a fault to generate the trust degree and likelihood degree corresponding to each electronic device.

[0180] Further, the data fusion module 740 is further configured to: Determine the conflict amount based on the trust degree of each electronic device having a fault; Calculate the redistribution factor by using a preset multi-evidence theory algorithm based on the conflict amount; Determine the trust degree and likelihood degree corresponding to each electronic device based on the redistribution factor, the first redistribution factor allocation ratio, and the trust degree of each electronic device having a fault.

[0181] Further, it further includes a confirmation module, which is configured to: Determine the trust degree of global uncertainty when the fault type cannot be judged based on the redistribution factor and the second redistribution factor allocation ratio.

[0182] Further, it further includes a trust degree adjustment module, which is configured to: If the conflict amount is greater than the first conflict amount threshold, load a preset weight allocation rule and adjust the trust degree corresponding to the electronic device based on the preset weight allocation rule; If the conflict amount is greater than the second conflict amount threshold, increase the trust degree of global uncertainty according to a preset adjustment method.

[0183] Further, the fault diagnosis module 750 is configured to: For any fault type A, if the confidence level of the fault type A is higher than a preset confidence threshold, and the difference between the likelihood and the confidence level of the fault type A is less than a first preset difference threshold, and the ratio of the confidence level of the fault type A to the confidence level of the global uncertainty is greater than a preset ratio threshold, then it is determined that the fault type of the storage link is the fault type A; where the fault type A is a hard disk fault, a controller fault, a backplane fault, or the fault type cannot be determined.

[0184] Further, it further includes a threshold adjustment module for: If the difference between the likelihood and the confidence level of the fault type A is greater than or equal to the first preset difference threshold, then according to the system operation state, it is determined whether the storage link is in a high-load scenario; If the storage link is in the high-load scenario, then according to a preset adjustment rule, the preset confidence threshold is decreased; If the storage link is in a low-load scenario, then according to the preset adjustment rule, the preset confidence threshold is increased.

[0185] Further, it further includes a fault determination module for: If the difference between the likelihood and the confidence level of the fault type A is greater than a second preset difference threshold, then an expert system is called to determine the fault type of the storage link based on historical fault data.

[0186] Further, it further includes an evaluation module for: Evaluating the misjudgment rate of the storage link fault diagnosis method through the Monte Carlo simulation method.

[0187] Further, it further includes a preprocessing module for: Performing data preprocessing on the multi-source data; where the data preprocessing includes at least one of data cleaning, noise removal, missing value and outlier processing, normalization processing, and data dimensionality reduction processing.

[0188] It should be noted that the description of the features in the embodiments corresponding to the storage link fault diagnosis device can refer to the relevant descriptions in the embodiments corresponding to the storage link fault diagnosis method, and will not be elaborated here one by one.

[0189] An embodiment of the present disclosure further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the embodiments of the above storage link fault diagnosis method.

[0190] Embodiments of the present disclosure also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the storage link fault diagnosis method when running.

[0191] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0192] Embodiments of the present disclosure also provide a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the storage link fault diagnosis method.

[0193] Embodiments of the present disclosure also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the storage link fault diagnosis method.

[0194] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present disclosure.

[0195] The above has introduced in detail a target detection method provided by the present disclosure. Specific examples are used herein to elaborate on the principles and implementation manners of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principles of the present disclosure, several improvements and modifications can be made to the present disclosure, and these improvements and modifications also fall within the protection scope of the claims of the present disclosure.

Claims

1. A method for diagnosing storage link faults, characterized in that, Including: Collecting multi-source data of different electronic devices on the acquisition and storage link; Analyzing the multi-source data to construct feature vectors of each of the electronic devices and correlation feature vectors between each of the electronic devices; Adjusting the contribution weight corresponding to each feature vector according to the operating state of each of the electronic devices; Performing multi-source evidence fusion processing on the contribution weight corresponding to each feature vector and each feature vector to generate a confidence level and a likelihood level corresponding to each of the electronic devices; Determining the fault type of the storage link according to the confidence level and the likelihood level corresponding to each of the electronic devices.

2. The method according to claim 1, characterized in that, The analyzing the multi-source data to construct feature vectors of each of the electronic devices and correlation feature vectors between each of the electronic devices includes: Inputting the multi-source data into a preset fault correlation degree model; Quantifying the correlation between each of the electronic devices based on the multi-source data through the preset fault correlation degree model to obtain feature vectors of each of the electronic devices; Determining correlation feature vectors between each of the electronic devices based on the feature vectors of each of the electronic devices; Constructing a feature vector of the multi-source data based on the feature vectors of each of the electronic devices and the correlation feature vectors between each of the electronic devices.

3. The method according to claim 2, wherein The electronic devices at least include a hard disk, a controller, and a backplane.

4. The method according to claim 3, wherein The data of the hard disk includes at least one of differential signal amplitude data, clock jitter data, bit error rate, spindle motor current ripple detection data, head suspension vibration spectrum analysis data, disc medium temperature gradient monitoring data, voltage value, and the log information of the hard disk; The data of the controller includes at least one of a voltage value and the log information of the controller; The data of the backplane includes at least one of clock jitter data, spindle motor current ripple detection data, and voltage value.

5. The method according to claim 3, characterized in that The quantifying the correlation between each of the electronic devices based on the multi-source data through the preset fault correlation degree model to obtain feature vectors of each of the electronic devices includes: Quantifying the correlation between the hard disk, the controller, and the backplane based on the multi-source data through the preset fault correlation degree model to obtain a feature vector of the hard disk, a feature vector of the controller, and a feature vector of the backplane; The determining correlation feature vectors between each of the electronic devices based on the feature vectors of each of the electronic devices includes: Constructing a correlation feature vector between the hard disk and the backplane based on the feature vector of the hard disk and the feature vector of the backplane; The constructing a feature vector of the multi-source data based on the feature vectors of each of the electronic devices and the correlation feature vectors between each of the electronic devices includes: Constructing a feature vector of the multi-source data based on the feature vector of the hard disk, the feature vector of the controller, the feature vector of the backplane, and the correlation feature vector between the hard disk and the backplane.

6. The method according to claim 3, wherein The adjusting the contribution weight corresponding to each feature vector according to the operating state of each of the electronic devices includes: Obtaining the system operating state; wherein, the system operating state at least includes the load state, temperature, and task type of each of the electronic devices; Determine the feature significance score according to the historical normal baseline, trend sensitivity coefficient, and cross-feature correlation weight; Adjust the feature type sensitivity coefficient according to the system operating state and the preset sensitivity coefficient matrix; Input the feature type sensitivity coefficient and the feature significance score into a preset weight factor calculation model; Through the preset weight factor calculation model, generate the normalized weight corresponding to each feature vector based on the feature type sensitivity coefficient and the feature significance score; By analyzing historical fault data, statistically calculate the contribution degree of each system operating state to fault diagnosis; Based on the contribution degree of each system operating state to fault diagnosis, adjust the normalized weight of each feature vector to obtain the contribution weight corresponding to each feature vector.

7. The method according to claim 6, wherein The adjusting the normalized weight of each feature vector based on the contribution degree of each system operating state to fault diagnosis to obtain the contribution weight corresponding to each feature vector includes: Perform median filtering on the derivative term of each feature vector; Adopt the method of exponential weighted moving average to smooth the normalized weight of each feature vector after median filtering to obtain the contribution weight corresponding to each feature vector.

8. The method according to claim 1, wherein The performing multi-source evidence fusion processing on the contribution weight corresponding to each feature vector and each feature vector to generate the trust degree and likelihood degree corresponding to each electronic device includes: Based on each feature vector and the contribution weight corresponding to each feature vector, generate the trust degree and likelihood degree of each electronic device having a fault through an adaptive membership function; Adopt a redistribution factor to perform fusion processing on the trust degree and likelihood degree of each electronic device having a fault to generate the trust degree and likelihood degree corresponding to each electronic device.

9. The method according to claim 8, characterized in that, The adopting a redistribution factor to perform fusion processing on the trust degree and likelihood degree of each electronic device having a fault to generate the trust degree and likelihood degree corresponding to each electronic device includes: Based on the trust degree of each electronic device having a fault, determine the conflict amount; Based on the conflict amount, calculate the redistribution factor using a preset multi-evidence theory algorithm; Based on the redistribution factor, the first redistribution factor allocation ratio, and the trust degree of each electronic device having a fault, determine the trust degree and likelihood degree corresponding to each electronic device.

10. The method according to claim 9, wherein It also includes: Based on the redistribution factor and the allocation ratio of the second redistribution factor, determine the trust degree of global uncertainty when the fault type cannot be judged.

11. The method according to claim 10, wherein It also includes: If the conflict amount is greater than the first conflict amount threshold, load a preset weight allocation rule and adjust the trust degree corresponding to the electronic device based on the preset weight allocation rule; If the conflict amount is greater than the second conflict amount threshold, increase the trust degree of global uncertainty according to a preset adjustment method.

12. The method according to claim 3, characterized in that, The determining the fault type of the storage link according to the trust degree and likelihood degree corresponding to each electronic device includes: For any fault type A, if the confidence level of the fault type A is higher than a preset confidence threshold, and the difference between the likelihood and the confidence level of the fault type A is less than a first preset difference threshold, and the ratio of the confidence level of the fault type A to the confidence level of the global uncertainty is greater than a preset ratio threshold, then determine that the fault type of the storage link is the fault type A; wherein, the fault type A is a hard disk fault, a controller fault, a backplane fault, or the fault type cannot be determined.

13. The method according to claim 12, wherein Further included are: If the difference between the likelihood and the confidence level of the fault type A is greater than or equal to the first preset difference threshold, then determine whether the storage link is in a high-load scenario according to the system operation state; If the storage link is in the high-load scenario, then reduce the preset confidence threshold according to a preset adjustment rule; If the storage link is in a low-load scenario, then increase the preset confidence threshold according to the preset adjustment rule.

14. The method according to claim 12, wherein Further included are: If the difference between the likelihood and the confidence level of the fault type A is greater than a second preset difference threshold, then call an expert system to determine the fault type of the storage link based on historical fault data.

15. The method according to claim 11, wherein Further included are: Evaluate the misjudgment rate of the storage link fault diagnosis method through a Monte Carlo simulation method.

16. The method according to claim 1, characterized in that, Before analyzing the multi-source data and constructing the feature vectors of each of the electronic devices and the associated feature vectors between each of the electronic devices, further included is: Perform data preprocessing on the multi-source data; wherein, the data preprocessing includes at least one of data cleaning, noise removal, missing value and outlier processing, normalization processing, and data dimensionality reduction processing.

17. A storage link fault diagnosis device, characterized in that Included are: A data acquisition module, configured to acquire multi-source data of different electronic devices on the storage link; A vector construction module, configured to analyze the multi-source data and construct the feature vectors of each of the electronic devices and the associated feature vectors between each of the electronic devices; A weight adjustment module, configured to adjust the contribution weight corresponding to each feature vector according to the operation state of each of the electronic devices; A data fusion module, configured to perform multi-source evidence fusion processing on the contribution weight corresponding to each feature vector and each feature vector to generate the confidence level and likelihood corresponding to each of the electronic devices; A fault diagnosis module, configured to determine the fault type of the storage link according to the confidence level and likelihood corresponding to each of the electronic devices.

18. An electronic device, characterized in that, Included are: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-16.

19. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-16.

20. A computer program product, characterized in that, Included is a computer program, which implements the method according to any one of claims 1-16 when executed by a processor.

Citation Information

Patent Citations

  • Collaborative fault diagnosis method, system and device for multi-source solid state disk and medium

    CN115543702A

  • Method and device for automatically detecting malicious codes and computer readable storage medium

    CN116010945A

  • Multi-element load prediction method and device for integrated energy system

    CN117035154A

  • SAS link fault diagnosis method and device, equipment and storage medium

    CN117811903A

  • Fault locating method and system based on multi-layer evaluation model

    US20210003640A1