Storage hard disk remote diagnosis system and method based on Internet of Things

Through IoT technology, real-time acquisition and analysis of hard disk data, combined with multi-scale convolutional networks and Bayesian causal networks, the problem of low efficiency of traditional hard disk failure detection is solved, real-time monitoring and early warning of hard disk failures is realized, and fault recovery time and cost are reduced.

CN120256178AInactive Publication Date: 2025-07-04SHENZHEN SANSHANG SCIENCE & TECHNOLOGY CO LTD
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510306994.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-15
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional hard disk fault detection methods are inefficient and cannot achieve real-time monitoring and early warning. Especially in large-scale storage systems, it cannot meet the simultaneous monitoring and diagnosis requirements of hundreds or even thousands of hard disks, resulting in long recovery time and high cost when the fault occurs.

Method used

The Internet of Things-based storage hard disk remote diagnosis system collects SMART parameters, IO operation timing data and ambient temperature and humidity data in real time through the data fusion module, and uses attention mechanism to dynamically allocate multi-source data weights. Combined with multi-scale timing convolution network and Bayesian causal network, it identifies hard disk failure mode and generates a root cause analysis report for failure, calculates the fault risk index through federated incremental learning, and generates a visual operation and maintenance map.

Benefits of technology

Real-time monitoring and early warning of hard disk failures are realized, comprehensiveness and accuracy of fault detection are improved, manual intervention costs are reduced, and business continuity and data security are ensured in complex storage environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256178A_ABST
    Figure CN120256178A_ABST
Patent Text Reader

Abstract

The invention discloses a storage hard disk remote diagnosis system and method based on the Internet of Things, and relates to the technical field of health management of the Internet of Things and storage equipment. The system comprises a data fusion module, a causal analysis module, a risk analysis module and a map construction module. The data fusion module collects SMART parameters, IO operation time sequence data and environment data through Internet of Things equipment, dynamically distributes multi-source data weights by using an attention mechanism, and extracts hard disk health state features. And the causal analysis module is combined with the multi-scale time sequence convolutional network and the Bayesian causal network to identify periodic abnormal fluctuation and generate a fault root cause analysis report. The risk analysis module matches historical cases through federal incremental learning, calculates a hard disk fault risk index and generates an early warning signal. And the atlas construction module optimizes resource isolation, data migration and response paths according to the risk indexes, generates a visual operation and maintenance atlas, provides fault positioning, risk links and repair priorities, and improves the operation and maintenance management efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Internet of Things and storage device health management, and particularly to a remote diagnosis system and method for storage hard disks based on the Internet of Things. Background Art

[0002] With the rapid development of information technology, the Internet of Things (IoT) has become an important part of modern society, especially in the fields of smart home, industrial automation, medical health, etc. As the core storage component of computers and various electronic devices, the reliability and stability of hard disks directly affect the overall performance of the system. However, with the sharp increase in data storage volume, the problems brought by hard disk failures have become increasingly serious. Especially in large-scale storage systems, traditional hard disk fault detection and diagnosis methods can no longer meet the growing demands. For this reason, a remote diagnosis system based on the Internet of Things has emerged. It monitors the running status of hard disks in real time, uses Internet of Things technology to transmit fault information to the cloud, and then realizes remote diagnosis and maintenance. This system can not only improve the efficiency of hard disk fault diagnosis, but also effectively reduce the costs of manual inspection and on-site maintenance, providing technical support for the efficient operation of data centers and large-scale storage systems.

[0003] Currently, traditional hard disk diagnosis methods mainly rely on manual inspection, regular maintenance, and simple hard disk health detection tools. Manual inspection not only has low efficiency, but is also easily affected by human factors, missing potential hard disk fault signals. Regular maintenance often cannot give real-time early warnings and early identifications of faults for some time before the hard disk fails. In addition, although some hard disk health detection tools can monitor parameters such as the temperature and read / write speed of hard disks, due to the lack of intelligent analysis and remote diagnosis functions, their roles are relatively limited and they cannot comprehensively and accurately evaluate the health status of hard disks. Especially in large-scale storage environments, traditional methods cannot meet the needs of simultaneously monitoring and diagnosing hundreds or even thousands of hard disks. When a fault occurs, it often has to rely on manual on-site detection, resulting in long recovery time and high costs. Therefore, there is an urgent need for an intelligent diagnosis system based on the Internet of Things to realize real-time monitoring, early warning, and remote diagnosis of hard disk faults. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a remote diagnosis system and method for storage hard disks based on the Internet of Things, which solves the problems in the above background art.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A remote diagnostic system for storage hard disks based on the Internet of Things, including the following modules: a data fusion module, a causal analysis module, a risk analysis module, and a graph construction module; the data fusion module is used to collect the SMART parameters, IO operation timing data, and environmental temperature and humidity data of the storage hard disk in real time through Internet of Things devices, dynamically allocate multi-source data weights through the attention mechanism, and extract the hard disk health state feature vector to characterize the coupling relationship between the physical loss and logical abnormality of the hard disk; the causal analysis module is used to extract the long-term dependence pattern of SMART parameters and IO delay through a multi-scale time series convolutional network according to the hard disk health state feature vector, identify periodic abnormal fluctuations, and at the same time construct a Bayesian causal network to quantify the contribution of environmental temperature and humidity to the firmware error rate, and generate a root cause analysis report of the fault; the risk analysis module is used to match the historical fault case library through the federated incremental learning framework based on the root cause analysis report of the fault, calculate the hard disk fault risk index, predict the remaining life and generate a hierarchical warning signal; the graph construction module is used to coordinate the resource isolation and data migration strategies of the storage array according to the fault risk index, and at the same time optimize the response path based on reinforcement learning to generate a visual operation and maintenance graph to display the fault location, risk link, and repair priority.

[0006] Further, the specific process of collecting the SMART parameters, IO operation timing data, and environmental temperature and humidity data of the storage hard disk in real time through Internet of Things devices is as follows: Read the SMART parameters through the hard disk controller interface, including the number of remapped sectors, read / write error rate, and media wear indicator; Capture the IO operation timing data through the operating system kernel module, record the read / write delay, queue depth, and completion status; Collect environmental data through the cabinet temperature and humidity sensor, normalize the SMART parameters, IO operation timing data, and environmental temperature and humidity data of the storage hard disk, convert them into a standard format, and upload them to the cloud after encapsulation and encryption.

[0007] Further, the specific process of dynamically allocating multi-source data weights through the attention mechanism and extracting the hard disk health state feature vector is as follows: Map the collected SMART parameters into a query vector to characterize the internal state of the hard disk; Map the IO operation timing data into a key vector to reflect the performance fluctuation characteristics of the hard disk; Map the environmental temperature and humidity data into a value vector to characterize the external environmental impact; Calculate the attention weight through the dynamic weight allocation mechanism, and generate the hard disk health state feature vector through weighted fusion to comprehensively characterize the coupling relationship between the physical loss and logical abnormality of the hard disk.

[0008] Furthermore, the specific process of extracting the long-term dependence patterns of SMART parameters and IO latency through a multi-scale temporal convolutional network and identifying periodic abnormal fluctuations is as follows: Design convolutional kernels with dilation rates to capture hourly, half-dayly, and daily periodic patterns respectively, and identify long-term dependence relationships; Input the hard disk health status feature vector, perform multi-scale temporal convolutional calculations, fuse features of each scale through residual connections, output temporal dependence features, and detect periodic anomalies.

[0009] Furthermore, the specific process of constructing a Bayesian causal network to quantify the contribution of environmental temperature and humidity to the firmware error rate and generating a root cause analysis report for faults is as follows: Use environmental temperature and humidity data as exogenous variables and firmware error events as target nodes; According to the temporal dependence features, combine the joint distribution of environmental temperature and humidity data and firmware error events to construct a Bayesian causal network; Quantify the contribution weight of environmental temperature and humidity to the firmware error rate, generate a root cause analysis report for faults, and clearly label the primary and secondary inducements and the proportion of contribution.

[0010] Furthermore, the specific process of matching the historical fault case library through the federated incremental learning framework based on the root cause analysis report for faults is as follows: Use the root cause analysis report for faults as input to match similar fault patterns in the historical case library; Dynamically adjust the parameters of the pre-trained fault prediction model in the cloud through the federated learning framework, strengthen the sample weights of high-contribution inducements, and combine the hard disk health status feature vector to optimize the generalization ability of the fault prediction model for new hard disk protocols.

[0011] Furthermore, the specific process of calculating the hard disk fault risk index is as follows: Input the health status feature vector into the fault prediction model to output the fault probability; Combine the root cause contribution weight to dynamically calculate the hard disk fault risk index by weighted averaging; Divide the risk level according to a preset threshold and trigger a hierarchical warning signal.

[0012] Furthermore, the specific process of generating a visual operation and maintenance map is as follows: Define the map nodes as hard disk devices, environmental factors, and fault types, and the edges as causal relationships and risk propagation links; Optimize the response strategy according to reinforcement learning, label the shortest repair path, dynamically render the map, and display the fault location, risk link, and repair priority in real time.

[0013] A remote diagnosis method for storage hard disks based on the Internet of Things, comprising the following steps: S1. Real-time collection of SMART parameters, IO operation timing data, and environmental temperature and humidity data of the storage hard disk through Internet of Things devices. Dynamically allocate multi-source data weights through the attention mechanism, and extract the hard disk health status feature vector to characterize the coupling relationship between physical loss and logical anomalies of the hard disk; S2. According to the hard disk health status feature vector, extract the long-term dependence patterns of SMART parameters and IO delays through a multi-scale time series convolutional network, identify periodic abnormal fluctuations, and simultaneously construct a Bayesian causal network to quantify the contribution of environmental temperature and humidity to the firmware error rate, generating a root cause analysis report for faults; S3. Based on the root cause analysis report for faults, match the historical fault case library through the federated incremental learning framework, calculate the hard disk fault risk index, which is used to predict the remaining life and generate a hierarchical warning signal; S4. According to the fault risk index, coordinate the resource isolation and data migration strategies of the storage array, and simultaneously optimize the response path based on reinforcement learning, generating a visual operation and maintenance map for displaying fault location, risk links, and repair priorities.

[0014] The present invention has the following beneficial effects: (1) A remote diagnosis system for storage hard disks based on the Internet of Things, through multi-modal data fusion and dynamic weight allocation technology, realizes the adaptive fusion of SMART parameters, IO timing data, and environmental temperature and humidity, breaks through the limitations of traditional single-dimensional monitoring, and significantly improves the comprehensiveness and accuracy of anomaly detection. Combining multi-scale time series analysis and Bayesian causal reasoning, accurately identify the periodic fault patterns of the hard disk and the contribution weights of environmental incentives, upgrade from passive response to predictive maintenance, effectively reduce the risk of sudden failures, and optimize the operation and maintenance efficiency.

[0015] (2) A remote diagnosis method for storage hard disks based on the Internet of Things, realizes the cold start adaptation of cross-protocol hard disks based on the dynamic transfer learning framework, solves the problem of feature mismatch between new devices and historical data, and at the same time ensures data privacy and model generalization ability. Through intelligent decision-making driven by reinforcement learning and a visual operation and maintenance map, real-time optimize the fault response path, reduce the cost of manual intervention, and ensure business continuity and data security in complex storage environments.

[0016] Of course, it is not necessary for any product implementing the present invention to simultaneously achieve all the above-mentioned advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of a remote diagnosis system for storage hard disks based on the Internet of Things according to the present invention.

[0018] Figure 2 It is a flowchart of a remote diagnosis method for storage hard disks based on the Internet of Things according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0019] In the embodiments of the present application, a remote diagnosis system and method for storage hard disks based on the Internet of Things are provided. Through multimodal data fusion, Bayesian causal networks, and federated incremental learning, the problems of fragmented multi-source data, fuzzy root causes of failures, poor cross-protocol adaptability, and low response efficiency in the remote diagnosis of storage hard disks are solved, and a closed-loop operation and maintenance system from data perception to intelligent decision-making is realized.

[0020] The general idea of the solution in the embodiments of the present application is as follows: Real-time collect the SMART parameters, IO operation timing data, and environmental temperature and humidity data of the storage hard disk through Internet of Things devices, dynamically allocate the weights of multi-source data through the attention mechanism, and extract the hard disk health status feature vector to represent the coupling relationship between the physical loss and logical anomalies of the hard disk.

[0021] According to the hard disk health status feature vector, extract the long-term dependence patterns of SMART parameters and IO latency through a multi-scale temporal convolutional network, identify periodic abnormal fluctuations, and at the same time construct a Bayesian causal network to quantify the contribution of environmental temperature and humidity to the firmware error rate, and generate a root cause analysis report for failures.

[0022] Based on the root cause analysis report for failures, match the historical failure case library through the federated incremental learning framework, calculate the hard disk failure risk index, which is used to predict the remaining life and generate a hierarchical warning signal.

[0023] According to the failure risk index, coordinate the resource isolation and data migration strategies of the storage array, and at the same time optimize the response path based on reinforcement learning to generate a visual operation and maintenance map, which is used to display the fault location, risk link, and repair priority.

[0024] Please refer to Figure 1 , Figure 2, an embodiment of the present invention provides a technical solution: a remote diagnosis system for storage hard disks based on the Internet of Things, including the following modules: a data fusion module, a causal analysis module, a risk analysis module, and a graph construction module; the data fusion module is used to collect the SMART parameters, IO operation timing data, and environmental temperature and humidity data of the storage hard disk in real time through Internet of Things devices, dynamically allocate the weights of multi-source data through the attention mechanism, and extract the hard disk health state feature vector to represent the coupling relationship between the physical loss and logical abnormality of the hard disk; the causal analysis module is used to extract the long-term dependence pattern between the SMART parameters and the IO delay through a multi-scale time series convolutional network according to the hard disk health state feature vector, identify periodic abnormal fluctuations, and at the same time construct a Bayesian causal network to quantify the contribution of environmental temperature and humidity to the firmware error rate, and generate a root cause analysis report of the fault; the risk analysis module is used to match the historical fault case library through the federated incremental learning framework based on the root cause analysis report of the fault, calculate the hard disk fault risk index, predict the remaining life and generate a hierarchical warning signal; the graph construction module is used to coordinate the resource isolation and data migration strategies of the storage array according to the fault risk index, and at the same time optimize the response path based on reinforcement learning to generate a visual operation and maintenance graph for displaying fault location, risk link and repair priority.

[0025] In this implementation plan, there is a data fusion module. The main function of this module is to collect the SMART parameters of the hard disk (information such as the health status, temperature, and bad sectors of the hard disk), the IO operation timing data (performance data such as read / write latency and response time), and the environmental temperature and humidity data (the impact of the environment on the performance of the hard disk) in real time through Internet of Things devices. The multiple collected data will be processed through an attention mechanism, automatically assigning weights according to the importance of different data sources. This process generates a hard disk health status feature vector, which can reflect both the physical wear and logical anomalies of the hard disk, thus providing a comprehensive assessment of the overall state of the hard disk. There is a causal analysis module. In this module, the system will analyze the hard disk health status feature vector, use a multi-scale temporal convolutional network (MS-TCN) to extract the long-term dependence patterns between SMART parameters and IO latency, and identify the periodic abnormal fluctuations of the hard disk at different time periods. In addition, this module also constructs a Bayesian causal network (BCN), which can quantify the specific impact of environmental temperature and humidity on hard disk failures (such as firmware error rate). Through these analyses, the module can generate a root cause analysis report of the failure, providing a basis for subsequent failure diagnosis and prevention. There is a risk analysis module. Based on the root cause analysis report of the failure, the system will use a federated incremental learning framework to match similar failure cases from the historical failure case library and dynamically adjust the parameters of the pre-trained model. The core role of this module is to calculate the failure risk index of the hard disk, that is, to predict the remaining service life of the hard disk based on historical data and the current health status. Based on this index, the system will generate hierarchical warning signals to ensure the timely discovery of potential failures and avoid hard disk crashes or data loss. There is a graph construction module. After the failure risk index is calculated, the security decision module will adopt corresponding resource isolation and data migration strategies according to the warning signals to ensure that the data in the storage array will not be affected by a single hard disk failure. At the same time, based on reinforcement learning, this module will optimize the failure response path and generate a visual operation and maintenance graph. This graph clearly shows the fault location, risk link, and repair priorities, helping operation and maintenance personnel make quick and effective decisions to reduce maintenance costs and shorten repair time.

[0026] Specifically, the specific process of collecting the SMART parameters, IO operation timing data, and environmental temperature and humidity data of the storage hard disk in real time through Internet of Things devices is as follows: Read the SMART parameters through the hard disk controller interface, including the number of remapped sectors, read / write error rate, and media wear indicator; Capture the IO operation timing data through the operating system kernel module, record the read / write latency, queue depth, and completion status; Collect environmental data through the cabinet temperature and humidity sensor, normalize the SMART parameters, IO operation timing data, and environmental temperature and humidity data of the storage hard disk, convert them into a standard format, encapsulate and encrypt them, and then upload them to the cloud.

[0027] In this implementation scheme, for the SMART parameters, the remapped sector count: When the hard disk detects that a certain sector (storage block) is damaged, the hard disk will redirect the data to a spare sector. The remapped sector count represents the number of sectors that have failed and been replaced on the hard disk. This parameter can reflect the degree of physical damage to the hard disk. Read / Write Error Rate: It represents the number of errors that occur during the read / write process of the hard disk. A higher error rate usually indicates that the hard disk has a fault, which may affect the reliability of the stored data. Media Wear Indicator: For Solid State Drives (SSDs), this indicator represents the degree of wear of the storage medium. After long-term read / write operations, the durability of the storage cells will decline, and this parameter is used to monitor the wear status of the SSD. Read / Write Latency: When the hard disk performs a data read or write operation, it is the time difference from the request initiation to the data completion. A higher latency usually indicates a decline in the performance of the hard disk, which may be due to high hard disk load or hardware problems. Queue Depth: It indicates how many IO operations are currently pending in the system. A too high queue depth may indicate an overloaded hard disk or system, which may lead to increased latency. Completion Status: Records whether each IO operation is successfully completed. If an IO operation fails, it may mean that there is a potential fault in the hard disk. Environmental temperature and humidity data: Temperature: The operating temperature of the hard disk directly affects its performance and lifespan. Too high a temperature may cause damage to the hard disk components or accelerate wear, and may even cause the hard disk to completely fail. Humidity: Too high or too low humidity may cause rust or static electricity in the internal circuit of the hard disk, affecting the stability of the hard disk and the reliability of the data.

[0028] Specifically, the specific process of dynamically allocating multi-source data weights through the attention mechanism to extract the hard disk health status feature vector is as follows: Map the collected SMART parameters to a query vector to represent the internal state of the hard disk; Map the IO operation timing data to a key vector to reflect the hard disk performance fluctuation characteristics; Map the environmental temperature and humidity data to a value vector to represent the external environmental impact; Through the dynamic weight allocation mechanism, calculate the attention weights, and generate the hard disk health status feature vector through weighted fusion to comprehensively represent the coupling relationship between the physical loss and logical abnormality of the hard disk.

[0029] In this implementation scheme, map the SMART parameters to a query vector. Query Vector: The SMART parameters are used to represent the internal state of the hard disk and are mapped to a query vector (Q). The role of the query vector is to calculate the contribution degree of each data source to the hard disk health according to the internal state of the hard disk (such as health indicators, error rates, etc.). The formula is expressed as: ; where f represents a mapping function that converts SMART parameters into vectors suitable for calculation through an algorithm. The timing data of IO operations is mapped into a key vector. KeyVector: The timing data of IO operations is used to reflect the performance fluctuation characteristics of the hard disk, such as latency, queue depth, etc., and is mapped into a key vector (K). The key vector reflects the performance of the hard disk under specific operations. The formula is expressed as: ; where f represents a function that maps the timing data of IO operations into a feature vector. The environmental temperature and humidity data is mapped into a value vector; ValueVector: The environmental temperature and humidity data is used to characterize the impact of the external environment on the performance of the hard disk and is mapped into a value vector (V). The value vector reflects the possibility of hard disk failure due to temperature and humidity and the external influencing factors of hard disk health. The formula is expressed as: ; where f represents a function that maps the environmental temperature and humidity data into a value vector. Calculate the attention weight through a dynamic weight allocation mechanism. AttentionWeight: According to the attention mechanism, the contributions of the internal state of the hard disk, performance fluctuations, and environmental temperature effects to the health state will dynamically adjust the weights according to their importance. The weight calculation formula is as follows: ; where: : The i-th query vector. : The i-th key vector. : The attention weight of the i-th data source. represents the vector dot product operation. This formula uses the inner product of the query vector and the key vector to calculate the impact of each data source (such as SMART parameters, IO timing data, environmental temperature data) on the health state of the hard disk. The calculated attention weights will be normalized according to the importance of each data source. Weighted fusion generates the health state feature vector of the hard disk; finally, use the calculated attention weights to perform weighted fusion on the value vector to generate the final health state feature vector (H) of the hard disk. The fusion formula is as follows: ; where: : The attention weight of the i-th data source (such as SMART parameters, IO operations, environmental temperature data). : The i-th value vector. H: The final health state feature vector of the hard disk.

[0030] Specifically, the long-term dependence patterns of SMART parameters and IO latency are extracted through a multi-scale temporal convolutional network. The specific process of identifying periodic abnormal fluctuations is as follows: Design convolutional kernels with dilation rates to capture hourly, half-dayly, and daily periodic patterns respectively, and identify long-term dependence relationships; Input the health state feature vector of the hard disk, perform multi-scale temporal convolutional calculations, fuse the features of each scale through residual connections, output the temporal dependence features, and detect periodic abnormalities.

[0031] In this implementation, convolution kernels with designed dilation rates are used to capture hourly, semi-daily, and daily periodic patterns respectively: Dilated Convolution: Dilated convolution is a technique for expanding the receptive field of a convolution kernel, which can help the network capture features with a longer time span without increasing the computational complexity. By adjusting the dilation rate, some input data can be skipped during the convolution operation, thereby capturing longer temporal dependencies. The formula is expressed as: ; where: : The value of the input temporal data (SMART parameter or I / O latency data) at time step t. : The weight of the convolution kernel corresponding to time step t, which is the convolution kernel of the k-th layer of convolution. : The output result of the dilated convolution. : The length of the temporal data. By adjusting the dilation rate, the convolution kernel can maintain an appropriate time span when capturing hourly, semi-daily, and daily periodic patterns. For example: the hourly period captures the data fluctuations per hour; the semi-daily period captures the trend changes within half a day; the daily period captures the change patterns of the whole-day data. Input the hard disk health status feature vector and perform multi-scale temporal convolution calculation: Multi-scale Convolution: The multi-scale convolution network captures the temporal features of the hard disk health status simultaneously at multiple scales by using convolution kernels with different dilation rates. Taking the hard disk health status feature vector as the input, it is processed through multiple convolutional layers (each convolutional layer has a different dilation rate). The formula is expressed as: ; where: : The input hard disk health status feature vector (the previously calculated hard disk health status feature vector ). : The convolution kernel weight (k represents the label of the convolutional layer). : The dilation rate, which controls the size of the receptive field. Fusing features of each scale through residual connection Residual Connection: Through the residual connection, the convolution results of different scales are directly fused to avoid the problem of gradient disappearance and accelerate convergence. The residual connection enables the network to directly transmit features at different levels, thereby enhancing the expressive power of the network. The formula is expressed as: ; where: : The output result of the k-th layer of convolution. : The output result of the previous layer of convolution. Through the residual connection, features of each scale are fused to output a representation with multi-scale dependent features. Outputting temporal dependence features and detecting periodic anomalies: Temporal Dependence Features: The features extracted by the multi-scale convolutional layer comprehensively consider the long-term dependence patterns of SMART parameters and IO latency. Finally, the multi-scale features after fusion and residual connection output temporal dependence features (D). The formula is expressed as: ; where: D: the final temporal dependence feature. : The output feature after residual connection in the k-th layer. Periodic anomaly detection: Through the output temporal dependence feature, the network can identify potential periodic anomaly fluctuations in SMART parameters and IO latency, especially detecting regularly occurring patterns and fluctuations deviating from the normal pattern. Combining temporal analysis tools or threshold judgment methods, the moments of periodic anomalies can be automatically marked.

[0032] Specifically, the specific process of constructing a Bayesian causal network to quantify the contribution of environmental temperature and humidity to the firmware error rate and generating a root cause analysis report of faults is as follows: Taking environmental temperature and humidity data as exogenous variables and firmware error events as target nodes; According to the temporal dependence feature, combining the joint distribution of environmental temperature and humidity data and firmware error events, constructing a Bayesian causal network; Quantifying the contribution weight of environmental temperature and humidity to the firmware error rate, generating a root cause analysis report of faults, and clearly marking the primary and secondary inducements and the proportion of contribution.

[0033] In this implementation plan, taking environmental temperature and humidity data as exogenous variables and firmware error events as target nodes: Environmental temperature and humidity data: In this process, environmental temperature and humidity data are regarded as exogenous variables (i.e., data not affected by other variables in the network). Environmental temperature and humidity can affect the operating state of the hard disk and the firmware error rate. Firmware error events: Firmware error events are target nodes, indicating the occurrence of hard disk firmware failures or errors. The firmware error rate is affected by various factors, and environmental temperature and humidity are one of them. Constructing a Bayesian causal network, a Bayesian causal network is a graphical model where nodes represent variables and edges represent the causal relationships between variables. The conditional probability distribution on each edge represents the probability distribution of the child node given the parent node. The causal relationship between environmental temperature and humidity and the firmware error rate: We model the relationship between environmental temperature and humidity data (E) and firmware error events (F) through the joint distribution. The formula is expressed as: where: : The conditional probability of firmware error event F given environmental temperature and humidity data E. : The conditional probability of environmental temperature and humidity data E given firmware error event F, reflecting the correlation between temperature and humidity and error events. : The prior probability of the occurrence of firmware error events. : Prior distribution of environmental temperature and humidity data. By constructing a Bayesian causal network, we can use the joint distribution of environmental temperature and humidity data and firmware error events to infer and quantify the impact of environmental factors on firmware errors. Quantifying the contribution weight of environmental temperature and humidity to the firmware error rate: The Bayesian causal network can be used to quantify the contribution of environmental temperature and humidity to the firmware error rate. This can be achieved by calculating the prior probability distribution of environmental temperature and humidity on firmware error events, reflecting the impact degree of environmental factors on firmware errors. To quantify the impact degree, we can calculate the information gain or mutual information, which will help us evaluate the contribution of environmental temperature and humidity to the firmware error rate. The formula for mutual information (I(E,F)) is as follows: ; where: : Represents the mutual information between environmental temperature and humidity data and firmware error events, measuring the degree of information sharing between them. : Joint probability distribution of environmental temperature and humidity data e and firmware error event f. and : Are the marginal probabilities of environmental temperature and humidity data and firmware error events respectively. By calculating the mutual information, we can quantify the contribution of environmental temperature and humidity to firmware error events and generate a report based on their contribution weights. Generating a root cause analysis report for faults: Based on the above calculation results, we can generate a root cause analysis report for faults. The report will list all possible inducing factors and their contribution degrees, clearly marking the primary and secondary inducing factors and their corresponding contribution ratios. The report content includes the following: The influence weight of each factor on the firmware error rate (quantified by calculating mutual information or conditional probability). The proportion of the influence of environmental temperature and comparison with other potential factors (such as hard disk usage time, IO operations, etc.). The primary and secondary ranking of the inducing factors, clarifying which factors contribute the most to the firmware error rate. The report generated by the Bayesian causal network can clearly point out the role of environmental temperature and humidity in the occurrence of faults and provide decision-making support for subsequent maintenance and optimization.

[0034] Specifically, based on the root cause analysis report for faults, the specific process of matching the historical fault case library through the federated incremental learning framework is as follows: Use the root cause analysis report for faults as input to match similar fault patterns in the historical case library; Dynamically adjust the parameters of the pre-trained fault prediction model in the cloud through the federated learning framework, strengthen the sample weights of high-contribution inducing factors, and combine the hard disk health status feature vector to optimize the generalization ability of the fault prediction model for new hard disk protocols.

[0035] In this implementation plan, the root cause analysis report of the failure is used as the input to match the similar failure modes in the historical case database: Root cause analysis report of the failure: In the previous steps, the root cause analysis report of the failure has been generated through the Bayesian causal network. The report contains the main inducements of the hard disk failure and the relevant environmental variables. This report provides an accurate description of the failure mode for the model. Historical case database: The historical case database is a database containing a large number of previous failure cases and their related parameters (such as hard disk SMART parameters, environmental data, IO operation data, etc.). Each case contains the state at the time of the failure and the corresponding hard disk characteristics. Matching similar failure modes: By calculating the similarity between the root cause analysis report of the failure and each failure case in the historical case database, the system can identify the historical case that is most similar to the current failure situation. This process usually uses the vector space model or other similarity measurement methods (such as cosine similarity, Euclidean distance, etc.). Dynamically adjusting the parameters of the pre-trained failure prediction model in the cloud through the federated learning framework: Federated learning framework: Federated learning is a distributed machine learning method that enables multiple devices (such as different hard disks or monitoring devices) to perform training locally and upload the updated model parameters to the cloud for integration without directly transmitting local data. Through this framework, a general model can be jointly trained while ensuring data privacy. Dynamically adjusting the parameters of the pre-trained failure prediction model: The model usually exists in the cloud in a pre-trained form and is continuously optimized through the uploaded local device data. According to the similar failure modes matched in the historical case database, the federated learning framework will adjust the parameters of the failure prediction model in the cloud to enhance the model's prediction ability for the current failure type. Adjustment process: Each device (or node) uploads the updates calculated locally (such as gradients, weight updates) to the cloud. The cloud receives and integrates these parameter updates, and then performs a global update on the model. Through multiple iterations, the cloud can optimize the failure prediction model to make it adaptable to more failure modes. Strengthening the sample weights of high-contribution inducements: In the root cause analysis report of the failure, it has been determined through the Bayesian causal network which factors contribute the most to the hard disk failure. For example, certain environmental variables (such as temperature and humidity) may play a dominant role in certain failure situations. Strengthening the sample weights of high-contribution inducements: To enhance the model's sensitivity to high-contribution factors, the training process of the model can be optimized by increasing the weights of these factors. This means that during model training, high-contribution factors will have a greater impact on the final result, making the failure prediction more accurate. Calculating sample weights: The samples can be weighted according to the contribution degree in the root cause analysis report. Usually, a contribution-degree-based weight adjustment strategy is used to give higher priority to high-contribution factors during the training process.Combined with the hard disk health status feature vector, optimize the generalization ability of the fault prediction model for new hard disk protocols: Hard disk health status feature vector: The hard disk health status feature vector contains the health status information of the hard disk (such as SMART parameters, IO operations, environmental data, etc.), which is the input data of the fault prediction model. By combining these feature vectors with historical cases, the model can more accurately identify the potential fault risks of the hard disk. Optimize the generalization ability: With the continuous development of hard disk technology, new hard disk protocols and types have emerged. To ensure that the fault prediction model can adapt to new hard disk protocols, it is necessary to enhance the generalization ability of the model, that is, to enable the model to effectively predict the fault conditions of different types of hard disks. Optimize the generalization ability through federated incremental learning: The model is trained by combining data from multiple devices (i.e., different hard disk protocols and device types). Each local training and parameter update of the device will make the model more general, thereby enhancing its adaptability to new hard disk protocols.

[0036] Specifically, the specific process of calculating the hard disk fault risk index is as follows: Input the health status feature vector into the fault prediction model to output the fault probability; Combine the root cause contribution weight and dynamically weight and calculate the hard disk fault risk index; Divide the risk level according to the preset threshold and trigger a hierarchical warning signal.

[0037] In this implementation plan, input the health status feature vector into the fault prediction model to output the fault probability: Hard disk health status feature vector: This vector contains multi-dimensional health data of the hard disk, such as SMART parameters, IO operation timing data, environmental temperature and humidity data, etc. It is used as the input of the fault prediction model. Fault prediction model: Use a trained machine learning or deep learning model, usually a regression model or a classification model, which has learned the ability to infer hard disk faults from the health status feature vector. Output the fault probability: The model predicts the probability of the hard disk failing by inputting the health status feature vector, that is, the possibility of the fault occurring. The formula representation of the hard disk fault risk index: ; Parameter explanation: : Hard disk fault risk index. N: The number of influencing factors (for example, SMART parameters, IO latency, environmental temperature, etc.). : The th root cause contribution weight of the influencing factor, indicating the influence degree of each factor on the risk index. : Activation function, usually Sigmoid, used for non-linear transformation to ensure that the output is between 0 and 1 and simulate the sharp fluctuation of the fault probability. The formula is: ; M: The number of sub-features within the th factor, considering refined features such as different time scales, different attributes, etc. : The th under the The weight coefficient of sub - features, which is used to quantify the contribution of each sub - feature to the failure probability. : The failure probability or predicted value of the th sub - feature under the : The global adjustment factor, which is used to globally adjust the impact of health status features on the failure risk. : The hard - disk health status feature function, which is based on the input health status feature vector and preset parameters , and is a function for a deep neural network to extract higher - level risk features from the state of the hard disk. The system determines the thresholds of multiple failure risk indices by analyzing historical data and hard - disk failure modes. According to the risk index values, the system classifies the failure risk of the hard disk into different levels: low risk, medium risk, and high risk. Graded warning signal: According to the calculated risk index, the system triggers the corresponding warning signal. For example, if the risk index is higher than a certain threshold, a high - risk alarm is triggered to remind the administrator to perform maintenance preferentially.

[0038] Specifically, the specific process of generating the visual operation and maintenance map is as follows: Define the map nodes as hard - disk devices, environmental factors, and failure types, and the edges as causal relationships and risk propagation links; Optimize the response strategy according to reinforcement learning, mark the shortest repair path, dynamically render the map, and display the fault location, risk link, and repair priority in real - time.

[0039] In this implementation plan, the graph nodes are defined as follows: Hard disk device: One of the core nodes of the graph, representing the status information of each hard disk device, including its health status, fault prediction results, related SMART parameters, IO operation data, etc. Environmental factors: Include the parameters of the hard disk operating environment (such as temperature, humidity, vibration, etc.), and there may be a causal relationship between these environmental factors and hard disk failures or performance degradation. Fault types: Represent various types of faults that the hard disk may encounter, such as physical damage to the hard disk, firmware errors, data loss, etc. The graph edges are defined as follows: Causal relationship: The edges in the graph represent the causal relationships between different nodes. For example, environmental factors may affect the health status of the hard disk device, and the fault type of the hard disk device may affect the operating status of other devices. Through these edges, the potential causes of hard disk failures can be analyzed. Risk propagation link: The edges of the graph also represent how risks spread from one node to another. For example, a hard disk failure may trigger a chain reaction, affecting other devices or systems, and the risk propagation link can help identify potential fault propagation paths. Reinforcement learning to optimize response strategies: Reinforcement learning (RL) is used to optimize the response strategies for dealing with faults. Through the reinforcement learning model, the operation strategies can be dynamically adjusted based on real-time monitoring data and historical fault cases to optimize the maintenance and repair behaviors of the system. Reward function: The system evaluates the quality of the response strategy based on indicators such as repair efficiency, reduction of fault impact, and repair time. Reinforcement learning selects the optimal fault response measures by continuously exploring and utilizing feedback. Mark the shortest repair path: Through the nodes and edges in the graph, the system can calculate the shortest path from the current fault state to the completion of the repair. This path usually takes into account the causal relationships and priorities between each repair step. Annotation of the repair path: Each repair step is marked according to its priority and importance, and may combine factors such as the severity of the fault and its impact on the system to optimize the repair order. Dynamically render the graph: The rendering process of the graph is dynamically updated according to real-time data streams and system states. For example, when a hard disk device fails, the graph will display its health status, repair path, and scope of influence in real time. Dynamic rendering: Using data visualization technologies (such as graphical interfaces, interactive graphs, etc.), enables operation and maintenance personnel to intuitively view fault information and system operating conditions. Real-time display of fault location, risk link, and repair priority: The graph visually shows the specific location and scope of influence of the fault, helping operation and maintenance personnel quickly locate problems. At the same time, the display of the risk link can visually present the propagation path of the fault, reminding operation and maintenance personnel to pay attention to potential risk diffusion. The annotation of the repair priority will be dynamically adjusted based on real-time data to ensure that the most urgent faults can be repaired first.

[0040] A remote diagnosis method for storage hard disks based on the Internet of Things, comprising the following steps: S1. Real-time collection of SMART parameters, IO operation timing data, and environmental temperature and humidity data of the storage hard disk through Internet of Things devices, dynamically allocating multi-source data weights through an attention mechanism, and extracting a hard disk health status feature vector for characterizing the coupling relationship between physical loss and logical abnormality of the hard disk; S2. According to the hard disk health status feature vector, extracting the long-term dependence pattern of SMART parameters and IO delay through a multi-scale time series convolutional network, identifying periodic abnormal fluctuations, and simultaneously constructing a Bayesian causal network to quantify the contribution of environmental temperature and humidity to the firmware error rate, and generating a root cause analysis report of the fault; S3. Based on the root cause analysis report of the fault, matching the historical fault case library through a federated incremental learning framework, calculating the hard disk fault risk index for predicting the remaining life and generating a hierarchical warning signal; S4. According to the fault risk index, coordinating the resource isolation and data migration strategies of the storage array, and simultaneously optimizing the response path based on reinforcement learning, and generating a visual operation and maintenance map for displaying fault location, risk link, and repair priority.

[0041] In this implementation scheme, S1: Real-time collection of key data of the hard disk (such as SMART parameters, IO operation data, environmental temperature and humidity data) through Internet of Things devices, and dynamically adjusting the weights of different data sources through an attention mechanism to generate a hard disk health status feature vector. This feature vector comprehensively reflects the physical loss and logical abnormality of the hard disk. S2: Using a multi-scale time series convolutional network to analyze the hard disk health status feature vector, so as to extract the long-term dependence pattern of SMART parameters and IO delay, and identify periodic abnormal fluctuations. At the same time, construct a Bayesian causal network to analyze the impact of environmental factors (such as temperature and humidity) on the firmware error rate, and generate a root cause analysis report of the fault. S3: Based on the root cause analysis report of the fault, use a federated incremental learning framework to match the historical fault case library, calculate the hard disk fault risk index, predict the remaining life of the hard disk, and generate a hierarchical warning signal. S4: According to the calculated fault risk index, adjust the resource isolation and data migration strategies of the storage array, and optimize the response path through reinforcement learning. Finally, generate a visual operation and maintenance map to display the fault location, risk propagation path, and repair priority, helping the operation and maintenance personnel to quickly respond to the fault.

[0042] In summary, this application has at least the following effects: An Internet of Things-based remote diagnosis system and method for storage hard drives can collect SMART parameters, IO operation timing data, and environmental temperature and humidity data of the hard drive in real time through Internet of Things devices. By combining the attention mechanism and multi-scale convolutional networks, it can accurately extract the health status features of the hard drive and identify the coupling relationship between physical wear and logical anomalies of the hard drive. The long-term dependence patterns of SMART parameters and IO latency are extracted using a multi-scale temporal convolutional network to identify periodic abnormal fluctuations, providing a strong basis for fault warning and maintenance. The influence of environmental temperature and humidity on the firmware error rate is analyzed through a Bayesian causal network, quantifying the role of environmental factors in hard drive failures and generating a root cause analysis report for fault prevention. A fault risk index is calculated by matching the historical fault case library through a federated incremental learning framework to achieve fault prediction, remaining life estimation, and hierarchical warning, improving the accuracy and response efficiency of fault prediction. Based on the fault risk index, the storage array resource isolation and data migration strategies are intelligently coordinated, and at the same time, the fault response path is optimized through reinforcement learning to generate a visual operation and maintenance map to help operation and maintenance personnel quickly locate faults and optimize repair priorities.

[0043] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0044] The present invention is described with reference to the flowcharts and / or block diagrams of systems, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.

[0045] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions in Figure 1 one process or multiple processes and / or blocksFigure 1 The functions specified in one or more boxes.

[0046] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 or more boxes.

[0047] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0048] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A remote diagnosis system for storage hard disks based on the Internet of Things, characterized in that, It includes the following modules: a data fusion module, a causal analysis module, a risk analysis module, and a graph construction module; The data fusion module is used to collect the SMART parameters of the storage hard disk, the IO operation timing data, and the environmental temperature and humidity data in real time through Internet of Things devices, dynamically allocate the weights of multi-source data through an attention mechanism, and extract the hard disk health status feature vector to characterize the coupling relationship between the physical loss and logical anomalies of the hard disk; The causal analysis module is used to extract the long-term dependence pattern between SMART parameters and IO delay through a multi-scale time series convolutional network according to the hard disk health status feature vector, identify periodic abnormal fluctuations, and at the same time construct a Bayesian causal network to quantify the contribution of environmental temperature and humidity to the firmware error rate, and generate a root cause analysis report of the fault; The risk analysis module is used to match the historical fault case library through a federated incremental learning framework based on the root cause analysis report of the fault, calculate the hard disk fault risk index, predict the remaining life and generate a hierarchical warning signal; The graph construction module is used to coordinate the resource isolation and data migration strategies of the storage array according to the fault risk index, and at the same time optimize the response path based on reinforcement learning to generate a visual operation and maintenance graph to display the fault location, risk link and repair priority.

2. The remote diagnosis system for a storage hard disk based on the Internet of Things according to claim 1, wherein: The specific process of collecting the SMART parameters of the storage hard disk, the IO operation timing data, and the environmental temperature and humidity data in real time through Internet of Things devices is as follows: Read the SMART parameters through the hard disk controller interface, including the number of remapped sectors, read / write error rate, and media wear indicator; Capture the IO operation timing data through the operating system kernel module, and record the read / write delay, queue depth, and completion status; Collect environmental data through the cabinet temperature and humidity sensor, normalize the SMART parameters of the storage hard disk, the IO operation timing data, and the environmental temperature and humidity data, convert them into a standard format, encapsulate and encrypt them, and upload them to the cloud.

3. The remote diagnosis system for a storage hard disk based on the Internet of Things according to claim 2, wherein: The specific process of dynamically allocating the weights of multi-source data through an attention mechanism and extracting the hard disk health status feature vector is as follows: Map the collected SMART parameters into a query vector to characterize the internal state of the hard disk; Map the IO operation timing data into a key vector to reflect the hard disk performance fluctuation characteristics; Map the environmental temperature and humidity data into a value vector to characterize the external environmental impact; Through the dynamic weight allocation mechanism, calculate the attention weight, and generate the hard disk health status feature vector through weighted fusion to comprehensively characterize the coupling relationship between the physical loss and logical anomalies of the hard disk.

4. The remote diagnosis system for a storage hard disk based on the Internet of Things according to claim 3, characterized in that: The specific process of extracting the long-term dependence pattern between SMART parameters and IO delay through a multi-scale time series convolutional network and identifying periodic abnormal fluctuations is as follows: Design convolutional kernels with dilation rates to capture hourly, half-dayly, and daily periodic patterns respectively, and identify long-term dependence relationships; Input the hard disk health status feature vector, perform multi-scale time series convolutional calculations, fuse the features of each scale through residual connections, output the time series dependence features, and detect periodic anomalies.

5. The remote diagnosis system for a storage hard disk based on the Internet of Things according to claim 4, wherein: The specific process of constructing a Bayesian causal network to quantify the contribution of environmental temperature and humidity to the firmware error rate and generating a root cause analysis report of the fault is as follows: Take the environmental temperature and humidity data as exogenous variables, and the firmware error event as the target node; Construct a Bayesian causal network based on the time-series dependence characteristics, combined with the joint distribution of environmental temperature and humidity data and firmware error events; Quantify the contribution weight of environmental temperature and humidity to the firmware error rate, generate a root cause analysis report of the failure, and clearly mark the primary and secondary incentives and the proportion of contribution.

6. The remote diagnostic system for a storage hard disk based on the Internet of Things according to claim 5, wherein: The specific process of matching the historical failure case library through the federated incremental learning framework based on the root cause analysis report of the failure is as follows: Take the root cause analysis report of the failure as the input and match the similar failure modes in the historical case library; Dynamically adjust the parameters of the pre-trained failure prediction model in the cloud through the federated learning framework, strengthen the sample weights of the incentives with high contribution, and combine the hard disk health status feature vector to optimize the generalization ability of the failure prediction model for the new hard disk protocol.

7. The remote diagnosis system for a storage hard disk based on the Internet of Things according to claim 6, characterized in that: The specific process of calculating the hard disk failure risk index is as follows: Input the health status feature vector into the failure prediction model and output the failure probability; Combine the root cause contribution weight and dynamically calculate the hard disk failure risk index by weighted method; Divide the risk level according to the preset threshold and trigger a hierarchical warning signal.

8. The remote diagnosis system for storage hard disks based on the Internet of Things according to claim 7, characterized in that: The specific process of generating a visual operation and maintenance graph is as follows: Define the graph nodes as hard disk devices, environmental factors, and failure types, and the edges as causal relationships and risk propagation links; According to the reinforcement learning to optimize the response strategy, mark the shortest repair path, dynamically render the graph, and display the fault location, risk link and repair priority in real time.

9. A remote diagnosis method for a storage hard disk based on the Internet of Things, which is applied to a remote diagnosis system for a storage hard disk based on the Internet of Things described in any one of claims 1-8, characterized in that, It includes the following steps: S1. Real-time collect the SMART parameters, IO operation timing data and environmental temperature and humidity data of the storage hard disk through the Internet of Things device, dynamically allocate the weights of multi-source data through the attention mechanism, and extract the hard disk health status feature vector to characterize the coupling relationship between the physical loss and logical abnormality of the hard disk; S2. According to the hard disk health status feature vector, extract the long-term dependence patterns of SMART parameters and IO delay through the multi-scale time-series convolutional network, identify the periodic abnormal fluctuations, and at the same time construct a Bayesian causal network to quantify the contribution of environmental temperature and humidity to the firmware error rate, and generate a root cause analysis report of the failure; S3. Based on the root cause analysis report of the failure, match the historical failure case library through the federated incremental learning framework, calculate the hard disk failure risk index, and use it to predict the remaining life and generate a hierarchical warning signal; S4. According to the failure risk index, coordinate the resource isolation and data migration strategies of the storage array, and at the same time optimize the response path based on reinforcement learning to generate a visual operation and maintenance graph to display the fault location, risk link and repair priority.

Citation Information

Cited By

  • Hard disk fault detection method and electronic equipment

    CN120469846A

  • Deep learning-based unified information UOS system fault rapid repair method

    CN120523641A

  • Hard disk fault prediction method, electronic equipment and storage medium

    CN120849203A

  • Hard drive failure prediction methods, electronic devices and storage media

    CN120849203B

  • Intelligent monitoring system based on multi-sensor fusion technology

    CN120909194A