Network fault management system and method based on multi-source data

By building an adaptive sampling frequency control function and logistic regression model, combined with the hard disk life decay characteristics, the problems of unreasonable data acquisition and low prediction accuracy in server hard disk fault monitoring are solved, and intelligent and efficient fault management is achieved, reducing the risk of business interruption.

CN120429149AActive Publication Date: 2025-08-05EXANDS INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510489386.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-05
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

In the monitoring of server hard disk faults, the data acquisition frequency is fixed, resulting in waste of resources or missing key fault signals. The selection of key features depends on empirical judgment, and the lack of full utilization of business performance data and network traffic data, the prediction accuracy is low, and the lack of an intelligent decision-making mechanism leads to lag in fault response and increases the risk of business interruption.

Method used

By collecting hardware health data, business performance data and network traffic data, an adaptive sampling frequency control function is built, key features are selected, and failure probability prediction model is constructed through logistic regression, and combined with the hard disk life decay model, automated response actions are triggered.

Benefits of technology

It realizes dynamic adjustment of data acquisition frequency based on network traffic, improves fault prediction accuracy, reduces resource waste and troubleshooting time, reduces business interruption risk, and realizes intelligent and efficient network failure management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429149A_ABST
    Figure CN120429149A_ABST
Patent Text Reader

Abstract

The invention discloses a network fault management system and method based on multi-source data, and relates to the technical field of server hard disk fault monitoring, and the method specifically comprises the following steps: collecting hardware health data, service performance data, network flow data and historical fault records, and determining the collection frequency; key features of server hard disk faults are obtained, and data quantification processing is carried out on the key features; taking the preprocessed data as a reference, selecting key feature data values of a plurality of server hard disks, and constructing a server hard disk fault probability prediction model through logistic regression; obtaining real-time key feature data, calculating the fault occurrence probability of the server hard disk through the server hard disk fault probability prediction model, and comparing a preset grading range to obtain a fault grade; and according to the fault level, triggering an automatic response action. According to the invention, the troubleshooting time and the business interruption risk are effectively reduced, and more intelligent and efficient network fault management is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of server hard disk failure monitoring, in particular to a network failure management system and method based on multi-source data. Background Art

[0002] In modern data centers and high-concurrency business environments, server hard drives, as the core storage, endure high read and write loads over extended periods. These hard drives are susceptible to physical damage, environmental factors, and software anomalies, leading to performance degradation and even failure, which in turn impacts business continuity. While widely used SMART technology provides basic health monitoring data such as drive temperature, bad sector count, and I / O speed, and AIOps technology uses machine learning models for fault prediction, existing methods still have numerous shortcomings. For one thing, traditional data collection typically uses a fixed frequency and fails to dynamically adjust based on network traffic. This results in over-sampling during low loads, wasting resources, and under-sampling during high loads, missing critical fault signals. Furthermore, the selection of key features often relies on empirical judgment, underutilizing business performance and network traffic data, which impacts prediction accuracy. Furthermore, existing failure probability calculation methods fail to fully account for the nonlinear lifespan of hard drives. Simple linear regression or statistical thresholds ignore the distinct characteristics of drives during early failure, stable, and wear-out periods, leading to significant prediction errors. Furthermore, traditional operations and maintenance rely primarily on manual monitoring and threshold alerts, lacking intelligent decision-making mechanisms. This leads to delayed fault response and increases the risk of business interruption. Summary of the Invention

[0003] The purpose of the present invention is to provide a network fault management system and method based on multi-source data to solve the problems raised in the prior art.

[0004] To achieve the above objectives, the present invention provides the following technical solution: a network fault management method based on multi-source data, the network fault management method specifically comprising the following steps:

[0005] Step S100: Collect hardware health data, service performance data, network traffic data, and historical fault records, construct an adaptive sampling frequency control function, determine the collection frequency, and update the historical fault records;

[0006] Step S200: Acquire key features of a server hard disk failure and perform data quantification processing on the key features;

[0007] Step S300: Based on the pre-processed data, select key feature data values of several server hard disks and construct a server hard disk failure probability prediction model through logistic regression;

[0008] Step S400: Acquire real-time key feature data, calculate the failure probability of the server hard disk using the server hard disk failure probability prediction model, and compare it with a preset classification range to obtain a failure level;

[0009] Step S500: triggering an automated response action according to the fault level.

[0010] In step S100, hardware health data, service performance data, network traffic data, and historical fault records are collected, an adaptive sampling frequency control function is constructed, the collection frequency of different data is determined, and the historical fault records are updated. Specifically, the following steps are included:

[0011] Step S101: Collect hardware health data, service performance data, network traffic data, and historical fault records; the hardware health data includes server hard disk temperature, response time, read error rate, and sector abnormality ratio; the service performance data includes database query latency, page load time, and request timeout rate; the network traffic data includes disk IOPS and switch port bandwidth utilization; the historical fault records include server hard disk model, fault time, fault type, and maintenance record;

[0012] Step S102: construct an adaptive sampling frequency control function to determine the acquisition frequency of different data;

[0013] Step S103: Update the historical fault records irregularly.

[0014] In step S102, an adaptive sampling frequency control function is constructed to determine the acquisition frequency, specifically:

[0015] Construct an adaptive sampling frequency control function based on changes in network traffic;

[0016] The adaptive sampling frequency control function includes a basic sampling frequency term, a traffic load perception term, and a traffic mutation response term.

[0017]

[0018] Where, f represents the adaptive sampling frequency; f base Indicates the basic sampling frequency; G indicates real-time network traffic; G max and G min They represent the historical peak and valley values of network traffic respectively; ΔG represents the rate of change of real-time network traffic; τ represents the traffic load control parameter.

[0019] In step S200, key features of the server hard disk failure are obtained, specifically:

[0020] Step S201: Arrange the hardware health data, service performance data, network traffic data, and the accumulated working time of the server hard disk in descending order, and record them as sequences r1 and r2;

[0021] Step S202: Calculate the rank of each data point in its respective sequence, i.e., find the position of the data point in the sequence. For repeated data points, use the average rank method.

[0022] Step S203: Calculate the rank difference d, specifically: the rank of the corresponding data point in r1 minus the rank of the corresponding data point in r2;

[0023] Step S204: Calculate the correlation coefficient between different data and the accumulated working time of the server hard disk through the change relationship of the rank difference;

[0024] As a preferred embodiment, the specific calculation formula is:

[0025]

[0026] Where ρ represents the correlation coefficient between the data and the cumulative working time of the server hard disk; N represents the number of data, N is a positive integer; d i Represents the rank difference of the i-th data;

[0027] Step S205: determine the data with a value greater than the average value of the correlation coefficient as key features, and record them as w1, w2, ..., w n ; Among them, w1, w2, ..., w n Indicates the 1st, 2nd, ..., n key features, n is a positive integer.

[0028] The key features are subjected to data quantification processing, specifically:

[0029] Step S211: Based on the obtained key feature, obtain the historical maximum value w of the key feature. max and the historical minimum value w min ;

[0030] Step S212: performing data quantification processing on the acquired key feature data, specifically:

[0031] δ(w)=(ww min ) / (w max -w min ), where δ(w) represents the normalization function of the key feature w.

[0032] In step S300, based on the pre-processed data, several key feature data values of server hard disks are selected, and a server hard disk failure probability prediction model is constructed through logistic regression, specifically:

[0033] Step S301: Based on historical fault records and pre-processed data, select key feature data values of several server hard disks as training data;

[0034] Step S302: performing data quantization processing on the training data;

[0035] Step S303: constructing a decay function of the accumulated running time of the server hard disk, combining the training data after data quantization processing, and constructing a server hard disk failure probability prediction model through logistic regression.

[0036] Preferably, the characterization formula of the server hard disk failure probability prediction model is:

[0037]

[0038] Where p represents the probability of server hard disk failure; T represents the cumulative running time of the server hard disk; w j Represents the jth key feature, j is the key feature label, j is a positive integer, j∈[1,n]; ρ j represents the weight coefficient of the jth key feature, which is represented by the correlation coefficient between the data and the cumulative working time of the server hard disk; e represents a natural constant; λ(t) represents the decay function of the cumulative running time of the server hard disk; δ(w j ) represents the jth key feature w j The normalization function of .

[0039] Specifically, in step S303, the decay function of the accumulated running time of the server hard disk is constructed as follows:

[0040] A three-stage hybrid model is used to combine and obtain the decay function of the cumulative operating time of the server hard disk. The three-stage hybrid model includes an early failure period, a stable period, and a wear period.

[0041] Data simulation is performed for three different periods to obtain a decay function of the cumulative operating time of the server hard disk.

[0042] Preferably, the decay function of the accumulated running time of the server hard disk is expressed as follows:

[0043] λ(t)=α×e -βt +γ+σ×(max(0,t-t0)) θ ; Where λ(t) represents the decay function value of the accumulated running time of the server hard disk; t represents the accumulated running time of the server hard disk;

[0044] The fitting expression of the early failure period is characterized by an exponential decay α×e-βt ;

[0045] α and β represent the initial risk and decay rate of the early failure period, respectively; β is the improvement rate of early failure; the fitting expression of the stable period is characterized by a constant γ; γ represents the random failure rate (environmental factor) in the stable period; the fitting expression of the wear-out period is characterized by a power-law growth σ×(max(0,t-t0)) θ ; The σ represents the intensity coefficient of failure growth during the wear-out period, which is usually obtained based on the manufacturer's reliability report; the θ represents the acceleration index of the wear-out period risk; and the t0 is the threshold time for the start of the wear-out period.

[0046] In step S400, real-time key feature data is obtained, and the failure probability of the server hard disk is calculated using the server hard disk failure probability prediction model. The failure level is compared with the preset classification range to obtain the failure level, which is specifically:

[0047] Step S401: Acquire real-time key feature data;

[0048] Step S402: performing data quantization processing using the method in step S212;

[0049] Step S403: Calculating the probability of failure of the server hard disk using the server hard disk failure probability prediction model;

[0050] Step S404: Compare with the preset classification range to obtain the fault level; the classification range is set according to the rules of normal distribution, and the classification range is preset based on the sum of the mean value and the integer multiple standard deviation by calculating the mean value and the standard deviation of the key features.

[0051] In step S500, an automated response action is triggered according to the fault level, specifically:

[0052] According to the historical fault records, the fuzzy control table is preset;

[0053] The response strategy is automatically executed according to the fuzzy control table.

[0054] A network fault management system based on multi-source data, comprising a data acquisition module, a data processing module, a fault prediction module, an automated operation and maintenance module, and an early warning and visualization module;

[0055] The data acquisition module is used to collect hardware health data, business performance data and network traffic data; construct an adaptive sampling frequency control function to adjust the sampling frequency;

[0056] The data processing module is used to determine the key features through the change relationship of the rank difference; and perform data quantification processing based on the historical maximum and historical minimum values of the key features;

[0057] The fault prediction module is used to build a server hard disk failure probability prediction model through logistic regression, and predict the probability of server hard disk failure based on real-time key feature data;

[0058] The automated operation and maintenance module is used to compare the probability of failure with a preset classification range to obtain the failure level; and automatically execute the response strategy according to the preset fuzzy control table;

[0059] The early warning and visualization module is used to issue an early warning according to the fault level and provide a visual display of network traffic.

[0060] The output end of the data acquisition module is connected to the input end of the data processing module; the output end of the data processing module is connected to the input end of the fault prediction module; the output end of the fault prediction module is connected to the input end of the automated operation and maintenance module; the output end of the automated operation and maintenance module is connected to the input end of the visualization module.

[0061] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention optimizes the data collection strategy by integrating hardware health data, business performance data, network traffic data and historical fault records, and improves the prediction accuracy of server hard disk failures based on logistic regression and hard disk life decay model. This method adopts an adaptive sampling frequency control function to dynamically adjust the data collection frequency according to network traffic, ensuring that key fault signals are not missed while avoiding waste of resources. In addition, the rank correlation analysis method is used to screen key features, and combined with standardization processing, the robustness and generalization ability of the model are improved. The fault prediction model integrates the three-stage decay characteristics of the hard disk life, so that the prediction results are more in line with the actual operating laws of the hard disk, thereby improving the reliability of the early warning. Based on the prediction results, this method can also provide an automated response strategy to trigger corresponding maintenance measures according to the preset fault level, reduce troubleshooting time and the risk of business interruption, and achieve more intelligent and efficient network fault management. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 A schematic diagram of the steps of a network fault management method based on multi-source data according to the present invention;

[0063] Figure 2 This is a structural diagram of a network fault management system based on multi-source data according to the present invention. DETAILED DESCRIPTION

[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0065] Example: Figure 1-Figure 2 As shown, the present invention provides a technical solution, a network fault management method based on multi-source data, and the network fault management method specifically includes the following steps:

[0066] Step S100: Collect hardware health data, service performance data, network traffic data, and historical fault records, construct an adaptive sampling frequency control function, determine the collection frequency, and update the historical fault records;

[0067] Step S200: Acquire key features of a server hard disk failure and perform data quantification processing on the key features;

[0068] Step S300: Based on the pre-processed data, select key feature data values of several server hard disks and construct a server hard disk failure probability prediction model through logistic regression;

[0069] Step S400: Acquire real-time key feature data, calculate the failure probability of the server hard disk using the server hard disk failure probability prediction model, and compare it with a preset classification range to obtain a failure level;

[0070] Step S500: triggering an automated response action according to the fault level.

[0071] In step S100, hardware health data, service performance data, network traffic data, and historical fault records are collected, an adaptive sampling frequency control function is constructed, the collection frequency of different data is determined, and the historical fault records are updated. Specifically, the following steps are included:

[0072] Step S101: Collect hardware health data, service performance data, network traffic data, and historical fault records; the hardware health data includes server hard disk temperature, response time, read error rate, and sector abnormality ratio; the service performance data includes database query latency, page load time, and request timeout rate; the network traffic data includes disk IOPS and switch port bandwidth utilization; the historical fault records include server hard disk model, fault time, fault type, and maintenance record;

[0073] Step S102: construct an adaptive sampling frequency control function to determine the acquisition frequency of different data;

[0074] Step S103: Update the historical fault records irregularly.

[0075] In step S102, an adaptive sampling frequency control function is constructed to determine the acquisition frequency, specifically:

[0076] Construct an adaptive sampling frequency control function based on changes in network traffic;

[0077] The adaptive sampling frequency control function includes a basic sampling frequency term, a traffic load perception term, and a traffic mutation response term.

[0078]

[0079] Where, f represents the adaptive sampling frequency; f base Indicates the basic sampling frequency; G indicates real-time network traffic; G max and G min They represent the historical peak and valley values of network traffic respectively; ΔG represents the rate of change of real-time network traffic; τ represents the traffic load control parameter.

[0080] In step S200, key features of the server hard disk failure are obtained, specifically:

[0081] Step S201: Arrange the hardware health data, service performance data, network traffic data, and the accumulated working time of the server hard disk in descending order, and record them as sequences r1 and r2;

[0082] Step S202: Calculate the rank of each data point in its respective sequence, i.e., find the position of the data point in the sequence. For repeated data points, use the average rank method.

[0083] Step S203: Calculate the rank difference d, specifically: the rank of the corresponding data point in r1 minus the rank of the corresponding data point in r2;

[0084] Step S204: Calculate the correlation coefficient between different data and the accumulated working time of the server hard disk through the change relationship of the rank difference;

[0085] As a preferred embodiment, the specific calculation formula is:

[0086]

[0087] Where ρ represents the correlation coefficient between the data and the cumulative working time of the server hard disk; N represents the number of data, N is a positive integer; d i Represents the rank difference of the i-th data;

[0088] Step S205: determine the data with a value greater than the average value of the correlation coefficient as key features, and record them as w1, w2, ..., w n; Among them, w1, w2, ..., w n Indicates the 1st, 2nd, ..., n key features, n is a positive integer.

[0089] The key features are subjected to data quantification processing, specifically:

[0090] Step S211: Based on the obtained key feature, obtain the historical maximum value w of the key feature. max and the historical minimum value w min ;

[0091] Step S212: performing data quantification processing on the acquired key feature data, specifically:

[0092] δ(w)=(ww min ) / (w max -w min ), where δ(w) represents the normalization function of the key feature w.

[0093] In step S300, based on the pre-processed data, several key feature data values of server hard disks are selected, and a server hard disk failure probability prediction model is constructed through logistic regression, specifically:

[0094] Step S301: Based on historical fault records and pre-processed data, select key feature data values of several server hard disks as training data;

[0095] Step S302: performing data quantization processing on the training data;

[0096] Step S303: constructing a decay function of the accumulated running time of the server hard disk, combining the training data after data quantization processing, and constructing a server hard disk failure probability prediction model through logistic regression.

[0097] Preferably, the characterization formula of the server hard disk failure probability prediction model is:

[0098]

[0099] Where p represents the probability of server hard disk failure; T represents the cumulative operating time of the server hard disk; w j Represents the jth key feature, j is the key feature label, j is a positive integer, j∈[1,n]; ρ j represents the weight coefficient of the jth key feature, which is represented by the correlation coefficient between the data and the cumulative working time of the server hard disk; e represents a natural constant; λ(t) represents the decay function of the cumulative running time of the server hard disk; δ(w j ) represents the jth key feature wj The normalization function of .

[0100] Specifically, in step S303, the decay function of the accumulated running time of the server hard disk is constructed as follows:

[0101] A three-stage hybrid model is used to combine and obtain the decay function of the cumulative operating time of the server hard disk. The three-stage hybrid model includes an early failure period, a stable period, and a wear period.

[0102] Data simulation is performed for three different periods to obtain a decay function of the cumulative operating time of the server hard disk.

[0103] Preferably, the decay function of the accumulated running time of the server hard disk is expressed as follows:

[0104] λ(t)=α×e -βt +γ+σ×(max(0,t-t0)) θ ; Where λ(t) represents the decay function value of the accumulated running time of the server hard disk; t represents the accumulated running time of the server hard disk;

[0105] The fitting expression of the early failure period is characterized by an exponential decay α×e -βt ;

[0106] α and β represent the initial risk and decay rate of the early failure period, respectively;

[0107] α is the initial failure rate due to manufacturing defects in the early stages of shipment, obtained by fitting the manufacturer's MTBF data;

[0108] β is the improvement rate of early-stage failures, preferably obtained by regression training based on the first 1000 hours of failure data;

[0109] The fitting expression of the stable period is characterized by a constant γ;

[0110] The γ represents the random failure rate (environmental factor) during the stable period; preferably, based on historical failure records, the long-term average value of the historical failure rate is used as the value of γ;

[0111] The fitting expression of the wear period is characterized by a power law growth σ×(max(0,t-t0)) θ ;

[0112] The σ represents the intensity coefficient of failure growth during the wear-out period, which is usually obtained based on the manufacturer's reliability report;

[0113] The θ represents the acceleration index of the loss period risk, which is obtained by fitting the shape parameters of the Weibull distribution;

[0114] The t0 is the threshold time for the start of the wear-out period, which is usually the manufacturer's nominal life multiplied by 0.7.

[0115] In step S400, real-time key feature data is obtained, and the failure probability of the server hard disk is calculated using the server hard disk failure probability prediction model. The failure level is compared with the preset classification range to obtain the failure level, which is specifically:

[0116] Step S401: Acquire real-time key feature data;

[0117] Step S402: performing data quantization processing using the method in step S212;

[0118] Step S403: Calculating the probability of failure of the server hard disk using the server hard disk failure probability prediction model;

[0119] Step S404: Compare with the preset classification range to obtain the fault level; the classification range is set according to the rules of normal distribution, and the classification range is preset based on the sum of the mean value and the integer multiple standard deviation by calculating the mean value and the standard deviation of the key features.

[0120] In step S500, an automated response action is triggered according to the fault level, specifically:

[0121] According to the historical fault records, the fuzzy control table is preset;

[0122] The response strategy is automatically executed according to the fuzzy control table.

[0123] Example 1,

[0124] By integrating SMART into the hard drive controller, data monitoring functions are implemented to continuously track key operating parameters and status information of the hard drive. Using sensors and logic circuits inside the hard drive, dozens of different data indicators such as head seek time, motor speed, temperature, data transfer rate, sector error rate, etc. are collected in real time.

[0125] Adaptive Data Collection:

[0126] Hardware parameters:

[0127] Hard disk model: HDD-X***;

[0128] Cumulative operating time: 20,00 hours;

[0129] Network bandwidth: 1Gbps;

[0130] Real-time data collection: (data type-indicator name-real-time value):

[0131] Hardware health data - temperature - 50°C;

[0132] Response time - 8ms;

[0133] Read error rate - 0.15%;

[0134] Sector anomaly ratio -0.02%;

[0135] Business performance data - database query latency - 120ms;

[0136] Page load time – 2.1s;

[0137] Request timeout rate - 1.8%;

[0138] Network traffic data - disk;

[0139] IOPS-8500;

[0140] Port bandwidth utilization - 85%;

[0141] Basic sampling frequency: f base =10Hz;

[0142] Network traffic parameters:

[0143] G max =900Mbps (historical peak);

[0144] G min =200Mbp (historical valley value);

[0145] Real-time traffic G = 850 Mbps;

[0146] According to the continuously collected traffic, the traffic change rate ΔG = 50 Mbps / s is obtained;

[0147] The flow load control parameter τ is set to 0.5;

[0148] The adaptive sampling frequency f is approximately 14 Hz. (The sampling frequency is adjusted based on real-time network traffic).

[0149] Example 2,

[0150] Through the change relationship of rank difference, the correlation coefficient between different data and the cumulative working time of the server hard disk is calculated. Some results are shown in Table 1:

[0151] feature Rank difference Correlation coefficient temperature 12 0.82 Read error rate 18 0.76 Sector Abnormal Ratio 25 0.63 Database query latency 30 0.55 Disk IOPS 8 0.89

[0152] Table 1

[0153] Example 3,

[0154] The classification range is set according to the rules of normal distribution, and the classification range is preset by calculating the average value and standard deviation of the key characteristics and taking the sum of the average value and the integer multiple standard deviation as the basis.

[0155] Taking the response time as an example (approximate value), by observing that the data points are approximately distributed in a straight line, it conforms to normality;

[0156] Average value μ = 50 ms; standard deviation σ' = 10 ms; (L represents the range identifier of the response time):

[0157] grade Scope Definition Threshold calculation Physical meaning normal L≤μ+1σ' ≤60ms Within the expected performance range focus on m+1s' <L≤μ+2σ’ 60<L≤70ms Potential performance bottlenecks warn m+2s' <L≤μ+3σ’ 70<L≤80ms Significantly affects user experience serious L>μ+3σ' >80ms The system is on the verge of failure or is unavailable

[0158] Table 2

[0159] The same analysis is performed on different key features, and the results are substituted into the probability model of server hard disk failure to calculate the classification range of different running times.

[0160] A network fault management system based on multi-source data, comprising a data acquisition module, a data processing module, a fault prediction module, an automated operation and maintenance module, and an early warning and visualization module;

[0161] The data acquisition module is used to collect hardware health data, business performance data and network traffic data; construct an adaptive sampling frequency control function to adjust the sampling frequency;

[0162] The data processing module is used to determine the key features through the change relationship of the rank difference; and perform data quantification processing based on the historical maximum and historical minimum values of the key features;

[0163] The fault prediction module is used to build a server hard disk failure probability prediction model through logistic regression, and predict the probability of server hard disk failure based on real-time key feature data;

[0164] The automated operation and maintenance module is used to compare the probability of failure with a preset classification range to obtain the failure level; and automatically execute the response strategy according to the preset fuzzy control table;

[0165] The early warning and visualization module is used to issue an early warning according to the fault level and provide a visual display of network traffic.

[0166] The output end of the data acquisition module is connected to the input end of the data processing module; the output end of the data processing module is connected to the input end of the fault prediction module; the output end of the fault prediction module is connected to the input end of the automated operation and maintenance module; the output end of the automated operation and maintenance module is connected to the input end of the visualization module.

[0167] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A network fault management method based on multi-source data, characterized by: The network fault management method specifically includes the following steps: Step S100: Collect hardware health data, service performance data, network traffic data, and historical fault records, construct an adaptive sampling frequency control function, determine the collection frequency, and update the historical fault records; Step S200: Acquire key features of a server hard disk failure and perform data quantification processing on the key features; Step S300: Based on the pre-processed data, select key feature data values of several server hard disks and construct a server hard disk failure probability prediction model through logistic regression; Step S400: Acquire real-time key feature data, calculate the failure probability of the server hard disk using the server hard disk failure probability prediction model, and compare it with a preset classification range to obtain a failure level; Step S500: triggering an automated response action according to the fault level.

2. The network fault management method based on multi-source data according to claim 1, characterized in that: In step S100, hardware health data, service performance data, network traffic data, and historical fault records are collected, an adaptive sampling frequency control function is constructed, the collection frequency of different data is determined, and the historical fault records are updated. Specifically, the following steps are included: Step S101: Collect hardware health data, service performance data, network traffic data, and historical fault records; the hardware health data includes server hard disk temperature, response time, read error rate, and sector abnormality ratio; the service performance data includes database query latency, page load time, and request timeout rate; the network traffic data includes disk IOPS and switch port bandwidth utilization; the historical fault records include server hard disk model, fault time, fault type, and maintenance record; Step S102: construct an adaptive sampling frequency control function to determine the acquisition frequency of different data; Step S103: Update the historical fault records irregularly.

3. The network fault management method based on multi-source data according to claim 2, characterized in that: In step S102, an adaptive sampling frequency control function is constructed to determine the acquisition frequency, specifically: Construct an adaptive sampling frequency control function based on changes in network traffic; The adaptive sampling frequency control function includes a basic sampling frequency term, a traffic load perception term, and a traffic mutation response term.

4. The network fault management method based on multi-source data according to claim 3, characterized in that: In step S200, key features of the server hard disk failure are obtained, specifically: Step S201: Arrange the hardware health data, service performance data, network traffic data, and the accumulated working time of the server hard disk in descending order, and record them as sequences r1 and r2; Step S202: Calculate the rank of each data point in its respective sequence, i.e., find the position of the data point in the sequence. For repeated data points, use the average rank method. Step S203: Calculate the rank difference d, specifically: the rank of the corresponding data point in r1 minus the rank of the corresponding data point in r2; Step S204: Calculate the correlation coefficient between different data and the accumulated working time of the server hard disk through the change relationship of the rank difference; Step S205: determine the data with a value greater than the average value of the correlation coefficient as key features, and record them as w1, w2, ..., w n ; Among them, w1, w2, ..., w n Indicates the 1st, 2nd, ..., n key features, n is a positive integer.

5. The network fault management method based on multi-source data according to claim 4, characterized in that: The key features are subjected to data quantification processing, specifically: Step S211: Based on the obtained key feature, obtain the historical maximum value w of the key feature. max and the historical minimum value w min ; Step S212: performing data quantification processing on the acquired key feature data, specifically: δ(w)=(ww min ) / (w max -w min ), where δ(w) represents the normalization function of the key feature w.

6. The network fault management method based on multi-source data according to claim 5, characterized in that: In step S300, based on the pre-processed data, several key feature data values of server hard disks are selected, and a server hard disk failure probability prediction model is constructed through logistic regression, specifically: Step S301: Based on historical fault records and pre-processed data, select key feature data values of several server hard disks as training data; Step S302: performing data quantization processing on the training data; Step S303: constructing a decay function of the accumulated running time of the server hard disk, combining the training data after data quantization processing, and constructing a server hard disk failure probability prediction model through logistic regression.

7. The network fault management method based on multi-source data according to claim 6, characterized in that: Specifically, in step S303, the decay function of the accumulated running time of the server hard disk is constructed as follows: A three-stage hybrid model is used to combine and obtain the decay function of the cumulative operating time of the server hard disk. The three-stage hybrid model includes an early failure period, a stable period, and a wear period. Data simulation is performed for three different periods to obtain a decay function of the cumulative operating time of the server hard disk.

8. The network fault management method based on multi-source data according to claim 7, characterized in that: In step S400, real-time key feature data is obtained, and the failure probability of the server hard disk is calculated using the server hard disk failure probability prediction model. The failure level is compared with the preset classification range to obtain the failure level, which is specifically: Step S401: Acquire real-time key feature data; Step S402: performing data quantization processing using the method in step S212; Step S403: Calculating the probability of failure of the server hard disk using the server hard disk failure probability prediction model; Step S404: Compare with the preset classification range to obtain the fault level; the classification range is set according to the rules of normal distribution, and the classification range is preset based on the sum of the mean value and the integer multiple standard deviation by calculating the mean value and the standard deviation of the key features.

9. The network fault management method based on multi-source data according to claim 7, characterized in that: In step S500, an automated response action is triggered according to the fault level, specifically: According to the historical fault records, the fuzzy control table is preset; The response strategy is automatically executed according to the fuzzy control table.

10. A network fault management system based on multi-source data, applying the network fault management method based on multi-source data according to any one of claims 1 to 9, characterized in that: The network fault management system includes a data acquisition module, a data processing module, a fault prediction module, an automated operation and maintenance module, and an early warning and visualization module; The data acquisition module is used to collect hardware health data, business performance data and network traffic data; construct an adaptive sampling frequency control function to adjust the sampling frequency; The data processing module is used to determine the key features through the change relationship of the rank difference; and perform data quantification processing based on the historical maximum and historical minimum values of the key features; The fault prediction module is used to build a server hard disk failure probability prediction model through logistic regression, and predict the probability of server hard disk failure based on real-time key feature data; The automated operation and maintenance module is used to compare the probability of failure with a preset classification range to obtain the failure level; and automatically execute the response strategy according to the preset fuzzy control table; The early warning and visualization module is used to issue an early warning according to the fault level and provide a visual display of network traffic; The output end of the data acquisition module is connected to the input end of the data processing module; the output end of the data processing module is connected to the input end of the fault prediction module; the output end of the fault prediction module is connected to the input end of the automated operation and maintenance module; the output end of the automated operation and maintenance module is connected to the input end of the visualization module.

Citation Information

Patent Citations

  • Disk fault prediction method based on SMART and performance logs

    CN111581072A

  • Disk fault prediction method and device, equipment and medium

    CN114610547A

  • Hard disk fault prediction method, system, equipment and medium

    CN115729761A

  • Model generation method and device, disk fault prediction method and device, equipment and medium

    CN116775437A

  • Regional air conditioner load prediction method and system based on large time sequence model

    CN119129838A

Cited By

  • Network equipment fault intelligent detection method and system based on data transmission

    CN120934995A