Fault root cause analysis method and system and computer readable storage medium
By generating multi-dimensional feature index and deep neural network model, the accuracy of communication equipment failure root cause positioning is solved, efficient and accurate fault detection and prediction are achieved, and troubleshooting efficiency and system stability are improved.
Patent Information
- Application Number
- CN202510312198.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-29
AI Technical Summary
The existing technology is difficult to quickly and accurately locate the root cause of communication equipment failures, especially when facing complex and diversified operating states and large-scale data, traditional methods lack accurate assessment of the operating state of the equipment, and a single-dimensional analysis cannot fully cover the overall operating state of the equipment, resulting in misdiagnosis and missed diagnosis.
By generating signal health index, data transmission performance index and hardware operation status index, combined with deep neural network model, a fault analysis model is established, the mean and standard deviation of the feature index are used to define the normal operating range, and the root cause positioning rules are established to achieve accurate fault prediction and positioning of communication equipment.
It realizes accurate detection and root cause positioning of communication equipment failures, improves the efficiency and accuracy of troubleshooting, can quickly identify the source of failures and take targeted measures to reduce the equipment failure rate, improve system stability and operation and maintenance efficiency.
Smart Images

Figure CN120386655A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of communication device fault analysis, and particularly to a method and system for fault root cause analysis and a computer-readable storage medium. Background Art
[0002] The method for fault root cause analysis is an important means to ensure the stable operation of complex systems, and is widely used in fields such as communication devices, industrial manufacturing, and energy systems. Traditional fault analysis methods mostly rely on expert experience and rule bases. This approach is relatively effective in dealing with simple problems, but in the face of the large-scale data, diverse operating states, and complex fault modes of modern devices, traditional methods often struggle to quickly and accurately locate the root cause of faults. With the development of big data technology and artificial intelligence, data-driven fault diagnosis and root cause analysis technologies have gradually emerged. Through real-time collection, modeling, and intelligent analysis of device operation data, more efficient and accurate fault prediction and root cause location can be achieved.
[0003] With the deepening of industrial digitization and intelligence, the fault root cause analysis technology will develop towards automation, intelligence, and precision. Combining advanced technologies such as the Internet of Things, big data analysis, and deep learning, future fault analysis methods will be able to process massive data more efficiently, accurately predict potential faults, and quickly locate the root cause. In addition, with the popularization of edge computing technology, fault analysis can be completed in real time at the device end, further improving the response speed, providing strong technical support for critical mission systems and high-reliability application scenarios (such as 5G communication, intelligent manufacturing, and autonomous driving, etc.), and having broad application prospects.
[0004] In the prior art, device fault analysis usually relies on traditional experience or rules, lacking an accurate assessment of the device operation state. Especially when facing device faults, it is difficult to quickly and accurately locate the root cause. Moreover, the prior art often only analyzes a single performance index or fault signal, separating the device fault analysis from different types of operating factors and ignoring their mutual influence, unable to comprehensively evaluate the overall operation state of the device. And currently, after detecting a fault anomaly, the existing technology often has difficulty in refining and judging the specific source of the fault, and can only determine the fault type through rough estimation, unable to accurately locate the fault type and location.
[0005] Therefore, it is necessary to provide a method and system for fault root cause analysis and a computer-readable storage medium to solve the above problems.
[0006] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0007] The purpose of the present invention is to provide a method, a system and a computer-readable storage medium for analyzing the root cause of a fault, so as to solve the problems raised in the above-mentioned background technology.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] A method for analyzing the root cause of a fault, the specific steps include:
[0010] Step 1: Obtain the operation data and corresponding operation labels of multiple communication devices. The operation data includes the signal strength, bit error rate, data transmission rate, signal-to-noise ratio, temperature, packet loss rate, and power supply voltage of the device, and the operation labels include the normal state and the fault state. The fault state includes;
[0011] Step 2: Generate a characteristic index that can represent the operation state of the communication device according to the obtained operation data of the communication device. The characteristic index includes a signal health index, a data transmission performance index, and a hardware operation state index, and based on the characteristic index in the normal state of the operation label, calculate the mean and standard deviation of each type of characteristic index, and establish the normal operation range of the communication device according to the calculated mean and standard deviation;
[0012] Step 3: Establish a fault analysis model, use each type of characteristic index as the input of the model, and use the corresponding fault prediction value of each type of characteristic index as the output of the model. Use normal data to train the device fault prediction analysis model, obtain the operation data when historical communication devices fail, and input the three types of characteristic indexes in the historical fault operation data into the trained model respectively to obtain three corresponding fault prediction values. The fault prediction values include a signal link fault prediction value, a data transmission fault prediction value, and a hardware operation fault prediction value;
[0013] Step 4: Based on the fault prediction value output by the fault analysis model, combine the normal operation range of the communication device to determine whether the communication device has a fault anomaly;
[0014] Step 5: Set up a root cause location rule. After detecting a fault anomaly of the communication device, determine the specific source of the fault according to the category of the abnormal characteristic index.
[0015] Further, the method for generating a characteristic index that can represent the operation state of the communication device, the characteristic index includes a signal health index, a data transmission performance index, and a hardware operation state index, is as follows:
[0016] The signal health index is used to measure the overall quality of the device's signal reception. The signal health index is calculated based on the signal strength, signal-to-noise ratio, and bit error rate in the operation data. The formula is as follows:
[0017]
[0018] Among them, represents the signal health index, is the signal strength, representing the signal reception power, is the signal-to-noise ratio, representing the clarity of the signal, is the bit error rate, representing the reliability of signal transmission;
[0019] The data transmission performance index is used to evaluate the efficiency and reliability of the device in data transmission. Based on the transmission rate, bit error rate, and packet loss rate in the operation data, the data transmission performance index is calculated. The formula is as follows:
[0020]
[0021] Among them, represents the data transmission performance index, is the data transmission rate, representing the amount of data transmitted per unit time, is the packet loss rate, representing the proportion of lost data packets;
[0022] The hardware operation status index is used to evaluate the health status of the device hardware. Based on the device temperature and power supply voltage in the operation data, the hardware operation status index is calculated. The formula is as follows:
[0023]
[0024] Among them, represents the hardware operation status index, , are the operating value and rated value of the power supply voltage respectively, , are the working temperature and normal operating temperature of the communication device respectively.
[0025] Furthermore, based on the calculated characteristic indices, calculate the mean and standard deviation of each type of characteristic index, and establish the normal operating range of the communication device according to the calculated mean and standard deviation. The formula is as follows:
[0026]
[0027] Among them, represents the mean of the th type of characteristic index data, is the total number of the th type of characteristic index data, is the th th characteristic index data in the th type of characteristic index data, is the index of the number of the ; When it is, it means that the mean value of the signal health index is calculated. When it is, it means that the mean value of the data transmission performance index is calculated. When it is, it means that the mean value of the hardware operating status index is calculated;
[0028]
[0029] Among them, represents the standard deviation of the characteristic index data of the When it is, it means that the standard deviation of the signal health index is calculated. When it is, it means that the standard deviation of the data transmission performance index is calculated. When it is, it means that the standard deviation of the hardware operating status index is calculated;
[0030] The normal range calculates an interval based on the mean and standard deviation, and the formula used is:
[0031]
[0032] Among them, represents the normal operating range of the characteristic index data of the is the normal adjustment parameter, and its value is 2 or 3. When , it means that the normal operating range of the characteristic index data of the class includes the characteristic index. When , it means that the normal operating range of the characteristic index data of the class includes the characteristic index. When it is, it means that the normal operating range of the signal health index is calculated. When it is, it means that the normal operating range of the data transmission performance index is calculated. When it is, it means that the normal operating range of the hardware operating status index is calculated.
[0033] Furthermore, a fault analysis model is established. The three types of characteristic indexes in the historical fault operation data are respectively input into the trained model to obtain three corresponding fault prediction values. The method used is:
[0034] The fault analysis model is established using a deep neural network structure. The deep neural network structure includes an input layer, a hidden layer, and an output layer. The input layer is used to receive each type of characteristic index, the hidden layer is used to process the data of each type of characteristic index, and the output layer is used to output the probability prediction values of the three types of faults;
[0035] Obtain the data of the historical communication device when it is operating normally, and based on expert scoring, evaluate the fault labels corresponding to the three types of characteristic indices. Use the mean square error as the loss function for model training, and use the backpropagation algorithm to minimize the loss function to train the fault analysis model. Obtain the data when the historical communication device fails, generate three characteristic indices, and input them into the trained model to obtain three fault prediction values, including the signal link fault prediction value, the data transmission fault prediction value, and the hardware operation fault prediction value.
[0036] Further, the logical formula for judging whether the communication device has a fault anomaly is:
[0037]
[0038] Among them, When, it means that the fault prediction value is not within the normal operation range, and the communication device has a fault anomaly; When, it means that the fault prediction value is within the normal operation range, and the communication device has no fault anomaly; is the fault prediction value. When When, it means the signal link fault prediction value, When, it means the data transmission fault prediction value, When, it means the hardware operation fault prediction value.
[0039] Further, establish a root cause location rule to determine the specific source of the fault according to the category of the abnormal characteristic index. The method is as follows:
[0040] After detecting an anomaly, judge the root cause of the fault through the characteristic index. The root cause location rule specifically includes that when the signal link fault prediction value is not within the normal operation range of the signal health index, the communication device fault causes the signal health index to be abnormal. The fault reasons are wireless signal interference or weak signal coverage, antenna or signal link hardware damage, and excessive environmental noise resulting in an increase in the bit error rate. At this time, the antenna connection status of the communication device and the network signal environment should be checked; when the data transmission fault prediction value is not within the normal operation range of the data transmission performance index, the communication device fault causes the data transmission performance index to be abnormal. The fault reasons are data interface or relay device failure, insufficient data transmission bandwidth or network congestion. At this time, check the status of the network interface device and whether there is a bottleneck in the network transmission path; when the hardware operation fault prediction value is not within the normal operation range of the hardware operation status index, the communication device fault causes the hardware operation status index to be abnormal. The fault reasons are overheating of the communication device, increased power supply voltage fluctuation, and abnormal heat dissipation system or power supply module. At this time, the hardware temperature monitoring system or the power supply module should be checked whether it is working stably.
[0041] The present invention also provides a fault root cause analysis system, which is used to execute the above-mentioned fault root cause analysis method, including:
[0042] An operation data acquisition module, which is used to obtain the operation data of multiple communication devices and the corresponding operation labels. The operation data includes the signal strength, bit error rate, data transmission rate, signal-to-noise ratio, temperature, packet loss rate, and power supply voltage of the device. The operation labels include normal state and fault state, and the fault state includes;
[0043] A feature index generation and normal range setting module, which is used to generate a feature index that can characterize the operation state of the communication device according to the obtained operation data of the communication device. The feature index includes a signal health index, a data transmission performance index, and a hardware operation state index, and based on the feature index in the normal state of the operation label, calculate the mean and standard deviation of each type of feature index, and establish the normal operation range of the communication device according to the calculated mean and standard deviation;
[0044] A fault prediction model construction module, which is used to establish a fault analysis model, with each type of feature index as the input of the model and the corresponding fault prediction value of each type of feature index as the output of the model. Use normal data to train the device fault prediction analysis model, obtain the operation data when historical communication devices fail, and input the three types of feature indexes in the historical fault operation data into the trained model respectively to obtain three corresponding fault prediction values. The fault prediction values include a signal link fault prediction value, a data transmission fault prediction value, and a hardware operation fault prediction value;
[0045] A fault judgment module, which judges whether the communication device has a fault anomaly based on the fault prediction value output by the fault analysis model and in combination with the normal operation range of the communication device;
[0046] A fault root cause location module, which is used to set up a root cause location rule, and after detecting a fault anomaly of the communication device, determine the specific source of the fault according to the category of the abnormal feature index.
[0047] The present invention also provides a computer-readable storage medium for fault root cause analysis. A computer program is stored on the storage medium, and when the computer program is executed by a processor, it is used to implement the above-mentioned fault root cause analysis method.
[0048] Compared with the prior art, the beneficial effects of the present invention are:
[0049] By introducing a deep neural network model and a feature index generation method, the present invention can analyze the operating state of a device in real time and accurately, generate more accurate fault prediction values, and solve the problem in the prior art of lacking accurate evaluation of the device operating state, especially the difficulty in quickly and accurately locating the root cause in the face of device failures;
[0050] In addition, by combining multi-dimensional feature indices such as signal health index, data transmission performance index, and hardware operating state index, the present invention comprehensively covers different operating states of communication devices, which enables fault diagnosis to consider more potential fault sources and avoids misdiagnosis and missed diagnosis caused by single-dimensional analysis in the prior art; moreover, the present invention also sets up specific root cause location rules, accurately determines the fault source according to the abnormal conditions of different feature indices, and takes corresponding fault troubleshooting measures according to different fault types, thereby greatly improving the efficiency and accuracy of fault troubleshooting;
[0051] By calculating the mean and standard deviation of feature indices, integrating multi-dimensional feature indices, and introducing a deep neural network structure, the present invention solves the problems of traditional fault analysis methods lacking dynamic adaptability, ignoring the comprehensive influence of multiple factors, and insufficient fault prediction accuracy, and realizes more flexible, comprehensive, and accurate fault detection and prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic diagram of the overall method flow of the present invention.
[0053] Figure 2 It is a schematic diagram of the system module flow of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0054] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments.
[0055] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs. The "first", "second", and similar terms used in the present invention do not indicate any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before this term cover the elements or items listed after this term and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0056] Example:
[0057] Please refer to Figure 1 , the present invention provides a technical solution:
[0058] A method for analyzing the root cause of faults, the specific steps include:
[0059] Step 1: Obtain the operation data and corresponding operation labels of multiple communication devices. The operation data includes the signal strength, bit error rate, data transmission rate, signal-to-noise ratio, temperature, packet loss rate, and power supply voltage of the device. The operation labels include normal state and fault state, and the fault state includes;
[0060] Step 2: Generate a characteristic index that can characterize the operation state of the communication device according to the obtained operation data of the communication device. The characteristic index includes a signal health index, a data transmission performance index, and a hardware operation state index. Based on the characteristic index in the normal state of the operation label, calculate the mean and standard deviation of each type of characteristic index, and establish the normal operation range of the communication device according to the calculated mean and standard deviation;
[0061] Step 3: Establish a fault analysis model, use each type of characteristic index as the input of the model, and use the corresponding fault prediction value of each type of characteristic index as the output of the model. Use normal data to train the device fault prediction analysis model, obtain the operation data when historical communication devices fail, and input the three types of characteristic indexes in the historical failed operation data into the trained model respectively to obtain three corresponding fault prediction values. The fault prediction values include a signal link fault prediction value, a data transmission fault prediction value, and a hardware operation fault prediction value;
[0062] Step 4: Based on the fault prediction value output by the fault analysis model, combine the normal operation range of the communication device to determine whether the communication device has a fault anomaly;
[0063] Step 5: Set up a root cause location rule. After detecting a fault anomaly of the communication device, determine the specific source of the fault according to the category of the abnormal characteristic index.
[0064] It should be noted that by defining the signal health index, the data transmission performance index, and the hardware operation state index, the operation state of the communication device is accurately quantified and characterized, providing a scientific basis and data support for subsequent fault analysis and root cause location. Through the comprehensive evaluation of key operation parameters such as signal strength, signal-to-noise ratio, bit error rate, data transmission rate, packet loss rate, device temperature, and power supply voltage, these characteristic indexes can comprehensively reflect the health status of the device in terms of signal reception, data transmission, and hardware operation, thus laying a foundation for achieving efficient and accurate fault prediction and location.
[0065] Therefore, it is necessary to generate characteristic indices that can characterize the operating state of the communication device. The characteristic indices include a signal health index, a data transmission performance index, and a hardware operating state index. The method is as follows:
[0066] The signal health index is used to measure the overall quality of the device's signal reception. Based on the signal strength, signal-to-noise ratio, and bit error rate in the operating data of the communication device, the signal health index is calculated. The formula is as follows:
[0067]
[0068] Among them, represents the signal health index, is the signal strength, indicating the signal reception power, is the signal-to-noise ratio, indicating the clarity of the signal, is the bit error rate, indicating the reliability of signal transmission; in the above formula, when increases, increases accordingly, meaning that the signal reception quality is better and the reliability of the communication link is higher; when increases, also increases, indicating that the signal clarity is higher and the communication quality is better; the smaller the better, decreases, will increase, meaning that the reliability of signal transmission increases and the errors in the communication process decrease. Therefore, the size comprehensively reflects the signal strength, clarity, and reliability. The larger the value, the better the overall signal quality.
[0069] The data transmission performance index is used to evaluate the efficiency and reliability of the device in data transmission. It is evaluated using the transmission rate, bit error rate, and packet loss rate of the communication device. The formula is as follows:
[0070]
[0071] Among them, represents the data transmission performance index, is the data transmission rate, indicating the amount of data transmitted per unit time, is the packet loss rate, indicating the proportion of lost data packets; in the above formula, when increases, the value increases, indicating higher data transmission performance. When increases, the value will decrease, indicating a decline in data transmission performance, indicating fewer errors in the device's data transmission process and higher transmission reliability, thus improving the value. When increases, it will also cause A decrease indicates a deterioration in data transmission performance, meaning fewer errors during the device's data transmission process, higher transmission reliability, fewer lost data packets, and better transmission integrity.
[0072] The hardware operation status index is used to evaluate the health status of the device's hardware and is related to the device temperature and power supply voltage of the communication device. The formula is as follows:
[0073]
[0074] Where, represents the hardware operation status index, and are the operating value and rated value of the power supply voltage respectively, and are the working temperature and normal operating temperature of the communication device respectively; in the above formula, when is close to , will become smaller, indicating that the device's power supply voltage is in a healthy state; when deviates from , will increase, indicating that the device's voltage operation status is abnormal; and are the same. When the two values are close, will become smaller, indicating that the device temperature is in a healthy state. When the two values deviate greatly, will increase, indicating that the device temperature operation status is abnormal. Therefore, the smaller the value of, the closer the hardware operation status is to normal and the better the health condition; the larger the value of, the more the hardware status deviates from normal and the worse the health condition.
[0075] It should be noted that statistical means provide an objective and quantitative evaluation standard for the device status, which can effectively define the boundary of "normal operation" and capture abnormal states. When or , it covers the data of or respectively, ensuring that most of the characteristic index data is within the normal range, making the evaluation both reliable and adaptable. The reason for this setting is that it can reasonably tolerate a small amount of random fluctuations and timely identify potential problems, helping the operation and maintenance personnel to warn of device failures and formulate precise maintenance plans, thereby improving the operation safety and efficiency of the communication device.
[0076] Therefore, based on the characteristic indexes generated by calculation, it is necessary to calculate the mean and standard deviation of each type of characteristic index, and establish the normal operation range of the communication device according to the calculated mean and standard deviation. The formula is as follows:
[0077]
[0078] in, Indicates the The mean of the class characteristic index data, For the The total number of class feature index data, For the Class characteristic index data Characteristic index data, For the Index of the number of class feature indices, ; When , it means that the mean value of the signal health index is calculated. When , it means that the data transmission performance index mean is calculated. When , it means that the mean value of the hardware operation status index is calculated;
[0079]
[0080] in, Indicates the The standard deviation of the class characteristic index data, When , it means that the standard deviation of the signal health index is calculated. When , it means that the standard deviation of the data transmission performance index is calculated. When , it means that the standard deviation of the hardware operation status index is calculated;
[0081] The normal range is calculated as an interval based on the mean and standard deviation, according to the formula:
[0082]
[0083] in, Indicates the Normal operating range of class characteristic index data, is the normal adjustment parameter, which takes the value of 2 or 3. , indicating the The normal operating range of the class characteristic index data includes The characteristic index of , indicating the The normal operating range of the class characteristic index data includes The characteristic index of When , it means that the normal operating range of the signal health index is calculated. When , it means that the data transmission performance index is calculated within the normal operating range. When , it indicates that the hardware operating status index is calculated within the normal operating range.
[0084] It should be noted that by inputting three types of characteristic indices in the historical data, including the signal health index, data transmission performance index, and hardware operation status index, the model can learn the complex non-linear relationship between normal operation and fault states, and generate probability prediction values for each type of fault, thus providing a scientific basis for the fault diagnosis of the device. The reason for choosing a deep neural network lies in its powerful feature extraction and pattern recognition capabilities, which can capture key features from high-dimensional data. The mean squared error is used as the loss function, and the model is optimized through the backpropagation algorithm to ensure the accuracy and robustness of model training. Through this setting, the fault analysis model can not only evaluate the device status in real time, but also early warn of potential problems, helping the operation and maintenance personnel to take targeted measures, reduce the device failure rate, and improve the operation stability and economic benefits of the system.
[0085] Therefore, it is necessary to establish a fault analysis model. Input the three types of characteristic indices in the historical fault operation data into the trained model respectively to obtain three corresponding fault prediction values. The method is as follows:
[0086] The fault analysis model is established using a deep neural network structure. The deep neural network structure includes an input layer, a hidden layer, and an output layer. The input layer is used to receive each type of characteristic index, the hidden layer is used to process the data of each type of characteristic index, and the output layer is used to output the probability prediction values of the three types of faults;
[0087] Obtain the data of the historical communication device in normal operation, and based on expert scoring, evaluate the fault labels corresponding to the three types of characteristic indices. Use the mean squared error as the loss function for model training, and use the backpropagation algorithm to minimize the loss function. Train the fault analysis model. Obtain the data when the historical communication device fails, generate three characteristic indices, input them into the trained model, and obtain three fault prediction values, including the signal link fault prediction value, the data transmission fault prediction value, and the hardware operation fault prediction value.
[0088] It should be noted that judging whether the communication device has a fault anomaly through a logical formula is an important part of realizing intelligent monitoring and operation and maintenance management. The core idea of this formula is to compare the fault prediction value whether it is within the normal operation range If it is not within the range, it is determined that the device has a fault anomaly; if it is within the range, it is determined that the device is operating normally. The importance of this method lies in that through quantitative logical judgment, it can quickly and clearly identify whether the device is abnormal, avoiding the error of subjective judgment, and thus improving the fault detection efficiency.
[0089] Therefore, it is necessary to judge whether the communication device has a fault anomaly. The logical formula is as follows:
[0090]
[0091] Among them, when it is the case, it indicates that the fault prediction value is not within the normal operation range, and there is a fault anomaly in the communication device; when it is the case, it indicates that the fault prediction value is within the normal operation range, and there is no fault anomaly in the communication device; is the fault prediction value. When it is the case, it indicates the fault prediction value of the signal link, when it is the case, it indicates the fault prediction value of data transmission, when it is the case, it indicates the fault prediction value of hardware operation.
[0092] It should be noted that establishing the root cause location rule and quickly determining the specific fault source of the communication device through the category of the abnormal feature index is the key means to improve the fault diagnosis accuracy and operation and maintenance efficiency. The importance of this rule lies in that it can accurately associate the abnormalities of the signal health index, data transmission performance index, and hardware operation status index with specific fault types, such as signal link problems, data transmission bottlenecks, or hardware operation faults, so as to achieve efficient and accurate fault location. The reason for its setting is that the rule logic is simple and targeted, covering most common fault scenarios. Combining with the device operation principle and industry experience, and integrating with the automated detection technology, it can quickly lock the root cause, reduce the maintenance cost, ensure the system stability, and provide data support for optimizing the device design and operation and maintenance strategy.
[0093] Therefore, it is necessary to establish the root cause location rule, and determine the specific source of the fault according to the category of the abnormal feature index. The method is as follows:
[0094] After detecting the abnormality, judge the root cause of the fault through the feature index. The root cause location rule specifically includes that when there is a signal link fault prediction value not within the normal operation range of the signal health index, the communication device fault causes the signal health index to be abnormal. The fault reasons are wireless signal interference or weak signal coverage, antenna or signal link hardware damage, and excessive environmental noise resulting in an increase in the bit error rate. At this time, the antenna connection status of the communication device and the network signal environment should be checked; when there is a data transmission fault prediction value not within the normal operation range of the data transmission performance index, the communication device fault causes the data transmission performance index to be abnormal. The fault reasons are data interface or relay device faults, insufficient data transmission bandwidth, or network congestion. At this time, check the status of the network interface device and whether there is a bottleneck in the network transmission path; when there is a hardware operation fault prediction value not within the normal operation range of the hardware operation status index, the communication device fault causes the hardware operation status index to be abnormal. The fault reasons are overheating of the communication device, increased power supply voltage fluctuation, and abnormal heat dissipation system or power supply module. At this time, the hardware temperature monitoring system or the power supply module should be checked whether it is working stably.
[0095] Please refer to Figure 2 , the present invention also provides a fault root cause analysis system, which is used to execute the above-mentioned fault root cause analysis method, including:
[0096] An operation data acquisition module, which is used to acquire the operation data and corresponding operation labels of multiple communication devices. The operation data includes signal strength, bit error rate, data transmission rate, signal-to-noise ratio, temperature, packet loss rate, and power supply voltage of the device. The operation labels include normal state and fault state, and the fault state includes;
[0097] A feature index generation and normal range setting module, which is used to generate feature indexes that can represent the operation state of communication devices according to the acquired operation data of communication devices. The feature indexes include signal health index, data transmission performance index, and hardware operation state index, and calculate the mean and standard deviation of each type of feature index based on the feature indexes in the normal state of the operation label. Establish the normal operation range of communication devices according to the calculated mean and standard deviation;
[0098] A fault prediction model construction module, which is used to establish a fault analysis model, with each type of feature index as the input of the model and the corresponding fault prediction value of each type of feature index as the output of the model. Use normal data to train the device fault prediction analysis model, obtain the operation data when historical communication devices fail, and input the three types of feature indexes in the historical failed operation data into the trained model respectively to obtain three corresponding fault prediction values. The fault prediction values include signal link fault prediction value, data transmission fault prediction value, and hardware operation fault prediction value;
[0099] A fault judgment module, which judges whether the communication device has a fault abnormality based on the fault prediction value output by the fault analysis model and in combination with the normal operation range of the communication device;
[0100] A fault root cause location module, which is used to establish root cause location rules. After detecting a fault abnormality of the communication device, determine the specific source of the fault according to the category of the abnormal feature index.
[0101] The present invention also provides a computer-readable storage medium for fault root cause analysis. A computer program is stored on the storage medium. When the computer program is executed by a processor, it is used to implement the above-mentioned fault root cause analysis method.
[0102] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0103] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed by hardware or software methods depends on the specific application and design constraints of the technical solution.
[0104] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. They may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0105] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application.
Claims
1. A method for analyzing the root cause of a fault, characterized in that, The specific steps include: Step 1: Obtain multiple sets of operation data of communication devices and their corresponding operation labels. The operation data includes signal strength, bit error rate, data transmission rate, signal-to-noise ratio, temperature, packet loss rate, and power supply voltage of the devices. The operation labels include normal state and fault state, and the fault state includes; Step 2: Generate a characteristic index that can represent the operation state of the communication device according to the obtained operation data of the communication device. The characteristic index includes a signal health index, a data transmission performance index, and a hardware operation state index. Based on the characteristic index in the normal state of the operation label, calculate the mean and standard deviation of each type of characteristic index, and establish the normal operation range of the communication device according to the calculated mean and standard deviation. Step 3: Establish a fault analysis model. Use each type of characteristic index as the input of the model and the corresponding fault prediction value of each type of characteristic index as the output of the model. Use normal data to train the device fault prediction analysis model. Obtain the operation data when historical communication devices fail, and input the three types of characteristic indexes in the historical failed operation data into the trained model respectively to obtain three corresponding fault prediction values. The fault prediction values include a signal link fault prediction value, a data transmission fault prediction value, and a hardware operation fault prediction value. Step 4: Based on the fault prediction value output by the fault analysis model, combine the normal operation range of the communication device to determine whether the communication device has a fault anomaly. Step 5: Set up a root cause location rule. After detecting a fault anomaly of the communication device, determine the specific source of the fault according to the category of the abnormal characteristic index.
2. The fault root cause analysis method according to claim 1, characterized in that Generate a characteristic index that can represent the operation state of the communication device. The characteristic index includes a signal health index, a data transmission performance index, and a hardware operation state index. The method is as follows: The signal health index is used to measure the overall quality of device signal reception. Calculate the signal health index based on the signal strength, signal-to-noise ratio, and bit error rate in the operation data. The formula is as follows: Among them, represents the signal health index, is the signal strength, representing the signal reception power, is the signal-to-noise ratio, representing the clarity of the signal, is the bit error rate, representing the reliability of signal transmission; The data transmission performance index is used to evaluate the efficiency and reliability of the device in data transmission. Calculate the data transmission performance index based on the transmission rate, bit error rate, and packet loss rate in the operation data. The formula is as follows: Among them, represents the data transmission performance index, is the data transmission rate, which represents the amount of data transmitted per unit time, is the packet loss rate, which represents the proportion of lost data packets; The hardware operation state index is used to evaluate the health state of the device hardware. Calculate the hardware operation state index based on the device temperature and power supply voltage in the operation data. The formula is as follows: Among them, represents the hardware operation status index, , are the operating value and the rated value of the power supply voltage respectively, , are the operating temperature and the normal operating temperature of the communication device respectively.
3. The fault root cause analysis method according to claim 1, characterized in that Based on the calculated characteristic index, calculate the mean and standard deviation of each type of characteristic index, and establish the normal operation range of the communication device according to the calculated mean and standard deviation. The formula is as follows: Among them, represents the mean of the characteristic index data of the th class, is the total number of the characteristic index data of the th class, is the th characteristic index data in the characteristic index data of the th class, is the index of the quantity of the characteristic index of the th class; ; When it is, it means that the mean of the signal health index is calculated, When it is, it means that the mean of the data transmission performance index is calculated, When it is, it means that the mean of the hardware operation status index is calculated; Among them, represents the standard deviation of the characteristic index data of the th class. When it represents that the standard deviation of the signal health index is calculated. When it represents that the standard deviation of the data transmission performance index is calculated. When The normal range is calculated to obtain an interval based on the mean and standard deviation. The formula is as follows: Among them, represents the normal operating range of the class characteristic index data, is the normal adjustment parameter, with a value of 2 or 3. When , it represents that the normal operating range of the class characteristic index data includes the characteristic index. When , it represents that the normal operating range of the class characteristic index data includes the characteristic index. When it is , it means that the normal operating range of the signal health index is calculated. When it is , it means that the normal operating range of the data transmission performance index is calculated. When it is , it means that the normal operating range of the hardware operating state index is calculated.
4. A method for analyzing the root cause of a fault according to claim 1, characterized in that, Establish a fault analysis model. Input the three types of characteristic indexes in the historical failed operation data into the trained model respectively to obtain three corresponding fault prediction values. The method is as follows: The fault analysis model is established using a deep neural network structure, which includes an input layer, a hidden layer, and an output layer. The input layer is used to receive each type of feature index, the hidden layer is used to process the data of each type of feature index, and the output layer is used to output the probability prediction values of three types of faults. Obtain the data of the historical communication equipment operating normally, and based on expert scoring, evaluate the fault labels corresponding to the three types of feature indexes. Use the mean square error as the loss function for model training, and use the backpropagation algorithm to minimize the loss function to train the fault analysis model. Obtain the data when the historical communication equipment fails, generate three types of feature indexes, and input them into the trained model to obtain three fault prediction values, including the signal link fault prediction value, the data transmission fault prediction value, and the hardware operation fault prediction value.
5. The failure root cause analysis method according to claim 3, characterized in that, Judge whether the communication equipment has a fault anomaly, and the logical formula is: Among them, When it is, it means that the fault prediction value is not within the normal operating range, and there is a fault abnormality in the communication device; When it is, it means that the fault prediction value is within the normal operating range, and there is no fault abnormality in the communication device; is the fault prediction value. When it is, it means the fault prediction value of the signal link, when it is, it means the fault prediction value of data transmission, when it is, it means the fault prediction value of hardware operation.
6. The failure root cause analysis method according to claim 1, wherein , Establish root cause location rules, and determine the specific source of the fault according to the category of the abnormal feature index. The method is: After detecting an anomaly, judge the root cause of the fault through the feature index. The root cause location rules specifically include that when a signal link fault prediction value is not within the normal operating range of the signal health index, the communication equipment fault causes the signal health index to be abnormal. The fault reasons are wireless signal interference or weak signal coverage, antenna or signal link hardware damage, and excessive environmental noise resulting in an increase in the bit error rate. At this time, the antenna connection status of the communication equipment and the network signal environment should be checked. When a data transmission fault prediction value is not within the normal operating range of the data transmission performance index, the communication equipment fault causes the data transmission performance index to be abnormal. The fault reasons are data interface or relay equipment failure, insufficient data transmission bandwidth, or network congestion. At this time, check the status of the network interface device and whether there is a bottleneck in the network transmission path. When a hardware operation fault prediction value is not within the normal operating range of the hardware operation status index, the communication equipment fault causes the hardware operation status index to be abnormal. The fault reasons are overheating of the communication equipment, increased power supply voltage fluctuation, and abnormal heat dissipation system or power supply module. At this time, the hardware temperature monitoring system or the power supply module should be checked whether it is working stably.
7. A root cause analysis system for faults, characterized in that, The analysis system is used to execute a fault root cause analysis method according to any one of claims 1-6, including: An operation data acquisition module, which is used to obtain the operation data and corresponding operation labels of multiple communication devices. The operation data includes the signal strength, bit error rate, data transmission rate, signal-to-noise ratio, temperature, packet loss rate, and power supply voltage of the device. The operation labels include the normal state and the fault state. The fault state includes; Feature Index Generation and Normal Range Setting Module, which is used to generate a feature index that can represent the operating state of the communication device according to the obtained operating data of the communication device. The feature index includes a signal health index, a data transmission performance index, and a hardware operating state index, and based on the feature index in the normal state with the running label, calculate the mean and standard deviation of each type of feature index, and establish the normal operating range of the communication device according to the calculated mean and standard deviation; Fault Prediction Model Construction Module, which is used to establish a fault analysis model, with each type of feature index as the input of the model and the corresponding fault prediction value of each type of feature index as the output of the model. Use normal data to train the device fault prediction analysis model, obtain the operating data when historical communication devices fail, and input the three types of feature indexes in the historical fault operating data into the trained model respectively to obtain three corresponding fault prediction values. The fault prediction values include a signal link fault prediction value, a data transmission fault prediction value, and a hardware operating fault prediction value; Fault Judgment Module, which judges whether the communication device has a fault abnormality based on the fault prediction value output by the fault analysis model and in combination with the normal operating range of the communication device; Fault Root Cause Location Module, which is used to establish root cause location rules, and after detecting a fault abnormality of the communication device, determine the specific source of the fault according to the category of the abnormal feature index.
8. A computer-readable storage medium for root cause analysis of faults, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by a processor, it is used to implement a fault root cause analysis method according to any one of claims 1-6.
Citation Information
Cited By
Semiconductor big data traceability analysis method and system
CN121032344A