A network base station fault diagnosis system and method based on multi-stage probability correction
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF COMPUTING TECH CHINESE ACAD OF SCI
- Filing Date
- 2026-04-10
- Publication Date
- 2026-07-10
AI Technical Summary
Existing network base station fault diagnosis systems lack differentiated calibration capabilities, distribution imbalance compensation mechanisms, and judgment mechanisms for easily confused fault categories when facing long-tailed fault categories, resulting in inaccurate diagnostic results, especially with a low success rate in identifying a few types of faults.
A multi-stage probabilistic correction method is adopted, including a temperature scaling correction module, a tilt correction module, and an arbitration module. By correcting the initial occurrence probability of network base station failures step by step, the accuracy of the final occurrence probability is improved. Combined with the arbitration mechanism, a final decision is made, forming a multi-stage collaborative optimization.
It improves the accuracy of the network base station fault diagnosis system, especially the success rate of identifying a few types of faults, and enhances the system's practicality under complex network conditions.
Smart Images

Figure CN122372393A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication networks, specifically to the field of network base station operation and maintenance, and more specifically, to a network base station fault diagnosis system and method based on multi-stage probability correction. Background Technology
[0002] In communication network operation and maintenance scenarios, network base station fault diagnosis systems can automatically determine the current fault category based on network base station performance indicators, network alarms, and other information, providing decision support for operation and maintenance personnel. However, in practical applications, the fault distribution of network base stations exhibits a significant long-tail distribution characteristic, meaning that a few head fault categories account for the vast majority, while most tail fault categories are extremely rare. This brings problems such as confidence bias, minority class suppression, and easily confused fault categories to network base station fault diagnosis systems, resulting in inaccurate diagnostic results.
[0003] To improve the diagnostic accuracy of network base station fault diagnosis systems, the first approach involves retraining the system using methods such as resampling, loss function adjustment, and feature space optimization. This approach requires a complete training, verification, and testing process before redeployment, consuming significant computational resources and time. However, in real-world applications, network base station fault diagnosis systems, once deployed, often need to maintain long-term stable operation. Therefore, taking the system offline, retraining, and then deploying it again often fails to meet the high availability requirements of communication networks. To avoid retraining the network base station fault diagnosis system, the second approach involves post-processing the system's fault diagnosis-related predictions using methods such as probability calibration, classifier adjustment, and integration, while maintaining the system's current predictive performance, to improve the system's diagnostic performance. However, the existing second approach has the following characteristics and problems:
[0004] Existing solutions typically use temperature scaling to uniformly calibrate parameters for all network base station fault categories. However, in real-world scenarios, different network base station fault categories often exhibit different confidence levels. If a uniform calibration strategy is adopted, it is difficult to achieve fine-grained probabilistic calibration, which will result in poor probabilistic calibration performance.
[0005] For minority network base station faults with few training samples, the probability predicted by the system is often suppressed. However, existing solutions do not guide the correction of the probability based on the actual recognition performance of the system, and therefore cannot effectively improve the recognition success rate of minority network base station faults.
[0006] Some network base station fault categories are highly similar in the feature space, causing the system to frequently make incorrect predictions among these network base station fault categories, giving "misattributed" prediction results. However, existing solutions lack a mechanism for identifying and judging these easily confused network base station faults, resulting in unstable prediction results for these network base station faults.
[0007] Existing solutions often optimize a single post-processing stage in isolation, lacking organic coordination between various correction methods. As a result, the probability deviation problem of various network base station failures has not been systematically solved, but can only be partially alleviated.
[0008] In summary, existing post-processing-based network base station fault diagnosis systems suffer from the following problems: lack of differentiated calibration capabilities for different network base station fault categories, lack of effective compensation mechanisms for the unbalanced distribution of network base station fault categories, lack of decision mechanisms for easily confused network base station fault categories, and lack of a collaborative structured correction framework. These problems all lead to insufficient accuracy of network base station fault diagnosis results, especially for a few types of network base station fault diagnosis results.
[0009] It should be noted that the background information presented here is only for illustrating relevant information about the present invention to aid in understanding the technical solution of the present invention, and does not imply that the relevant information is necessarily prior art. The relevant information was submitted and disclosed together with the present invention, and should not be considered prior art unless there is evidence that the relevant information was disclosed before the filing date of the present invention. Summary of the Invention
[0010] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a network base station fault diagnosis system and method based on multi-stage probability correction.
[0011] The objective of this invention is achieved through the following technical solution:
[0012] According to a first aspect of the present invention, a network base station fault diagnosis system based on multi-stage probability correction is provided, used to determine whether the network base station is faulty based on the current performance indicators and alarms of the network base station, and to output a fault category diagnosis result when a fault is found. The system includes:
[0013] The basic diagnostic unit is a pre-trained network base station fault diagnosis model. It is used to determine whether a fault has occurred in the network base station based on the current performance indicators and alarms of the network base station, and to predict multiple types of faults that may occur and the initial probability of occurrence of each type of fault when a fault occurs.
[0014] A correction unit is used to perform multi-stage correction on the initial occurrence probability of each type of fault predicted by the basic diagnostic unit to obtain the final occurrence probability of each type of fault, and to determine the fault category diagnosis result based on the final occurrence probability of each type of fault. The correction unit includes: a temperature scaling correction module, used to scale and renormalize the initial occurrence probability of each type of fault to obtain the scaled occurrence probability of each type of fault; a tilt correction module, used to tilt correct the scaled occurrence probability of each type of fault according to a pre-configured tilt weight for each type of fault to obtain the final occurrence probability of each type of fault, wherein the fault category with the highest final occurrence probability is the first candidate fault category, and the fault category with the second highest final occurrence probability is the second candidate fault category; and an arbitration module, used to determine whether to initiate arbitration according to a preset arbitration initiation rule, wherein if arbitration is not initiated, the first candidate fault category is used as the fault category diagnosis result; if arbitration is initiated, a judgment result is made according to a preset arbitration decision rule, and either the first candidate fault category or the second candidate fault category is used as the fault category diagnosis result based on the judgment result.
[0015] The output unit is used to acquire and output the fault category diagnosis results.
[0016] This scheme achieves at least the following beneficial technical effects: 1. It calibrates the prediction probability separately for each type of fault, breaking through the limitations of traditional schemes that do not distinguish between fault categories and uniformly calibrate probabilities. This improves the granularity and accuracy of calibration, laying the foundation for the final output of accurate fault category diagnosis results. 2. It cascades the temperature scaling correction module, tilt correction module, and arbitration module in sequence as a correction unit to achieve multi-stage collaborative prediction probability optimization, thereby enhancing the optimization effect of prediction probability and improving the accuracy of fault judgment in the network base station fault diagnosis system.
[0017] Optionally, the temperature scaling correction module scales and renormalizes the initial occurrence probability of each type of fault in the following manner:
[0018]
[0019] in, represent Class of faults, Representative to The result after scaling and renormalizing the initial occurrence probability of the fault class. represent The initial probability of occurrence of a fault class. represent The optimal temperature coefficient for this type of fault. Index representing fault categories, The total number of fault categories. Representing the The initial probability of occurrence of a fault class. Representing the The optimal temperature coefficient for this type of fault.
[0020] This scheme can achieve at least the following beneficial technical effects: by increasing or decreasing the information entropy of the probability distribution through temperature scaling, it can truly reflect the uncertainty of the prediction or suppress noise interference from non-dominant categories, thereby alleviating the system's overconfidence or underconfidence in a certain fault category.
[0021] Optionally, the optimal temperature coefficient for each type of fault comes from the same set of optimal temperature coefficients. This set is obtained by minimizing the objective function based on a pre-acquired validation dataset. The pre-acquired validation dataset includes multiple samples, each sample including network base station performance metrics and alarms, and fault category labels corresponding to the network base station performance metrics and alarms. Minimizing the objective function means:
[0022]
[0023] in, This represents the set of temperature coefficients to be optimized. This represents the number of samples in the validation dataset; Represents the first in the validation dataset Performance metrics and alarms of network base stations for each sample; Represents the first in the validation dataset Fault category labels for each sample; The basic diagnostic unit is based on Among the predicted possible faults, those related to the fault category label The result is the initial probability of occurrence of the same type of fault after scaling and renormalization.
[0024] This scheme can achieve at least the following beneficial technical effects: it ensures that the probability distribution after temperature scaling correction module is statistically most consistent with the actual validation set distribution.
[0025] Optionally, in the tilt correction module, the pre-configured tilt weight for each type of fault is:
[0026]
[0027] in, represent Class of faults, represent Fault class skew weights; The reference recall rate is the average recall rate of the basic diagnostic unit for selected multiple fault types. Represents basic diagnostic units Recall rate for this type of fault; This represents a small positive number that prevents the denominator from being zero; This represents the tilt strength coefficient.
[0028] Optional, Recall rate of class of faults The recall rate is calculated by having the basic diagnostic unit make predictions based on a pre-acquired validation dataset, and by calculating the number of correct and incorrect predictions made by the basic diagnostic unit. The pre-acquired validation dataset includes multiple samples, each sample including network base station performance metrics and alarms, and fault category labels corresponding to the network base station performance metrics and alarms; among which, the recall rate... The calculation method is as follows:
[0029]
[0030] in, This represents the number of times the basic diagnostic unit predicted correctly. This represents the number of prediction errors made by the basic diagnostic unit; where a correct prediction means that the fault category label in the sample is... In the case of a type of fault, among the various types of faults predicted by the basic diagnostic unit, there are... Class of faults, and The initial probability of this type of fault is the highest.
[0031] Optionally, the tilt correction module performs tilt correction on the scaled probability of occurrence for each type of fault in the following manner:
[0032]
[0033] in, represent The scaled probability of occurrence of this type of fault after skew correction. represent Scaling the probability of occurrence of class-specific faults. Represents pre-configuration Fault class skew weights, Index representing fault categories, The total number of fault categories. Representing the Scaling the probability of occurrence of class-specific faults Representing the pre-configured first Tilt weights for fault classes.
[0034] This scheme can achieve at least the following beneficial technical effects: by giving higher weight to a few fault categories with low recall, the recognition probability of these fault categories can be improved.
[0035] Optionally, in the arbitration module, the preset arbitration initiation rule is as follows: if the first candidate fault category and the second candidate fault category belong to a predefined easily confused fault pair, and the difference between the final occurrence probability of the first candidate fault category and the final occurrence probability of the second candidate fault category does not reach the set probability difference threshold, then arbitration is initiated; otherwise, arbitration is not initiated. Here, a predefined easily confused fault pair refers to two types of faults whose probability of being mispredicted as each other is greater than or equal to the preset misprediction probability threshold.
[0036] Optionally, the arbitration module is equipped with a pre-trained binary classifier, and the preset arbitration decision rule is: if The judgment result is as follows: Class of faults; if The judgment result is as follows: If the fault category is determined as type 1, then the result is classified as the first candidate fault category. The class of faults is the first type of fault in the predefined easily confused fault pairs. The type of fault is the second type of fault in the predefined easily confused fault pairs. The final occurrence probabilities of all types of faults are obtained by inputting them into a pre-trained binary classifier. Class of faults and Confidence level of decisions between different types of faults To make the pre-trained binary classifier accurate The classification accuracy of the fault types reached the preset confidence threshold. To make the pre-trained binary classifier accurate The classification accuracy of the fault types reached the preset confidence threshold.
[0037] This scheme can achieve at least the following beneficial technical effects: the pre-trained binary classifier only... or Only when the decision that may change the fault category of the network base station is made, so as to ensure that the final correction is only made when the pre-trained binary classifier is extremely confident, thereby avoiding the introduction of new errors.
[0038] According to a second aspect of the present invention, a method for diagnosing network base station faults is provided. The method includes: S1, acquiring the current performance indicators and alarms of the network base station; S2, performing fault diagnosis based on the acquired current performance indicators and alarms of the network base station using a network base station fault diagnosis system based on multi-stage probability correction as described in the first aspect of the present invention, so as to obtain a fault category diagnosis result.
[0039] Compared with the prior art, the advantages of the present invention are as follows: By introducing a multi-stage cascaded correction unit, the present invention enables the network base station fault diagnosis system to have differentiated calibration capabilities for different network base station fault categories, effective compensation capabilities for the unbalanced distribution of network base station fault categories, and the ability to make judgments on easily confused network base station fault categories. Without modifying the structure and parameters of the original pre-trained network base station fault diagnosis model, it can effectively improve the accuracy of network base station fault diagnosis results, especially the accuracy of fault diagnosis results for a few types of network base stations, thereby enhancing the practicality of the network base station fault diagnosis system under complex network conditions. Attached Figure Description
[0040] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:
[0041] Figure 1 This is a schematic diagram of the architecture of a network base station fault diagnosis system based on multi-stage probability correction according to an embodiment of the present invention;
[0042] Figure 2 This is a schematic diagram of the architecture of the correction unit of the network base station fault diagnosis system based on multi-stage probability correction according to an embodiment of the present invention;
[0043] Figure 3 The figure shows the ablation experiment results of a network base station fault diagnosis system based on multi-stage probability correction according to an embodiment of the present invention. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the invention.
[0045] As mentioned in the background section, existing post-processing-based network base station fault diagnosis systems have the following problems: lack of differentiated calibration capabilities for different network base station fault categories, lack of effective compensation mechanisms for the unbalanced distribution of network base station fault categories, lack of decision mechanisms for easily confused network base station fault categories, and lack of a collaborative structured correction framework. These problems will lead to insufficient accuracy of network base station fault diagnosis results, especially the accuracy of fault diagnosis results for a few types of network base stations.
[0046] To address the aforementioned issues, the inventors discovered that, firstly, existing solutions fail to differentiate calibration for different types of network base station faults because their initial design aims to solve the overall calibration problem, neglecting the heterogeneity of confidence characteristics among different types of network base station faults. For example, some high-frequency network base station fault categories tend to output excessively high confidence levels due to sufficient training samples, while some low-frequency network base station fault categories suffer from generally low confidence levels due to insufficient training. Secondly, the systematic suppression of the predicted probability of minority network base station faults is not entirely due to insufficient system discrimination ability, but rather related to the implicit encoding of the distribution of network base station fault categories into the system parameters during training. However, existing solutions only consider simple prior proportion compensation for the distribution of network base station fault categories, failing to guide the construction of compensation weights by accurately identifying differences in the system's recognition performance on actual verification data. Thirdly, easily confused network base station fault categories are not randomly distributed but concentrated among specific, paired network base station fault categories. Therefore, if these paired easily confused network base station fault categories can be identified in advance, they can be addressed in a targeted manner. Finally, the fundamental reason why existing solutions optimize a single post-processing stage in isolation is that they do not conduct an in-depth analysis of the logical order of the correction stages. In fact, there is an inherent connection between various probability deviation problems. That is, subsequent stages need to take the probability distribution corrected by the previous stages as input in order to form a progressive and collaborative optimization effect.
[0047] Based on the above research findings, this invention provides a network base station fault diagnosis system based on multi-stage probability correction. This system is used to determine whether a network base station is faulty based on its current performance indicators and alarms, and outputs a fault category diagnosis result when a fault is found. In summary, referring to... Figure 1 The system includes:
[0048] The basic diagnostic unit is a pre-trained network base station fault diagnosis model. It is used to determine whether a fault has occurred in the network base station based on the current performance indicators and alarms of the network base station, and to predict multiple types of faults that may occur and the initial probability of occurrence of each type of fault when a fault occurs.
[0049] The correction unit is used to perform multi-stage correction on the initial occurrence probability of each type of fault predicted by the basic diagnostic unit to obtain the final occurrence probability of each type of fault, and to determine the fault category diagnosis result based on the final occurrence probability of each type of fault; wherein, referring to Figure 2The correction unit includes: a temperature scaling correction module, used to scale and renormalize the initial occurrence probability of each type of fault to obtain the scaled occurrence probability of each type of fault; a tilt correction module, used to tilt correct the scaled occurrence probability of each type of fault according to the pre-configured tilt weight of each type of fault to obtain the final occurrence probability of each type of fault, wherein the fault category with the highest final occurrence probability is the first candidate fault category, and the fault category with the second highest final occurrence probability is the second candidate fault category; and an arbitration module, used to determine whether to start arbitration according to a preset arbitration start rule, wherein if arbitration is not started, the first candidate fault category is used as the fault category diagnosis result; if arbitration is started, a judgment result is made according to the preset arbitration judgment rule, and one of the first candidate fault category or the second candidate fault category is used as the fault category diagnosis result according to the judgment result.
[0050] The output unit is used to acquire and output the fault category diagnosis results.
[0051] As summarized above regarding the network base station fault diagnosis system based on multi-stage probability correction, this invention provides a systematic improvement solution for the multiple probability bias problem caused by long-tailed distributed network base station fault categories by introducing a multi-stage collaborative correction unit composed of a temperature scaling correction module, a tilt correction module, and an arbitration module connected in series. Furthermore, this invention overcomes the limitations of existing solutions that uniformly correct faults regardless of their type by applying temperature scaling and tilt correction separately for the probability of each type of network base station fault, and by using an arbitration mechanism for final judgment, achieving more refined differentiated adjustments. Therefore, compared to existing systems, the system provided by this invention improves the accuracy of network base station fault diagnosis, especially when dealing with a few types of network base station faults.
[0052] To better understand the present invention, a detailed description is provided below with reference to specific embodiments.
[0053] According to one embodiment of the present invention, the temperature scaling correction module scales and renormalizes the initial occurrence probability of each type of fault in the following manner:
[0054]
[0055] in, represent Class of faults, Representative to The result after scaling and renormalizing the initial occurrence probability of the fault class. represent The initial probability of occurrence of a fault class. represent The optimal temperature coefficient for this type of fault. Index representing fault categories, The total number of fault categories. Representing the The initial probability of occurrence of a fault class. Representing the The optimal temperature coefficient for this type of fault. It should be noted that if... This indicates that... The initial probability of occurrence of a fault type is softened, which increases the information entropy of the probability distribution formed by all network base station fault categories. This avoids excessive concentration of probability values in the fault category with the highest probability, thereby alleviating the overconfidence of the basic diagnostic unit in that type of network base station fault and truly reflecting the uncertainty of its prediction; if This indicates that... The initial probability distribution of fault types is sharpened, which reduces the information entropy of the probability distribution formed by all network base station fault types. This causes the probability values to converge to the network base station fault type with the highest probability, thereby suppressing noise interference from non-dominant network base station fault types and enhancing the system's certainty in distinguishing such network base station faults.
[0056] According to one embodiment of the present invention, the optimal temperature coefficient for each type of fault comes from the same set of optimal temperature coefficients. This set of optimal temperature coefficients is obtained by minimizing an objective function based on a pre-acquired validation dataset. The pre-acquired validation dataset includes multiple samples, each sample including network base station performance indicators and alarms, and fault category labels corresponding to the network base station performance indicators and alarms. The minimization of the objective function is expressed as follows:
[0057]
[0058] in, This represents the set of temperature coefficients to be optimized, which consists of the initial temperature coefficients for all types of faults. ; This represents the number of samples in the validation dataset; Represents the first in the validation dataset Performance metrics and alarms of network base stations for each sample; Represents the first in the validation dataset Fault category labels for each sample; The basic diagnostic unit is based on Among the predicted possible faults, those related to the fault category label The initial probability of occurrence of the same type of fault is scaled and renormalized. In this embodiment, the optimal temperature coefficient set is obtained by solving an optimization problem on the validation dataset. Specifically, the basic diagnostic unit parameters are fixed first, and then the negative log-likelihood (NLL) of the validation dataset is used as the objective function. Optimization algorithms such as the finite-memory quasi-Newton method are used to minimize NLL to obtain the optimal temperature coefficient set. In one implementation of this invention, the pre-acquired validation dataset contains six network base station fault categories: interference, coverage, hardware, resources, parameter configuration, and other. After solving the optimal temperature coefficient set according to this embodiment, the optimal temperature coefficients for the interference, coverage, hardware, resources, parameter configuration, and other categories are 1.07, 1.02, 1.00, 0.995, 0.8, and 0.8, respectively.
[0059] According to one embodiment of the present invention, in the tilt correction module, the pre-configured tilt weight for each type of fault is:
[0060]
[0061] in, represent Class of faults, represent Tilt weights for fault classes. The reference recall rate is the average recall rate of the basic diagnostic unit for selected fault classes. Specifically, the reference recall rate... The method for determining the reference recall rate is as follows: Prioritize the selection of top network base station fault categories whose sample size in the validation dataset exceeds a preset threshold (e.g., 10%), and then calculate the arithmetic mean of the recall rates of these top network base station fault categories on the validation dataset to determine the reference recall rate. The reason for prioritizing the selection of top network base station fault categories rather than all network base station fault categories to calculate the reference recall rate is that, generally speaking, the training samples of top network base station fault categories are relatively sufficient, so their recall rate can reflect the normal recognition level of the basic diagnostic unit under sufficient learning conditions. Using this as a benchmark allows the tail network base station fault categories to receive a larger bias weight because their recall rate is significantly lower than the benchmark, thereby achieving effective compensation for the minority network base station faults. Represents basic diagnostic units Recall rate for this type of fault. This represents a small positive number to prevent the denominator from being zero (generally taken as an example based on experience). ). The tilt intensity coefficient is used to control the tilt correction magnitude. Its value is determined by the hyperparameter search method, which searches for the optimal value in the interval [0,1] with the goal of maximizing the macro-average F1 score on the validation dataset, so as to achieve the best balance between the head network base station fault category and tail network base station fault category recognition performance.
[0062] According to one embodiment of the present invention, Recall rate of class of faults It involves having the basic diagnostic unit make predictions based on a pre-acquired validation dataset, and then calculating the recall rate based on the number of correct and incorrect predictions made by the basic diagnostic unit; among these, the recall rate is calculated. The calculation method is as follows:
[0063]
[0064] in, This represents the number of times the basic diagnostic unit predicted correctly. This represents the number of prediction errors made by the basic diagnostic unit; where a correct prediction means that the fault category label in the sample is... In the case of a type of fault, among the various types of faults predicted by the basic diagnostic unit, there are... Class of faults, and The initial probability of this type of fault is the highest.
[0065] According to one embodiment of the present invention, the tilt correction module performs tilt correction on the scaled occurrence probability of each type of fault in the following manner:
[0066]
[0067] in, represent The scaled probability of occurrence of this type of fault after skew correction. represent Scaling the probability of occurrence of class-specific faults Represents pre-configuration Fault class skew weights, Index representing fault categories, The total number of fault categories. Representing the Scaling the probability of occurrence of class-specific faults Representing the pre-configured first Tilting weights for fault types. According to the tilt correction method provided in this embodiment, the identification probability of a network base station fault type can be improved by giving higher tilt weights to tail-end network base station fault types with low recall.
[0068] According to an embodiment of the present invention, in the arbitration module, the preset arbitration initiation rule is as follows: if the first candidate fault category and the second candidate fault category belong to a predefined easily confused fault pair, and the difference between the final occurrence probability of the first candidate fault category and the final occurrence probability of the second candidate fault category does not reach a set probability difference threshold, then arbitration is initiated; otherwise, arbitration is not initiated. Herein, a predefined easily confused fault pair refers to two types of faults whose probability of being mispredicted as each other is greater than or equal to a preset misprediction probability threshold. Specifically, a predefined easily confused fault pair includes a first type of fault and a second type of fault corresponding to the predefined easily confused fault pair, and is predefined as follows: the basic diagnostic unit makes predictions based on a pre-acquired validation dataset, and the mutual misprediction probability of the basic diagnostic unit for the first type of fault and the second type of fault is calculated. If the mutual misprediction probability exceeds a preset misprediction probability threshold, then the first type of fault and the second type of fault constitute the predefined easily confused fault pair. Mutual misprediction means that the fault category label of the current sample is the first type of fault, but the fault category with the highest initial occurrence probability predicted by the basic diagnostic unit is the second type of fault, or the fault category label of the current sample is the second type of fault, but the fault category with the highest initial occurrence probability predicted by the basic diagnostic unit is the first type of fault. To better understand the predefined easily confused fault pairs in this invention, a specific example is given below. Suppose the basic diagnostic unit makes predictions based on 100 samples labeled as hardware faults in the validation dataset. However, among the 15 samples predicted, the initial probability of a interference fault is the highest, resulting in a 15% confusion rate for this prediction. Next, the basic diagnostic unit makes predictions based on another 100 samples labeled as interference faults. Again, among the 12 samples predicted, the initial probability of a hardware fault is the highest, resulting in a 12% confusion rate for this prediction. The sum of these two confusion rates is 27%. Therefore, 27% is the probability of the basic diagnostic unit mispredicting both hardware and interference faults. If the preset misprediction probability threshold is 25%, then hardware and interference faults constitute a predefined easily confused fault pair because 27% > 25%.
[0069] According to one embodiment of the present invention, the arbitration module is equipped with a pre-trained binary classifier, which can output the decision confidence between easily confused fault pairs based on the skew-corrected probability vector (i.e., the final occurrence probability of all types of faults). The pre-defined arbitration award rule is: if The judgment result is as follows: Class of faults; if The judgment result is as follows: If the fault category is determined as type 1, then the result is classified as the first candidate fault category. The class of faults is the first type of fault in the predefined easily confused fault pairs. The type of fault is the second type of fault in the predefined easily confused fault pairs. The final occurrence probabilities of all types of faults are obtained by inputting them into a pre-trained binary classifier. Class of faults and Confidence level of decisions between different types of faults To make the pre-trained binary classifier accurate The classification accuracy of the fault types reached the preset confidence threshold. To make the pre-trained binary classifier accurate The classification accuracy of the fault class reaches a preset confidence threshold. In this embodiment, the pre-trained binary classifier only... or The decision that might change the fault category of the network base station is made only when the pre-trained binary classifier is extremely confident, thus avoiding the introduction of new errors.
[0070] According to one embodiment of the present invention, the present invention also provides a network base station fault diagnosis method, the method comprising: S1, obtaining the current performance indicators and alarms of the network base station; S2, performing fault diagnosis based on the obtained current performance indicators and alarms of the network base station using a network base station fault diagnosis system based on multi-stage probability correction as described in any of the above embodiments, so as to obtain a fault category diagnosis result.
[0071] In addition, to verify the effectiveness of the present invention, the inventors also conducted an ablation experiment. This experiment used network base station performance indicators and alarm data as input, allowing the network base station fault diagnosis system to provide fault category diagnosis results. By comparing the performance of fault diagnosis systems composed of different combinations of correction modules (i.e., temperature scaling correction module, tilt correction module, and arbitration module) for network base station fault diagnosis (using Macro-F1 as the metric), the contribution of each module to the system's fault diagnosis performance was verified. Specifically, the dataset used in the experiment consisted of network base station performance indicators and alarm data from a wireless cell in a certain city provided by a certain operator, as well as corresponding poor quality fault root cause data (which can be understood as fault category labels). The poor quality fault root causes included six network base station fault categories: hardware (39.2%), interference (19.4%), coverage (16.7%), resources (15.8%), parameter configuration (1.5%), and other (7.6%), exhibiting typical long-tail distribution characteristics. (Refer to...) Figure 3The experimental results show that the Macro-F1 is 0.610 when no correction module is used for post-processing; the Macro-F1 is improved when any two correction modules are combined (two-stage cascade) in post-processing, verifying the effectiveness of each correction module; the Macro-F1 reaches 0.672 when all three correction modules are combined (three-stage cascade) in post-processing, achieving the best performance, a 10.16% improvement compared to when no correction module is used for post-processing. Furthermore, the performance of the diagnostic system with three-stage cascade is better than that with two-stage cascade, indicating that the correction stages do not act independently, but rather form a progressive collaborative optimization through the cascade structure. The experimental results fully demonstrate that the network base station fault diagnosis system based on multi-stage probability correction provided by this invention can effectively improve the problems of confidence bias, minority class suppression, and easily confused network base station fault categories faced by existing solutions in scenarios with long-tailed distribution of network base station fault categories, proving the practical value of this invention.
[0072] In summary, by introducing a multi-stage cascaded correction unit, this invention enables the network base station fault diagnosis system to possess differentiated calibration capabilities for different network base station fault categories, effective compensation capabilities for imbalanced distribution of network base station fault categories, and the ability to determine easily confused network base station fault categories. Without modifying the structure and parameters of the original pre-trained network base station fault diagnosis model, it effectively improves the accuracy of network base station fault diagnosis results, especially the accuracy of fault diagnosis results for a few types of network base stations, thereby enhancing the practicality of the network base station fault diagnosis system under complex network conditions.
[0073] It should be noted that the application of the present invention is not limited to network base station fault diagnosis scenarios. In fact, any classification application involving long-tail distribution can utilize the multi-stage probability correction in the present invention for post-processing to improve the accuracy of classification results, especially the accuracy of classification results for a few categories. Furthermore, the present invention can be a computer program product, which mainly refers to software products that implement the present invention through computer programs.
[0074] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A network base station fault diagnosis system based on multi-stage probability correction, used to determine whether a network base station is faulty based on its current performance indicators and alarms, and to output a fault category diagnosis result when a fault is found, characterized in that, The system includes: The basic diagnostic unit is a pre-trained network base station fault diagnosis model, which is used to determine whether the network base station is currently faulty based on the current performance indicators and alarms of the network base station, and predict multiple types of faults that may occur and the initial probability of occurrence of each type of fault when a fault occurs. A correction unit is used to perform multi-stage correction on the initial occurrence probability of each type of fault predicted by the basic diagnostic unit to obtain the final occurrence probability of each type of fault, and to determine the fault category diagnosis result based on the final occurrence probability of each type of fault; wherein, the correction unit includes: The temperature scaling correction module is used to scale and renormalize the initial occurrence probability of each type of fault to obtain the scaled occurrence probability of each type of fault. The tilt correction module is used to tilt correct the scaling probability of each type of fault according to the pre-configured tilt weight of each type of fault, so as to obtain the final probability of each type of fault. The fault category with the highest final probability is the first candidate fault category, and the fault category with the second highest final probability is the second candidate fault category. The arbitration module is used to determine whether to initiate arbitration based on preset arbitration initiation rules. If arbitration is not initiated, the first candidate fault category is used as the fault category diagnosis result. If arbitration is initiated, a judgment result is made according to preset arbitration judgment rules, and the first candidate fault category or the second candidate fault category is used as the fault category diagnosis result based on the judgment result. The output unit is used to acquire and output the fault category diagnosis results.
2. The network base station fault diagnosis system based on multi-stage probability correction according to claim 1, characterized in that, The temperature scaling correction module scales and renormalizes the initial occurrence probability of each type of fault in the following way: in, represent Class of faults, Representative to The result after scaling and renormalizing the initial occurrence probability of the fault class. represent The initial probability of occurrence of a fault class. represent The optimal temperature coefficient for this type of fault. Index representing fault categories, The total number of fault categories. Representing the The initial probability of occurrence of a fault class. Representing the The optimal temperature coefficient for this type of fault.
3. The network base station fault diagnosis system based on multi-stage probability correction according to claim 2, characterized in that, The optimal temperature coefficient for each type of fault comes from the same set of optimal temperature coefficients. This set is obtained by minimizing the objective function based on a pre-acquired validation dataset. The pre-acquired validation dataset includes multiple samples, each of which includes network base station performance indicators and alarms, and fault category labels corresponding to the network base station performance indicators and alarms. Minimizing the objective function means: in, This represents the set of temperature coefficients to be optimized. This represents the number of samples in the validation dataset; Represents the first in the validation dataset Performance metrics and alarms of network base stations for each sample; Represents the first in the validation dataset Fault category labels for each sample; The basic diagnostic unit is based on Among the predicted possible faults, those related to the fault category label The result is the initial probability of occurrence of the same type of fault after scaling and renormalization.
4. The network base station fault diagnosis system based on multi-stage probability correction according to claim 1, characterized in that, In the tilt correction module, the pre-configured tilt weight for each type of fault is as follows: in, represent Class of faults, represent Fault class skew weights; The reference recall rate is the average recall rate of the basic diagnostic unit for selected multiple fault types. Represents basic diagnostic units Recall rate for this type of fault; This represents a small positive number that prevents the denominator from being zero; This represents the tilt strength coefficient.
5. The network base station fault diagnosis system based on multi-stage probability correction according to claim 4, characterized in that, Recall rate of class of faults The recall rate is calculated by having the basic diagnostic unit make predictions based on a pre-acquired validation dataset, and by calculating the number of correct and incorrect predictions made by the basic diagnostic unit. The pre-acquired validation dataset includes multiple samples, each sample including network base station performance metrics and alarms, and fault category labels corresponding to the network base station performance metrics and alarms; among which, the recall rate... The calculation method is as follows: in, This represents the number of times the basic diagnostic unit predicted correctly. This represents the number of prediction errors made by the basic diagnostic unit; where a correct prediction means that the fault category label in the sample is... In the case of a type of fault, among the various types of faults predicted by the basic diagnostic unit, there are... Class of faults, and The initial probability of this type of fault is the highest.
6. The network base station fault diagnosis system based on multi-stage probability correction according to claim 1, characterized in that, The tilt correction module performs tilt correction based on the scaling probability of occurrence for each type of fault as follows: in, represent The scaled probability of occurrence of this type of fault after skew correction. represent Scaling the probability of occurrence of class-specific faults. Represents pre-configuration Fault class skew weights, Index representing fault categories, The total number of fault categories. Representing the Scaling the probability of occurrence of class-specific faults Representing the pre-configured first Tilt weights for fault classes.
7. The network base station fault diagnosis system based on multi-stage probability correction according to claim 1, characterized in that, In the arbitration module, the preset arbitration initiation rule is as follows: if the first candidate fault category and the second candidate fault category belong to a predefined easily confused fault pair, and the difference between the final occurrence probability of the first candidate fault category and the final occurrence probability of the second candidate fault category does not reach the set probability difference threshold, then arbitration is initiated; otherwise, arbitration is not initiated. Here, a predefined easily confused fault pair refers to two types of faults whose probability of being mispredicted as each other is greater than or equal to the preset misprediction probability threshold.
8. The network base station fault diagnosis system based on multi-stage probability correction according to claim 7, characterized in that, The arbitration module is equipped with a pre-trained binary classifier, and the preset arbitration decision rules are as follows: like The judgment result is as follows: Class of faults; like The judgment result is as follows: Class of faults; Otherwise, the decision will be the first candidate fault category; in, The class of faults is the first type of fault in the predefined easily confused fault pairs. The type of fault is the second type of fault in the predefined easily confused fault pairs. The final occurrence probabilities of all types of faults are obtained by inputting them into a pre-trained binary classifier. Class of faults and Confidence level of decisions between different types of faults To make the pre-trained binary classifier accurate The classification accuracy of the fault types reached the preset confidence threshold. To make the pre-trained binary classifier accurate The classification accuracy of the fault types reached the preset confidence threshold.
9. A method for diagnosing network base station faults, characterized in that, The method includes: S1. Obtain the current performance indicators and alarms of the network base station; S2. The network base station fault diagnosis system based on multi-stage probability correction as described in any one of claims 1-8 performs fault diagnosis based on the current performance indicators and alarms of the acquired network base station to obtain fault category diagnosis results.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 9.