Cell fault condition prediction method and device and related equipment
By combining principal component analysis and machine learning model of cell performance index data, advance prediction of cell failures is achieved, and the problem of failures in the existing technology can only be troubleshooted after a failure occurs, improving network stability and user experience.
Patent Information
- Application Number
- CN202510252254.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The lack of accurate and effective fault prediction technology in the prior art, resulting in network failures that can only be troubleshooted after the failure occurs, and cannot be predicted in advance, affecting network stability and user experience.
By obtaining multiple performance indicator data of the target cell, performing principal component analysis and processing, transforming the data into low-dimensional space, combining the pre-trained lightweight gradient hoist model and the random forest model, output the fault probability and determine the prediction result.
It realizes advance prediction of cell failures, improves network stability and reliability, reduces user complaints, and reduces the time and cost of troubleshooting.
Smart Images

Figure CN120111547A_ABST
Abstract
Description
Background Art
[0002] The wireless network fault detection and processing technology in the related art has the following problems:
[0003] 1) Due to the lack of accurate and effective fault prediction technology, troubleshooting can only be performed after a fault occurs, but it is impossible to predict a fault before it occurs. At this time, the network has already been affected, which will lead to deterioration of network quality and increase in user complaints. If a fault can be predicted before it occurs and the root cause of the fault can be eliminated in advance, the stability and reliability of the network will be improved, effectively reducing user complaints.
[0004] 2) The fault handling method in the relevant technology requires manual troubleshooting of various network elements and indicators. However, the number of existing network indicators is huge, so troubleshooting requires a lot of manpower and time costs. It is difficult to locate the root cause of the fault and perform corresponding repairs in a short time, which will further amplify the impact of the fault.
[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0006] The present disclosure provides a method, an apparatus and related equipment for predicting cell fault conditions, which at least to a certain extent overcome the problem that related technologies lack advance prediction of cell faults.
[0007] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.
[0008] According to one aspect of the present disclosure, a method for predicting a cell fault condition is provided, comprising: acquiring multiple performance indicator data of a target cell; performing principal component analysis on the multiple performance indicator data of the target cell, transforming the multiple performance indicator data of the target cell into a low-dimensional space, and obtaining multiple target performance indicator data; inputting the multiple target performance indicator data of the target cell into a pre-trained lightweight gradient boosting machine model, and outputting the failure probability of each performance indicator data; inputting the failure probability of each performance indicator data into a pre-trained random forest model, and outputting the failure probability of the target cell; and determining a prediction result of the target cell according to the failure probability of the target cell.
[0009] In some exemplary embodiments of the present disclosure, based on the aforementioned scheme, the prediction result includes: whether the target cell will fail or not fail, and determining the prediction result of the target cell according to the failure probability of the target cell includes: comparing the failure probability of the target cell with a first preset threshold; when the failure probability of the target cell is greater than or equal to the first preset threshold, determining that the target cell will fail; when the failure probability of the target cell is less than the first preset threshold, determining that the target cell will not fail.
[0010] In some exemplary embodiments of the present disclosure, based on the aforementioned scheme, the random forest model, when outputting the failure probability of the target cell, is also used to output the contribution value of each performance indicator data among multiple target performance indicator data to the prediction result. After determining that the target cell will fail, the method also includes: locating the fault of the target cell according to the contribution value of each performance indicator data to the prediction result.
[0011] In some exemplary embodiments of the present disclosure, based on the aforementioned scheme, principal component analysis is performed on multiple performance indicator data of the target cell, and the multiple performance indicator data of the target cell are transformed into a low-dimensional space to obtain multiple target performance indicator data, including: calculating the cross-correlation between the multiple performance indicator data of the target cell; screening out performance indicator data that is less than a second preset threshold according to the cross-correlation; and performing principal component analysis on the performance indicator data that is screened out and is less than the second preset threshold to obtain multiple target performance indicator data of the target cell.
[0012] In some exemplary embodiments of the present disclosure, based on the aforementioned scheme, the cross-correlation between the performance indicator data of the target cell is calculated, including: inputting multiple performance indicator data of the target cell into a pre-trained cross-correlation calculation model, and outputting the cross-correlation between the performance indicator data of the target cell.
[0013] In some exemplary embodiments of the present disclosure, based on the aforementioned scheme, before inputting multiple performance indicator data of the target cell into a pre-trained cross-correlation calculation module and outputting the cross-correlation between the performance indicator data of the target cell, the method also includes: preprocessing the performance indicator data of the target cell, and the preprocessing includes at least one of the following: missing value filling processing, data smoothing processing and threshold processing.
[0014] According to another aspect of the present disclosure, a cell fault condition prediction device is also provided, including: a performance indicator data acquisition module, used to acquire multiple performance indicator data of a target cell; a principal component analysis processing module, used to perform principal component analysis on the multiple performance indicator data of the target cell, transform the multiple performance indicator data of the target cell into a low-dimensional space, and obtain multiple target performance indicator data; a performance indicator data failure probability output module, used to input the multiple target performance indicator data of the target cell into a pre-trained lightweight gradient boosting machine model, and output the failure probability of each performance indicator data; a target cell failure probability output module, used to input the failure probability of each performance indicator data into a pre-trained random forest model, and output the failure probability of the target cell; a prediction result determination module, used to determine the prediction result of the target cell according to the failure probability of the target cell.
[0015] According to another aspect of the present disclosure, an electronic device is provided, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any one of the above-mentioned cell fault condition prediction methods by executing the executable instructions.
[0016] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any one of the above-mentioned methods for predicting a cell fault condition is implemented.
[0017] According to another aspect of the present disclosure, a computer program product is further provided, including: a computer program or instructions, wherein when the computer program or instructions are executed by a processor, any one of the above-mentioned methods for predicting a cell fault condition is implemented.
[0018] A method, apparatus and related equipment for predicting cell fault conditions are provided in the embodiments of the present disclosure. Multiple performance indicator numbers of a target cell are processed to obtain multiple target performance indicator data. A lightweight gradient boosting machine model is trained separately for the multiple target performance indicator data to output the failure probability of each target performance indicator data. Finally, the failure probabilities of all target performance indicator data are input into a pre-trained random forest model, thereby predicting whether a failure will occur in the target cell. Furthermore, the embodiments of the present disclosure combine the fine analysis of the lightweight gradient boosting machine model with the decision-making advantages of the random forest model to predict the failure of the target cell, which can improve the prediction accuracy and reduce misjudgment.
[0019] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0021] Figure 1 A schematic diagram showing an exemplary application system architecture of a method for predicting a cell fault condition in an embodiment of the present disclosure;
[0022] Figure 2 A schematic diagram of a method for predicting a cell failure condition in an embodiment of the present disclosure is shown;
[0023] Figure 3 A schematic diagram of a method for determining whether a target cell will fail by combining a lightweight gradient boosting machine model and a random forest model in an embodiment of the present disclosure is shown;
[0024] Figure 4 A schematic diagram of a cell fault condition prediction process in an embodiment of the present disclosure is shown;
[0025] Figure 5 A schematic diagram of a cell fault condition prediction device in an embodiment of the present disclosure is shown;
[0026] Figure 6 A schematic diagram of an electronic device to which a method for predicting a cell fault condition is applied in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0027] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0028] In addition, the described features, structures or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the present disclosure.
[0029] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0030] like Figure 1 As shown, the system architecture includes a terminal device 101 , a network 102 and a network side device 103 .
[0031] The network 102 is a medium for providing a communication link between the terminal device 101 and the network-side device 103 , and may be a wired network or a wireless network.
[0032] Optionally, the wireless network or wired network described above uses standard communication technology and / or protocol. The network is usually the Internet, but it can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a dedicated network or any combination of a virtual private network). In some embodiments, the data exchanged through the network is represented by technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPSec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.
[0033] Optionally, the terminal device in the embodiment of the present disclosure may also be referred to as UE (User Equipment). In a specific implementation, the terminal device may be a mobile phone, a tablet computer (Tablet Personal Computer), a laptop computer (Laptop Computer), a personal digital assistant (Personal Digital Assistant, PDA), a mobile Internet device (Mobile Internet Device, MID), a wearable device (Wearable Device) or a vehicle-mounted device, etc. It should be noted that the specific type of the terminal device is not limited in the embodiment of the present invention.
[0034] The network side device may be a base station, a relay or an access point, etc. The base station may be a base station of 5G or later versions (e.g., 5G NR NB), or a base station in other communication systems (e.g., an eNB base station). It should be noted that the specific type of the network side device is not limited in the embodiments of the present disclosure.
[0035] Those skilled in the art will know that Figure 1 The number of terminals, networks, and network-side devices in the example is only illustrative, and any number of terminals, networks, and network-side devices may be provided according to actual needs. The embodiments of the present disclosure are not limited to this.
[0036] In order to more clearly describe the embodiments of the present disclosure, the following professional terms are explained:
[0037] Feature Engineering is an important step in machine learning and data science. It involves converting raw data into features that can better reflect the nature of the problem. Good feature engineering can significantly improve the performance of the model. The main tasks of feature engineering include feature selection, feature construction, and feature transformation.
[0038] Machine Learning is an important branch of Artificial Intelligence (AI), which enables computers to learn from data and make predictions or decisions without explicit programming. The core of machine learning is to build algorithmic models that enable computers to automatically improve their performance and increase accuracy as experience accumulates.
[0039] Principal Component Analysis (PCA) is a commonly used data dimensionality reduction technique, which is widely used in data preprocessing, feature extraction and visualization. The main goal of PCA is to convert high-dimensional data into low-dimensional data through linear transformation while retaining the variance information of the original data as much as possible.
[0040] Light Gradient Boosting Machine (LightGBM): An efficient machine learning framework based on the Gradient Boosting Decision Tree (GBDT) algorithm, developed by Microsoft Research and open sourced in 2016.
[0041] Random Forest (RF): An ensemble learning method that improves prediction accuracy and prevents overfitting by constructing multiple decision trees and averaging their results. Random Forest has a high accuracy rate for predicting high-dimensional data and can effectively eliminate the impact of noise and outliers.
[0042] Figure 2 A schematic diagram of a method for predicting a cell fault condition in an embodiment of the present disclosure is shown, and the method comprises the following steps:
[0043] S202: Acquire multiple performance indicator data of the target cell.
[0044] It should be noted that the target cell in the embodiment of the present disclosure can be any geographical area covered by one base station or multiple base stations working in coordination. For example, the geographical area covered by base station A includes cell 1, cell 2, and cell 3, and the target cell can be cell 1, cell 2, or cell 3. In addition, multiple performance indicator data of the cell are used to evaluate and optimize the operation of the network. For example, the performance indicator data include: connection success rate, call drop rate / line drop rate, switching success rate, uplink / downlink throughput, etc. In more detail, the connection success rate is a measure of the success rate of user equipment attempting to access the network, usually expressed as a percentage, and a higher success rate means better network coverage and service quality; the call drop rate / line drop rate refers to the proportion of unexpected interruptions during ongoing calls or data transmission. A low call drop rate is an important indicator of network stability; the switching success rate refers to the ratio of users being able to successfully switch from one cell to another when moving between different cells. An efficient switching process is essential to maintaining call quality and data service continuity; the uplink / downlink throughput reflects the speed of data transmission on the network, corresponding to the data transmission rate from the user to the network (uplink) and from the network to the user (downlink), respectively.
[0045] S204, performing principal component analysis on the multiple performance indicator data of the target cell, transforming the multiple performance indicator data of the target cell into a low-dimensional space, and obtaining multiple target performance indicator data.
[0046] It should be noted that the principal component analysis process in the embodiment of the present disclosure is used to reduce the number of dimensions of the performance indicator data. Specifically, the principal component analysis process includes: data standardization, calculation of the covariance matrix or the correlation coefficient matrix, calculation of the eigenvalues and eigenvectors, selection of principal components, and conversion of the original data. In more detail, since different performance indicator data may have different dimensions and scales, it is first necessary to standardize the original multiple performance indicator data so that each variable has a zero mean and unit variance; the covariance matrix shows the degree of covariance between different performance indicator data, while the correlation coefficient matrix measures the strength and direction of this relationship and is not affected by the scale; by solving the eigenvalues and eigenvectors of the covariance matrix or the correlation coefficient matrix, it can be determined which directions (i.e., principal components) explain the largest data variation; sorting according to the size of the eigenvalues, selecting the first few eigenvectors that explain most of the data variation as principal components, usually, when the cumulative contribution rate reaches 70%-95%, it can be considered that these principal components can better represent the information of the original data set; projecting the original multiple performance indicator data onto the selected principal components to obtain multiple target performance indicator data after dimensionality reduction.
[0047] S206, input multiple target performance indicator data of the target cell into a pre-trained lightweight gradient boosting machine model, and output the failure probability of each performance indicator data.
[0048] It should be noted that the lightweight gradient boosting machine model in the embodiment of the present disclosure can be a model obtained by pre-training various artificial intelligence algorithm models (for example, neural network models) or their combination models through machine learning. The model can calculate multiple target performance indicator data of the target cell to obtain the failure probability of each performance indicator data. The input data of the model is the multiple target performance indicator data of the target cell, and the output data is the failure probability of each performance indicator data.
[0049] Through the above embodiment, the lightweight gradient boosting machine model obtained through machine learning training is used to calculate the failure probability of multiple target performance indicator data of the target cell to obtain the failure probability of each corresponding performance indicator data, which can realize the rapid calculation of the failure probability of the target performance indicator data.
[0050] In some embodiments, the lightweight gradient boosting machine model in the embodiments of the present disclosure has a significant improvement in speed compared to the traditional gradient boosting method, especially on large data sets; the lightweight gradient boosting machine model in the embodiments of the present disclosure adopts a histogram algorithm instead of a traditional pre-sorting algorithm, which can significantly reduce memory consumption and make the model training process more lightweight; and it can effectively handle problems with high feature dimensions and large sample numbers, and provide accurate prediction results; moreover, the lightweight gradient boosting machine model supports multi-threaded parallel computing and distributed training, which further speeds up the training speed, and is particularly suitable for large-scale data processing tasks in industrial-grade applications.
[0051] S208, input the failure probability of each performance indicator data into a pre-trained random forest model, and output the failure probability of the target cell.
[0052] It should be noted that the random forest model in the embodiment of the present disclosure can be a model obtained by pre-training various artificial intelligence algorithm models (for example, a neural network model) or their combination models through machine learning. The model can calculate the failure probability of each performance indicator data of the target cell to obtain the failure probability of the target cell. The input data of the model is the failure probability of each performance indicator data of the target cell, and the output data is the failure probability of the target cell.
[0053] In some embodiments, the random forest model in the embodiments of the present disclosure is a powerful integrated learning method that improves the performance of the model by constructing multiple decision trees and aggregating their prediction results. In more detail, the random forest model reduces the variance of a single model by integrating multiple decision trees, thereby reducing the risk of overfitting and improving the prediction accuracy of unseen data; when faced with a data set with a large number of features, the random forest model does not require feature selection or dimensionality reduction to work effectively; because each decision tree can be trained independently, the random forest model can use multi-core processors for parallel computing to speed up training.
[0054] S210: Determine a prediction result of the target cell according to the failure probability of the target cell.
[0055] The method for predicting a cell fault condition provided in an embodiment of the present disclosure first obtains a plurality of performance indicator data of a target cell; secondly, performs principal component analysis on the plurality of performance indicator data of the target cell, transforms the plurality of performance indicator data of the target cell into a low-dimensional space, and obtains a plurality of target performance indicator data; then, inputs the plurality of target performance indicator data of the target cell into a pre-trained lightweight gradient boosting machine model, and outputs the fault probability of each performance indicator data; thereafter, inputs the fault probability of each performance indicator data into a pre-trained random forest model, and outputs the fault probability of the target cell; finally, determines the prediction result of the target cell according to the fault probability of the target cell. Compared with the problem of lack of advance prediction of cell failures in related technologies, the embodiment of the present disclosure processes multiple performance indicator numbers of the target cell to obtain multiple target performance indicator data, trains a lightweight gradient boosting machine model for the multiple target performance indicator data separately, outputs the failure probability of each target performance indicator data, and finally inputs the failure probabilities of all target performance indicator data into the pre-trained random forest model, so as to predict whether the target cell will have a failure. Furthermore, the embodiment of the present disclosure combines the fine analysis of the lightweight gradient boosting machine model with the decision-making advantages of the random forest model to predict failures of the target cell, which can improve the prediction accuracy and reduce misjudgment.
[0056] In some embodiments, the prediction results in the embodiments of the present disclosure include: whether the target cell will fail or will not fail, and the prediction results of the target cell are determined according to the failure probability of the target cell, including: comparing the failure probability of the target cell with a first preset threshold; when the failure probability of the target cell is greater than or equal to the first preset threshold, determining that the target cell will fail; when the failure probability of the target cell is less than the first preset threshold, determining that the target cell will not fail. Specifically, after the random forest model outputs the failure probability of the target cell, if the failure probability is greater than or equal to the first preset threshold, it is determined that the target cell will fail; if the failure probability is less than the first preset threshold, it is determined that the target cell will not fail. In more detail, the first preset threshold in the embodiments of the present disclosure is a threshold value obtained after multiple experiments. For example, the embodiments of the present disclosure can set the first preset threshold value to 30%. It should be noted that the embodiments of the present disclosure do not specifically limit the first preset threshold value, and the first preset threshold value in the embodiments of the present disclosure can be flexibly set according to actual conditions.
[0057] In some embodiments, the random forest model in the disclosed embodiment is also used to output the contribution value of each performance indicator data in multiple target performance indicator data to the prediction result when outputting the failure probability of the target cell. After determining that the target cell will fail, the cell failure situation prediction method in the disclosed embodiment also includes: locating the fault of the target cell according to the contribution value of each performance indicator data to the prediction result. Specifically, the random forest model in the disclosed embodiment can output the contribution value of each performance indicator data to the prediction result during prediction. The performance indicator data with a large contribution value is the root cause of the fault. The root cause of the fault is determined based on the fault prediction, and the root cause of the fault can be quickly identified and located, reducing operation and maintenance costs and improving network stability.
[0058] In some embodiments, the embodiment of the present disclosure performs principal component analysis on multiple performance indicator data of the target cell, transforms the multiple performance indicator data of the target cell into a low-dimensional space, and obtains multiple target performance indicator data, including: calculating the mutual correlation between the multiple performance indicator data of the target cell; filtering out the performance indicator data less than the second preset threshold according to the mutual correlation; performing principal component analysis on the performance indicator data filtered out less than the second preset threshold, and obtaining multiple target performance indicator data of the target cell. Specifically, the embodiment of the present disclosure first calculates the correlation between the multiple performance indicator data of the target cell, retains the performance indicator data with a mutual correlation less than the second preset threshold, and then performs principal component analysis on the retained performance indicator data to achieve data dimensionality reduction, improve the training and reasoning speed of the subsequent model, and reduce the influence of noise; in more detail, the second preset threshold in the embodiment of the present disclosure is a threshold obtained after multiple experiments. For example, the embodiment of the present disclosure can set the second preset threshold to 60%. It should be noted that the embodiment of the present disclosure does not specifically limit the second preset threshold, and the second preset threshold in the embodiment of the present disclosure can be flexibly set according to actual conditions.
[0059] In some embodiments, the disclosed embodiments calculate the mutual correlation between the performance indicator data of the target cell, including: inputting multiple performance indicator data of the target cell into a pre-trained mutual correlation calculation model, and outputting the mutual correlation between the performance indicator data of the target cell. It should be noted that mutual correlation is a statistical method for measuring the degree of similarity between two performance indicator data. Specifically, the neural network model can also be used to calculate or estimate mutual correlation in certain specific scenarios, especially in complex data pattern recognition and high-dimensional data analysis. The mutual correlation calculation model in the disclosed embodiments can be obtained after training the neural network model. For example, the convolutional neural network is good at capturing local features. The disclosed embodiments can indirectly estimate the mutual correlation by learning the features of the input data; for data with time series, a recurrent neural network, a long short-term memory network or a gated recurrent unit can be used to capture time dependencies, and the mutual correlation can be estimated by the learned features.
[0060] In some embodiments, before inputting multiple performance indicator data of the target cell into a pre-trained cross-correlation calculation module and outputting the cross-correlation between the performance indicator data of the target cell, the method for predicting a cell fault condition in the embodiment of the present disclosure also includes: preprocessing the performance indicator data of the target cell, and the preprocessing includes at least one of the following: missing value filling processing, data smoothing processing, and threshold processing. Specifically, preprocessing the performance indicator data of the target cell includes: supplementing missing values; detecting discontinuous zero values in the data, and smoothing with adjacent data; detecting abnormal values that exceed the normal range, and replacing them with the upper and lower limits of the data. By preprocessing the performance indicator data of the target cell, abnormal values can be removed to ensure the authenticity and reliability of the data. Missing data can be processed by interpolation, mean filling, etc. to avoid analysis deviations caused by missing values. Preprocessing the performance indicator data is an important step in the data analysis and optimization process. Preprocessing can significantly improve the quality of the data, thereby making subsequent analysis, modeling, and decision-making more accurate and effective.
[0061] In some embodiments, Figure 3 As shown, the process of combining the lightweight gradient boosting machine model and the random forest model to obtain the prediction result of whether the target cell will fail includes:
[0062] S302, obtaining a plurality of performance indicator data of the target cell in a preset time period in the past, including historical data of performance indicator 1, historical data of performance indicator 2, ..., historical data of performance indicator n.
[0063] S304, input the acquired multiple performance indicator historical data into the corresponding pre-trained lightweight gradient boosting machine model, and output the failure probability of each performance indicator data, including the failure probability of performance indicator 1, the failure probability of performance indicator 2, ..., the failure probability of performance indicator n.
[0064] S306, inputting the failure probability of each performance indicator data into a pre-trained random forest model, and outputting the failure probability of the target cell.
[0065] S308: Determine a prediction result of whether a failure will occur in the target cell according to the failure probability of the target cell.
[0066] In some embodiments, the disclosed embodiments construct a correspondence between performance indicator data and a lightweight gradient boosting machine model, input the performance indicator data to obtain the fault probability, and then input all fault probabilities into the random forest model, and compare the first preset threshold to determine the fault condition of the target cell.
[0067] In some embodiments, Figure 4 As shown, the method for predicting a cell fault condition in an embodiment of the present disclosure specifically includes:
[0068] S402, obtaining a plurality of performance indicator data of the target cell at an hourly granularity within a preset time period in the past, and arranging them in chronological order.
[0069] S404: pre-process the acquired historical data (that is, a plurality of performance indicator data at an hourly granularity within a preset time period in the past).
[0070] S406, calculating the mutual correlation between each performance indicator data, selecting one of the two performance indicator data whose mutual correlation exceeds the second preset threshold according to the preset retention condition and retaining it, until the mutual correlation between all performance indicator data is less than the second preset threshold.
[0071] S408, performing principal component analysis on the retained performance indicator data, discarding the performance indicator data whose eigenvalue is less than the second preset threshold, transforming the original performance indicator data into a low-dimensional space, and obtaining a plurality of target performance indicator data.
[0072] S410, input each target performance indicator data into a pre-trained lightweight gradient boosting machine model to obtain a failure probability of each target performance indicator data failing.
[0073] S412, input the failure probability of all target performance indicator data into the pre-trained random forest model, output the failure probability of the target cell and compare it with the first preset threshold to determine whether the target cell will fail in the future; repeat S410 and S412 when training the model until the error meets the requirement or the number of iterations reaches the upper limit.
[0074] S414: If it is determined that the target cell will fail in the future, the contribution of each performance indicator data to the prediction result is calculated and sorted, and the performance indicator data causing the failure is output to assist in locating the root cause of the failure.
[0075] In some embodiments, the embodiments of the present disclosure control the optimization model according to the error and the number of iterations through repeated training, calculates and outputs key fault performance indicator data for the predicted target cell to assist in root cause location. The random forest model in the embodiments of the present disclosure can cope with complex environments and output root cause indicators based on fault prediction, thereby reducing operation and maintenance costs and improving network stability.
[0076] Based on the same inventive concept, the disclosed embodiment also provides a cell fault condition prediction device, such as the following embodiment. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0077] Figure 5 A schematic diagram of a cell fault condition prediction device in an embodiment of the present disclosure is shown, the device comprising:
[0078] The performance indicator data acquisition module 501 is used to acquire multiple performance indicator data of the target cell;
[0079] The principal component analysis processing module 502 is used to perform principal component analysis on the multiple performance indicator data of the target cell, transform the multiple performance indicator data of the target cell into a low-dimensional space, and obtain multiple target performance indicator data;
[0080] The performance indicator data failure probability output module 503 is used to input multiple target performance indicator data of the target cell into a pre-trained lightweight gradient boosting machine model, and output the failure probability of each performance indicator data;
[0081] The target cell failure probability output module 504 is used to input the failure probability of each performance indicator data into a pre-trained random forest model and output the failure probability of the target cell;
[0082] The prediction result determination module 505 is used to determine the prediction result of the target cell according to the failure probability of the target cell.
[0083] A cell fault condition prediction device provided in an embodiment of the present disclosure obtains multiple performance indicator data of a target cell through a performance indicator data acquisition module; performs principal component analysis on the multiple performance indicator data of the target cell through a principal component analysis processing module, transforms the multiple performance indicator data of the target cell into a low-dimensional space, and obtains multiple target performance indicator data; inputs the multiple target performance indicator data of the target cell into a pre-trained lightweight gradient boosting machine model through a performance indicator data failure probability output module, and outputs the failure probability of each performance indicator data; inputs the failure probability of each performance indicator data into a pre-trained random forest model through a target cell failure probability output module, and outputs the failure probability of the target cell; and determines the prediction result of the target cell according to the failure probability of the target cell through a prediction result determination module. Compared with the problem of lack of advance prediction of cell failures in related technologies, the embodiment of the present disclosure processes multiple performance indicator numbers of the target cell to obtain multiple target performance indicator data, trains a lightweight gradient boosting machine model for the multiple target performance indicator data separately, outputs the failure probability of each target performance indicator data, and finally inputs the failure probabilities of all target performance indicator data into the pre-trained random forest model, so as to predict whether the target cell will have a failure. Furthermore, the embodiment of the present disclosure combines the fine analysis of the lightweight gradient boosting machine model with the decision-making advantages of the random forest model to predict failures of the target cell, which can improve the prediction accuracy and reduce misjudgment.
[0084] In some embodiments, the prediction results in the embodiments of the present disclosure include: the target cell will fail or will not fail. The prediction result determination module in the embodiments of the present disclosure is also used to compare the failure probability of the target cell with a first preset threshold; when the failure probability of the target cell is greater than or equal to the first preset threshold, it is determined that the target cell will fail; when the failure probability of the target cell is less than the first preset threshold, it is determined that the target cell will not fail.
[0085] In some embodiments, the random forest model in the disclosed embodiment is also used to output the contribution value of each performance indicator data in multiple target performance indicator data to the prediction result when outputting the failure probability of the target cell. The cell failure situation prediction device in the disclosed embodiment also includes: a fault location module, which is used to locate the fault of the target cell according to the contribution value of each performance indicator data to the prediction result after determining that the target cell will fail.
[0086] In some embodiments, the principal component analysis processing module in the embodiments of the present disclosure is also used to calculate the cross-correlation between multiple performance indicator data of the target cell; filter out performance indicator data that is less than a second preset threshold value based on the cross-correlation; perform principal component analysis on the performance indicator data that is less than the second preset threshold value that is filtered out to obtain multiple target performance indicator data of the target cell.
[0087] In some embodiments, the principal component analysis processing module in the embodiments of the present disclosure is also used to input multiple performance indicator data of the target cell into a pre-trained cross-correlation calculation model, and output the cross-correlation between the performance indicator data of the target cell.
[0088] In some embodiments, the cell fault condition prediction device in the embodiments of the present disclosure further includes: a performance indicator data preprocessing module, which is used to preprocess the performance indicator data of the target cell before inputting multiple performance indicator data of the target cell into a pre-trained cross-correlation calculation module and outputting the cross-correlation between the performance indicator data of the target cell, and the preprocessing includes at least one of the following: missing value filling processing, data smoothing processing and threshold processing.
[0089] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods or program products. Therefore, various aspects of the present disclosure may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as "circuits", "modules" or "systems".
[0090] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present disclosure, the electronic device comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any one of the above-mentioned methods for predicting a cell fault condition by executing the executable instructions. Since the principle of solving the problem in the electronic device embodiment is similar to that in the above-mentioned method embodiment, the implementation of the electronic device embodiment can refer to the implementation of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0091] Refer to the following Figure 6 The electronic device 600 according to this embodiment of the present disclosure is described. Figure 6 The electronic device 600 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0092] like Figure 6As shown, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 may include but are not limited to: at least one processing unit 601, at least one storage unit 602, and a bus 603 connecting different system components (including the storage unit 602 and the processing unit 601).
[0093] The storage unit stores program codes, which can be executed by the processing unit 601, so that the processing unit 601 executes the steps described in the above “exemplary method” section of this specification according to various exemplary embodiments of the present disclosure.
[0094] In some embodiments, when the electronic device is used to control, for example, the cell fault condition prediction method disclosed above, the processing unit 601 may perform the following steps of the above method embodiment:
[0095] Acquire multiple performance indicator data of the target cell; perform principal component analysis on the multiple performance indicator data of the target cell, transform the multiple performance indicator data of the target cell into a low-dimensional space, and obtain multiple target performance indicator data; input the multiple target performance indicator data of the target cell into a pre-trained lightweight gradient boosting machine model, and output the failure probability of each performance indicator data; input the failure probability of each performance indicator data into a pre-trained random forest model, and output the failure probability of the target cell; determine the prediction result of the target cell according to the failure probability of the target cell.
[0096] The storage unit 602 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6021 and / or a cache storage unit 6022 , and may further include a read-only storage unit (ROM) 6023 .
[0097] The storage unit 602 may also include a program / utility 6024 having a set (at least one) of program modules 6025, such program modules 6025 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0098] Bus 603 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0099] The electronic device 600 may also communicate with one or more external devices 604 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 605. Furthermore, the electronic device 600 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 606. As shown, the network adapter 606 communicates with other modules of the electronic device 600 via a bus 603. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0100] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0101] Based on the same inventive concept, a computer-readable storage medium is also provided in an embodiment of the present disclosure, on which a computer program is stored, and when the computer program is executed by a processor, any of the above-mentioned methods for predicting cell fault conditions is implemented. Since the principle of solving the problem in the computer-readable storage medium embodiment is similar to that in the above-mentioned method embodiment, the implementation of the computer-readable storage medium embodiment can refer to the implementation of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0102] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0103] In the present disclosure, a computer readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, wherein a readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0104] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0105] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0106] Based on the same inventive concept, a computer program product is also provided in an embodiment of the present disclosure, including a computer program product, including: a computer program or an instruction, wherein when the computer program or the instruction is executed by a processor, the method for predicting a cell fault condition in any one of the above method embodiments is implemented. Since the principle of solving the problem in the computer program product embodiment is similar to that in the above method embodiment, the implementation of the computer program product embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0107] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0108] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.
[0109] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0110] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.
Claims
1. A method for predicting cell failure conditions, characterized in that: include: Obtain multiple performance indicator data of the target cell; Performing principal component analysis on the multiple performance indicator data of the target cell, transforming the multiple performance indicator data of the target cell into a low-dimensional space, and obtaining multiple target performance indicator data; Inputting multiple target performance indicator data of the target cell into a pre-trained lightweight gradient boosting machine model, and outputting the failure probability of each performance indicator data; Inputting the failure probability of each performance indicator data into a pre-trained random forest model, and outputting the failure probability of the target cell; A prediction result of the target cell is determined according to the failure probability of the target cell.
2. The method for predicting cell fault conditions according to claim 1, characterized in that: The prediction result includes: whether the target cell will fail or not fail, and determining the prediction result of the target cell according to the failure probability of the target cell includes: Comparing the failure probability of the target cell with a first preset threshold; When the failure probability of the target cell is greater than or equal to the first preset threshold, determining that the target cell will fail; When the failure probability of the target cell is less than the first preset threshold, it is determined that no failure will occur in the target cell.
3. The method for predicting cell fault conditions according to claim 2, characterized in that: When outputting the failure probability of the target cell, the random forest model is further used to output a contribution value of each performance indicator data in a plurality of target performance indicator data to the prediction result. After determining that the target cell will fail, the method further includes: The target cell is fault located according to the contribution value of each performance indicator data to the prediction result.
4. The method for predicting a cell fault condition according to claim 1, characterized in that: The plurality of performance indicator data of the target cell are subjected to principal component analysis, and the plurality of performance indicator data of the target cell are transformed into a low-dimensional space to obtain a plurality of target performance indicator data, including: Calculating the mutual correlation between the multiple performance indicator data of the target cell; Filter out performance indicator data that is less than a second preset threshold value according to the mutual correlation; The performance indicator data that are screened out and are less than the second preset threshold are processed by principal component analysis to obtain multiple target performance indicator data of the target cell.
5. The method for predicting cell fault conditions according to claim 4, characterized in that: Calculating the mutual correlation between the performance indicator data of the target cell includes: The multiple performance indicator data of the target cell are input into a pre-trained mutual correlation calculation model, and the mutual correlation between the performance indicator data of the target cell is output.
6. The method for predicting cell fault conditions according to claim 4, characterized in that: Before inputting the plurality of performance indicator data of the target cell into a pre-trained mutual correlation calculation module and outputting the mutual correlation between the performance indicator data of the target cell, the method further includes: The performance indicator data of the target cell is preprocessed, and the preprocessing includes at least one of the following: missing value filling processing, data smoothing processing and threshold processing.
7. A cell fault condition prediction device, characterized in that: include: A performance indicator data acquisition module, used to acquire multiple performance indicator data of the target cell; A principal component analysis processing module, used to perform principal component analysis on the multiple performance indicator data of the target cell, transform the multiple performance indicator data of the target cell into a low-dimensional space, and obtain multiple target performance indicator data; A performance indicator data failure probability output module, used to input multiple target performance indicator data of the target cell into a pre-trained lightweight gradient boosting machine model, and output the failure probability of each performance indicator data; A target cell failure probability output module, used to input the failure probability of each performance indicator data into a pre-trained random forest model, and output the failure probability of the target cell; The prediction result determination module is used to determine the prediction result of the target cell according to the failure probability of the target cell.
8. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the method for predicting a cell fault condition according to any one of claims 1 to 6 by executing the executable instructions.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting a cell failure condition according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising: A computer program or instruction, characterized in that when the computer program or instruction is executed by a processor, the method for predicting a cell fault condition as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Power distribution network fault outage rate prediction method and system based on improved random forest
CN111027629A
Big data-based optimization method for fault probability prediction of power grid equipment
CN113516280A
Cloud host fault prediction method and device and medium
CN114840402A
Network hidden fault monitoring method and device
CN115208773A
Intranet service quality optimization method and system based on deep reinforcement learning
CN119496716A