Communication equipment performance data risk prediction method, system and application system
By preprocessing the performance data of communication equipment and comparative learning of multi-algorithm models, and automated trend and risk prediction, the problem of low risk prediction efficiency in the existing technology that relies on manual experience is solved, and efficient risk warning and equipment management are achieved.
Patent Information
- Application Number
- CN202510342978.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-11
AI Technical Summary
The risk prediction of existing communication equipment depends on manual analysis, requiring operation and maintenance personnel to have high level of experience, resulting in low risk prediction efficiency and relying on manual experience, and it is impossible to efficiently detect equipment problems in advance.
By preprocessing the performance data of communication equipment, time series data is generated, and multiple algorithm models in the model tree are used for comparison and learning, the final prediction model is determined, and trend prediction and risk assessment is carried out based on the prediction model, including Lasso polynomial regression, a combined prediction model of the mlxtend framework, and NBeats neural network prediction model.
It realizes automated risk prediction, reduces the workload of manual screening of data, can predict equipment performance trends and risks in advance, helps operation and maintenance personnel to quickly understand the equipment status and prioritize potential risks.
Smart Images

Figure CN120296694A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence big data analysis, and particularly relates to a method, a system and an application system for predicting risks of communication device performance data. Background Art
[0002] When existing communication devices are used in actual engineering construction and operation and maintenance, they completely rely on the alarms issued by the devices and the situation of daily inspections to evaluate the probability of future risks of the devices. It completely relies on manual analysis and may only be discovered when problems occur. In addition, during the operation and maintenance of communication devices, according to a large number of alarms generated by the communication devices and the set threshold values, it is determined one by one according to manual experience whether the alarms generated by the devices need to be processed and whether there are risks in this device. That is, manual analysis requires the operation and maintenance personnel to have sufficient processing experience to analyze device problems and possible risks, which has relatively high requirements for the operation and maintenance personnel. Further, the collected performance data needs to be stored and managed, and then manually analyzed in depth through query statistical analysis tools to identify performance trends and potential problems. At the same time, daily inspections and patrols of the devices are also required. During this process, device problems are discovered, and it is determined according to experience whether to repair the problems that have occurred or not to process them. Therefore, the existing methods are all based on the manual experience of the operation and maintenance personnel, and require the operation and maintenance personnel to process with high timeliness and high efficiency to ensure that risks can be discovered and solved in advance.
[0003] Regarding communication devices, how to provide an efficient risk prediction method has increasingly become a technical problem to be solved urgently. Summary of the Invention
[0004] In view of the above problems, the purpose of the present invention is to provide a method for predicting risks of communication device performance data, including:
[0005] Preprocessing the obtained communication device performance data to obtain preprocessed performance data;
[0006] Connecting the preprocessed performance data to a model tree for performance prediction, and performing comparative learning on the algorithm models in the model tree to determine the final prediction model;
[0007] Based on the final prediction model, input real-time communication device performance data to predict the performance trend of the communication device;
[0008] Based on the performance trend prediction result, perform risk prediction according to the set performance index threshold.
[0009] Further, the communication device performance data includes performance data of performance indicators in communication devices corresponding to the transmission network system, data network system, LTE-R system, and / or optical cable monitoring system in the communication system.
[0010] Further, the preprocessing of the obtained communication device performance data includes
[0011] Based on the obtained communication device performance data, time series data is generated;
[0012] According to the preset prediction step and the result of data periodicity judgment, the time series data is trimmed to obtain the preprocessed performance data.
[0013] Further, the model tree includes multiple algorithm models. Connecting the preprocessed performance data to the model tree for performance prediction includes
[0014] For each algorithm model among the multiple algorithm models, the Bayesian optimization method is respectively used for hyperparameter optimization to obtain the corresponding hyperparameter combinations of the multiple algorithm models;
[0015] The multiple algorithm models respectively predict the communication device performance data based on the preprocessed performance data and the obtained hyperparameter combinations, and respectively obtain the prediction data corresponding to each of the multiple algorithm models.
[0016] Further, the multiple algorithm models include a polynomial regression model based on Lasso, a combined prediction model based on the mlxtend framework, and a neural network prediction model based on NBeats.
[0017] Further, connecting the preprocessed performance data to the model tree for performance prediction, and performing comparative learning on the algorithm models in the model tree to determine the final prediction model includes, based on the comparative learning model, performing comparative learning on the algorithm models in the model tree to determine the final prediction model, specifically including
[0018] Using a loss function to optimize the parameters of the multiple algorithm models respectively;
[0019] Using a fitting evaluation method and an anomaly monitoring evaluation method to evaluate the accuracy of each algorithm model among the multiple algorithm models respectively, and respectively obtaining the accuracy evaluation scores of the prediction data corresponding to each algorithm model;
[0020] The algorithm model with the highest accuracy evaluation score is the final prediction model.
[0021] Further, using a fitting evaluation method and an anomaly monitoring evaluation method to evaluate the accuracy of each algorithm model among the multiple algorithm models respectively, and respectively obtaining the accuracy evaluation scores of the prediction data corresponding to each algorithm model includes
[0022] Based on the trend shown by the obtained communication device performance data, the obtained communication device performance data is divided into stable performance data and non-stable performance data;
[0023] Based on piecewise R2 Perform a linear regression analysis on the predicted data and the obtained communication device performance data. Among them,
[0024] When the obtained communication device performance data is stationary performance data, the fitting evaluation satisfies:
[0025]
[0026] Among them, R 2 is the coefficient of determination, and the value of R 2 is between 0 and 1; pred i represents the predicted value of the i-th sample; real i represents the true value of the i-th sample; n represents the number of samples;
[0027] When the obtained communication device performance data is non-stationary performance data, the fitting evaluation satisfies:
[0028]
[0029] Among them:
[0030]
[0031] Among them, SSR is the regression sum of squares, SST is the total sum of squares, SSE is the sum of squares of errors, n is the number of samples, y i is the true value of the i-th sample, is the predicted value of the i-th sample, is the average value of the true values of all samples;
[0032] Perform normalization processing on the calculated R 2 to obtain the percentage score of the predicted data and the obtained communication device performance data in terms of the goodness of fit in R 2 ;
[0033] Respectively obtain the proportions of the predicted data within multiple preset error intervals, and perform weighted calculation on the multiple proportion values according to the corresponding multiple weights to obtain the percentage score of the predicted data in the standard deviation σ anomaly detection;
[0034] Obtain the mean value of the percentage score of the predicted data and the obtained communication device performance data in terms of the goodness of fit in R 2 and the percentage score of the predicted data in the standard deviation σ anomaly detection to obtain the accuracy evaluation score of the predicted data of each algorithm model.
[0035] Furthermore, predicting the communication device performance trend includes predicting the performance data of the communication device after a certain period of time.
[0036] Further, based on the performance trend prediction result, risk prediction according to the set performance metric threshold includes
[0037] Determining the time when the risk occurs according to the set performance metric threshold;
[0038] Determining the risk level according to the time difference between the time when the risk occurs and the current time.
[0039] Further, it also includes timing scheduling management, and the timing scheduling management includes training the final prediction model regularly, updating and saving the final prediction model; and
[0040] Performing online prediction based on the final prediction model and outputting the prediction result.
[0041] Another object of the present invention is to provide a communication device performance data risk prediction system, including
[0042] A preprocessing module for preprocessing the obtained communication device performance data to obtain preprocessed performance data;
[0043] A model determination module for accessing the preprocessed performance data into a model tree for performance prediction, and performing comparative learning on the algorithm models in the model tree to determine the final prediction model;
[0044] A prediction module for predicting the communication device performance trend based on the final prediction model by inputting real-time communication device performance data; and performing risk prediction according to the set performance metric threshold based on the performance trend prediction result.
[0045] Another object of the present invention is to provide a communication device performance data risk prediction application system, including a communication system, the above-mentioned risk prediction system, and an intelligent operation and maintenance system, wherein
[0046] The communication system is used to obtain communication device performance data and send it to the risk prediction system;
[0047] The risk prediction system is used to execute the above-mentioned communication device performance data risk prediction method;
[0048] The intelligent operation and maintenance system is used to maintain the communication device based on the risk prediction result.
[0049] The method of the present invention collects performance data of communication device performance indicators, reduces manual massive screening of indicator data to view changes in performance data, predicts trends in communication device performance data, allows viewing of future prediction trends, anticipates future situations in advance, helps operation and maintenance personnel quickly grasp the status of device performance indicators, and buys sufficient time for handling device problems. In addition, based on the prediction results of the predicted trends of performance indicators, a risk level assessment of device performance indicators is carried out to help operation and maintenance personnel quickly understand the future risk level of the device, and they can prioritize and focus on the device to prevent or avoid risks.
[0050] Other features and advantages of the present invention will be described in the following specification, and some of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 FIG. shows a schematic flowchart of a method for predicting risks of communication device performance data in an embodiment of the present invention;
[0053] Figure 2 FIG. shows a framework diagram of a risk prediction of communication device performance data in an embodiment of the present invention;
[0054] Figure 3 FIG. shows another framework diagram of a risk prediction of communication device performance data in an embodiment of the present invention;
[0055] Figure 4 FIG. shows a schematic structural diagram of a neural network prediction model based on NBeats in an embodiment of the present invention;
[0056] Figure 5 FIG. shows a schematic structural diagram of a system for predicting risks of communication device performance data in an embodiment of the present invention;
[0057] Figure 6 FIG. shows a schematic diagram of an application system for predicting risks of communication device performance data in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] As Figure 1 shown, an embodiment of the present invention discloses a method for predicting risks of communication device performance data. The method is implemented based on AI (Artificial Intelligence), and includes: First, preprocess the obtained communication device performance data to obtain preprocessed performance data; Second, connect the preprocessed performance data to a model tree for performance prediction, and perform comparative learning on the algorithm models in the model tree to determine the final prediction model; Then, based on the final prediction model, input real-time communication device performance data to predict the performance trend of the communication device; Finally, based on the performance trend prediction result, perform risk prediction according to the set performance index threshold. The above method performs risk prediction on the performance data of the communication device performance indicators collected based on the algorithm model, reduces the manual massive screening of indicator data to view the changes in performance data, predicts the trend of the communication device performance data, can view the future prediction trend, anticipate future situations in advance, helps the operation and maintenance personnel quickly master the status of the device performance indicators, and buys enough time for device problem handling. In addition, according to the trend prediction result, the risk level is evaluated, which helps the operation and maintenance personnel quickly understand the future risk level of the device, and can give priority to focusing on the device to prevent or avoid risks.
[0060] Specifically, the communication device performance data is the performance indicator data that focuses on collecting the communication system, including the performance data of the performance indicators of the communication devices corresponding to the transmission network system, data network system, LTE-R system (Long Term Evolution-Railway), and / or optical cable monitoring system in the railway communication system. However, it is not limited to this. The performance data of the communication devices in other systems in the railway system is also applicable to the present invention. The communication device performance data includes, but is not limited to, the performance data of performance indicators such as CPU utilization rate, memory usage rate, network throughput, latency, and packet loss rate. Further, as Figure 2As shown in the figure, the performance data of the communication device obtained is preprocessed, specifically including: First, based on the obtained performance data of the communication device, time series data is generated; Exemplarily, preprocessing steps such as cleaning and standardizing the obtained performance data of the communication device are performed to generate time series data. Then, according to the preset prediction step length and the result of data periodicity judgment, the time series data is cropped to obtain the preprocessed performance data. Among them, the prediction step length is the prediction length determined according to business needs, that is, the prediction step length is to predict multiple future data. The preset prediction step length for communication device performance prediction can be 96, that is, to predict 96 future values. The result of periodicity judgment will obtain the period duration of the data according to the periodicity of the data. For example, if the data collection period is once every 1h (hour), then the prediction step length of 96 can be understood as predicting 96 point data in the future, that is, one prediction value per hour within the next 4 days. At this time, the periodicity of the data is once every hour, and the period duration is 4 days, but it is not limited to this. For example, if the periodicity of the data can be once every 30min, and 96 point data is collected, then the period duration is 2 days, etc., all of which are applicable to the present invention. Thus, by cropping the obtained performance data and selecting an appropriate length of historical monitoring data, the training data is simplified on the premise of ensuring that the periodic characteristics of the data are not lost, and the algorithm performance is improved.
[0061] In the embodiment of the present invention, the model tree includes multiple algorithm models. Specifically, exemplarily, such as Figure 2 As shown in the figure, the model tree includes three algorithm models, including: a polynomial regression model based on Lasso, a combined prediction model based on the mlxtend framework, and a neural network prediction model based on NBeats. However, it is not limited to this. For example, Figure 3 in which, the model tree may include n models, where n is an integer greater than 1. For example, using two or four or more than two algorithm models are all applicable to the present invention. In the following content of the embodiment of the present invention, the above three algorithm models are used for exemplary description. Specifically, the polynomial regression model based on Lasso (Least Absolute Shrinkage and Selection Operator: a linear regression model) satisfies:
[0062] Adding the square of the L2-norm of the w vector with a penalty coefficient λ to the cost function of ridge regression:
[0063]
[0064] Where N represents the number of samples, that is, the number of data in the dataset. That is, under one communication device, a data set formed by collecting data of one performance metric (one type of data in the communication device performance data) at the same time granularity. Exemplarily, for communication device A, the performance metric is CPU utilization rate, with one collection value per hour. Then the data set of the most recent 24 hours is 24 sample numbers; w represents the weight vector, and λ represents the penalty coefficient, which is used to control the weight of the regularization term in the cost function, y i is the true target value of the i-th sample, x i is the feature vector of the i-th sample, w T x i represents the dot product of the weight vector w and the feature vector x i That is, the predicted value of the algorithm model for the i-th sample.
[0065] The Lasso regression algorithm also adds a regularization term like ridge regression, except that it adds the L1-norm of the w vector with a penalty coefficient λ as the penalty term (the meaning of the L1-norm is the sum of the absolute values of each element of the vector w). Therefore, this regularization method is also called L1 regularization.
[0066]
[0067] Similarly, it is to find the magnitude of w when the cost function is minimized:
[0068]
[0069] Since the added is the L1-norm of the vector, and there are absolute values in it, resulting in the cost function not being differentiable everywhere. So, it is impossible to directly obtain the analytical solution of w by the direct derivative method. In the embodiments of the present invention, the Least Angle Regression (LARS) method is mainly used to solve the weight coefficient w. The least angle regression method is a feature selection method that calculates the correlation of each feature and gradually calculates the optimal solution through mathematical formulas. Specifically, it includes the following steps:
[0070] Step 1: Initialize the weight coefficient w, for example, initialize it as a zero vector.
[0071] Step 2: Initialize the residual vector residual as the target label vector y - Xw, where X represents the feature matrix. Since w is a zero vector at this time, the residual vector is equal to the target label vector y;
[0072] Step 3: Select a feature vector x i ,, along the direction of the feature vector x i to find a set of weight coefficients w, and another feature vector x that has the largest correlation with the residual vector appearsj Make the new residual vector residual equal to the correlation between the two eigenvectors (that is, the residual vector is equal to the angular bisector vector of the two eigenvectors), and recalculate the residual vector, where x j is the feature vector of the jth sample.
[0073] Step 4: Repeat step 3 to find a set of weight coefficients w so that the third eigenvector x with the greatest correlation with the residual vector k Make the correlation between the new residual vector residual and the three eigenvectors equal (that is, the residual vector is equal to the equiangular vector of the three eigenvectors), and so on, where x k is the feature vector of the kth sample.
[0074] Step 5: When the residual vector residual is small enough or all feature vectors have been selected, the iteration ends and the predicted data is output.
[0075] The combined prediction model based on the mlxtend framework includes a sequential feature selection algorithm. The prediction methods used in the combined prediction include KNN (K-Nearest Neighbors: K nearest neighbor algorithm), linear regression, support vector machine, simple neural network regression, and finally random forest is used to select the model, and then the prediction data is obtained based on the selected model.
[0076] The neural network prediction model based on NBeats has created a new time series prediction backbone (backbone network), which can achieve time series prediction only through full connection, such as Figure 4 As shown in the figure, the whole model includes multiple stacks, that is, M stacks, M is an integer, and M≥1, each stack includes multiple blocks, that is, K blocks, K is an integer, and K≥1, each block is the most basic structural module of Nbeats, consisting of multiple fully connected layers. Each block contains two main parts. The first part maps the input time series into expansion coefficients, and the second part maps the expansion coefficients back to the time series, and finally obtains the predicted data.
[0077] Therefore, the preprocessed time series data is connected to the model tree for training, which also includes,
[0078] For each of the multiple algorithm models, the Bayesian optimization method is used to optimize the hyperparameters respectively, and the corresponding hyperparameter combinations of the multiple algorithm models are obtained. The Bayesian optimization method: Bayesian optimization is used for machine learning hyperparameter tuning and was proposed by J. Snoek (2012). Given the objective function to be optimized (a general function, only the input and output need to be specified, without knowing the internal structure and mathematical properties), the posterior distribution of the objective function (Gaussian process) is updated by continuously adding sample points until the posterior distribution basically fits the true distribution. It has the advantages of fewer iteration times, faster speed, and being able to find the global minimum from as few steps as possible. Thus, through Bayesian optimization, better algorithm parameters can be obtained with fewer iteration times.
[0079] The multiple algorithm models respectively predict the communication device performance data based on the preprocessed time series data and the obtained hyperparameter combinations, and respectively obtain the prediction data corresponding to each of the multiple algorithm models. Exemplarily, the above three algorithm models are used, and each algorithm model is independent of each other. The Bayesian optimization method optimizes the hyperparameters of the three algorithm models respectively, so as to finally obtain the hyperparameter combinations of the three algorithm models respectively.
[0080] As Figure 2 、 Figure 3 shown, based on the contrast learning model, contrast learning is performed on the multiple algorithm models in the model tree, and it is determined that the final prediction model (i.e., the prediction model in the figure) includes
[0081] using the loss function to optimize the parameters of the multiple algorithm models respectively; the loss function satisfies:
[0082]
[0083] where L represents the value of the loss function, which is the objective to be minimized during the model optimization process, N represents the total number of samples, z i and z′ i respectively represent the representations of the i-th sample and its positive sample in the feature space. In contrast learning, the positive sample is usually a sample similar to the i-th sample. For example, they may come from the same category. z j represents the representation of the j-th sample in the feature space, where j is not equal to i, representing negative samples, and exp is the exponential function used to calculate the exponential decay of the distance. |z i -z′ i | 2 represents the square of the distance between the i-th sample and its positive sample in the feature space, and this distance is usually the Euclidean distance. |z j -z j | 2Denotes the square of the distance between the i-th sample and the j-th sample in the feature space.
[0084] The loss function is adopted to make the distance between similar sample pairs in the feature space smaller. By minimizing the loss function, a better data representation can be obtained, which helps the generalization ability of the model to prevent overfitting.
[0085] The fitting evaluation method and the anomaly monitoring evaluation method are used to evaluate the accuracy of each algorithm model in multiple algorithm models, and the accuracy evaluation scores of the predicted data corresponding to each algorithm model are obtained respectively. Specifically, the fitting evaluation method includes
[0086] Based on the trend reflected by the obtained communication device performance data, the obtained communication device performance data is divided into stable performance data and non-stable performance data. Based on piecewise R 2 (That is, by analyzing the two cases of R 2 ) linear regression analysis is performed on the predicted data and the obtained communication device performance data, where
[0087] For some collected communication device performance data (i.e., real-time data), if its variance is less than one-thousandth of its mean, it is stable performance data. Thus, when the obtained communication device performance data is stable performance data, the fitting evaluation satisfies:
[0088]
[0089] Among them, R 2 is the coefficient of determination, representing the proportion of the variability explained by the model (the degree to which the model can capture the changes in the data). The value of R 2 is between 0 and 1. The closer the value is to 1, the better the model fits. pred i represents the predicted value of the i-th sample (i.e., the predicted data); real i represents the true value of the i-th sample (the obtained real data); n represents the number of samples.
[0090] If the variance of the collected communication device performance data is not less than one-thousandth of its mean, it is non-stable performance data. That is, when the obtained communication device performance data is non-stable performance data, the fitting evaluation satisfies:
[0091]
[0092] Among them:
[0093]
[0094] Among them, SSR (Regression Sum of Squares) is the regression sum of squares, representing the variance explained by the model. SST (Total Sum of Squares) is the total sum of squares, representing the total variance of the performance data. SSE (Error Sum of Squares) is the error sum of squares, representing the variance not explained by the model. n is the number of samples, and y i is the true value of the i-th sample, is the predicted value of the i-th sample, and is the average value of the true values of all samples.
[0095] For the calculated R 2 perform normalization to obtain the percentage score of the predicted data and the true data (i.e., the obtained communication device performance data) in terms of the goodness of fit in R 2 ; specifically, for the calculated R 2 lying in the interval (-∞, 1], for the convenience of subsequent calculations, normalize the score. According to general experience, normalize the scores in the interval [-5, 1] to the interval [0, 100], and calculate all scores in the interval (-∞, -5) as 0 points. Finally, obtain the percentage score of the predicted data and the true data in terms of the goodness of fit in R 2 .
[0096] The abnormal monitoring and evaluation method includes respectively obtaining the proportions of the predicted data within multiple preset error intervals, and performing weighted calculation on the multiple proportion values according to the corresponding multiple weights to obtain the percentage score of the predicted data in the abnormal detection of the standard deviation σ; specifically, use value 真实 ±σ 真实 (i.e., the true value plus or minus the standard deviation calculated based on the true value) as the acceptable error range, and count the proportion of the number of predicted data falling outside the error range. The formula for the standard deviation σ is as follows:
[0097]
[0098] where σ is the standard deviation, n is the number of samples, x i is the i-th sample, and μ is the average value, that is, the average value of the sum of all samples. Thus, according to the above formula, calculate the standard deviation σ using the true value 真实 .
[0099] The multiple preset error intervals can be but are not limited to value 真实 ±1σ 真实 、value 真实 ±2σ 真实 、value 真实 ±3σ 真实Three intervals, and the weights of the above three intervals can be, but are not limited to, 20%, 30%, and 50% respectively.
[0100] Finally, the accuracy of each algorithm model among multiple algorithm models is evaluated separately, including obtaining the average of the percentage score of the predicted data and the real data in terms of the R 2 fitting degree and the percentage score of the predicted data in terms of the standard deviation σ anomaly detection, to obtain the accuracy evaluation score of the predicted data. Among them, the higher the accuracy score, the higher the accuracy of this model. Therefore, the algorithm model with the highest accuracy evaluation score is the final prediction model.
[0101] In the embodiments of the present invention, predicting the performance trend of a communication device includes predicting the performance data of the communication device after a certain period of time. The certain period of time can be 24 hours, or a month, etc., which are all applicable to the present invention. Thus, based on the performance trend prediction result, risk prediction is performed according to the set performance index threshold, including determining the time when the risk occurs according to the set performance index threshold; and determining the risk level according to the time difference between the time when the risk occurs and the current time. The setting of the performance index threshold, time difference, and risk level can be set according to the application scenario. Exemplarily, for communication device A, its performance index is the CPU usage rate. Thus, the data period of the performance data of the performance index: one acquisition value every 30 minutes, the CPU usage rate performance index threshold is 65, the current time is 8:30, and the prediction start time is 9:00, that is, the corresponding time of the first data value in the prediction set is 9:00. Assuming the prediction step is 12, the prediction result is a set of 12 values {55, 58, 62, 70, 72, 72, 72, 72, 72, 72, 72, 72}. Then, starting from the 4th value, it exceeds the threshold of 65. Then, the device is considered a risk device. The time of the 4th value is 10:30, the time when the risk point appears is 10:30, and the time difference from the current time 8:30 is 2h. Then, the risk level is determined according to the time difference. Further exemplarily, the shorter the time difference, the higher the risk level. For example, if the time difference is 1h, it is risk level I, and if the time difference is 2h, it is risk level II, etc.
[0102] In the embodiments of the present invention, the method further includes timing scheduling management, timing training to update and save the training model and online prediction, and outputting the prediction result. After accumulating a certain amount of performance data, the prediction module is trained regularly, and the algorithm model is updated regularly, making the prediction of the communication device performance data more accurate and ensuring the reliability of the operation of the communication device and the corresponding system.
[0103] Such as Figure 5As shown, in an embodiment of the present invention, a risk prediction system for communication device performance data capable of executing the above method is further introduced. The system includes a preprocessing module, a model determination module, and a prediction module. Among them, the preprocessing module is used to preprocess the obtained communication device performance data to obtain the preprocessed performance data; the model determination module is used to access the preprocessed performance data into a model tree for performance prediction, and perform comparative learning on the algorithm models in the model tree to determine the final prediction model; the prediction module is used to input real-time communication device performance data based on the final prediction model to predict the performance trend of the communication device; and based on the performance trend prediction result, perform risk prediction according to the set performance index threshold.
[0104] As Figure 6 shown, in an embodiment of the present invention, a risk prediction application system for communication device performance data is further introduced. The application system includes a communication system, the above-mentioned risk prediction system, and an intelligent operation and maintenance system. Among them, the communication system is used to obtain communication device performance data and send it to the risk prediction system; the risk prediction system is used to execute the above method; the intelligent operation and maintenance system is used to maintain the communication device based on the risk prediction result. Figure 6 Among them, the algorithm ability is to realize the risk prediction of the communication device based on the communication device performance data, and finally apply the prediction result to the intelligent operation and maintenance system for operation and maintenance management and service management.
[0105] Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting risks of communication device performance data, characterized in that, including preprocessing the obtained communication device performance data to obtain preprocessed performance data; connecting the preprocessed performance data to a model tree for performance prediction, and performing comparative learning on the algorithm models in the model tree to determine the final prediction model; based on the final prediction model, inputting real-time communication device performance data to predict the performance trend of the communication device; based on the performance trend prediction result, performing risk prediction according to the set performance index threshold.
2. The method for predicting the risk of communication device performance data according to claim 1, wherein The communication device performance data includes the performance data of the performance indicators in the communication devices corresponding to the transmission network system, data network system, LTE-R system, and / or optical cable monitoring system in the communication system.
3. The method for predicting the risk of communication device performance data according to claim 1, wherein Preprocessing the obtained communication device performance data includes generating time series data based on the obtained communication device performance data; clipping the time series data according to the preset prediction step and the data periodicity judgment result to obtain the preprocessed performance data.
4. The method for predicting the risk of communication device performance data according to claim 1, wherein The model tree includes multiple algorithm models. Connecting the preprocessed performance data to the model tree for performance prediction includes for each algorithm model in the multiple algorithm models, respectively using the Bayesian optimization method to optimize the hyperparameters, and obtaining the corresponding hyperparameter combinations of the multiple algorithm models; the multiple algorithm models respectively predict the communication device performance data based on the preprocessed performance data and the obtained hyperparameter combinations, and respectively obtain the prediction data corresponding to each of the multiple algorithm models.
5. The method for predicting the risk of communication device performance data according to claim 4, wherein The multiple algorithm models include a polynomial regression model based on Lasso, a combined prediction model based on the mlxtend framework, and a neural network prediction model based on NBeats.
6. The method for predicting the risk of communication device performance data according to claim 4 or 5, characterized in that Connecting the preprocessed performance data to the model tree for performance prediction, and performing comparative learning on the algorithm models in the model tree to determine the final prediction model includes: performing comparative learning on the algorithm models in the model tree based on the comparative learning model to determine the final prediction model, specifically including using a loss function to optimize the parameters of the multiple algorithm models respectively; using a fitting evaluation method and an anomaly monitoring evaluation method to evaluate the accuracy of each algorithm model in the multiple algorithm models respectively, and respectively obtaining the accuracy evaluation scores of the prediction data corresponding to each algorithm model; the algorithm model with the highest accuracy evaluation score is the final prediction model.
7. The method for predicting the risk of communication device performance data according to claim 6, wherein Using a fitting evaluation method and an anomaly monitoring evaluation method to evaluate the accuracy of each algorithm model in the multiple algorithm models respectively, and respectively obtaining the accuracy evaluation scores of the prediction data corresponding to each algorithm model includes based on the trend shown by the obtained communication device performance data, dividing the obtained communication device performance data into stable performance data and non-stable performance data; Based on segmented R 2 perform a linear regression analysis on the predicted data and the obtained communication device performance data, where when the obtained communication device performance data is stable performance data, the fitting evaluation satisfies: Among them, R 2 is the coefficient of determination, and the value of R 2 is between 0 and 1; pred i represents the predicted value of the i-th sample; real i represents the true value of the i-th sample; n represents the number of samples; when the obtained communication device performance data is non-stable performance data, the fitting evaluation satisfies: Wherein: Among them, SSR is the regression sum of squares, SST is the total sum of squares, SSE is the error sum of squares, n is the number of samples, and y i is the true value of the i-th sample, is the predicted value of the i-th sample, is the average of the true values of all samples; Normalize the calculated R 2 to obtain the percentage score of the predicted data and the obtained communication device performance data in terms of the fitting degree at R 2 ; respectively obtaining the proportions of the prediction data within multiple preset error intervals, and performing weighted calculation on the multiple proportion values according to the corresponding multiple weights to obtain the percentage score of the prediction data in the standard deviation σ anomaly detection; Obtain the mean of the percentage scores of the predicted data and the obtained communication device performance data in terms of the goodness of fit in R 2 and the percentage scores of the predicted data in terms of standard deviation σ anomaly detection, to obtain the accuracy evaluation score of the predicted data for each algorithm model.
8. The method for predicting risks of communication device performance data according to claim 7, characterized in that Predicting the performance trend of the communication device includes predicting the performance data of the communication device after a certain period of time.
9. The method for predicting the risk of communication device performance data according to claim 8, wherein Based on the performance trend prediction results, risk prediction is carried out according to the set performance index thresholds, including determining the time when the risk occurs according to the set performance index thresholds; determining the risk level according to the time difference between the time when the risk occurs and the current time.
10. The method for predicting the risk of communication device performance data according to claim 9, characterized in that, It also includes timing scheduling management, which includes timing training of the final prediction model and updating and saving the final prediction model; and, online prediction based on the final prediction model and outputting the prediction results.
11. A communication device performance data risk prediction system, characterized in that, Including, a preprocessing module for preprocessing the obtained communication device performance data to obtain preprocessed performance data; a model determination module for accessing the preprocessed performance data into a model tree for performance prediction and performing comparative learning on the algorithm models in the model tree to determine the final prediction model; a prediction module for predicting the performance trend of a communication device by inputting real-time communication device performance data based on the final prediction model; and risk prediction according to the set performance index thresholds based on the performance trend prediction results.
12. A communication device performance data risk prediction application system, characterized in that, Including a communication system, the risk prediction system described in claim 11, and an intelligent operation and maintenance system, where the communication system is used to obtain communication device performance data and send it to the risk prediction system; the risk prediction system is used to execute the communication device performance data risk prediction method described in any one of claims 1-10; the intelligent operation and maintenance system is used to maintain the communication device based on the risk prediction results.