Abnormality detection method and device, nonvolatile storage medium and electronic equipment
By obtaining network performance data for feature extraction and using pre-trained machine learning model classification, the problem of inability to distinguish the causes of network performance data abnormalities in the prior art is solved, and fast and accurate cause identification and optimization are achieved.
Patent Information
- Application Number
- CN202510458969.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art cannot accurately distinguish the specific causes of abnormal network performance data, resulting in the inability to optimize for specific reasons.
By acquiring network performance data, performing feature extraction, using historical data to determine target thresholds, and using pre-trained machine learning models to classify the causes of exceptions.
It accurately and quickly distinguishes the specific causes of network performance data abnormalities, can optimize specific reasons, and improves the stability of network performance and user experience.
Smart Images

Figure CN120378281A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of communication networks, and in particular, to an anomaly detection method and device, a non-volatile storage medium, and an electronic device. Background Art
[0002] In modern communication networks, the real-time monitoring and analysis of network performance data are key aspects to ensure service quality and user experience. Network operators usually rely on a series of performance metrics, such as packet loss rate, latency, throughput, call success rate, etc., to evaluate the network health status. However, when these metrics show anomalies, such as a sudden increase in packet loss rate or an increase in latency, the network monitoring system can only detect the anomalies in the performance data, but cannot further distinguish whether the anomalies are caused by hardware failures, user behaviors, or other factors. In such cases, network operation and maintenance personnel need to manually check a large amount of device logs, alarm information, and network traffic data to try to locate the root cause of the problem. This process is not only time-consuming and laborious, but also inefficient, and sometimes even causes the root cause of the problem to be ignored, delaying necessary repair or optimization measures.
[0003] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0004] This application provides an anomaly detection method and device, a non-volatile storage medium, and an electronic device, so as to at least solve the technical problem that the specific reasons for the anomalies in network performance data cannot be distinguished due to related technologies, resulting in the inability to optimize network performance according to specific reasons.
[0005] According to one aspect of this application, an anomaly detection method is provided, including: obtaining network performance data, and performing feature extraction on the network performance data to obtain network performance features; determining a target threshold for evaluating whether the network performance data is abnormal according to historical network performance data and historical connection success rate data; comparing the network performance features with the target threshold, and determining whether the network performance data is abnormal according to the comparison result; in the case where the network performance data is abnormal, using a pre-trained machine learning model to classify the reasons for the anomalies in the network performance data.
[0006] Optionally, determining a target threshold for evaluating whether network performance data is abnormal according to historical network performance data and historical call connection rate data includes: obtaining a network operation dataset, where the network operation dataset includes: a first time series corresponding to the historical network performance data and a second time series corresponding to the historical call connection rate data; performing collaborative preprocessing on the data in the network operation dataset to obtain multiple time series matrices; using the multiple time series matrices to train an autoregressive integrated moving average model until a target parameter combination is determined, obtaining a trained autoregressive integrated moving average model, where the target parameter combination is a parameter combination composed of the autoregressive term order, the differencing order, the moving average term order, and the seasonal parameter that minimizes a first parameter, and the first parameter is used to characterize the fitting degree of the autoregressive integrated moving average model; the seasonal parameter is dynamically set according to the periodic analysis result of the user behavior pattern feature vector, where the user behavior pattern feature vector includes: the periodic session frequency and the service request type; determining the target threshold based on the trained autoregressive integrated moving average model.
[0007] Optionally, after generating the target threshold based on the trained autoregressive integrated moving average model, the method further includes: in the case of detecting a target event of a change in the network topology structure, triggering a model parameter retraining process and a target threshold reconstruction process for the autoregressive integrated moving average model.
[0008] Optionally, the machine learning model is trained by the following method: obtaining a training set, where the training set includes: network element device alarm log text streams, user session traffic time series, and network performance metrics; generating a device semantic encoding vector corresponding to the network element device alarm log text stream through domain adaptation fine-tuning, where in the domain adaptation fine-tuning process, the semantic similarity between network element devices is constrained by a network element device topology connection graph, and the network element device topology connection graph includes physical connections and logical connections between network element devices; performing vectorization processing on the user session traffic time series to obtain a first vector, and performing vectorization processing on the network performance metrics to obtain a second vector; determining a first weight corresponding to the first vector, a second weight corresponding to the second vector, and a third weight corresponding to the device semantic encoding vector according to the first quantity of network element devices, the second quantity of terminal devices, the third quantity of network element devices with faults, and the target ratio, where the target ratio is the ratio of the number of times the network element devices successfully establish connections to the number of times the terminal devices request to establish connections with the network element devices; training a preset machine learning model based on the first vector, the first weight, the second vector, the second weight, the device semantic encoding vector, and the third weight, and obtaining a trained machine learning model when a preset stop condition is satisfied.
[0009] Optionally, determining the first weight corresponding to the first vector, the second weight corresponding to the second vector, and the third weight corresponding to the device semantic coding vector according to the number of network element devices, the number of terminal devices, the number of network element devices with faults, and the target ratio of the terminal device includes: calculating the device health factor through the following formula where α is the device health factor, F e is the third quantity, N e is the first quantity; calculating the terminal behavior factor through the following formula where β is the terminal behavior factor, R c is the target ratio, N t is the second quantity; determining the first weight W1 through the following formula: W1 = σ(0.6α + 0.4β), where σ is the normalization function; determining the second weight W2 through the following formula: W2 = σ(0.3(1 - α) + 0.7β); determining the third weight W3 through the following formula: W3 = σ(0.4α + 0.6(1 - β)).
[0010] Optionally, performing feature extraction on the network performance data to obtain network performance features, including: removing invalid information and duplicate information in the network performance data through a data cleaning algorithm to obtain first data, and removing invalid information and duplicate information in the first data through a regular expression to obtain second data; calculating the correlation index between each feature in the second data and the target label through a preset statistical correlation algorithm, and determining an effective feature set based on the correlation index; performing weight correction and logical combination on the features in the effective feature set through a preset rule library to obtain empirically enhanced features; and fusing the effective feature set with the empirically enhanced features to obtain network performance features.
[0011] Optionally, after classifying the cause of the abnormal network performance data by using a pre-trained machine learning model, the method further includes: in the case where the classification result is a network element device failure, locating the fault propagation path and sending the fault propagation path to the network automation operation and maintenance system; in the case where the classification result is an abnormal user behavior, adding the identification information of the target device that causes the abnormal network performance data to a preset database.
[0012] According to another aspect of the present application, there is also provided an anomaly detection device, including: an acquisition module, configured to acquire network performance data and perform feature extraction on the network performance data to obtain network performance features; a first determination module, configured to determine a target threshold for evaluating whether the network performance data is abnormal according to historical network performance data and historical connection rate data; a second determination module, configured to compare the network performance features with the target threshold and determine whether there is an anomaly in the network performance data according to the comparison result; a classification module, configured to, when there is an anomaly in the network performance data, classify the reasons for the anomaly in the network performance data by using a pre-trained machine learning model.
[0013] According to another aspect of the present application, there is also provided a non-volatile storage medium, where the storage medium includes a stored program, and when the program runs, it controls the device where the storage medium is located to execute the above anomaly detection method.
[0014] According to another aspect of the present application, there is also provided an electronic device, including: a memory and a processor, where the processor is configured to run a program stored in the memory, and when the program runs, it executes the above anomaly detection method.
[0015] According to another aspect of the present application, there is also provided a computer program, where when the computer program is executed by a processor, it implements the above anomaly detection method.
[0016] According to another aspect of the present application, there is also provided a computer program product, where the computer program product includes a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above anomaly detection method.
[0017] In the present application, by acquiring network performance data, performing feature extraction on the network performance data to obtain network performance features; determining a target threshold for evaluating whether the network performance data is abnormal according to historical network performance data and historical connection rate data; comparing the network performance features with the target threshold and determining whether there is an anomaly in the network performance data according to the comparison result; and when there is an anomaly in the network performance data, classifying the reasons for the anomaly in the network performance data by using a pre-trained machine learning model, the purpose of accurately and quickly distinguishing the specific reasons for the anomaly in the network performance data is achieved, thereby realizing the technical effect of optimizing the network performance for specific reasons, and further solving the technical problem that the related art cannot distinguish the specific reasons for the anomaly in the network performance data, resulting in the inability to optimize the network performance for specific reasons. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:
[0019] Figure 1 is a flowchart of an anomaly detection method according to an embodiment of the present application;
[0020] Figure 2 is a flowchart of another anomaly detection method according to an embodiment of the present application;
[0021] Figure 3 is a business flowchart of an anomaly detection method according to an embodiment of the present application;
[0022] Figure 4 is a structural diagram of an anomaly detection device according to an embodiment of the present application;
[0023] Figure 5 is a hardware structure block diagram of a computer terminal of an anomaly detection method according to an embodiment of the present application. Detailed implementation manners
[0024] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] According to an embodiment of the present application, a method embodiment of an anomaly detection method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0027] Figure 1 is a flowchart of an anomaly detection method according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0028] Step S102, obtain network performance data, and perform feature extraction on the network performance data to obtain network performance features.
[0029] For example, the anomaly detection method provided in this embodiment can be applied to an IP Multimedia Subsystem (IMS). IMS is a set of standard network architectures formulated by 3GPP (Third Generation Partnership Project), mainly used to provide high-quality multimedia services such as voice, video, and data, and is the core network platform for realizing multimedia services in the next-generation network and mobile communication system. The introduction of the IMS architecture is to build unified, IP-based communication services on top of various access technologies (such as broadband, wireless networks). It allows users to seamlessly switch between different networks while maintaining communication quality and service continuity.
[0030] In step S102, first, it is necessary to collect network performance data from the network regularly or in real time. This part of the data includes but is not limited to key performance indicators such as network latency, packet loss rate, throughput, CPU utilization, memory occupancy, disk I / O, call establishment success rate, and the number of online users.
[0031] Data collection can be achieved through various channels such as network monitoring software, sensors, and log files. The key is to ensure the comprehensiveness and real-time nature of the data. The collected data needs to cover all levels of the network, from the physical link to the application layer, in order to evaluate the network health status from multiple perspectives.
[0032] Next is to extract features from the network performance data. The purpose of feature extraction is to reduce the data dimension, eliminate noise, and retain the information most valuable for judging the network state. Specifically, feature extraction can be carried out from the following aspects: 1. Statistical features: Calculate statistics such as the mean, median, variance, maximum, and minimum of each performance index. These statistics can reflect the stability and fluctuation degree of network performance. 2. Time series features: Extract time-related features, such as the trend, periodicity, and seasonal changes of performance indicators. This is crucial for identifying the long-term change trend and short-term fluctuations of network performance. 3. Correlation features: Analyze the mutual relationship between each index and identify which combination of indicators can best reflect the abnormal network state. For example, the simultaneous occurrence of high CPU utilization and high packet loss rate may indicate serious network congestion. 4. Behavioral features: Extract features based on user behavior patterns to understand whether user activities during a specific period affect network performance. For example, large-scale user call activities during the business peak period may temporarily reduce the connection rate.
[0033] Step S104: Determine a target threshold for evaluating whether the network performance data is abnormal based on the historical network performance data and historical connection rate data.
[0034] Step S104 involves determining a dynamic or intelligent target threshold based on the historical network performance data and historical connection rate data to evaluate whether the current network performance data is abnormal. The related method is to use a static threshold, that is, a value is preset in advance, and if this value is exceeded, it is considered abnormal. This method is simple but not accurate enough because it does not consider the natural fluctuations of network performance and the changes in the external environment. In this embodiment, a time series analysis method is used to predict future performance indicators based on historical data, and a confidence interval including upper and lower limits is calculated as the target threshold.
[0035] Step S106: Compare the network performance features with the target threshold, and determine whether the network performance data is abnormal according to the comparison result.
[0036] Step S106 is a process of comparing the network performance features obtained through feature extraction with the target threshold determined in step S104. The comparison method can be based on statistical principles, such as Z-score, IQR (interquartile range) detection, etc., to judge whether the current network performance index exceeds the normal range. Preferably, if any index exceeds the target threshold, the system will mark the current network performance data as abnormal.
[0037] Step S108: In the case where the network performance data is abnormal, use a pre-trained machine learning model to classify the reasons for the abnormal network performance data.
[0038] Step S108 classifies and identifies the reasons for the abnormal network performance data using a pre-trained machine learning model. The machine learning model here can be a classifier, such as Random Forest (RF), Logistic Regression (LR), Support Vector Machine (SVM), etc., or a deep learning model, such as Convolutional Neural Network (CNN), Long Short-Term Memory Network (LSTM), etc.
[0039] The purpose of model classification is to attribute the anomalies to specific categories such as hardware failures, software defects, network attacks, user behaviors, or external factors. This process involves model training and testing. Specifically, a historical dataset (including labeled normal and abnormal data samples) is used to train the model. The training data should contain examples of various types of anomalies to ensure that the model can identify and distinguish different reasons for anomalies. Before model training, the feature set that has the greatest impact on the model classification effect is selected through feature importance analysis to improve the accuracy and operating efficiency of the model. The performance of the model is evaluated through techniques such as cross-validation and grid search, and the model parameters are continuously adjusted and optimized until satisfactory accuracy and recall rates are achieved. Finally, the trained and optimized model is used to analyze network performance data in real time, quickly identify the type of anomaly, provide timely guidance for network operation and maintenance personnel, quickly locate the problem, and take corresponding measures for repair or optimization, thereby minimizing the impact of anomalies on the network and ensuring the stable operation of the network and the service experience of users.
[0040] According to the above steps, network performance data is obtained, feature extraction is performed on the network performance data to obtain network performance features; according to historical network performance data and historical connection success rate data, a target threshold for evaluating whether the network performance data is abnormal is determined; the network performance features are compared with the target threshold, and based on the comparison result, it is determined whether there is an abnormality in the network performance data; in the case where the network performance data is abnormal, by using the method of classifying the reasons for the abnormal network performance data with a pre-trained machine learning model, the purpose of accurately and quickly distinguishing the specific reasons for the abnormal network performance data is achieved, thereby realizing the technical effect of optimizing the network performance for specific reasons.
[0041] The following gives Figure 1 an exemplary illustration and explanation of the
[0042] According to some alternative embodiments of the present application, determining a target threshold for evaluating whether network performance data is abnormal based on historical network performance data and historical call connection rate data can be achieved through the following method: Obtain a network operation dataset, where the network operation dataset includes: a first time series corresponding to the historical network performance data and a second time series corresponding to the historical call connection rate data; perform collaborative preprocessing on the data in the network operation dataset to obtain multiple time series matrices; use the multiple time series matrices to train an autoregressive integrated moving average model until a target parameter combination is determined, obtaining a trained autoregressive integrated moving average model, where the target parameter combination is a parameter combination composed of the autoregressive order, differencing order, moving average order, and seasonal parameter that minimizes a first parameter, and the first parameter is used to characterize the fitting degree of the autoregressive integrated moving average model; the seasonal parameter is dynamically set according to the periodic analysis result of the user behavior pattern feature vector, where the user behavior pattern feature vector includes: periodic session frequency and service request type; based on the trained autoregressive integrated moving average model, determine the target threshold.
[0043] In the above embodiment, the ARIMA model is introduced. The ARIMA model is a statistical model widely used in time series prediction and is particularly suitable for analyzing data with trends, seasonality, and autocorrelation. When training the ARIMA model, it is necessary to determine the autoregressive order (p), differencing order (d), moving average order (q), and seasonal parameter (s). The selection of these parameters directly affects the fitting degree and prediction ability of the model.
[0044] The goal of parameter optimization is to find a set of parameter combinations that minimize the first parameter (such as the AIC or BIC information criterion, or the sum of squared prediction errors) of the model, thereby obtaining the best model fitting effect. This process is achieved through grid search, random search, or gradient-based optimization algorithms, continuously adjusting the parameter combination, and evaluating the performance of the model on the validation set until the optimal combination is found.
[0045] The selection of the seasonal parameter (s) is based on the periodic analysis result of the user behavior pattern feature vector. In the user behavior pattern feature vector, the periodic session frequency and service request type are two key indicators, which can reveal the regularity of user activities and the seasonal changes in network load. For example, there may be significant differences in the session frequency and service request type during different time periods such as weekdays and weekends, day and night, commercial promotion periods and non-promotion periods. By analyzing these periodic characteristics, the seasonal pattern in the network performance data can be determined, providing a basis for setting the seasonal parameter of the ARIMA model.
[0046] After completing the model training, based on the trained ARIMA model, the normal range of network performance data can be predicted. The determination of the target threshold depends on the prediction results of the model and is achieved through the following steps: Use the trained ARIMA model to predict the network performance data for a period of time in the future. Based on the uncertainty of the model prediction, calculate the confidence interval of the predicted value, such as the 95% confidence interval. Take the upper and lower bounds of the prediction interval as the target threshold, and any real-time network performance data that exceeds this range is regarded as abnormal. As the model continues to learn and the network environment changes, the target threshold should be updated regularly to reflect the latest network status and user behavior patterns.
[0047] Through the above method, an ARIMA model can be constructed from historical network performance data and historical call connection rate data. This model can dynamically set seasonal parameters according to the periodicity of user behavior patterns, thereby accurately predicting the normal range of network performance. Based on the target threshold determined by the prediction results of this model, the network monitoring system can intelligently distinguish normal fluctuations and abnormal situations, improving the efficiency and accuracy of anomaly detection.
[0048] It should be explained that the seasonal parameters are dynamically set according to the periodic analysis results of the feature vectors of user behavior patterns. Among them, the periodic session frequency refers to the frequency of users communicating or using services in the network, showing a pattern according to a specific time period. Sessions can be any form of network interaction, including but not limited to phone calls, video conferences, instant messages, web browsing, online payments, etc. Here, the time period can be a certain time period of each day (such as working hours or after-school hours), certain days of each week (such as weekends or weekdays), specific dates of each month (such as pay days), or specific annual events (such as holidays, promotional activities, etc.).
[0049] The type of business request refers to different types of requests or services initiated by users to the network or service provider. In a network environment, different types of business requests have different demands and impacts on network resources. For example, video streaming service requests usually require higher bandwidth and lower latency, while text email service requests have relatively lower requirements for the network. For example, enterprise users have a higher session frequency during working hours on weekdays, while it is relatively lower at night and on weekends, forming an obvious periodic pattern. Another example is that during shopping festivals such as "Double 11" on e-commerce platforms, the session frequency increases significantly, showing obvious seasonal fluctuations compared to normal times.
[0050] The above dynamic setting method can be specifically implemented through the following method: By analyzing the user behavior pattern feature vectors, such as the periodic session frequency and the business request type, the periodic pattern of user activities is identified. For example, by analyzing the periodic session frequency, it can be found that the session volume of the user significantly increases during specific time periods of each day (such as 9 am to 11 am, 2 pm to 4 pm), which may be related to the working hours. By observing the business request type, it can be identified that during special festivals (such as "618", "Double 11") or promotion periods, the request volume of specific services (such as online payment, shopping cart query) will show a periodic surge.
[0051] Through the above analysis, the periodic pattern of user behavior is identified, which forms the basis for setting the seasonal parameters. Then, the seasonal parameter (s) of the ARIMA model is dynamically set according to the identified periodic pattern. For example, if it is found through user behavior analysis that the network performance data shows an obvious periodic change every 24 hours, then the seasonal parameter (s) of the ARIMA model may be set to 24, telling the model that the pattern repeating every 24 time points (such as every 24 hours) is seasonal, and the model should consider this periodicity for prediction.
[0052] Optionally, after generating the target threshold based on the trained autoregressive integrated moving average model, the method further includes: in the case of detecting a target event of a change in the network topology structure, triggering a retraining process of the model parameters of the autoregressive integrated moving average model and a reconstruction process of the target threshold.
[0053] The change in the network topology structure, whether due to the addition of network devices, the adjustment of connection methods, or the upgrade of network configurations, will affect the network performance, and the original ARIMA model and the target threshold set based on historical data may no longer be applicable. Therefore, timely retraining of the model and adjustment of the threshold are the keys to ensuring the continuous effectiveness of the network anomaly detection system.
[0054] According to some other alternative embodiments of the present application, the machine learning model is trained by the following method: obtaining a training set, where the training set includes: network element device alarm log text streams, user session traffic time series, and network performance metrics; generating a device semantic encoding vector corresponding to the network element device alarm log text stream through domain adaptation fine-tuning, where during the domain adaptation fine-tuning process, the semantic similarity between network element devices is constrained by using the network element device topology connection graph, and the network element device topology connection graph includes physical connections and logical connections between network element devices; vectorizing the user session traffic time series to obtain a first vector, and vectorizing the network performance metrics to obtain a second vector; determining a first weight corresponding to the first vector, a second weight corresponding to the second vector, and a third weight corresponding to the device semantic encoding vector according to the first quantity of network element devices, the second quantity of terminal devices, the third quantity of network element devices with faults, and the target ratio of the terminal devices, where the target ratio is the ratio of the number of times the network element devices successfully establish connections to the number of times the terminal devices request to establish connections with the network element devices; training a preset machine learning model based on the first vector, the first weight, the second vector, the second weight, the device semantic encoding vector, and the third weight, and obtaining a trained machine learning model when the preset stop condition is satisfied.
[0055] The above-mentioned domain adaptation fine-tuning to generate the device semantic encoding vector uses a pre-trained natural language processing model (such as BERT, RoBERTa, etc.) to further train and adjust the model within a specific network operation and maintenance domain to better understand and encode the information in the network element device alarm log text stream. This process can more accurately capture and encode the semantic information of the device alarm text by considering the topology connection graph between network element devices, while maintaining the coherence and consistency of the semantics between devices, so that it can not only reflect the state of a single device, but also reflect the interdependence and influence relationship between devices. Among them, the network element device topology connection graph contains physical connections and logical connections between devices. The physical connection refers to the actual lines or interface connections between devices, and the logical connection covers abstract connection methods such as network protocols and data flow paths. These connection information are used as additional constraint conditions during the domain adaptation fine-tuning process to help the model understand the context of the device alarm text, that is, the alarm of one device is related to the states of other devices directly or indirectly connected to it.
[0056] Generate a device semantic encoding vector corresponding to the alarm log text stream of the network element device through domain adaptive fine-tuning. During the domain adaptive fine-tuning process, the semantic similarity between network element devices is constrained by using the topological connection graph of the network element device. The topological connection graph of the network element device includes physical connections and logical connections between network element devices. It can be achieved through the following method: In the fine-tuning stage, design a multi-task learning framework. On the one hand, maintain the integrity of the original semantic information through the text reconstruction task. On the other hand, construct a contrastive learning task or a similarity constraint loss using the edge relationship of the topological graph to force the distance between the encoding vectors of adjacent devices to be less than that of non-adjacent devices. For the problem of domain differences, an adversarial training strategy can be adopted to align the feature distributions of the source domain and the target domain, and at the same time, combine the domain-invariant features of the topological graph as an auxiliary supervision signal to enhance the generalization ability of the model to devices in the new domain. The finally generated semantic encoding vector can not only reflect the actual content of the alarm text but also imply the connection relationship between devices, providing a feature representation that is more in line with the network characteristics for subsequent fault analysis, root cause location, and other tasks.
[0057] Specifically, according to the number of network element devices, the number of terminal devices, the number of network element devices with faults, and the target ratio of the terminal devices, determine the first weight corresponding to the first vector, the second weight corresponding to the second vector, and the third weight corresponding to the device semantic encoding vector, including the following steps: Calculate the device health factor through the following formula where α is the device health factor, F e is the third quantity, N e is the first quantity; Calculate the terminal behavior factor through the following formula where β is the terminal behavior factor, R c is the target ratio, N t is the second quantity; Determine the first weight W1 through the following formula: W1 = σ(0.6α + 0.4β), where σ is the normalization function; Determine the second weight W2 through the following formula: W2 = σ(0.3(1 - α) + 0.7β); Determine the third weight W3 through the following formula: W3 = σ(0.4α + 0.6(1 - β)).
[0058] For the device health factor α, the ratio of the number of faulty devices to the total number of devices reflects the prevalence of faults. The lower the ratio, the better the health status of the device. Secondly, the total number of devices is adjusted through the log function, considering the impact of network scale on fault perception. In a large-scale network, even a small number of device failures may cause serious impacts. Therefore, the log(1 + N e ) part ensures that as the network scale increases, the evaluation of the device health status will not increase linearly.
[0059] For the terminal behavior factor β, the impact degree of the terminal device behavior on the network performance is evaluated. When the target ratio is low, that is, the request ratio of the network element device to successfully connect is small, the terminal device may impose additional pressure on the network, so the value of β will be high. In addition, Partially reflects the proportional relationship between the number of terminals and the number of network elements, ensuring that the model will not be overly biased towards terminal behavior even in scenarios where the number of terminal devices is much larger than the number of network element devices.
[0060] Furthermore, σ is a normalization function used to map the calculation results to a specific range (such as between 0 and 1) for easy comparison and use of weights.
[0061] W1 is the combined weight based on the device health factor α and the terminal behavior factor β, reflecting the importance of the user session traffic time series sequence (the first vector). 0.6α + 0.4β means that in weight determination, more importance is attached to the device health status, but the impact of terminal behavior is not ignored. The normalization function ensures a reasonable range of weights.
[0062] Similar to W1, W2 takes into account the device health status and terminal behavior, but focuses on the impact of terminal behavior. The 0.3(1 - α) part reflects the contribution of the weight when the device health status is poor, which means that when the device health status in the network is relatively poor, the weight of the network performance indicator (the second vector) will increase to reflect potential performance problems.
[0063] W3 represents the weight of the device semantic encoding vector (the third vector), comprehensively considering the impacts of the device health status and terminal behavior. W3 finds a balance between the device health status and terminal behavior, ensuring the importance of the device semantic encoding vector in the model.
[0064] Through the above weight determination method, the network anomaly detection system can comprehensively consider the health status of devices in the network, the behavior of terminal devices, and network performance indicators, providing a data fusion strategy for model training, enabling the model to more accurately evaluate the network status, distinguish abnormal user behavior from network device failures, and effectively improve the efficiency and accuracy of network operation and maintenance. The dynamic adjustment mechanism of weights can adapt to changes in the network environment, ensuring that the model can maintain good prediction performance under different network scales and user activity levels.
[0065] As some alternative embodiments of the present application, feature extraction is performed on network performance data to obtain network performance features, which can be achieved through the following methods: removing invalid information and duplicate information in the network performance data through a data cleaning algorithm to obtain first data, and removing invalid information and duplicate information in the first data through regular expressions to obtain second data; calculating the correlation index between each feature in the second data and the target label through a preset statistical correlation algorithm, and determining an effective feature set based on the correlation index; performing weight correction and logical combination on the features in the effective feature set through a preset rule library to obtain empirically enhanced features; and fusing the effective feature set with the empirically enhanced features to obtain network performance features.
[0066] In the above embodiment, data cleaning is a key step before data analysis, and its purpose is to remove noise and redundancy in the data and improve data quality. Removing invalid information and duplicate information in the network performance data through a data cleaning algorithm, the invalid information includes data that cannot be used for analysis, records with incorrect formats, etc., and the duplicate information refers to the same data that appears multiple times in the data set. The first data obtained in this step is a network performance data set that has removed preliminary redundancy and errors. Then, the first data is further cleaned using regular expressions. Regular expressions are a powerful text processing tool that can accurately match and extract specific patterns or structures in the text. Through regular expressions, unstructured or non-standard text in the data can be further removed to ensure data consistency and readability. After removing invalid information and duplicate information in this step, second data is obtained, that is, a cleaner, more structured and clearer network performance data set.
[0067] Use a preset statistical correlation algorithm to calculate the correlation index between each feature in the data set and the target label (such as network anomaly status). The correlation index can be the Pearson correlation coefficient, Spearman rank correlation coefficient or mutual information, etc. These indexes can quantify the relationship strength between the feature and the target, and help us identify which features have a significant impact on predicting network anomaly status. By setting a correlation threshold, features with a high correlation with the target label can be screened out, and these features form an effective feature set, which have higher value in describing network status and predicting anomalies.
[0068] Using a preset rule library, the features in the effective feature set are corrected in weight and logically combined. The rule library can be based on the experience of network operation and maintenance experts or historical fault data. It contains some rules for identifying the logical relationship between a specific combination of features and network anomalies, as well as the importance level of the features. By correcting the feature weights, it can be ensured that the features that contribute more to network anomaly detection occupy a more important position in subsequent analysis. The logical combination is to combine multiple features according to a certain logical relationship to reflect the change pattern of complex network states. For example, the combination of high CPU usage and high memory usage may indicate server overload, while the combination of low connection rate and high packet loss rate may mean network link problems. Through logical combination, the model can utilize the interaction between features to improve the ability to identify network anomalies.
[0069] The effective feature set is fused with experience-enhanced features to obtain network performance features. Experience-enhanced features refer to the feature set obtained according to the logical combination and weight correction in the rule library. They contain the experience of network operation and maintenance experts and the wisdom of historical data. The fusion process may be achieved through feature weighting, feature splicing, or other techniques in feature engineering to ensure that the final network performance feature set can comprehensively and accurately reflect the network state and has good predictability.
[0070] Through the above steps, not only the invalid and duplicate information in the data is removed, but also the most valuable feature set for network anomaly detection is identified and optimized through statistical analysis and expert experience. The final network performance feature set will be used as an important input for the network anomaly detection model to help the model more accurately identify and locate anomalies in the network, improving the efficiency and intelligent level of network operation and maintenance.
[0071] Optionally, after classifying the reasons for the abnormal network performance data using a pre-trained machine learning model, the following steps can also be performed: in the case where the classification result is a network element device failure, locate the fault propagation path and send the fault propagation path to the network automated operation and maintenance system; in the case where the classification result is an abnormal user behavior, add the identification information of the target device that causes the abnormal network performance data to the preset database.
[0072] In the above steps, when the classification result indicates a network element device failure, a fault propagation path localization program will be initiated. This program typically relies on the network topology diagram and the connection relationships of the faulty devices. By tracing the associations between the faulty device and its adjacent devices, it analyzes how the fault gradually spreads in the IMS network. The localization algorithm can employ depth-first search (DFS), breadth-first search (BFS), or other graph traversal techniques to determine the scope and path of the fault's impact downstream from the initial node. After localizing the fault propagation path, a detailed fault report is generated, including the faulty device ID, fault type, list of affected devices, and the specific propagation path of the fault. This report is sent to the network automated operation and maintenance system, which, based on this information, automatically or semi-automatically performs actions such as fault isolation, resource scheduling, and fault recovery, accelerating the fault handling speed, reducing network interruption time, and enhancing the continuity of network services and user satisfaction. If the detection result shows that the anomaly is caused by user behavior, identifying the specific target devices that lead to abnormal network performance data may involve analyzing which devices have generated abnormal call patterns, data traffic, or other metrics that deviate from normal behavior. The identification information of the target devices (such as device ID, MAC address, IP address, etc.) is recorded in a specially established preset database. This database is used to store all known user behavior anomaly events, including details such as the identification of the abnormal device, the specific manifestations of the abnormal behavior, and the occurrence timestamp. The abnormal device information in the preset database can be used for subsequent in-depth analysis to help the operation and maintenance team understand the patterns and root causes of abnormal behavior, and then formulate targeted preventive measures and response strategies. For example, if it is found that a certain type of user behavior often causes network congestion, traffic control strategies can be considered, or network resource allocation can be optimized to mitigate the impact of this behavior on network performance.
[0073] Through the above steps, not only can device failures and user behavior anomalies be accurately distinguished, but also corresponding measures can be taken to effectively localize faults and record abnormal behaviors, thereby enhancing the efficiency and accuracy of network operation and maintenance. This differentiation and processing mechanism play an important role in ensuring network stability and optimizing the user experience. At the same time, through continuous learning and database accumulation, the system continuously evolves and improves its anomaly detection and response capabilities to adapt to the ever-changing network environment.
[0074] Figure 2 is a flowchart of another anomaly detection method according to an embodiment of the present application, Figure 3 is a business flowchart of an anomaly detection method according to an embodiment of the present application, where, Figure 2 the method shown is applied to Figure 3 the four modules in
[0075] First, the data collection and preprocessing stage starts with the setting of network performance metrics. Through in-depth docking with the network management system, key performance data of the IMS network is automatically collected, covering important metrics such as call establishment success rate and the number of online users, and stored in a time series database for subsequent data query and analysis in the time dimension. At the same time, the collected data is labeled as abnormal or normal according to network alarm situations, forming the basis for supervised learning. When encountering abnormal labels, more extensive performance metric data and network element alarm information related to this abnormality are further collected, including but not limited to CPU load, memory usage, network traffic, etc., providing more basis for accurately identifying the source of the abnormality. The application of regular expressions and data cleaning algorithms ensures the integrity and consistency of the data, eliminating invalid and duplicate information and laying a solid data foundation for subsequent analysis.
[0076] Secondly, the feature engineering module is responsible for extracting key features from the time series database and the preprocessed data to reflect the network operation status and user behavior patterns. This includes two in-depth operations: on the one hand, further removing outliers and noise through data cleaning to purify the data set; on the other hand, conducting feature correlation analysis based on label information to identify the features most influential for anomaly detection, and combining the experience of network operation and maintenance experts to screen out the metrics and strategies often used in fault troubleshooting to construct empirical features. The integration of statistical features and empirical features not only considers the objective analysis driven by data but also incorporates the subjective judgment of human experience, greatly improving the comprehensiveness and effectiveness of the feature set and providing strong support for distinguishing subsequent user behavior and network anomalies.
[0077] Then, the time series prediction and dynamic threshold adjustment link introduces the ARIMA model, trains based on historical connection rate data, captures the internal law of network performance changing over time, and generates predicted values of future network performance according to this law. Combining historical data trends and statistical analysis, the anomaly detection threshold is dynamically adjusted according to the seasonal changes of network load and user activities, ensuring that the threshold setting can not only respond to real network faults in a timely manner but also avoid false alarms caused by regular user behavior, greatly improving the accuracy and flexibility of anomaly detection.
[0078] Next, the anomaly detection engine triggers the anomaly detection mechanism based on the comparison between the real-time performance data and the predicted dynamic threshold. In the initial stage, due to the lack of sufficient large amounts of data required for model training, the system temporarily relies on the association rule model generated by combining expert experience with clustering analysis to initially identify abnormal user behaviors, especially those call fluctuations caused by external factors such as promotional activities and holiday effects in the short term. As data accumulates, the metrics, alarm information, and occurrence times during abnormal behaviors are stored in the anomaly location database as the training materials for the supervised learning of the random forest model. Through the efficient key features obtained by feature engineering, the random forest classification model trained later can more accurately distinguish between abnormal user behaviors and network device failures, significantly improving the intelligence level of anomaly detection.
[0079] Finally, the result feedback and optimization loop function runs through the entire system. Through the real-time data processing framework Kafka, the system can quickly capture network alarm data, such as warnings of a decrease in the connection rate metric. The anomaly detection results are compared with the actual network status to verify the effectiveness of the model and feedback the results to the system for continuously optimizing the model parameters and adjusting the algorithm strategy, forming a closed loop of self-learning and improvement, and continuously enhancing the accuracy and efficiency of anomaly detection. When an anomaly is detected, the system can quickly send an alarm to network administrators to facilitate rapid response and repair, ensuring network stability and user experience.
[0080] Figure 4 is a structural diagram of an anomaly detection device according to an embodiment of the present application, as Figure 4 shown, the device includes:
[0081] An acquisition module 42, configured to acquire network performance data and perform feature extraction on the network performance data to obtain network performance features.
[0082] A first determination module 44, configured to determine a target threshold for evaluating whether the network performance data is abnormal according to historical network performance data and historical connection rate data.
[0083] A second determination module 46, configured to compare the network performance features with the target threshold and determine whether there is an abnormality in the network performance data according to the comparison result.
[0084] A classification module 48, configured to classify the reasons for the abnormality of the network performance data by using a pre-trained machine learning model when the network performance data is abnormal.
[0085] Optionally, the first determination module 44 is further configured to perform the following steps: obtain a network operation dataset, where the network operation dataset includes: a first time series corresponding to historical network performance data and a second time series corresponding to historical connection rate data; perform collaborative preprocessing on the data in the network operation dataset to obtain multiple time series matrices; use the multiple time series matrices to train an autoregressive integrated moving average model until a target parameter combination is determined, obtaining a trained autoregressive integrated moving average model, where the target parameter combination is a parameter combination composed of the autoregressive term order, the differencing order, the moving average term order, and the seasonal parameter that minimizes a first parameter, and the first parameter is used to characterize the fitting degree of the autoregressive integrated moving average model; the seasonal parameter is dynamically set according to the periodic analysis result of the user behavior pattern feature vector, where the user behavior pattern feature vector includes: the periodic session frequency and the service request type; based on the trained autoregressive integrated moving average model, determine a target threshold.
[0086] Optionally, the anomaly detection device is further configured to, after generating a target threshold based on the trained autoregressive integrated moving average model, perform the following steps: in the case of detecting a target event of a change in the network topology structure, trigger a model parameter retraining process and a target threshold reconstruction process for the autoregressive integrated moving average model.
[0087] Optionally, the machine learning model is trained by the following method: obtain a training set, where the training set includes: network element device alarm log text streams, user session traffic time series, and network performance metrics; generate a device semantic encoding vector corresponding to the network element device alarm log text stream through domain adaptation fine-tuning, where in the domain adaptation fine-tuning process, the semantic similarity between network element devices is constrained by using a network element device topology connection graph, and the network element device topology connection graph includes physical connections and logical connections between network element devices; perform vectorization processing on the user session traffic time series to obtain a first vector, and perform vectorization processing on the network performance metrics to obtain a second vector; determine a first weight corresponding to the first vector, a second weight corresponding to the second vector, and a third weight corresponding to the device semantic encoding vector according to a first quantity of network element devices, a second quantity of terminal devices, a third quantity of network element devices with faults, and a target ratio of the terminal devices, where the target ratio is the ratio of the number of times the network element devices successfully establish connections to the number of times the terminal devices request to establish connections with the network element devices; based on the first vector, the first weight, the second vector, the second weight, the device semantic encoding vector, and the third weight, train a preset machine learning model, and obtain a trained machine learning model when a preset stop condition is satisfied.
[0088] Optionally, according to the number of network element devices, the number of terminal devices, the number of network element devices with faults, and the target ratio of the terminal devices, determining the first weight corresponding to the first vector, the second weight corresponding to the second vector, and the third weight corresponding to the device semantic coding vector includes: calculating the device health factor through the following formula where α is the device health factor, F e is the third quantity, N e is the first quantity; calculating the terminal behavior factor through the following formula where β is the terminal behavior factor, R c is the target ratio, N t is the second quantity; determining the first weight W1 through the following formula: W1 = σ(0.6α + 0.4β), where σ is the normalization function; determining the second weight W2 through the following formula: W2 = σ(0.3(1 - α) + 0.7β); determining the third weight W3 through the following formula: W3 = σ(0.4α + 0.6(1 - β)).
[0089] Optionally, the obtaining module 42 is further configured to perform the following steps: removing invalid information and duplicate information in the network performance data through a data cleaning algorithm to obtain first data, and removing invalid information and duplicate information in the first data through a regular expression to obtain second data; calculating the correlation index between each feature in the second data and the target label through a preset statistical correlation algorithm, and determining an effective feature set based on the correlation index; performing weight correction and logical combination on the features in the effective feature set through a preset rule library to obtain empirically enhanced features; and fusing the effective feature set with the empirically enhanced features to obtain network performance features.
[0090] Optionally, after using a pre-trained machine learning model to classify the reasons for the abnormality of the network performance data, the anomaly detection device further performs the following steps: in the case where the classification result is a network element device failure, locating the fault propagation path and sending the fault propagation path to the network automation operation and maintenance system; in the case where the classification result is an abnormal user behavior, adding the identification information of the target device that causes the abnormality of the network performance data to a preset database.
[0091] It should be noted that the above Figure 4 each module may be a program module (for example, a set of program instructions for implementing a specific function), or a hardware module. For the latter, it may be presented in the following forms, but not limited to: the above each module is presented as a processor, or the functions of the above each module are implemented by a processor.
[0092] It should be noted that Figure 4 the preferred implementation manners of the illustrated embodiments can be referred to Figure 1The related descriptions of the illustrated embodiments are not repeated herein.
[0093] Figure 5 The hardware structure block diagram of a computer terminal for implementing an anomaly detection method is shown. As Figure 5 shown, the computer terminal 50 may include one or more processors 502 (shown as 502a, 502b, ……, 502n in the figure) (the processor 502 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 504 for storing data, and a transmission module 506 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 5 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 50 may further include more or fewer components than Figure 5 shown, or have a different configuration from Figure 5 shown.
[0094] It should be noted that the above one or more processors 502 and / or other data processing circuits are generally referred to as "data processing circuits" herein. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 50. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).
[0095] The memory 504 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the anomaly detection method in the embodiments of the present application. The processor 502 executes various functional applications and data processing by running the software programs and modules stored in the memory 504, that is, implements the above-mentioned anomaly detection method. The memory 504 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 504 may further include a memory remotely set relative to the processor 502, and these remote memories can be connected to the computer terminal 50 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0096] The transmission module 506 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the computer terminal 50. In one example, the transmission module 506 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network element devices through a base station so as to communicate with the Internet. In one example, the transmission module 506 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0097] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 50.
[0098] It should be noted here that in some alternative embodiments, the above-mentioned Figure 5 The computer terminal shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 5 is only an example of a specific specific instance and is intended to show the types of components that may exist in the above computer terminal.
[0099] It should be noted that Figure 5 The computer terminal shown is used to execute Figure 1 the anomaly detection method shown, so the relevant explanations in the execution method of the above command also apply to this electronic device, which will not be elaborated here.
[0100] The embodiment of the present application also provides a non-volatile storage medium. The non-volatile storage medium includes a stored program. Wherein, when the program runs, it controls the device where the storage medium is located to execute the above anomaly detection method.
[0101] The program executed by the non-volatile storage medium has the following functions: obtaining network performance data, extracting features from the network performance data to obtain network performance features; determining a target threshold for evaluating whether the network performance data is abnormal according to historical network performance data and historical connection rate data; comparing the network performance features with the target threshold, and determining whether there is an anomaly in the network performance data according to the comparison result; in the case where the network performance data is abnormal, using a pre-trained machine learning model to classify the reasons for the abnormal network performance data.
[0102] The embodiment of the present application also provides an electronic device, including: a memory and a processor. The processor is used to run the program stored in the memory. Wherein, when the program runs, it executes the above anomaly detection method.
[0103] The processor is used to run a program that performs the following functions: obtain network performance data, perform feature extraction on the network performance data to obtain network performance features; determine a target threshold for evaluating whether the network performance data is abnormal based on historical network performance data and historical call connection rate data; compare the network performance features with the target threshold, and determine whether there is an abnormality in the network performance data according to the comparison result; in the case where the network performance data is abnormal, use a pre-trained machine learning model to classify the reasons for the abnormal network performance data.
[0104] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0105] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0106] In the above embodiments of the present application, the information collected is information and data authorized by the user or fully authorized by all parties, and the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application complies with relevant laws, regulations, and standards, takes necessary protection measures, does not violate public order and good customs, and provides a corresponding operation entry for the user to choose to authorize or refuse.
[0107] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0108] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0109] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0110] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network element device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0111] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. An anomaly detection method, characterized in that, Including: Obtain network performance data, and perform feature extraction on the network performance data to obtain network performance features; Determine a target threshold for evaluating whether the network performance data is abnormal according to historical network performance data and historical call connection rate data; Compare the network performance features with the target threshold, and determine whether there is an abnormality in the network performance data according to the comparison result; When the network performance data is abnormal, use a pre-trained machine learning model to classify the reasons for the abnormality of the network performance data.
2. The method according to claim 1, wherein Determine a target threshold for evaluating whether the network performance data is abnormal according to historical network performance data and historical call connection rate data, including: Obtain a network operation data set, where the network operation data set includes: a first time series corresponding to the historical network performance data and a second time series corresponding to the historical call connection rate data; Perform collaborative preprocessing on the data in the network operation data set to obtain multiple time series matrices; Use the multiple time series matrices to train an autoregressive integrated moving average model until a target parameter combination is determined, and obtain a trained autoregressive integrated moving average model, where the target parameter combination is an autoregressive term order, a differencing order, a moving average term order, and a seasonal parameter that minimize a first parameter, and the first parameter is used to characterize the fitting degree of the autoregressive integrated moving average model; the seasonal parameter is dynamically set according to the periodic analysis result of the user behavior pattern feature vector, where the user behavior pattern feature vector includes: periodic session frequency and service request type; Determine the target threshold based on the trained autoregressive integrated moving average model.
3. The method according to claim 2, wherein After generating the target threshold based on the trained autoregressive integrated moving average model, the method further includes: When a target event of detecting a change in the network topology structure occurs, trigger a model parameter retraining process and a target threshold reconstruction process for the autoregressive integrated moving average model.
4. The method according to claim 1, wherein The machine learning model is trained by the following method: Obtain a training set, where the training set includes: network element device alarm log text streams, user session traffic time series, and network performance metrics; Generate a device semantic coding vector corresponding to the network element device alarm log text stream through domain adaptation fine-tuning, where in the domain adaptation fine-tuning process, the semantic similarity between network element devices is constrained by a network element device topology connection graph, and the network element device topology connection graph includes physical connections and logical connections between network element devices; Perform vectorization processing on the user session traffic time series to obtain a first vector, and perform vectorization processing on the network performance metrics to obtain a second vector; Determine the first weight corresponding to the first vector, the second weight corresponding to the second vector, and the third weight corresponding to the device semantic coding vector according to the first quantity of network elements, the second quantity of terminal devices, the third quantity of network elements with faults, and the target ratio of the terminal devices, where the target ratio is the ratio of the number of times the network elements successfully establish connections and the number of times the terminal devices request to establish connections with the network elements; Based on the first vector, the first weight, the second vector, the second weight, the device semantic coding vector, and the third weight, train a preset machine learning model, and when the preset stop condition is satisfied, obtain the trained machine learning model.
5. The method according to claim 4, wherein Determine the first weight corresponding to the first vector, the second weight corresponding to the second vector, and the third weight corresponding to the device semantic coding vector according to the quantity of network elements, the quantity of terminal devices, the quantity of network elements with faults, and the target ratio of the terminal devices, including: Calculate the device health factor through the following formula where α is the device health factor, F e is the third quantity, N e is the first quantity; Calculate the terminal behavior factor through the following formula where β is the terminal behavior factor, R c is the target ratio, N t is the second quantity; Determine the first weight W1 through the following formula: W1 = σ(0.6α + 0.4β), where σ is a normalization function; determine the second weight W2 through the following formula: W2 = σ(0.3(1 - α) + 0.7β); determine the third weight W3 through the following formula: W3 = σ(0.4α + 0.6(1 - β)).
6. The method according to claim 1, wherein Extract features from the network performance data to obtain network performance features, including: Remove invalid information and duplicate information in the network performance data through a data cleaning algorithm to obtain first data, and remove invalid information and duplicate information in the first data through a regular expression to obtain second data; Calculate the correlation index between each feature in the second data and the target label through a preset statistical correlation algorithm, and determine an effective feature set based on the correlation index; Modify the weights and perform logical combination on the features in the effective feature set through a preset rule library to obtain empirically enhanced features; Fuse the effective feature set with the empirically enhanced features to obtain the network performance features.
7. The method according to claim 1, characterized in that, After classifying the reasons for the abnormal network performance data using a pre-trained machine learning model, the method further includes: When the classification result is a network element fault, locate the fault propagation path and send the fault propagation path to the network automation operation and maintenance system; When the classification result is an abnormal user behavior, add the identification information of the target device that causes the abnormal network performance data to a preset database.
8. An anomaly detection device, characterized in that, Including: An acquisition module, configured to acquire network performance data and extract features from the network performance data to obtain network performance features; A first determination module, configured to determine a target threshold for evaluating whether the network performance data is abnormal according to historical network performance data and historical connection success rate data; A second determination module, configured to compare the network performance features with the target threshold and determine whether the network performance data is abnormal according to the comparison result; A classification module, configured to, when there is an abnormality in the network performance data, classify the cause of the abnormality in the network performance data by using a pre-trained machine learning model.
9. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored program, wherein when the program runs, it controls the device where the non-volatile storage medium is located to execute the abnormality detection method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Comprising: A memory and a processor, the processor is configured to run a program stored in the memory, wherein when the program runs, it executes the abnormality detection method according to any one of claims 1 to 7.
11. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the abnormality detection method according to any one of claims 1 to 7.
Citation Information
Cited By
Memory overflow program exception positioning method and device based on intelligent analysis
CN120762952A
Detecting anomalous resource distribution patterns in distributed artificial intelligence-based agent networks
US12592897B2
Detecting anomalous resource distribution patterns in distributed artificial intelligence-based agent networks
US20260012432A1