Core network element state diagnosis method and device based on dynamic AI algorithm

By adopting the core network element state diagnosis method based on dynamic AI algorithm in mobile communication networks, the problem that traditional methods are difficult to cope with complex network environments is solved, and the refined monitoring and fault prevention of core network element states is achieved.

CN120110876AActive Publication Date: 2025-06-06HANGZHOU EASTCOM SOFTWARE TECH
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510291049.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-06
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

In current mobile communication networks, due to data dispersion, incomplete perception of fault elements, and low fault analysis and positioning efficiency, traditional threshold monitoring methods based on expert experience are difficult to effectively deal with complex and changeable network environments.

Method used

The core network element state diagnosis method based on dynamic AI algorithm is adopted, and real-time state data is obtained by connecting with the data source of the third-party platform, an isolated forest model is established to calculate outlier scores, dynamic security thresholds are output, abnormal information in the device log and alarms are identified, and the final health score of the network element is obtained, which triggers the fault event processing process.

Benefits of technology

It realizes refined monitoring of the status of core network elements, improves fault prevention and accuracy, reduces emergency response needs, and ensures immediate visible and controllable problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120110876A_ABST
    Figure CN120110876A_ABST
Patent Text Reader

Abstract

The invention relates to the field of core network equipment operation and maintenance, and provides a core network element state diagnosis method and device based on a dynamic AI algorithm, and the method comprises the steps: carrying out the butt joint with a third-party platform data source, and obtaining the real-time state data of a core network element; outputting an abnormal point of the real-time performance index data of the network element based on the dynamic security threshold; for the detected abnormal log data and alarm information, outputting abnormal information of the abnormal log data and the alarm information; and obtaining a final health degree score of the network element based on the abnormal point, the abnormal log data, the abnormal information of the alarm information, the dial test result information and the complaint information, and triggering a hidden danger checking task. According to the method, the data are analyzed through the AI algorithm based on the performance, alarm and log multi-dimensional data, the network element state is finally obtained through the network element health degree comprehensive evaluation algorithm, network changes can be automatically adapted, the method has the characteristics of high intelligence and automation, and the fault detection accuracy and the positioning efficiency are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of core network equipment operation and maintenance, and proposes a core network element status diagnosis method and device based on a dynamic AI algorithm. Background Art

[0002] In recent years, with the widespread popularity of 5G mobile communication networks and the rapid development of technology, the field of mobile communications has ushered in unprecedented changes. In this process, traditional 2G / 4G networks coexist with emerging virtualized 5G networks, jointly building a complex and ever-changing network environment. However, this increase in complexity has also brought many challenges, among which the most significant problems are bottlenecks such as data scattered across different platforms, incomplete perception of faulty network elements, and low efficiency in fault analysis and location.

[0003] Currently, mobile operators generally adopt dynamic and static threshold monitoring methods based on expert experience to deal with the poor quality of Internet access services of communication equipment. However, with the expansion of mobile communication network scale and the evolution of technology, the network environment has become increasingly complex and changeable. The traditional threshold setting method or single indicator analysis method can no longer effectively meet the production support requirements.

[0004] In actual operation, this fixed threshold exposes many limitations. On the one hand, the core network has a huge fault perception system, and manual analysis of this data is a difficult task, which is time-consuming and prone to excessive warnings or omissions of critical faults; on the other hand, effective identification and positioning of problem nodes requires deep professional knowledge and practical experience, and the threshold is high, which poses a severe test for the front-line operation and maintenance team; finally, the network is dynamic and complex, and the communication network is in a state of continuous evolution, which requires sufficient flexibility to respond to emergencies and long-term trends rather than fixed thresholds. Establishing a comprehensive and accurate set of rules to cover all possible network states requires relying on expert experience, extremely high human cost investment, and often lags behind the speed of network development. Therefore, building a core network element status analysis and diagnosis method and device with a dynamic AI algorithm that can automatically adapt to network changes and has highly intelligent and automated characteristics has become an important issue that current operators need to solve urgently. Summary of the invention

[0005] In order to solve the above problems, the present invention discloses a core network element status diagnosis method based on a dynamic AI algorithm, comprising:

[0006] Connecting with the third-party platform data source to obtain real-time status data of core network elements; the real-time status data of core network elements includes real-time performance indicator data of network elements, equipment log data, network element alarm information, network element dialing result information and complaint information; the core network runs on the equipment;

[0007] An isolation forest model is established based on historical network element performance indicator data to calculate the outlier score of each performance indicator of each network element. Based on the outlier score, a dynamic safety threshold of the network element performance indicator is output; based on the dynamic safety threshold, an outlier point of the network element real-time performance indicator data is output;

[0008] Perform feature extraction on the processed device log data and alarm information, identify key information and patterns in the logs and alarms, determine whether the input device log data and alarm information are abnormal, and output abnormal information of the abnormal log data and alarm information for the detected abnormal log data and alarm information;

[0009] The final health score of the network element is obtained based on the abnormal points of the network element's real-time performance indicators output by the performance indicator analysis model, abnormal log data and abnormal information of alarm information, dial test result information and complaint information;

[0010] Based on the final health score of the network element, the fault event handling process is triggered.

[0011] The steps of triggering the fault event processing process based on the final health score of the network element specifically include:

[0012] The final health score of the network element has an initial score, and points are deducted based on the abnormal points of the network element real-time performance indicator data output by the isolation forest model, the abnormal information of the abnormal device log data and network element alarm information, the network element dialing result information and the complaint information;

[0013] If S1≥the final health score of the network element>S2, the first fault event processing task is started;

[0014] If S2≥the final health score of the network element>S3, the second fault event processing task is started;

[0015] If S3≥the final health score of the network element>S4, the third fault event processing task is started;

[0016] If Sn≥the final health score of the network element>0, the nth fault event processing task is started;

[0017] The S1, S2, S3, ..., Sn are all between the initial score of the final health score of the network element and 0.

[0018] Also includes:

[0019] Define the mapping relationship between different types of data sources and Doris tables; the data source types include network element real-time performance indicator data, equipment log data, alarm information, dial test result information and complaint information;

[0020] Based on the Doris table model mapping configuration, the data hierarchy is divided into the original data layer ODS, the detailed data layer DWD, the wide table data DWS, and the application layer ADS;

[0021] Connect with various data sources to obtain different types of data sources;

[0022] Collect data files provided by data sources and parse the data files into structured data streams;

[0023] The data stream is pushed to the Kafka distributed message queue, and based on Flink's real-time stream processing capabilities, the data is processed separately, and different data streams are pushed to the original data layer ODS, the detailed data layer DWD, the wide table data DWS, and the application layer ADS to perform data standardization and data aggregation processing;

[0024] The processed data is efficiently stored and managed based on Doris.

[0025] The step of outputting a dynamic safety threshold value of a network element performance indicator based on an outlier score specifically includes:

[0026] For the indicator data identified as normal, a normal indicator data collection A and a nearly abnormal indicator data collection B are extracted;

[0027] Perform mean processing on the normal indicator data set A to obtain the normal value a;

[0028] Compare the indicator data set B with the normal value a, and output the dynamic safety threshold of the network element performance indicator:

[0029] If the indicator data set B is greater than the normal value a, the maximum value of the indicator data set X is taken as the safety boundary value, and the dynamic safety threshold is obtained as "≤safety boundary value";

[0030] If the indicator data set B contains both indicator data greater than the normal value a and less than the normal value a, the maximum value of the indicators greater than the normal value a is taken as the safety boundary data, and the minimum value of the indicators less than the normal value a is taken as the safety boundary data, that is, the dynamic safety threshold is "≤ the maximum safety indicator value, and ≥ the minimum safety indicator value";

[0031] If the indicator data set B is all smaller than the normal value a, take the minimum value in the indicator data set X as the safety boundary value, that is, the dynamic safety threshold is "≥ minimum safety indicator value".

[0032] The step of outputting abnormal points of the real-time performance indicator data of the network element based on the dynamic safety threshold specifically includes: marking abnormal points of the indicator based on the performance indicator safety threshold provided by the isolation forest model. If the indicator value of a certain device is not within the dynamic safety threshold at a certain time point, one abnormal point is marked.

[0033] The step of deducting points based on abnormal points of the real-time performance indicators of the network element output by the performance indicator analysis model, abnormal log data, abnormal information of the alarm information, dialing test result information and complaint information specifically includes:

[0034] If no indicator is marked as an abnormal point within the preset n consecutive time points, no points will be deducted;

[0035] If only one indicator is marked as an abnormal point within the preset n consecutive time points, it is regarded as a "single point abnormality" and 2 points are deducted;

[0036] If only one indicator is marked with more than one abnormal point within the preset n consecutive time points, it is regarded as "continuous abnormality" and 5m points are deducted, where m is the number of abnormal points;

[0037] If multiple indicators have "single point anomalies" within the preset n consecutive time points, 2p points will be deducted, where p is the number of indicators;

[0038] If multiple indicators show "continuous abnormalities" within the preset n consecutive time points, 40 points will be deducted and it will be considered unhealthy.

[0039] The step of establishing an isolation forest model based on historical network element performance indicator data also includes:

[0040] Standardize the date format of historical network element performance indicator data to keep the time interval of all data consistent;

[0041] Determine whether there are missing timestamp points. If so, use linear interpolation to fill in the missing time points and their corresponding indicator data;

[0042] According to the starting time and the established time interval of the data, the data is sorted in ascending order of date to construct a complete and continuous time series;

[0043] Integrate historical network element performance indicator data into the newly constructed time series and reset the data index;

[0044] Get a new quality indicator dataset X;

[0045] Split the total quality indicator data X into n quality indicator data subsets according to the dimension of (network element, indicator);

[0046] Based on the subset of quality indicator data, an isolation forest model is built for each indicator.

[0047] In a second aspect, a core network element status diagnosis device based on a dynamic AI algorithm is disclosed, comprising:

[0048] The multivariate data integration module processes data from different data sources based on the Doris table model mapping configuration, and is used to perform data collection, data analysis, cleaning and other processing tasks, and store and manage data;

[0049] The performance indicator analysis module is used to build an isolation forest model, calculate the outlier score of each performance indicator of each network element, and output the dynamic safety threshold of the network element performance indicator;

[0050] The log alarm analysis module is used to learn the features of log and alarm data, identify the key information and patterns in logs and alarms, determine whether the input logs are abnormal, and output abnormal logs and alarm information;

[0051] The forward-looking warning module is used to predict faults by using a deduction system to score the health of network elements, and provide the health scores of network elements to the emergency response module as the triggering basis for the fault process response mechanism;

[0052] The emergency response module is used to respond to different processing processes based on the network element health score.

[0053] The present invention has the following advantages:

[0054] 1. The present invention collects data in real time and uses diversified data integration technology to integrate cross-type data sources, making data analysis more accurate and in-depth;

[0055] 2. The present invention uses the small isolation forest model to calculate the outlier score of each indicator of each network element, and uses the large qwen2.5 model to analyze log and alarm abnormal information. The large model is used to perform macro trend prediction and pattern recognition, while the small model focuses on the refined analysis of specific problems. It can handle more complex scenarios than a single small model analysis, provides multi-dimensional result interpretation, and enhances the flexibility and accuracy of decision-making;

[0056] 3. The present invention uses a prediction mechanism based on historical data and a comprehensive health assessment algorithm to significantly improve the effectiveness and accuracy of fault prevention and reduce emergency response requirements;

[0057] 4. The present invention uses real-time monitoring and emergency response mechanisms to achieve instant warnings based on health status, ensuring that problems are immediately visible and controllable. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 This is a flow chart of a method for diagnosing core network element status in one embodiment of the present invention;

[0059] Figure 2This is a structural diagram of a core network element status diagnosis device in one embodiment of the present invention;

[0060] Figure 3 A schematic diagram of a core network element status data integration process in one embodiment of the present invention;

[0061] Figure 4 A performance indicator data cleaning flow chart in one embodiment of the present invention;

[0062] Figure 5 It is a schematic diagram of the isolation forest model training of the core network element status diagnosis method in one embodiment of the present invention;

[0063] Figure 6 The figure is a schematic diagram of indicator abnormality diagnosis of a core network element status diagnosis method in one embodiment of the present invention. DETAILED DESCRIPTION

[0064] The best mode of implementation of the present invention is described below by way of example. It should be understood that the specific embodiments herein are used to explain the present invention in detail and should not be construed as limiting the present invention. It should be noted that various changes and modifications may be made under the premise of following the principles and core scope of the present invention, and these changes should all be deemed to be within the scope of protection of the present invention. In conjunction with the accompanying drawings, the specific implementation steps of the present invention are described in detail.

[0065] like Figure 1 FIG. 1 is a flow chart of a method for diagnosing the status of a core network element in an embodiment of the present invention, which specifically includes the following steps:

[0066] Step 1: Connect with the third-party platform data source to obtain real-time status data of core network elements;

[0067] In a specific embodiment, the third-party platform data source may be a core network maintenance data center.

[0068] In a specific embodiment, S1 includes the following steps:

[0069] S101: Obtain the Doris table model mapping configuration and define the mapping relationship between different types of data sources and the Doris table;

[0070] In a specific embodiment, the mapping relationship between different types of data sources and Doris tables includes information such as data source type, data source format, and the corresponding relationship between data fields and Doris table columns.

[0071] The data source types include network element real-time performance indicator data, equipment log data, network element alarm information, network element dialing result information and complaint information, etc.

[0072] S102: Based on the Doris table model mapping configuration, the data layer is divided into the original data layer ODS, the detailed data layer DWD, the wide table data DWS, and the application layer ADS according to the big data layered architecture.

[0073] In a specific embodiment, the data structure and granularity of the original data layer are consistent with the content provided by the original file; the detailed data layer is responsible for cleaning, conversion and integration operations to generate detailed data tables, mainly for standardization of data; the wide table data layer performs aggregation operations to generate wide table data to meet general business needs. The application layer is used to store data and provide query interfaces for other modules.

[0074] S103: Connect with various data sources to obtain different types of data sources;

[0075] In a specific embodiment, the docking method can be active or passive, depending on the characteristics of the data platform and the frequency of data updates.

[0076] S104: After collecting the data files provided by each data source, the data files are parsed into a structured data stream. The parsing process includes operations such as format recognition, field extraction, and data type conversion of the data files.

[0077] S105: Push the parsed data stream to the Kafka distributed message queue, and based on Flink's real-time stream processing capabilities, process the data separately and push different data streams to each data layer to perform data standardization, data aggregation and other processing tasks.

[0078] S106: After the data processing task is completed, efficient data storage and management are achieved based on Doris.

[0079] like Figure 3 As shown, it is a schematic diagram of the core network element status data integration process in one embodiment of the present invention. The sources of collected indicators are OMC equipment and third-party data platforms. The sources of indicator data can be HTTP Adapter, SNMP Trap Adapter and Syslog Adapter, among which HTTP adapter is used to receive real-time performance indicators, SNMP Trap processes alarm information, and Syslog processes device logs. These adapters can automatically push indicator data.

[0080] In a specific embodiment, the data source may also be an SFTP adapter, an SNMP adapter, an SSH adapter, a RESTful adapter, or a WebService adapter, wherein the SFTP adapter is used to obtain files from a remote server or device via a secure file transfer protocol, the SNMP adapter is used to obtain performance indicators from a network device in real time, and the SSH adapter is connected to a network device via a secure SSH protocol to execute commands and obtain data collected by command output. The RESTful adapter is used to obtain indicator data via HTTP requests, and the WebService adapter is used to initiate requests to the target system via a WebService interface to obtain the required monitoring data. In a specific embodiment, indicator data may be obtained by active collection methods such as initiating a connection, requesting a command, or sending a query operation.

[0081] In a specific embodiment, the indicator data may also be collected by issuing instructions.

[0082] In a specific embodiment, different types of data sources obtain data in different formats and types, so the data need to be processed separately to perform processing tasks such as data standardization and data aggregation.

[0083] Step 2: Establish an isolation forest model based on historical network element performance indicator data, calculate the outlier score of each performance indicator of each network element, and output the dynamic safety threshold of the network element performance indicator based on the outlier score; based on the dynamic safety threshold, output the outlier points of the network element real-time performance indicator data;

[0084] In a specific embodiment, the specific implementation steps of S2 are as follows:

[0085] S201. Construct a historical performance indicator data subset Xn based on the historical performance indicator data of the core network element:

[0086] Collect and analyze the historical indicator data of core network elements to extract the key original performance indicator data set Z; data set Z should include core information such as network element name, indicator name, corresponding indicator value and collection time.

[0087] The extracted original performance indicator data set Z is cleaned and processed, and after a series of standardization processes, the total performance indicator data set X is obtained;

[0088] Furthermore, due to multiple factors such as equipment hardware, software version and configuration, operating environment, application scenarios and requirements, maintenance and management, the values ​​of the same performance indicators in different network elements are different; therefore, it is believed that each indicator of each device needs to be analyzed one by one; therefore, the total performance indicator data set X needs to be split into n performance indicator data subsets according to the dimension of (network element, indicator) to form Xn (n=1, 2, 3...n).

[0089] In a specific embodiment, Figure 4 As shown, S201 includes the following steps:

[0090] S2011. Standardize the format and date of historical network element performance indicator data to ensure that the time interval of all data is consistent;

[0091] S2012: determine whether there are accurate timestamp points. If so, use linear interpolation to fill in the missing time points and their corresponding indicator data;

[0092] S2013. Sort the data in ascending order of date based on the start time of the data and the established time interval to construct a complete and continuous time series;

[0093] S2014, integrate the original data into the newly constructed time series and reset the data index;

[0094] S2015, obtain a new quality indicator data set X;

[0095] S2016. Split the total quality indicator data set X into n quality indicator data subsets Xn according to the dimension of (network element, indicator);

[0096] S202, constructing an isolation forest model based on the isolation forest algorithm and using the data subset Xn;

[0097] In a specific embodiment, the isolation forest algorithm can be roughly divided into two stages. The first stage requires training t isolated trees (iTree) to form an isolation forest (iForest). Then, each sample point is brought into each isolated tree in the forest, the average height is calculated, and then the outlier score of each sample point is calculated. The training process includes the following steps:

[0098] S2021. Split each performance indicator data set Xn into j sample points, i.e., Xn = {x n1 ,...,x nj}, randomly set the dimension d (feature), each sample point has d dimensions (features), that is Randomly select φ sample points from Xn as the sample subset Xn' and put them into the root node of the tree.

[0099] S2022. Randomly specify a dimension q (i.e., feature) from the d dimensions in the sample subset Xn', and randomly generate a cutting point p in the current node data. The cutting point is generated between the maximum value and the minimum value of the specified dimension in the current node data, i.e.

[0100] S2023. A hyperplane is generated based on this cutting point, and the data space of the current node is divided into two subspaces: the data less than p in the specified dimension is placed in the left child node of the current node, and the data greater than or equal to p is placed in the right child node of the current node.

[0101] S2024. Recursively perform steps S2022 and S2023 in the child node, and continuously construct new child nodes until there is only one data in the child node (no further cutting is possible) or the number of isolated data has reached a limited height.

[0102] S2025. Each subset Xn repeats the above steps until each Xn generates t isolated trees (iTree) to form n isolated forests (iForest).

[0103] like Figure 5 As shown, it is a training graph of an isolation forest in one embodiment of the present invention, wherein the points marked in black have the shortest path, are isolated points, and are closest to anomalies.

[0104] S203. Based on the generated n iForest models, output the outlier score of the performance indicator:

[0105] In a specific embodiment, for each data point x ni , let it traverse each iTree corresponding to iForest and calculate the point x ni The average height h in the corresponding forest n (x ni ), normalize the average height of all points, and calculate the outlier score of the performance indicator.

[0106] The outlier score is calculated as follows:

[0107]

[0108] in,

[0109]

[0110] E(h n (x ni )) is the data point x ni The average path length among all isolation forests, c(φ) is the normalization factor of the expected path length.

[0111] Anomaly score The value range is (0,1). The closer the score is to 1, the more likely the data point is an outlier; the closer the score is to 0, the more likely the data point is a normal point.

[0112] S204. Finally, based on the outlier score, the dynamic safety threshold of the network element performance indicator is output, and the outlier point of the network element real-time performance indicator data is output.

[0113] In a specific embodiment, the specific implementation steps of S204 are as follows:

[0114] S2041, separately saving the indicator data identified as normal (i.e., the indicator data with an abnormality score close to 0), and extracting a normal indicator data collection A and a nearly abnormal indicator data collection B;

[0115] S2042, performing mean processing on the normal indicator data set A to obtain a normal value a;

[0116] S2043. Compare the indicator data set B with the normal value a, and output the dynamic safety threshold of the network element performance indicator:

[0117] If the indicator data set B is greater than the normal value a, the maximum value of the indicator data set B is taken as the safety boundary value, and the dynamic safety threshold is obtained as "≤safety boundary value";

[0118] If the indicator data set B contains both indicator data greater than the normal value a and less than the normal value a, the maximum value of the indicators greater than the normal value a is taken as the safety boundary data, and the minimum value of the indicators less than the normal value a is taken as the safety boundary data, that is, the dynamic safety threshold is "≤ the maximum safety indicator value, and ≥ the minimum safety indicator value";

[0119] If the indicator data set B is all less than the normal value a, take the minimum value in the indicator data set X as the safety boundary value, that is, the dynamic safety threshold is "≥ minimum safety indicator value"

[0120] S2044. Anomaly points are marked based on the performance indicator safety threshold provided by the isolation forest model. If the indicator value of a device is not within the safety threshold at a certain time point, an anomaly point is marked.

[0121] Step 3: Extract features from the device log data and alarm information, identify key information and patterns in the logs and alarms, determine whether the input device log data and alarm information are abnormal, and output abnormal information of the abnormal log data and alarm information for the detected abnormal log data and alarm information;

[0122] In a specific embodiment, S3 includes the following steps:

[0123] S301, collect real-time syslog log data and alarm information, and perform cleaning processing such as removing noise and invalid data on the log and alarm data; extract key information such as timestamp, event type, event description, etc. in the log and perform structured processing.

[0124] S302: Label the training log data and use the labeled log and alarm data to train the Qwen2.5 model. During the training process, the model will learn the characteristics and patterns of normal logs and alarms, and improve the accuracy and generalization ability of the model by adjusting the model parameters and optimizing the algorithm.

[0125] S303, input the real-time syslog logs and alarms to be identified into the trained Qwen2.5 model, the model will extract features from the input log and alarm data, identify the key information and patterns in the logs and alarms, and judge whether the input logs are abnormal based on the learned normal log and alarm features and patterns. If the logs and alarms are significantly different from the normal pattern, they will be marked as abnormal logs and alarms.

[0126] S304. For detected abnormal logs and alarms, the model will output detailed abnormal information in the form of alarm triggering, including network element abnormal conditions (such as equipment failure, performance degradation, etc.), abnormality levels (such as severe, general, etc.), impact scope (such as which services or systems are affected) and impact on services (such as causing business interruption, service degradation, etc.) and other related content.

[0127] Step 4: Based on the abnormal points of the network element real-time performance indicators output by the performance indicator analysis model, abnormal information of abnormal log data and alarm information, dial test result information and complaint information, the final health score of the network element is obtained;

[0128] In a specific embodiment, a deduction system is used for fault prediction, and the final health score of the network element is a full score of 100 points, and the deduction ends when all the points are deducted, and no negative number can appear;

[0129] The final health score of the network element is mainly related to five dimensions: performance indicators, golden indicators, alarms, logs, complaints, etc. Weights are set based on each dimension, and the final health score of the network element is obtained by combining the weights.

[0130] In a specific embodiment, the following deduction system may be used to calculate the final health score of the network element.

[0131] For performance indicator data, such as Figure 6 The figure shows a schematic diagram of indicator abnormality diagnosis. If no indicator of a device is marked with abnormal points within the preset n consecutive time points, it is considered "no abnormality" and no points are deducted for this indicator;

[0132] If a device has only one indicator marked as an abnormal point within the preset n consecutive time points, it is considered a "single point abnormality" and will be scored -2 ​​points;

[0133] If a device has only one indicator marked with more than one abnormal point within the preset n consecutive time points, it is considered "continuous abnormality" and -5m points (m is the number of abnormal points);

[0134] If a device has "single-point anomalies" in multiple indicators within the preset n consecutive time points, then -2p points (p is the number of indicators);

[0135] If a device has "continuous abnormalities" in multiple indicators within the preset n consecutive time points, it will be directly deducted 40 points and considered unhealthy.

[0136] For abnormal logs and alarm information, based on the alarm information extracted by qwen2.5 model training, the first-level alarm business node is -5 points, and the second-level alarm business node is -3 points. Different deduction points are given based on the abnormal log data screened out and combined with the analysis of the large model.

[0137] For dial test information, if the dial test fails more than 5 times within 15 minutes, the business node will be directly deducted 30 points, and 6 points will be deducted for each failure within 15 minutes.

[0138] Regarding complaint information, the number of complaints in the area covered by the network element has increased, with more than 100 cases, and the trend is rising. It is very easy to reproduce through dialing tests, and is concentrated in the area covered by the faulty network element. It is directly deducted by 40 points and is considered unhealthy.

[0139] In summary, the scores and weights of the five aspects are provided to the emergency response module as the basis for triggering the fault process response mechanism.

[0140] It should be noted that the above algorithm and deduction rules are only used as an example and do not limit the flexibility and scalability of the solution of the present invention.

[0141] S5. Based on the final health score of the network element, a fault event processing task is triggered.

[0142] The final health score of the network element has an initial score, and points are deducted based on the abnormal points of the network element real-time performance indicator data output by the isolation forest model, the abnormal information of the abnormal device log data and network element alarm information, the network element dialing result information and the complaint information;

[0143] If S1≥the final health score of the network element>S2, the first fault event processing task is started;

[0144] If S2≥the final health score of the network element>S3, the second fault event processing task is started;

[0145] If S3≥the final health score of the network element>S4, the third fault event processing task is started;

[0146] If Sn≥the final health score of the network element>0, the nth fault event processing task is started;

[0147] The S1, S2, S3, ..., Sn are all between the initial score of the final health score of the network element and 0.

[0148] In a possible embodiment, the initial score of the final health score of the network element is 100 points. When the final health score of the network element is in [70, 85), it is determined as a hidden danger event, and the corresponding hidden danger troubleshooting task is triggered immediately;

[0149] If the final health score of the network element is in [60,70), it is considered a major fault event and the fault event handling process is immediately initiated;

[0150] If the final health score of the network element is in [0,60), it is judged as a major fault event and triggers the highest priority fault event processing task.

[0151] The present invention also discloses a core network element status analysis and diagnosis device based on a dynamic AI algorithm, such as Figure 2 As shown, specifically including:

[0152] The multivariate data integration module processes data from different data sources based on the Doris table model mapping configuration, and is used to perform data collection, data analysis, cleaning and other processing tasks, and store and manage data;

[0153] The performance indicator analysis module is used to build an isolation forest model, calculate the outlier score of each performance indicator of each network element, and output the dynamic safety threshold of the network element performance indicator;

[0154] The log alarm analysis module is used to learn the features of log and alarm data, identify the key information and patterns in logs and alarms, determine whether the input logs are abnormal, and output abnormal logs and alarm information;

[0155] The forward-looking warning module is used to predict faults by using a deduction system to score the health of network elements, and provide the health scores of network elements to the emergency response module as the triggering basis for the fault process response mechanism;

[0156] The emergency response module is used to respond to different processing processes based on the network element health score.

[0157] The present invention proposes a core network element status analysis and diagnosis method and device based on a dynamic AI algorithm, which calculates the health of the network element, as well as the real-time monitoring and emergency response mechanism through massive data real-time and diversified data integration technology, dual model collaboration strategy, forward-looking warning algorithm and comprehensive health assessment algorithm, to achieve refined monitoring of the status of core network network elements and automatic and accurate capture of fault events.

[0158] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A core network element status diagnosis method based on dynamic AI algorithm, characterized in that: include: Connect with third-party platform data sources to obtain real-time status data of core network elements; The core network element real-time status data includes network element real-time performance indicator data, equipment log data, network element alarm information, network element dialing result information and complaint information; The core network element runs on the device; An isolation forest model is established based on historical network element performance indicator data to calculate the outlier score of each performance indicator of the core network element, and the dynamic safety threshold of the network element performance indicator is output based on the outlier score; Based on dynamic security thresholds, output abnormal points of network element real-time performance indicator data; Extract features from the device log data and network element alarm information, identify key information and patterns in the device log data and network element alarm information, determine whether the input device log data and network element alarm information are abnormal, and output abnormal information of the abnormal device log data and network element alarm information for the detected abnormal device log data and network element alarm information; The final health score of the network element is obtained based on the abnormal points of the network element real-time performance indicator data output by the isolation forest model, the abnormal information of the abnormal device log data and the network element alarm information, the network element dialing result information and the complaint information; Based on the final health score of the network element, the fault event processing task is triggered.

2. The method according to claim 1, characterized in that The steps of triggering the fault event processing task based on the final health score of the network element specifically include: The final health score of the network element has an initial score, and points are deducted based on the abnormal points of the network element real-time performance indicator data output by the isolation forest model, the abnormal information of the abnormal device log data and network element alarm information, the network element dialing result information and the complaint information; If S1≥the final health score of the network element>S2, the first fault event processing task is started; If S2≥the final health score of the network element>S3, the second fault event processing task is started; If S3≥the final health score of the network element>S4, the third fault event processing task is started; If Sn≥the final health score of the network element>0, the nth fault event processing task is started; The S1, S2, S3, and Sn are all between the initial score of the final health score of the network element and 0.

3. The method according to claim 1, characterized in that Also includes: Define the mapping relationship between different types of data sources and Doris tables; The data source types include network element real-time performance indicator data, equipment log data, network element alarm information, network element dialing result information and complaint information; Based on the Doris table model mapping configuration, the data hierarchy is divided into the original data layer ODS, the detailed data layer DWD, the wide table data DWS, and the application layer ADS; Connect with various data sources to obtain different types of data sources; Collect data files provided by data sources and parse the data files into structured data streams; The data stream is pushed to the Kafka distributed message queue, and based on Flink's real-time stream processing capabilities, the data is processed separately, and different data streams are pushed to the original data layer ODS, the detailed data layer DWD, the wide table data DWS, and the application layer ADS to perform data standardization and data aggregation processing; The processed data is stored and managed based on Doris.

4. The method according to claim 1, characterized in that The step of outputting a dynamic safety threshold value of a network element performance indicator based on an outlier score specifically includes: For the indicator data identified as normal, a normal indicator data collection A and a nearly abnormal indicator data collection B are extracted; Perform mean processing on the normal indicator data set A to obtain the normal value a; Compare the indicator data set B with the normal value a, and output the dynamic safety threshold of the network element performance indicator: If the indicator data set B is greater than the normal value a, the maximum value of the indicator data set X is taken as the safety boundary value, and the dynamic safety threshold is ≤ the safety boundary value; If the indicator data set B contains both indicator data greater than the normal value a and less than the normal value a, the maximum value of the indicators greater than the normal value a is taken as the safety boundary data, and the minimum value of the indicators less than the normal value a is taken as the safety boundary data, that is, the dynamic safety threshold is ≤ the maximum safety indicator value and ≥ the minimum safety indicator value; If the indicator data set B is all smaller than the normal value a, the minimum value in the indicator data set X is taken as the safety boundary value, that is, the dynamic safety threshold is ≥ the minimum safety indicator value.

5. The method according to claim 1, characterized in that: The step of outputting abnormal points of network element real-time performance indicator data based on the dynamic security threshold specifically includes: The outlier points of the indicators are marked based on the performance indicator safety threshold provided by the isolation forest model. If the indicator value of a device is not within the dynamic safety threshold at a certain time point, an outlier point is marked.

6. The method according to claim 2, characterized in that The step of deducting points based on the abnormal points of the network element real-time performance indicator data output by the isolation forest model, the abnormal information of the abnormal device log data and the network element alarm information, the network element dialing result information and the complaint information specifically includes: If no indicator is marked as an abnormal point within the preset n consecutive time points, no points will be deducted; If only one indicator is marked as an abnormal point within the preset n consecutive time points, it is regarded as a single point abnormality and 2 points are deducted; If only one indicator is marked with more than one abnormal point within the preset n consecutive time points, it is regarded as a continuous abnormality and 5m points are deducted, where m is the number of abnormal points; If a single abnormality occurs in multiple indicators within the preset n consecutive time points, 2p points will be deducted, where p is the number of indicators; If multiple indicators are abnormal for consecutive n preset time points, 40 points will be deducted and the system will be considered unhealthy.

7. The method according to claim 1, characterized in that The step of establishing an isolation forest model based on historical network element performance indicator data also includes: Standardize the date format of historical network element performance indicator data to keep the time interval of all data consistent; Determine whether there are missing timestamp points. If so, use linear interpolation to fill in the missing time points and their corresponding indicator data; According to the starting time and the established time interval of the data, the data is sorted in ascending order of date to construct a complete and continuous time series; Integrate the historical network element performance indicator data into the newly constructed time series and reset the data index to obtain a new quality indicator data set X; Split the total quality indicator data X into n quality indicator data subsets according to the dimensions of network elements and indicators; Based on the subset of quality indicator data, an isolation forest model is built for each indicator.

8. A core network element status analysis and diagnosis device based on dynamic AI algorithm, characterized in that: The method for implementing any one of the above methods 1-7 includes: The multivariate data integration module processes data from different data sources based on the Doris table model mapping configuration, and is used to perform data collection, data analysis, cleaning and other processing tasks, and store and manage data; The performance indicator analysis module is used to build an isolation forest model, calculate the outlier score of each performance indicator of each network element, and output the dynamic safety threshold of the network element performance indicator; The log alarm analysis module is used to learn the features of log and alarm data, identify the key information and patterns in logs and alarms, determine whether the input logs are abnormal, and output abnormal logs and alarm information; The forward-looking warning module is used to predict faults by using a deduction system to score the health of network elements, and provide the health scores of network elements to the emergency response module as the triggering basis for the fault process response mechanism; The emergency response module is used to respond to different processing processes based on the network element health score.

Citation Information

Patent Citations

  • Automatic network element abnormal behavior detection method and device

    CN105871879A

  • Defense method for federal learning poisoning attack based on isolated forest

    CN114565106A

  • Intelligent operation and maintenance management control method and system and storage medium

    CN115865649A

  • Intelligent guarding method and device, electronic equipment and storage medium

    CN115996414A

  • Fault network element identification method and device

    CN116112963A