Intelligent value watch method and device, electronic equipment and storage medium
By collecting and aggregating monitoring data from core network element devices and using a device health assessment model for status prediction, the problem of low efficiency and security risks in operator network element cutover monitoring in existing technologies has been solved, achieving fully automated device monitoring and accurate anomaly detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA MOBILE GROUP JIANGSU
- Filing Date
- 2021-10-20
- Publication Date
- 2026-07-21
AI Technical Summary
In existing technologies, the manual operation of network element cutover monitoring by operators is inefficient and costly. Furthermore, semi-automated systems cannot achieve fine-grained and real-time monitoring, posing network security risks.
Collect monitoring data from core network elements, generate monitoring factors through data aggregation, use equipment health assessment models for status prediction, and combine fuzzification processing to achieve fully automated cutover monitoring.
It enables fine-grained and real-time monitoring of devices, reduces operation and maintenance costs, improves the accuracy of anomaly detection and security, and enhances the automation of network maintenance.
Smart Images

Figure CN115996414B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communications, and more particularly to an intelligent monitoring method and apparatus, electronic device and storage medium. Background Technology
[0002] With the rapid development of mobile 5G, 5G construction has entered its second peak period of large-scale deployment. To support 4G / 5G convergence, the upgrading and transformation of existing networks are proceeding at a rapid pace. However, frequent network adjustments not only bring a massive workload to maintenance personnel but also pose risks to network security. Therefore, automated and intelligent security verification and monitoring are particularly urgent. How to efficiently operate and maintain multiple networks of different standards, how to use artificial intelligence for unified health management of the core network and peripheral equipment, how to make maintenance work readily accessible and convenient, and how to continuously reduce operation and maintenance costs while achieving energy conservation and emission reduction are currently major issues in network work. In existing technologies, post-operator network element cutover monitoring mainly uses the following two methods:
[0003] Manual monitoring not only requires a significant amount of manpower for basic tasks such as checking metrics, performance analysis, and service testing, but also carries the risk of incomplete service coverage during testing. Frequent network adjustments impose a massive workload on maintenance personnel. Relying on manual monitoring solutions results in high maintenance costs and low efficiency. Furthermore, monitoring is limited by employee experience, and insufficient experience or errors can lead to an inability to quickly and effectively detect network anomalies, posing risks to network security.
[0004] The semi-automatic monitoring system inputs the performance difference of network elements before and after a cutover into an anomaly detection model trained on an autoencoder neural network, and obtains the output difference reconstruction data. It then compares the reconstruction error between the performance difference data and the reconstructed difference data with an error threshold to determine whether the network element is in an abnormal operating state after the cutover. Because it uses deep learning to compare pre- and post-cutover performance data for monitoring, it improves the accuracy of determining whether network elements are operating abnormally compared to manual monitoring, and reduces the false alarm rate to some extent. However, this monitoring method only targets data at a specific moment before and after the cutover, and cannot achieve fine-grained and real-time monitoring. Furthermore, anomalies in certain operating parameters do not necessarily indicate an anomaly in the overall state of the existing network equipment. If multiple operating parameters are abnormal, further manual assessment of the fault condition of the existing network equipment is still required, thus failing to achieve automated monitoring and exhibiting low accuracy. Summary of the Invention
[0005] This invention provides an intelligent monitoring method and device, electronic device and storage medium to solve the technical defects existing in the prior art.
[0006] This invention provides an intelligent monitoring method, comprising:
[0007] Collect monitoring data from core network element devices, including monitoring reports;
[0008] The monitored data is aggregated to obtain the monitored factor;
[0009] The monitoring factors are input into the equipment health assessment model, and the status prediction results of the core network element equipment are output.
[0010] The equipment health assessment model is obtained by training based on the monitored sample factor data and pre-determined health labels.
[0011] According to the intelligent monitoring method of the present invention, the collection of monitoring data of core network element devices includes:
[0012] Customized monitoring tasks, including: dial-up testing tasks, high-frequency monitoring tasks, and periodic monitoring tasks;
[0013] Based on the duty tasks, collect the duty data of core network element devices.
[0014] According to the intelligent monitoring method of the present invention, the collection of monitoring data of core network element devices includes:
[0015] Obtain the duty reports of core network element devices;
[0016] Based on the equipment specifications, selectively collect equipment specifications.
[0017] Based on the device alarm status, selectively collect device alarm information.
[0018] According to the intelligent monitoring method of the present invention, the selective collection of equipment indicators based on equipment indicator status and the selective collection of equipment alarm information based on equipment alarm status include:
[0019] After the duty task is executed, the system automatically checks whether there are equipment indicators in the duty report;
[0020] If equipment specifications exist, then collect the equipment specifications.
[0021] Check if there are any equipment alarms in the duty report; if there are equipment alarms, collect the equipment alarm information.
[0022] According to the intelligent monitoring method of the present invention, after inputting the monitoring factor into the equipment health assessment model and outputting the state prediction result of the core network element equipment, the process includes:
[0023] The names of the core network element devices are obfuscated.
[0024] According to the intelligent monitoring method of the present invention, the equipment health assessment model is obtained by training based on monitoring sample factor data and pre-determined health labels, and includes:
[0025] A health index is constructed based on the impact of various monitoring factors of core network elements on health. The corresponding health level of core network elements is determined based on the health index. The various monitoring factors of the core network elements are used as monitoring sample factor data, and the health level is used as a pre-determined health label.
[0026] The loss function of the health assessment model is:
[0027]
[0028] Among them, w k Let C represent the result obtained in this iteration, C represent a vector, and m represent the number of guard sample factor data in this training.
[0029] According to the intelligent monitoring method of the present invention, the method further includes:
[0030] The accuracy of the state prediction results of the core network element devices is obtained by testing the output state prediction results using a sliding window.
[0031] The present invention also provides an intelligent monitoring device, comprising:
[0032] The duty data acquisition module is used to collect duty data from core network element devices, including duty reports.
[0033] The data aggregation module is used to aggregate the guard data to obtain the guard factor;
[0034] The duty prediction module is used to input the duty factors into the equipment health assessment model and output the status prediction results of the core network element equipment;
[0035] The equipment health assessment model is obtained by training based on the monitored sample factor data and pre-determined health labels.
[0036] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the above-described intelligent monitoring methods.
[0037] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described intelligent watchkeeping methods.
[0038] This invention, based on monitoring data from core network elements, enables fully automated cutover monitoring, significantly improving execution efficiency while reducing maintenance costs. It allows for fine-grained and real-time monitoring of devices, enabling customized monitoring solutions based on different needs. Utilizing a device health assessment model, it achieves comprehensive multi-dimensional analysis of anomalies in the overall status of existing network devices, resulting in higher accuracy, stronger real-time performance, and enhanced security. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0040] Figure 1 This is a flowchart illustrating the intelligent monitoring method provided by the present invention;
[0041] Figure 2 This is a schematic diagram of the intelligent monitoring device provided by the present invention;
[0042] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0044] The following is combined Figure 1 The present invention describes an intelligent monitoring method, the method comprising:
[0045] S1. Collect monitoring data from core network element devices, including monitoring reports;
[0046] Core network elements include related equipment and peripheral equipment in the core network. The duty report itself has many inspection items, including indicators and alarms, so it is necessary to aggregate the data in the duty report.
[0047] S2. Aggregate the guarded data to obtain the guarded factor;
[0048] Based on the test execution time, collect test execution time period data and subsequent batches of equipment indicator data and equipment alarm data. Aggregate this data into a timeline format using timestamps. Use the C4.5 algorithm to construct a classifier in the form of a decision tree for data convergence. (Timestamps are a standard time format for all data types; data of each data type forms an "event time series" based on the order of timestamps. Device topology association determines logical relationships through topological relationships, such as related devices, directly connected devices, and business-related devices. This guides which data containing timestamps are aggregated into an "event time series.") Specific steps may include:
[0049] (1) Pattern / Ontology Alignment: The inherent hardware and software attributes (name, type, value range) of network elements and the adjacency relationships between attributes are used to find the correspondence between source patterns and intermediate patterns, and to determine the topological and service relationships between network elements. Furthermore, a general model of the core network equipment classifier is established based on the type of network element using the C4.5 algorithm, and data from different data sources is uniformly labeled and classified in the form of a decision tree.
[0050] (2) Entity Linking: The key lies in entity identification, mainly identifying similar entities (similar: multiple named entities can correspond to one real entity) and eliminating entity ambiguity (one entity can correspond to multiple real entities). This technology performs uniqueness processing in a unified standard for network element devices, eliminating ambiguity in description. For all names of the same entity, an additional identifier is added, and the identifier uses UUID to ensure uniqueness.
[0051] (3) Conflict resolution: Identify the correct value from all conflicts, mainly to resolve data differences caused by the differences in sampling time of various types of data collection.
[0052] For example, performance metrics are typically collected directly from counters on the devices, resulting in 5-minute granularity statistical files. These files are then aggregated at the network management system, creating 15-minute granularity statistical files. This can lead to discrepancies in KPI calculations when both data sources are at 15:00:00, as one uses 5-minute statistics and the other uses 15-minute statistics. This step unifies the performance metric data from both data sources.
[0053] (4) Relationship deduction: The isolated slice data is merged using the C4.5 algorithm to achieve data convergence. The test report, equipment indicators, equipment alarm information, and various types of inspection data are unified and associated according to equipment topology relationships, event time sequence, etc. For example, the associated data is identified according to the time sequence. If the equipment indicators are statistically analyzed in 15-minute intervals, then the alarms are alarms from the same or similar times.
[0054] For example, in a large number of alarms, we can deduce the initial alarm: based on the inherent correlation of alarms and the order of their times, we can find the initial alarm, which is the cause of a series of subsequent problems. For instance, our call testing data may affect call traffic and services. We can correlate performance indicators exceeding limits and alarms within a certain range based on the device or time of the call test.
[0055] The frequency of use for different data types can be dynamically adjusted as needed, for example:
[0056] Detail: For network elements with deteriorating indicators, add automatic inspection and dynamically increase the detail of the inspection. For example, if only 5 items are inspected, expand to inspect 20 items.
[0057] Multi-scenario: For example, in equipment indicator monitoring, the threshold values of the indicators can be automatically matched for multiple scenarios, such as major holidays, school opening / closing, and other specific application scenarios, and different threshold schemes can be adopted.
[0058] Correlation: For example, after the success rate metric deteriorates, extract the relevant user count from the inspection. It may be a normal deterioration caused by a decrease in the number of users, so exclude low-level and minor alarms.
[0059] S3. Input the monitoring factor into the equipment health assessment model and output the status prediction results of the core network element equipment;
[0060] Because the system adopts a comprehensive data collection strategy, data collection and statistical analysis are conducted as thoroughly as possible from various dimensions during actual construction. To address the challenges of fragmentation, independence, and consistency arising from this, the system employs data fusion to establish multi-dimensional and multi-granular relationships between data, information, and knowledge fragments. This enables more layers of information interaction, thereby converging the intrinsic relationships between data and abstracting the relationships between data to establish a unique evaluation system. The core of this evaluation system is the equipment health assessment model. Using the equipment health assessment model, the state prediction results of core network elements can be obtained.
[0061] The equipment health assessment model is obtained by training based on the monitored sample factor data and pre-determined health labels.
[0062] This invention, based on monitoring data from core network elements, enables fully automated cutover monitoring, significantly improving execution efficiency while reducing maintenance costs. It allows for fine-grained and real-time monitoring of devices, enabling customized monitoring solutions based on different needs. Utilizing a device health assessment model, it achieves comprehensive multi-dimensional analysis of anomalies in the overall status of existing network devices, resulting in higher accuracy, stronger real-time performance, and enhanced security.
[0063] This invention establishes a complete equipment health assessment system for centralized data processing and analysis. The assessment presents equipment status as a percentage, highlighting the severity of problems through the percentage. The importance assessment of problems is flexibly defined using a combination of fixed and dynamic percentages. The fixed percentage employs a forward calculation method based on a built-in network structure model, directly transmitting the state of the input unit to the computation unit for weighted and biased calculations. The dynamic percentage uses an AI backpropagation optimization algorithm, leveraging the chain rule to calculate the derivatives of two or more conformal functions, backpropagating the gradient from the computation unit back to the input unit. Based on the calculated gradient, the learnable parameters of the network are adjusted to achieve real-time dynamic assessment of the core network's overall status. This dynamic percentage adjustment scheme can comprehensively capture all actual problems, further refining the granularity of problems and providing favorable support for the correlation of problems.
[0064] The equipment health assessment model refers to the specific implementation, including a series of behaviors such as model training, model use, and model feedback. The equipment health assessment system is constructive and guiding for the equipment health assessment model. The equipment health assessment system can guide the behavior of the equipment health assessment model, such as how the model should be built, how the model's data should be processed, and how the output should be calculated.
[0065] According to the intelligent monitoring method of the present invention, the collection of monitoring data of core network element devices includes:
[0066] Customized monitoring tasks, including: dial-up testing tasks, high-frequency monitoring tasks, and periodic monitoring tasks;
[0067] Based on the duty tasks, collect the duty data of core network element devices.
[0068] High-frequency monitoring tasks refer to periodic slice inspections performed at fixed intervals, such as 5 minutes, 15 minutes, 30 minutes, 1 hour, etc., without relying on the dialing system.
[0069] Regularly scheduled inspection tasks refer to conducting slice inspections at fixed times, such as 1 hour, 3 hours, and 5 hours after a network cutover or upgrade. This method tends to be more comprehensive and involves a single, long-term inspection with complex inspection logic.
[0070] The duty report includes both test results and inspection results. Specifically, when performing test tasks, the duty report is the test result, and when performing high-frequency duty tasks and regular duty tasks, the duty report is the inspection result.
[0071] Inspection refers to the centralized checking of equipment status, health, and whether it is working properly, which can be queried through commands.
[0072] "Monitoring" describes the entire process. Inspections can be considered one data source, while testing, performance indicators, and equipment alarms are different data sources. Monitoring continuously collects this data, provides feedback on problems, and achieves the goal of automated monitoring.
[0073] According to the intelligent monitoring method of the present invention, the collection of monitoring data of core network element devices includes:
[0074] Obtain the duty reports of core network element devices;
[0075] Based on the equipment specifications, selectively collect equipment specifications.
[0076] Equipment metrics refer to the various performance indicators of related equipment in the core network, such as UDC, which includes multiple performance indicators such as HLR-FE, HSS-FE, CUDB, and PG.
[0077] Based on the device alarm status, selectively collect device alarm information.
[0078] Equipment alarms refer to the process of setting thresholds for various equipment operating parameters to detect system anomalies and generate corresponding equipment alarms.
[0079] According to the intelligent monitoring method of the present invention, the selective collection of equipment indicators based on equipment indicator status and the selective collection of equipment alarm information based on equipment alarm status include:
[0080] After the duty task is executed, the system automatically checks whether there are equipment indicators in the duty report;
[0081] If equipment metrics exist, collect them. Duty reports are necessary, but equipment metrics and alarms are not. If the current report contains equipment metrics, proceed to the next step; otherwise, perform data aggregation directly.
[0082] The system checks whether there are any device alarms in the duty report; if there are device alarms, it collects the device alarm information. If there are no device alarms, it directly performs data aggregation.
[0083] According to the intelligent monitoring method of the present invention, after inputting the monitoring factor into the equipment health assessment model and outputting the state prediction result of the core network element equipment, the process includes:
[0084] The names of the core network element devices are obfuscated.
[0085] Network information security is a critical aspect of network maintenance, and security is a key consideration and comprehensive assessment during system construction. This system utilizes fuzzy control technology to automatically perform fuzzification of network element names. This fuzzification process is unrelated to the health prediction mentioned above.
[0086] While retaining some identification information, the original naming logic is eliminated, and a completely new naming expression that only maintenance personnel can understand is adopted. The naming expression can be changed periodically or irregularly, and the fuzzification of network element names is automatically completed. The specific steps are as follows:
[0087] Step 101: Obfuscate the interface
[0088] Convert the real, definite input into a fuzzy vector.
[0089] For example, if we have network element names such as TEST001-000-01, TEST001-000-02, TEST002-000-01, and TEST002-000-02, we can obtain the generation rules for the network element names (e.g., TEST sequence number + 000 + sequence number). With the rules, we can perform fuzzy processing, thus converting a definite quantity into a vector.
[0090] Step 102: Knowledge Base Establishment
[0091] The knowledge base consists of two parts: a database and a rule base.
[0092] Database: Record all fuzzy subsets of all input and output variables, such as network element names, alarm names, and indicator values, into the database. This database is used to provide data to the inference engine during the process of solving the fuzzy relation equations in rule-based reasoning.
[0093] Rule base: Based on the expert knowledge base or the long-term experience accumulated by manual operators, a series of relational terms are established. The relational terms are "translated" with the help of the expert knowledge base, and finally the fuzzy rules are quantified.
[0094] For example, we propose a rule based on date-time-network element name-serial number. Date, network element name, etc., are relational terms, each of which is a separate matching item. Finally, we translate and combine each matching item to generate a numerical rule.
[0095] Step 103: Blur Processing
[0096] While retaining some identification information, the original naming logic is removed, and a brand-new naming expression that can only be understood by maintenance personnel is adopted. The network element naming expression can be changed periodically or irregularly, and the network element name fuzzification is automatically completed.
[0097] In one specific embodiment, the blurring process includes:
[0098] Step 1: Enter the network element name;
[0099] Step 2: Parse the network element name, extract the network element format, and perform matching;
[0100] Determine the identifier information of the user command [e; f; g; h].
[0101] User instruction information [e; f; g; h]: e represents the user's permission information, f represents the network element device health score, g represents the current data type, and h represents the beginning and end type information of the processed data.
[0102] Based on the identifier information e, the user's permission level is determined. When the user's level is greater than L, the data sent by the user can be obfuscated.
[0103] An assessment is conducted based on the health status of the network element device and the current data type. If the health status of the network element device (see above for the specific calculation method) is >80%, and the current data type is data or character, then proceed to the next step.
[0104] Determine the identifier information h. If d is a type that starts with a letter and ends with a letter, then perform semi-fuzzy processing (e.g., TEST001A). If d is a type that starts with a letter and ends with a number, then perform full fuzzy processing (e.g., TEST0011).
[0105] Step 3: If full fuzzy processing is performed, determine whether the last digit of the network element name is even. If it is even, select rule A; if it is odd, select rule B. If semi-fuzzy processing is performed, parse the network element name and separate the letter segment from the number segment. Select rule C for the letter segment and rule D for the number segment.
[0106] Rule A: Convert the entire data to BASE format and extract the first 5 characters; encrypt the entire data using MD5 and extract the first 5 characters; then concatenate the two results.
[0107] Rule B: Extract the numbers from the filename, multiply each number by 10 and add them together to get the result, then convert the result to hexadecimal; concatenate the letter segments of the filename and convert them to BASE format, then extract the first 5 characters and concatenate them into a hexadecimal result.
[0108] Rule C: Concatenate the ASCII codes (decimal) of each letter and convert the final result to hexadecimal.
[0109] Rule D: Calculate the sum of the squares of all numbers and perform a hexadecimal conversion.
[0110] According to the intelligent monitoring method of the present invention, the equipment health assessment model is obtained by training based on monitoring sample factor data and pre-determined health labels, and includes:
[0111] A health index is constructed based on the impact of various monitoring factors of core network elements on health. The corresponding health level of core network elements is determined based on the health index. The various monitoring factors of the core network elements are used as monitoring sample factor data, and the health level is used as a pre-determined health label.
[0112] The loss function of the health assessment model is:
[0113]
[0114] Among them, w k Let C represent the result obtained in this iteration, C represent a vector, and m represent the number of guard sample factor data in this training.
[0115] The model is evaluated using a loss function; a lower loss function value indicates better robustness. Once trained, the model can be used directly. However, if special cases arise, such as models not being covered during training, iterative optimization is necessary.
[0116] Based on the data aggregation and analysis results, an equipment health assessment system is constructed.
[0117] Construct a health assessment system for network element devices and obtain training samples.
[0118] First, based on factors affecting health such as core network element equipment performance indicators, equipment alarm information, inspection results, and testing results, a network element equipment health assessment system is constructed by progressively subdividing the system into four levels, as detailed below:
[0119] The first-level judgment criterion is the test report. The value of this level will be obtained based on the alarm situation. The lower the alarm value, the lower the starting value of this level.
[0120] The second level of judgment is based on the alarm status. The value for this level will be obtained based on the inspection results, and the standard will be divided according to the median value of the inspection results.
[0121] The third level of judgment is based on the inspection results, and the value of this level will be determined by the number of abnormalities in the inspection results.
[0122] The fourth level is judged by performance indicators. The closer the indicator value is to the threshold, the lower the score.
[0123] The scores of the health factors in the four levels mentioned above are related to the duty reports, network element status parameters, etc., and are determined based on experience to obtain the health level of a certain device. For example, if the scores of a certain device for the first, second, third, and fourth levels are 8, 5, 4, and 3 respectively, then the final health level is (8*a+5*b+4*c+3*d) / ((a+b+c+d)*10), which gives a final health index of 50% (the weights can be adjusted according to the actual situation; in this calculation, the weights are 1:1:1:1); a, b, c, and d represent the weights of the four items: (8*1+5*1+4*1+3*1) / (1+1+1+1)*10 gives a result of 50%; this result will be persistently saved and used as a sample in model training.
[0124] The health level of the aforementioned equipment (50%) is used to form a training sample along with the corresponding equipment indicators, alarm information, inspection results, and test results. This process is automated and requires no manual intervention. The only manual task is to adjust the weighting of each item based on the actual situation.
[0125] Specifically, an example of a network element device health assessment system is shown in Table 1 below:
[0126] Table 1
[0127]
[0128]
[0129] The structure of the network element health assessment model based on neural networks is as follows:
[0130] Data processing. Performance indicators: 271, standardized to between 0 and 1. Equipment alarm information and inspection items: 413, standardized to between 0 and 1. Test result items: 100, standardized to between 0 and 1.
[0131] The network structure is determined as follows: This embodiment adopts a three-layer neural network structure. The input layer has 784 nodes, corresponding to performance indicators, alarm / inspection items, and test results. The input layer passes the corresponding information to the hidden layer, which then passes the corresponding data to the output layer. The number of nodes in the hidden layer is adjusted according to the running efficiency, training iterations, and total network error. The output layer has 1 node, representing the health rating.
[0132] Model training:
[0133] The activation value of each neuron in the next layer is equal to the weighted sum of all activation values in the previous layer, plus some weight bias. The weight bias is 16+16+10, so 784*16+16*16+16*10=13002. Here, the health rating is based on a 10-level standard.
[0134] There are more than 13,002 weight bias values for the overall network health, which we can adjust to get infinitely closer to the true health status of the network as a whole.
[0135] In this embodiment, when training the deep neural network, all weights and biases are randomly initialized. The neural network is trained to optimize each weight and bias. When the error between the actual output and the expected output of the neural network is less than a preset error threshold, training stops, and a network element health assessment model is obtained; otherwise, the parameters of each layer are updated, and training is repeated. The algorithm that can be used in this "neural network" training process is "gradient descent". The transfer function is a hardlim threshold function, chosen because a threshold function is selected because a certain condition is met to output a specified value. The error calculation method and parameter update method are existing technologies and will not be elaborated upon further in this invention.
[0136] According to the intelligent monitoring method of the present invention, the method further includes:
[0137] The accuracy of the state prediction results of the core network element devices is obtained by testing the output state prediction results using a sliding window.
[0138] In actual training using supervised learning, a sliding window is considered for testing. Sliding windows are more suitable for processing time series data. Here, previous data (health status) is used as the input parameter, and the predicted data is used as the output parameter to monitor the accuracy of the prediction model's output.
[0139] See Figure 2 The intelligent monitoring device provided by the present invention is described below. The intelligent monitoring device described below can be referred to in correspondence with the intelligent monitoring method described above. The intelligent monitoring device includes:
[0140] The duty data acquisition module 10 is used to collect duty data from core network element devices, including duty reports.
[0141] Core network elements include related equipment and peripheral equipment in the core network. The duty report itself has many inspection items, including indicators and alarms, so it is necessary to aggregate the data in the duty report.
[0142] Data aggregation module 20 is used to aggregate the guard data to obtain guard factors;
[0143] The duty prediction module 30 is used to input the duty factor into the equipment health assessment model and output the status prediction results of the core network element equipment;
[0144] Based on the test execution time, collect test execution time period data and subsequent batches of equipment indicator data and equipment alarm data. Aggregate this data into a timeline format using timestamps. Use the C4.5 algorithm to construct a classifier in the form of a decision tree for data convergence. (Timestamps are a standard time format for all data types; data of each data type forms an "event time series" based on the order of timestamps. Device topology association determines logical relationships through topological relationships, such as related devices, directly connected devices, and business-related devices. This guides which data containing timestamps are aggregated into an "event time series.") Specific steps may include:
[0145] (1) Pattern / Ontology Alignment: The inherent hardware and software attributes (name, type, value range) of network elements and the adjacency relationships between attributes are used to find the correspondence between source patterns and intermediate patterns, and to determine the topological and service relationships between network elements. Furthermore, a general model of the core network equipment classifier is established based on the type of network element using the C4.5 algorithm, and data from different data sources is uniformly labeled and classified in the form of a decision tree.
[0146] (2) Entity Linking: The key lies in entity identification, mainly identifying similar entities (similar: multiple named entities can correspond to one real entity) and eliminating entity ambiguity (one entity can correspond to multiple real entities). This technology performs uniqueness processing in a unified standard for network element devices, eliminating ambiguity in description. For all names of the same entity, an additional identifier is added, and the identifier uses UUID to ensure uniqueness.
[0147] (3) Conflict resolution: Identify the correct value from all conflicts, mainly to resolve data differences caused by the differences in sampling time of various types of data collection.
[0148] For example, performance metrics are typically collected directly from counters on the devices, resulting in 5-minute granularity statistical files. These files are then aggregated at the network management system, creating 15-minute granularity statistical files. This can lead to discrepancies in KPI calculations when both data sources are at 15:00:00, as one uses 5-minute statistics and the other uses 15-minute statistics. This step unifies the performance metric data from both data sources.
[0149] (4) Relationship deduction: The isolated slice data is merged using the C4.5 algorithm to achieve data convergence. The test report, equipment indicators, equipment alarm information, and various types of inspection data are unified and associated according to equipment topology relationships, event time sequence, etc. For example, the associated data is identified according to the time sequence. If the equipment indicators are statistically analyzed in 15-minute intervals, then the alarms are alarms from the same or similar times.
[0150] For example, in a large number of alarms, we can deduce the initial alarm: based on the inherent correlation of alarms and the order of their times, we can find the initial alarm, which is the cause of a series of subsequent problems. For instance, our call testing data may affect call traffic and services. We can correlate performance indicators exceeding limits and alarms within a certain range based on the device or time of the call test.
[0151] The frequency of use for different data types can be dynamically adjusted as needed, for example:
[0152] Detail: For network elements with deteriorating indicators, add automatic inspection and dynamically increase the detail of the inspection. For example, if only 5 items are inspected, expand to inspect 20 items.
[0153] Multi-scenario: For example, in equipment indicator monitoring, the threshold values of the indicators can be automatically matched for multiple scenarios, such as major holidays, school opening / closing, and other specific application scenarios, and different threshold schemes can be adopted.
[0154] Correlation: For example, after the success rate metric deteriorates, extract the relevant user count from the inspection. It may be a normal deterioration caused by a decrease in the number of users, so exclude low-level and minor alarms.
[0155] Because the system adopts a comprehensive data collection strategy, data collection and statistical analysis are conducted as thoroughly as possible from various dimensions during actual construction. To address the challenges of fragmentation, independence, and consistency arising from this, the system employs data fusion to establish multi-dimensional and multi-granular relationships between data, information, and knowledge fragments. This enables more layers of information interaction, thereby converging the intrinsic relationships between data and abstracting the relationships between data to establish a unique evaluation system. The core of this evaluation system is the equipment health assessment model. Using the equipment health assessment model, the state prediction results of core network elements can be obtained.
[0156] The equipment health assessment model is obtained by training based on the monitored sample factor data and pre-determined health labels.
[0157] According to the intelligent monitoring device of the present invention, the monitoring data acquisition module 10 is used for:
[0158] Customized monitoring tasks, including: dial-up testing tasks, high-frequency monitoring tasks, and periodic monitoring tasks;
[0159] Based on the duty tasks, collect the duty data of core network element devices.
[0160] "Monitoring" describes the entire process. Inspections can be considered one data source, while testing, performance indicators, and equipment alarms are different data sources. Monitoring continuously collects this data, provides feedback on problems, and achieves the goal of automated monitoring.
[0161] According to the intelligent monitoring device of the present invention, the monitoring data acquisition module 10 is used for:
[0162] Obtain the duty reports of core network element devices;
[0163] Based on the equipment specifications, selectively collect equipment specifications.
[0164] Equipment metrics refer to the various performance indicators of related equipment in the core network, such as UDC, which includes multiple performance indicators such as HLR-FE, HSS-FE, CUDB, and PG.
[0165] Based on the device alarm status, selectively collect device alarm information.
[0166] According to the intelligent monitoring device of the present invention, the monitoring data acquisition module 10 is used for:
[0167] After the duty task is executed, the system automatically checks whether there are equipment indicators in the duty report;
[0168] If equipment metrics exist, collect them. Duty reports are necessary, but equipment metrics and alarms are not. If the current report contains equipment metrics, proceed to the next step; otherwise, perform data aggregation directly.
[0169] The system checks whether there are any device alarms in the duty report; if there are device alarms, it collects the device alarm information. If there are no device alarms, it directly performs data aggregation.
[0170] The intelligent monitoring device according to the present invention further includes a fuzzification processing module, the fuzzification processing module being used for:
[0171] The names of the core network element devices are obfuscated.
[0172] Network information security is a critical aspect of network maintenance, and security is a key consideration and comprehensive assessment during system construction. This system utilizes fuzzy control technology to automatically perform fuzzification of network element names. This fuzzification process is unrelated to the health prediction mentioned above.
[0173] While retaining some identification information, the original naming logic is eliminated, and a completely new naming expression that only maintenance personnel can understand is adopted. The naming expression can be changed periodically or irregularly, and the fuzzification of network element names is automatically completed. The specific steps are as follows:
[0174] According to the intelligent monitoring method of the present invention, the equipment health assessment model is obtained by training based on monitoring sample factor data and pre-determined health labels, and includes:
[0175] A health index is constructed based on the impact of various monitoring factors of core network elements on health. The corresponding health level of core network elements is determined based on the health index. The various monitoring factors of the core network elements are used as monitoring sample factor data, and the health level is used as a pre-determined health label.
[0176] The loss function of the health assessment model is:
[0177]
[0178] Among them, w k Let C represent the result obtained in this iteration, C represent a vector, and m represent the number of guard sample factor data in this training.
[0179] The model is evaluated using a loss function; a lower loss function value indicates better robustness. Once trained, the model can be used directly. However, if special cases arise, such as models not being covered during training, iterative optimization is necessary.
[0180] According to the intelligent monitoring device of the present invention, the device further includes a testing module, which is used for:
[0181] The accuracy of the state prediction results of the core network element devices is obtained by testing the output state prediction results using a sliding window.
[0182] In actual training using supervised learning, a sliding window is considered for testing. Sliding windows are more suitable for processing time series data. Here, previous data (health status) is used as the input parameter, and the predicted data is used as the output parameter to monitor the accuracy of the prediction model's output.
[0183] Figure 3A schematic diagram of the physical structure of an electronic device is provided. This electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. The processor 310, communication interface 320, and memory 330 communicate with each other via the communication bus 340. The processor 310 can invoke logical instructions from the memory 330 to execute an intelligent monitoring method, which includes:
[0184] S1. Collect monitoring data from core network element devices, including monitoring reports;
[0185] S2. Aggregate the guarded data to obtain the guarded factor;
[0186] S3. Input the monitoring factor into the equipment health assessment model and output the status prediction results of the core network element equipment;
[0187] The equipment health assessment model is obtained by training based on the monitored sample factor data and pre-determined health labels.
[0188] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0189] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the intelligent monitoring method provided by the above methods, the method comprising:
[0190] S1. Collect monitoring data from core network element devices, including monitoring reports;
[0191] S2. Aggregate the guarded data to obtain the guarded factor;
[0192] S3. Input the monitoring factor into the equipment health assessment model and output the status prediction results of the core network element equipment;
[0193] The equipment health assessment model is obtained by training based on the monitored sample factor data and pre-determined health labels.
[0194] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the aforementioned intelligent watchkeeping methods, the method comprising:
[0195] S1. Collect monitoring data from core network element devices, including monitoring reports;
[0196] S2. Aggregate the guarded data to obtain the guarded factor;
[0197] S3. Input the monitoring factor into the equipment health assessment model and output the status prediction results of the core network element equipment;
[0198] The equipment health assessment model is obtained by training based on the monitored sample factor data and pre-determined health labels.
[0199] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0200] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intelligent monitoring method, characterized in that, include: Collect monitoring data from core network elements, including monitoring reports; The monitored data is aggregated to obtain the monitored factor; The monitoring factors are input into the equipment health assessment model, and the status prediction results of the core network element equipment are output. The equipment health assessment model is obtained by training based on the on-duty sample factor data and the pre-determined health labels. After inputting the monitoring factor into the equipment health assessment model and outputting the state prediction results of the core network element equipment, the process includes: The names of the core network element devices are obfuscated; The equipment health assessment model is obtained by training based on the monitored sample factor data and pre-determined health labels, including: A health index is constructed based on the impact of various monitoring factors of core network elements on health. The corresponding health level of core network elements is determined based on the health index. The various monitoring factors of the core network elements are used as monitoring sample factor data, and the health level is used as a pre-determined health label. The loss function of the health assessment model is: ; in, Let C represent the result obtained in this iteration, C represent a vector, and m represent the number of guard sample factor data in this training.
2. The intelligent monitoring method according to claim 1, characterized in that, The data collected from the core network element devices includes: Customized monitoring tasks, including: dial-up testing tasks, high-frequency monitoring tasks, and periodic monitoring tasks; Based on the duty tasks, collect the duty data of core network element devices.
3. The intelligent monitoring method according to claim 1, characterized in that, The data collected from the core network element devices includes: Obtain the duty reports of core network element devices; Based on the equipment specifications, selectively collect equipment specifications. Based on the device alarm status, selectively collect device alarm information.
4. The intelligent monitoring method according to claim 3, characterized in that, The selective collection of equipment metrics based on equipment performance indicators and the selective collection of equipment alarm information based on equipment alarm status include: After the duty task is executed, the system automatically checks whether there are equipment indicators in the duty report; If equipment specifications exist, then collect the equipment specifications. Check if there are any equipment alarms in the duty report; if there are equipment alarms, collect the equipment alarm information.
5. The intelligent monitoring method according to claim 1, characterized in that, The method further includes: The accuracy of the state prediction results of the core network element devices is obtained by testing the output state prediction results using a sliding window.
6. An intelligent monitoring device, characterized in that, include: The duty data acquisition module is used to collect duty data from core network element devices, including duty reports. The data aggregation module is used to aggregate the guard data to obtain the guard factor; The duty prediction module is used to input the duty factors into the equipment health assessment model and output the status prediction results of the core network element equipment; The equipment health assessment model is obtained by training based on the on-duty sample factor data and the pre-determined health labels. The obfuscation processing module is used to obfuscate the names of the core network element devices. The equipment health assessment model is obtained by training based on the monitored sample factor data and pre-determined health labels, including: A health index is constructed based on the impact of various monitoring factors of core network elements on health. The corresponding health level of core network elements is determined based on the health index. The various monitoring factors of the core network elements are used as monitoring sample factor data, and the health level is used as a pre-determined health label. The loss function of the health assessment model is: ; in, Let C represent the result obtained in this iteration, C represent a vector, and m represent the number of guard sample factor data in this training.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the intelligent watchdog method as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the intelligent watchdog method as described in any one of claims 1 to 5.