A communication network management platform
By generating a topology diagram through a unified topology discovery unit, and combining it with a network-wide monitoring and fault management unit, the anomaly detection model solves the problem of difficulty in real-time monitoring of device operation status in the network, enabling rapid detection and timely alarm of device faults, and ensuring network stability.
Patent Information
- Application Number
- CN202411609894.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-11-12
AI Technical Summary
Existing communication network management platforms are unable to monitor the operational status of all devices in real time, leading to network outages and untimely fault handling.
A unified topology discovery unit generates a topology diagram, which works in conjunction with the network-wide monitoring unit and the fault management unit to monitor the device status in real time. The anomaly detection model quickly locates the fault point and uses machine learning algorithms to build the anomaly detection model, enabling automatic monitoring and alarms.
It enables real-time monitoring of devices in the communication network and rapid fault detection, preventing network interruptions and ensuring network stability and continuity.
Smart Images

Figure CN119484295B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication network technology, and more specifically to a communication network management platform. Background Technology
[0002] A communication network management platform is a system that integrates multiple functions such as network monitoring, resource management, and fault diagnosis. By monitoring and collecting data from various devices in the communication network in real time, it helps operators comprehensively understand the network's operational status, promptly identify and address potential problems, thereby ensuring the continuity and stability of communication services.
[0003] In the operation of various communication networks, the communication network management platform plays a crucial role. However, facing increasingly complex IT environments, network administrators often find it difficult to monitor the operational status of all devices in real time, making it challenging to quickly address network outages and malfunctions. Therefore, a communication network management platform capable of promptly identifying device faults and providing timely alerts is needed. Summary of the Invention
[0004] The purpose of this invention is to provide a communication network management platform that helps network administrators monitor the operating status of all devices in real time, quickly detect device faults, and prevent potential device problems.
[0005] To achieve the above objectives, the present invention provides a communication network management platform, comprising:
[0006] The unified topology discovery unit collects and analyzes the connection relationships between devices in the communication network and generates a topology diagram.
[0007] The network-wide monitoring unit is connected to the unified topology discovery unit to monitor the operating status of the devices in real time.
[0008] The fault management unit is connected to the network monitoring unit to monitor the fault status of the device in real time; and the fault management unit can work in conjunction with the network monitoring unit to monitor the device.
[0009] The performance management unit is connected to the unified topology discovery unit and the fault management unit respectively, and collects the performance index data of the device to provide a means of large-scale IP network performance monitoring.
[0010] The report management unit is connected to the performance management unit via a signal, receives performance indicator data provided by the performance management unit, and develops reports; it can also manage access permissions for the generated reports.
[0011] Optionally, the monitoring methods of the performance management unit include:
[0012] Equipment performance monitoring to obtain equipment performance information;
[0013] Link layer monitoring obtains information on the connection relationships between devices;
[0014] Path layer monitoring obtains the logical paths and subnetting of the communication network.
[0015] Optionally, the method for generating a topology graph using the unified topology discovery unit includes:
[0016] Step 1: Obtain device performance information using the device performance monitoring of the performance management unit;
[0017] Step 2: Use the link layer monitoring and path layer monitoring of the performance management unit to obtain device connection relationship information, logical paths of the communication network, and subnetting.
[0018] Step 3: Based on the device information collected in Step 1 and the device connection relationships collected in Step 2, the devices in the communication network are graphically represented to form a topology diagram.
[0019] Optionally, step 3 further includes: dynamically updating the topology diagram based on the changes in devices in the communication network monitored in real time by the network-wide monitoring unit; wherein the changes in devices include: device operating status and replacement of new / old devices.
[0020] Optionally, step 2 further includes: for device connection relationships that cannot be discovered through automatic scanning, the network administrator can manually input and verify the connection relationships between devices to complete the supplementation of device connection relationships.
[0021] Optionally, in step 1, the method for obtaining device performance information includes:
[0022] Broadcast detection involves sending broadcast messages within a communication network system to request all responding devices to report their performance and location information.
[0023] Protocol analysis identifies device performance information within a communication network by analyzing the protocols used in that network.
[0024] Simple Network Management Protocol (SMMP) query: Use the Simple Network Management Protocol to send a query request to the device to obtain detailed configuration, performance, and status information of the device.
[0025] Optionally, the method for the network-wide monitoring unit and the fault management unit to work together to complete equipment monitoring includes:
[0026] Step A: The network-wide monitoring unit collects real-time monitoring performance indicator data from multiple data sources;
[0027] Step B involves setting alarm thresholds for the metric data collected in Step A through the performance management unit, and defining the requirements and alarm levels for triggering alarms through the fault management unit.
[0028] Step C: Based on the alarm threshold and alarm triggering requirements set in Step B, the performance management unit compares the alarm thresholds or uses an anomaly detection model to determine and evaluate whether the indicator data collected in Step A meets the alarm triggering requirements.
[0029] When the indicator data reaches the requirements to trigger an alarm, the fault management unit will issue an alarm according to the set alarm level, promptly reminding the network administrator to handle the equipment fault problem, and transmitting the equipment fault status to the network-wide monitoring unit to update the equipment operating status in the topology diagram; when the indicator data does not reach the requirements to trigger an alarm, it indicates that the equipment is operating normally.
[0030] Optionally, in step C, the method for establishing the anomaly detection model includes:
[0031] Step C1: Collect the dataset for the anomaly detection model;
[0032] Step C2: Extract the feature indicators that can be used to judge equipment abnormalities from the various indicators in the dataset, enter the data in the dataset into the feature indicators, and preprocess the missing data, abnormal data and duplicate data in the entry process to construct the feature dataset.
[0033] Step C3: Divide the feature dataset into training set, validation set, and test set; use the training set to establish the relationship between feature data and equipment fault status, and establish a loss function; use the validation set to iteratively optimize the loss function and optimize the anomaly detection model.
[0034] An anomaly detection model is used to calculate data from the test set to evaluate the accuracy of the anomaly test model in predicting equipment operating status.
[0035] Optionally, in step B, the alarm threshold can be set by user-defined settings, based on the fault cause provided by the fault management unit, or by obtaining historical trends of performance indicator data through the report management unit.
[0036] Optionally, in step A, the data source includes: devices, servers, and applications, and the performance indicator data includes: system resources, application performance, and network latency.
[0037] Compared with the prior art, the technical solution of the present invention has at least the following beneficial effects:
[0038] This invention automatically scans devices in the network through a unified topology discovery unit, collects and analyzes the connection relationships between devices, establishes a complete network topology diagram, and dynamically updates the network topology diagram to ensure its real-time performance and accuracy. This helps network administrators monitor the operational status of all devices in real time. Simultaneously, this invention utilizes a network-wide monitoring unit and a fault management unit working collaboratively to monitor the operational status of all devices in real time, identify and resolve potential device problems, and prevent network outages and other issues.
[0039] This invention collects real-time indicator data from multiple data sources through a fault management unit, sets thresholds for specific indicators to define when to trigger alarms, and facilitates timely reminders to network administrators to handle device faults.
[0040] This invention establishes an anomaly detection model using machine learning algorithms to automatically monitor the fault status of all network devices, quickly locate device fault points, and thus prevent network outages and other failures. Training, validation, and test sets are constructed using multiple data sources, and the collected data is processed using linear regression to establish a hypothesis model and loss function, thereby obtaining the parameters of the anomaly detection model and enabling rapid detection of device faults. Attached Figure Description
[0041] Figure 1 This is a schematic diagram of the communication network management platform of the present invention.
[0042] Figure 2 This is an architecture diagram for establishing the topology in the communication network management platform of the present invention.
[0043] Figure 3 This is an architecture diagram of the device monitoring implementation in the communication network management platform of the present invention.
[0044] Figure 4 This is an architecture diagram of the anomaly detection model built in the communication network management platform of the present invention. Detailed Implementation
[0045] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0046] In the description of this invention, it should be noted that the terms "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0047] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0048] like Figure 1 As shown, this invention provides a communication network management platform, including: a unified topology discovery unit, a network-wide monitoring unit, a fault management unit, a performance management unit, and a report management unit. The unified topology discovery unit is signal-connected to both the network-wide monitoring unit and the performance management unit, receiving real-time device operating status data from the network-wide monitoring unit and device performance index data collected by the performance management unit. The fault management unit is signal-connected to the network-wide monitoring unit and can determine whether a device is faulty based on the real-time monitored device performance data. The performance management unit is signal-connected to the report management unit and can transmit and store the collected performance index data in the report management unit.
[0049] The unified topology discovery unit can collect and analyze the connection relationships between devices in the communication network, generate and display the topology diagram, and provide a visual operation interface to help network administrators monitor the operation status of all devices in real time. Users can also directly click on the nodes in the topology diagram through the visual operation interface to enter the corresponding device operation panel.
[0050] The network-wide monitoring unit is signal-connected to the unified topology discovery unit, enabling real-time monitoring of the operational status of all devices in the communication network and sending the real-time monitoring results to the unified topology discovery unit to update the topology diagram. The monitoring results include: device online status and performance indicator data.
[0051] The fault management unit and the network monitoring unit are connected by a signal, enabling real-time monitoring of the equipment's fault status and providing alarms and audible / visual alerts when a fault is detected. The fault management unit allows users to define alarm levels for different faults, masking duplicate alarms and intermittent alarms to reduce false alarms. Simultaneously, the fault management unit provides an alarm handling experience storage function, storing the cause and handling process of each alarm for future equipment maintenance and facilitating automated troubleshooting.
[0052] Furthermore, the fault management unit can work in conjunction with the network-wide monitoring unit to jointly monitor the operating status of devices in the communication network and promptly issue alarms for detected device faults, reminding network administrators to handle the faults.
[0053] The performance management unit is connected to the unified topology discovery unit and the fault management unit, respectively. It can collect device performance data, provide large-scale IP network performance monitoring methods, and support threshold alarms for network performance indicators. The monitoring methods include device performance monitoring, link-layer monitoring, and path-layer monitoring, which can identify network performance bottlenecks at different layers and provide data support for communication network optimization.
[0054] Specifically, device performance information can be obtained through device performance monitoring, connection relationship information between devices can be obtained through link layer monitoring, and logical paths and subnetting of the communication network can be obtained through path layer monitoring.
[0055] The report management unit is signal-connected to the performance management unit, enabling it to receive performance indicator data provided by the performance management unit and develop reports based on this data to obtain historical trends. The report management unit also manages access permissions for the generated reports. Specifically, the report development function provides template-based report development and web-based report generation, distribution, and management capabilities; the report access management function manages report permissions to meet the operational needs of different users.
[0056] The performance management system automatically scans network device information, device connection relationships, and network path information, and sends this information to the unified topology discovery unit, establishing a complete network topology diagram. This diagram is continuously and dynamically updated based on device changes, changes in device connections, or changes in network paths, ensuring its real-time accuracy and helping network administrators monitor the operational status of all devices in real time. Simultaneously, the network-wide monitoring unit and the fault management unit work together to monitor the operational status of all devices in real time and detect any faults in the network. The network-wide monitoring unit collects real-time device indicator data from multiple data sources, filters characteristic indicators from this data as indicators reflecting device faults, sets alarm thresholds for these indicators, and defines the requirements for triggering alarms. By comparing whether the characteristic indicator data meets the alarm requirements, the system determines the device's operational status, quickly locates the fault point, and effectively prevents network outages and other device malfunctions.
[0057] Specifically, the method for generating a topology graph using the unified topology discovery unit includes:
[0058] Step 1: Automatically scan for devices in the communication network.
[0059] Device performance information is obtained through device performance monitoring using the performance management unit. This is achieved by broadcasting messages to the communication network system, requesting all responding devices to report their performance and location information, thus performing broadcast probing of devices in the communication network. Furthermore, device performance information is identified by analyzing the protocols used in the communication network (including Address Resolution Protocol, Dynamic Host Configuration Protocol, and Internet Control Message Protocol), enabling protocol analysis of the devices. Simultaneously, Simple Network Management Protocol (SNMP) queries are sent to devices to obtain detailed configuration, performance, and status information, enabling SNMP queries of the communication network.
[0060] Step 2: Determine the connection relationships between the devices.
[0061] The performance management unit utilizes link-layer and path-layer monitoring to acquire device connectivity information, logical paths, and subnetting of the communication network. By analyzing device MAC address tables and port statuses, the physical connections between devices are determined, completing link-layer monitoring of the communication network. Logical paths and subnetting are obtained by parsing routing tables in routers, completing routing information parsing. Simultaneously, the unified topology discovery unit supports user configuration; for device connectivity relationships that cannot be automatically discovered, network administrators can manually input and verify the connectivity relationships between devices, supplementing the device connectivity information.
[0062] Step 3: Generate and dynamically refresh the topology diagram.
[0063] Based on the device performance information collected in step 1 and the device connection relationships collected in step 2, the devices in the communication network are graphically represented, and a topology diagram is constructed through a unified topology discovery unit.
[0064] Based on the actual structure and complexity of the communication network, the topology diagram is displayed in layers to facilitate better understanding and analysis.
[0065] Meanwhile, the topology diagram can be dynamically updated based on changes in devices in the communication network monitored in real time by the network-wide monitoring unit (including device operating status and replacement of new / old devices). When a new device is added or an old device is removed from the communication network, the topology diagram will be automatically updated based on the new connection relationships added by the new device or the connection relationships removed by the old device.
[0066] Specifically, methods for coordinating the network-wide monitoring unit and the fault management unit to jointly complete equipment monitoring include:
[0067] Step A: Monitoring data collection.
[0068] The network-wide monitoring unit collects real-time performance metrics data from multiple data sources. These data sources include: devices (e.g., routers, switches, firewalls), servers, and applications; the performance metrics data include: system resources (e.g., CPU utilization, memory usage), application performance (e.g., response time, throughput), and network latency.
[0069] Step B: Set the threshold for the indicator data.
[0070] The alarm thresholds for the metric data collected in step A are set through the performance management unit, and the requirements for triggering alarms and the alarm levels are defined through the fault management unit. These alarm thresholds can be set by the user, based on the fault causes provided by the fault management unit, or by obtaining historical trends of performance metric data through the report management unit.
[0071] Step C: Perform anomaly detection on each indicator.
[0072] Based on the alarm thresholds and alarm triggering requirements set in step B, the performance management unit compares the alarm thresholds to determine and evaluate whether the indicator data collected in step A meets the alarm triggering requirements. When the indicator data meets the alarm triggering requirements, the fault management unit issues an alarm according to the set alarm level, promptly reminding the network administrator to handle the device fault and transmitting the device fault status to the network-wide monitoring unit to update the device operating status in the topology diagram. When the indicator data does not meet the alarm triggering requirements, it indicates that the device is operating normally.
[0073] Furthermore, an anomaly detection model can be established to predict and alert on device status. The performance management unit automatically evaluates whether the performance metrics data collected in step A triggers an alarm, thus automatically optimizing the anomaly detection process. Data from various sources, including network devices, security systems, application logs, and traffic analysis tools, is collected to obtain a dataset containing multiple performance metrics. This dataset is then divided into training, experimental, and test sets. A linear regression algorithm is used to train the training set, obtaining the hypothesis function of the anomaly detection model. Simultaneously, an appropriate loss function (e.g., mean squared error) is selected to measure the difference between the model's predicted and actual values. Based on the experimental set, optimization algorithms (e.g., gradient descent) are used to optimize the anomaly detection model parameters. By minimizing the loss function, the parameters of the anomaly detection model are estimated, thus obtaining the optimal solution for the model's parameters. The test set is used to evaluate the model's performance, calculating the difference between predicted and actual values to verify whether the anomaly detection model can determine the device's operating status based on performance metrics data. The tested anomaly detection model is then deployed in practical applications to achieve automatic detection of device faults.
[0074] Specifically, the method for establishing the anomaly detection model includes:
[0075] Step C1: Collect the dataset of the anomaly detection model.
[0076] The required data set includes not only the real-time monitoring performance metrics data obtained in step A, but also historical data such as network traffic data, session logs, system logs, and user behavior data obtained from self-security systems (e.g., intrusion detection systems IDS / IPS), log analysis tools, and traffic analysis tools.
[0077] Step C2 involves preprocessing the data in the dataset.
[0078] The system extracts characteristic indicators from various metrics in the dataset that can be used to determine device anomalies. These characteristic indicators include: traffic, number of data packets, connection duration, protocol type, and source / destination IP address.
[0079] The data in the dataset is entered into the feature index. At the same time, the feature data in the entered feature index is preprocessed, duplicate data in the feature data is deleted, missing data in the feature data is handled by interpolation, and outlier handling methods such as Z-score or interquartile range (IQR) are used to handle outlier data, thereby improving the quality of the feature data and obtaining the preprocessed feature dataset.
[0080] Linear regression is used to linearize several sets of feature data in the feature dataset, which facilitates the establishment of subsequent anomaly detection models.
[0081] The data in the feature dataset is numericalized and normalized, and then transformed into a format that machine learning can process.
[0082] Step C3: Establish an anomaly detection model.
[0083] The feature dataset in step C2 is divided into a training set, a validation set, and a test set, with each set containing multiple sets of feature index data.
[0084] The anomaly detection model is trained using a training set to establish the relationship between feature data and equipment fault states, constructing the hypothesis function of the anomaly detection model, and subsequently, its loss function. Simultaneously, the loss function is iteratively optimized using a validation set, adjusting the parameters of the anomaly detection model (including weights and biases) to obtain the minimum loss function value. The parameters of the machine learning model at which the loss function reaches its minimum value are taken as the current optimal parameters, thus optimizing the anomaly detection model.
[0085] Step C4: Evaluate the anomaly detection model.
[0086] An anomaly detection model is used to calculate the data in the test set to obtain the prediction results of the feature data detection in the test set. The prediction results are then compared with the actual equipment status reflected by the feature data in the test set to evaluate the accuracy of the anomaly test model in predicting the equipment operating status.
[0087] When the anomaly detection model can accurately determine the device status corresponding to the performance index data in the test set, the anomaly detection model can be put into actual use; and based on the actual collected performance index data, it can determine whether the device has failed, and then issue an alarm based on the device failure status and its corresponding alarm level to remind the network administrator to deal with the device failure problem in a timely manner.
[0088] In summary, this invention constructs and refreshes the topology diagram in real time through a unified topology discovery unit, ensuring the real-time nature and accuracy of the topology diagram and helping network administrators monitor the operating status of all devices in real time. At the same time, through the collaborative work of the network-wide monitoring unit and the fault management unit, the operating status of devices is monitored in real time, and the anomaly detection model is constructed to automatically detect whether there are faults in the device status, thereby achieving rapid detection of device faults.
[0089] Although the present invention has been described in detail through the preferred embodiments above, it should be understood that the above description should not be considered as a limitation of the present invention. Various modifications and substitutions to the present invention will be apparent to those skilled in the art after reading the above description. Therefore, the scope of protection of the present invention should be defined by the appended claims.
Claims
1. A communication network management platform, characterized in that, include: The unified topology discovery unit collects and analyzes the connection relationships between devices in the communication network and generates a topology diagram. The network-wide monitoring unit is connected to the unified topology discovery unit to monitor the operating status of the devices in real time. The fault management unit is connected to the network monitoring unit to monitor the fault status of the equipment in real time. Furthermore, the fault management unit can work in conjunction with the network-wide monitoring unit to monitor the equipment; The performance management unit is connected to the unified topology discovery unit and the fault management unit respectively, and collects the performance index data of the device to provide a means of large-scale IP network performance monitoring. The report management unit is connected to the performance management unit via a signal, receives performance indicator data provided by the performance management unit, and performs report development. It can also manage access permissions for the generated reports; The method for monitoring the device by the collaborative operation of the network-wide monitoring unit and the fault management unit includes: Step A: The network-wide monitoring unit collects real-time monitoring performance indicator data from multiple data sources; Step B involves setting alarm thresholds for the metric data collected in Step A through the performance management unit, and defining the requirements and alarm levels for triggering alarms through the fault management unit. Step C: Based on the alarm threshold and alarm triggering requirements set in Step B, the performance management unit compares the alarm thresholds or uses an anomaly detection model to determine and evaluate whether the indicator data collected in Step A meets the alarm triggering requirements. When the indicator data reaches the requirements to trigger an alarm, the fault management unit will issue an alarm according to the set alarm level, promptly reminding the network administrator to handle the equipment fault problem, and transmitting the equipment fault status to the network-wide monitoring unit to update the equipment operating status in the topology diagram; when the indicator data does not reach the requirements to trigger an alarm, it indicates that the equipment is operating normally.
2. The communication network management platform according to claim 1, characterized in that, The monitoring methods of the performance management unit include: Equipment performance monitoring to obtain equipment performance information; Link layer monitoring obtains information on the connection relationships between devices; Path layer monitoring obtains the logical paths and subnetting of the communication network.
3. The communication network management platform according to claim 2, characterized in that, The method for generating a topology graph using the unified topology discovery unit includes: Step 1: Obtain device performance information using the device performance monitoring of the performance management unit; Step 2: Use the link layer monitoring and path layer monitoring of the performance management unit to obtain device connection relationship information, logical paths of the communication network, and subnetting. Step 3: Based on the device information collected in Step 1 and the device connection relationships collected in Step 2, the devices in the communication network are graphically represented to form a topology diagram.
4. The communication network management platform according to claim 3, characterized in that, Step 3 further includes: dynamically updating the topology diagram based on the changes in devices in the communication network monitored in real time by the network-wide monitoring unit; wherein the changes in devices include: device operating status and replacement of new / old devices.
5. The communication network management platform according to claim 3, characterized in that, Step 2 further includes: for device connection relationships that cannot be discovered through automatic scanning, the network administrator can manually input and verify the connection relationships between devices to complete the supplementation of device connection relationships.
6. The communication network management platform according to claim 3, characterized in that, In step 1, the method for obtaining device performance information includes: Broadcast detection involves sending broadcast messages within a communication network system to request all responding devices to report their performance and location information. Protocol analysis identifies device performance information within a communication network by analyzing the protocols used in that network. Simple Network Management Protocol (SMMP) query: Use the Simple Network Management Protocol to send a query request to the device to obtain detailed configuration, performance, and status information of the device.
7. The communication network management platform according to claim 1, characterized in that, In step C, the method for establishing the anomaly detection model includes: Step C1: Collect the dataset for the anomaly detection model; Step C2: Extract the feature indicators that can be used to judge equipment abnormalities from the various indicators in the dataset, enter the data in the dataset into the feature indicators, and preprocess the missing data, abnormal data and duplicate data in the entry process to construct the feature dataset. Step C3: Divide the feature dataset into training set, validation set, and test set; use the training set to establish the relationship between feature data and equipment fault status, and establish a loss function; use the validation set to iteratively optimize the loss function and optimize the anomaly detection model. An anomaly detection model is used to calculate data from the test set to evaluate the accuracy of the anomaly test model in predicting equipment operating status.
8. The communication network management platform according to claim 1, characterized in that, In step B, the alarm threshold can be set by user-defined settings, based on the fault cause provided by the fault management unit, or by obtaining historical trends of performance indicator data through the report management unit.
9. The communication network management platform according to claim 1, characterized in that, In step A, the data source includes: devices, servers, and applications, and the performance indicator data includes: system resources, application performance, and network latency.
Citation Information
Patent Citations
Network equipment monitoring management method and system
CN111953530A
Method and system for monitoring and analysing a telecommunication transport network.
EP2928119A1