PTP network performance monitoring method and system based on big data analysis, and medium
Through the PTP network performance monitoring method based on big data analysis, the problem that traditional monitoring systems are difficult to build a global perspective and intelligent optimization suggestions is solved, and intelligent abnormality detection, accurate fault location and predictive maintenance of PTP network are realized, which improves the continuous improvement of network performance.
Patent Information
- Application Number
- CN202510585856.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-08
AI Technical Summary
Traditional PTP network monitoring systems lack a global perspective, it is difficult to build a complete PTP clock distribution topology, it cannot intuitively present the overall PTP state of the system, and it lacks intelligent optimization suggestions, making it difficult to effectively locate the fault source and perform predictive maintenance.
The PTP network performance monitoring method based on big data analysis is adopted, and PTP packets are collected through multiple data acquisition points, data preprocessing and clock error calculation are carried out, PTP network performance timing database is built, and a multi-dimensional PTP network performance analysis model is established to perform real-time status monitoring, abnormal detection, machine learning analysis and network performance optimization.
It realizes global status monitoring, intelligent abnormality detection, accurate fault location and predictive maintenance of PTP networks, improves the ability to continuously improve network performance, and reduces the difficulty of troubleshooting and deployment costs.
Smart Images

Figure CN120110951A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a PTP network performance monitoring method, system and medium based on big data analysis. Background Art
[0002] With the widespread application of IP technology in the radio and television industry, network synchronization technology based on PTP (Precision Time Protocol) has become a key infrastructure of modern radio and television production and broadcasting systems. The PTP protocol is defined by IEEE1588-2008 and can achieve high-precision time synchronization of less than 1μs in distributed network communications. In the broadcasting field, the SMPTE ST 2059-2 standard is based on the IEEE 1588 standard and provides a nanosecond-accurate solution for time and frequency synchronization in professional broadcasting environments. Traditional PTP network monitoring systems mainly rely on basic network management protocols such as SNMP and NetConf to obtain device PTP status information, or analyze PTP messages in specific network segments by capturing packets. These systems usually use a threshold alarm mechanism to trigger an alarm when a PTP parameter is detected to be out of the preset range, and provide basic status display and historical data query functions.
[0003] However, with the continuous growth of the scale and complexity of IP broadcasting systems, traditional PTP network monitoring methods face many challenges. First, existing monitoring systems generally lack a global perspective, making it difficult to build a complete PTP clock distribution topology and unable to intuitively present the overall PTP status of the system. Second, traditional monitoring methods use simple threshold judgments and fail to make full use of historical data for in-depth analysis, resulting in limited ability to predict potential problems. Furthermore, in a complex network environment, a single anomaly may trigger a chain reaction, and traditional monitoring systems are difficult to effectively locate the root cause of the fault, increasing the difficulty of troubleshooting. In addition, the existing system lacks intelligent optimization suggestion functions and cannot automatically generate optimization plans based on historical data and current status, which restricts the continuous improvement of PTP network performance. Finally, traditional monitoring systems usually rely on dedicated hardware implementation, with high deployment costs and poor scalability, making it difficult to adapt to the rapidly changing broadcast technology environment. Summary of the invention
[0004] The present application provides a PTP network performance monitoring method, system and medium based on big data analysis, which is used to achieve global status monitoring, intelligent anomaly detection, precise fault location and predictive maintenance of the PTP network by establishing a PTP network performance monitoring method based on big data analysis, thereby ensuring the high-precision time synchronization requirements of the radio and television IP production and broadcasting system.
[0005] In the first aspect, the present application provides a PTP network performance monitoring method based on big data analysis, and the PTP network performance monitoring method based on big data analysis includes: collecting PTP message data, network equipment PTP status data and RTP stream data packets in the PTP network through multiple data collection points to obtain a structured data stream; performing data preprocessing and clock error calculation on the structured data stream to obtain a standardized data set containing PTP network status parameters; constructing a PTP network performance time series database based on the standardized data set, and establishing a multidimensional PTP network performance analysis model to obtain a network performance analysis result; performing real-time status monitoring and anomaly detection according to the network performance analysis result to obtain a PTP network abnormal event record; performing machine learning analysis based on the PTP network abnormal event record and the network performance analysis result to obtain a network performance score and a fault location result; performing network performance optimization and predictive maintenance according to the network performance score and the fault location result, and generating optimization suggestions and maintenance work orders.
[0006] In a second aspect, the present application provides a PTP network performance monitoring system based on big data analysis, and the PTP network performance monitoring system based on big data analysis includes:
[0007] The collection module is used to collect PTP message data, network device PTP status data and RTP stream data packets in the PTP network through multiple data collection points to obtain a structured data stream;
[0008] A calculation module, used for performing data preprocessing and clock error calculation on the structured data stream to obtain a standardized data set including PTP network status parameters;
[0009] Establish a module for constructing a PTP network performance time series database based on the standardized data set, and establishing a multidimensional PTP network performance analysis model to obtain network performance analysis results;
[0010] A detection module, used to perform real-time status monitoring and anomaly detection according to the network performance analysis results, and obtain a record of abnormal events of the PTP network;
[0011] An execution module, configured to perform machine learning analysis based on the PTP network abnormal event record and the network performance analysis result to obtain a network performance score and a fault location result;
[0012] A maintenance module is used to perform network performance optimization and predictive maintenance according to the network performance score and the fault location result, and generate optimization suggestions and maintenance work orders.
[0013] In a third aspect, a PTP network performance monitoring device based on big data analysis is provided, comprising: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor calls the instructions in the memory so that the PTP network performance monitoring device based on big data analysis executes the above-mentioned PTP network performance monitoring method based on big data analysis.
[0014] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, and when the computer-readable storage medium is run on a computer, the computer executes the above-mentioned PTP network performance monitoring method based on big data analysis.
[0015] In the technical solution provided by the present application, the PTP message data, the PTP status data of the network equipment and the RTP stream data packet in the PTP network are collected through multiple data collection points to obtain a structured data stream, which realizes comprehensive and accurate data collection and provides a high-quality data basis for subsequent analysis. Data preprocessing and clock error calculation are performed for the collected structured data stream to obtain a standardized data set containing PTP network status parameters, which effectively eliminates the influence of network jitter and improves the accuracy of clock deviation calculation. A PTP network performance time series database is constructed based on a standardized data set, and a multidimensional PTP network performance analysis model is established to obtain network performance analysis results. Through a distributed storage architecture and a time span hierarchical storage strategy, efficient storage and query of massive time series data are realized, and the long short-term memory network model gives full play to the advantages of artificial intelligence algorithms in time series data prediction, accurately captures the long-term dependencies and short-term fluctuation characteristics in the PTP network, and significantly improves the accuracy and timeliness of performance prediction. According to the network performance analysis results, real-time status monitoring and anomaly detection are performed to obtain PTP network abnormal event records, realize intuitive visualization and accurate anomaly detection of network status, and the correlation analysis algorithm effectively filters short-term anomalies and reduces the false alarm rate. Machine learning analysis is performed based on the PTP network abnormal event records and network performance analysis results to obtain network performance scores and fault location results. The machine learning algorithm greatly improves the accuracy and efficiency of fault root location by learning historical abnormal patterns, and converts complex multi-dimensional data into intuitive performance scores, which is convenient for management decisions. Network performance optimization and predictive maintenance are performed based on network performance scores and fault location results, and optimization suggestions and maintenance work orders are generated. With the help of time series prediction algorithms, network performance is predicted in a forward-looking manner, so that maintenance is transformed from passive response to active prevention. The simulation evaluation mechanism of optimization suggestions ensures the effectiveness and safety of optimization measures. The overall solution deeply integrates big data collection, artificial intelligence analysis and PTP network monitoring. Especially in the analysis of time series data, the application of long short-term memory networks gives full play to the advantages of deep learning algorithms in processing time series data, and can accurately capture complex patterns and long-term dependencies in PTP clock data; in the anomaly detection link, unsupervised learning algorithms can automatically identify abnormal patterns from massive data without pre-defining all abnormal types; in fault root cause analysis, causal reasoning algorithms can accurately trace abnormal propagation paths and locate the root causes. The characteristics of these artificial intelligence algorithms and models directly improve the accuracy, foresight and automation of PTP network monitoring, and realize the transformation from passive monitoring to active prevention. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0017] Figure 1 A schematic diagram of an embodiment of a PTP network performance monitoring method based on big data analysis in an embodiment of the present application;
[0018] Figure 2 A schematic diagram of an embodiment of a PTP network performance monitoring system based on big data analysis in an embodiment of the present application;
[0019] Figure 3 It is a structural schematic block diagram of a PTP network performance monitoring device based on big data analysis in an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The embodiments of the present application provide a PTP network performance monitoring method, system and medium based on big data analysis. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0021] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 , an embodiment of the PTP network performance monitoring method based on big data analysis in the embodiment of the present application includes:
[0022] Step S101, collecting PTP message data, network device PTP status data and RTP stream data packets in the PTP network through multiple data collection points to obtain a structured data stream;
[0023] Step S102: performing data preprocessing and clock error calculation on the structured data stream to obtain a standardized data set including PTP network status parameters;
[0024] Step S103: construct a PTP network performance time series database based on the standardized data set, and establish a multi-dimensional PTP network performance analysis model to obtain network performance analysis results;
[0025] Step S104: Perform real-time status monitoring and anomaly detection according to the network performance analysis results to obtain a PTP network abnormal event record;
[0026] Step S105: Perform machine learning analysis based on the PTP network abnormal event records and network performance analysis results to obtain network performance scores and fault location results;
[0027] Step S106: Execute network performance optimization and predictive maintenance according to the network performance score and fault location results, and generate optimization suggestions and maintenance work orders.
[0028] It is understandable that the execution subject of the present application can be a PTP network performance monitoring system based on big data analysis, or a terminal or a server, which is not limited here. The embodiment of the present application is described by taking the server as the execution subject as an example.
[0029] Specifically, the PTP network performance monitoring method based on big data analysis first collects PTP message data, network equipment PTP status data and RTP stream data packets in the PTP network through multiple data collection points to obtain a structured data stream. In this step, the collection server is deployed at a strategic position in the network, accesses the IP data stream through the optical fiber network card, and directly obtains the hardware timestamp of the network data packet using the MII interface. PTP message collection is performed on four key messages: synchronization message (Sync), follow message (Follow_Up), delay request message (Delay_Req) and delay response message (Delay_Resp), and extracts data fields such as timestamp information, message sequence number, source address and destination address from them. The PTP status data of the network device is obtained through NetConf and SNMP protocols, and the data content includes parameters such as PTP lock status, Grandmaster ID, ClockID, PTP working mode and step mode. For devices that do not support standard management protocols, the SSH protocol is used to obtain the original data, and then converted into standard JSON format through regular expressions. RTP stream data packet collection focuses on extracting source address, source port, destination address, destination port, sequence number, tag information and RTP timestamp. All the collected data are attached with collection timestamp and collection point identification information to form a complete structured data stream.
[0030] Data preprocessing and clock error calculation are performed on the structured data stream to obtain a standardized data set containing PTP network status parameters. In the data preprocessing stage, four types of timestamps are first extracted from the structured data stream: the sending timestamp t1 and receiving timestamp t2 of the synchronization message, as well as the sending timestamp t3 of the delay request message and the receiving timestamp t4 of the delay response message. Based on these four timestamps, the network path delay value delay = [(t2-t1)+(t4-t3)] / 2 is calculated, and then the clock offset offset = (t2-t1-delay) is calculated. In order to eliminate the error caused by network jitter, the Kalman filter algorithm is used to de-jitter the original clock offset value. The Kalman filter algorithm smoothes the clock offset value through two stages: state prediction and measurement update, effectively removing short-term jitter interference. Data preprocessing also includes outlier identification and removal. Through statistical analysis, deviation values beyond the normal range are identified and removed, and data from different sources and formats are unified into a standard format, ultimately forming a standardized data set containing key parameters such as the PTP lock status of network equipment, clock source information, clock deviation value, jitter value, and media delay value.
[0031] Based on the standardized data set, a PTP network performance time series database is constructed, and a multi-dimensional PTP network performance analysis model is established to obtain the network performance analysis results. The time series database adopts a distributed storage architecture, which is optimized for the characteristics of time series data, and supports nanosecond time accuracy and efficient time interval query. Data is stored in layers according to the time span. The high-precision original data of the last 7 days is stored in the high-speed storage layer, and the historical data is stored in the low-cost storage layer after downsampling according to the time period. On this basis, the feature vectors of clock deviation, jitter value and network delay data are extracted from the PTP data storage structure, and a multi-dimensional PTP network performance analysis model based on the long short-term memory network (LSTM) is constructed. The model includes an input layer, two long short-term memory hidden layers and a fully connected output layer. The mean square error loss function and Adam optimization algorithm are used to iteratively optimize the model parameters. The trained model can predict clock stability and network performance indicators based on the input real-time standardized data set, and generate complete network performance analysis results in combination with PTP clock distribution topology information.
[0032] According to the results of network performance analysis, real-time status monitoring and anomaly detection are performed to obtain the abnormal event record of PTP network. First, a complete PTP distribution topology diagram is constructed based on the performance analysis results to display the Grandmaster device, boundary clock device and ordinary clock device and their master-slave relationship in real time. In the topology diagram, the PTP lock status, clock deviation value and jitter value of each device are marked in color coding to intuitively reflect the device status. Multi-level threshold judgment criteria are set to monitor the clock deviation change value in real time. When the change value exceeds the preset threshold (such as 500 nanoseconds), a clock jump abnormality mark is generated. Monitor the PTP lock status change of the device or the long-term unlocking situation, and generate a lock status abnormality alarm. Analyze the clock deviation trend, identify the situation where the unidirectional continuous growth trend exceeds the preset drift rate, and generate a clock drift abnormality mark. Apply correlation analysis to filter short-term anomalies caused by network jitter, reduce the false alarm rate, and finally form an accurate record of PTP network abnormal events.
[0033] Machine learning analysis is performed based on the PTP network abnormal event records and network performance analysis results to obtain network performance scores and fault location results. First, clock stability data (including clock deviation standard deviation, maximum deviation value), synchronization accuracy data (average deviation value, deviation distribution), network transmission data (PTP message packet loss rate, delay change rate) and system reliability data (device state transition frequency, abnormal event frequency) are extracted and normalized to obtain standardized feature data. A weighted calculation method is applied to the standardized feature data to generate a network performance score of 0-100 points. The abnormal occurrence time, abnormal type, involved equipment and abnormal parameter value are extracted from the abnormal event record to construct an abnormal feature vector. The device or link where the abnormality first occurred is identified through root cause analysis. Combined with the historical fault processing database, the root cause of the fault is accurately located.
[0034] Perform network performance optimization and predictive maintenance based on network performance scores and fault location results, and generate optimization suggestions and maintenance work orders. Apply time series prediction algorithms to analyze historical data and current network status, and predict future trends in device clock stability, network delay, and jitter. Compare the prediction results with the preset performance thresholds, identify potential performance degradation risks, and generate a list of predictive maintenance requirements. According to the fault location results and low-scoring items in the network performance score, extract the corresponding optimization solutions from the optimization knowledge base, and generate targeted PTP configuration parameter optimization suggestions, including adjusting parameters such as announce interval, sync interval, and delay_req interval. Simulate and evaluate the optimization suggestions, and calculate the performance improvement effect and potential risks after parameter adjustment. Analyze the long-term trend of device clock stability, identify signs of hardware aging, and generate device health assessment results. Finally, based on the optimization assessment report and device health assessment results, generate standardized maintenance work orders, including maintenance objects, maintenance content, and priority information.
[0035] For example, in a certain production and broadcasting network system, after the method is deployed, the sending timestamp t1 of the PTP synchronization message collected by the data collection point from the core switch is 1613455782.123456789 seconds, the receiving timestamp t2 is 1613455782.123458789 seconds, the sending timestamp t3 of the delay request message is 1613455782.223456789 seconds, and the receiving timestamp t4 of the delay response message is 1613455782.223458789 seconds. The path delay value delay = [(t2-t1)+(t4-t3)] / 2 = [(123458789-123456789)+(223458789-223456789)] / 2 = 1000 nanoseconds is calculated, and the clock offset offset = (t2-t1-delay) = 1000 nanoseconds is calculated. After Kalman filtering, the stable clock offset value is 985 nanoseconds. The system stores these data in the timing database. The multi-dimensional PTP network performance analysis model analyzes historical data and finds that the clock deviation of the device shows a slow growth trend. It predicts that the deviation will exceed 1500 nanoseconds after 24 hours. The real-time monitoring system generates a clock drift anomaly mark. The machine learning analysis module identifies the root cause of the problem as the aging of the crystal oscillator of the device based on the abnormal feature vector. The system automatically generates a maintenance work order including a suggestion to replace the crystal oscillator, thereby preventing potential synchronization problems in advance.
[0036] In the embodiment of the present application, the PTP message data, the network equipment PTP status data and the RTP stream data packet in the PTP network are collected through multiple data collection points to obtain a structured data stream, which realizes comprehensive and accurate data collection and provides a high-quality data basis for subsequent analysis. Data preprocessing and clock error calculation are performed for the collected structured data stream to obtain a standardized data set containing PTP network status parameters, which effectively eliminates the influence of network jitter and improves the accuracy of clock deviation calculation. A PTP network performance time series database is constructed based on a standardized data set, and a multidimensional PTP network performance analysis model is established to obtain network performance analysis results. Through a distributed storage architecture and a time span hierarchical storage strategy, efficient storage and query of massive time series data are realized, and the long short-term memory network model gives full play to the advantages of artificial intelligence algorithms in time series data prediction, accurately captures the long-term dependencies and short-term fluctuation characteristics in the PTP network, and significantly improves the accuracy and timeliness of performance prediction. According to the network performance analysis results, real-time status monitoring and anomaly detection are performed to obtain PTP network abnormal event records, realize intuitive visualization and accurate anomaly detection of network status, and the correlation analysis algorithm effectively filters short-term anomalies and reduces the false alarm rate. Machine learning analysis is performed based on the PTP network abnormal event records and network performance analysis results to obtain network performance scores and fault location results. The machine learning algorithm greatly improves the accuracy and efficiency of fault root location by learning historical abnormal patterns, and converts complex multi-dimensional data into intuitive performance scores, which is convenient for management decisions. Network performance optimization and predictive maintenance are performed based on network performance scores and fault location results, and optimization suggestions and maintenance work orders are generated. With the help of time series prediction algorithms, network performance is predicted in a forward-looking manner, so that maintenance is transformed from passive response to active prevention. The simulation evaluation mechanism of optimization suggestions ensures the effectiveness and safety of optimization measures. The overall solution deeply integrates big data collection, artificial intelligence analysis and PTP network monitoring. Especially in the analysis of time series data, the application of long short-term memory networks gives full play to the advantages of deep learning algorithms in processing time series data, and can accurately capture complex patterns and long-term dependencies in PTP clock data; in the anomaly detection link, unsupervised learning algorithms can automatically identify abnormal patterns from massive data without pre-defining all abnormal types; in fault root cause analysis, causal reasoning algorithms can accurately trace abnormal propagation paths and locate the root causes. The characteristics of these artificial intelligence algorithms and models directly improve the accuracy, foresight and automation of PTP network monitoring, and realize the transformation from passive monitoring to active prevention.
[0037] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0038] The deployed collection server is connected to the optical fiber network card to access the IP data stream, and the hardware timestamp of the network data packet is obtained through the MII interface to obtain the timestamp information with nanosecond accuracy.
[0039] Parse the PTP synchronization message, follow-up message, delay request message and delay response message, extract the timestamp information, message sequence number and address information therein, and obtain the PTP message data;
[0040] Collect data from network devices through NetConf and SNMP protocols to obtain PTP lock status, Grandmaster ID and Clock ID, and obtain PTP status data of network devices;
[0041] For devices that do not support standard management protocols, data is obtained through the SSH protocol, and the data is converted into JSON format through regular expressions to obtain standardized device status data;
[0042] Collect the RTP stream data packets received by the terminal device, extract the source address, destination address, sequence number and RTP timestamp, and obtain the media stream clock data;
[0043] The PTP message data, the network device PTP status data, the standardized device status data and the media stream clock data are added with the collection timestamp and the collection point identification information to obtain the structured data stream.
[0044] Specifically, the acquisition server is deployed to connect the optical fiber network card to access the IP data stream, and the hardware timestamp of the network data packet is obtained through the MII interface to obtain the timestamp information with nanosecond accuracy. MII interface is the abbreviation of Media Independent Interface, which is defined by the IEEE-802.3 standard to describe the interface between the Ethernet transceiver and the network controller. The acquisition server is equipped with an optical fiber network card, through which the IP data is accessed. The data packet is first sent to the server IP decapsulation module, and then further processed. In order to obtain the timestamp with nanosecond accuracy, the acquisition server modifies the network driver and uses the MII interface timestamp to use an independent channel to send and receive messages, ensuring the high accuracy of timestamp capture. This process actually directly stamps the hardware-level timestamp when the network data packet just enters the physical layer, avoiding the problem of inaccurate timestamps caused by operating system scheduling delays. The acquisition server also sets a performance optimization module to increase the priority of synchronization processes and synchronization threads, and uses a high-precision time timer to obtain the current time value to ensure the accuracy of timestamps.
[0045] Parse the PTP synchronization message, follow message, delay request message and delay response message, extract the timestamp information, message sequence number and address information, and obtain the PTP message data. The PTP protocol is a precise time protocol defined by IEEE 1588-2008, and the SMPTE ST 2059-2 standard is a time synchronization standard for audio and video systems based on the IEEE 1588 standard. After receiving the PTP message, the acquisition server first identifies the message type, which is divided into four basic messages: the synchronization message (Sync) is used by the master clock to send time information to the slave clock, and the synchronization message contains the sending timestamp of the master clock; the follow message (Follow_Up) is used to transmit more accurate timestamps in the two-step mode; the delay request message (Delay_Req) is sent by the slave clock to measure the network transmission delay; the delay response message (Delay_Resp) is the master clock's response to the delay request, and contains the timestamp of the received delay request. When parsing these messages, the collection server extracts key fields from each message: timestamp information records the exact time when the message is sent or received, sequence numbers are used to identify and match corresponding request and response messages, and source and destination addresses identify the network location of the PTP device. These extracted fields form structured PTP message data.
[0046] Data collection of network devices is performed through NetConf and SNMP protocols to obtain PTP lock status, Grandmaster ID and Clock ID, and obtain PTP status data of network devices. NetConf is a network configuration protocol that uses XML formatted data to interact with network devices through a secure transmission protocol; SNMP is a simple network management protocol used for monitoring and management of network devices. The collection system first determines the devices in the network that support NetConf or SNMP, establishes connections with these devices, and then sends standardized query requests to obtain PTP-related information. The PTP lock status indicates whether the device is successfully synchronized with the master clock. The Grandmaster ID is the identifier of the most authoritative clock source in the entire PTP network, and the Clock ID is the PTP clock identifier of the device itself. The collection system also obtains information such as PTP mode (master clock, slave clock or boundary clock), step mode (one-step method or two-step method), current PTP system time, latest synchronization time and PTP port status to form complete PTP status data of network devices.
[0047] For devices that do not support standard management protocols, data is obtained through the SSH protocol, and the data is converted into JSON format through regular expressions to obtain standardized device status data. Some special devices or old devices may not support NetConf or SNMP protocols. The collection system uses the SSH protocol (Secure Shell Protocol) to directly log in to the command line interface of the device and execute commands to query the PTP status, such as "show ptp", "display ptp status", etc. The acquired data is usually unstructured text output, and the collection system uses pre-defined regular expression patterns to match relevant information. A regular expression is an expression used to match a specific pattern in a string, such as using pattern="GM ID: ([0-9a-fA-F:]+)" to match the Grandmaster ID. The information extracted after matching is then converted into a standard JSON format, such as {"device_id": "switch01", "gm_id": "01:02:03:04:05:06:07:08", "clock_state": "locked"}, to achieve the unification and standardization of data formats for subsequent processing.
[0048] The RTP stream data packets received by the acquisition terminal device are extracted, and the source address, destination address, sequence number and RTP timestamp are extracted to obtain the media stream clock data. RTP (Real-time Transport Protocol) is a protocol for transmitting audio and video data on an IP network and is widely used in IP production and broadcasting systems. The acquisition system captures the RTP data packets passing through the network, parses its header information, and extracts key fields: the source address refers to the source IP address of the streaming device, the source port is the source port of the streaming device, the destination address is usually a multicast address, the destination port is the destination port for sending traffic, the sequence number (Sequence Number) is used to identify the order of the data packets, the tag information (M and F) represents the mark tag and the field sequence tag respectively, and the most important is the RTP timestamp, which records the time point of media sampling. The acquisition system extracts these fields to form the media stream clock data, which will be used for comparative analysis with the PTP clock to evaluate the synchronization accuracy of the media stream and the PTP clock.
[0049] Add the acquisition timestamp and acquisition point identification information to the PTP message data, network device PTP status data, standardized device status data and media stream clock data to obtain a structured data stream. In this step, the acquisition system integrates and marks all the data obtained previously. First, add an acquisition timestamp to each piece of data to record the exact time of data acquisition, using a nanosecond time format such as "1613455782.123456789". Secondly, add the acquisition point identification information, such as "collector_node_001", to clarify the source of the data. Then, the system organizes different types of data according to a unified data structure template to form a standardized structured data stream, which includes fields such as data type identification, acquisition time, acquisition point, and raw data. The structured data stream uses JSON or similar formats to facilitate network transmission and database storage. Finally, these structured data are transmitted to the central data processing platform in real time, providing a basis for subsequent data preprocessing and analysis.
[0050] For example, an IP production and broadcasting center deployed this method for PTP network monitoring. The acquisition server is connected to the mirror port of the core switch and receives the PTP synchronization message through the fiber optic network card. The MII interface immediately stamps the message with a hardware timestamp "1613455782.123456789". The PTP synchronization message is parsed to extract the master clock sending time "1613455782.123455789", the message sequence number "45678", the source address "10.1.1.1" (master clock) and the destination address "224.0.1.129" (PTP multicast address). At the same time, the acquisition system queries the core switch through the SNMP protocol and obtains its PTP lock status as "Locked", Grandmaster ID as "00:01:02:03:04:05:06:07", and Clock ID as "00:A1:A2:A3:A4:A5:A6:A7". For an encoder that does not support SNMP, the system executes the command through the SSH protocol to obtain the text output "PTP Status: Synchronized to GM 00:01:02:03:04:05:06:07", uses regular expressions to extract information and converts it to JSON format. The system also captures the RTP stream data packets sent from the encoder, extracts the source address "10.2.2.2", the destination address "239.1.1.1", the sequence number "12345" and the RTP timestamp "9876543210". Finally, the collection timestamp "1613455782.123456789" and the collection point identifier "collector_001" are added to all data to form a complete structured data stream, which is transmitted to the central processing platform in real time for subsequent analysis.
[0051] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0052] Extracting the sending timestamp and receiving timestamp of the PTP synchronization message, as well as the sending timestamp of the delay request message and the receiving timestamp of the delay response message from the structured data stream to obtain a timestamp data set;
[0053] Calculate the path delay value based on the timestamp data set, add the synchronization message time difference and the delay message time difference and average them to obtain the network path delay parameter;
[0054] The clock offset is calculated based on the network path delay parameter, and the network path delay parameter is subtracted from the synchronization message time difference to obtain the original clock offset value;
[0055] Apply Kalman filter algorithm to the original clock offset value to perform de-jitter processing to obtain a stable clock deviation value;
[0056] Identify and remove outliers in structured data streams, unify data from different sources into a standard format, and obtain a cleaned data set;
[0057] The cleaned data set is associated and integrated with the stable clock deviation value to form a standardized data set that includes the PTP lock status of network equipment, clock source information, clock deviation value, jitter value and media delay value.
[0058] Specifically, the sending timestamp and receiving timestamp of the PTP synchronization message, the sending timestamp of the delay request message and the receiving timestamp of the delay response message are extracted from the structured data stream to obtain a timestamp data set. In this process, the processing program traverses each record in the structured data stream and filters out the PTP synchronization message (Sync), follow-up message (Follow_Up), delay request message (Delay_Req) and delay response message (Delay_Resp) according to the message type identifier. For each message, its key timestamp information is extracted: the master clock sending time t1 (sending timestamp) and the slave clock receiving time t2 (receiving timestamp) are extracted from the synchronization message. In the two-step method mode, the sending timestamp t1 is extracted from the follow message with a more accurate value; the slave clock sending time t3 (sending timestamp) is extracted from the delay request message, and the master clock receiving time t4 (receiving timestamp) is extracted from the delay response message. These four timestamps together constitute a complete PTP time synchronization cycle, forming a timestamp data set, and each data set contains four timestamp values associated with the same sequence number. Calculate the path delay value based on the timestamp data set, add the synchronization message time difference and the delay message time difference and average them to get the network path delay parameter. The specific calculation process is as follows: first calculate the time difference of the synchronization message, that is, the time t2 when the slave clock receives the synchronization message minus the time t1 when the master clock sends the synchronization message, and get the transmission time from the master to the slave plus the clock deviation; then calculate the time difference of the delay message, that is, the time t4 when the master clock receives the delay request message minus the time t3 when the slave clock sends the delay request message, and get the transmission time from the slave to the master minus the clock deviation; finally, add these two time differences and divide them by 2, that is, (t2-t1)+(t4-t3) / 2, to get the average network path delay value. This calculation method assumes that the network transmission delay is symmetrical in both directions. By averaging, the influence of clock deviation is eliminated, and the pure network transmission delay is obtained.
[0059] The clock offset is calculated based on the network path delay parameter. The network path delay parameter is subtracted from the synchronization message time difference to obtain the original clock offset value. The clock offset calculation formula is: offset = (t2-t1-delay), where t2 is the time when the slave clock receives the synchronization message, t1 is the time when the master clock sends the synchronization message, and delay is the network path delay parameter calculated in the previous step. This calculation process actually strips the pure network transmission delay from the total time difference between the master and slave clocks, and what remains is the actual deviation between the two clocks. The original clock offset value represents the deviation of the slave clock relative to the master clock. A positive value indicates that the slave clock is faster than the master clock, and a negative value indicates that the slave clock is slower than the master clock.
[0060] Apply the Kalman filter algorithm to the original clock offset value to remove jitter and obtain a stable clock offset value. Kalman filtering is a recursive optimal estimation algorithm that is particularly suitable for processing noisy time series data. In the PTP network, due to factors such as network jitter and load fluctuations, the original clock offset value often contains noise and instantaneous fluctuations. The Kalman filter algorithm works in two stages: prediction and update. The prediction stage predicts the state at the next moment based on the current state and error covariance; the update stage adjusts the prediction result based on the actual measurement value to obtain the optimal estimate. In the specific implementation, the clock state model is set to include two state variables, the offset value and the drift rate, and the measurement model is the direct observation of the offset value. The algorithm recursively calculates the Kalman gain and dynamically adjusts the degree of trust of the predicted value to the measured value, thereby smoothing the jitter in the original offset value and obtaining a more stable clock offset estimate.
[0061] Identify and remove outliers in structured data streams, unify data from different sources into a standard format, and obtain a cleaned data set. Multiple strategies are used to identify outliers: the statistical threshold method identifies deviation values that exceed the normal range, such as setting deviations exceeding ±100 microseconds as abnormal; mutation detection identifies situations where deviation values change dramatically in a short period of time; consistency checks verify the logical relationship between related data, such as the deviation values of different devices in the same period should have a certain correlation. Identified outliers are marked as invalid or replaced with interpolated estimates. Data format unification converts data obtained from different sources and different protocols into a unified structure, including standardized time format, unified deviation value units (nanoseconds), unified device identifier format, etc., to ensure data consistency and comparability.
[0062] The cleaned data set is associated and integrated with the stable clock deviation value to form a standardized data set containing the PTP lock status, clock source information, clock deviation value, jitter value and medium delay value of the network equipment. The association and integration process first matches the data according to the device identifier and timestamp, and associates different types of data of the same device at similar time points; then the stable clock deviation value processed by Kalman filtering is merged with other status data of the corresponding device; finally, additional indicators such as jitter value (standard deviation of clock deviation) and stability indicators (such as Allan variance) are calculated. The standardized data set adopts a unified data structure, contains rich metadata and indicators, and provides a comprehensive data foundation for subsequent PTP network performance analysis.
[0063] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0064] The standardized data set is stored in a time series database with a distributed storage architecture, and stored in layers according to the time span to obtain an efficient PTP data storage structure;
[0065] Extract features of clock deviation, jitter value and network delay data in the PTP data storage structure, construct input feature vectors, and obtain training data sets;
[0066] Based on the training data set, a multidimensional PTP network performance analysis model based on the long short-term memory network is constructed. The multidimensional PTP network performance analysis model includes an input layer, two long short-term memory hidden layers and a fully connected output layer, and an initial model structure is obtained;
[0067] The initial model structure is trained using historical data, and the mean square error loss function and Adam optimization algorithm are used to iteratively optimize the model parameters to obtain a trained multi-dimensional PTP performance analysis model.
[0068] Input the real-time standardized data set into the trained multi-dimensional PTP performance analysis model, and obtain the clock stability prediction value and network performance prediction value through forward propagation calculation;
[0069] Based on the clock stability prediction value and network performance prediction value, combined with the PTP clock distribution topology information, a network performance analysis result including clock stability evaluation, performance trend prediction and abnormal risk evaluation is generated.
[0070] Specifically, the standardized data set is stored in a time series database with a distributed storage architecture, and hierarchical storage is performed according to the time span to obtain an efficient PTP data storage structure. The time series database is a type of database designed specifically for processing time series data, and it is optimized for data with timestamps. In the PTP monitoring scenario, the distributed storage architecture disperses the storage pressure through multiple nodes and supports high-concurrency write and query operations. A hierarchical storage strategy is adopted in the data storage process, and the data is divided into different levels according to the time of oldness: the original high-precision data of the last 7 days is stored in a high-speed storage layer, such as SSD or memory database, retaining full accuracy for detailed analysis; data from 7 days to 30 days is moderately downsampled, aggregated once a minute, and stored in a medium-speed storage layer; historical data over 30 days is further downsampled, aggregated once an hour, and stored in a low-cost storage layer. Each piece of data has a time index and a multi-dimensional tag index, such as device ID, PTP role, etc., which is convenient for quick query and screening.
[0071] The clock deviation, jitter value and network delay data in the PTP data storage structure are feature extracted to construct input feature vectors and obtain training data sets. Feature extraction is the process of converting raw data into more meaningful feature representations. First, the time series features are extracted. For the clock deviation data of each device, the statistical features in the sliding window are calculated, including the average value, standard deviation, maximum value, minimum value, peak-to-peak value, rate of change, etc.; similarly, the statistical features of the jitter value and network delay data are extracted. Secondly, the frequency domain features are extracted. The time series is converted to the frequency domain through fast Fourier transform, and the main frequency components, power spectrum density and other features are extracted to identify periodic patterns and abnormal oscillations. The device topology features are also extracted, including the hierarchical position of the device in the PTP network, the master-slave relationship, the stability of the upstream clock source, etc. After feature extraction, normalization is performed to scale the features of different magnitudes to the same range to avoid the dimensional differences between the features affecting the model training. A training sample containing multi-dimensional features is formed, and each sample is associated with the network status at a time point and the performance indicators at subsequent time points as labels.
[0072] Based on the training data set, a multidimensional PTP network performance analysis model based on long short-term memory network is constructed. The multidimensional PTP network performance analysis model includes an input layer, two long short-term memory hidden layers and a fully connected output layer to obtain the initial model structure. Long short-term memory network (LSTM) is a special recurrent neural network that can effectively handle long-term dependencies in time series data. The model structure is designed as follows: the input layer receives feature vectors, and the dimension is consistent with the number of features; the first layer of LSTM contains 128 neurons, uses the tanh activation function, and sets the dropout rate to 0.2 to prevent overfitting; the second layer of LSTM contains 64 neurons, and also uses the tanh activation function and dropout mechanism; followed by a fully connected layer, which maps the LSTM output to the prediction target, such as clock stability indicators and network performance indicators; finally, the output layer sets different activation functions according to the prediction task type, using linear activation functions for regression tasks and sigmoid or softmax activation functions for classification tasks. The model design takes into account the temporal characteristics of PTP network data. The two-layer LSTM structure can capture short-term fluctuations and long-term trends at the same time, providing a basis for accurately predicting network performance.
[0073] The initial model structure was trained using historical data, and the mean square error loss function and Adam optimization algorithm were used to iteratively optimize the model parameters to obtain a trained multi-dimensional PTP performance analysis model. The model training process first divided the historical data set into a training set (70%), a validation set (15%), and a test set (15%) in chronological order to ensure the consistency of the data distribution of each set. The mean square error (MSE) was used as the loss function, which calculates the average of the sum of the squares of the difference between the predicted value and the true value. It is sensitive to outliers and helps to reduce large prediction errors. The Adam optimization algorithm combines the advantages of the momentum method and RMSProp, adaptively adjusts the learning rate, and accelerates the convergence process. The batch processing mechanism is used in the training process. The size of each batch of data is set to 64. The forward propagation is performed to calculate the loss, and then the model parameters are updated by back propagation. At the same time, the early stopping strategy is implemented to monitor the performance of the validation set. When the validation loss no longer decreases for 10 consecutive epochs, the training is stopped to prevent overfitting. The training process also includes learning rate scheduling. The initial learning rate is set to 0.001. When the validation loss does not improve for 3 consecutive epochs, the learning rate is reduced to the original 0.5.
[0074] The real-time standardized data set is input into the trained multi-dimensional PTP performance analysis model, and the clock stability prediction value and network performance prediction value are obtained through forward propagation calculation. In the real-time prediction process, the newly collected standardized data is first subjected to the same feature extraction and normalization processing as the training data to ensure the consistency of the data format. The prediction calculation adopts a sliding window mechanism, and the data of the most recent N time points are taken as input each time. The N value is consistent with the setting during model training (usually 48 points, corresponding to the data of the past 24 hours). The input data is calculated in sequence through each layer of the model: first, it is processed by two layers of LSTM to capture the timing features, and then mapped to the specific prediction target through the fully connected layer. The model output includes multiple dimensions: the clock stability prediction value includes the trend change of clock deviation in the next 24 hours, the expected maximum deviation value and the degree of deviation fluctuation; the network performance prediction value includes the trend of network delay change, the expected jitter level and the link stability evaluation. These prediction values serve as an important basis for real-time decision-making and guide network operation and optimization.
[0075] Based on the clock stability prediction value and the network performance prediction value, combined with the PTP clock distribution topology information, a network performance analysis result including clock stability evaluation, performance trend prediction and abnormal risk evaluation is generated. In this step, the prediction result is first compared with the historical baseline data to determine the percentile of the current performance level in the historical performance. The clock stability evaluation is based on the predicted deviation trend and fluctuation degree, calculates professional evaluation indicators such as Allan variance and maximum time interval error (MTIE), and compares them with the PTP specification requirements to form a compliance rating. Performance trend prediction analyzes the direction and rate of change of the predicted value, identifies the performance improvement or degradation trend, and estimates the time point when the critical threshold is reached. The abnormal risk assessment combines the PTP clock distribution topology to analyze potential fault points, simulates the propagation effect of single-point faults, evaluates the impact of different node or link faults on the performance of the entire network, identifies key nodes and calculates the risk level. The final generated network performance analysis result adopts a multi-level structure, which includes both a macro evaluation of the entire network and a detailed analysis of each device node, providing comprehensive support for management decisions.
[0076] For example, a certain ultra-high-definition production and broadcasting center has deployed a PTP network performance monitoring method based on big data analysis. Through preliminary data collection and preprocessing, a standardized data set of three months has been accumulated, with sampling once per second, forming about 7.8 million records. These data are first stored in a distributed time series database according to the time span: about 600,000 raw data in the last 7 days are stored in high-speed SSD nodes, retaining complete nanosecond-level accuracy; data from 7 to 30 days are aggregated once a minute, and about 33,000 data are stored in medium-speed storage nodes; older historical data are aggregated once an hour, and about 2,000 data are stored in low-cost storage nodes. The database creates a time index and device label index for each record to achieve a millisecond-level query response speed. Feature extraction is performed on the data of the core switch, and 20 statistical features such as the mean, standard deviation, maximum value, minimum value, and linear regression slope of the clock deviation are calculated within a 48-hour sliding window. Features are also extracted for jitter values and network delays, and finally a feature vector of about 60 dimensions is obtained at each time point. Based on these feature vectors, a deep learning model with two layers of LSTM with 128 nodes and 64 nodes was constructed. The model was trained using the data from the first two months, and the model performance was verified using the data from the last month. The training used the Adam optimizer with a batch size of 64 and an initial learning rate of 0.001. After about 500 iterations, the error of the validation set was stabilized within 50 nanoseconds. After the model was put into use, the clock stability change trend of the device in the next 24 hours was predicted in real time. When the clock deviation prediction curve of a certain boundary clock device was monitored to show a slow but continuous upward trend, and it was expected to exceed the 500 nanosecond threshold after 72 hours, the system immediately generated an early warning message and determined through topological analysis that the device was the upstream clock source of 20 terminal nodes, with a risk rating of "high". Based on this, the maintenance personnel arranged an equipment inspection in advance and found that the abnormal temperature control system of the device caused the crystal oscillator frequency to drift. The faulty component was replaced in time to avoid potential large-scale synchronization interruption events.
[0077] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0078] Based on the network performance analysis results, a complete PTP distribution topology diagram is constructed to display the Grandmaster devices, boundary clock devices, and ordinary clock devices in the network and their master-slave relationships in real time, and obtain a real-time status diagram of the PTP network;
[0079] Color-code the PTP lock status, clock deviation value, and jitter value of each device in the real-time status diagram of the PTP network to obtain an intuitive device status view;
[0080] According to the network performance analysis results, a multi-level threshold judgment standard is set to perform real-time detection when the clock deviation change value exceeds the preset threshold, and obtain the clock jump abnormality mark;
[0081] Based on the lock status information in the network performance analysis results, identify the change of the device PTP lock status or the situation of being unlocked for a long time, and obtain an abnormal lock status alarm;
[0082] Judge the clock deviation trend in the network performance analysis results, identify the situation where the unidirectional continuous growth trend exceeds the preset drift rate, and obtain the clock drift abnormality mark;
[0083] The clock jump abnormality mark, lock state abnormality alarm and clock drift abnormality mark are integrated and recorded in time series, and correlation analysis is applied to filter short-term abnormalities to obtain the PTP network abnormal event record.
[0084] Specifically, the Grandmaster devices, boundary clock devices and ordinary clock devices in the network and their master-slave relationships are displayed in real time to obtain a real-time status diagram of the PTP network. The PTP distribution topology diagram is a network structure diagram that reflects the hierarchical relationship of PTP clock synchronization. It intuitively shows how the clock signal is distributed from the master clock source to each node in the network. During the construction process, the basic information of all PTP devices is first extracted from the network performance analysis results, including device identifiers, IP addresses, MAC addresses, device types, and PTP roles. Then the master-slave relationship data between the devices is extracted, and the connection relationship is established by analyzing the Grandmaster ID and upstream clock source information of each device to determine the flow direction of the synchronization signal. For the Grandmaster device, that is, the device that serves as the master clock source, it is placed at the top level of the topology diagram by matching its Clock ID with the Grandmaster ID reported by other devices; for the boundary clock device, its upstream clock source and downstream slave devices are identified and placed in the middle layer of the topology diagram; for the ordinary clock device, its upstream clock source is determined and placed at the end of the topology diagram. Through this data association and hierarchical analysis, a complete tree or mesh topology is formed, which reflects the synchronization path and clock distribution status of the entire PTP network in real time.
[0085] The PTP lock status, clock deviation value, and jitter value of each device in the real-time status diagram of the PTP network are color-coded to obtain an intuitive view of the device status. Color coding is a method of converting numerical values or states into visual colors, which facilitates intuitive identification of device status. For the PTP lock status, the typical color coding scheme is: green is used for the locked state, indicating that the device has successfully synchronized with the master clock; red is used for the unlocked state, indicating that the device has failed to establish synchronization with the master clock; yellow is used for the locked state, indicating that the device is in the synchronization process but has not yet stabilized. For the clock deviation value, a gradient color spectrum is used for coding: green indicates that the deviation value is within ±100 nanoseconds, meeting high-precision requirements; yellow indicates that the deviation value is between ±100 nanoseconds and ±500 nanoseconds, which is in the warning range; orange indicates that the deviation value is between ±500 nanoseconds and ±1 microseconds, close to the critical value; red indicates that the deviation value exceeds ±1 microsecond and has exceeded the normal range. For jitter values, a gradient color spectrum is also used: blue to green indicates that the jitter value is less than 50 nanoseconds, which is in good condition; yellow to orange indicates that the jitter value is between 50 nanoseconds and 200 nanoseconds, which requires attention; red indicates that the jitter value exceeds 200 nanoseconds, which may affect the synchronization quality. Through this color coding, network administrators can identify problematic devices and links at a glance.
[0086] According to the results of network performance analysis, multi-level threshold judgment criteria are set, and real-time detection is performed for situations where the clock deviation change value exceeds the preset threshold, and a clock jump abnormality mark is obtained. Clock jump refers to the phenomenon that the clock deviation of the PTP device changes significantly in a short period of time, which is usually caused by network congestion, hardware failure or external interference. In order to detect clock jumps, multi-level threshold judgment criteria are first set for different types of devices: the jump threshold of core devices (such as backbone switches) is set lower, usually 100 nanoseconds, because such devices have higher requirements for time accuracy; the jump threshold of edge devices is relatively loose and can be set to 500 nanoseconds. During the real-time detection process, the clock deviation of each device is continuously monitored, and the difference between two adjacent measurement values, that is, the deviation change rate, is calculated. When the deviation change rate exceeds the preset threshold, the clock jump abnormality mark is triggered. To reduce false alarms, the duration of the change must also be considered: instantaneous jumps (duration less than 100 milliseconds) are marked as low-level abnormalities; continuous jumps (duration longer than 100 milliseconds but less than 1 second) are marked as medium-level abnormalities; long-term jumps (duration longer than 1 second) are marked as high-level abnormalities. Each anomaly tag contains information such as device identification, occurrence time, duration, jump amplitude and severity.
[0087] Based on the lock status information in the network performance analysis results, the PTP lock status change or long-term unlocking of the device is identified to obtain a lock status abnormality alarm. The PTP lock status refers to the synchronization status of the device and the master clock, including locked, unlocked, and locked intermediate states. The lock status abnormality is divided into two categories: one is the lock status change, that is, the device changes from the locked state to the unlocked state, indicating that the synchronization is interrupted; the other is long-term unlocking, that is, the device fails to establish synchronization within the expected time. The identification process first extracts the lock status history of each device from the network performance analysis results to establish a state time series. For the detection of lock status changes, the state values at adjacent time points are compared. When the state changes from locked to unlocked, a state change abnormality mark is generated. For the detection of long-term unlocking, the cumulative time that the device is in the unlocked state is calculated. When it exceeds the preset time threshold (usually 30 seconds to 5 minutes depending on the device type), a long-term unlocking abnormality mark is generated. The abnormality judgment also considers the importance of the device in the PTP network: the abnormal lock status of the device on the core path will trigger a high priority alarm; the abnormal lock status of the edge device will trigger a low priority alarm. Each lock status abnormality alarm contains information such as device identification, abnormality type, occurrence time, duration and impact range.
[0088] The clock deviation trend in the network performance analysis results is judged, and the situation where the unidirectional continuous growth trend exceeds the preset drift rate is identified to obtain a clock drift anomaly mark. Clock drift refers to the phenomenon that the clock deviation of the device shows a continuous unidirectional change, which is usually caused by crystal oscillator aging, temperature change or hardware failure. The process of identifying clock drift first extracts the clock deviation time series data of each device from the network performance analysis results. For each device, the deviation data of the last N time points (usually the past 1 hour) is analyzed using the linear regression method to calculate the drift rate (slope). The linear regression equation is y=ax+b, where y represents the deviation value, x represents the time, and a is the drift rate. When the absolute value of the calculated drift rate exceeds the preset threshold (usually 100 nanoseconds per hour) and the correlation coefficient R² is greater than 0.7 (indicating a good fit), it is determined that there is a drift trend. In order to improve the accuracy of the judgment, the persistence of the drift also needs to be considered: multiple consecutive time windows are analyzed, and a clock drift anomaly mark is generated only when multiple consecutive windows (usually 3) detect a drift trend in the same direction. Each drift anomaly mark contains information such as device identification, occurrence time, drift rate, fit, and estimated time to reach the critical deviation value.
[0089] The clock jump abnormality mark, lock state abnormality alarm and clock drift abnormality mark are integrated and recorded in time series, and the correlation analysis is applied to filter the short-term abnormality to obtain the PTP network abnormality event record. The integrated record is to merge different types of abnormal information into a unified time series database to facilitate global analysis and association mining. The integration process first establishes a unified abnormal event data structure, which contains fields such as event ID, device ID, abnormality type, start time, end time, severity and detailed parameters. Then all abnormality marks are inserted into this unified structure in chronological order to form a complete abnormal event timeline. In order to filter short-term abnormalities and reduce false alarms, correlation analysis technology is applied: first, time correlation analysis is performed to identify abnormalities that appear multiple times and disappear quickly within a short time window (usually 1 second), and regard them as instantaneous fluctuations rather than real abnormalities; secondly, topological correlation analysis is performed. When an abnormality occurs in the upstream device of a device, similar abnormalities that occur subsequently in the device are regarded as cascading effects rather than independent events; finally, pattern correlation analysis is performed to match historically known abnormal patterns with current abnormalities to improve judgment accuracy. After correlation analysis and filtering, the retained abnormal events form the final PTP network abnormal event records, each of which contains a complete abnormal description and context information.
[0090] For example, a certain ultra-high-definition video live broadcast system runs a PTP clock synchronization network and deploys a PTP network performance monitoring method based on big data analysis. Before a live broadcast event, the monitoring system extracted the PTP network structure information including 1 Grandmaster device, 8 boundary clock devices and 42 ordinary clock devices from the network performance analysis results, and generated a complete distribution topology map. The topology map clearly shows the path of the master clock signal from the Grandmaster device through the core switch and the secondary switch to each terminal device. Color coding is applied to the topology map, and the 41 devices in normal operation are displayed in green, of which 1 boundary clock device is displayed in yellow (indicating that the clock deviation is about 300 nanoseconds), and 2 ordinary clock devices are displayed in red (indicating an unlocked state). The monitoring system performs clock jump detection on the boundary clock device marked in yellow and finds that its clock deviation suddenly jumps from 50 nanoseconds to 300 nanoseconds in the past 10 minutes. The change of 250 nanoseconds exceeds the preset threshold of 200 nanoseconds for the device. The system immediately generates a clock jump abnormality mark. At the same time, the lock status of the two red devices was analyzed and it was found that they had been unlocked for 45 consecutive minutes, exceeding the preset threshold of 30 minutes, and the system generated a long-term unlock alarm. The clock deviation trend of the yellow device was further analyzed, and the drift rate of the last 60 data points was calculated using linear regression. The results showed that the deviation value continued to increase at a rate of about 120 nanoseconds per hour, and the fit R² reached 0.85, exceeding the drift threshold of 100 nanoseconds per hour. The system generated a clock drift anomaly mark. These three types of abnormal information were integrated and recorded by time. Through correlation analysis, it was found that the abnormality of the yellow device was highly correlated with the abnormality of the two red devices downstream, and finally a complete abnormal event record was generated: "The clock drift of the boundary clock device BC03 caused the synchronization failure of the downstream OC12 and OC17 devices." Based on this precise abnormal location, the maintenance personnel quickly checked the BC03 device and found that its temperature abnormality caused the crystal oscillator frequency to drift. After replacing the relevant components, the synchronization status of the entire network returned to normal, ensuring the smooth progress of the live broadcast event.
[0091] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0092] Extract clock stability data, synchronization accuracy data, network transmission data and system reliability data from the network performance analysis results, perform normalization processing, and obtain standardized feature data;
[0093] The stability score is calculated for the standard deviation and maximum deviation value of the clock deviation in the standardized feature data, and the accuracy score is calculated for the average deviation value and deviation distribution to obtain the sub-item score of the clock performance;
[0094] The transmission score is calculated based on the PTP message packet loss rate and delay change rate in the standardized feature data, and the reliability score is calculated based on the device state transition frequency and abnormal event frequency to obtain the network performance sub-item score;
[0095] The clock performance sub-item score and the network performance sub-item score are calculated comprehensively through a weighted calculation method to obtain a network performance score of 0-100;
[0096] Extract the abnormal occurrence time, abnormal type, involved equipment and abnormal parameter value from the abnormal event record of the PTP network, generate an abnormal feature vector, and obtain abnormal feature data;
[0097] Perform root cause analysis based on abnormal feature data, identify the device or link where the abnormality first occurred, and query the historical fault processing database to obtain the fault location result.
[0098] Specifically, clock stability data, synchronization accuracy data, network transmission data and system reliability data are extracted from the network performance analysis results, and normalized to obtain standardized feature data. Clock stability data includes indicators such as the standard deviation of clock deviation, maximum deviation value, Allan variance, deviation drift rate, etc., which reflect the stability of the device clock operation; synchronization accuracy data includes indicators such as average deviation value, deviation distribution, maximum time interval error (MTIE), etc., which reflect the accuracy of device clock synchronization; network transmission data includes indicators such as PTP message packet loss rate, delay change rate, message round-trip time, etc., which reflect the quality of network transmission; system reliability data includes indicators such as device state transition frequency, abnormal event frequency, lock stabilization time, etc., which reflect the overall reliability of the system. Normalization is the process of converting data of different dimensions and orders of magnitude into a unified range to ensure that the weight of each indicator is reasonable. For each indicator, we first determine its theoretical maximum and minimum values, and then use a linear normalization formula or interval mapping method to convert it to the [0,1] interval. For example, for the clock deviation standard deviation, when the original value is 5 nanoseconds, it is normalized to 0.95; when the original value is 500 nanoseconds, it is normalized to 0.1, reflecting the evaluation standard of "the smaller the value, the better".
[0099] The stability score is calculated for the standard deviation and maximum deviation of the clock deviation in the standardized feature data, and the precision score is calculated for the average deviation and deviation distribution to obtain the sub-item score of the clock performance. In the process of calculating the clock stability score, the weight of each indicator is first determined: the standard deviation of the clock deviation has a higher weight (such as 0.6) because it directly reflects the clock stability; the maximum deviation value has the second highest weight (such as 0.3), reflecting the performance in the worst case; the weight of supplementary indicators such as Allan variance is lower (such as 0.1) as an auxiliary evaluation basis. Then the corresponding weight is applied to each indicator, and the weighted sum is obtained to obtain the original stability score in the interval [0,1]. In order to better meet the actual evaluation needs, a nonlinear transformation function such as the S-type function is applied to the original score, so that the performance close to the ideal state obtains a higher score, while the performance score close to the critical value drops rapidly. Similarly, in the process of calculating the precision score, the average deviation value has a higher weight (such as 0.5), reflecting the overall synchronization accuracy; the deviation distribution has the second highest weight (such as 0.3), reflecting the stability of the accuracy; the maximum time interval error has a lower weight (such as 0.2). Through weighted calculation and nonlinear adjustment, the accuracy score in the interval [0,1] is finally obtained. These two scores are weighted and combined again to form a clock performance sub-score, which reflects the overall performance level of the device clock.
[0100] The transmission score is calculated for the PTP message packet loss rate and delay change rate in the standardized feature data, and the reliability score is calculated for the device state transition frequency and abnormal event frequency to obtain the network performance sub-item score. In the transmission score calculation, the PTP message packet loss rate has the highest weight (such as 0.5), because packet loss directly affects the synchronization accuracy; the delay change rate has the second highest weight (such as 0.3), reflecting the stability of network transmission; the message round trip time has a lower weight (such as 0.2) as an auxiliary evaluation basis. These indicators are weighted and converted into a transmission score in the [0,1] interval by applying a logarithmic transformation function. The score approaches 1 when the packet loss rate is 0, and the score drops rapidly to below 0.5 when the packet loss rate exceeds 1%. In the reliability score calculation, the device state transition frequency and abnormal event frequency are both important indicators. The former reflects the stability of the device state, and the latter reflects the frequency of abnormal occurrence. For the device state transition frequency, the fewer the number of state changes per hour, the higher the score; for the abnormal event frequency, the number of various abnormal events in the past 24 hours is counted, and the fewer the number, the higher the score. These two indicators are weighted to form a reliability score in the [0,1] interval. The transmission score and reliability score are weighted again to form a network performance sub-score, which comprehensively reflects the network transmission and system reliability status.
[0101] The clock performance sub-item score and the network performance sub-item score are calculated by weighted calculation method to obtain a network performance score of 0-100. In the weighted calculation process, the weights of the two major categories of scores are first determined: in high-precision PTP application scenarios, the clock performance weight is usually higher (such as 0.7) because clock synchronization accuracy is the core goal; the network performance weight is lower (such as 0.3) as a supporting factor. The two scores are weighted and summed to obtain a comprehensive score in the [0,1] interval. This score is then converted to the 0-100 interval through linear mapping to form the final network performance score. This score adopts a multi-level structure, which can not only show the overall score of the entire network, but also be refined to the single score at the device level, which is convenient for accurately locating problems. The scoring system is also designed to take into account the differences in requirements of different application scenarios: the broadcast production environment has extremely high requirements for clock performance, and the weight setting is biased towards clock stability and accuracy; while the general industrial control environment has high requirements for network reliability, and the weight setting is biased towards transmission quality and system reliability. The scoring results are displayed through an intuitive dashboard, with different score ranges corresponding to different colors: 90-100 points are green, indicating excellent; 70-89 points are yellow, indicating good; 50-69 points are orange, indicating average; 0-49 points are red, indicating poor.
[0102] The abnormal occurrence time, abnormality type, involved devices and abnormal parameter values are extracted from the abnormal event records of the PTP network to generate abnormal feature vectors and obtain abnormal feature data. The abnormal feature vector is a structured data representation used to describe the key features of abnormal events. First, the abnormal event records are extracted in chronological order, and each record contains complete abnormal description information. For each record, the following key information is extracted: abnormal occurrence time, accurate to milliseconds; abnormality type, such as clock jump abnormality, lock state abnormality, clock drift abnormality, etc.; involved devices, including device identifiers, device types, PTP roles, etc.; abnormal parameter values, such as clock deviation change amplitude, drift rate, lock state duration, etc. This information is organized into a fixed-dimensional feature vector according to a predefined format, such as [timestamp, abnormality type code, device ID, parameter value 1, parameter value 2, ...]. In order to improve processing efficiency, the abnormality type is encoded: clock jump abnormality is encoded as 1, lock state abnormality is encoded as 2, clock drift abnormality is encoded as 3, and compound abnormality is encoded as a combination of them. In addition, for consecutive anomalies of the same type, time compression processing is performed to record the start time, end time and peak parameters of the anomaly to reduce data redundancy. The final anomaly feature data is a multidimensional matrix, where each row represents an abnormal event and each column represents a feature dimension, providing structured input for subsequent root cause analysis.
[0103] Based on the abnormal feature data, root cause analysis is performed to identify the device or link that first experienced an abnormality and query the historical fault processing database to obtain the fault location result. Root cause analysis is a process of tracing back from the appearance to the essential cause. For PTP network abnormalities, the time series analysis method is first applied to sort the abnormal feature data according to the time of abnormality occurrence to find the device or link that first experienced an abnormality, which is usually the origin of the fault. Then, the topological association analysis is applied to identify the propagation path of the abnormality in the network based on the PTP clock distribution topology: if an abnormality occurs in the upstream device, and then the downstream devices also experience abnormalities one after another, the upstream device is likely to be the root cause; if a device is abnormal, but its upstream and downstream devices are normal, the device itself is likely to have a problem. For complex situations, causal reasoning algorithms are applied to construct a causal graph of abnormal events, and the most likely root source device or link is inferred based on the Bayesian network model. After determining the potential root cause, the system queries the historical fault processing database to match similar abnormal patterns and their solutions. The database contains past fault cases, each of which contains abnormal features, root causes, and treatment methods. Through similarity calculation, the historical case most similar to the current anomaly is found, and its fault cause and solution are extracted to form a complete fault location result, including the faulty device or link, fault type, possible cause and recommended processing method.
[0104] For example, during the operation of the PTP network of a certain ultra-high-definition live broadcast system, the monitoring system extracted multiple performance data of the core switch HC-01 from the network performance analysis results: the standard deviation of the clock deviation is 15 nanoseconds, the maximum deviation value is 120 nanoseconds, the average deviation value is 35 nanoseconds, the deviation distribution presents a normal distribution, the PTP message packet loss rate is 0.02%, the delay change rate is 0.5%, the device state conversion frequency is once every 24 hours, and the abnormal event frequency is twice a week. These raw data are normalized and converted into standardized feature data: the normalized value of the clock deviation standard deviation is 0.85 (the smaller the original value, the better), the normalized value of the maximum deviation value is 0.78, the normalized value of the average deviation value is 0.82, the normalized value of the deviation distribution is 0.9, the normalized value of the PTP message packet loss rate is 0.95, the normalized value of the delay change rate is 0.85, the normalized value of the device state conversion frequency is 0.8, and the normalized value of the abnormal event frequency is 0.75. Based on these normalized data, the scores of each item are calculated: the clock stability score is 0.85×0.6 + 0.78×0.3 + 0.88×0.1 = 0.829, the accuracy score is 0.82×0.5 + 0.9×0.3 + 0.85×0.2 = 0.85, the transmission score is 0.95×0.5 + 0.85×0.3 + 0.9×0.2 = 0.915, and the reliability score is 0.8×0.5 + 0.75×0.5 = 0.775. The clock performance score is further calculated to be 0.829×0.5 + 0.85×0.5 = 0.84, and the network performance score is 0.915×0.6 + 0.775×0.4 = 0.859. The final comprehensive network performance score is (0.84×0.7 +0.859×0.3)×100 = 84.6 points, which is in the good range. However, a few days later, the system extracted an anomaly from the PTP network abnormal event record: the boundary clock device BC-02 had a clock drift abnormality at 10:15:30.500, with a drift rate of 150 nanoseconds per hour. After 30 minutes, the two downstream terminal devices OC-05 and OC-06 successively had lock state abnormalities at around 10:45:20. The system generated abnormal feature vectors [1677489330.5, 3, "BC-02", 150, 1800] and [1677491120, 2, "OC-05", 2700], [1677491150, 2, "OC-06", 2715]. The root cause analysis, based on the time series and topological relationship, determined that BC-02 was the first device to experience an abnormality, and its abnormal time was 30 minutes earlier than that of the downstream devices, which highly matched the fault propagation pattern. The system queried the historical fault database and found three historical cases with similar patterns, of which two were caused by temperature fluctuations causing crystal frequency drift.The final fault location result pointed out that the BC-02 device had a crystal oscillator frequency drift due to abnormal temperature, and it was recommended to check the device's cooling system and ambient temperature control. Based on this location check, the maintenance personnel found that the air outlet of the air conditioner in the cabinet where the BC-02 device was located was partially blocked. After cleaning, the problem was solved and the network performance score returned to normal levels.
[0105] In a specific embodiment, the process of executing step S106 may specifically include the following steps:
[0106] Based on historical data and current network status, a time series prediction algorithm is applied to predict future trends in device clock stability, network delay, and jitter, and obtain network performance prediction results.
[0107] Compare the network performance prediction results with the preset performance threshold, and when the prediction value is lower than the performance threshold, obtain a predictive maintenance demand list;
[0108] According to the fault location results and the low-scoring items in the network performance score, the corresponding optimization solutions are extracted from the optimization knowledge base to generate PTP configuration parameter optimization suggestions;
[0109] Conduct simulation evaluation on PTP configuration parameter optimization suggestions, calculate the performance improvement effect and potential risks after parameter adjustment, and obtain an optimization solution evaluation report;
[0110] Based on the predictive maintenance requirements list and fault location results, analyze the long-term trend of device clock stability, identify signs of hardware aging, and obtain device health assessment results;
[0111] Based on the optimization plan evaluation report and equipment health assessment results, a standardized maintenance work order containing maintenance objects, maintenance content and priority is generated.
[0112] Specifically, based on historical data and current network status, a time series prediction algorithm is applied to predict future trends in device clock stability, network delay, and jitter, and obtain network performance prediction results. A time series prediction algorithm is a collection of methods for analyzing time patterns in historical data and predicting future trends. In PTP network monitoring, a variety of prediction algorithms are used to process different types of time series data: for clock deviation data that exhibits obvious periodic or seasonal changes, a seasonal ARIMA (autoregressive integrated moving average) model is used, which can capture trends, periodicity, and random fluctuations in the data; for relatively stable network delay data, exponential smoothing methods are used, including simple exponential smoothing, linear trend method, and seasonal trend method; for jitter data that exhibits nonlinear characteristics, an LSTM (long short-term memory) neural network model is used. The prediction process first extracts the historical data of each device in the past 30 days from the time series database, and performs data cleaning and preprocessing, including removing outliers, filling missing values and standardization. Then, an appropriate prediction model is selected according to the data characteristics, and the model parameters are optimized, such as the difference order, autoregressive order and moving average order of the ARIMA model, or the number of network layers and neurons of the LSTM model. Finally, the trained model is used to predict the clock stability, network delay and jitter change trends of the device in the next 7 days to form a complete network performance prediction result.
[0113] The network performance prediction results are compared with the preset performance thresholds. When the predicted value is lower than the performance threshold, a predictive maintenance demand list is obtained. The performance threshold is the critical value of the performance indicator set according to different equipment types and application scenarios, including the clock deviation threshold (such as ±500 nanoseconds), jitter threshold (such as 200 nanoseconds) and network delay threshold (such as 5 milliseconds). The comparison process first analyzes the prediction results and extracts the predicted values at key time points, especially the turning points and peak points where the predicted values change significantly; then the predicted values of these key points are compared with the corresponding performance thresholds to analyze when the predicted values will exceed the threshold and the degree of breakthrough; finally, a risk assessment is performed, considering the duration, scope of impact and severity of the breakthrough threshold. When the predicted value of a device shows that it will be lower than (that is, the performance degrades to) the preset threshold in the future, the system adds the device to the predictive maintenance demand list. The list is sorted in order of the time when the threshold is expected to be broken. Each record contains the device identification, performance indicator name, predicted breakthrough time, predicted value change trend and recommended processing time window. This prediction-based maintenance demand identification method transforms maintenance work from passive response to active prevention, effectively avoiding system failures caused by sudden performance degradation.
[0114] According to the fault location results and the low-scoring items in the network performance score, the corresponding optimization solutions are extracted from the optimization knowledge base to generate PTP configuration parameter optimization suggestions. The optimization knowledge base is a structured database containing solutions to various problems in the PTP network, which consists of historical optimization experience, expert advice and best practices. The extraction process first analyzes the fault location results to identify the fault type and potential causes; then queries the items in the network performance score that are lower than the threshold (such as 70 points), such as the clock stability sub-item, the transmission quality sub-item, etc.; then uses this information as a search condition to query the optimization knowledge base to find the optimization solution with the highest matching degree. Different optimization suggestions are generated for different types of problems: For clock jump problems, it is recommended to adjust parameters such as the PTP announce interval, sync interval, and delay_req interval, such as reducing the announce interval from the default 2 seconds to 1 second to increase the frequency of clock status updates; for lock stability problems, it is recommended to adjust the PTP domain priority setting to ensure that a stable clock source is selected as the master clock; for synchronization problems caused by network congestion, it is recommended to apply QoS (quality of service) policies to PTP messages to increase their transmission priority. The generated PTP configuration parameter optimization suggestions are presented in a structured form, including parameter name, current value, recommended value, adjustment reason and expected effect, which is convenient for maintenance personnel to understand and implement. The PTP configuration parameter optimization suggestions are simulated and evaluated, and the performance improvement effect and potential risks after parameter adjustment are calculated to obtain an optimization scheme evaluation report. Simulation evaluation is a method of predicting the optimization effect through mathematical models or simulation systems before the actual implementation of optimization. The evaluation process first establishes a PTP network performance influencing factor model, which describes the relationship between each configuration parameter and performance index, such as the relationship between the announce interval and locking stability, the sync interval and synchronization accuracy, and the delay_req interval and delay measurement accuracy; then the parameter adjustment values in the optimization suggestions are input into the model to calculate the expected performance indicators after adjustment; then compared with the current performance indicators to quantify the improvement effect; finally, risk analysis is performed to evaluate the possible negative effects of parameter adjustment. For example, excessively reducing the announce interval will increase the network load, and excessively increasing the delay_req interval will reduce the delay measurement accuracy. The evaluation report contains analysis results from multiple dimensions: performance improvement estimation, which indicates the expected changes in various performance indicators after parameter adjustment; resource consumption evaluation, which analyzes the impact of parameter adjustment on network bandwidth, processor load and other resources; compatibility analysis, which evaluates the compatibility of parameter adjustment with existing devices; and stable period estimation, which predicts the time required for the system to reach a new stable state. These analysis results form a complete optimization solution evaluation report, providing a scientific basis for decision-making.
[0115] Based on the predictive maintenance demand list and fault location results, the long-term change trend of the device clock stability is analyzed to identify signs of hardware aging and obtain the device health assessment results. Hardware aging refers to the phenomenon that the performance of the device gradually degrades as the use time increases. In the PTP network, it is mainly manifested as the frequency drift of the clock crystal oscillator increases and the synchronization accuracy decreases. The analysis process first extracts the long-term (such as half a year or longer) clock stability data of the device from the timing database, including indicators such as clock deviation, drift rate and Allan variance; then, trend analysis methods such as linear regression or curve fitting are applied to identify the long-term change trend of the indicators; then, it is compared with the normal aging model of the device to determine whether there are signs of abnormal aging. Normal aging is usually manifested as slow and steady performance degradation, while abnormal aging is manifested as accelerated degradation or mutation. In addition, the analysis will also consider the impact of environmental factors on device performance, such as temperature fluctuations and power supply stability. Based on these analyses, a health assessment result is generated for each device that appears in the predictive maintenance demand list or fault location results, including the equipment health status classification (such as normal, attention, warning, dangerous), the estimated remaining service life, the main manifestations of aging and recommended treatment methods. These health assessment results provide data support for equipment lifecycle management and help formulate reasonable update and maintenance plans.
[0116] Based on the optimization plan evaluation report and the equipment health evaluation results, a standardized maintenance work order containing maintenance objects, maintenance content and priority is generated. The maintenance work order is a standardized document that guides maintenance personnel to perform specific maintenance tasks, and contains all necessary information and processes. The generation process first comprehensively analyzes the optimization plan evaluation report and the equipment health evaluation results to determine the type of maintenance operation that needs to be performed, such as configuration optimization, hardware inspection, component replacement, etc.; then, based on the operation type and equipment characteristics, standard operating procedures and technical requirements are extracted from the maintenance operation library; then the maintenance priority is determined, considering factors such as problem severity, impact scope, expected failure time and equipment importance; finally, a complete standardized maintenance work order is generated. The standardized work order contains the following contents: maintenance object information, which describes the equipment or system to be maintained in detail, including equipment identification, type, location and key configuration; maintenance content, which specifies the operation steps, technical requirements and precautions to be performed; a list of required tools and materials; expected working time and impact scope; maintenance priority and recommended execution time window; acceptance criteria and recovery process; associated fault analysis results and historical maintenance record references. Through this standardized maintenance work order management, the standardization, traceability and effectiveness of maintenance work are ensured, and experience data is accumulated for the handling of similar problems in the future.
[0117] For example, a PTP synchronization system running in an ultra-high-definition production and broadcasting network deployed a PTP network performance monitoring method based on big data analysis. The monitoring system continuously collected and stored network performance data for three months. In the most recent analysis, the ARIMA (2,1,1) time series prediction model was applied to the boundary clock device BC-05 in the core area to process its clock deviation historical data. The model first performed first-order difference processing on the original deviation data to obtain a stationary series, and then combined the second-order autoregressive term and the first-order moving average term to establish a prediction model. The prediction results show that the clock deviation value of BC-05 will gradually increase from the current 200 nanoseconds to 550 nanoseconds in the next five days, exceeding the 500 nanosecond performance threshold set by the system. At the same time, the exponential smoothing method was applied to the network delay data of the device, showing that the delay will remain within the normal range. The system added BC-05 to the predictive maintenance needs list, marked it as "clock stability warning", and recommended that it be processed within three days. Combined with the previous fault location results, the clock performance score of the device was only 65 points, which is a low score. The system retrieved two matching optimization solutions from the optimization knowledge base: one is to adjust the PTP configuration parameters, shortening the sync interval from 1 second to 0.5 seconds and the delay_req interval from 2 seconds to 1 second; the other is to adjust the temperature control parameters of the device, because historical data shows that the performance fluctuation of the device is highly correlated with the ambient temperature change. The simulation evaluation of these two optimization solutions is expected to reduce the deviation value to about 150 nanoseconds after the configuration parameter adjustment, but increase the network traffic by about 1%; the temperature control optimization is expected to reduce the deviation fluctuation by 50%, but requires physical operation and a short device offline. At the same time, the long-term trend analysis shows that the frequency drift rate of the clock crystal oscillator of BC-05 has accelerated in the past three months, from 50 nanoseconds per month to the current 120 nanoseconds per month, which is in line with the typical characteristics of crystal oscillator aging. The equipment health assessment results classify BC-05 as "warning" state, and the remaining effective life of the crystal oscillator is expected to be less than 6 months. Based on these analysis results, the system generated a standardized maintenance work order: the maintenance object was "Boundary Clock Device BC-05", the maintenance content included "PTP parameter optimization configuration" and "Crystal oscillator detection and scheduled replacement", the priority was set to "medium-high", and it was recommended to complete the first parameter optimization within 48 hours and replace the crystal oscillator component in the next scheduled maintenance window. The maintenance personnel performed parameter optimization according to the work order requirements, and the clock deviation immediately dropped to 140 nanoseconds. The crystal oscillator was replaced in the maintenance window one week later, effectively avoiding potential synchronization failures and ensuring stable system operation.
[0118] The above describes the PTP network performance monitoring method based on big data analysis in the embodiment of the present application. The following describes the PTP network performance monitoring system based on big data analysis in the embodiment of the present application. Figure 2In the embodiment of the present application, an embodiment of a PTP network performance monitoring system based on big data analysis includes:
[0119] The collection module 201 is used to collect PTP message data, network device PTP status data and RTP stream data packets in the PTP network through multiple data collection points to obtain a structured data stream;
[0120] A calculation module 202, configured to perform data preprocessing and clock error calculation on the structured data stream to obtain a standardized data set including PTP network status parameters;
[0121] Establishing module 203, used to build a PTP network performance time series database based on the standardized data set, and establish a multi-dimensional PTP network performance analysis model to obtain network performance analysis results;
[0122] A detection module 204 is used to perform real-time status monitoring and anomaly detection according to the network performance analysis result to obtain a PTP network abnormal event record;
[0123] An execution module 205 is used to perform machine learning analysis based on the PTP network abnormal event record and the network performance analysis result to obtain a network performance score and a fault location result;
[0124] The maintenance module 206 is used to perform network performance optimization and predictive maintenance according to the network performance score and the fault location result, and generate optimization suggestions and maintenance work orders.
[0125] Through the collaboration of the above components, PTP message data, PTP status data of network equipment and RTP stream data packets in the PTP network are collected through multiple data collection points to obtain structured data streams, realize comprehensive and accurate data collection, and provide a high-quality data foundation for subsequent analysis. Data preprocessing and clock error calculation are performed on the collected structured data streams to obtain a standardized data set containing PTP network status parameters, which effectively eliminates the impact of network jitter and improves the accuracy of clock deviation calculation. Based on the standardized data set, a PTP network performance time series database is constructed, and a multi-dimensional PTP network performance analysis model is established to obtain network performance analysis results. Through the distributed storage architecture and time span hierarchical storage strategy, efficient storage and query of massive time series data are achieved, and the long short-term memory network model gives full play to the advantages of artificial intelligence algorithms in time series data prediction, accurately captures the long-term dependencies and short-term fluctuation characteristics in the PTP network, and significantly improves the accuracy and timeliness of performance prediction. According to the network performance analysis results, real-time status monitoring and anomaly detection are performed to obtain PTP network abnormal event records, realize intuitive visualization of network status and accurate anomaly detection, and the correlation analysis algorithm effectively filters short-term anomalies and reduces the false alarm rate. Machine learning analysis is performed based on the PTP network abnormal event records and network performance analysis results to obtain network performance scores and fault location results. The machine learning algorithm greatly improves the accuracy and efficiency of fault root location by learning historical abnormal patterns, and converts complex multi-dimensional data into intuitive performance scores, which is convenient for management decisions. Network performance optimization and predictive maintenance are performed based on network performance scores and fault location results, and optimization suggestions and maintenance work orders are generated. With the help of time series prediction algorithms, network performance is predicted in a forward-looking manner, so that maintenance is transformed from passive response to active prevention. The simulation evaluation mechanism of optimization suggestions ensures the effectiveness and safety of optimization measures. The overall solution deeply integrates big data collection, artificial intelligence analysis and PTP network monitoring. Especially in the analysis of time series data, the application of long short-term memory networks gives full play to the advantages of deep learning algorithms in processing time series data, and can accurately capture complex patterns and long-term dependencies in PTP clock data; in the anomaly detection link, unsupervised learning algorithms can automatically identify abnormal patterns from massive data without pre-defining all abnormal types; in fault root cause analysis, causal reasoning algorithms can accurately trace abnormal propagation paths and locate the root causes. The characteristics of these artificial intelligence algorithms and models directly improve the accuracy, foresight and automation of PTP network monitoring, and realize the transformation from passive monitoring to active prevention.
[0126] above Figure 2The PTP network performance monitoring system based on big data analysis in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The PTP network performance monitoring device based on big data analysis in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0127] Figure 3 It is a structural schematic diagram of a PTP network performance monitoring device based on big data analysis provided by an embodiment of the present invention. The PTP network performance monitoring device 300 based on big data analysis may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 (for example, one or more mass storage device terminals) storing application programs 333 or data 332. Among them, the memory 320 and the storage medium 330 can be short-term storage or permanent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the PTP network performance monitoring device 300 based on big data analysis. Furthermore, the processor 310 can be configured to communicate with the storage medium 330, and execute a series of instruction operations in the storage medium 330 on the PTP network performance monitoring device 300 based on big data analysis to implement the steps of the above-mentioned PTP network performance monitoring method based on big data analysis.
[0128] The PTP network performance monitoring device 300 based on big data analysis may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will appreciate that Figure 3 The structure of the PTP network performance monitoring device based on big data analysis shown does not constitute a limitation on the PTP network performance monitoring device based on big data analysis provided by the present invention, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.
[0129] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the PTP network performance monitoring method based on big data analysis.
[0130] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0131] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or partly contributed to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a PTP network performance monitoring device based on big data analysis (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.
[0132] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A PTP network performance monitoring method based on big data analysis, characterized in that: The method comprises: Collect PTP message data, network device PTP status data and RTP stream data packets in the PTP network through multiple data collection points to obtain a structured data stream; Performing data preprocessing and clock error calculation on the structured data stream to obtain a standardized data set including PTP network status parameters; Building a PTP network performance time series database based on the standardized data set, and establishing a multidimensional PTP network performance analysis model to obtain network performance analysis results; Perform real-time status monitoring and anomaly detection according to the network performance analysis results to obtain a PTP network abnormal event record; Perform machine learning analysis based on the PTP network abnormal event record and the network performance analysis result to obtain a network performance score and a fault location result; Network performance optimization and predictive maintenance are performed according to the network performance score and the fault location result, and optimization suggestions and maintenance work orders are generated.
2. The PTP network performance monitoring method based on big data analysis according to claim 1, characterized in that, The method collects PTP message data, network device PTP status data and RTP stream data packets in the PTP network through multiple data collection points to obtain a structured data stream, including: The deployed collection server is connected to the optical fiber network card to access the IP data stream, and the hardware timestamp of the network data packet is obtained through the MII interface to obtain the timestamp information with nanosecond accuracy. Parse the PTP synchronization message, follow-up message, delay request message and delay response message, extract the timestamp information, message sequence number and address information therein, and obtain the PTP message data; Collect data from network devices through NetConf and SNMP protocols to obtain PTP lock status, GrandmasterID and Clock ID, and obtain PTP status data of network devices; For devices that do not support standard management protocols, data is obtained through the SSH protocol, and the data is converted into JSON format through regular expressions to obtain standardized device status data; Collect the RTP stream data packets received by the terminal device, extract the source address, destination address, sequence number and RTP timestamp, and obtain the media stream clock data; The PTP message data, the network device PTP status data, the standardized device status data and the media stream clock data are added with a collection timestamp and a collection point identification information to obtain a structured data stream.
3. The PTP network performance monitoring method based on big data analysis according to claim 1, characterized in that, The data preprocessing and clock error calculation are performed on the structured data stream to obtain a standardized data set containing PTP network status parameters, including: Extracting the sending timestamp and receiving timestamp of the PTP synchronization message, as well as the sending timestamp of the delay request message and the receiving timestamp of the delay response message from the structured data stream to obtain a timestamp data set; Calculating the path delay value according to the timestamp data set, adding and averaging the synchronization message time difference and the delay message time difference to obtain the network path delay parameter; Calculating a clock offset based on the network path delay parameter, and subtracting the network path delay parameter from the synchronization message time difference to obtain an original clock offset value; Applying a Kalman filter algorithm to the original clock offset value to perform a de-jitter process to obtain a stable clock deviation value; Identify and remove outliers in the structured data stream, unify data from different sources into a standard format, and obtain a cleaned data set; The cleaned data set is associated and integrated with the stable clock deviation value to form a standardized data set including the PTP locking state of the network device, clock source information, clock deviation value, jitter value and medium delay value.
4. The PTP network performance monitoring method based on big data analysis according to claim 1, characterized in that, The PTP network performance time series database is constructed based on the standardized data set, and a multi-dimensional PTP network performance analysis model is established to obtain network performance analysis results, including: The standardized data set is stored in a time series database using a distributed storage architecture, and hierarchical storage is performed according to the time span to obtain an efficient PTP data storage structure; Perform feature extraction on the clock deviation, jitter value and network delay data in the PTP data storage structure, construct an input feature vector, and obtain a training data set; Based on the training data set, a multidimensional PTP network performance analysis model based on a long short-term memory network is constructed, wherein the multidimensional PTP network performance analysis model comprises an input layer, two long short-term memory hidden layers and a fully connected output layer, to obtain an initial model structure; The initial model structure is trained using historical data, and the model parameters are iteratively optimized using a mean square error loss function and an Adam optimization algorithm to obtain a trained multidimensional PTP performance analysis model; Inputting the real-time standardized data set into the trained multi-dimensional PTP performance analysis model, and obtaining a clock stability prediction value and a network performance prediction value through forward propagation calculation; According to the clock stability prediction value and the network performance prediction value, combined with the PTP clock distribution topology information, a network performance analysis result including clock stability evaluation, performance trend prediction and abnormal risk evaluation is generated.
5. The PTP network performance monitoring method based on big data analysis according to claim 1, characterized in that, The real-time status monitoring and anomaly detection are performed according to the network performance analysis result to obtain a PTP network abnormal event record, including: Based on the network performance analysis results, a complete PTP distribution topology diagram is constructed to display the Grandmaster devices, boundary clock devices and ordinary clock devices in the network and their master-slave relationships in real time, and obtain a real-time status diagram of the PTP network; The PTP lock status, clock deviation value and jitter value of each device in the PTP network real-time status diagram are color-coded to obtain an intuitive device status view; According to the network performance analysis result, a multi-level threshold judgment standard is set to perform real-time detection on the situation where the clock deviation change value exceeds the preset threshold, and obtain a clock jump abnormality mark; Based on the locking status information in the network performance analysis result, the change of the PTP locking status of the device or the situation of being unlocked for a long time is identified, and an abnormal locking status alarm is obtained; Judging the clock deviation trend in the network performance analysis result, identifying the situation where the unidirectional continuous growth trend exceeds the preset drift rate, and obtaining a clock drift abnormality mark; The clock jump abnormal mark, the lock state abnormal alarm and the clock drift abnormal mark are integrated and recorded in time series, and correlation analysis is applied to filter short-term abnormalities to obtain the PTP network abnormal event record.
6. The PTP network performance monitoring method based on big data analysis according to claim 1, characterized in that: The performing of machine learning analysis based on the PTP network abnormal event record and the network performance analysis result to obtain a network performance score and a fault location result includes: Extracting clock stability data, synchronization accuracy data, network transmission data and system reliability data from the network performance analysis results, performing normalization processing, and obtaining standardized feature data; Calculate the stability score for the clock deviation standard deviation and the maximum deviation value in the standardized characteristic data, and calculate the accuracy score for the average deviation value and the deviation distribution to obtain the clock performance sub-item score; The transmission score is calculated for the PTP message packet loss rate and delay change rate in the standardized feature data, and the reliability score is calculated for the device state transition frequency and abnormal event frequency to obtain the network performance sub-item score; The clock performance sub-item score and the network performance sub-item score are comprehensively calculated by a weighted calculation method to obtain a network performance score of 0-100; Extract the abnormal occurrence time, abnormal type, involved equipment and abnormal parameter value from the PTP network abnormal event record, generate an abnormal feature vector, and obtain abnormal feature data; A root cause analysis is performed based on the abnormal feature data to identify the device or link where the abnormality first occurred and query the historical fault processing database to obtain a fault location result.
7. The PTP network performance monitoring method based on big data analysis according to claim 1, characterized in that: The performing of network performance optimization and predictive maintenance according to the network performance score and the fault location result, and generating optimization suggestions and maintenance work orders, includes: Based on historical data and current network status, a time series prediction algorithm is applied to predict future trends in device clock stability, network delay, and jitter, and obtain network performance prediction results. Comparing the network performance prediction result with a preset performance threshold, and obtaining a predictive maintenance requirements list when the predicted value is lower than the performance threshold; According to the fault location result and the low-scoring items in the network performance score, extract the corresponding optimization solution from the optimization knowledge base to generate a PTP configuration parameter optimization suggestion; Conduct simulation evaluation on the PTP configuration parameter optimization suggestion, calculate the performance improvement effect and potential risks after parameter adjustment, and obtain an optimization solution evaluation report; Based on the predictive maintenance requirements list and the fault location results, analyze the long-term change trend of the device clock stability, identify signs of hardware aging, and obtain device health assessment results; A standardized maintenance work order including maintenance objects, maintenance contents and priorities is generated according to the optimization scheme evaluation report and the equipment health evaluation results.
8. A PTP network performance monitoring system based on big data analysis, characterized in that: Used to implement the PTP network performance monitoring method based on big data analysis as described in any one of claims 1 to 7, the PTP network performance monitoring system based on big data analysis includes: The collection module is used to collect PTP message data, network device PTP status data and RTP stream data packets in the PTP network through multiple data collection points to obtain a structured data stream; A calculation module, used for performing data preprocessing and clock error calculation on the structured data stream to obtain a standardized data set including PTP network status parameters; Establish a module for constructing a PTP network performance time series database based on the standardized data set, and establishing a multidimensional PTP network performance analysis model to obtain network performance analysis results; A detection module, used to perform real-time status monitoring and anomaly detection according to the network performance analysis results, and obtain a record of abnormal events of the PTP network; An execution module, configured to perform machine learning analysis based on the PTP network abnormal event record and the network performance analysis result to obtain a network performance score and a fault location result; A maintenance module is used to perform network performance optimization and predictive maintenance according to the network performance score and the fault location result, and generate optimization suggestions and maintenance work orders.
9. A PTP network performance monitoring device based on big data analysis, characterized in that: It includes a memory and a processor, the memory stores a computer program that can be run on the processor, and the processor implements the PTP network performance monitoring method based on big data analysis described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor executes the PTP network performance monitoring method based on big data analysis as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Network traffic modeling and predicting method and device based on mixed integer programming
CN115941511A
Method for optimizing performance of indoor industrial network, TSN bridge and industrial network
CN119094347A
Dynamic multi-link intelligent management and scheduling system
CN119420691A
Detecting oscillation anomalies in a mesh network using machine learning
US20170078170A1
Determining the impact of network events on network applications
US20210176114A1
Cited By
Communication dynamic analysis method for virtual network security
CN120415917A
A communication dynamic analysis method for virtual network security
CN120415917B
Multi-dimensional time sequence equipment abnormal state prediction method and system
CN120639589A
A method and system for multi-dimensional time-series equipment anomaly state prediction
CN120639589B
Link aggregation control protocol function test method and device, equipment and storage medium
CN120658653A