PTP Network Performance Monitoring Method, System and Medium Based on Big Data Analysis

The big data-driven PTP network monitoring system addresses the limitations of traditional methods by offering real-time state monitoring and predictive maintenance, enhancing network visibility and fault detection in complex broadcast environments.

CN120110951BActive Publication Date: 2025-07-15BEIJING COOLSHARK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510585856.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-15
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Traditional PTP network monitoring systems lack a global perspective, struggle with limited predictive capabilities, are inefficient in fault localization, and are costly and inflexible, making them unsuitable for complex broadcast environments.

Method used

A method and system utilizing big data analysis for PTP network performance monitoring, including data collection, preprocessing, and machine learning to provide real-time state monitoring, fault localization, and predictive maintenance, enhancing network visibility and fault detection.

Benefits of technology

The solution provides comprehensive, accurate, and proactive network monitoring, reducing fault localization time and costs while improving network performance and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120110951B_ABST
    Figure CN120110951B_ABST
Patent Text Reader

Abstract

This application relates to the field of data processing technologies, and discloses a PTP network performance monitoring method, system and medium based on big data analysis. The method includes: collecting PTP packet data, device status data and RTP stream data; preprocessing the data and calculating clock error; constructing a time series database and an analysis model; performing real-time monitoring and anomaly detection; applying machine learning for performance scoring and fault location; and accordingly performing network optimization and predictive maintenance. This application realizes global status monitoring, intelligent anomaly detection, accurate fault location and predictive maintenance of the PTP network by establishing a PTP network performance monitoring method based on big data analysis, so as to ensure the high-precision time synchronization requirements of the radio and television IP production and broadcast system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a PTP network performance monitoring method, system, and medium based on big data analysis. Background Art

[0002] With the wide application of IP technology in the radio and television industry, the network synchronization technology based on PTP (Precision Time Protocol) has become a key infrastructure for modern radio and television production and broadcast systems. The PTP protocol is defined by IEEE1588-2008 and can achieve high-precision time synchronization with a precision of less than 1 μs in distributed network communications. In the broadcast field, the SMPTE ST 2059-2 standard is based on the IEEE 1588 standard and provides a solution accurate to the nanosecond level for time and frequency synchronization in professional broadcast environments. Traditional PTP network monitoring systems mainly rely on basic network management protocols such as SNMP and NetConf to obtain the PTP status information of devices, or analyze PTP packets in a specific network segment through packet capture. These systems usually adopt a threshold alarm mechanism, which triggers an alarm when it detects that the PTP parameters exceed the preset range, and provides basic status display and historical data query functions.

[0003] However, with the continuous growth of the scale and complexity of IP broadcast systems, traditional PTP network monitoring methods face many challenges. First, existing monitoring systems generally lack a global perspective and are difficult to construct a complete PTP clock distribution topology, making it impossible to intuitively present the overall PTP status of the system. Second, traditional monitoring methods use simple threshold judgments and do not fully utilize historical data for in-depth analysis, resulting in limited predictive ability for potential problems. Third, in a complex network environment, a single anomaly may trigger a chain reaction, and traditional monitoring systems are difficult to effectively locate the root cause of the failure, increasing the difficulty of troubleshooting. In addition, existing systems lack an intelligent optimization suggestion function and cannot automatically generate optimization solutions based on historical data and the current state, restricting the continuous improvement of PTP network performance. Finally, traditional monitoring systems usually rely on dedicated hardware for implementation, with high deployment costs and poor scalability, making it difficult to adapt to the rapidly changing broadcast technology environment. Summary of the Invention

[0004] This application provides a PTP network performance monitoring method, system, and medium based on big data analysis, which is used to realize the global status monitoring, intelligent anomaly detection, accurate fault location, and predictive maintenance of the PTP network by establishing a PTP network performance monitoring method based on big data analysis, so as to ensure the high-precision time synchronization requirements of the radio and television IP production and broadcast system.

[0005] In a first aspect, the present application provides a PTP network performance monitoring method based on big data analysis. The PTP network performance monitoring method based on big data analysis includes: collecting PTP message data, network device PTP status data, and RTP flow data packets in a PTP network through multiple data collection points to obtain a structured data stream; performing data preprocessing and clock error calculation on the structured data stream to obtain a standardized data set containing PTP network status parameters; constructing a PTP network performance time series database based on the standardized data set, and establishing a multi-dimensional PTP network performance analysis model to obtain a network performance analysis result; performing real-time status monitoring and anomaly detection based on the network performance analysis result to obtain PTP network anomaly event records; performing machine learning analysis based on the PTP network anomaly event records and the network performance analysis result to obtain a network performance score and a fault location result; performing network performance optimization and predictive maintenance based on the network performance score and the fault location result to generate optimization suggestions and maintenance work orders.

[0006] In a second aspect, the present application provides a PTP network performance monitoring system based on big data analysis. The PTP network performance monitoring system based on big data analysis includes:

[0007] A collection module, configured to collect PTP message data, network device PTP status data, and RTP flow data packets in a PTP network through multiple data collection points to obtain a structured data stream;

[0008] A calculation module, configured to perform data preprocessing and clock error calculation on the structured data stream to obtain a standardized data set containing PTP network status parameters;

[0009] An establishment module, configured to construct a PTP network performance time series database based on the standardized data set, and establish a multi-dimensional PTP network performance analysis model to obtain a network performance analysis result;

[0010] A detection module, configured to perform real-time status monitoring and anomaly detection based on the network performance analysis result to obtain PTP network anomaly event records;

[0011] An execution module, configured to perform machine learning analysis based on the PTP network anomaly event records and the network performance analysis result to obtain a network performance score and a fault location result;

[0012] A maintenance module, configured to perform network performance optimization and predictive maintenance based on the network performance score and the fault location result to generate optimization suggestions and maintenance work orders.

[0013] In a third aspect, there is provided a PTP network performance monitoring device based on big data analysis, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor invokes the instructions in the memory to cause the PTP network performance monitoring device based on big data analysis to execute the above-mentioned PTP network performance monitoring method based on big data analysis.

[0014] In a fourth aspect, there is provided a computer-readable storage medium, in which instructions are stored, and when it runs on a computer, it causes the computer to execute the above-mentioned PTP network performance monitoring method based on big data analysis.

[0015] In the technical solution provided by this application, PTP message data, network device PTP status data, and RTP stream data packets in the PTP network are collected through multiple data collection points to obtain a structured data stream, achieving comprehensive and accurate data collection and providing a high-quality data basis for subsequent analysis. Data preprocessing and clock error calculation are performed on the collected structured data stream to obtain a standardized data set containing PTP network status parameters, effectively eliminating the influence of network jitter and improving the accuracy of clock deviation calculation. A PTP network performance time series database is constructed based on the standardized data set, and a multi-dimensional PTP network performance analysis model is established to obtain network performance analysis results. Through a distributed storage architecture and a time-span hierarchical storage strategy, efficient storage and query of massive time series data are realized. Moreover, the long short-term memory network model gives full play to the advantages of artificial intelligence algorithms in time series data prediction, accurately capturing long-term dependence relationships and short-term fluctuation characteristics in the PTP network, and significantly improving the accuracy and timeliness of performance prediction. Real-time status monitoring and anomaly detection are performed according to the network performance analysis results to obtain PTP network anomaly event records, realizing intuitive visualization of the network status and accurate anomaly detection. The correlation analysis algorithm effectively filters out short-term anomalies and reduces the false alarm rate. Machine learning analysis is performed based on the PTP network anomaly event records and network performance analysis results to obtain network performance scores and fault location results. Among them, the machine learning algorithm greatly improves the accuracy and efficiency of fault root cause location through learning historical anomaly patterns, and converts complex multi-dimensional data into intuitive performance scores for convenient management decision-making. Network performance optimization and predictive maintenance are performed according to the network performance scores and fault location results to generate optimization suggestions and maintenance work orders. With the help of time series prediction algorithms, forward-looking prediction of network performance is carried out, transforming maintenance from passive response to active prevention. The simulation evaluation mechanism of the optimization suggestions ensures the effectiveness and safety of the optimization measures. The overall solution deeply integrates big data collection, artificial intelligence analysis, and PTP network monitoring. Especially in time series data analysis, the application of the long short-term memory network gives full play to the advantages of deep learning algorithms in processing time series data, and can accurately capture complex patterns and long-term dependence relationships in PTP clock data; in the anomaly detection link, unsupervised learning algorithms can automatically identify anomaly patterns from massive data without pre-defining all anomaly types; in the fault root cause analysis, causal inference algorithms can accurately trace the anomaly propagation path and locate the root cause. These characteristics of artificial intelligence algorithms and models directly improve the accuracy, forward-looking, and automation of PTP network monitoring, realizing the transformation from passive monitoring to active prevention. Brief Description of the Drawings

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0017] Figure 1 FIG. is a schematic diagram of an embodiment of a PTP network performance monitoring method based on big data analysis in an embodiment of the present application;

[0018] Figure 2 FIG. is a schematic diagram of an embodiment of a PTP network performance monitoring system based on big data analysis in an embodiment of the present application;

[0019] Figure 3 FIG. is a structural schematic block diagram of a PTP network performance monitoring device based on big data analysis in an embodiment of the present invention. Detailed implementation manners

[0020] The embodiments of the present application provide a PTP network performance monitoring method, system and medium based on big data analysis. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and accompanying drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0021] For ease of understanding, the following describes the specific process of the embodiments of the present application. Please refer to Figure 1 , an embodiment of the PTP network performance monitoring method based on big data analysis in the embodiments of the present application includes:

[0022] Step S101: Collect PTP message data, network device PTP status data, and RTP stream data packets in the PTP network through multiple data collection points to obtain a structured data stream;

[0023] Step S102: Perform data preprocessing and clock error calculation on the structured data stream to obtain a standardized data set containing PTP network status parameters;

[0024] Step S103: Build a timing database for the PTP network performance based on the standardized dataset, and establish a multi-dimensional PTP network performance analysis model to obtain the network performance analysis results;

[0025] Step S104: Perform real-time status monitoring and anomaly detection based on the network performance analysis results to obtain the PTP network anomaly event records;

[0026] Step S105: Perform machine learning analysis based on the PTP network anomaly event records and the network performance analysis results to obtain the network performance score and the fault location result;

[0027] Step S106: Perform network performance optimization and predictive maintenance based on the network performance score and the fault location result to generate optimization suggestions and maintenance work orders.

[0028] It can be understood that the execution subject of this application can be a PTP network performance monitoring system based on big data analysis, or a terminal or a server. Specifically, it is not limited here. In the embodiments of this application, the server is used as the execution subject for illustration.

[0029] Specifically, the PTP network performance monitoring method based on big data analysis first collects PTP message data, network device PTP status data, and RTP flow data packets in the PTP network through multiple data collection points to obtain a structured data stream. In this step, the collection server is deployed at a strategic location in the network, accesses the IP data stream through a fiber optic network card, and directly obtains the hardware timestamp of the network data packet using the MII interface. The PTP message collection parses four key messages: Sync message, Follow_Up message, Delay_Req message, and Delay_Resp message, and extracts data fields such as timestamp information, message sequence number, source address, and destination address. The network device PTP status data is obtained through the NetConf and SNMP protocols, and the data content includes parameters such as PTP lock status, Grandmaster ID, ClockID, PTP working mode, and step mode. For devices that do not support the standard management protocol, the SSH protocol is used to obtain the raw data and then converted into the standard JSON format through regular expressions. The RTP flow data packet collection focuses on extracting the source address, source port, destination address, destination port, sequence number, marking information, and RTP timestamp. All the collected data is appended with the collection timestamp and the collection point identification information to form a complete structured data stream.

[0030] Perform data preprocessing and clock error calculation on the structured data stream to obtain a standardized data set containing PTP network status parameters. In the data preprocessing stage, first extract four types of timestamps from the structured data stream: the transmission timestamp t1 and reception timestamp t2 of the synchronization message, and the transmission timestamp t3 and reception timestamp t4 of the delay request message. Based on these four timestamps, calculate the network path delay value delay = [(t2 - t1) + (t4 - t3)] / 2, and then calculate the clock offset offset = (t2 - t1 - delay). To eliminate the error impact caused by network jitter, use the Kalman filter algorithm to de-jitter the original clock offset value. The Kalman filter algorithm smooths the clock offset value through two stages: state prediction and measurement update, effectively removing short-term jitter interference. Data preprocessing also includes outlier identification and removal. Identify and remove deviation values outside the normal range through statistical analysis, unify data from different sources and in different formats into a standard format, and finally form a standardized data set containing key parameters such as the PTP locking status of network devices, clock source information, clock deviation value, jitter value, and medium delay value.

[0031] Build a PTP network performance time series database based on the standardized data set, and establish a multi-dimensional PTP network performance analysis model to obtain network performance analysis results. The time series database adopts a distributed storage architecture, which is optimized for the characteristics of time series data, supports nanosecond-level time accuracy and efficient time interval queries. Store the data in layers according to the time span. The high-precision original data for the most recent 7 days is stored in the high-speed storage layer, and the historical data is stored in the low-cost storage layer after downsampling according to time periods. On this basis, extract the feature vectors of clock deviation, jitter value, and network delay data from the PTP data storage structure, and build a multi-dimensional PTP network performance analysis model based on the long short-term memory network (LSTM). The model includes an input layer, two long short-term memory hidden layers, and a fully connected output layer. Use the mean squared error loss function and the Adam optimization algorithm to iteratively optimize the model parameters. The trained model can predict clock stability and network performance indicators based on the input real-time standardized data set, and combine the PTP clock distribution topology information to generate complete network performance analysis results.

[0032] Execute real-time status monitoring and anomaly detection based on the network performance analysis results to obtain PTP network anomaly event records. First, construct a complete PTP distribution topology diagram based on the performance analysis results, and display the Grandmaster device, boundary clock devices, and ordinary clock devices and their master-slave relationships in real time. In the topology diagram, use color coding to mark the PTP locking status, clock deviation values, and jitter values of each device, intuitively reflecting the device status. Set multi-level threshold judgment criteria to monitor the change value of the clock deviation in real time. When the change value exceeds the preset threshold (such as 500 nanoseconds), generate a clock jump anomaly mark. Monitor the change in the PTP locking status of the device or the situation of not being locked for a long time, and generate a locking status anomaly alarm. Analyze the clock deviation trend, identify the situation where the one-way continuous growth trend exceeds the preset drift rate, and generate a clock drift anomaly mark. Apply correlation analysis to filter out short-term anomalies caused by network jitter, reduce the false alarm rate, and finally form accurate PTP network anomaly event records.

[0033] Execute machine learning analysis based on the PTP network anomaly event records and network performance analysis results to obtain network performance scores and fault location results. First, extract clock stability data (including standard deviation of clock deviation, maximum deviation value), synchronization accuracy data (average deviation value, deviation distribution), network transmission data (PTP packet loss rate, delay change rate), and system reliability data (device status conversion frequency, anomaly event frequency), and perform normalization processing to obtain standardized feature data. Apply a weighted calculation method to the standardized feature data to generate a network performance score ranging from 0 to 100. Extract the anomaly occurrence time, anomaly type, involved devices, and anomaly parameter values from the anomaly event records to construct an anomaly feature vector, identify the device or link where the anomaly first appears through root cause analysis, and combine with the historical fault handling database to accurately locate the root cause of the fault.

[0034] Execute network performance optimization and predictive maintenance based on the network performance scores and fault location results to generate optimization suggestions and maintenance work orders. Apply time series prediction algorithms to analyze historical data and the current network status, and predict the future change trends of device clock stability, network delay, and jitter. Compare the prediction results with the preset performance thresholds to identify potential performance degradation risks and generate a list of predictive maintenance requirements. According to the fault location results and low-scoring items of the network performance scores, extract corresponding optimization solutions from the optimization knowledge base to generate targeted PTP configuration parameter optimization suggestions, including adjusting parameters such as announce interval, sync interval, and delay_req interval. Conduct a simulation evaluation of the optimization suggestions, calculate the performance improvement effect and potential risks after parameter adjustment. Analyze the long-term change trend of device clock stability, identify signs of hardware aging, and generate device health assessment results. Finally, generate standardized maintenance work orders based on the optimization evaluation report and device health assessment results, including maintenance objects, maintenance contents, and priority information.

[0035] For example, in a certain production and broadcasting network system, after deploying this method, the sending timestamp t1 of the PTP synchronization message collected by the data acquisition point from the core switch is 1613455782.123456789 seconds, the receiving timestamp t2 is 1613455782.123458789 seconds, the sending timestamp t3 of the delay request message is 1613455782.223456789 seconds, and the receiving timestamp t4 of the delay response message is 1613455782.223458789 seconds. By calculation, the path delay value delay = [(t2 - t1) + (t4 - t3)] / 2 = [(123458789 - 123456789) + (223458789 - 223456789)] / 2 = 1000 nanoseconds, and then the clock offset offset = (t2 - t1 - delay) = 1000 nanoseconds. After Kalman filtering, a stable clock deviation value of 985 nanoseconds is obtained. The system stores these data in the time series database. The multi-dimensional PTP network performance analysis model discovers through analyzing historical data that the clock deviation of this device shows a slow growth trend, and predicts that the deviation will exceed 1500 nanoseconds after 24 hours. The real-time monitoring system generates a clock drift anomaly flag, and the machine learning analysis module identifies the root cause of the problem as the aging of the crystal oscillator of this device based on the anomaly feature vector. The system automatically generates a maintenance work order including a suggestion to replace the crystal oscillator, thus preventing potential synchronization problems in advance.

[0036] In the embodiments of the present application, PTP message data, network device PTP status data, and RTP stream data packets in the PTP network are collected through multiple data collection points to obtain a structured data stream, achieving comprehensive and accurate data collection and providing a high-quality data foundation for subsequent analysis. Data preprocessing and clock error calculation are performed on the collected structured data stream to obtain a standardized data set containing PTP network status parameters, effectively eliminating the influence of network jitter and improving the accuracy of clock deviation calculation. A PTP network performance time series database is constructed based on the standardized data set, and a multi-dimensional PTP network performance analysis model is established to obtain network performance analysis results. Through a distributed storage architecture and a time-span hierarchical storage strategy, efficient storage and query of massive time series data are realized. Moreover, the long short-term memory network model gives full play to the advantages of artificial intelligence algorithms in time series data prediction, accurately capturing long-term dependencies and short-term fluctuation characteristics in the PTP network, and significantly improving the accuracy and timeliness of performance prediction. Real-time status monitoring and anomaly detection are performed according to the network performance analysis results to obtain PTP network anomaly event records, realizing intuitive visualization of the network status and accurate anomaly detection. The correlation analysis algorithm effectively filters out short-term anomalies and reduces the false alarm rate. Machine learning analysis is performed based on the PTP network anomaly event records and network performance analysis results to obtain network performance scores and fault location results. Among them, the machine learning algorithm greatly improves the accuracy and efficiency of fault root cause location through learning historical anomaly patterns, and converts complex multi-dimensional data into intuitive performance scores for convenient management decision-making. Network performance optimization and predictive maintenance are performed according to the network performance scores and fault location results to generate optimization suggestions and maintenance work orders. With the help of time series prediction algorithms, forward-looking prediction of network performance is carried out, transforming maintenance from passive response to active prevention. The simulation evaluation mechanism of the optimization suggestions ensures the effectiveness and safety of the optimization measures. The overall solution deeply integrates big data collection, artificial intelligence analysis, and PTP network monitoring. Especially in time series data analysis, the application of the long short-term memory network gives full play to the advantages of deep learning algorithms in processing time series data and can accurately capture complex patterns and long-term dependencies in PTP clock data; in the anomaly detection link, unsupervised learning algorithms can automatically identify anomaly patterns from massive data without pre-defining all anomaly types; in the fault root cause analysis, causal inference algorithms can accurately trace the anomaly propagation path and locate the root cause. These characteristics of artificial intelligence algorithms and models directly improve the accuracy, forward-looking, and automation of PTP network monitoring, realizing the transformation from passive monitoring to active prevention.

[0037] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0038] Connect to the IP data stream through the deployed acquisition server connected to the fiber optic network card, and obtain the hardware timestamp of the network packet through the MII interface to obtain timestamp information with nanosecond-level accuracy;

[0039] Parse the PTP synchronization message, follow-up message, delay request message, and delay response message, and extract the timestamp information, message sequence number, and address information therein to obtain PTP message data;

[0040] Collect data from network devices through the NetConf and SNMP protocols, and obtain the PTP lock status, Grandmaster ID, and Clock ID to obtain PTP status data of network devices;

[0041] Obtain data from devices that do not support standard management protocols through the SSH protocol, and convert the data into JSON format through regular expressions to obtain standardized device status data;

[0042] Collect RTP stream data packets received by the terminal device, and extract the source address, destination address, sequence number, and RTP timestamp to obtain media stream clock data;

[0043] Add the acquisition timestamp and acquisition point identification information to the PTP message data, network device PTP status data, standardized device status data, and media stream clock data to obtain structured data streams.

[0044] Specifically, connect to the IP data stream through the deployed acquisition server connected to the fiber optic network card, and obtain the hardware timestamp of the network packet through the MII interface to obtain timestamp information with nanosecond-level accuracy. The MII interface is the abbreviation of Media Independent Interface, which is an interface defined by the IEEE-802.3 standard to describe the interface between an Ethernet transceiver and a network controller. The acquisition server is configured with a fiber optic network card, and the IP data is accessed through the fiber optic network card. The data packet is first sent to the server IP decapsulation module and then further processed. In order to obtain timestamp information with nanosecond-level accuracy, the acquisition server modifies the network driver, uses the MII interface timestamp to send and receive messages through an independent channel, and ensures the high accuracy of timestamp capture. This process actually directly stamps the hardware-level timestamp when the network packet just enters the physical layer, avoiding the problem of inaccurate timestamps caused by operating system scheduling delays. The acquisition server also sets a performance optimization module to increase the priority of the synchronization process and synchronization thread, and uses a high-precision time timer to obtain the current time value to ensure timestamp accuracy.

[0045] Parse the PTP synchronization messages, follow-up messages, delay request messages, and delay response messages, extract the timestamp information, message sequence number, and address information therein to obtain the PTP message data. The PTP protocol is the Precision Time Protocol defined by IEEE 1588-2008, and the SMPTE ST 2059-2 standard is the audio and video system time synchronization standard based on the IEEE 1588 standard. After receiving the PTP message, the acquisition server first identifies the message type, which is divided into four basic messages: The synchronization message (Sync) is used by the master clock to send time information to the slave clock, and the synchronization message contains the transmission timestamp of the master clock; the follow-up message (Follow_Up) is used to transfer a more accurate timestamp in the two-step mode; the delay request message (Delay_Req) is sent by the slave clock to measure the network transmission delay; the delay response message (Delay_Resp) is the response of the master clock to the delay request and contains the timestamp of receiving the delay request. When parsing these messages, the acquisition server extracts the key fields in each message: The timestamp information records the exact time when the message is sent or received, the sequence number is used to identify and match the corresponding request and response messages, and the source address and destination address identify the network location of the PTP device. These extracted fields form the structured PTP message data.

[0046] Collect data from network devices through the NetConf and SNMP protocols, obtain the PTP lock status, Grandmaster ID, and Clock ID to get the PTP status data of the network devices. NetConf is a network configuration protocol that uses data in XML format to interact with network devices through a secure transport protocol; SNMP is the Simple Network Management Protocol used for the monitoring and management of network devices. The acquisition system first determines the devices in the network that support NetConf or SNMP, establishes connections with these devices, and then sends standardized query requests to obtain PTP-related information. The PTP lock status indicates whether the device is successfully synchronized with the master clock. The Grandmaster ID is the identifier of the most authoritative clock source in the entire PTP network, and the Clock ID is the PTP clock identifier of the device itself. The acquisition system also obtains information such as the PTP mode (master clock, slave clock, or boundary clock), step mode (one-step or two-step), the current PTP system time, the latest synchronization time, and the PTP port status, etc., to form the complete PTP status data of the network devices.

[0047] Data is obtained from devices that do not support the standard management protocol via the SSH protocol, and the data is converted into JSON format through regular expressions to obtain standardized device status data. Some dedicated devices or old devices may not support the NetConf or SNMP protocol. The acquisition system then uses the SSH protocol (Secure Shell Protocol) to directly log in to the command-line interface of the device and execute commands to query the PTP status, such as "show ptp", "display ptp status", etc. The data obtained is usually unstructured text output, and the acquisition system uses pre-defined regular expression patterns to match relevant information. Regular expressions are expressions used to match specific patterns in strings. For example, the pattern "GM ID: ([0-9a-fA-F:]+)" is used to match the Grandmaster ID. The information extracted by matching is then converted into the standard JSON format, such as {"device_id": "switch01", "gm_id": "01:02:03:04:05:06:07:08", "clock_state": "locked"}, achieving the unification and standardization of the data format for subsequent processing.

[0048] The RTP stream data packets received by the acquisition terminal device are used to extract the source address, destination address, sequence number, and RTP timestamp to obtain media stream clock data. RTP (Real-Time Transport Protocol) is a protocol used to transmit audio and video data over IP networks and is widely used in IP production and broadcast systems. The acquisition system captures the RTP data packets passing through the network, parses their header information, and extracts the key fields: The source address refers to the source IP address of the streaming device, the source port is the source port of the streaming device, the destination address is usually a multicast address, the destination port is the destination port for sending traffic, the sequence number (Sequence Number) is used to identify the order of the data packets, the marker information (M and F) represents the mark flag and the field order flag respectively, and most importantly, the RTP timestamp, which records the time point of media sampling. The acquisition system forms media stream clock data after extracting these fields, and this data will be used for comparative analysis with the PTP clock to evaluate the synchronization accuracy between the media stream and the PTP clock.

[0049] Add the collection timestamp and collection point identification information to the PTP message data, network device PTP status data, standardized device status data, and media stream clock data to obtain a structured data stream. In this step, the collection system integrates and tags all the data obtained previously. First, add the collection timestamp to each piece of data to record the exact time of data collection, using a nanosecond-level time format such as "1613455782.123456789". Second, add the collection point identification information, such as "collector_node_001", to clarify the data source. Then, the system organizes different types of data according to a unified data structure template to form a standardized structured data stream, which includes fields such as data type identification, collection time, collection point, and original data. The structured data stream uses JSON or a similar format, which is convenient for network transmission and database storage. Finally, these structured data are transmitted to the central data processing platform in real time to provide a basis for subsequent data preprocessing and analysis.

[0050] For example, a certain IP-based broadcast production center deploys this method for PTP network monitoring. The collection server is connected to the mirror port of the core switch and receives the PTP synchronization message through the fiber optic network card. The MII interface immediately adds the hardware timestamp "1613455782.123456789" to this message. Parse the PTP synchronization message to extract the master clock sending time "1613455782.123455789", the message sequence number "45678", the source address "10.1.1.1" (master clock), and the destination address "224.0.1.129" (PTP multicast address). At the same time, the collection system queries the core switch through the SNMP protocol to obtain its PTP lock status as "Locked", the Grandmaster ID as "00:01:02:03:04:05:06:07", and the Clock ID as "00:A1:A2:A3:A4:A5:A6:A7". For an encoder that does not support SNMP, the system executes commands through the SSH protocol to obtain the text output "PTP Status: Synchronized to GM 00:01:02:03:04:05:06:07", and uses regular expressions to extract information and convert it into JSON format. The system also captures the RTP stream data packets sent from this encoder, extracts the source address "10.2.2.2", the destination address "239.1.1.1", the sequence number "12345", and the RTP timestamp "9876543210". Finally, add the collection timestamp "1613455782.123456789" and the collection point identification "collector_001" to all the data to form a complete structured data stream, which is transmitted to the central processing platform in real time for subsequent analysis.

[0051] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0052] Extract the transmission timestamp and reception timestamp of the PTP synchronization message, as well as the transmission timestamp of the delay request message and the reception timestamp of the delay response message from the structured data stream to obtain a timestamp data set;

[0053] Calculate the path delay value according to the timestamp data set, add the synchronization message time difference and the delay message time difference and then average to obtain the network path delay parameter;

[0054] Calculate the clock offset based on the network path delay parameter, subtract the network path delay parameter from the synchronization message time difference to obtain the original clock offset value;

[0055] Apply the Kalman filtering algorithm to the original clock offset value for jitter reduction processing to obtain a stable clock deviation value;

[0056] Identify and remove outliers in the structured data stream, unify data from different sources into a standard format to obtain a cleaned data set;

[0057] Associate and integrate the cleaned data set with the stable clock deviation value to form a standardized data set including the PTP locking state of the network device, clock source information, clock deviation value, jitter value, and medium delay value.

[0058] Specifically, the transmission timestamp and reception timestamp of the PTP synchronization message, as well as the transmission timestamp of the delay request message and the reception timestamp of the delay response message, are extracted from the structured data stream to obtain a timestamp data set. In this process, the handler filters out the PTP synchronization message (Sync), follow-up message (Follow_Up), delay request message (Delay_Req), and delay response message (Delay_Resp) according to the message type identifier by traversing each record in the structured data stream. For each type of message, its key timestamp information is extracted: the master clock transmission time t1 (transmission timestamp) and the slave clock reception time t2 (reception timestamp) are extracted from the synchronization message. In the two-step mode, a more accurate value of the transmission timestamp t1 is extracted from the follow-up message; the slave clock transmission time t3 (transmission timestamp) is extracted from the delay request message, and the master clock reception time t4 (reception timestamp) is extracted from the delay response message. These four timestamps together form a complete PTP time synchronization cycle, forming a timestamp data set, and each set of data sets contains four timestamp values associated with the same sequence number. The path delay value is calculated based on the timestamp data set, and the time difference of the synchronization message and the time difference of the delay message are added and averaged to obtain the network path delay parameter. The specific calculation process is as follows: First, calculate the time difference of the synchronization message, that is, the time t2 when the slave clock receives the synchronization message minus the time t1 when the master clock sends the synchronization message, to obtain the transmission time in the master-to-slave direction plus the clock offset; then calculate the time difference of the delay message, that is, the time t4 when the master clock receives the delay request message minus the time t3 when the slave clock sends the delay request message, to obtain the transmission time in the slave-to-master direction minus the clock offset; finally, add these two time differences and divide by 2, that is, (t2 - t1) + (t4 - t3) / 2, to obtain the average network path delay value. This calculation method assumes that the network transmission delay is symmetric in both directions, and the influence of the clock offset is eliminated by averaging, and the obtained is the pure network transmission delay.

[0059] The clock offset is calculated based on the network path delay parameter. Subtract the network path delay parameter from the time difference of the synchronization message to obtain the original clock offset value. The clock offset calculation formula is: offset = (t2 - t1 - delay), where t2 is the time when the slave clock receives the synchronization message, t1 is the time when the master clock sends the synchronization message, and delay is the network path delay parameter calculated in the previous step. This calculation process is actually to strip the pure network transmission delay from the total time difference between the master and slave clocks, and the remaining is the actual deviation between the two clocks. The original clock offset value represents the deviation of the slave clock relative to the master clock. A positive value indicates that the slave clock is faster than the master clock, and a negative value indicates that the slave clock is slower than the master clock.

[0060] Apply the Kalman filter algorithm to the original clock offset value for jitter removal, obtaining a stable clock deviation value. The Kalman filter is a recursive optimal estimation algorithm, especially suitable for processing time series data with noise. In a PTP network, due to factors such as network jitter and load fluctuations, the original clock offset value often contains noise and instantaneous fluctuations. The Kalman filter algorithm works through two stages: the prediction stage and the update stage. In the prediction stage, the state at the next moment is predicted based on the current state and error covariance; in the update stage, the prediction result is adjusted by combining the actual measurement value to obtain the optimal estimate. When specifically implemented, the clock state model is set to include two state variables: the offset value and the drift rate, and the measurement model is the direct observation of the offset value. The algorithm calculates the Kalman gain recursively, dynamically adjusting the trust degree of the prediction value for the measurement value, thereby smoothing the jitter in the original offset value and obtaining a more stable clock deviation estimate value.

[0061] Identify and remove outliers in the structured data stream, unify data from different sources into a standard format, and obtain a cleaned data set. Multiple strategies are adopted for outlier identification: the statistical threshold method identifies deviation values beyond the normal range, such as setting a deviation exceeding ±100 microseconds as an outlier; mutation detection identifies situations where the deviation value changes violently within a short period; consistency check verifies the logical relationship between related data, such as the deviation values of different devices within the same time period should have a certain correlation. The identified outliers are marked as invalid or replaced with interpolation estimated values. The data format unification process converts data obtained from different sources and different protocols into a unified structure, including a standardized time format, a unified deviation value unit (nanosecond), a unified device identifier format, etc., ensuring data consistency and comparability.

[0062] Associate and integrate the cleaned data set with the stable clock deviation value to form a standardized data set containing the PTP locking state of network devices, clock source information, clock deviation value, jitter value, and medium delay value. In the association and integration process, the data is first matched according to the device identifier and timestamp to associate different types of data of the same device at similar time points; then the stable clock deviation value processed by the Kalman filter is merged with other state data of the corresponding device; finally, additional metrics are calculated, such as the jitter value (standard deviation of the clock deviation) and the stability metric (such as Allan variance). The standardized data set adopts a unified data structure, containing rich metadata and metrics, providing a comprehensive data basis for subsequent PTP network performance analysis.

[0063] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0064] Store the standardized data set in a time series database using a distributed storage architecture, and perform hierarchical storage according to the time span to obtain an efficient PTP data storage structure;

[0065] Extract features from the clock deviation, jitter value, and network latency data in the PTP data storage structure, construct an input feature vector, and obtain a training dataset;

[0066] Build a multi-dimensional PTP network performance analysis model based on the long short-term memory network using the training dataset. The multi-dimensional PTP network performance analysis model includes an input layer, two long short-term memory hidden layers, and a fully connected output layer to obtain an initial model structure;

[0067] Train the initial model structure using historical data, and use the mean squared error loss function and the Adam optimization algorithm to iteratively optimize the model parameters to obtain a trained multi-dimensional PTP performance analysis model;

[0068] Input the real-time standardized dataset into the trained multi-dimensional PTP performance analysis model, and through forward propagation calculation, obtain the clock stability prediction value and the network performance prediction value;

[0069] According to the clock stability prediction value and the network performance prediction value, combined with the PTP clock distribution topology information, generate a network performance analysis result including clock stability evaluation, performance trend prediction, and abnormal risk assessment.

[0070] Specifically, store the standardized dataset in a time series database using a distributed storage architecture, and perform hierarchical storage according to the time span to obtain an efficient PTP data storage structure. The time series database is a type of database specifically designed for processing time series data, and it is optimized for data with timestamps. In the PTP monitoring scenario, the distributed storage architecture distributes the storage pressure through multiple nodes and supports high-concurrency write and query operations. During the data storage process, a hierarchical storage strategy is adopted, and the data is divided into different levels according to the time freshness: the original high-precision data in the recent 7 days is stored in the high-speed storage layer, such as an SSD or a memory database, to retain the full precision for detailed analysis; the data from 7 days to 30 days is moderately downsampled, and the data is aggregated once per minute and stored in the medium-speed storage layer; the historical data over 30 days is further downsampled, aggregated once per hour, and stored in the low-cost storage layer. Each piece of data is accompanied by a time index and a multi-dimensional label index, such as device ID, PTP role, etc., for fast query and screening.

[0071] Extract features from the clock deviation, jitter value, and network latency data in the PTP data storage structure, construct an input feature vector, and obtain a training dataset. Feature extraction is the process of converting raw data into a more meaningful feature representation. First, extract time series features. For the clock deviation data of each device, calculate the statistical features within the sliding window, including the mean, standard deviation, maximum value, minimum value, peak-to-peak value, rate of change, etc.; similarly, extract statistical features for the jitter value and network latency data. Second, extract frequency domain features. Convert the time series to the frequency domain through the fast Fourier transform, and extract features such as the main frequency components and power spectral density to identify periodic patterns and abnormal oscillations. Also extract device topology features, including the hierarchical position of the device in the PTP network, master-slave relationship, stability of the upstream clock source, etc. After feature extraction, perform normalization processing to scale features of different magnitudes to the same range, avoiding the influence of dimensional differences between features on model training. Form training samples containing multi-dimensional features, and each sample is associated with the network state at a time point and the performance metrics at subsequent time points as labels.

[0072] Construct a multi-dimensional PTP network performance analysis model based on the long short-term memory network using the training dataset. The multi-dimensional PTP network performance analysis model includes an input layer, two long short-term memory hidden layers, and a fully connected output layer to obtain the initial model structure. The long short-term memory network (LSTM) is a special type of recurrent neural network that can effectively handle long-term dependencies in time series data. The model structure is designed as follows: The input layer receives the feature vector, and the dimension is the same as the number of features; the first layer of LSTM contains 128 neurons, uses the tanh activation function, and sets the dropout rate to 0.2 to prevent overfitting; the second layer of LSTM contains 64 neurons, also uses the tanh activation function and the dropout mechanism; then comes the fully connected layer, which maps the LSTM output to the prediction target, such as the clock stability metric and network performance metrics; finally, there is the output layer, and different activation functions are set according to the type of prediction task. The linear activation function is used for regression tasks, and the sigmoid or softmax activation function is used for classification tasks. The model design takes into account the temporal characteristics of the PTP network data. The double-layer LSTM structure can capture both short-term fluctuations and long-term trends, providing a basis for accurately predicting network performance.

[0073] The initial model structure is trained using historical data. The mean squared error loss function and the Adam optimization algorithm are used to iteratively optimize the model parameters, resulting in a trained multi-dimensional PTP performance analysis model. In the model training process, the historical dataset is first split into a training set (70%), a validation set (15%), and a test set (15%) in chronological order to ensure the consistency of the data distribution in each set. The mean squared error (MSE) is used as the loss function, which calculates the average of the sum of the squares of the differences between the predicted values and the true values. It is sensitive to outliers and helps reduce large prediction errors. The Adam optimization algorithm combines the advantages of the momentum method and RMSProp, adaptively adjusts the learning rate, and accelerates the convergence process. During the training process, a batch processing mechanism is adopted, with the size of each batch of data set to 64. The loss is calculated through forward propagation, and then the model parameters are updated through backpropagation. At the same time, an early stopping strategy is implemented to monitor the performance of the validation set. When the validation loss does not decrease for 10 consecutive epochs, the training is stopped to prevent overfitting. The training process also includes learning rate scheduling. The initial learning rate is set to 0.001. When the validation loss does not improve for 3 consecutive epochs, the learning rate is reduced to 0.5 of the original value.

[0074] The real-time standardized dataset is input into the trained multi-dimensional PTP performance analysis model, and through forward propagation calculation, the clock stability prediction value and the network performance prediction value are obtained. In the real-time prediction process, the newly collected standardized data is first subjected to the same feature extraction and normalization processing as the training data to ensure the consistency of the data format. The prediction calculation adopts a sliding window mechanism, and each time the data of the nearest N time points is taken as the input, and the value of N is the same as that set during model training (usually 48 points, corresponding to the data of the past 24 hours). The input data is sequentially calculated through each layer of the model: first, it is processed by two layers of LSTM to capture the temporal features, and then it is mapped to the specific prediction target through the fully connected layer. The model output includes multiple dimensions: the clock stability prediction value includes the trend change of the clock deviation within the next 24 hours, the expected maximum deviation value, and the deviation fluctuation degree; the network performance prediction value includes the network delay change trend, the expected jitter level, and the link stability evaluation. These prediction values serve as important bases for real-time decision-making, guiding network operation and maintenance and optimization.

[0075] Based on the predicted values of clock stability and network performance, combined with the PTP clock distribution topology information, generate network performance analysis results that include clock stability assessment, performance trend prediction, and anomaly risk assessment. In this step, first compare the prediction results with the historical baseline data to determine the percentile of the current performance level in the historical performance. The clock stability assessment is based on the predicted deviation trend and fluctuation degree, calculate professional assessment indicators such as Allan variance and maximum time interval error (MTIE), and compare with the PTP specification requirements to form a compliance rating. The performance trend prediction identifies the performance improvement or degradation trend by analyzing the change direction and rate of the predicted values, and estimates the time point to reach the key threshold. The anomaly risk assessment combines the PTP clock distribution topology to analyze potential fault points, and evaluates the impact of different node or link failures on the overall network performance by simulating the single-point fault propagation effect, identifies the key nodes and calculates the risk level. The finally generated network performance analysis results adopt a multi-level structure, which includes both the macro assessment of the overall network and the detailed analysis refined to each device node, providing comprehensive support for management decisions.

[0076] For example, a UHD production and broadcast center has deployed a PTP network performance monitoring method based on big data analysis. Through preliminary data collection and preprocessing, a standardized data set of 3 months has been accumulated, sampled once per second, forming approximately 7.8 million records. These data are first stored in a distributed time series database according to the time span: approximately 600,000 original data records in the most recent 7 days are stored in high-speed SSD nodes, retaining the full nanosecond-level accuracy; data from 7 to 30 days are aggregated every minute, and approximately 33,000 data records are stored in medium-speed storage nodes; older historical data are aggregated every hour, and approximately 2,000 data records are stored in low-cost storage nodes. The database creates a time index and a device label index for each record, achieving a query response speed of milliseconds. Feature extraction is performed on the data of the core switch, and 20 statistical features such as the mean, standard deviation, maximum value, minimum value, and linear regression slope of the clock deviation are calculated within a 48-hour sliding window. Similarly, features are extracted for the jitter value and network latency, and finally a feature vector of approximately 60 dimensions is obtained for each time point. Based on these feature vectors, a deep learning model with two layers of LSTM with 128 nodes and 64 nodes is constructed. The model is trained using the data of the first two months, and the performance of the model is verified using the data of the last month. The training uses the Adam optimizer with a batch size of 64 and an initial learning rate of 0.001. After approximately 500 rounds of iteration, the validation set error stabilizes within 50 nanoseconds. After the model is put into use, it can predict the future 24-hour clock stability change trend of the device in real time. When it is monitored that the clock deviation prediction curve of a certain boundary clock device shows a slow but continuous upward trend and is expected to exceed the 500-nanosecond threshold after 72 hours, the system immediately generates a warning message and determines that the device is the upstream clock source of 20 terminal nodes through topology analysis, with a risk rating of "high". Based on this, the maintenance personnel arranged equipment inspection in advance and found that the crystal oscillator frequency drift was caused by the abnormal temperature control system of the device. The faulty components were replaced in time, avoiding a potential large-scale synchronization interruption event.

[0077] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0078] Construct a complete PTP distribution topology map based on the network performance analysis results, and display the Grandmaster device, boundary clock device, and ordinary clock device in the network and their master-slave relationships in real time to obtain the real-time state map of the PTP network;

[0079] Perform color-coding marking on the PTP locking status, clock deviation value, and jitter value of each device in the real-time state map of the PTP network to obtain an intuitive device status view;

[0080] Set a multi-level threshold judgment standard according to the network performance analysis results, and perform real-time detection for the situation where the clock deviation change value exceeds the preset threshold to obtain a clock jump anomaly mark;

[0081] Based on the locking state information in the network performance analysis results, identify the changes in the PTP locking state of the device or the situation of no locking for a long time, and obtain an abnormal locking state alarm;

[0082] Judge the clock deviation trend in the network performance analysis results, identify the situation where the unidirectional continuous growth trend exceeds the preset drift rate, and obtain an abnormal clock drift mark;

[0083] Integrate and record the abnormal clock jump mark, abnormal locking state alarm, and abnormal clock drift mark in time series, and apply correlation analysis to filter out short-term anomalies to obtain the PTP network abnormal event record.

[0084] Specifically, the Grandmaster device, boundary clock device, and ordinary clock device in the network and their master-slave relationships are displayed in real time to obtain the real-time state diagram of the PTP network. The PTP distribution topology diagram is a network structure diagram that reflects the PTP clock synchronization hierarchy relationship, which intuitively shows how the clock signal is distributed layer by layer from the master clock source to each node in the network. During the construction process, first extract the basic information of all PTP devices from the network performance analysis results, including device identifiers, IP addresses, MAC addresses, device types, and PTP roles. Then extract the master-slave relationship data between devices, establish connection relationships by analyzing the Grandmaster ID and upstream clock source information of each device, and determine the flow direction of the synchronization signal. For the Grandmaster device, that is, the device serving as the master clock source, place it at the top layer of the topology diagram by matching its Clock ID with the Grandmaster ID reported by other devices; for the boundary clock (Boundary Clock) device, identify its upstream clock source and downstream slave devices, and place it in the middle layer of the topology diagram; for the ordinary clock (Ordinary Clock) device, determine its upstream clock source and place it at the end of the topology diagram. Through this data association and hierarchical analysis, a complete tree-like or network topology structure is formed to reflect the synchronization path and clock distribution status of the entire PTP network in real time.

[0085] Color-code the PTP lock status, clock deviation value, and jitter value of each device in the PTP network real-time status diagram to obtain an intuitive device status view. Color coding is a method of converting numerical values or statuses into visual colors for easy visual identification of device status. For the PTP lock status, a typical color coding scheme is as follows: the locked status is represented by green, indicating that the device has successfully synchronized with the master clock; the unlocked status is represented by red, indicating that the device has failed to establish synchronization with the master clock; the locking status is represented by yellow, indicating that the device is in the synchronization process but has not yet stabilized. For the clock deviation value, a gradient color spectrum coding is adopted: green indicates that the deviation value is within ±100 nanoseconds, meeting the high-precision requirements; yellow indicates that the deviation value is between ±100 nanoseconds and ±500 nanoseconds, within the warning range; orange indicates that the deviation value is between ±500 nanoseconds and ±1 microsecond, approaching the critical value; red indicates that the deviation value exceeds ±1 microsecond, which has exceeded the normal range. For the jitter value, a gradient color spectrum is also adopted: blue to green indicates that the jitter value is less than 50 nanoseconds, with good status; yellow to orange indicates that the jitter value is between 50 nanoseconds and 200 nanoseconds, which requires attention; red indicates that the jitter value exceeds 200 nanoseconds, which may affect the synchronization quality. Through this color coding, network administrators can identify problematic devices and links at a glance.

[0086] Set multi-level threshold judgment criteria according to the network performance analysis results, and perform real-time detection for the situation where the change value of the clock deviation exceeds the preset threshold to obtain a clock jump anomaly mark. Clock jump refers to the phenomenon that the clock deviation of a PTP device changes significantly in a short period of time, usually caused by network congestion, hardware failure, or external interference. To detect clock jumps, first set multi-level threshold judgment criteria for different types of devices: the jump threshold for core devices (such as backbone switches) is set relatively low, usually 100 nanoseconds, because such devices have higher requirements for time accuracy; the jump threshold for edge devices is relatively loose and can be set to 500 nanoseconds. During the real-time detection process, continuously monitor the clock deviation of each device and calculate the difference between two adjacent measurement values, that is, the deviation change rate. When the deviation change rate exceeds the preset threshold, trigger a clock jump anomaly mark. To reduce false alarms, the duration of the change also needs to be considered: instantaneous jumps (duration less than 100 milliseconds) are marked as low-level anomalies; continuous jumps (duration exceeding 100 milliseconds but less than 1 second) are marked as medium-level anomalies; long-term jumps (duration exceeding 1 second) are marked as high-level anomalies. Each anomaly mark contains information such as device identification, occurrence time, duration, jump amplitude, and severity.

[0087] Based on the locking status information in the network performance analysis results, identify the changes in the PTP locking status of the device or the situation of not being locked for a long time, and obtain the locking status exception alarm. The PTP locking status refers to the synchronization status between the device and the master clock, including locked, unlocked, locking, etc. The locking status exceptions are divided into two categories: one is the change in the locking status, that is, the device changes from the locked state to the unlocked state, indicating that the synchronization is interrupted; the other is not being locked for a long time, that is, the device fails to establish synchronization within the expected time. The identification process first extracts the locking status history records of each device from the network performance analysis results and establishes a status time series. For the detection of changes in the locking status, compare the status values at adjacent time points. When the status changes from locked to unlocked, generate a status change exception flag. For the detection of not being locked for a long time, calculate the cumulative time that the device is in the unlocked state. When it exceeds the preset time threshold (usually 30 seconds to 5 minutes depending on the device type), generate a long-time unlocking exception flag. The exception judgment also considers the importance of the device in the PTP network: the locking status exception of the device on the core path will trigger a high-priority alarm; the locking status exception of the edge device will trigger a low-priority alarm. Each locking status exception alarm contains information such as device identification, exception type, occurrence time, duration, and impact range.

[0088] Judge the clock deviation trend in the network performance analysis results, identify the situation where the unidirectional continuous growth trend exceeds the preset drift rate, and obtain the clock drift exception flag. Clock drift refers to the phenomenon that the clock deviation of the device shows a continuous unidirectional change, usually caused by crystal oscillator aging, temperature change, or hardware failure. The process of identifying clock drift first extracts the clock deviation time series data of each device from the network performance analysis results. For each device, use the linear regression method to analyze the deviation data of the nearest N time points (usually the past 1 hour) and calculate the drift rate (slope). The linear regression equation is y = ax + b, where y represents the deviation value, x represents time, and a is the drift rate. When the absolute value of the calculated drift rate exceeds the preset threshold (usually 100 nanoseconds per hour) and the correlation coefficient R² is greater than 0.7 (indicating a good fit), it is determined that there is a drift trend. To improve the accuracy of the judgment, the persistence of the drift also needs to be considered: analyze multiple consecutive time windows. When the same-direction drift trend is detected in multiple consecutive windows (usually 3), generate the clock drift exception flag. Each drift exception flag contains information such as device identification, occurrence time, drift rate, goodness of fit, and the estimated time to reach the critical deviation value.

[0089] Integrate and record clock jump anomaly markers, locked state anomaly alarms, and clock drift anomaly markers in a time series, and apply correlation analysis to filter out transient anomalies to obtain PTP network anomaly event records. The integration record merges different types of anomaly information into a unified time series database for global analysis and correlation mining. The integration process first establishes a unified anomaly event data structure, including fields such as event ID, device ID, anomaly type, start time, end time, severity, and detailed parameters. Then, all anomaly markers are inserted into this unified structure in chronological order to form a complete anomaly event timeline. To filter out transient anomalies and reduce false alarms, correlation analysis techniques are applied: First, perform time correlation analysis to identify anomalies that appear multiple times and disappear quickly within a short time window (usually 1 second), and regard them as transient fluctuations rather than real anomalies; Second, perform topological correlation analysis. When an upstream device of a certain device has an anomaly, subsequent similar anomalies of this device are regarded as cascading effects rather than independent events; Finally, perform pattern correlation analysis to match the known anomaly patterns in history with the current anomaly to improve the judgment accuracy. After filtering by correlation analysis, the remaining anomaly events form the final PTP network anomaly event records, and each record contains a complete anomaly description and context information.

[0090] For example, a certain ultra-high-definition video live broadcast system runs a PTP clock synchronization network and deploys a PTP network performance monitoring method based on big data analysis. Before a live broadcast event, the monitoring system extracts the PTP network structure information including 1 Grandmaster device, 8 boundary clock devices, and 42 ordinary clock devices from the network performance analysis results and generates a complete distribution topology map. This topology map clearly shows the path of the master clock signal from the Grandmaster device through the core switch and the secondary switch to each terminal device. Applying color coding to the topology map, 41 devices operating normally are shown in green, among which 1 boundary clock device is shown in yellow (indicating a clock deviation of about 300 nanoseconds), and 2 ordinary clock devices are shown in red (indicating an unlocked state). The monitoring system performs clock jump detection on the boundary clock device marked in yellow and finds that its clock deviation suddenly jumps from 50 nanoseconds to 300 nanoseconds within the past 10 minutes, and the change amplitude of 250 nanoseconds exceeds the preset threshold of 200 nanoseconds for this device. The system immediately generates a clock jump anomaly mark. At the same time, analyzing the locked state of the two red devices, it is found that they have been in an unlocked state for 45 consecutive minutes, exceeding the preset threshold of 30 minutes. The system generates a long-term unlocked alarm. Further analyzing the clock deviation trend of the yellow device, using linear regression to calculate the drift rate of the last 60 data points, the result shows that the deviation value continues to increase at a rate of about 120 nanoseconds per hour, and the goodness of fit R² reaches 0.85, exceeding the drift threshold of 100 nanoseconds per hour. The system generates a clock drift anomaly mark. These three types of anomaly information are integrated and recorded by time. Through correlation analysis, it is found that the anomaly of the yellow device is highly correlated with the anomalies of the two red devices downstream. Finally, a complete anomaly event record is generated: "The boundary clock device BC03 has a clock drift, resulting in the synchronization failure of the downstream OC12 and OC17 devices". The maintenance personnel quickly check the BC03 device based on this precise anomaly location and find that its temperature anomaly causes the crystal oscillator frequency to drift. After replacing the relevant components, the entire network synchronization state returns to normal, ensuring the smooth progress of the live broadcast event.

[0091] In a specific embodiment, the process of executing step S105 may specifically include the following steps:

[0092] Extract clock stability data, synchronization accuracy data, network transmission data, and system reliability data from the network performance analysis results, perform normalization processing, and obtain standardized feature data;

[0093] Calculate the stability score for the standard deviation of clock deviation and the maximum deviation value in the standardized feature data, and calculate the accuracy score for the average deviation value and deviation distribution to obtain the sub-item scores of clock performance;

[0094] Calculate the transmission score based on the PTP packet loss rate and delay change rate in the standardized feature data, and calculate the reliability score based on the device state transition frequency and abnormal event frequency to obtain the sub - scores of network performance.

[0095] Comprehensively calculate the sub - scores of clock performance and network performance through a weighted calculation method to obtain a network performance score ranging from 0 to 100.

[0096] Extract the abnormal occurrence time, abnormal type, affected devices, and abnormal parameter values from the PTP network abnormal event records to generate an abnormal feature vector and obtain abnormal feature data.

[0097] Conduct root cause analysis based on the abnormal feature data, identify the device or link where the abnormality first occurred, and query the historical fault handling database to obtain the fault location result.

[0098] Specifically, extract clock stability data, synchronization accuracy data, network transmission data, and system reliability data from the network performance analysis results, and perform normalization processing to obtain standardized feature data. Clock stability data includes indicators such as the standard deviation of clock deviation, maximum deviation value, Allan variance, and deviation drift rate, which reflect the stability of the device clock operation; synchronization accuracy data includes indicators such as the average deviation value, deviation distribution, and maximum time interval error (MTIE), which reflect the accuracy of device clock synchronization; network transmission data includes indicators such as PTP packet loss rate, delay change rate, and packet round - trip time, which reflect the network transmission quality; system reliability data includes indicators such as device state transition frequency, abnormal event frequency, and lock - in stability time, which reflect the overall reliability of the system. Normalization processing is a process of converting data with different dimensions and magnitudes into a unified interval to ensure reasonable weights for each indicator. For each indicator, first determine its theoretical maximum and minimum values, and then use a linear normalization formula or interval mapping method to convert it to the [0, 1] interval. For example, for the standard deviation of clock deviation, when the original value is 5 nanoseconds, it is normalized to 0.95; when the original value is 500 nanoseconds, it is normalized to 0.1, reflecting the evaluation criterion of "the smaller the value, the better".

[0099] Calculate the stability score for the clock deviation standard deviation and the maximum deviation value in the standardized feature data, and calculate the accuracy score for the average deviation value and the deviation distribution to obtain the sub - score of clock performance. During the calculation of the clock stability score, first determine the weights of each index: the weight of the clock deviation standard deviation is relatively high (e.g., 0.6) because it directly reflects the clock stability; the weight of the maximum deviation value is the second (e.g., 0.3), reflecting the performance in the worst - case scenario; the weights of supplementary indicators such as Allan variance are relatively low (e.g., 0.1), serving as an auxiliary evaluation basis. Then apply the corresponding weights to each index and sum them up through weighted calculation to obtain the raw stability score within the range of [0, 1]. To better meet the actual evaluation requirements, apply a non - linear transformation function, such as the S - shaped function, to the raw score, so that the performance closer to the ideal state obtains a higher score, while the performance closer to the critical value has a rapid decline in score. Similarly, during the calculation of the accuracy score, the weight of the average deviation value is relatively high (e.g., 0.5), reflecting the overall synchronization accuracy; the weight of the deviation distribution is the second (e.g., 0.3), reflecting the accuracy stability; the weight of the maximum time interval error is relatively low (e.g., 0.2). Through weighted calculation and non - linear adjustment, the accuracy score within the range of [0, 1] is finally obtained. These two scores are weighted and combined again to form the sub - score of clock performance, reflecting the overall performance level of the device clock.

[0100] Calculate the transmission score for the PTP packet loss rate and the delay change rate in the standardized feature data, and calculate the reliability score for the device state transition frequency and the abnormal event frequency to obtain the sub - score of network performance. In the calculation of the transmission score, the weight of the PTP packet loss rate is the highest (e.g., 0.5) because packet loss directly affects the synchronization accuracy; the weight of the delay change rate is the second (e.g., 0.3), reflecting the network transmission stability; the weight of the round - trip time of the packet is relatively low (e.g., 0.2), serving as an auxiliary evaluation basis. These indicators are weighted and calculated and then applied with a logarithmic transformation function to be converted into the transmission score within the range of [0, 1]. When the packet loss rate is 0, the score approaches 1, and when the packet loss rate exceeds 1%, the score quickly drops below 0.5. In the calculation of the reliability score, both the device state transition frequency and the abnormal event frequency are important indicators. The former reflects the device state stability, and the latter reflects the frequency of abnormal occurrences. For the device state transition frequency, the fewer the number of state changes per hour, the higher the score; for the abnormal event frequency, count the number of various abnormal events in the past 24 hours, and the fewer the number, the higher the score. These two indicators are weighted and calculated to form the reliability score within the range of [0, 1]. The transmission score and the reliability score are weighted again to form the sub - score of network performance, comprehensively reflecting the network transmission and system reliability status.

[0101] The clock performance sub-item score and the network performance sub-item score are calculated by weighted calculation method to obtain a network performance score of 0-100. In the weighted calculation process, the weights of the two major categories of scores are first determined: in high-precision PTP application scenarios, the clock performance weight is usually higher (such as 0.7) because clock synchronization accuracy is the core goal; the network performance weight is lower (such as 0.3) as a supporting factor. The two scores are weighted and summed to obtain a comprehensive score in the [0,1] interval. This score is then converted to the 0-100 interval through linear mapping to form the final network performance score. This score adopts a multi-level structure, which can not only show the overall score of the entire network, but also be refined to the single score at the device level, which is convenient for accurately locating problems. The scoring system is also designed to take into account the differences in requirements of different application scenarios: the broadcast production environment has extremely high requirements for clock performance, and the weight setting is biased towards clock stability and accuracy; while the general industrial control environment has high requirements for network reliability, and the weight setting is biased towards transmission quality and system reliability. The scoring results are displayed through an intuitive dashboard, with different score ranges corresponding to different colors: 90-100 points are green, indicating excellent; 70-89 points are yellow, indicating good; 50-69 points are orange, indicating average; 0-49 points are red, indicating poor.

[0102] The abnormal occurrence time, abnormality type, involved devices and abnormal parameter values are extracted from the abnormal event records of the PTP network to generate abnormal feature vectors and obtain abnormal feature data. The abnormal feature vector is a structured data representation used to describe the key features of abnormal events. First, the abnormal event records are extracted in chronological order, and each record contains complete abnormal description information. For each record, the following key information is extracted: abnormal occurrence time, accurate to milliseconds; abnormality type, such as clock jump abnormality, lock state abnormality, clock drift abnormality, etc.; involved devices, including device identifiers, device types, PTP roles, etc.; abnormal parameter values, such as clock deviation change amplitude, drift rate, lock state duration, etc. This information is organized into a fixed-dimensional feature vector according to a predefined format, such as [timestamp, abnormality type code, device ID, parameter value 1, parameter value 2, ...]. In order to improve processing efficiency, the abnormality type is encoded: clock jump abnormality is encoded as 1, lock state abnormality is encoded as 2, clock drift abnormality is encoded as 3, and compound abnormality is encoded as a combination of them. In addition, for consecutive anomalies of the same type, time compression processing is performed to record the start time, end time and peak parameters of the anomaly to reduce data redundancy. The final anomaly feature data is a multidimensional matrix, where each row represents an abnormal event and each column represents a feature dimension, providing structured input for subsequent root cause analysis.

[0103] Perform root cause analysis based on abnormal feature data, identify the device or link where the abnormality first occurred, and query the historical fault handling database to obtain the fault location result. Root cause analysis is a process of tracing from the appearance to the essential cause. For PTP network abnormalities, first, apply the time series analysis method to sort the abnormal feature data according to the time of abnormality occurrence, and find the device or link where the abnormality first occurred, which are usually the origin points of the fault. Then, apply topological association analysis. Based on the PTP clock distribution topology, identify the propagation path of the abnormality in the network: if an upstream device has an abnormality and then downstream devices also have abnormalities one after another, the upstream device is very likely to be the root cause; if a device has an abnormality but its upstream and downstream devices are normal, there is likely a problem with the device itself. For complex situations, apply the causal reasoning algorithm to construct a causal graph of abnormal events and infer the most likely root cause device or link based on the Bayesian network model. After determining the potential root cause, the system queries the historical fault handling database to match similar abnormal patterns and their solutions. This database contains past fault cases, and each case includes abnormal features, root causes, and handling methods. Through similarity calculation, find the historical case most similar to the current abnormality, extract its fault cause and solution, and form a complete fault location result, including the faulty device or link, fault type, possible cause, and recommended handling method.

[0104] For example, during the operation of the PTP network of a certain ultra-high-definition live broadcast system, the monitoring system extracted multiple performance data of the core switch HC-01 from the network performance analysis results: the standard deviation of clock deviation was 15 nanoseconds, the maximum deviation value was 120 nanoseconds, the average deviation value was 35 nanoseconds, the deviation distribution showed a normal distribution, the PTP packet loss rate was 0.02%, the delay change rate was 0.5%, the device status conversion frequency was 1 time per 24 hours, and the abnormal event frequency was 2 times per week. These original data were normalized and converted into standardized feature data: the normalized value of the standard deviation of clock deviation was 0.85 (the smaller the original value, the better), the normalized value of the maximum deviation value was 0.78, the normalized value of the average deviation value was 0.82, the normalized value of the deviation distribution was 0.9, the normalized value of the PTP packet loss rate was 0.95, the normalized value of the delay change rate was 0.85, the normalized value of the device status conversion frequency was 0.8, and the normalized value of the abnormal event frequency was 0.75. Based on these normalized data, the scores of each item were calculated: the score of clock stability was 0.85×0.6 + 0.78×0.3 + 0.88×0.1 = 0.829, the score of accuracy was 0.82×0.5 + 0.9×0.3 + 0.85×0.2 = 0.85, the score of transmission was 0.95×0.5 + 0.85×0.3 + 0.9×0.2 = 0.915, and the score of reliability was 0.8×0.5 + 0.75×0.5 = 0.775. Further calculate the sub-item score of clock performance as 0.829×0.5 + 0.85×0.5 = 0.84, and the sub-item score of network performance as 0.915×0.6 + 0.775×0.4 = 0.859. The final comprehensive network performance score was (0.84×0.7 + 0.859×0.3)×100 = 84.6 points, which belonged to the good range. However, a few days later, the system extracted an anomaly from the PTP network abnormal event record: the boundary clock device BC-02 had a clock drift anomaly at 10:15:30.500, with a drift rate of 150 nanoseconds per hour. After 30 minutes, the two downstream terminal devices OC-05 and OC-06 successively showed abnormal locking states around 10:45:20. The system generated abnormal feature vectors [1677489330.5, 3, "BC-02", 150, 1800] and [1677491120, 2, "OC-05", 2700], [1677491150, 2, "OC-06", 2715]. Root cause analysis was based on time series and topological relationships, and it was determined that BC-02 was the device that first showed the anomaly, and its anomaly time was 30 minutes earlier than that of the downstream devices, highly matching the fault propagation pattern. The system queried the historical fault database and found 3 historical cases of similar patterns, and the root causes of 2 of these cases were that temperature fluctuations caused the crystal oscillator frequency to drift.The final fault location result indicates that the crystal oscillator frequency of the BC-02 device drifts due to abnormal temperature. It is recommended to check the device's heat dissipation system and environmental temperature control. According to this location, the maintenance personnel found that the air outlet of the air conditioner in the cabinet where the BC-02 device is located was partially blocked. After cleaning, the problem was solved and the network performance score returned to the normal level.

[0105] In a specific embodiment, the process of executing step S106 may specifically include the following steps:

[0106] Based on historical data and the current network status, apply the time series prediction algorithm to predict the future change trends of device clock stability, network delay, and jitter, and obtain the network performance prediction result;

[0107] Compare the network performance prediction result with the preset performance threshold. When the predicted value is lower than the performance threshold, obtain the predictive maintenance requirement list;

[0108] According to the fault location result and the low-scoring items in the network performance score, extract the corresponding optimization solutions from the optimization knowledge base and generate PTP configuration parameter optimization suggestions;

[0109] Conduct a simulation evaluation of the PTP configuration parameter optimization suggestions, calculate the performance improvement effect and potential risks after parameter adjustment, and obtain the optimization solution evaluation report;

[0110] Based on the predictive maintenance requirement list and the fault location result, analyze the long-term change trend of device clock stability, identify signs of hardware aging, and obtain the device health assessment result;

[0111] Generate a standardized maintenance work order including the maintenance object, maintenance content, and priority according to the optimization solution evaluation report and the device health assessment result.

[0112] Specifically, based on historical data and the current network status, a time series prediction algorithm is applied to predict the future trends of device clock stability changes, network latency, and jitter changes, obtaining network performance prediction results. The time series prediction algorithm is a set of methods for analyzing time patterns in historical data and predicting future trends. In PTP network monitoring, multiple prediction algorithms are used to process different types of time series data: for clock deviation data showing obvious periodic or seasonal variations, a seasonal ARIMA (Autoregressive Integrated Moving Average) model is used, which can capture trends, periodicity, and random fluctuations in the data; for relatively stable network latency data, exponential smoothing methods are used, including simple exponential smoothing, linear trend method, and seasonal trend method; for jitter data showing non-linear characteristics, an LSTM (Long Short-Term Memory) neural network model is used. The prediction process first extracts the historical data of each device in the past 30 days from the time series database and performs data cleaning and preprocessing, including removing outliers, filling missing values, and standardization; then selects an appropriate prediction model according to the data characteristics and optimizes the model parameters, such as the differencing order, autoregressive order, and moving average order of the ARIMA model, or the number of network layers and neurons of the LSTM model; finally, applies the trained model to predict the trends of device clock stability, network latency, and jitter changes within the next 7 days, forming a complete network performance prediction result.

[0113] Compare the network performance prediction results with the preset performance thresholds. When the predicted value is lower than the performance threshold, a predictive maintenance requirement list is obtained. The performance threshold is a critical value of the performance index set according to different device types and application scenarios, including clock deviation thresholds (such as ±500 nanoseconds), jitter thresholds (such as 200 nanoseconds), and network latency thresholds (such as 5 milliseconds), etc. The comparison process first analyzes the prediction results, extracts the predicted values at key time points, especially the turning points and peak points where the predicted values change significantly; then compares the predicted values of these key points with the corresponding performance thresholds, analyzes when and to what extent the predicted values will break through the thresholds; finally, conducts a risk assessment, considering the duration, scope of influence, and severity of breaking through the thresholds. When the predicted value of a certain device shows that it will be lower than (i.e., the performance deteriorates to) the preset threshold within a certain period in the future, the system adds this device to the predictive maintenance requirement list. The list is sorted in ascending order of the predicted time to break through the threshold, and each record includes the device identifier, the name of the performance index, the predicted breakthrough time, the trend of the predicted value change, and the recommended processing time window. This predictive-based maintenance requirement identification method transforms the maintenance work from passive response to active prevention, effectively avoiding system failures caused by sudden performance degradation.

[0114] According to the fault location results and the low-scoring items in the network performance score, extract the corresponding optimization solutions from the optimization knowledge base to generate PTP configuration parameter optimization suggestions. The optimization knowledge base is a structured database containing solutions to various PTP network problems, which is composed of historical optimization experiences, expert suggestions, and best practices. The extraction process first analyzes the fault location results to identify the fault type and potential causes; then queries the items in the network performance score with scores lower than the threshold (such as 70 points), such as the clock stability item, the transmission quality item, etc.; then uses this information as the retrieval condition to query the optimization knowledge base to find the optimization solution with the highest matching degree. For different types of problems, different optimization suggestions are generated: for the clock jump problem, it is recommended to adjust parameters such as the PTP announce interval, sync interval, and delay_req interval. For example, reduce the announce interval from the default 2 seconds to 1 second to increase the clock state update frequency; for the lock stability problem, it is recommended to adjust the PTP domain priority setting to ensure that a stable clock source is selected as the master clock; for the synchronization problem caused by network congestion, it is recommended to apply QoS (Quality of Service) policies to PTP packets to improve their transmission priority. The generated PTP configuration parameter optimization suggestions are presented in a structured form, including parameter name, current value, suggested value, adjustment reason, and expected effect, which is convenient for maintenance personnel to understand and implement. Conduct a simulation evaluation of the PTP configuration parameter optimization suggestions, calculate the performance improvement effect and potential risks after parameter adjustment, and obtain an optimization solution evaluation report. Simulation evaluation is a method to predict the optimization effect through a mathematical model or a simulation system before actual implementation of the optimization. The evaluation process first establishes a PTP network performance impact factor model, which describes the relationship between each configuration parameter and performance indicators, such as the relationship between the announce interval and lock stability, the sync interval and synchronization accuracy, and the delay_req interval and delay measurement accuracy; then inputs the parameter adjustment values in the optimization suggestions into the model to calculate the expected performance indicators after adjustment; then compares with the current performance indicators to quantify the improvement effect; finally, conducts a risk analysis to evaluate the possible negative impacts of parameter adjustment. For example, overly reducing the announce interval will increase the network load, and overly increasing the delay_req interval will reduce the delay measurement accuracy. The evaluation report includes analysis results in multiple dimensions: performance improvement estimation, indicating the expected changes in each performance indicator after parameter adjustment; resource consumption assessment, analyzing the impact of parameter adjustment on resources such as network bandwidth and processor load; compatibility analysis, evaluating the compatibility of parameter adjustment with existing devices; stable period estimation, predicting the time required for the system to reach a new stable state. These analysis results form a complete optimization solution evaluation report, providing a scientific basis for decision-making.

[0115] Based on the predictive maintenance requirements list and the fault location results, analyze the long-term change trend of the device clock stability, identify signs of hardware aging, and obtain the device health assessment results. Hardware aging refers to the phenomenon that the performance of the device gradually degrades with the increase of usage time. In the PTP network, it is mainly manifested as the aggravation of the clock oscillator frequency drift and the decrease of the synchronization accuracy. The analysis process first extracts the long-term (such as half a year or longer) clock stability data of the device from the time series database, including indicators such as clock deviation, drift rate, and Allan variance; then applies trend analysis methods, such as linear regression or curve fitting, to identify the long-term change trend of the indicators; then compares with the normal aging model of the device to judge whether there are signs of abnormal aging. Normal aging usually shows a slow and steady performance decline, while abnormal aging shows an accelerated degradation or mutation. In addition, the analysis will also consider the impact of environmental factors on the device performance, such as temperature fluctuations, power supply stability, etc. Based on these analyses, generate health assessment results for each device that appears in the predictive maintenance requirements list or the fault location results, including the device health status classification (such as normal, attention required, warning, dangerous), the estimated remaining service life, the main manifestations of aging, and the recommended handling methods. These health assessment results provide data support for device life cycle management and help formulate reasonable renewal and maintenance plans.

[0116] Generate a standardized maintenance work order containing the maintenance object, maintenance content, and priority according to the optimization plan evaluation report and the device health assessment results. The maintenance work order is a standardized document that guides maintenance personnel to perform specific maintenance tasks and contains all necessary information and processes. The generation process first comprehensively analyzes the optimization plan evaluation report and the device health assessment results to determine the type of maintenance operations to be performed, such as configuration optimization, hardware inspection, component replacement, etc.; then extracts the standard operation procedures and technical requirements from the maintenance operation library according to the operation type and device characteristics; then determines the maintenance priority, considering factors such as problem severity, impact scope, estimated failure time, and device importance; finally generates a complete standardized maintenance work order. The standardized work order includes the following contents: maintenance object information, which details the device or system to be maintained, including device identification, type, location, and key configurations; maintenance content, which specifically describes the operation steps, technical requirements, and precautions to be performed; a list of required tools and materials; the estimated working time and impact scope; maintenance priority and recommended execution time window; acceptance criteria and recovery process; references to associated fault analysis results and historical maintenance records. Through this standardized maintenance work order management, ensure the standardization, traceability, and effectiveness of maintenance work, and at the same time accumulate experience data for the handling of future similar problems.

[0117] For example, a PTP synchronization system operating in a UHD production and broadcast network deploys a PTP network performance monitoring method based on big data analysis. The monitoring system continuously collects and stores network performance data for 3 months. In the most recent analysis, an ARIMA(2,1,1) time series prediction model was applied to the boundary clock device BC-05 in the core area to process its historical clock deviation data. The model first performed a first-order difference operation on the original deviation data to obtain a stationary sequence, and then established a prediction model by combining 2nd-order autoregressive terms and 1st-order moving average terms. The prediction results show that the clock deviation value of BC-05 will gradually increase from the current 200 nanoseconds to 550 nanoseconds within the next 5 days, exceeding the system-set performance threshold of 500 nanoseconds. At the same time, applying the exponential smoothing method to predict the network delay data of this device shows that the delay will remain within the normal range. The system adds BC-05 to the predictive maintenance requirement list, marks it as "clock stability warning", and recommends processing within 3 days. Combining the previous fault location results, the clock performance score of this device is only 65 points, belonging to the low-scoring items. The system retrieves two matching optimization solutions from the optimization knowledge base: one is to adjust the PTP configuration parameters, shorten the sync interval from 1 second to 0.5 seconds, and shorten the delay_req interval from 2 seconds to 1 second; the other is to adjust the temperature control parameters of the device because historical data shows that the performance fluctuation of this device is highly correlated with environmental temperature changes. Simulating and evaluating these two optimization solutions, it is expected that the deviation value will drop to about 150 nanoseconds after the configuration parameter adjustment, but it will increase the network traffic by about 1%; the temperature control optimization is expected to reduce the deviation fluctuation by 50%, but it requires physical operation and a short device offline. At the same time, long-term trend analysis shows that the clock oscillator frequency drift rate of BC-05 has shown an accelerating growth trend in the past 3 months, rising from 50 nanoseconds per month to the current 120 nanoseconds per month, which conforms to the typical characteristics of oscillator aging. The device health assessment result classifies BC-05 as the "warning" state, and it is expected that the remaining effective life of the oscillator is less than 6 months. Based on these analysis results, the system generates a standardized maintenance work order: the maintenance object is "boundary clock device BC-05", the maintenance content includes "optimized configuration of PTP parameters" and "oscillator detection and scheduled replacement", the priority is set as "medium-high", and it is recommended to complete the first parameter optimization within 48 hours and replace the oscillator component during the next planned maintenance window. The maintenance personnel performed the parameter optimization according to the work order requirements, and the clock deviation immediately dropped to 140 nanoseconds, and the oscillator replacement was scheduled during the maintenance window one week later, effectively avoiding potential synchronization failures and ensuring the stable operation of the system.

[0118] The above describes the PTP network performance monitoring method based on big data analysis in the embodiments of the present application. Next, the PTP network performance monitoring system based on big data analysis in the embodiments of the present application will be described. Please refer to Figure 2, an embodiment of the PTP network performance monitoring system based on big data analysis in the embodiments of the present application includes:

[0119] A collection module 201, configured to collect PTP message data, network device PTP status data, and RTP stream data packets in the PTP network through multiple data collection points, and obtain a structured data stream;

[0120] A calculation module 202, configured to perform data preprocessing and clock error calculation on the structured data stream, and obtain a standardized data set including PTP network status parameters;

[0121] A building module 203, configured to build a PTP network performance time series database based on the standardized data set, and establish a multi-dimensional PTP network performance analysis model to obtain a network performance analysis result;

[0122] A detection module 204, configured to perform real-time status monitoring and anomaly detection according to the network performance analysis result, and obtain a PTP network anomaly event record;

[0123] An execution module 205, configured to perform machine learning analysis based on the PTP network anomaly event record and the network performance analysis result, and obtain a network performance score and a fault location result;

[0124] A maintenance module 206, configured to perform network performance optimization and predictive maintenance according to the network performance score and the fault location result, and generate optimization suggestions and maintenance work orders.

[0125] Through the collaborative cooperation of the above - mentioned components, PTP packet data, network device PTP status data, and RTP stream data packets in the PTP network are collected through multiple data collection points to obtain a structured data stream, achieving comprehensive and accurate data collection and providing a high - quality data basis for subsequent analysis. Data pre - processing and clock error calculation are performed on the collected structured data stream to obtain a standardized data set containing PTP network status parameters, effectively eliminating the influence of network jitter and improving the accuracy of clock deviation calculation. Based on the standardized data set, a PTP network performance time - series database is constructed, and a multi - dimensional PTP network performance analysis model is established to obtain network performance analysis results. Through a distributed storage architecture and a time - span hierarchical storage strategy, efficient storage and query of massive time - series data are realized. Moreover, the long short - term memory network model gives full play to the advantages of artificial intelligence algorithms in time - series data prediction, accurately capturing long - term dependence relationships and short - term fluctuation characteristics in the PTP network, and significantly improving the accuracy and timeliness of performance prediction. According to the network performance analysis results, real - time status monitoring and anomaly detection are performed to obtain PTP network anomaly event records, achieving intuitive visualization of the network status and accurate anomaly detection. The correlation analysis algorithm effectively filters out short - term anomalies and reduces the false alarm rate. Machine learning analysis is performed based on the PTP network anomaly event records and network performance analysis results to obtain network performance scores and fault location results. Among them, machine learning algorithms, by learning historical anomaly patterns, greatly improve the accuracy and efficiency of fault root cause location, transforming complex multi - dimensional data into intuitive performance scores for convenient management decision - making. According to the network performance scores and fault location results, network performance optimization and predictive maintenance are performed to generate optimization suggestions and maintenance work orders. With the help of time - series prediction algorithms, forward - looking prediction of network performance is carried out, transforming maintenance from passive response to active prevention. The simulation evaluation mechanism of the optimization suggestions ensures the effectiveness and safety of the optimization measures. The overall solution deeply integrates big data collection, artificial intelligence analysis, and PTP network monitoring. Especially in time - series data analysis, the application of the long short - term memory network gives full play to the advantages of deep learning algorithms in processing time - series data, and can accurately capture complex patterns and long - term dependence relationships in PTP clock data; in the anomaly detection link, unsupervised learning algorithms can automatically identify anomaly patterns from massive data without pre - defining all anomaly types; in the fault root cause analysis, causal inference algorithms can accurately trace the anomaly propagation path and locate the root cause. These characteristics of artificial intelligence algorithms and models directly improve the accuracy, forward - looking, and automation of PTP network monitoring, realizing the transformation from passive monitoring to active prevention.

[0126] Above Figure 2The PTP network performance monitoring system based on big data analysis in the embodiments of the present invention is described in detail from the perspective of modular functional entities. Next, the PTP network performance monitoring device based on big data analysis in the embodiments of the present invention is described in detail from the perspective of hardware processing.

[0127] Figure 3 FIG. 4 is a schematic structural diagram of a PTP network performance monitoring device based on big data analysis provided by an embodiment of the present invention. The PTP network performance monitoring device 300 based on big data analysis may vary greatly due to configuration or performance differences, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 333 or data 332 (for example, one or more mass storage device terminals). Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the PTP network performance monitoring device 300 based on big data analysis. Further, the processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the PTP network performance monitoring device 300 to implement the steps of the above-mentioned PTP network performance monitoring method based on big data analysis.

[0128] The PTP network performance monitoring device 300 based on big data analysis may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 3 the shown structural diagram of the PTP network performance monitoring device based on big data analysis does not limit the PTP network performance monitoring device provided by the present invention, and may include more or fewer components than shown, or combine some components, or have different component arrangements.

[0129] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium, and when the instructions are run on a computer, the computer is made to execute the steps of the PTP network performance monitoring method based on big data analysis.

[0130] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0131] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a PTP network performance monitoring device based on big data analysis (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0132] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.

Claims

1. A method for monitoring the performance of a PTP network based on big data analysis, characterized in that, The method includes: Collecting PTP message data, network device PTP status data, and RTP stream data packets in the PTP network through multiple data collection points to obtain a structured data stream; Performing data preprocessing and clock error calculation on the structured data stream to obtain a standardized data set containing PTP network status parameters; Constructing a PTP network performance time series database based on the standardized data set and establishing a multi-dimensional PTP network performance analysis model to obtain network performance analysis results; Performing real-time status monitoring and anomaly detection according to the network performance analysis results to obtain PTP network anomaly event records; Performing machine learning analysis based on the PTP network anomaly event records and the network performance analysis results to obtain network performance scores and fault location results; Performing network performance optimization and predictive maintenance according to the network performance scores and the fault location results to generate optimization suggestions and maintenance work orders.

2. The PTP network performance monitoring method based on big data analysis according to claim 1, wherein The collecting PTP message data, network device PTP status data, and RTP stream data packets in the PTP network through multiple data collection points to obtain a structured data stream includes: Connecting a fiber optic network card through a deployed collection server to access the IP data stream and obtaining the hardware timestamp of the network data packet through the MII interface to obtain timestamp information with nanosecond-level accuracy; Parsing PTP synchronization messages, follow-up messages, delay request messages, and delay response messages, and extracting the timestamp information, message sequence numbers, and address information therein to obtain PTP message data; Collecting data from network devices through the NetConf and SNMP protocols to obtain the PTP lock status, Grandmaster ID, and Clock ID to obtain network device PTP status data; Obtaining data from devices that do not support the standard management protocol through the SSH protocol and converting the data into JSON format through regular expressions to obtain standardized device status data; Collecting RTP stream data packets received by the terminal device, and extracting the source address, destination address, sequence number, and RTP timestamp to obtain media stream clock data; Adding a collection timestamp and collection point identification information to the PTP message data, the network device PTP status data, the standardized device status data, and the media stream clock data to obtain a structured data stream.

3. The PTP network performance monitoring method based on big data analysis according to claim 1, wherein The performing data preprocessing and clock error calculation on the structured data stream to obtain a standardized data set containing PTP network status parameters includes: Extracting the transmission timestamp and reception timestamp of the PTP synchronization message, and the transmission timestamp of the delay request message and the reception timestamp of the delay response message from the structured data stream to obtain a timestamp data set; Calculating the path delay value according to the timestamp data set, and averaging the sum of the synchronization message time difference and the delay message time difference to obtain a network path delay parameter; Calculating the clock offset based on the network path delay parameter, and subtracting the network path delay parameter from the synchronization message time difference to obtain an original clock offset value; Applying the Kalman filter algorithm to the original clock offset value for jitter removal processing to obtain a stable clock deviation value; Identify and remove outliers in the structured data stream, unify data from different sources into a standard format, and obtain a cleaned data set. Associate and integrate the cleaned data set with the stable clock deviation value to form a standardized data set containing the PTP locking status of network devices, clock source information, clock deviation value, jitter value, and medium delay value.

4. The PTP network performance monitoring method based on big data analysis according to claim 1, wherein Build a PTP network performance time series database based on the standardized data set, and establish a multi-dimensional PTP network performance analysis model to obtain network performance analysis results, including: Store the standardized data set in a time series database using a distributed storage architecture, and perform hierarchical storage according to the time span to obtain an efficient PTP data storage structure. Extract features from the clock deviation, jitter value, and network delay data in the PTP data storage structure, construct an input feature vector, and obtain a training data set. Build a multi-dimensional PTP network performance analysis model based on the long short-term memory network. The multi-dimensional PTP network performance analysis model includes an input layer, two long short-term memory hidden layers, and a fully connected output layer to obtain an initial model structure. Train the initial model structure using historical data, and use the mean squared error loss function and the Adam optimization algorithm to iteratively optimize the model parameters to obtain a trained multi-dimensional PTP performance analysis model. Input the real-time standardized data set into the trained multi-dimensional PTP performance analysis model, and obtain the clock stability prediction value and the network performance prediction value through forward propagation calculation. According to the clock stability prediction value and the network performance prediction value, combined with the PTP clock distribution topology information, generate a network performance analysis result including clock stability evaluation, performance trend prediction, and abnormal risk assessment.

5. The PTP network performance monitoring method based on big data analysis according to claim 1, characterized in that Execute real-time status monitoring and anomaly detection according to the network performance analysis result to obtain PTP network anomaly event records, including: Build a complete PTP distribution topology map based on the network performance analysis result, and display the Grandmaster device, boundary clock device, and ordinary clock device in the network and their master-slave relationships in real time to obtain a PTP network real-time status map. Color-code and mark the PTP locking status, clock deviation value, and jitter value of each device in the PTP network real-time status map to obtain an intuitive device status view. Set a multi-level threshold judgment standard according to the network performance analysis result, and perform real-time detection for the case where the clock deviation change value exceeds the preset threshold to obtain a clock jump anomaly mark. Based on the locking status information in the network performance analysis result, identify the change in the PTP locking status of the device or the situation of being unlocked for a long time to obtain a locking status anomaly alarm. Judge the clock deviation trend in the network performance analysis result, and identify the situation where the unidirectional continuous growth trend exceeds the preset drift rate to obtain a clock drift anomaly mark. Integrate and record the clock jump anomaly mark, the locking status anomaly alarm, and the clock drift anomaly mark in time series, and apply correlation analysis to filter out short-term anomalies to obtain PTP network anomaly event records.

6. The PTP network performance monitoring method based on big data analysis according to claim 1, characterized in that Performing machine learning analysis based on the PTP network anomaly event record and the network performance analysis result to obtain a network performance score and a fault location result, including: Extracting clock stability data, synchronization accuracy data, network transmission data, and system reliability data from the network performance analysis result, and performing normalization processing to obtain standardized feature data; Calculating a stability score for the standard deviation of clock deviation and the maximum deviation value in the standardized feature data, and calculating an accuracy score for the average deviation value and the deviation distribution to obtain a sub - score for clock performance; Calculating a transmission score for the PTP packet loss rate and the delay variation rate in the standardized feature data, and calculating a reliability score for the device state transition frequency and the anomaly event frequency to obtain a sub - score for network performance; Comprehensively calculating the sub - score for clock performance and the sub - score for network performance through a weighted calculation method to obtain a network performance score ranging from 0 to 100; Extracting the anomaly occurrence time, anomaly type, affected devices, and anomaly parameter values from the PTP network anomaly event record to generate an anomaly feature vector and obtain anomaly feature data; Performing root cause analysis based on the anomaly feature data, identifying the device or link where the anomaly first occurred, and querying the historical fault handling database to obtain a fault location result.

7. The method for monitoring the performance of a PTP network based on big data analysis according to claim 1, wherein Performing network performance optimization and predictive maintenance according to the network performance score and the fault location result, and generating optimization suggestions and maintenance work orders, including: Applying a time - series prediction algorithm based on historical data and the current network state to predict the future change trends of device clock stability, network delay, and jitter to obtain a network performance prediction result; Comparing the network performance prediction result with a preset performance threshold, and obtaining a list of predictive maintenance requirements when the predicted value is lower than the performance threshold; Extracting the corresponding optimization plan from the optimization knowledge base according to the fault location result and the low - scoring item in the network performance score to generate an optimization suggestion for PTP configuration parameters; Performing a simulation evaluation on the optimization suggestion for PTP configuration parameters, calculating the performance improvement effect and potential risks after parameter adjustment to obtain an optimization plan evaluation report; Analyzing the long - term change trend of device clock stability based on the list of predictive maintenance requirements and the fault location result, and identifying signs of hardware aging to obtain a device health assessment result; Generating a standardized maintenance work order including maintenance objects, maintenance content, and priorities according to the optimization plan evaluation report and the device health assessment result.

8. A PTP network performance monitoring system based on big data analysis, characterized in that, For implementing the PTP network performance monitoring method based on big data analysis as described in any one of claims 1 - 7, the PTP network performance monitoring system based on big data analysis includes: An acquisition module, configured to acquire PTP packet data, PTP status data of network devices, and RTP stream data packets in the PTP network through multiple data acquisition points to obtain a structured data stream; A calculation module, configured to perform data pre - processing and clock error calculation on the structured data stream to obtain a standardized data set including PTP network state parameters; A building module, configured to build a PTP network performance time-series database based on the standardized data set, and establish a multi-dimensional PTP network performance analysis model to obtain network performance analysis results; A detection module, configured to perform real-time status monitoring and anomaly detection according to the network performance analysis results to obtain PTP network anomaly event records; An execution module, configured to perform machine learning analysis based on the PTP network anomaly event records and the network performance analysis results to obtain network performance scores and fault location results; A maintenance module, configured to perform network performance optimization and predictive maintenance according to the network performance scores and the fault location results, and generate optimization suggestions and maintenance work orders.

9. A PTP network performance monitoring device based on big data analysis, characterized in that, It includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the PTP network performance monitoring method based on big data analysis according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program runs on the processor, it causes the processor to execute the PTP network performance monitoring method based on big data analysis according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Network traffic modeling and predicting method and device based on mixed integer programming

    CN115941511A

  • Dynamic multi-link intelligent management and scheduling system

    CN119420691A