A real-time quality monitoring and self-repair system and control method for IPTV edge nodes
Through the real-time quality monitoring system of edge nodes, using One-ClassSVM and LSTM models for anomaly identification and trend prediction, and dynamically adjusting buffers and protocols, the detection delay and response delay problems of the IPTV system are solved, and adaptive self-repair and low-cost deployment are achieved.
Patent Information
- Application Number
- CN202510642040.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-05-19
AI Technical Summary
Existing IPTV systems suffer from high detection delays, weak prediction and prevention capabilities, passive alarms, and a lack of real-time decision-making capabilities on the end side in ensuring service quality. This results in fault response delays of up to minutes and an inability to achieve active self-healing.
A real-time quality monitoring system based on edge nodes is adopted, including data collection, anomaly detection, trend prediction and adaptive decision-making modules. One-Class SVM and LSTM models are used for anomaly identification and trend prediction, and buffers and protocols are dynamically adjusted to achieve self-repair.
It achieves millisecond-level fault response capability, proactive prediction and prevention, adaptive dynamic optimization, reduced bandwidth consumption, user privacy protection, and low-cost deployment.
Smart Images

Figure CN120186141B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous intelligent operation and maintenance of streaming media, and in particular to an edge node-based IPTV real-time playback quality monitoring and self-repair system and control method. Background Art
[0002] In current IPTV systems, service quality assurance primarily relies on centralized cloud-based monitoring systems. However, these systems often fail to fully utilize the computing and storage capabilities of edge devices (such as set-top boxes). Furthermore, existing centralized cloud-based monitoring systems often rely on centralized processing, lacking device-side autonomous intelligence and failing to integrate proactive protection technologies like machine learning and trend prediction into the quality assurance process.
[0003] This leads to the following problems in the service quality assurance of existing IPTV systems:
[0004] 1. High detection delay: The terminal needs to upload data to the cloud for analysis, which usually results in a delay of several seconds to tens of seconds.
[0005] 2. Weak prediction and prevention capabilities: Traditional solutions can only rely on “discovery-alarm-manual processing” and cannot achieve automatic predictive intervention;
[0006] 3. Passive alarms: Traditional systems rely on user complaints or cloud-based batch analysis to generate alarms. They lack real-time decision-making capabilities on the device side, resulting in fault response delays of up to minutes and an inability to achieve proactive self-healing.
[0007] Therefore, there is an urgent need for an edge node-based IPTV real-time playback quality monitoring and self-repair system and control method to solve the above problems. Summary of the Invention
[0008] The purpose of the present invention is to provide an IPTV edge node real-time quality monitoring, self-repair system and control method, which improves the fault response capability of the IPTV system while realizing adaptive dynamic optimization of the system.
[0009] To achieve the above-mentioned purpose, the present invention is implemented through the following technical solutions:
[0010] On the one hand, a real-time quality monitoring and self-repair system for IPTV edge nodes is provided, comprising:
[0011] Data acquisition module, used to: collect network status data and playback status data in real time;
[0012] Anomaly detection module, which is used to identify anomalies using threshold detection, time series analysis algorithms, and One-Class SVM.
[0013] Trend prediction module, used to: use ARIMA model to predict network status;
[0014] Adaptive decision module, used to: Equipped with QoE model, dynamically optimize bitrate;
[0015] Executor module, used to: dynamically switch protocols and adjust buffer size.
[0016] On the other hand, a control method for the above-mentioned IPTV edge node real-time quality monitoring and self-repair system is provided, characterized in that it includes the following steps:
[0017] Step S1: The data acquisition module sets the acquisition frequency, collects network status data and playback status data, and pre-processes the collected network status data and playback status data;
[0018] Step S2: The anomaly monitoring module sets a static threshold and constructs a One-Class SVM based on the pre-processed network status data and playback status data to identify anomalies. It then classifies and prioritizes the identified anomalies and generates a unique event ID.
[0019] Step S3: The trend prediction module uses sliding window mean prediction and LSTM long-term trend prediction to perform trend prediction based on the unique event ID;
[0020] Step S4: The adaptive decision module sets the optimization target of the QoE model and sets the bitrate switching rules based on the trend prediction results;
[0021] Step S5: The executor module adjusts the buffer size according to the pre-processed network status data and the playback status dynamic data switching protocol.
[0022] Preferably, the step S1 is specifically:
[0023] The data acquisition module collects network status data and playback status data in real time at a frequency of 10 times per second. The network status data includes: packet loss rate ,bandwidth , round-trip delay The playback status data includes: the number of freezes , buffering time , Startup Delay ; Use Z-Score to normalize the collected raw network status data and playback status data to eliminate dimensional differences:
[0024] ;
[0025] in, is the mean of historical data, is the standard deviation, set and The update duration is used to adapt to network fluctuations;
[0026] Use DBSCAN clustering algorithm to mark and remove outliers. If the data point Dissatisfied , then mark this data point as an outlier.
[0027] Preferably, the step S2 includes:
[0028] Step S21: Perform anomaly detection using static thresholds. or If it lasts for two seconds, it will trigger the network abnormality judgment; when or , then the playback abnormality judgment is triggered;
[0029] Step S22: construct a One-ClassSVM dynamic learning model, use historical normal data to train the model, and use the RBF kernel function;
[0030] Step S23: Establish an inference model and calculate anomaly scores in real time:
[0031] ;
[0032] Among them, the weight coefficient , threshold Dynamic adjustment;
[0033] Step S24: Classify and prioritize the abnormal situations and generate a unique event ID.
[0034] Preferably, in step S24, the abnormal situations are classified and prioritized, including:
[0035] Network anomaly: The trigger condition is bandwidth drop or sudden increase in packet loss rate, and the priority is high;
[0036] Player abnormality: The trigger condition is that the number of freezes exceeds the limit, and the priority is medium;
[0037] Abnormal user operation: The trigger condition is that the channel switching frequency exceeds the threshold, and the priority is low.
[0038] Preferably, in step S3, the sliding window mean prediction is specifically:
[0039] The input data is the bandwidth sequence of the last 5 seconds , the prediction process is:
[0040] ;
[0041] LSTM long-term trend prediction is specifically as follows:
[0042] The model input layer is historical data of 5 time steps, the model hidden layer is 32 LSTM units, the activation function is ReLU, the Dropout rate is 0.2, and the output layer is the bandwidth prediction value for the next 3 seconds. ; Use the mean square error loss function to train the model; the model optimizer is Adam, and the learning rate , training cycle ,The data set takes 7 days of historical data, 80% of the data is used as the training set and 20% of the data is used as the validation set for training.
[0043] Preferably, the step S4 includes:
[0044] Step S41: Set the QoE optimization target to minimize the weighted sum of the freeze rate and image quality loss. The calculation process is:
[0045] ;
[0046] Wherein, StallRate(r) represents the playback stall rate at bit rate r, and is calculated as follows: ; QualityLoss(r) represents the image quality loss caused by reducing the bit rate r, using the peak signal-to-noise ratio The drop value is quantified using the formula: ; 0.7 and 0.3 are weight coefficients;
[0047] Step S42: Set the code rate switching rule: If , switch to 2Mbps bit rate; if the predicted bandwidth continues to decrease, reduce the bit rate to the next gradient in advance.
[0048] Preferably, the step S5 includes:
[0049] Step S51: Packet loss rate or bandwidth fluctuations When switching from HTTP-FLV to HLS protocol, the switching delay , users perceive lag ;
[0050] Bandwidth fluctuations is the bandwidth variation per unit time, Switching delay The time difference between the triggering and taking effect of the protocol switching operation, including but not limited to the time taken for protocol stack reloading and buffer reset.
[0051] Step S52: The process of dynamically adjusting the capacity of the buffer zone is as follows:
[0052] ;
[0053] in, To predict the delay increment, specifically: predict the difference between the future network delay and the current delay, which is used to dynamically adjust the buffer size. The formula is: ,when , restores the default buffer size.
[0054] Compared with the prior art, the beneficial effects of the present invention are:
[0055] 1. Millisecond-level fault response capability: Since all monitoring, analysis, decision-making, and execution are completed locally on the set-top box, without uploading to the cloud, end-to-end latency is less than 100ms. This is achieved through closed-loop control using the MAPE-K model (Monitor-Analyze-Plan-Execute-Knowledge).
[0056] 2. Proactive prediction and prevention: The sliding window averaging method and LSTM prediction model are introduced to predict bandwidth decline trends before failures occur, allowing for proactive adjustments to playback strategies.
[0057] 3. Adaptive dynamic optimization: Dynamically adjust system strategies based on user experience feedback (such as reduced playback interruption rate and reduced buffering times), and the self-learning mechanism continuously iterates and improves the repair effect;
[0058] 4. Bandwidth saving and privacy protection: Only summary logs are reported, without transmitting original playback data; ensuring the privacy of user viewing behavior and meeting data compliance requirements;
[0059] 5. Low-cost deployment: It can be run on existing low-end set-top boxes (such as RK3228); no hardware modification is required, and it can be achieved only through software upgrades. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a schematic diagram of the system structure of the present invention;
[0061] Figure 2 It is a flow chart of the control method of the present invention. DETAILED DESCRIPTION
[0062] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall within the scope limited by the application equally.
[0063] In the present invention, terms such as "upper", "lower", "left", "right", "front", "back", "vertical", "horizontal", "side", "bottom", etc. indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. They are relational words determined only for the convenience of describing the structural relationships of the various parts or elements of the present invention, and do not specifically refer to any part or element in the present invention, and should not be understood as limiting the present invention.
[0064] Example:
[0065] like Figure 1 As shown, this embodiment provides an IPTV edge node real-time quality monitoring and self-repair system, including:
[0066] Data acquisition module, used to: collect network status data and playback status data in real time;
[0067] Anomaly detection module, which is used to identify anomalies using threshold detection, time series analysis algorithms, and One-Class SVM.
[0068] Trend prediction module, used to: use ARIMA model to predict network status;
[0069] Adaptive decision module, used to: Equipped with QoE model, dynamically optimize bitrate;
[0070] Executor module, used to: dynamically switch protocols and adjust buffer size.
[0071] In this embodiment, the test scenario is:
[0072] The test was conducted in residential area A, a neighborhood with high user density and frequent network interference. A mainstream set-top box (STB) running Android 4.4 and 2GB of RAM was used as the test site. The network environment used was home broadband (bandwidth fluctuating between 2 and 12 Mbps) with a peak latency of 280ms. The test involved watching live sports events at a bitrate gradient of [8M, 6M, 4M, 2M].
[0073] like Figure 2 As shown, this embodiment also provides a control method for the above-mentioned IPTV edge node real-time quality monitoring and self-repair system, which, combined with the test scenario, specifically includes the following steps:
[0074] Step S1: The data acquisition module sets the acquisition frequency, collects network status data and playback status data, and pre-processes the collected network status data and playback status data;
[0075] Specifically:
[0076] The data acquisition module collects network status data and playback status data in real time at a frequency of 10 times per second. The network status data includes: packet loss rate ,bandwidth , round-trip delay , playback status data includes: number of freezes , buffering time , Startup Delay ; Use Z-Score to normalize the collected raw network status data and playback status data to eliminate dimensional differences:
[0077] ;
[0078] in, is the mean of historical data, is the standard deviation, set and The update time is 5 seconds to adapt to network fluctuations;
[0079] Using DBSCAN clustering algorithm (parameters: neighborhood radius , minimum number of samples ), mark and remove abnormal points, if the data point Dissatisfied , then mark this data point as an outlier.
[0080] Step S2: The anomaly monitoring module sets a static threshold and constructs a One-Class SVM based on the pre-processed network status data and playback status data to identify anomalies. It then classifies and prioritizes the identified anomalies and generates a unique event ID. Specifically:
[0081] Anomaly detection via static thresholds, or If it lasts for two seconds, it will trigger the network abnormality judgment. or , then the playback abnormality judgment is triggered;
[0082] Construct a One-ClassSVM dynamic learning model, use historical normal data (window is the last 30 minutes) to train the model, and use the kernel function RBF ( );
[0083] Model inference, real-time calculation of anomaly scores:
[0084] ;
[0085] Among them, the weight coefficient , threshold Dynamic adjustment;
[0086] Abnormal situations are classified and prioritized. Network abnormalities are triggered by bandwidth drop or sudden increase in packet loss rate, with a high priority (needing immediate action). Player abnormalities are triggered by excessive freezes, with a medium priority (needing a response within 5 seconds). User operation abnormalities are triggered by channel switching frequency exceeding a threshold (e.g., 10 times / minute), with a low priority (only logging).
[0087] In this embodiment, the calculation result of the abnormal scoring model is , exceeding the threshold of 0.75; the initial abnormality type was judged to be "network fluctuation-type lag", and the system recorded the abnormal event IDEX_20250428_#1; the user end has not yet triggered a complaint event, but the early warning system has pushed it to the front-end status bar.
[0088] Step S3: The trend prediction module uses sliding window mean prediction and LSTM long-term trend prediction to perform trend prediction based on the unique event ID. Specifically:
[0089] Sliding window mean prediction: input data is the bandwidth sequence of the last 5 seconds , the prediction formula is:
[0090] ;
[0091] LSTM long-term trend prediction: The model input layer is historical data of 5 time steps, the model hidden layer is 32 LSTM units, the activation function is ReLU, the dropout rate is 0.2, and the output layer is the bandwidth forecast value for the next 3 seconds , model training: loss function mean square error (MSE), optimizer: Adam, learning rate , training cycle ,The data set takes 7 days of historical data, with a ratio of 80% for training and 20% for verification;
[0092] Based on the event ID from step S2 above:
[0093] Input data: The bandwidth sequence in the last 5 seconds is: [6.1, 5.3, 4.6, 3.9, 3.3] Mbps;
[0094] Model parameters: p=2, d=1, q=2, determined by minimizing the AIC criterion;
[0095] Prediction result: trend average for the next 10 seconds ;
[0096] The system has determined that there is a "continuously downward bandwidth" trend.
[0097] Step S4: The adaptive decision module sets the optimization target of the QoE model and the bitrate switching rules based on the trend prediction results; specifically:
[0098] The QoE optimization goal is to minimize the weighted sum of the jamming rate and image quality loss. The calculation formula is:
[0099] ;
[0100] Set the bitrate switching rule: If , switch to 2Mbps bitrate; if the predicted bandwidth continues to decrease, reduce the bitrate to the next gradient in advance.
[0101] Step S5: The executor module adjusts the buffer size according to the pre-processed network status data and the playback status dynamic data switching protocol; specifically:
[0102] Protocol switching operation, packet loss rate or bandwidth fluctuations When switching from HTTP-FLV to HLS protocol, the switching delay , ;
[0103] Dynamic adjustment of buffer zone: Expansion formula:
[0104] Where ΔT is the predicted delay increment (e.g. ), when the network is stable ( ), gradually restore the default buffer size.
[0105] Step S6: Kernel thread priority adjustment: Raise the decoding thread priority to real-time level (RTpriority99) (RTpriority stands for real-time thread priority, ranging from 0 to 99 (higher values mean higher priority)). Limit the CPU usage of non-critical threads (such as log upload) to ≤10%;
[0106] Feedback optimization and reinforcement learning, configure Q-Learning model training: State definition:
[0107] ;
[0108] Action Space , reward function design:
[0109] ;
[0110] ;
[0111] ;
[0112] Synchronize cloud log data at 3:00 AM every day and retrain the LSTM and Q-Learning models;
[0113] Update the terminal policy library and push it to the set-top box through differential update (DeltaUpdate);
[0114] Local logs record event IDs, timestamps, exception types, remediation actions, execution delays, and QoE indicator changes. Using JSON compression for storage, a single log entry is ≤1KB in size.
[0115] Asynchronous reporting uses the MQTT protocol (QoS=1) to ensure at least one delivery. Data is compressed using the LZ4 algorithm with a compression rate of ≥80%. The transmission rate is limited to 10kbps to avoid affecting playback traffic.
[0116] After analyzing the logs on the cloud, global optimization instructions are generated (such as adjusting the threshold or update model parameters);
[0117] The instructions are sent to the terminal through the HTTPS encrypted channel, and the terminal performs the update during its idle period.
[0118] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A control method for a real-time quality monitoring and self-repair system for an IPTV edge node, characterized in that: The following steps are involved: Step S1: Setting the collection frequency, collecting network status data and playback status data, and pre-processing the collected network status data and playback status data; Step S2: Based on the pre-processed network status data and playback status data, a static threshold is set and a One-Class SVM is constructed to identify anomalies. After classifying and prioritizing the identified anomalies, a unique event ID is generated. Step S3: Based on the unique event ID, trend prediction is performed using sliding window mean prediction and LSTM long-term trend prediction; Step S4: According to the trend prediction results, set the optimization target of the QoE model and set the bitrate switching rules; Step S5: adjusting the buffer size according to the pre-processed network status data and the playback status dynamic data switching protocol; Step S2 includes: Step S21: Perform anomaly detection using static thresholds. or If it lasts for two seconds, it will trigger the network abnormality judgment; when or , then the playback abnormality judgment is triggered; Step S22: construct a One-ClassSVM dynamic learning model, use historical normal data to train the model, and use the RBF kernel function; Step S23: Establish an inference model and calculate anomaly scores in real time: ; Among them, the weight coefficient , threshold Dynamic adjustment; Step S24: Classify and prioritize the abnormal situations and generate a unique event ID; Step S4 includes: Step S41: Set the QoE optimization target to minimize the weighted sum of the freeze rate and image quality loss. The calculation process is: ; Wherein, StallRate(r) represents the playback stall rate at bit rate r, and is calculated as follows: ; QualityLoss(r) represents the image quality loss caused by reducing the bit rate r, using the peak signal-to-noise ratio The drop value is quantified using the formula: ; 0.7 and 0.3 are weight coefficients; Step S42: Set the code rate switching rule: If , switch to 2Mbps bitrate; if the predicted bandwidth continues to decrease, reduce the bitrate to the next gradient in advance; In step S24, the abnormal situations are classified and prioritized, including: Network anomaly: The trigger condition is bandwidth drop or sudden increase in packet loss rate, and the priority is high; Player abnormality: The trigger condition is that the number of freezes exceeds the limit, and the priority is medium; Abnormal user operation: The trigger condition is that the channel switching frequency exceeds the threshold, and the priority is low; The step S5 comprises: Step S51: Packet loss rate or bandwidth fluctuations When switching from HTTP-FLV to HLS protocol, the switching delay , users perceive lag ; Bandwidth fluctuations is the bandwidth variation per unit time, Switching delay The time difference between the triggering and taking effect of the protocol switching operation, including but not limited to the time taken for protocol stack reloading and buffer reset; Step S52: The process of dynamically adjusting the capacity of the buffer zone is as follows: ; in, To predict the delay increment, specifically: predict the difference between the future network delay and the current delay, which is used to dynamically adjust the buffer size. The formula is: ,when , restores the default buffer size.
2. A control method according to claim 1, characterized in that: The step S1 is specifically as follows: The data acquisition module collects network status data and playback status data in real time at a frequency of 10 times per second. The network status data includes: packet loss rate ,bandwidth , round-trip delay The playback status data includes: the number of freezes , buffering time , Startup Delay ; Use Z-Score to normalize the collected raw network status data and playback status data to eliminate dimensional differences: ; in, is the mean of historical data, is the standard deviation, set and The update duration is used to adapt to network fluctuations; Use DBSCAN clustering algorithm to mark and remove outliers. If the data point Dissatisfied , then mark this data point as an outlier.
3. A control method according to claim 1, characterized in that: In step S3, the sliding window mean prediction is specifically as follows: The input data is the bandwidth sequence of the last 5 seconds , the prediction process is: ; LSTM long-term trend prediction is specifically as follows: The model input layer is historical data of 5 time steps, the model hidden layer is 32 LSTM units, the activation function is ReLU, the Dropout rate is 0.2, and the output layer is the bandwidth prediction value for the next 3 seconds. , ; Use the mean square error loss function to train the model; The model optimizer is Adam, and the learning rate , training cycle ,The data set takes 7 days of historical data, 80% of the data is used as the training set and 20% of the data is used as the validation set for training.
4. A real-time quality monitoring and self-repair system for IPTV edge nodes, used to implement the control method according to any one of claims 1 to 3, characterized in that: include: Data acquisition module, used to: collect network status data and playback status data in real time; Anomaly detection module, which is used to identify anomalies using threshold detection, time series analysis algorithms, and One-Class SVM. Trend prediction module, used to: use ARIMA model to predict network status; Adaptive decision module, used to: Equipped with QoE model, dynamically optimize bitrate; Executor module, used to: dynamically switch protocols and adjust buffer size.
Citation Information
Patent Citations
Method and device for switching audio channels of communication module
CN118538247A
Wide fixed network user optical link quality standard analysis and monitoring method
CN119853796A
Network anomaly detection and automatic processing method and system, storage medium and equipment
CN119892476A
Dynamic data flow monitoring system and method for data center
CN119945949A
Video transmission method, apparatus and device, and storage medium
WO2024120134A1