An audio and video code stream real-time monitoring and diagnosing method and system
By constructing causal graph models and probabilistic graph models, multi-dimensional indicator data of audio and video transmission systems are collected in real time, thresholds are dynamically adjusted, and fault propagation paths are identified. This solves the problem of low efficiency in monitoring and diagnosis of audio and video transmission systems in existing technologies, and realizes real-time and accurate fault diagnosis and preventive measures.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-04-14
AI Technical Summary
Existing audio and video transmission system monitoring and diagnostic methods are inefficient, unable to detect faults in real time, and easily affected by human factors. Timed polling or passive alarm mechanisms are insufficient in terms of the timeliness of anomaly detection, false alarm rate control, and fault root cause analysis, and cannot meet the needs of real-time monitoring and rapid response.
A causal graph model is constructed to collect multi-dimensional indicator data of the audio and video transmission system in real time and add semantic tags. Real-time analysis is performed through the causal graph model and the probabilistic graph model to identify abnormal nodes and fault propagation paths. Thresholds are dynamically adjusted to adapt to system changes, thereby achieving end-to-end fault delimitation.
It enables real-time monitoring and accurate diagnosis of audio and video streams, improves troubleshooting efficiency, accurately identifies the root cause of faults and generates targeted preventive inspection instructions, adapts to system changes, and improves the system's practicality and accuracy.
Smart Images

Figure CN121603712B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data monitoring, and in particular to a method and system for real-time monitoring and diagnosis of audio and video streams. Background Technology
[0002] In today's digital age, audio and video transmission systems are widely used in various fields, such as broadcasting, video conferencing, and live streaming. The prevalence of these applications makes the quality and stability of audio and video transmission crucial. The normal operation of audio and video transmission systems not only affects the user's viewing experience but also directly impacts the conduct of related business operations.
[0003] In the monitoring and diagnosis of audio and video transmission systems, the traditional approach is experience-based manual inspection. Technicians periodically check and test each device in the transmission system, checking its operating status and performance indicators, and judging whether there are any faults through observation and manual operation. Another approach relies on timed polling or passive alarm mechanisms to detect the bitstream quality, such as evaluating the bitstream status by periodically checking the key parameters of the TS (Transport Stream).
[0004] However, these existing technologies have significant drawbacks. Experience-based manual inspections are inefficient, unable to detect faults in real time, and susceptible to human error, potentially overlooking some potential faults. Timed polling or passive alarm mechanisms fall short in terms of timeliness of anomaly detection, false alarm rate control, and root cause analysis, failing to meet the demands for real-time monitoring and rapid response.
[0005] Therefore, a method is needed to monitor and diagnose audio and video stream faults in real time and accurately. Summary of the Invention
[0006] This application provides a method and system for real-time monitoring and diagnosis of audio and video streams, which can monitor audio and video streams in real time and accurately diagnose the root cause of abnormalities.
[0007] Firstly, this application provides a method for real-time monitoring and diagnosis of audio and video streams, the method comprising:
[0008] A causal graph model is constructed to represent the causal relationships between physical entities and logical indicators in an audio-visual transmission system. Nodes in the causal graph model represent the physical entity status indicators and audio-visual stream quality indicators in the audio-visual transmission system, and edges represent potential fault propagation paths between nodes. Based on the causal graph model, corresponding multi-dimensional indicator data are collected in real time from encoders, network links, and decoders, and semantic tags are added to the multi-dimensional indicator data. The semantic tags include the device identifier of the data source and the node identifier corresponding to the device identifier in the causal graph model.
[0009] The multi-dimensional indicator data is calculated in real time to obtain the quantitative feature value of each indicator at the current time. The quantitative feature value is compared with the dynamic threshold determined based on the state of the parent node in the causal graph model to determine whether there are any preliminary abnormal nodes.
[0010] When a preliminary abnormal node exists, the state combination between two nodes directly connected by an edge in the causal graph model is detected to see if it violates the preset causal logic, so as to identify the joint abnormal pattern and form an abnormal state set based on the preliminary abnormal node and the joint abnormal pattern.
[0011] The set of abnormal states is input into the causal graph model. Through the probabilistic graph model inference algorithm, the probability of all upstream nodes becoming the root cause nodes of the set of abnormal states is calculated. The node or group of nodes with the highest probability is determined as the root cause of this abnormality.
[0012] By employing the above technical solution and constructing a causal graph model, we not only focus on the data itself but also integrate domain knowledge (physical entities, logical indicators, and fault propagation paths) into the model, laying a theoretical foundation for accurate diagnosis. We utilize dynamic thresholds based on parent node states to detect initial anomalies; furthermore, by verifying the causal logic between nodes, we identify associated anomalies, thereby uncovering more hidden and complex fault patterns. When an anomaly occurs, probabilistic graphical reasoning automatically locates the root cause of the problem. It comprehensively analyzes the entire set of abnormal states and reverse-engineers the initial fault point most likely to trigger all phenomena, greatly improving troubleshooting efficiency. By unifying the model and linking indicators across all stages, we can break down data silos and accurately determine whether the problem originates at the source, network, or receiver, achieving true end-to-end fault delimitation.
[0013] In some embodiments, constructing a causal graph model representing the causal relationships between physical entities and logical indicators in an audio / video transmission system includes:
[0014] Based on prior knowledge in the field of audio and video transmission, the core entities and key performance indicators preset in the system are defined as initial nodes, and based on known fault propagation paths, directed edges are established between nodes with direct causal relationships to form the initial causal graph skeleton.
[0015] Collect multi-dimensional indicator data from historical normal and abnormal periods, calculate the transfer entropy value of each directed edge in the initial causal graph skeleton based on a preset transfer entropy algorithm, quantify the transfer entropy value into an initial weight value, assign it to the corresponding directed edge, and form an initial causal graph model.
[0016] Multi-dimensional indicator data is collected in real time. Based on the multi-dimensional indicator data within the sliding time window, the conditional independence between nodes of the target directed edge is continuously calculated, and the weight of the target directed edge is dynamically adjusted according to the conditional independence. The target directed edge can be any directed edge.
[0017] By employing the aforementioned technical solution, an initial causal graph skeleton is constructed based on domain prior knowledge. This ensures the logical correctness of the model structure and significantly accelerates the convergence speed, avoiding absurd correlations that might occur with purely data-driven methods. By introducing transfer entropy, an information-theoretic metric, the causal relationships between nodes are elevated from a qualitative "existence or non-existence" to a quantitative "strong or weak" relationship. This provides crucial, quantifiable evidence for subsequent root cause probability calculations, making inference more accurate. The model is not static. It dynamically adjusts edge weights by continuously calculating conditional independence, enabling the model to adapt to changes in the system itself and new, unknown fault propagation patterns, ensuring the model's long-term effectiveness and accuracy. This dynamic weighting mechanism allows the model to capture weak causal relationships or indirect influences that were not fully recognized in the early stages of system design or only manifest under specific conditions. This enhances the ability to detect and analyze complex, cascading faults.
[0018] In some embodiments, comparing the quantized feature value with a dynamic threshold determined based on the parent node state in the causal graph model includes:
[0019] Traversing the causal graph model, for each current target node to be detected, identify all directly upstream parent nodes of the target node, obtain the real-time state of all parent nodes of the target node, the real-time state includes discrete state and continuous state, normalize the quantized feature value of the parent node in the continuous state to the tension index, and map the quantized feature value of the parent node in the discrete state to the tension index.
[0020] Based on the tension index of each parent node, the overall tension level of the local context in which the target node is located is calculated using a preset aggregation function. The aggregation function is either a maximum value function or a weighted average function, where the weights correspond to the weights of the edges in the causal graph model.
[0021] Obtain the static threshold baseline of the target node under historical normal operating conditions. Based on the overall stress level, adjust the upper and lower bounds of the static threshold baseline using a preset scaling rule to obtain the dynamic threshold.
[0022] By employing the above technical solution, the "stress level" of the local system is quantified into an overall index through comprehensive analysis of the real-time status of all parent nodes of the target node. This allows the threshold to be dynamically adjusted according to the current system context, better reflecting the actual operating conditions of complex systems. Transforming parent node indices of different natures (continuous and discrete states) into a stress index solves the problem of comprehensively evaluating multi-source heterogeneous data, providing a unified scale and foundation for subsequent aggregation calculations. The calculation of the overall stress level (especially using a weighted average, with weights corresponding to the strength of causal edges) directly quantifies the potential pressure brought about by upstream fault propagation to the current target node. This enables the threshold to not only perceive the state but also predict risks, realizing a shift from state monitoring to stress early warning. The final dynamic threshold is obtained by scaling the overall stress level based on a historical static baseline. This ensures that the threshold is based on normal operating conditions while allowing it to adapt sensitively to changes in system pressure. Furthermore, since the stress originates from a clearly defined parent node, any change in the threshold has good interpretability, facilitating the tracing of causes.
[0023] In some embodiments, the step of adjusting the upper and lower bounds of the static threshold baseline to obtain a dynamic threshold based on the overall tension level using a preset scaling rule specifically includes:
[0024] Based on the static attributes of the target node, a target scaling rule is matched from a preset strategy library. The static attributes include the data type of the indicator and its topological importance in the causal graph model. The target scaling rule is a function that takes the static threshold baseline to be adjusted and the overall tension level as input and the dynamically adjusted threshold boundary as output.
[0025] The static threshold baseline and the overall tension level are input into the target scaling rule to obtain the positive relaxation amount that the original upper bound of the static threshold baseline should be expanded and the negative relaxation amount that the original lower bound should be expanded, so as to obtain the preliminary dynamic threshold boundary.
[0026] The preliminary dynamic threshold boundary is compared with the absolute threshold safety range defined by the system physical limits and service level agreement. A boundary pruning algorithm is then used to prune the portion that exceeds the absolute threshold safety range to obtain the dynamic threshold.
[0027] By employing the above technical solution, different scaling rules are matched based on the static attributes of nodes (data type, topological importance), allowing for more sensitive or conservative adjustment strategies for key nodes or indicators of specific data types, thus achieving differentiated and refined management. The relaxation amount explicitly quantifies the tolerance (positive relaxation) or tightening standard (negative relaxation) required based on system stress, making the adjustment process logically clear, controllable, and predictable. Introducing an absolute threshold safety range for boundary pruning is a crucial step. It ensures that the dynamically adjusted threshold does not violate physical laws (e.g., CPU utilization cannot exceed 100%) or exceed business contracts (e.g., the maximum latency specified in the SLA), thereby preventing technically impossible or business-unacceptable alarms and greatly improving the system's usability.
[0028] In some embodiments, detecting whether the state combination between two nodes directly connected by an edge in the causal graph model violates a preset causal logic in order to identify a joint anomaly pattern specifically includes:
[0029] For each directed edge in the causal graph model, a causal expectation table is predefined. The causal expectation table enumerates the range of normal states that the child nodes of the directed edge should be in when the parent node of the directed edge is in different states.
[0030] The current states of the two nodes of the target directed edge are obtained in real time to form a real-time state combination. The current state of the parent node of the target directed edge is used as the query key to retrieve the expected state range of the corresponding child node in the causal expectation table.
[0031] The expected state range of the child node is compared with the actual current state of the child node of the target directed edge. If the actual current state falls outside the expected state range of the child node, it is determined that the state combination of the node pair of the target directed edge violates the preset causal logic.
[0032] Based on the judgment result, a joint anomaly pattern instance is generated. The joint anomaly pattern instance includes the identifier of the target directed edge, the real-time state combination of the parent and child nodes of the target directed edge, and the range of the expected state that is violated. The joint anomaly pattern instance is then added to the anomaly state set.
[0033] By employing the above technical solution, the normal correspondence between "cause and effect" between nodes is clearly defined through a causal expectation table. This enables the system to detect hidden faults where individual node indicators appear normal, but the combination violates causal logic, greatly improving the depth and accuracy of detection. When a node pair violating causal logic is detected, the system can immediately and accurately pinpoint the specific link in the fault propagation path (directed edge) where the anomaly occurred, providing extremely valuable and direct contextual information for subsequent root cause analysis.
[0034] In some embodiments, calculating the probability that all upstream nodes are the root cause nodes leading to the set of abnormal states using a probabilistic graphical model inference algorithm specifically includes:
[0035] For each first node involved in the set of abnormal states, the state probability distribution of the first node is set as an abnormal certain state in the causal graph model. The abnormal certain state means that the probability value representing the abnormal state is set to a preset maximum value, and the probability value representing the normal state is set to a preset minimum value.
[0036] Each second node in the causal graph model is assigned a corresponding initial state probability value calculated based on historical statistical data. The second node is any node in the causal graph model other than the first node.
[0037] Based on the initial state probability value, according to the direction and weight of the edges in the causal graph model, multiple rounds of iterative probability message passing are performed between nodes. In each round of iteration, the state probability value of the first node remains unchanged, and the state probability value of the second node is updated according to the probability messages received from all parent and child nodes of the second node.
[0038] The iteration terminates when all the state probability values of the second node converge, and the current state probability value of each second node in the causal graph model is taken as the probability.
[0039] By employing the aforementioned technical solution, the detected anomalies (set of anomaly states) are transformed into hard evidence that can be processed within the probabilistic graphical model. Specifically, by setting the state probability distribution of the first node as the anomalous certainty state, the established observational facts (such as "node A is anomalous") serve as the fixed anchor and starting point for inference, providing a reliable factual basis for the entire probabilistic reasoning process. Through multi-round iterative probabilistic message passing, the algorithm is not limited to local analysis but allows the anomalous evidence to propagate throughout the network along the relationships (edge direction and weight) defined in the causal graphical model. This enables the system to assess the global impact of a local anomaly on all relevant upstream nodes, thereby capturing indirect and distant causal relationships. By keeping the probability of anomalous nodes constant while updating the probabilities of other nodes, the probability of all other nodes in the system being "potentially anomalous" given that some nodes are known to be "definitely anomalous" is simulated. This process quantifies the likelihood of each upstream node as a potential root cause.
[0040] In some embodiments, the method further includes:
[0041] The system receives the handling results for the root cause and converts the handling results into a quantified feedback signal, wherein if the abnormal state set is eliminated after handling, a positive feedback signal is generated; if the abnormal state set continues or expands after handling, a negative feedback signal is generated.
[0042] For edges on the causal path from the root cause to nodes in the set of abnormal states, the weight is increased if a positive feedback signal is received, and decreased if a negative feedback signal is received.
[0043] Based on the adjusted causal graph model, the state value of the third node corresponding to the currently monitored root cause is used as the initial input for the simulation. The preset deterioration process is simulated by changing the state value of the third node, and the impact of the state change on the downstream node is calculated using the causal graph model.
[0044] Based on the impact, the downstream target node that first reaches the abnormal threshold is identified, a preventive check instruction is generated for the downstream target node, and an elastic expansion plan is generated for the business resources associated with the downstream target node.
[0045] By employing the above technical solution, and transforming the handling results into feedback signals to adjust the weights of causal edges, the system forms a learning loop. Successful diagnostic experiences are reinforced, and erroneous diagnoses are corrected, enabling the causal graph model to continuously self-calibrate over time and with system changes, becoming increasingly accurate and reliable, and truly possessing adaptive capabilities. Using the validated model for simulation, the simulation process is not merely a general warning, but can accurately calculate the first downstream target node affected. This is equivalent to finding the weakest link in the entire fault propagation chain, allowing preventative measures to be targeted, directly generating preventative inspection instructions and elastic scaling plans for that node, greatly improving the efficiency and accuracy of operational actions.
[0046] In a second aspect, embodiments of this application provide a computer system including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program to implement the steps of the method described in any possible implementation of the first aspect.
[0047] Thirdly, embodiments of this application provide a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implement the steps of the method described in any possible implementation of the first aspect.
[0048] Fourthly, embodiments of this application provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the method described in any possible implementation of the first aspect.
[0049] It is understood that the computer system provided in the second aspect, the storage medium provided in the third aspect, and the computer program product provided in the fourth aspect are all used to execute the method provided in this application. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.
[0050] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0051] 1. By constructing a causal graph model and adding semantic tags containing device and node identifiers to the collected multi-dimensional data, data from isolated links such as encoders, networks, and decoders are uniformly mapped into a global, interconnected knowledge model. This breaks down data silos and enables all subsequent analyses to be based on a context with clear physical and logical meaning, laying the foundation for accurate diagnosis.
[0052] 2. By using the parent node's state to determine the dynamic threshold of the current node, anomaly detection can take into account the current context pressure of the system. When the system load is high, the standard can be relaxed appropriately to reduce false alarms, and the standard can be tightened when the system is calm to avoid missed alarms, making detection more intelligent and accurate.
[0053] 3. All detected anomalies are treated as a whole and input into the causal graph model. Through the probabilistic graph model inference algorithm, the uncertainty of the fault propagation along the causal path can be simulated, and the initial root cause most likely to trigger all observed evidence can be calculated in reverse, which greatly improves the troubleshooting efficiency.
[0054] 4. Because the cause-effect graph model covers all key entities and metrics from encoders and networks to decoders, it can perform comprehensive analysis of the entire transmission link within a framework. When a problem occurs, it can clearly determine whether the root cause is in the encoder at the source, the network link in the middle, or the decoder at the terminal, thus achieving true end-to-end fault localization. Attached Figure Description
[0055] Figure 1 This is a flowchart illustrating a real-time monitoring and diagnosis method for audio and video streams according to an embodiment of this application.
[0056] Figure 2 This is a schematic diagram of the process of constructing a cause-effect graph model in an embodiment of this application;
[0057] Figure 3 This is a schematic diagram of an exemplary hardware structure of a computer system in an embodiment of this application. Detailed Implementation
[0058] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations including one or more of the listed items.
[0059] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0060] The following is combined Figure 1 The method of the embodiments of this application will be described below.
[0061] Figure 1 This is a flowchart illustrating a real-time monitoring and diagnosis method for audio and video streams according to an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following steps:
[0062] S101. Construct a causal graph model to represent the causal relationship between physical entities and logical indicators in the audio and video transmission system. The nodes in the causal graph model represent the physical entity status indicators and audio and video bitstream quality indicators in the audio and video transmission system, and the edges represent potential fault propagation paths between nodes. Based on the causal graph model, collect corresponding multi-dimensional indicator data in real time from the encoder, network link and decoder, and add semantic tags to the multi-dimensional indicator data. The semantic tags include the device identifier of the data source and the node identifier corresponding to the device identifier in the causal graph model.
[0063] S102. Perform real-time calculation on the multi-dimensional indicator data to obtain the quantitative feature value of each indicator at the current moment, and compare the quantitative feature value with the dynamic threshold determined based on the parent node state in the causal graph model to determine whether there are any preliminary abnormal nodes.
[0064] S103. When there is a preliminary abnormal node, detect whether the state combination between two nodes directly connected by an edge in the causal graph model violates the preset causal logic, so as to identify the joint abnormal pattern and form an abnormal state set based on the preliminary abnormal node and the joint abnormal pattern.
[0065] S104. Input the set of abnormal states into the causal graph model, and calculate the probability that all upstream nodes are the root cause nodes of the set of abnormal states through the probabilistic graph model inference algorithm. Then, determine the node or group of nodes with the highest probability as the root cause of this abnormality.
[0066] A cause-effect graph model consists of nodes and edges. Nodes represent key entities and metrics in the system, while edges represent potential fault propagation paths between nodes. Edges generally exist as directed edges, which are directional arrows defining the causal propagation path of a fault or state. The starting point of a directed edge is called the parent node, representing the cause or upstream influencing factor, and the ending point is called the child node, representing the result or downstream affected object. For example, a directed edge from the "network packet loss rate" node to the "video stuttering rate" node indicates that network packet loss may cause video stuttering. Nodes are divided into physical entity status metrics and audio / video stream quality metrics. Physical entity status metrics include: CPU utilization of encoder A, network packet loss rate of gateway B, and memory usage of decoder C. Audio / video stream quality metrics include: video stuttering rate, audio distortion, and end-to-end latency. Various metric data (such as CPU, bandwidth, bitrate, cache status, etc.) are collected in real time from sources such as encoders, network devices, and decoders. Each piece of raw data is labeled with two types of tags: device identifier and node identifier. The device identifier indicates the source of the data (e.g., "Beijing Data Center - Encoder - 01"), and the node identifier indicates the specific node corresponding to this data in the causal graph model (e.g., "Node_12: Network Packet Loss Rate"). The collected multi-dimensional indicators undergo preliminary processing to obtain standardized, current-time quantified feature values (e.g., normalized Z-score, mean / variance over one minute, etc.). For any target node to be detected in the causal graph (which is now a child node), the system traces all directed edges pointing to it to find all parent nodes that directly affect it. For example, if the goal is to detect the target node / child node "video stuttering rate," the system will find its parent nodes, which may include network packet loss rate and decoder cache saturation. The system will comprehensively analyze the current state of these parent nodes to calculate a stress index representing the operational status of that local area. For example, if the parent node's "network latency" and "encoder CPU" are both high, then the threshold for the child node's "video stuttering" should be appropriately relaxed, because slight stuttering at this time may be a normal manifestation of system stress rather than a real failure. Based on this index representing the tension of the local area, the static threshold baseline derived from the historical normal data of the target node is scaled to generate a dynamic threshold specific to the current moment. The current quantized feature value of the target node is compared with this dynamic threshold, and if it exceeds, it is marked as a preliminary abnormal node. Pre-defined causal logic: configured directly for each directed edge. A causal expectation table is defined for the edge "parent node A → child node B". This table specifies that when parent node A is in state X, child node B should be in state Y. For example, for the directed edge "parent node: network packet loss rate → child node: video stuttering rate", its causal expectation table may specify that when the parent node's network packet loss rate > 5%, the expected normal range of the child node's video stuttering rate is [0.1%, 10%]. The system traverses each directed edge and checks the real-time state combination of its parent and child nodes.If the actual state of a child node deviates significantly from the range expected based on the state of its parent node, the node pair is considered to have a "joint anomaly." For example, the parent node may have a high network packet loss rate (10%), but the child node may have an extremely low video stuttering rate (0.01%). This violates the causal logic that "high packet loss usually leads to high stuttering," and the two nodes associated with this edge are jointly marked as anomalies. All initially found anomalous nodes and all joint anomaly patterns are merged to form a comprehensive set of anomalous states, which describes the full picture of the current system anomaly. The state probabilities of the nodes involved in the set of anomalous states are fixed as "anomalous" (i.e., given 100% confidence). The state probabilities of other nodes are initialized based on historical statistics. The algorithm performs probabilistic message passing along the direction (and weight) of the edges in the causal graph in multiple iterations. Simply put, it simulates the process of fault propagation and backtracking along all possible paths. If a node is identified by multiple anomalous child nodes, or by a high-weight anomalous child node, then the probability that it is the root cause increases. The algorithm stops when the probability values of all nodes stabilize (converge). At this point, each node receives a probability value indicating it is the root cause of the anomaly. The system ultimately outputs the node or group of nodes with the highest probability as the root cause of the anomaly.
[0067] Figure 2 This is a schematic diagram of the process of constructing a cause-effect graph model in an embodiment of this application, such as... Figure 2 As shown, constructing a causal graph model representing the causal relationships between physical entities and logical indicators in an audio / video transmission system includes the following steps:
[0068] S201. Based on prior knowledge in the field of audio and video transmission, define the core entities and key performance indicators preset in the system as initial nodes, and based on the known fault propagation path, establish directed edges between nodes with direct causal relationships to form the initial causal graph skeleton.
[0069] S202. Collect multi-dimensional indicator data of historical normal and abnormal periods, calculate the transfer entropy value of each directed edge in the initial causal graph skeleton based on the preset transfer entropy algorithm, quantify the transfer entropy value into an initial weight value, assign it to the corresponding directed edge, and form an initial causal graph model.
[0070] S203. Collect multi-dimensional indicator data in real time, continuously calculate the conditional independence between nodes of the target directed edge based on the multi-dimensional indicator data within the sliding time window, and dynamically adjust the weight of the target directed edge according to the conditional independence. The target directed edge can be any directed edge.
[0071] Core entities: These refer to specific hardware or software components in the system, such as "video encoder A in the Beijing data center," "core switch B," and "end-user decoding App C." Key performance indicators (KPIs): These refer to the measurable state or output quality indicators of these entities, such as "CPU utilization of encoder A," "port packet loss rate of switch B," and "number of video stutters in App C." These entities and KPIs are abstracted as nodes in a causal graph. Based on known fault propagation paths, directed arrows (directed edges) are drawn between nodes with direct causal relationships. The direction of the edge represents the direction of the causal relationship, from cause (parent node) to effect (child node). For example, based on the knowledge that "network congestion leads to video stuttering," a directed edge is established from the parent node "network packet loss rate" to the child node "video stuttering rate." Multi-dimensional indicator data, including data from normal and abnormal system operation periods, are collected to provide sufficient basis for calculations. Transfer entropy is a metric from information theory used to measure the degree of uncertainty reduction of one time series over another. For each directed edge X→Y in the skeleton, the transfer entropy value from X to Y is calculated using historical data. For example, looking only at the historical data of Y (video buffering rate), we find that the future value of Y can be largely predicted by its values from the past few moments (e.g., the value from the past second), which constitutes the basic uncertainty in predicting Y. Now, we consider not only Y's past but also the past values of X (network packet loss rate) from the past few moments, and try to predict Y's future value again. Compare the effects of the two prediction methods: Method A: Predicting the future of Y using only Y's past. Method B: Predicting the future of Y using both Y's past and X's past. Calculate how much more accurate Method B is in predicting Y's future compared to Method A. This "improvement in accuracy" is the transfer entropy to be calculated. High transfer entropy value: indicates that the parent node X has a strong information flow or causal influence on the child node Y. For example, historical data of "network packet loss rate" is very helpful in predicting the future value of "video buffering rate," resulting in a high calculated transfer entropy. Low transfer entropy value: indicates a weak relationship between the two; the causal relationship defined by experts may not be significant in reality, or it may be dominated by other factors. The calculated transition entropy values are normalized and quantified into specific weight values, which are then assigned to the corresponding directed edges. For example, if the transition entropy of one edge is 10 times that of another, its initial weight can be set to 10 times. The system continuously collects real-time indicator data and always analyzes it within a recent sliding time window (e.g., the past 5 minutes) to ensure that the focus is on the current system state. Conditional independence is a core concept in probabilistic graphical models. Here, it is used to test whether two nodes (X and Y) still have a dependency relationship given the states of other related nodes (Z). If, within a time window, the calculation finds that the parent node X and child node Y are still highly correlated under other conditions, then the weight of this directed edge is increased.This indicates that the causal relationship is very active and important in the current context. Conversely, if computation finds that they become conditionally independent (i.e., X cannot provide additional information about Y after knowing other factors), the weight of this edge is weakened. This may mean that, under the current operating mode, the causal relationship has been replaced or masked by other factors. This process is continuously performed on any directed edge in the model, ensuring that the entire model is in a state of continuous fine-tuning and adaptation.
[0072] In some embodiments, comparing the quantized feature value with a dynamic threshold determined based on the parent node state in the causal graph model includes:
[0073] Traversing the causal graph model, for each current target node to be detected, identify all directly upstream parent nodes of the target node, obtain the real-time state of all parent nodes of the target node, the real-time state includes discrete state and continuous state, normalize the quantized feature value of the parent node in the continuous state to the tension index, and map the quantized feature value of the parent node in the discrete state to the tension index.
[0074] Based on the tension index of each parent node, the overall tension level of the local context in which the target node is located is calculated using a preset aggregation function. The aggregation function is either a maximum value function or a weighted average function, where the weights correspond to the weights of the edges in the causal graph model.
[0075] Obtain the static threshold baseline of the target node under historical normal operating conditions. Based on the overall stress level, adjust the upper and lower bounds of the static threshold baseline using a preset scaling rule to obtain the dynamic threshold.
[0076] The system traverses every node in the causal graph model (as the target node). For each target node, the system locates all directly upstream parent nodes (i.e., the starting points of all directed edges pointing directly to the target node). The states of parent nodes of different types and scales are uniformly transformed into a standardized index representing the "stress level," namely the tension index. This index is typically normalized to the range [0, 1], where 0 represents completely normal / no stress, and 1 represents extreme stress / near the limit. For parent nodes with continuous states (such as CPU utilization, network latency), a normalization function is used for transformation. For example, a common practice is to use the cumulative distribution function. For instance, assuming the parent node is network latency, historical data shows that 95% of the time latency is less than 50ms. Therefore, when the real-time latency is 50ms, its tension index might be mapped to 0.95; when the latency is 100ms (an extreme value), the tension index might approach 1.0. This indicates that the higher the latency, the higher the tension. For parent nodes with discrete states (e.g., "service status": running / stopped, "switch": on / off): direct mapping is used for transformation. For example, if a parent node is a load balancer in a healthy state, its state can be either normal or faulty. A preset mapping rule can be used: normal → 0, faulty → 1. If another parent node's encoder master / standby state is either master or standby, it can be mapped to master → 0, standby → 0.5 (because switching to a standby machine inherently implies a certain level of stress). The overall stress level of the target node's local environment is determined by combining the stress of all parent nodes. The maximum value among all parent node stress indices is used as the overall stress level. This is suitable for systems with a significant "bottleneck effect," where a serious problem in any parent node is enough to cause abnormalities in child nodes. This is the most sensitive and conservative strategy. For example, if parent node A (CPU) has a stress level of 0.2 and parent node B (memory) has a stress level of 0.9, then the overall stress level is 0.9. This means the system considers the current environment to be very stressful. Alternatively, a weighted average can be calculated using the weights of the directed edges in the causal graph model as the weights of each parent node's stress index. This approach more precisely reflects the impact of different causes on the results, with parent nodes (higher edge weights) contributing more to the overall stress level. For example, a parent node with a network packet loss rate (edge weight = 0.9) has a stress level of 1.0, while a parent node with an encoder configuration (edge weight = 0.1) has a stress level of 0.5. The overall stress level = (1.0 × 0.9 + 0.5 × 0.1) / (0.9 + 0.1) = 0.95. Obtaining a static threshold baseline: This is a threshold range [lower limit, upper limit] learned from historical normal operating data. When the overall stress level is high, the system allows for a larger fluctuation range in the target node's metrics; therefore, the threshold should be relaxed (i.e., increase the upper limit and decrease the lower limit).Example of scaling rules (linear scaling): New upper limit = static upper limit + (overall tension level × relaxation coefficient), new lower limit = static lower limit - (overall tension level × relaxation coefficient).
[0077] In some embodiments, the step of adjusting the upper and lower bounds of the static threshold baseline to obtain a dynamic threshold based on the overall tension level using a preset scaling rule specifically includes:
[0078] Based on the static attributes of the target node, a target scaling rule is matched from a preset strategy library. The static attributes include the data type of the indicator and its topological importance in the causal graph model. The target scaling rule is a function that takes the static threshold baseline to be adjusted and the overall tension level as input and the dynamically adjusted threshold boundary as output.
[0079] The static threshold baseline and the overall tension level are input into the target scaling rule to obtain the positive relaxation amount that the original upper bound of the static threshold baseline should be expanded and the negative relaxation amount that the original lower bound should be expanded, so as to obtain the preliminary dynamic threshold boundary.
[0080] The preliminary dynamic threshold boundary is compared with the absolute threshold safety range defined by the system physical limits and service level agreement. A boundary pruning algorithm is then used to prune the portion that exceeds the absolute threshold safety range to obtain the dynamic threshold.
[0081] Static attributes include the data type of the metric and its topological importance in the causal graph model. Metric data types include: rate-based (e.g., bitrate, frame rate), utilization-based (e.g., CPU, memory), error-based (e.g., number of stutters, packet loss), and state-based (e.g., latency). Topological importance: the position and criticality of the node in the causal graph. For example, is it a core quality node representing the end-user experience (e.g., video stuttering rate) or an intermediate metric node within the system (e.g., encoder cache size)? The policy library is a pre-defined set of functions, each encapsulating an adjustment logic. For example, for core quality nodes (high topological importance): a "conservative" scaling rule might be matched. Even under system stress, only a small relaxation of the threshold is allowed because it needs to remain highly sensitive to ensure user experience. For intermediate metric nodes: a "relaxed" scaling rule might be matched. The threshold is allowed to adjust more significantly with stress levels to avoid unnecessary internal alarm interference. For rate-based metrics (e.g., output bitrate): an "asymmetric scaling rule" might be matched. When under stress, the upper limit is relaxed (allowing higher bitrate fluctuations), while the criteria for the lower limit (too low bitrate) may remain unchanged or even tighten. Input: Static threshold baseline [Ls, Us] and overall stress level T. Output: Positive relaxation R+ and negative relaxation R-. Positive relaxation: The amount by which the original upper limit should be adjusted upwards under the current stress level. Negative relaxation: The amount by which the original lower limit should be adjusted downwards under the current stress level. Example of a rule function (linear rule): R+ = T × (S+) (S+ is the preset positive relaxation coefficient), R- = T × (S-) (S- is the preset negative relaxation coefficient). The initial dynamic threshold boundaries are calculated as follows: Initial upper limit = Us+(R+); Initial lower limit = Ls-(R-). For example, target node: video stuttering rate. Static threshold baseline [Ls, Us]: [0%, 0.5%]. Overall stress level T: 0.8. Matching rule relaxation coefficients: S+ = 1.0%, S- = 0% (because the stuttering rate should not be negative). R+ = 0.8 × 1.0% = 0.8%; R- = 0.8 × 0% = 0%; Initial upper bound = 0.5% + 0.8% = 1.3%; Initial lower bound = 0% - 0% = 0%; The initial dynamic threshold boundary is obtained as: [0%, 1.3%]. The absolute threshold safety range is defined based on system physical limits and service level agreements (SLAs). System physical limits: Hard boundaries determined by technical principles. For example, the absolute range of CPU utilization is [0%, 100%]; the absolute range of network port speed is [0, 10Gbps]. Service level agreements (SLAs): Business boundaries defined by business contracts or commitments. For example, an SLA signed with a customer stipulates that the end-to-end latency must be less than 100ms. Therefore, its absolute safety upper bound is 100ms. Even under severe system stress, the abnormal threshold cannot be set wider than this, otherwise it would mean violating the SLA.The boundary clipping algorithm is adopted: Final upper bound = min(preliminary upper bound, absolute safety upper bound); Final lower bound = max(preliminary lower bound, absolute safety lower bound). For example, the preliminary dynamic threshold boundary for video stuttering rate is [0%, 1.3%]. Assuming that the business SLA stipulates that a stuttering rate exceeding 1.0% is considered a fault, then its absolute safety upper bound is 1.0%. Final upper bound = min(1.3%, 1.0%) = 1.0%; Final lower bound = max(0%, 0%) = 0%, resulting in the final dynamic threshold used for judgment: [0%, 1.0%.
[0082] In some embodiments, detecting whether the state combination between two nodes directly connected by an edge in the causal graph model violates a preset causal logic in order to identify a joint anomaly pattern specifically includes:
[0083] For each directed edge in the causal graph model, a causal expectation table is predefined. The causal expectation table enumerates the range of normal states that the child nodes of the directed edge should be in when the parent node of the directed edge is in different states.
[0084] The current states of the two nodes of the target directed edge are obtained in real time to form a real-time state combination. The current state of the parent node of the target directed edge is used as the query key to retrieve the expected state range of the corresponding child node in the causal expectation table.
[0085] The expected state range of the child node is compared with the actual current state of the child node of the target directed edge. If the actual current state falls outside the expected state range of the child node, it is determined that the state combination of the node pair of the target directed edge violates the preset causal logic.
[0086] Based on the judgment result, a joint anomaly pattern instance is generated. The joint anomaly pattern instance includes the identifier of the target directed edge, the real-time state combination of the parent and child nodes of the target directed edge, and the range of the expected state that is violated. The joint anomaly pattern instance is then added to the anomaly state set.
[0087] In a causal graph model, each directed edge has its own independent causal expectation table. The query key is the state range of the parent node. This is typically achieved by discretizing the continuous values of the parent node into several meaningful levels (e.g., "low," "medium," "high"), or by directly using its discrete state. The output value is the expected normal state range of the child node. This is based on historical data and expert experience, representing the numerical range that the child node should fall into when the parent node is in a certain state. For example, the directed edge: network packet loss rate → video stuttering rate. The causal expectation table is as follows: Low network packet loss rate (0% < packet loss rate ≤ 0.1%) → video stuttering rate [0%, 0.05%]; Medium network packet loss rate (0.1% < packet loss rate ≤ 1%) → video stuttering rate [0.02%, 0.2%]; High network packet loss rate (packet loss rate > 1%) → video stuttering rate [0.1%, 1.5%]. The causal expectation table stipulates that when the network packet loss rate is in a "high" state, the expected video stuttering rate will increase accordingly. As long as it falls within the range of [0.1%, 1.5%], it is considered a normal performance that conforms to causal logic. The system obtains the current values of the two nodes on the target directed edge in real time. For example, it obtains: parent node - network packet loss rate = 2.5% (belonging to the "high" state), child node - video stuttering rate = 0.8%. Retrieving the expected state range: using the current state of the parent node 2.5% as the query key, in the causal expectation table of the edge network packet loss rate → video stuttering rate, it finds that the state interval of 2.5% is "high", and retrieves the corresponding expected state range of the child node as [0.1%, 1.5%]. The actual current state of the child node is compared with the retrieved expected state range ([0.1%, 1.5%]). Case 1 (consistent with logic): The actual current state falls within the range of [0.1%, 1.5%]. It is judged as normal, indicating that "high packet loss leads to high stuttering", and the system behavior conforms to expectations. Scenario 2 (Violation of Logic): If the actual current state falls outside the expected range, it is judged as violating causal logic. For example, if the actual stuttering rate is 0.01%, it means that the video is almost uninterrupted even with severe network packet loss, which defies common sense. The possible reason is that the stuttering rate statistics module itself is malfunctioning. Or, if the actual stuttering rate is 5.0%, it means that the stuttering level is far beyond what the current network packet loss level can explain. There may be other unknown reasons (such as decoder failure) working together to amplify the stuttering effect. When an edge is judged to violate causal logic, the system will create a structured joint anomaly pattern instance. This instance contains the identifier of the target directed edge, the real-time state combination, and the range of the violated expected state. Identifier of the target directed edge: indicates the causal relationship where the problem occurred (e.g., "edge_123: packet loss rate → stuttering rate"). Real-time state combination: records the specific data when the anomaly occurred (parent node = 2.5%, child node = 0.01%). Range of the violated expected state: records the standard that should have been normal ([0.1%, 1.5%]). Add this instance to the global collection of exception states.
[0088] In some embodiments, calculating the probability that all upstream nodes are the root cause nodes leading to the set of abnormal states using a probabilistic graphical model inference algorithm specifically includes:
[0089] For each first node involved in the set of abnormal states, the state probability distribution of the first node is set as an abnormal certain state in the causal graph model. The abnormal certain state means that the probability value representing the abnormal state is set to a preset maximum value, and the probability value representing the normal state is set to a preset minimum value.
[0090] Each second node in the causal graph model is assigned a corresponding initial state probability value calculated based on historical statistical data. The second node is any node in the causal graph model other than the first node.
[0091] Based on the initial state probability value, according to the direction and weight of the edges in the causal graph model, multiple rounds of iterative probability message passing are performed between nodes. In each round of iteration, the state probability value of the first node remains unchanged, and the state probability value of the second node is updated according to the probability messages received from all parent and child nodes of the second node.
[0092] The iteration terminates when all the state probability values of the second node converge, and the current state probability value of each second node in the causal graph model is taken as the probability.
[0093] The first node refers to those nodes in the set of abnormal states, which have been confirmed as abnormal, such as the video stuttering rate node and the decoder buffer overflow node. In the probabilistic graphical model, each node has a state probability distribution, for example, P(node = normal) = 0.95, P(node = abnormal) = 0.05. The probability distribution of the first node is set to the state of certainty of abnormality, that is: P(node = normal) is set to a preset minimum value (e.g., 0.1), and P(node = abnormal) is set to a preset maximum value (e.g., 0.9). The second node is all other nodes in the causal graphical model besides the first node; their states are unknown and need to be inferred. The initial state probability value of the second node is not arbitrarily set, but calculated based on historical statistical data. It is usually used to determine the percentage of time the node was in a normal or abnormal state in historical data. For example, for a core switch load node, analyzing data from the past month reveals that it was normal 99.5% of the time and abnormal 0.5% of the time. Then its initial probability value can be set as: P(normal) = 0.995, P(abnormal) = 0.005. The message passing mechanism is a core iterative process. In each iteration, each second node performs the following operations: (1) Receive messages. It receives probability messages from all its parent nodes and all its child nodes. This probability message contains the probability influence that the sender node currently believes the receiver node is in a certain state. (2) Integrate information. The node integrates all messages from the parent and child nodes. The direction and weight of the edges play a key role here: Messages from the parent node (forward propagation): When a parent node is calculated to have a high probability of an abnormal state, it will pass a probability evidence supporting that "the child node is in an abnormal state" to the child node according to the conditional probability table defined by the directed edges between them. The strength of this evidence is positively correlated with the weight value of the directed edges. The higher the weight value, the greater the conditional probability influence of the parent node's state on the child node's state, and therefore the stronger the evidence passed. Messages from the child node (backward propagation): The fixed abnormal state of the child node, as observation evidence, will update the posterior probability of each of its parent nodes along the reverse direction of the directed edges. The weight of the directed edge connecting the child node and the parent node determines the distribution ratio of the observed evidence among different parent nodes; the higher the weight, the greater the conditional probability that the parent node is the cause of the child node's abnormal state, and therefore the parent node will be allocated a larger proportion of the abnormal probability update. (3) Update its own probability. Based on all the integrated messages, the node recalculates its own probability of being in a normal or abnormal state using the inference algorithm of the probabilistic graphical model (such as the belief propagation algorithm). If a node is "accused" by multiple abnormal child nodes at the same time, or by a high-weight abnormal child node, then its own probability of being updated as abnormal will increase significantly.Maintaining the first nodes unchanged: Throughout the iteration process, the probability values of the first nodes, which have been solidified into anomaly-certain states, remain constant. They act as sources of evidence, continuously injecting information into the network without being altered themselves. The system continuously monitors the state probability values of all second nodes. After one iteration, if the change in the probability values of all second nodes is less than a preset threshold, the model is considered to have converged, and the iteration terminates. At this point, each second node obtains a converged, stable state probability value P (anomaly). This probability value is calculated after comprehensively considering all observed anomaly evidence, the causal graph topology, and edge weights, indicating the probability that the node is the root cause of the current failure.
[0094] In some embodiments, the method further includes:
[0095] The system receives the handling results for the root cause and converts the handling results into a quantified feedback signal, wherein if the abnormal state set is eliminated after handling, a positive feedback signal is generated; if the abnormal state set continues or expands after handling, a negative feedback signal is generated.
[0096] For edges on the causal path from the root cause to nodes in the set of abnormal states, the weight is increased if a positive feedback signal is received, and decreased if a negative feedback signal is received.
[0097] Based on the adjusted causal graph model, the state value of the third node corresponding to the currently monitored root cause is used as the initial input for the simulation. The preset deterioration process is simulated by changing the state value of the third node, and the impact of the state change on the downstream node is calculated using the causal graph model.
[0098] Based on the impact, the downstream target node that first reaches the abnormal threshold is identified, a preventive check instruction is generated for the downstream target node, and an elastic expansion plan is generated for the business resources associated with the downstream target node.
[0099] Positive Feedback Signal: This signal is generated when maintenance personnel take action based on the root cause identified by the system. If the original set of abnormal states in the system is eliminated and the system performance returns to normal, this confirms that the root cause identification was accurate and the causal path was correct. Negative Feedback Signal: This signal is generated if the original anomaly is not eliminated or even a new anomaly appears after the action is taken. This indicates that the root cause identification may be biased, and the strength or direction of the causal path needs to be corrected. Causal Path Edge Weight Adjustment: This refers to all directed edges along the propagation path from the node identified as the root cause to all nodes affected by it in the set of abnormal states. Receiving positive feedback: Increases the weight of these edges. This means the system has confirmed that these causal relationships are real and the main propagation paths in this failure, and should be given higher confidence in similar scenarios in the future. Receiving negative feedback: Decreases the weight of these edges. This indicates that the causal relationships represented by these edges were not dominant in this scenario, their strength may have been overestimated, or there may be other paths not captured by the model. The third node: This is the node identified as the root cause in the current cycle. For example, if the root cause of this failure is excessive memory utilization of database server A, then this node is the third node, and its current state value (e.g., 85%) is used as the starting point for the simulation. Simulating the deterioration process: The system will preset one or a series of deterioration scenarios, such as: "If the memory utilization of database server A gradually increases from 85% to 95%..." This process is simulated and does not actually occur; it aims to predict the system's behavior under stress. Calculating downstream impact: Using the current (adjusted through feedback) causal graph model, a forward probability or state propagation calculation is performed. The model will calculate, based on the causal relationships and edge weights between nodes, how the memory utilization deteriorates from 85% to 95% and affects its downstream nodes (e.g., "database query latency" → "application server response time" → "frontend API error rate"). Identifying downstream target nodes: In the simulation, the system will identify the downstream node that first reaches its abnormal threshold when the root cause node deteriorates. This node is the weakest link in the entire fault propagation chain and is also the downstream target node that needs priority protection. For example, simulation results show that the "application server response time" will be the first to rise to an unacceptable level. The system will automatically generate an instruction to recommend that operations and maintenance personnel "strengthen the health status check of the 'application server' cluster." This is a low-cost, high-efficiency preventive measure. The system will pre-generate an elastic scaling plan for the business resources associated with this target node. For example, an automated script or policy can be generated with the logic: "When it is detected that the memory utilization of 'database server A' is consistently higher than a% and the 'application server response time' shows an upward trend, automatically trigger the scaling up of the 'application server' cluster by b instances."
[0100] The above describes a method for real-time monitoring and diagnosis of audio and video streams in the embodiments of this application. The computer system in the embodiments of this application will be described in detail below in conjunction with the above-mentioned method for real-time monitoring and diagnosis of audio and video streams.
[0101] Please see Figure 3 This is a schematic diagram of an exemplary hardware structure of a computer system in an embodiment of this application.
[0102] In some embodiments, the computer system 300 includes a computer device, which may be a terminal device. The computer device includes a processor 301, a memory 302, a sensor module 303, a communication module 304, an input device 305, and an output device 306 connected via a system bus. The processor 301 of the computer device provides computing and control capabilities. The memory 302 of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database is used to store data.
[0103] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0104] In some embodiments of this application, a computer-readable storage medium is provided, including instructions that, when executed on the computer system 300, cause the computer system 300 to perform a real-time monitoring and diagnosis method for audio and video streams according to an embodiment of this application.
[0105] In some embodiments of this application, a computer program product is also provided, which, when run on a computer system 300, causes the computer system 300 to execute a real-time monitoring and diagnosis method for audio and video streams according to an embodiment of this application.
[0106] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0107] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0108] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for real-time monitoring and diagnosis of audio and video bitstreams, characterized in that, include: A causal graph model is constructed to represent the causal relationships between physical entities and logical indicators in an audio-visual transmission system. Nodes in the causal graph model represent the physical entity status indicators and audio-visual stream quality indicators in the audio-visual transmission system, and edges represent potential fault propagation paths between nodes. Based on the causal graph model, corresponding multi-dimensional indicator data are collected in real time from encoders, network links, and decoders, and semantic tags are added to the multi-dimensional indicator data. The semantic tags include the device identifier of the data source and the node identifier corresponding to the device identifier in the causal graph model. The multi-dimensional indicator data is calculated in real time to obtain the quantitative feature value of each indicator at the current time. The quantitative feature value is compared with the dynamic threshold determined based on the state of the parent node in the causal graph model to determine whether there are any preliminary abnormal nodes. When a preliminary abnormal node exists, the state combination between two nodes directly connected by an edge in the causal graph model is detected to see if it violates the preset causal logic, so as to identify the joint abnormal pattern and form an abnormal state set based on the preliminary abnormal node and the joint abnormal pattern. The set of abnormal states is input into the causal graph model. Through the probabilistic graph model inference algorithm, the probability of all upstream nodes becoming the root cause nodes of the set of abnormal states is calculated. The node or group of nodes with the highest probability is determined as the root cause of this abnormality.
2. The method for real-time monitoring and diagnosis of audio and video streams according to claim 1, characterized in that, The causal graph model for constructing the causal relationships between physical entities and logical indicators in the audio and video transmission system includes: Based on prior knowledge in the field of audio and video transmission, the core entities and key performance indicators preset in the system are defined as initial nodes, and based on known fault propagation paths, directed edges are established between nodes with direct causal relationships to form the initial causal graph skeleton. Collect multi-dimensional indicator data from historical normal and abnormal periods, calculate the transfer entropy value of each directed edge in the initial causal graph skeleton based on a preset transfer entropy algorithm, quantify the transfer entropy value into an initial weight value, assign it to the corresponding directed edge, and form an initial causal graph model. Multi-dimensional indicator data is collected in real time. Based on the multi-dimensional indicator data within the sliding time window, the conditional independence between nodes of the target directed edge is continuously calculated, and the weight of the target directed edge is dynamically adjusted according to the conditional independence. The target directed edge can be any directed edge.
3. The method for real-time monitoring and diagnosis of audio and video streams according to claim 1, characterized in that, The step of comparing the quantized feature value with a dynamic threshold determined based on the state of the parent node in the causal graph model includes: Traversing the causal graph model, for each current target node to be detected, identify all directly upstream parent nodes of the target node, obtain the real-time state of all parent nodes of the target node, the real-time state includes discrete state and continuous state, normalize the quantized feature value of the parent node in the continuous state to the tension index, and map the quantized feature value of the parent node in the discrete state to the tension index. Based on the tension index of each parent node, the overall tension level of the local context in which the target node is located is calculated using a preset aggregation function. The aggregation function is either a maximum value function or a weighted average function, where the weights correspond to the weights of the edges in the causal graph model. Obtain the static threshold baseline of the target node under historical normal operating conditions. Based on the overall stress level, adjust the upper and lower bounds of the static threshold baseline using a preset scaling rule to obtain the dynamic threshold.
4. The method for real-time monitoring and diagnosis of audio and video streams according to claim 3, characterized in that, The process of adjusting the upper and lower bounds of the static threshold baseline to obtain a dynamic threshold based on the overall tension level and using a preset scaling rule specifically includes: Based on the static attributes of the target node, a target scaling rule is matched from a preset strategy library. The static attributes include the data type of the indicator and its topological importance in the causal graph model. The target scaling rule is a function that takes the static threshold baseline to be adjusted and the overall tension level as input and the dynamically adjusted threshold boundary as output. The static threshold baseline and the overall tension level are input into the target scaling rule to obtain the positive relaxation amount that the original upper bound of the static threshold baseline should be expanded and the negative relaxation amount that the original lower bound should be expanded, so as to obtain the preliminary dynamic threshold boundary. The preliminary dynamic threshold boundary is compared with the absolute threshold safety range defined by the system physical limits and service level agreement. A boundary pruning algorithm is then used to prune the portion that exceeds the absolute threshold safety range to obtain the dynamic threshold.
5. The method for real-time monitoring and diagnosis of audio and video streams according to claim 1, characterized in that, The detection of whether the state combination between two nodes directly connected by an edge in the causal graph model violates a preset causal logic, in order to identify joint abnormal patterns, specifically includes: For each directed edge in the causal graph model, a causal expectation table is predefined. The causal expectation table enumerates the range of normal states that the child nodes of the directed edge should be in when the parent node of the directed edge is in different states. The current states of the two nodes of the target directed edge are obtained in real time to form a real-time state combination. The current state of the parent node of the target directed edge is used as the query key to retrieve the expected state range of the corresponding child node in the causal expectation table. The expected state range of the child node is compared with the actual current state of the child node of the target directed edge. If the actual current state falls outside the expected state range of the child node, it is determined that the state combination of the node pair of the target directed edge violates the preset causal logic. Based on the judgment result, a joint anomaly pattern instance is generated. The joint anomaly pattern instance includes the identifier of the target directed edge, the real-time state combination of the parent and child nodes of the target directed edge, and the range of the expected state that is violated. The joint anomaly pattern instance is then added to the anomaly state set.
6. The method for real-time monitoring and diagnosis of audio and video streams according to claim 1, characterized in that, The step of calculating the probability that all upstream nodes are the root cause nodes leading to the abnormal state set using a probabilistic graphical model inference algorithm specifically includes: For each first node involved in the set of abnormal states, the state probability distribution of the first node is set as an abnormal certain state in the causal graph model. The abnormal certain state means that the probability value representing the abnormal state is set to a preset maximum value, and the probability value representing the normal state is set to a preset minimum value. Each second node in the causal graph model is assigned a corresponding initial state probability value calculated based on historical statistical data. The second node is any node in the causal graph model other than the first node. Based on the initial state probability value, according to the direction and weight of the edges in the causal graph model, multiple rounds of iterative probability message passing are performed between nodes. In each round of iteration, the state probability value of the first node remains unchanged, and the state probability value of the second node is updated according to the probability messages received from all parent and child nodes of the second node. The iteration terminates when all the state probability values of the second node converge, and the current state probability value of each second node in the causal graph model is taken as the probability.
7. The method for real-time monitoring and diagnosis of audio and video streams according to claim 1, characterized in that, The method further includes: The system receives the handling results for the root cause and converts the handling results into a quantified feedback signal, wherein if the abnormal state set is eliminated after handling, a positive feedback signal is generated; if the abnormal state set continues or expands after handling, a negative feedback signal is generated. For edges on the causal path from the root cause to nodes in the set of abnormal states, the weight is increased if a positive feedback signal is received, and decreased if a negative feedback signal is received. Based on the adjusted causal graph model, the state value of the third node corresponding to the currently monitored root cause is used as the initial input for the simulation. The preset deterioration process is simulated by changing the state value of the third node, and the impact of the state change on the downstream node is calculated using the causal graph model. Based on the impact, the downstream target node that first reaches the abnormal threshold is identified, a preventive check instruction is generated for the downstream target node, and an elastic expansion plan is generated for the business resources associated with the downstream target node.
8. A computer system comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1-7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1-7.
Citation Information
Patent Citations
Short video production method and system based on artificial intelligence
CN120547417A
Video stream real-time coding and decoding transmission method under cluster
CN120812312A