Method, apparatus, medium, and program product for monitoring a node

By real-time monitoring and analysis of the status data of edge computing nodes, the problems of excessive granularity and alarm delay in stability monitoring of edge computing nodes are solved, and accurate monitoring and low-latency alarms for a single live stream are achieved, improving the stability of the live stream and user experience.

CN118827328BActive Publication Date: 2025-09-30SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410799154.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-19
Publication Date
2025-09-30
Estimated Expiration
2044-06-19

AI Technical Summary

Technical Problem

In existing technologies, the stability monitoring granularity of edge computing nodes is too large to be specific to the monitoring of a single live stream on the node. The real-time performance is poor, the alarm delay is high, problems cannot be discovered in a timely manner, and there is a lack of automatic processing procedures, which affects the user experience.

Method used

By periodically acquiring the status data of edge computing nodes, down to a single push-pull stream action, we can monitor the transmission problems of live streams in real time, analyze node status data, quickly locate problem nodes, perform low-latency alarm processing, and set up monitoring and alarm configurations for important streams.

Benefits of technology

It achieves real-time and accurate monitoring and alarming of edge computing nodes, reduces alarm delays, ensures the stability of important flows, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118827328B_ABST
    Figure CN118827328B_ABST
Patent Text Reader

Abstract

The present application provides a method, device, medium and program product for monitoring nodes. The method according to the present application includes: in response to detecting that a transmission problem occurs in the live stream of a target node, obtaining the transmission action information of the target node for the live stream; obtaining the status monitoring data of the target node and the previous node of the target node; based on the transmission action information, determining the problem node by analyzing the corresponding status monitoring data of the target node and the previous node of the target node. When a transmission problem occurs, the present application quickly and accurately locates the node where the problem actually occurs by analyzing the aggregated status detection information of the relevant nodes, and obtains the push-pull flow action that causes the transmission problem on the problem node; performs stability analysis on each node based on the status data on the edge computing node obtained in real time, and promptly performs alarm processing on unstable nodes, thereby reducing alarm delay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device, computer-readable medium, and computer program product for monitoring a node. Background Art

[0002] Based on the existing technology solution, in the traditional live broadcast architecture, the anchor uses the broadcast tool to obtain the push stream address from the scheduling system and push the live stream to the cloud edge computing cluster. The edge computing cluster includes upstream streaming, recording, screenshots, transcoding and other services. The various service modules cooperate with each other to form an overall live broadcast link. If a module in the link fails or is unavailable, it will cause the frame rate of the live stream to jitter or be disconnected, resulting in lag or black screen in the live broadcast room, affecting the audience's viewing experience.

[0003] The edge computing cluster includes multiple services, each of which deploys a large number of edge computing nodes. Each node operates normally and cooperates with each other to form a live cloud link. Because live streaming is a long link, if the edge computing node encounters network instability, packet loss, and other problems, it will cause the live streaming service on the node to experience frame loss or even interruption, affecting the user experience. Therefore, a complete edge computing node quality management system is needed to achieve node stability monitoring and alarm. However, there are still some problems in the existing system, mainly including:

[0004] 1) The granularity of stability monitoring for edge computing nodes is too large to monitor a specific push or pull action on a single live stream on the node. Instead of focusing on the stream itself, the monitoring can only focus on the stream but not on a specific action within the stream. Furthermore, the real-time monitoring is poor, as the need to aggregate data results in significant latency. This prevents timely and accurate monitoring of the stability of the stream on the edge computing node and the overall stability of the node.

[0005] 2) The stability alarm delay of edge computing nodes is too high, which cannot help R&D personnel identify and solve problems in a timely and accurate manner. Aggregating all stream data is prone to data loss, resulting in delayed or lost alarms.

[0006] 3) It is unable to focus on important and special flows, and cannot issue alerts and handle data transmission changes of important flows. For example, when the frame rate of a live broadcast stream of a certain event fluctuates, it cannot issue an alert in time, and it is impossible to monitor the stability of important flows and dynamically configure alerts.

[0007] 4) When edge computing nodes have stability issues, there is a lack of automatic processing procedures and excessive reliance on manual processing. Manual processing may cause omissions, causing the problem to continue to expand and affecting the user experience. Summary of the Invention

[0008] Various aspects of the present application provide a method, apparatus, computer-readable medium, and computer program product for monitoring a node.

[0009] In one aspect of the present application, a method for monitoring a node is provided, wherein the method comprises:

[0010] In response to detecting that a transmission problem occurs in the live stream of the target node, obtaining transmission action information of the target node for the live stream;

[0011] Acquire status monitoring data of the target node and the node previous to the target node;

[0012] Based on the transmission action information, the problem node is determined by analyzing the corresponding status monitoring data of the target node and the previous node of the target node.

[0013] In one aspect of the present application, a method for monitoring a node is provided, wherein the method comprises:

[0014] means for obtaining transmission action information of the target node for the live stream in response to detecting a transmission problem of the live stream of the target node;

[0015] means for acquiring status monitoring data of the target node and the node preceding the target node;

[0016] A device for determining a problem node by analyzing the corresponding status monitoring data of the target node and the previous node of the target node based on the transmission action information.

[0017] Another aspect of the present application provides an electronic device, comprising:

[0018] at least one processor; and

[0019] a memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method of the embodiment of the application.

[0021] In another aspect of the present application, a computer-readable storage medium is provided, on which computer program instructions are stored. The computer program instructions can be executed by a processor to implement the method of the embodiment of the application.

[0022] In another aspect of the present application, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the method of the embodiment of the application is implemented.

[0023] The solution provided in the embodiment of the present application realizes real-time monitoring of the push-pull flow actions of the live stream on a single node by periodically acquiring the status data on the edge computing nodes and performing aggregation processing according to the push-pull flow actions. Therefore, when a transmission problem occurs, the node where the problem actually occurs can be quickly and accurately located by analyzing the aggregated status detection information of the relevant nodes, and the push-pull flow actions that cause the transmission problem on the problem node can be known; based on the status data on the edge computing nodes obtained in real time, the stability of each node is analyzed, and the unstable nodes are promptly processed, thereby reducing the alarm delay; by setting the corresponding monitoring configuration information and alarm configuration information for special live streams, the data transmission changes of important streams can be alarmed and processed. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, a brief introduction will be given below to the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0025] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0026] Figure 1 A schematic diagram of a process for monitoring a node provided by an embodiment of the present application is shown;

[0027] Figure 2 A schematic diagram of the structure of a device for monitoring a node provided in an embodiment of the present application is shown;

[0028] Figure 3 A structural diagram of a device suitable for implementing the solution in the embodiments of the present application is shown.

[0029] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION

[0030] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0031] In a typical configuration of the present application, the terminal and the equipment of the service network each include one or more processors (CPUs), input / output interfaces, network interfaces and memories.

[0032] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0033] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer program instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc-read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0034] The following describes the terms involved in the embodiments of the present application.

[0035] Live streaming: A data stream that transmits audio and video content in real time. It is real-time and interactive. The live streaming process is a long link. Unless the host actively interrupts the stream, the live streaming will continue.

[0036] Live Edge Computing: A technology that pushes live stream processing and distribution to the edge of the network. It combines edge computing and live streaming to improve the live streaming experience and reduce the impact of latency and network congestion.

[0037] Edge computing nodes are an important part of the edge computing architecture. They are located at the edge of the network and are used to perform computing tasks, process data, and provide services.

[0038] Frame rate: refers to the number of frames displayed per second in video processing, usually expressed in fps.

[0039] Bit rate: refers to the amount of data transmitted per unit time in digital communication or data transmission, usually expressed in bits per second (bps) or bytes per second (Bps).

[0040] Figure 1The flowchart of a method provided in an embodiment of the present application is shown. The method at least includes step S101, step S102 and step S103.

[0041] In the live broadcast scenario, the method according to this embodiment can be executed by a server for monitoring the service quality of edge computing nodes, which interacts with each edge computing node in the edge computing cluster through a network connection. The server can collect the status indication data reported by the edge computing nodes in real time, and determine whether there is a transmission problem with the live broadcast stream on each edge computing node based on the collected information. The resources of the edge computing nodes can be deployed using containers. Edge computing resources can be a POD composed of one or more containers. For example, it includes a recording POD, a screenshot POD, etc. related to the live broadcast scenario. The POD is the smallest scheduling unit in kubernetes, and the containers in the POD share network and storage resources.

[0042] Moreover, according to the method of the embodiment of the present application, the quality monitoring of the edge computing node is refined to a single push-pull stream action, so as to realize real-time supervision of the push or pull stream action of the live stream on a single node, thereby quickly and accurately locating the node where the problem actually occurs and the push-pull stream action that causes the transmission problem on the problematic node when a transmission problem occurs.

[0043] exist Figure 1 Before the step of, the method further includes step S104, step S105 and step S106.

[0044] In step S104, status indication data on each edge computing node is periodically obtained.

[0045] The status indication data includes various data that can be used to indicate the transmission stability of the live stream of the edge computing node. Optionally, the status indication data includes one or more preset monitoring indicators, including but not limited to frame rate or bit rate.

[0046] In step S105 , for each edge computing node, the acquired status indication data of the edge computing node is aggregated according to a plurality of predetermined transmission action types, and the aggregated data is used as status monitoring data.

[0047] According to one embodiment, the transmission action types may include active push streaming, active pull streaming, passive push streaming, and passive pull streaming. The aggregation processing aggregates the status indication data of each edge computing node according to active push streaming, active pull streaming, passive push streaming, and passive pull streaming to obtain status indication data corresponding to the four transmission action types, respectively, as the status monitoring data of each edge computing node.

[0048] In step S106, the status monitoring data of each edge computing node is stored.

[0049] Reference Figure 1 To illustrate, in step S101, in response to detecting that a transmission problem occurs in a live stream of a target node, transmission action information of the target node for the live stream is obtained.

[0050] The transmission action information is used to indicate the transmission action performed by the target node on the DC. The transmission action may include four types: active push flow, active pull flow, passive push flow, and passive pull flow.

[0051] The transmission problem includes various types of transmission problems, such as frame loss or interruption.

[0052] The method for determining whether a transmission problem occurs in the live stream of the target node includes but is not limited to at least any of the following:

[0053] 1) In response to the fault information reported by the node, determine that a transmission problem occurs in the live stream;

[0054] 2) By analyzing the status monitoring data of each edge computing node obtained in real time, it is determined whether there is a transmission problem with the live stream. Based on this approach, the method further includes step S107.

[0055] In step S107, based on one or more monitoring indicators contained in the status monitoring data of each edge computing node obtained in real time, it is determined according to a predetermined problem judgment rule whether a transmission problem occurs.

[0056] The monitoring indicators are one or more preset indicators for determining whether a live streaming transmission problem occurs. Optionally, the monitoring indicators include but are not limited to frame rate or bit rate.

[0057] Those skilled in the art should be familiar with the fact that a variety of problem judgment rules can be used to determine whether a live stream transmitted on a certain node has a transmission problem. For example, the frame rate or bit rate of the current node obtained in real time is compared with a predetermined threshold value. If the monitoring indicator value is lower than the predetermined threshold value, it is determined that a transmission problem has occurred. For another example, the monitoring indicators of the current node at input and output are compared. If the difference exceeds a predetermined threshold value, it is determined that a transmission problem has occurred. Those skilled in the art can set appropriate problem judgment rules based on actual needs, and the embodiments of the present application do not make specific limitations.

[0058] In step S102, status monitoring data of the target node and the node immediately preceding the target node are obtained.

[0059] The previous node transmits the live stream to the target node.

[0060] In step S103, based on the transmission action information, the problem node is determined by analyzing the corresponding status monitoring data of the target node and the previous node of the target node.

[0061] Specifically, if the live stream is pushed to the target node by the previous node, the problem node is located based on the status monitoring data of the previous node corresponding to active pushing and the status monitoring data of the target node corresponding to passive pushing; if the live stream is pulled from the previous node to itself by the target node, the problem node is located based on the status monitoring data of the previous node corresponding to passive pulling and the status monitoring data of the target node corresponding to active pulling.

[0062] When locating the problem node, similar to the aforementioned step S107, the method can determine whether a transmission problem occurs in the target node and the previous node respectively according to predetermined judgment rules based on one or more monitoring indicators contained in the corresponding status monitoring data of the target node and the previous node of the target node. The specific process will not be repeated here.

[0063] According to one embodiment, the method further includes step S108.

[0064] In step S108, problem analysis information corresponding to the live stream having the transmission problem is generated and stored.

[0065] The problem analysis information includes information indicating the live stream where the problem occurs, the node, and the corresponding transmission action on the node.

[0066] For example, if live stream stream_1 is pushed from node node_A to node node_B, and node_B detects frame drops in the live stream, then step S103 determines that the actual problem node causing the frame drops is node_B. The generated problem analysis information may include: transmission problem type: "frame drops"; live stream stream_1 where the problem occurred; node ID of the problem node: node_B; and the transmission action that caused the problem: "passive push."

[0067] According to one embodiment, the method further includes step S109.

[0068] In step S109 , stability analysis is performed on each edge computing node based on the periodically acquired status monitoring data of each edge computing node.

[0069] The analysis results of the stability analysis include an indication of whether there is an instability problem in the live stream transmitted by the edge computing node.

[0070] Optionally, the stability conditions can be pre-classified and criteria for each level can be set. For example, the DC transmission stability conditions on the node can be classified as "excellent," "qualified," and "unqualified" based on multiple frame rate value ranges. The analysis results of the stability analysis can include information on the transmission stability level of the edge computing node.

[0071] Those skilled in the art should be familiar with the fact that whether the live stream transmitted on a node is stable can be determined based on various indicators that can reflect the quality of live stream transmission, such as frame rate or bit rate.

[0072] Optionally, the analysis results obtained through the stability analysis correspond to the identifiers of the live streams transmitted on the edge computing nodes.

[0073] Optionally, if the difference in bitrate or frame rate of the current edge computing node within a preset time period exceeds a preset threshold, it is determined that the edge computing node has a stability issue with the live stream transmission. Because the status monitoring data corresponds to four transmission action types: active push, active pull, passive push, and passive pull, the stability analysis can further determine the transmission action experiencing stability issues.

[0074] Optionally, you can configure the time interval for stability analysis. For example, if you configure the stability analysis interval to 20 seconds, the frame rate and bit rate of all live streams on the node within 20 seconds will be integrated. The data will be divided according to the node dimension and the push and pull operation of the live stream. This group of live stream data will be analyzed to see if there are any sudden increases or decreases in fps or bps. If so, it is determined that the push and pull operation has an instability issue.

[0075] According to one embodiment, the method further includes step S110.

[0076] In step S110, based on the stability analysis results of each node, if the predetermined alarm processing conditions are met, corresponding alarm processing is performed.

[0077] Specifically, it is first determined whether a predetermined alarm processing condition is met based on the stability analysis result of each node, and alarm processing is performed accordingly when the alarm processing condition is met.

[0078] Optionally, the alarm processing condition includes but is not limited to at least one of the following:

[0079] 1) The number of live streams experiencing instability on the node exceeds a predetermined threshold;

[0080] 2) The ratio of the number of live streams with unstable problems on the node to the number of full streams on the node exceeds a predetermined ratio threshold; for example, the predetermined ratio threshold may be 10%.

[0081] The method may dynamically adjust the quantity threshold or ratio threshold.

[0082] By continuously adjusting the data aggregation interval and alarm threshold, and aggregating live stream stability data in multiple push-pull methods through a single node, low-latency, high-precision node quality, stability monitoring and alarming are achieved.

[0083] According to one embodiment, the method further includes step S111 , step S112 and step S113 .

[0084] In step S111 , it is determined whether the current live stream is a pre-configured target live stream.

[0085] The method uses one or more key live streams as target live streams, records and stores the identification information corresponding to the pre-configured one or more key live streams and their respective monitoring configuration information and alarm configuration information. For example, the live stream corresponding to the live broadcast room of a major event can be set as the key live stream.

[0086] The monitoring configuration information includes a time interval for indicating executing the above steps S104 to S106 to obtain and store the status monitoring data of each node.

[0087] Optionally, the monitoring configuration information may include information indicating monitoring indicators for determining whether the live stream transmitted on the node is stable, so that the method can analyze the stability of the node based on different monitoring indicators for different key live streams.

[0088] The alarm configuration information includes alarm processing conditions, so that the method can perform alarm processing based on different conditions for different key live streams.

[0089] In step S112, for the target live stream, the pre-stored monitoring configuration information and alarm configuration information corresponding to the target live stream are obtained.

[0090] In step S113, based on the acquired stability, status monitoring and alarm processing are performed on the target live stream accordingly.

[0091] For example, based on the settings of relevant personnel, the names of live streams that require important attention can be added or deleted, and the corresponding monitoring indicators and time intervals for data aggregation analysis to obtain status monitoring data can be configured for these live streams. For example, in the live broadcast room of a major event, the live stream is configured to perform data aggregation analysis every 1 second to obtain status monitoring data. The edge computing node pulls data from the central configuration center at regular intervals based on the configured time intervals and monitoring indicators. If a live stream that requires special attention is found on the node itself, data aggregation analysis is performed according to the configured time interval to obtain status monitoring data. Combined with the push-pull action mode of the live stream at the node, if a sharp change in frame rate or bit rate is found, an alarm will be issued and pushed to the peer node of the push-pull action, so as to ensure the key protection of important rooms and facilitate R&D personnel to locate problems in a timely and accurate manner.

[0092] Optionally, the method of this embodiment further includes step S114.

[0093] In step S114, based on the stability analysis results of each node, a scheduling setting is performed on the node whose stability does not meet the predetermined requirements, so that the node no longer receives new live streams.

[0094] For example, if a node experiences instability, reconnection, and other alarms more than a threshold number of times within a certain period of time, the node will be set to unschedulable, so that no new live traffic will be scheduled to it.

[0095] Optionally, the method synchronizes information of unstable nodes to the broadcast side device, which determines whether to continue transmitting the live stream at the node, thereby ensuring that losses can be stopped in time when the node is abnormal to avoid expanding the magnitude of the impact.

[0096] According to the method of the embodiment of the present application, by periodically obtaining the status data on the edge computing nodes and performing aggregation processing according to the push-pull flow actions, real-time monitoring of the push-pull flow actions of the live stream on a single node is achieved, so that when a transmission problem occurs, the node where the problem actually occurs can be quickly and accurately located by analyzing the aggregated status detection information of the relevant nodes, and the push-pull flow actions that cause the transmission problem on the problem node can be known; based on the status data on the edge computing nodes obtained in real time, the stability of each node is analyzed, and alarm processing is performed on unstable nodes in a timely manner, thereby reducing the alarm delay; by setting corresponding monitoring configuration information and alarm configuration information for special live streams, alarms and processing are performed on data transmission changes of important streams.

[0097] In addition, the embodiment of the present application also provides a device for monitoring nodes, the structure of the device is as follows: Figure 2 shown.

[0098] The device includes: a device for obtaining the transmission action information of the target node for the live stream in response to detecting a transmission problem in the live stream of the target node (hereinafter referred to as "action information acquisition device 101"), a device for obtaining status monitoring data of the target node and the previous node of the target node (hereinafter referred to as "monitoring data acquisition device 102"), and a device for determining the problem node by analyzing the corresponding status monitoring data of the target node and the previous node of the target node based on the transmission action information (hereinafter referred to as "problem node device 103").

[0099] The device also includes a state acquisition device, an information processing device and a monitoring data storage device.

[0100] The status acquisition device periodically acquires status indication data on each edge computing node.

[0101] The status indication data includes various data that can be used to indicate the transmission stability of the live stream of the edge computing node. Optionally, the status indication data includes one or more preset monitoring indicators, including but not limited to frame rate or bit rate.

[0102] For each edge computing node, the information processing device aggregates the acquired status indication data of the edge computing node according to a plurality of predetermined transmission action types, and uses the data obtained by the aggregation as status monitoring data.

[0103] According to one embodiment, the transmission action types may include active push streaming, active pull streaming, passive push streaming, and passive pull streaming. The aggregation processing aggregates the status indication data of each edge computing node according to active push streaming, active pull streaming, passive push streaming, and passive pull streaming to obtain status indication data corresponding to the four transmission action types, respectively, as the status monitoring data of each edge computing node.

[0104] The monitoring data storage device stores the status monitoring data of each edge computing node.

[0105] Reference Figure 2 In response to detecting that a transmission problem occurs in the live stream of the target node, the action information acquisition device 101 acquires the transmission action information of the target node for the live stream.

[0106] The transmission action information is used to indicate the transmission action performed by the target node on the DC. The transmission action may include four types: active push flow, active pull flow, passive push flow, and passive pull flow.

[0107] The transmission problem includes various types of transmission problems, such as frame loss or interruption.

[0108] The method for the device to determine whether a transmission problem occurs in the live stream of the target node includes but is not limited to at least any one of the following:

[0109] 1) In response to the fault information reported by the node, determine that a transmission problem occurs in the live stream;

[0110] 2) By analyzing the status monitoring data of each edge computing node obtained in real time, it is determined whether there is a transmission problem in the live stream. Based on this method, the device also includes a problem judgment device.

[0111] The problem judgment device determines whether a transmission problem occurs according to a predetermined problem judgment rule based on one or more monitoring indicators contained in the status monitoring data of each edge computing node obtained in real time.

[0112] The monitoring indicators are one or more preset indicators for determining whether a live streaming transmission problem occurs. Optionally, the monitoring indicators include but are not limited to frame rate or bit rate.

[0113] Those skilled in the art should be familiar with the fact that a variety of problem judgment rules can be used to determine whether a live stream transmitted on a certain node has a transmission problem. For example, the frame rate or bit rate of the current node obtained in real time is compared with a predetermined threshold value. If the monitoring indicator value is lower than the predetermined threshold value, it is determined that a transmission problem has occurred. For another example, the monitoring indicators of the current node at input and output are compared. If the difference exceeds a predetermined threshold value, it is determined that a transmission problem has occurred. Those skilled in the art can set appropriate problem judgment rules based on actual needs, and the embodiments of the present application do not make specific limitations.

[0114] The monitoring data acquisition device 102 acquires the status monitoring data of the target node and the node immediately preceding the target node.

[0115] The previous node transmits the live stream to the target node.

[0116] The problem node device 103 determines the problem node by analyzing the corresponding status monitoring data of the target node and the previous node of the target node based on the transmission action information.

[0117] Specifically, if the live stream is pushed to the target node by the previous node, the problem node device 103 locates the problem node based on the status monitoring data of the previous node corresponding to active push and the status monitoring data of the target node corresponding to passive push; if the live stream is pulled from the previous node to itself by the target node, the problem node device 103 locates the problem node based on the status monitoring data of the previous node corresponding to passive pull and the status monitoring data of the target node corresponding to active pull.

[0118] When locating the problem node, similar to the operation of the aforementioned problem judgment device, the problem node device 103 can determine whether a transmission problem occurs in the target node and the previous node respectively according to predetermined judgment rules based on one or more monitoring indicators contained in the corresponding status monitoring data of the target node and the previous node of the target node. The specific process will not be repeated here.

[0119] According to one embodiment, the apparatus further comprises a result generating device.

[0120] The result generating device generates and stores problem analysis information corresponding to the live stream having the transmission problem.

[0121] The problem analysis information includes information indicating the live stream where the problem occurs, the node, and the corresponding transmission action on the node.

[0122] For example, if live stream stream_1 is pushed from node node_A to node node_B, and node_B detects frame drops in the live stream, then step S103 determines that the actual problem node causing the frame drops is node_B. The generated problem analysis information may include: transmission problem type: "frame drops"; live stream stream_1 where the problem occurred; node ID of the problem node: node_B; and the transmission action that caused the problem: "passive push."

[0123] According to one embodiment, the method also stabilizes the analytical device.

[0124] The stability analysis device performs stability analysis on each edge computing node based on the status monitoring data of each edge computing node obtained periodically.

[0125] The analysis results of the stability analysis include an indication of whether there is an instability problem in the live stream transmitted by the edge computing node.

[0126] Optionally, the device may pre-classify the stability conditions and set criteria for each level. For example, the DC transmission stability conditions at the node may be classified as "excellent," "qualified," or "unqualified" based on multiple frame rate value ranges. The analysis results of the stability analysis may include information on the transmission stability level of the edge computing node.

[0127] Those skilled in the art should be familiar with the fact that whether the live stream transmitted on a node is stable can be determined based on various indicators that can reflect the quality of live stream transmission, such as frame rate or bit rate.

[0128] Optionally, the analysis results obtained through the stability analysis correspond to the identifiers of the live streams transmitted on the edge computing nodes.

[0129] Optionally, if the difference in bitrate or frame rate of the current edge computing node during a preset time period exceeds a preset threshold, the stability analysis device determines that the edge computing node has a stability issue with the live stream being transmitted. Because the status monitoring data corresponds to four transmission action types: active push, active pull, passive push, and passive pull, the stability analysis can further identify the transmission action experiencing a stability issue.

[0130] Optionally, the device can configure a time interval for performing stability analysis. For example, the time interval for performing stability analysis is configured to be 20 seconds. Based on this time interval, the frame rate and bit rate of all live streams on the node within 20 seconds are integrated, divided according to the node dimension and the push and pull actions of the live stream, and the live stream data is analyzed to see whether there are sudden increases or decreases in fps, bit rates, etc. If so, it is determined that the push and pull action has an instability issue.

[0131] According to one embodiment, the method further alerts the processing device.

[0132] Based on the stability analysis results of each node, if the predetermined alarm processing conditions are met, the alarm processing device performs corresponding alarm processing.

[0133] Specifically, the alarm processing device first determines whether a predetermined alarm processing condition is met based on the stability analysis results of each node, and performs alarm processing accordingly when the alarm processing condition is met.

[0134] Optionally, the alarm processing condition includes but is not limited to at least one of the following:

[0135] 1) The number of live streams experiencing instability on the node exceeds a predetermined threshold;

[0136] 2) The ratio of the number of live streams with unstable problems on the node to the number of full streams on the node exceeds a predetermined ratio threshold; for example, the predetermined ratio threshold may be 10%.

[0137] The device may dynamically adjust the quantity threshold or ratio threshold.

[0138] By continuously adjusting the data aggregation interval and alarm threshold, and aggregating live stream stability data in multiple push-pull methods through a single node, low-latency, high-precision node quality, stability monitoring and alarming are achieved.

[0139] According to one embodiment, the apparatus further includes a target flow determining means, a configuration acquiring means, and a configuration processing means.

[0140] The target stream determining device determines whether the current live stream is a preconfigured target live stream.

[0141] The method uses one or more key live streams as target live streams, records and stores the identification information corresponding to the pre-configured one or more key live streams and their respective monitoring configuration information and alarm configuration information. For example, the live stream corresponding to the live broadcast room of a major event can be set as the key live stream.

[0142] The monitoring configuration information includes a time interval for indicating executing the above steps S104 to S106 to obtain and store the status monitoring data of each node.

[0143] Optionally, the monitoring configuration information may include information indicating monitoring indicators for determining whether the live stream transmitted on the node is stable, so that the method can analyze the stability of the node based on different monitoring indicators for different key live streams.

[0144] The alarm configuration information includes alarm processing conditions, so that the method can perform alarm processing based on different conditions for different key live streams.

[0145] The configuration acquisition device acquires pre-stored monitoring configuration information and alarm configuration information corresponding to the target live stream.

[0146] The configuration processing device performs status monitoring and alarm processing on the target live stream based on the acquired stability situation.

[0147] For example, based on the settings of relevant personnel, the names of live streams that require important attention can be added or deleted, and the corresponding monitoring indicators and time intervals for data aggregation and analysis to obtain status monitoring data can be configured for these live streams. For example, in the live broadcast room of a major event, the live stream is configured to perform data aggregation and analysis every 1 second to obtain status monitoring data. Based on the configured time intervals and monitoring indicators, the edge computing node regularly pulls data from the central configuration center. If a live stream that requires special attention is found on the node itself, data aggregation and analysis are performed according to the configured time interval to obtain status monitoring data. Combined with the push-pull action mode of the live stream on the node, if a sharp change in frame rate or bit rate is found, an alarm will be issued and pushed to the peer node of the push-pull action, thereby ensuring the key protection of important rooms and facilitating R&D personnel to locate problems in a timely and accurate manner.

[0148] Optionally, the method of this embodiment further includes a node scheduling device.

[0149] The node scheduling device performs scheduling settings on the nodes whose stability does not meet the predetermined requirements based on the stability analysis results of each node, so that the node no longer receives new live streams.

[0150] For example, if a node has instability, reconnection, and other alarms greater than a threshold number within a certain time period, the node scheduling device will set the node as unschedulable, so that no new live traffic will be scheduled.

[0151] Optionally, the node scheduling device synchronizes the information of the unstable node to the broadcast side device, and the broadcast side device determines whether to continue transmitting the live stream at the node, thereby ensuring that the loss can be stopped in time when the node is abnormal to avoid expanding the magnitude of the impact.

[0152] According to the device of the embodiment of the present application, by periodically obtaining the status data on the edge computing nodes and performing aggregation processing according to the push-pull flow actions, real-time monitoring of the push-pull flow actions of the live stream on a single node is achieved. Therefore, when a transmission problem occurs, the node where the problem actually occurs is quickly and accurately located by analyzing the aggregated status detection information of the relevant nodes, and the push-pull flow actions that cause the transmission problem on the problem node are known; based on the status data on the edge computing nodes obtained in real time, the stability of each node is analyzed, and alarm processing is performed on unstable nodes in a timely manner, thereby reducing the alarm delay; by setting corresponding monitoring configuration information and alarm configuration information for special live streams, alarms and processing are performed on data transmission changes of important streams.

[0153] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present application. The method corresponding to the electronic device may be the method for monitoring a node in the aforementioned embodiment, and its principle of solving the problem is similar to that of the method. The electronic device provided in an embodiment of the present application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the methods and / or technical solutions of the aforementioned multiple embodiments of the present application.

[0154] The electronic device may be a user device, or a device formed by integrating a user device and a network device via a network, or an application running on the above device. The user device includes but is not limited to various terminal devices such as computers, mobile phones, tablets, smart watches, and bracelets. The network device includes but is not limited to network hosts, single network servers, multiple network server sets, or cloud computing-based computer collections, and can be used to implement some of the processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing (Cloud Computing), where cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computers.

[0155] Figure 3The structure of a device suitable for implementing the method and / or technical solution in the embodiment of the present application is shown. The device 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage part 1208 into the random access memory (RAM) 1203. Various programs and data required for system operation are also stored in RAM 1203. CPU 1201, ROM 1202 and RAM 1203 are connected to each other through a bus 1204. Input / output (I / O) interface 1205 is also connected to bus 1204.

[0156] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, a touch screen, a microphone, an infrared sensor, and the like; an output section 1207 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), an LED display, an OLED display, and a speaker; a storage section 1208 including one or more computer-readable media such as a hard disk, an optical disk, a magnetic disk, and a semiconductor memory; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication section 1209 performs communication processing via a network such as the Internet.

[0157] In particular, the methods and / or embodiments of the present application can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 1201, the above-mentioned functions defined in the method of the present application are performed.

[0158] Another embodiment of the present application further provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of the present application.

[0159] Specifically, the present embodiment can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.

[0160] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0161] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0162] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0163] The flow chart or block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the equipment, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code include one or more executable instructions for realizing the logical function of the specification. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs the function or operation of the specification, or can be implemented with a combination of dedicated hardware and computer instructions.

[0164] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0165] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or page components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0166] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0167] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0168] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute some steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

[0170] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device through software or hardware. Terms such as "first" and "second" are used to indicate names and do not imply any particular order.

Claims

1. A method for monitoring a node, wherein: The method comprises: In response to detecting that a transmission problem occurs in the live stream of the target node, obtaining transmission action information of the target node for the live stream; Acquire status monitoring data of the target node and the node previous to the target node; Based on the transmission action information, determining the problem node by analyzing the corresponding status monitoring data of the target node and the previous node of the target node; The determining of the problem node by analyzing the corresponding status monitoring data of the target node and the previous node of the target node based on the transmission action information includes: If the live stream is pushed to the target node by the previous node, the problem node is located based on the status monitoring data of the previous node corresponding to active push and the status monitoring data of the target node corresponding to passive push; if the live stream is pulled from the previous node to itself by the target node, the problem node is located based on the status monitoring data of the previous node corresponding to passive pull and the status monitoring data of the target node corresponding to active pull.

2. The method according to claim 1, wherein The method further comprises: Periodically obtain status indication data on each edge computing node; For each edge computing node, the acquired status indication data of the edge computing node is aggregated according to a plurality of predetermined transmission action types, and the aggregated data is used as the status monitoring data; Stores status monitoring data of each edge computing node.

3. The method according to claim 1 or 2, wherein: The method further comprises: Based on one or more monitoring indicators contained in the status monitoring data of each edge computing node obtained in real time, it is determined whether a transmission problem occurs according to predetermined problem judgment rules.

4. The method according to claim 1 or 2, wherein: The method further comprises: Based on the status monitoring data of each edge computing node obtained periodically, the stability of each edge computing node is analyzed.

5. The method according to claim 4, wherein The method further comprises: Based on the stability analysis results of each node, if the predetermined alarm processing conditions are met, corresponding alarm processing will be performed.

6. The method according to claim 4, wherein: The method further comprises: Based on the stability analysis results of each node, the node whose stability does not meet the predetermined requirements is scheduled and set so that the node no longer receives new live streams.

7. The method according to claim 1 or 2, wherein: The method further comprises: Determine whether the current live stream is a pre-configured target live stream; For the target live stream, obtain the pre-stored monitoring configuration information and alarm configuration information corresponding to the target live stream; Based on the acquired stability, the target live stream is monitored and alarms are processed accordingly.

8. A method for monitoring a node, wherein: The method comprises: means for obtaining transmission action information of the target node for the live stream in response to detecting a transmission problem of the live stream of the target node; means for acquiring status monitoring data of the target node and the node preceding the target node; means for determining a problem node by analyzing corresponding status monitoring data of the target node and a node preceding the target node based on the transmission action information; The device for determining the problem node by analyzing the corresponding status monitoring data of the target node and the previous node of the target node based on the transmission action information is used to: If the live stream is pushed to the target node by the previous node, the problem node is located based on the status monitoring data of the previous node corresponding to active push and the status monitoring data of the target node corresponding to passive push; if the live stream is pulled from the previous node to itself by the target node, the problem node is located based on the status monitoring data of the previous node corresponding to passive pull and the status monitoring data of the target node corresponding to active pull.

9. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.