Fault rapid identification and repair method and device based on streaming media service, and medium

By automatically analyzing faults in the streaming media service cluster through a health monitoring system and a fault information analysis system, and combining function call chains and code change records, code defects can be quickly identified and repaired, solving the problem of faulty nodes in large-scale streaming media service clusters and improving the efficiency of fault analysis and repair.

CN120416082APending Publication Date: 2025-08-01RINGSLINK XIAMEN NETWORK COMM TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510426100.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In large-scale real-time audio and video streaming service clusters, due to the complexity of the logic and frequent version changes, code defects are easily introduced, leading to faulty nodes and affecting user experience. Existing technologies rely on manual collection and analysis of fault information, which is inefficient.

Method used

By introducing a health monitoring system, a cluster health management system, and a fault information analysis system, the system automatically collects the operating status and fault information of streaming media service nodes. Combined with function call chains and code change records, it automatically analyzes code defects through a fault probability algorithm to guide developers to quickly fix them.

Benefits of technology

It enables rapid identification and repair of code defects, improves fault analysis efficiency, ensures service availability and user experience, and reduces resource waste and misoperation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416082A_ABST
    Figure CN120416082A_ABST
Patent Text Reader

Abstract

The invention discloses a rapid fault identification and repair method and device based on streaming media service, and a medium. The method comprises the following steps: detecting the operation state of a streaming media service node in a streaming media service cluster; when the streaming media service node has a fault, reporting a fault event of the target service node to a server node of the cluster health management system; the server node reports the fault event to a target agent node of a cluster health management system where the target service node is located; the target agent node collects service process stack information and service log information where the target service node is located, and uploads the information to a fault information analysis system; the fault information analysis system pulls a code change record according to the information reported by the target agent node; the fault information analysis system automatically analyzes code defects in combination with service process stack information, service log information and code change records; and a developer locates and repairs the code defect according to the defect analysis result. According to the invention, the availability of the streaming media service can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of streaming media services, and particularly to a method, device, and medium for quickly identifying and repairing faults based on streaming media services. Background Art

[0002] In a large-scale real-time audio and video streaming media service cluster, due to the interaction between the signaling module and the real-time audio and video transmission module, the logic is complex and the function iteration is frequent, which easily leads to code defects being introduced during the version change process. Specifically, the business logic operation does not meet the expectations, resulting in the problem that one or several cluster nodes in the large-scale real-time audio and video streaming media service cluster always have faults. As a result, the real-time audio and video sessions scheduled to these nodes cannot work as expected, which has a great impact on the user experience. Currently, in the industry, in the face of these problems, most still manually collect relevant information during faults and then analyze it manually, with relatively low efficiency. Therefore, ensuring that code defects can be quickly identified and repaired under a large-scale real-time audio and video streaming media service cluster is an important task for building the overall service availability.

[0003] The Chinese invention patent with the application number CN202210195665.7 discloses a method, device, equipment, and storage medium for fault detection of a streaming media system. The patent includes: obtaining the system metric values of the streaming media system at the current moment; clustering the system metric values at the current moment and the system metric values at the previous N moments, where N is a positive integer; if the system metric values at the current moment belong to the minority class, it is determined that the streaming media system has a fault. In this way, it is possible to quickly determine whether the streaming media system has a fault by clustering the system metric values of the streaming media system at the current moment and the system metric values at the previous N moments, improving the fault detection efficiency of the streaming media system. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to propose a method for quickly identifying and repairing faults based on streaming media services, which can greatly improve the availability of large-scale streaming media services.

[0005] To achieve the above technical objectives, the technical solution adopted by the present invention is as follows:

[0006] The present invention provides a method for quickly identifying and repairing faults based on streaming media services, which is applied to a health monitoring system, a streaming media service cluster, a cluster health management system, and a fault information analysis system, and includes the following steps:

[0007] Step 1: Use the health monitoring system to detect the running status of each streaming media service node in the streaming media service cluster in real time;

[0008] Step 2: When the health monitoring system detects a failure of a certain streaming media service node, the faulty streaming media service node is taken as the target service node, and the failure event of this target service node is reported to the server node of the cluster health management system;

[0009] Step 3: The server node of the cluster health management system reports the failure event of this target service node to the target proxy node of the cluster health management system where the target service node is located;

[0010] Step 4: The target proxy node of the cluster health management system collects the service process stack information and business log information of the target service node, and packs and uploads them to the failure information analysis system;

[0011] Step 5: The failure information analysis system pulls the records of the most recent multiple code changes according to the service process stack information reported by the target proxy node;

[0012] Step 6: The failure information analysis system combines the service process stack information, business log information, and code change records at the time of failure to automatically analyze the code defects and obtain the defect analysis result;

[0013] Step 7: The developer locates and fixes the code defects according to the defect analysis result.

[0014] Further, Step 1 is specifically: The health monitoring system detects through the rtsp protocol of the streaming media service, and monitors the running status of each streaming media service node in the streaming media service cluster in real time.

[0015] Further, Step 3 specifically includes:

[0016] Step 31: Deploy multiple streaming media service nodes and multiple proxy nodes of the cluster health management system in the same streaming media service cluster, and the streaming media service nodes and proxy nodes correspond one by one;

[0017] Step 32: Establish a communication connection between the server node of the cluster health management system and multiple proxy nodes;

[0018] Step 33: After the server node of the cluster health management system receives the failure event of the target service node reported for the first time, it enters Step 34;

[0019] Step 34: The server node of the cluster health management system selects the proxy node of the cluster health management system corresponding to this target service node as the target proxy node according to the target service node;

[0020] Step 35: The server node of the cluster health management system reports the failure event of this target service node to the target proxy node.

[0021] Further, after step 35, it further includes: recording the time of the first reported fault, observing whether a second reported fault occurs within a set time interval. If so, it indicates that there is a network anomaly between the health monitoring system and the streaming service cluster, and it is determined that the second to the nth reported faults occurring within the set time interval are false alarms, and the fault events reported from the second to the nth time are ignored; where n represents an integer and n>2.

[0022] Further, after step 4, it further includes: the target proxy node of the cluster health management system sends a restart instruction to the service process of the streaming service cluster. After receiving the restart instruction, the service process performs a restart of the service process to resume the service.

[0023] Further, step 5 specifically includes:

[0024] Step 51: After the fault information analysis system receives the service process stack information and business log information reported by the target proxy node, it obtains the function call stack of the critical business thread from the service process stack information and updates the occurrence times of the historical function call chain in the fault information analysis system in real time;

[0025] Step 52: The fault information analysis system pulls the code change records of the service in the last K times in real time. The value range of K is set by the user according to requirements, and the changed functions in the code change records are extracted.

[0026] Further, step 6 specifically includes:

[0027] Step 61: The fault information analysis system calculates the probability of each function call chain causing a fault according to the occurrence times of the function call chain and the changed functions in the code change records; the calculation formula for the probability Q of each function call chain causing a fault is: Q = M * N, where M represents the occurrence times of the function call chain, and N represents whether the code has changed; if the function call chain does not contain a changed function, it means that the code has not changed, and N = 1 is taken. If the function call chain contains a changed function, it means that the code has changed, and N>1 is taken;

[0028] Step 62: Sort the probabilities of each function call chain causing a fault calculated from large to small to obtain a probability sorting diagram;

[0029] Step 63: Download the corresponding business log information from the fault information analysis system to the local;

[0030] Step 64: The fault information analysis system performs automatic analysis of code defects according to the probability sorting diagram and the downloaded business log information, and performs analysis in the order of the probability magnitudes in the probability sorting diagram.

[0031] Further, step 7 is specifically as follows: Developers locate the function where the problem code is located according to the probability sorting diagram and the downloaded log information, and fix the code defect.

[0032] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for fast identification and repair of faults based on streaming media services as described above.

[0033] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for fast identification and repair of faults based on streaming media services as described above.

[0034] Adopting the above technical solutions, compared with the prior art, the present invention has the following beneficial effects:

[0035] By introducing a health monitoring system, a streaming media service cluster, a cluster health management system, and a fault information analysis system, the present invention collects the function call chain information of each thread of the service at the first time when a fault occurs in the service. This information is the call mirror point where the service process fails, that is, the service fails when it runs to this mirror point. This information has important reference value for analyzing faults. Then, by automatically collecting the function call chain information at the moment of the fault, combining the call chains before and after and the code change situation, and applying a set of fault probability algorithms, the code defect point most likely to cause the fault is calculated, shortening the troubleshooting range of the defective code, and then being able to guide the developers to repair the defects, helping to realize the process of fast code defect identification and repair, and ensuring the continuous improvement of the code defect analysis efficiency. Description of the Drawings

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0037] Figure 1 It is a flowchart of a method for fast identification and repair of faults based on streaming media services provided by an embodiment of the present invention.

[0038] Figure 2 It is a schematic diagram of an electronic device provided by an embodiment of the present invention.

[0039] Figure 3 It is a schematic diagram of a computer-readable storage medium provided by an embodiment of the present invention. Detailed implementation manners

[0040] The following combines the accompanying drawings and embodiments to make a further detailed description of the present invention. It should be specifically pointed out that the following embodiments are only used to illustrate the present invention, but do not limit the scope of the present invention. Similarly, the following embodiments are only partial embodiments of the present invention rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0041] Please refer to Figure 1 , a method for rapid fault identification and repair based on streaming media services of the present invention is applied to a health monitoring system, a streaming media service cluster, a cluster health management system, and a fault information analysis system, and includes the following steps:

[0042] Step 1: The health monitoring system detects the running status of each streaming media service node in the streaming media service cluster in real time;

[0043] In this embodiment, step 1 is specifically: the health monitoring system detects through the rtsp protocol of the streaming media service and monitors the running status of each streaming media service node in the streaming media service cluster in real time. This step is to monitor the running status of each streaming media service node in the streaming media service cluster in real time through the rtsp protocol of the streaming media service, ensure that when a streaming media service node fails, it can be discovered and notified to the cluster health management system in the first time, provide instant data support for subsequent fault handling and repair, and avoid the spread of faults affecting the user experience.

[0044] Step 2: When the health monitoring system detects that a certain streaming media service node fails, the failed streaming media service node is used as the target service node, and the fault event of the target service node is reported to the server node of the cluster health management system;

[0045] Step 3: The server node of the cluster health management system reports the fault event of the target service node to the target proxy node of the cluster health management system where the target service node is located;

[0046] In this embodiment, step 3 specifically includes:

[0047] Step 31: Deploy multiple streaming media service nodes and multiple proxy nodes of the cluster health management system in the same streaming media service cluster, and the streaming media service nodes and proxy nodes correspond one by one;

[0048] Step 32: Establish a communication connection between the server node of the cluster health management system and multiple proxy nodes;

[0049] Step 33: After the server node of the cluster health management system receives the fault event of the target service node reported for the first time, it proceeds to Step 34;

[0050] Step 34: The server node of the cluster health management system selects the proxy node of the cluster health management system corresponding to the target service node as the target proxy node according to the target service node;

[0051] Step 35: The server node of the cluster health management system reports the fault event of the target service node to the target proxy node.

[0052] In this embodiment, after Step 35, it further includes: recording the time of the first reported fault, observing whether a second reported fault occurs within a set time interval. If so, it indicates that there is a network anomaly between the health monitoring system and the streaming media service cluster, and it is determined that the second to the nth reported faults occurring within the set time interval are false alarms, and the fault events reported from the second to the nth time are ignored; where n represents an integer and n > 2. The built-in rule engine is used to confirm the faults reported from the second time onwards, eliminate false alarms or transient problems, exclude false alarm situations, ensure the authenticity of the faults, improve the accuracy of fault identification, and avoid unnecessary resource waste and misoperations.

[0053] Step 4: The target proxy node of the cluster health management system collects the service process stack information and business log information of the target service node, and packages and uploads them to the fault information analysis system; ensuring the complete preservation of fault information and providing a basis for subsequent analysis;

[0054] In this embodiment, after Step 4, it further includes: the target proxy node of the cluster health management system sends a restart instruction to the service process of the streaming media service cluster. After receiving this restart instruction, the service process performs a restart of the service process to restore the service. The self-healing process quickly restores the service, ensuring that it returns to a healthy state in the first place and reducing the impact on users.

[0055] Step 5: The fault information analysis system pulls the recent multiple code change records according to the service process stack information reported by the target proxy node;

[0056] In this embodiment, Step 5 specifically includes:

[0057] Step 51. After the fault information analysis system receives the service process stack information and business log information reported by the target proxy node, it obtains the function call stack of the key business threads from the service process stack information and updates the occurrence times of the historical function call chains in the fault information analysis system in real time; the function call stack is a data structure during program operation, recording all active function calls in the current execution thread (i.e., the hierarchical relationship of function nested calls). When a fault occurs in the streaming media service node, the proxy node of the cluster health management system will capture the service process stack information (i.e., the function call stack) at the moment of the fault. By analyzing the function names and call sequences in the function call stack, the specific location where the program is executed when the fault occurs can be located.

[0058] Step 52. The fault information analysis system pulls the code change records of the service in the most recent K times in real time. The value range of K is set by the user according to requirements. For example, K takes 3 - 5 times; and the changed functions in the code change records are extracted. The changed function refers to the code modification history in the code change record, including newly added, deleted, or modified function codes.

[0059] Step 6. The fault information analysis system combines the service process stack information, business log information, and code change records at the time of the fault to automatically analyze code defects and obtain defect analysis results;

[0060] In this embodiment, the specific steps of Step 6 include:

[0061] Step 61. The fault information analysis system calculates the probability of each function call chain causing a fault based on the occurrence times of the function call chain and the changed functions in the code change record; the calculation formula for the probability Q of each function call chain causing a fault is: Q = M * N, where M represents the occurrence times of the function call chain, and N represents whether the code has changed; if the function call chain does not contain a changed function, it means the code has not changed, and N = 1 is taken. If the function call chain contains a changed function, it means the code has changed, and N > 1 is taken. For example, if 1 < N ≤ 3, then the probability value M * N will increase significantly and is marked as a high - risk defect point; through automated analysis using the fault probability algorithm, the defect location time can be significantly shortened and the repair efficiency can be improved.

[0062] Step 62. Sort the probabilities of each function call chain causing a fault calculated from large to small to obtain a probability sorting graph;

[0063] Step 63. Download the corresponding business log information to the local from the fault information analysis system;

[0064] Step 64: The fault information analysis system automatically analyzes code defects based on the probability sorting graph and the downloaded service log information. The analysis is carried out in the order of the probabilities in the probability sorting graph, ensuring that the function call chains with higher probabilities of causing faults can be processed first. Developers can start troubleshooting from the places with higher probabilities, which helps to quickly locate problems.

[0065] Step 7: Developers locate and fix code defects according to the defect analysis results.

[0066] In this embodiment, Step 7 is specifically as follows: Developers locate the function (i.e., a certain fragment of the code) where the problem code is located based on the probability sorting graph and the downloaded log information, and fix the code defects. Combining the automated analysis results and manual judgment ensures the accuracy and efficiency of defect repair.

[0067] As Figure 2 shown, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned method for quickly identifying and repairing faults based on streaming media services is implemented.

[0068] As Figure 3 shown, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-mentioned method for quickly identifying and repairing faults based on streaming media services is implemented.

[0069] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0070] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0071] The above are only some embodiments of the present invention, and thus do not limit the protection scope of the present invention. Any equivalent device or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for fast fault identification and repair based on streaming media services, characterized in that Applied to a health monitoring system, a streaming media service cluster, a cluster health management system, and a fault information analysis system, it includes the following steps: Step 1: The health monitoring system detects the running status of each streaming media service node in the streaming media service cluster in real time; Step 2: When the health monitoring system detects a fault in a certain streaming media service node, the faulty streaming media service node is used as the target service node, and the fault event of this target service node is reported to the server node of the cluster health management system; Step 3: The server node of the cluster health management system reports the fault event of this target service node to the target proxy node of the cluster health management system where the target service node is located; Step 4: The target proxy node of the cluster health management system collects the service process stack information and business log information of the service process where the target service node is located, and packs and uploads them to the fault information analysis system; Step 5: The fault information analysis system pulls the records of the most recent multiple code changes according to the service process stack information reported by the target proxy node; Step 6: The fault information analysis system automatically analyzes the code defects by combining the service process stack information, business log information, and code change records at the time of the fault to obtain the defect analysis result; Step 7: Developers locate and fix the code defects according to the defect analysis result.

2. The method for rapid fault identification and repair based on streaming media services according to claim 1, wherein The specific content of Step 1 is: The health monitoring system detects through the rtsp protocol of the streaming media service to monitor the running status of each streaming media service node in the streaming media service cluster in real time.

3. The method for fast fault identification and repair based on streaming media services according to claim 1, characterized in that, The specific content of Step 3 includes: Step 31: Deploy multiple streaming media service nodes and multiple proxy nodes of the cluster health management system in the same streaming media service cluster, and the streaming media service nodes and proxy nodes correspond one by one; Step 32: Establish a communication connection between the server node of the cluster health management system and multiple proxy nodes; Step 33: After the server node of the cluster health management system receives the fault event of the target service node reported for the first time, it enters Step 34; Step 34: The server node of the cluster health management system selects the proxy node of the cluster health management system corresponding to this target service node as the target proxy node according to the target service node; Step 35: The server node of the cluster health management system reports the fault event of this target service node to the target proxy node.

4. The method for rapid fault identification and repair based on streaming media service according to claim 3, wherein After Step 35, it further includes: Record the time of the first reported fault, and observe whether a second reported fault occurs within a set time interval. If so, it means that there is a network anomaly between the health monitoring system and the streaming media service cluster, and it is determined that the second to the nth reported faults occurring within the set time interval are false alarms, and the fault events reported from the second to the nth time are ignored; where n represents an integer and n>2.

5. The method for fast fault identification and repair based on streaming media services according to claim 1, wherein After Step 4, it further includes: The target proxy node of the cluster health management system sends a restart instruction to the service process of the streaming media service cluster. After receiving this restart instruction, the service process performs a restart of the service process to restore the service.

6. The method for rapid fault identification and repair based on streaming media services according to claim 1, characterized in that The specific content of Step 5 includes: Step 51: After the fault information analysis system receives the service process stack information and business log information reported by the target proxy node, obtain the function call stack of the key business threads from the service process stack information, and update the occurrence times of the historical function call chain in the fault information analysis system in real time; Step 52: The fault information analysis system pulls the code change records of the service in the most recent K times in real time, where the value range of K is set by the user according to requirements, and extracts the changed functions in the code change records.

7. The method for fast fault identification and repair based on streaming media service according to claim 1, characterized in that The specific steps of step 6 are as follows: Step 61: The fault information analysis system calculates the probability of each function call chain causing a fault according to the occurrence times of the function call chain and the changed functions in the code change records; the calculation formula for the probability Q of each function call chain causing a fault is: Q = M * N, where M represents the occurrence times of the function call chain, and N represents whether the code has changed; if the function call chain does not contain a changed function, it means that the code has not changed, and N = 1 is taken; if the function call chain contains a changed function, it means that the code has changed, and N > 1 is taken; Step 62: Sort the probabilities of each function call chain causing a fault calculated from large to small to obtain a probability sorting graph; Step 63: Download the corresponding business log information to the local from the fault information analysis system; Step 64: The fault information analysis system performs automatic analysis of code defects according to the probability sorting graph and the downloaded business log information, and performs analysis in the order of the probability magnitudes in the probability sorting graph.

8. The method for fast fault identification and repair based on streaming media service according to claim 6, characterized in that The specific steps of step 7 are as follows: The developer locates the function where the problem code is located according to the probability sorting graph and the downloaded log information, and repairs the code defect.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the fault rapid identification and repair method based on streaming media service according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, it implements the fault rapid identification and repair method based on streaming media service according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Fault detection method and device for streaming media system, equipment and storage medium

    CN114697247A