Media big data hadoop cluster monitoring method

By building a layered monitoring architecture and improved data acquisition and processing strategies, the problem of the inability to comprehensively and accurately monitor the media big data Hadoop cluster in the existing technology is solved, and efficient monitoring and data analysis capabilities of the Hadoop cluster are achieved.

CN120144403APending Publication Date: 2025-06-13LINYI ZHONGKE YINGTAI INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510236829.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing monitoring technology cannot comprehensively and accurately monitor the media big data Hadoop cluster, and it lacks the ability to process monitoring data, so it cannot detect and handle task abnormalities and node failures in a timely manner.

Method used

Build a layered monitoring architecture, including data acquisition layer, data processing layer, data storage layer and monitoring display layer. The improved Gmond client uses dynamic sampling strategy to collect data, Gmetad summarizes and initially processes the data, Nagios performs threshold monitoring of key indicators, combines RRDTool and HBase database to store monitoring data, and displays monitoring data through Gweb and customized web interfaces.

Benefits of technology

It realizes comprehensive and accurate monitoring of Hadoop clusters, promptly detects and handles task abnormalities and node failures, improves the data storage and analysis capabilities of the monitoring system, and meets the media industry's demand for in-depth mining of historical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144403A_ABST
    Figure CN120144403A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data processing and monitoring, and discloses a media big data hadoop cluster monitoring method which comprises the following steps: S1, constructing a layered monitoring architecture which is composed of a data acquisition layer, a data processing layer, a data storage layer and a monitoring display layer; according to the media big data hadoop cluster monitoring method, a layered monitoring architecture is constructed, so that the problem of incomplete functions when existing monitoring software is independently used is effectively solved; in a data acquisition layer, an improved Gmond client adopts a dynamic sampling strategy, and the time interval and range of data acquisition are adjusted in real time, so that data acquisition is more frequent in high load, and overhead is reduced in low load; the data processing layer is used for summarizing and primarily processing data through Gmetad, and carrying out threshold monitoring on key indexes by utilizing Nagios; the multi-layer structure enables the system to accurately monitor various resources of the Hadoop cluster, and timely discovers and processes task abnormity and node fault problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data processing and monitoring, and specifically to a method for monitoring a media big data Hadoop cluster. Background Art

[0002] With the rapid development of the media industry, the amount of media big data has increased sharply. The Hadoop cluster, with its distributed storage and computing capabilities, has become a key platform for processing media big data. However, the existing Hadoop cluster monitoring technologies have obvious deficiencies.

[0003] Current monitoring software such as Nagios, Cacti, Ganglia, etc., when used alone, cannot fully meet the monitoring requirements of the media big data Hadoop cluster. Although Ganglia is suitable for monitoring large-scale clusters, its alarm function is weak; Nagios has a strong alarm function but poor graphing function; Cacti also has defects in terms of alarm.

[0004] In addition, with the continuous increase of media big data, the complexity of data processing has risen. The existing monitoring systems are difficult to accurately monitor various resources of the Hadoop cluster during the process of media big data processing, and cannot timely discover and handle problems such as task anomalies and node failures caused by the sharp increase in data volume. At the same time, the traditional monitoring systems have limited capabilities in storing and analyzing monitoring data, and cannot meet the needs of the media industry for in-depth mining and utilization of historical data. Therefore, we propose a method for monitoring a media big data Hadoop cluster. Summary of the Invention

[0005] (I) Technical Problems to be Solved Aiming at the deficiencies of the existing technology, the present invention provides a method for monitoring a media big data Hadoop cluster, which has the advantages of comprehensively and accurately monitoring the running state of the cluster, efficiently processing and storing monitoring data, timely and accurately alarming, and deeply mining the data value, and solves the problems that the existing monitoring technologies cannot comprehensively and accurately monitor the media big data Hadoop cluster and have insufficient capabilities in processing monitoring data.

[0006] (II) Technical Solutions

[0007] To achieve the above purposes of comprehensively and accurately monitoring the running state of the cluster, efficiently processing and storing monitoring data, timely and accurately alarming, and deeply mining the data value, the present invention provides the following technical solutions: A method for monitoring a media big data Hadoop cluster, including the following steps: S1. Construct a hierarchical monitoring architecture, which is composed of a data acquisition layer, a data processing layer, a data storage layer, and a monitoring display layer; S2. At the data collection layer, an improved Gmond client is used to collect node data. The improved Gmond client adopts a dynamic sampling strategy to collect data, and adjusts the data collection time interval and scope by monitoring the change in the number of media big data processing tasks and the data traffic fluctuation. S3. At the data processing layer, Gmetad is used to summarize and preliminarily process the collected data, and Nagios is used to monitor the key indicators for thresholds. S4. At the data storage layer, RRDTool and HBase database are combined to store the monitoring data. S5. At the monitoring and display layer, the monitoring data is displayed through Gweb and a customized Web interface.

[0008] Preferably, the rules for adjusting the data collection time interval and scope by monitoring the change in the number of media big data processing tasks and the data traffic fluctuation are as follows: When the number of media big data processing tasks increases by 20% within 5 minutes, and at the same time the data traffic fluctuates by 30% within 1 minute, the data collection time interval is shortened from the default 5 minutes to 1 minute. When the number of tasks decreases by 30% within 10 minutes, and the data traffic fluctuates by 10% within 5 minutes, the collection time interval is extended to 10 minutes; if the data traffic ratio of the video processing task reaches 50%, the task progress, data processing rate, and data transmission delay indicators of the video processing task are key collected.

[0009] Preferably, the use of Nagios to monitor the key indicators for thresholds also includes improving the alarm mechanism of Nagios, adding the alarm functions of WeChat and DingTalk instant messaging tools, setting different alarm levels and notification methods, and having an intelligent alarm filtering function.

[0010] Preferably, the alarm functions of WeChat and DingTalk instant messaging tools are realized by calling the Webhook addresses of the WeChat enterprise account and the DingTalk robot. The Webhook addresses of the WeChat enterprise account and the DingTalk robot are configured in Nagios. When an alarm is triggered, the alarm information is encapsulated in the data format specified by the WeChat and DingTalk platforms and sent to the corresponding Webhook addresses.

[0011] Preferably, the setting of different alarm levels and notification methods is as follows: The faults in the media big data processing process are divided into serious faults, moderate faults, and minor faults. When the data block loss ratio in the Hadoop cluster reaches 10%, it is determined as a serious fault, and the administrator is notified simultaneously by SMS, email, WeChat, and DingTalk. When the failure rate of MapReduce tasks reaches 20%, it is determined as a medium - level fault, and the administrator is notified via email and WeChat. When the cache hit rate of HBase is lower than 60%, it is determined as a minor fault, and the administrator is only notified via WeChat.

[0012] Preferably, the intelligent alarm filtering function is as follows: When the number of alarm times for the same fault reaches 3 times within 15 minutes, the subsequent alarm messages are merged into one, and the number of fault occurrences is noted in the alarm message; If the fault duration is less than 5 minutes and it does not occur again within 1 hour, the fault alarm is automatically ignored.

[0013] Preferably, in the data storage layer, the data storage strategy of the RRDTool database is adjusted, and the table structure design of the HBase database is optimized.

[0014] Preferably, for the adjustment of the data storage strategy of the RRDTool database, the data storage period is set to 2 hours, and the archiving strategy is to archive once a week; For the optimization of the table structure design of the HBase database, the row key adopts a combination of timestamp and task ID, and the column families are divided into task information column family, data traffic column family, and fault information column family; The column qualifiers of the task information column family are task type and task progress, the column qualifiers of the data traffic column family are input data traffic and output data traffic, and the column qualifiers of the fault information column family are fault type and fault occurrence time.

[0015] Preferably, it also includes using Hive and Pig data analysis tools to deeply analyze the historical monitoring data stored in HBase, mining the resource usage patterns and fault occurrence rules of media big data processing tasks; Using Hive to write query statements to count the data traffic distribution of video processing, audio processing, and image processing tasks during the morning peak, evening peak, and late night time periods, as well as the distribution of corresponding fault types when tasks fail; Using Pig to write scripts to analyze the correlation between the task failure rate and data traffic fluctuations, so as to obtain the fault occurrence rules.

[0016] Compared with the prior art, the present invention provides a method for monitoring a media big data hadoop cluster, having the following beneficial effects: 1. The method for monitoring the media big data Hadoop cluster effectively solves the problem of incomplete functions when existing monitoring software is used alone by constructing a hierarchical monitoring architecture, including a data collection layer, a data processing layer, a data storage layer, and a monitoring and display layer. In the data collection layer, the improved Gmond client adopts a dynamic sampling strategy to adjust the time interval and range of data collection in real time, ensuring more frequent data collection under high load and reducing overhead under low load. The data processing layer aggregates and preliminarily processes data through Gmetad and uses Nagios to monitor the thresholds of key indicators. This multi-layer structure enables the system to accurately monitor various resources of the Hadoop cluster and promptly discover and handle task anomalies and node failures.

[0017] 2. In the data storage layer of the method for monitoring the media big data Hadoop cluster, the storage strategy of monitoring data is optimized by combining RRDTool and HBase database. RRDTool is used for the efficient storage and chart drawing of short-term periodic data, while HBase is used for the storage of long-term massive data to meet the need for in-depth mining of historical data. By setting the data storage period of RRDTool to 2 hours and the archiving strategy to once a week, and optimizing the table structure design of HBase, with the row key using a combination of timestamp and task ID, and the column families divided into task information, data traffic, and fault information, etc., it ensures the orderly storage and convenient query of data. These measures improve the data storage and analysis capabilities of the monitoring system and provide strong support for the media industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the embodiments and drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] Please refer to Figure 1 , a method for monitoring a media big data Hadoop cluster, including the following steps: S1. Construct a hierarchical monitoring architecture, which consists of a data collection layer, a data processing layer, a data storage layer, and a monitoring and display layer; S2. At the data collection layer, use the improved Gmond client to collect node data. The improved Gmond client adopts a dynamic sampling strategy to collect data, and adjusts the time interval and scope of data collection by monitoring the change in the number of media big data processing tasks and the fluctuation of data traffic. S3. At the data processing layer, use Gmetad to summarize and preliminarily process the collected data, and use Nagios to monitor the thresholds of key indicators. S4. At the data storage layer, combine RRDTool and HBase database to store monitoring data. S5. At the monitoring and display layer, display the monitoring data through Gweb and a customized web interface. Embodiment 1:

[0021] The hierarchical monitoring architecture constructed by the present invention consists of a data collection layer, a data processing layer, a data storage layer, and a monitoring and display layer.

[0022] Data collection layer: At this layer, use the improved Gmond client to collect node data. The traditional Gmond client usually adopts a fixed collection strategy during data collection and cannot well adapt to the dynamic changes of media big data processing tasks. The improved Gmond client of the present invention adopts a dynamic sampling strategy to collect data, and adjusts the time interval and scope of data collection by monitoring the change in the number of media big data processing tasks and the fluctuation of data traffic. Its principle is to obtain the information of the number of tasks and data traffic in real time through the built-in task number monitoring module and data traffic monitoring module. When the number of media big data processing tasks increases by 20% within 5 minutes and the data traffic fluctuates by 30% within 1 minute, it indicates that the system load and data transmission situation change greatly. At this time, shorten the data collection time interval from the default 5 minutes to 1 minute, so that data can be collected more frequently, and the change of the system operation status can be obtained in time to discover potential problems in time. When the number of tasks decreases by 30% within 10 minutes and the data traffic fluctuates by 10% within 5 minutes, it indicates that the system load is relatively low and the data change is not frequent. Extend the collection time interval to 10 minutes to reduce unnecessary data collection and reduce system overhead. If the data traffic ratio of the video processing task reaches 50%, then focus on collecting the task progress, data processing rate, and data transmission delay indicators of the video processing task. This is because the video processing task usually consumes a large amount of resources and has high real-time requirements. Focusing on collecting these indicators can better monitor the execution of the video processing task and ensure the quality and efficiency of video processing.

[0023] Data processing layer: Gmetad is used to summarize and preliminarily process the collected data; Gmetad collects the data collected by the improved Gmond client in a polling manner, and integrates and preliminarily analyzes this data, such as calculating statistical information such as average value, maximum value, and minimum value, so as to store and display data more efficiently in the follow-up; at the same time, Nagios is used to monitor the thresholds of key indicators; Nagios will monitor the key indicators of the Hadoop cluster in real time according to the pre-set thresholds; when the indicators exceed the threshold range, the corresponding alarm mechanism will be triggered; in the present invention, the alarm mechanism of Nagios is also improved, adding alarm functions for WeChat and DingTalk instant messaging tools, setting different alarm levels and notification methods, and having an intelligent alarm filtering function.

[0024] Data storage layer: Combine RRDTool and HBase database to store monitoring data; The RRDTool database is suitable for storing short-term and periodic data, and it can efficiently store and draw charts; while the HBase database is used to store long-term and massive data to meet the storage requirements of historical data; in order to better play the advantages of both, adjust the data storage strategy for the RRDTool database, set the data storage period to 2 hours, and the archiving strategy to archive once a week, so that while ensuring data timeliness, the storage space can be reasonably managed; optimize the table structure design for the HBase database, and use a combination of timestamp and task ID for the row key, so that data can be quickly queried according to the time sequence and task ID; the column families are divided into task information column family, data traffic column family, and fault information column family; the column qualifiers of the task information column family are task type and task progress, which are used to store the basic information and execution progress of the task; the column qualifiers of the data traffic column family are input data traffic and output data traffic, which are used to record the data traffic situation during the execution of the task; the column qualifiers of the fault information column family are fault type and fault occurrence time, which are used to store the fault information that occurs during the execution of the task.

[0025] Monitoring and display layer: Display the monitoring data through Gweb and a customized Web interface; Gweb can display the data collected and processed by Ganglia, presenting the running time series diagram of the system in the form of charts, which is convenient for administrators to intuitively view the running status of the system; the customized Web interface displays specific monitoring data according to the needs of media big data Hadoop cluster monitoring, such as highlighting the relevant indicators of video processing tasks and displaying fault information according to the alarm level, so that administrators can obtain key information more quickly. Embodiment 2:

[0026] Add instant messaging tool alarm function: The alarm functions of WeChat and DingTalk instant messaging tools send alarm information by calling the Webhook addresses of WeChat Enterprise Account and DingTalk robot. Configure the Webhook addresses of WeChat Enterprise Account and DingTalk robot in Nagios. When an alarm is triggered, Nagios will encapsulate the alarm information according to the data format specified by the WeChat and DingTalk platforms. For WeChat Enterprise Account, it is necessary to construct JSON format data containing fault information, time, cluster nodes, etc. according to its requirements. For DingTalk robot, data encapsulation should also be carried out according to its specific format requirements, and then the encapsulated alarm information is sent to the corresponding Webhook address to ensure that administrators can receive alarm information in a timely manner on the WeChat and DingTalk clients.

[0027] Set different alarm levels and notification methods: Classify the faults in the media big data processing process into severe faults, moderate faults and minor faults. When the data block loss ratio in the Hadoop cluster reaches 10%, it is determined as a severe fault. The loss of data blocks may lead to incomplete data and seriously affect the normal operation of the system. Therefore, SMS, email, WeChat and DingTalk are used to notify the administrator at the same time to ensure that the administrator can be aware of and handle the fault in the first time. When the failure rate of MapReduce tasks reaches 20%, it is determined as a moderate fault. At this time, the administrator is notified by email and WeChat. The email can describe the fault information in detail, and WeChat can remind the administrator to check the email in time. When the HBase cache hit rate is lower than 60%, it is determined as a minor fault, and only the administrator is notified by WeChat to avoid disturbing the administrator with excessive notifications.

[0028] Intelligent alarm filtering function: When the number of alarm times for the same fault reaches 3 times within 15 minutes, it indicates that the fault persists and is relatively serious. The subsequent alarm information is merged into one, and the number of fault occurrences is noted in the alarm information to facilitate the administrator to understand the severity and duration of the fault. If the fault duration is less than 5 minutes and it does not occur again within 1 hour, the fault alarm is automatically ignored. This is because a fault that appears briefly and does not occur again may be an accidental fluctuation of the system. Automatically ignoring it can reduce unnecessary alarm information and improve the efficiency of the administrator in handling alarm information. Example 3:

[0029] Optimization of Data Storage Strategy: Adjust the data storage strategy for the RRDTool database, set the data storage period to 2 hours, and the archiving strategy to once a week. Such settings can ensure the timeliness of short-term data while saving data within a certain period through archiving for subsequent data analysis and troubleshooting. Optimize the table structure design for the HBase database. The row key adopts a combination of timestamp and task ID. This design can ensure that data is stored in an orderly manner according to time sequence and task ID, facilitating quick location and query of monitoring data for specific times and tasks. The design of column families and column qualifiers classifies relevant information according to the characteristics of media big data processing tasks, improving the efficiency of data storage and query.

[0030] In-depth Analysis of Historical Data: Use Hive and Pig data analysis tools to conduct in-depth analysis of historical monitoring data stored in HBase. Write query statements using Hive to count the data traffic distribution of video processing, audio processing, and image processing tasks during morning rush hour, evening rush hour, and late at night. The principle is to use the query function of Hive to screen data traffic data for specific time periods and task types from the HBase database and conduct statistical analysis. By analyzing the data traffic distribution, the resource requirements of different types of tasks at different times can be understood, providing a basis for the reasonable allocation of cluster resources. At the same time, count the distribution of corresponding failure types when tasks fail to better understand the failure types prone to different task types and take preventive measures in advance. Write a script using Pig to analyze the correlation between task failure rate and data traffic fluctuations. The Pig script can process and analyze a large amount of historical data, and through establishing mathematical models or statistical methods, find the potential laws between task failure rate and data traffic fluctuations, thus obtaining the failure occurrence rules, providing strong support for the optimization and management of the cluster, and improving the efficiency and quality of media big data processing. Example 4:

[0031] System Deployment: Deploy the improved Gmond client on each node of the media big data Hadoop cluster to ensure that it can collect node data normally. Select appropriate nodes to deploy Gmetad and Nagios, configure the data transmission and communication parameters between them to ensure that data can be accurately summarized and processed. Deploy the RRDTool and HBase databases and configure them according to the optimized storage strategy and table structure design. Build Gweb and a customized web interface to ensure that it can display monitoring data normally.

[0032] Daily monitoring: The improved Gmond client collects data according to the dynamic sampling strategy and transmits the data to Gmetad. After Gmetad summarizes and preliminarily processes the data, Nagios monitors the key indicators for thresholds. When a failure occurs, Nagios notifies the administrator via text message, email, WeChat, DingTalk, etc. according to the severity of the failure. The administrator can view the monitoring data in real time through Gweb and the customized web interface to understand the running status of the cluster.

[0033] Data analysis and optimization: Regularly use Hive and Pig to deeply analyze the historical monitoring data stored in HBase. According to the analysis results, adjust the resource allocation strategy of the cluster and optimize the task scheduling algorithm to improve the efficiency and quality of media big data processing. For example, if it is found that the data traffic of video processing tasks is large during the evening peak period, resulting in an increase in the task failure rate, more resources can be allocated to video processing tasks during the evening peak period, or the scheduling order of video processing tasks can be optimized to avoid task conflicts and resource competition.

[0034] In summary, for the method of monitoring the media big data hadoop cluster, by constructing a hierarchical monitoring architecture, including a data collection layer, a data processing layer, a data storage layer, and a monitoring and display layer, the present invention effectively solves the problem of incomplete functions when existing monitoring software is used alone. In the data collection layer, the improved Gmond client adopts a dynamic sampling strategy to adjust the time interval and range of data collection in real time, ensuring more frequent data collection under high load and reducing overhead under low load. The data processing layer summarizes and preliminarily processes the data through Gmetad and uses Nagios to monitor the thresholds of key indicators. This multi-layer structure enables the system to accurately monitor various resources of the Hadoop cluster, and promptly discover and handle task anomalies and node failure problems.

[0035] Moreover, for the method of monitoring the media big data hadoop cluster, in the data storage layer, the storage strategy of monitoring data is optimized by combining RRDTool and the HBase database. RRDTool is used for the efficient storage and chart drawing of short-term periodic data, while HBase is used for the storage of long-term massive data to meet the need for in-depth mining of historical data. By setting the data storage period of RRDTool to 2 hours, the archiving strategy to once a week, and optimizing the table structure design of HBase, with the row key combined with the timestamp and task ID, and the column families divided into task information, data traffic, and failure information, etc., it ensures the orderly storage and convenient query of data. These measures improve the data storage and analysis capabilities of the monitoring system, provide strong support for the media industry, and solve the problems that existing monitoring technologies cannot comprehensively and accurately monitor the media big data Hadoop cluster and have insufficient data processing capabilities for monitoring data.

[0036] All related modules involved in this system are hardware system modules or functional modules that combine computer software programs or protocols in the prior art with hardware. The computer software programs or protocols themselves involved in this functional module are all well-known technologies to those skilled in the art and are not the improvements of this system. The improvement of this system lies in the interaction relationship or connection relationship between modules, that is, the overall structure of the system is improved to solve the corresponding technical problems to be solved by this system.

[0037] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for monitoring a media big data Hadoop cluster, characterized in that: The following steps are involved: S1. Build a hierarchical monitoring architecture, which consists of a data collection layer, a data processing layer, a data storage layer, and a monitoring display layer; S2. At the data collection layer, node data is collected using an improved Gmond client, which adopts a dynamic sampling strategy to collect data and adjusts the time interval and range of data collection by monitoring the task quantity changes and data flow fluctuations of the media big data processing tasks; S3. In the data processing layer, Gmetad is used to aggregate and preliminarily process the collected data, and Nagios is used to monitor the thresholds of key indicators. S4. In the data storage layer, RRDTool and HBase database are combined to store monitoring data; S5. In the monitoring display layer, monitoring data is displayed through Gweb and customized web interface.

2. A method for monitoring a media big data Hadoop cluster according to claim 1, characterized in that: The rules for adjusting the time interval and range of data collection by monitoring the task quantity changes and data flow fluctuations of the media big data processing tasks are as follows: When the number of media big data processing tasks increases by 20% within 5 minutes and the data traffic fluctuates by 30% within 1 minute, shorten the data collection interval from the default 5 minutes to 1 minute; When the number of tasks decreases by 30% within 10 minutes and the data traffic fluctuates by 10% within 5 minutes, the collection time interval will be extended to 10 minutes. If the data traffic proportion of the video processing task reaches 50%, the focus will be on collecting the task progress, data processing rate, and data transmission delay indicators of the video processing task.

3. The method for monitoring a media big data Hadoop cluster according to claim 1, characterized in that: The use of Nagios to perform threshold monitoring on key indicators also includes improving the alarm mechanism of Nagios, adding alarm functions for WeChat and DingTalk instant messaging tools, setting different alarm levels and notification methods, and having intelligent alarm filtering functions.

4. A method for monitoring a media big data Hadoop cluster according to claim 3, characterized in that: The alarm function of the WeChat and DingTalk instant messaging tools sends alarm information by calling the Webhook addresses of the WeChat enterprise account and the DingTalk robot. The Webhook addresses of the WeChat enterprise account and the DingTalk robot are configured in Nagios. When the alarm is triggered, the alarm information is packaged in the data format specified by the WeChat and DingTalk platforms and sent to the corresponding Webhook address.

5. The method for monitoring a media big data Hadoop cluster according to claim 3, characterized in that: The settings for different alarm levels and notification methods are as follows: Classify the failures in the media big data processing process into serious failures, moderate failures, and minor failures; When the data block loss ratio in the Hadoop cluster reaches 10%, it is considered a serious failure and the administrator is notified simultaneously via SMS, email, WeChat, and DingTalk. When the MapReduce task failure rate reaches 20%, it is considered a moderate failure and the administrator is notified via email and WeChat; When the HBase cache hit rate is lower than 60%, it is considered a minor failure and the administrator is notified only through WeChat.

6. The method for monitoring a media big data Hadoop cluster according to claim 3, characterized in that: The intelligent alarm filtering function is as follows: when the same fault alarms three times within 15 minutes, the subsequent alarm information will be merged into one, and the number of fault occurrences will be noted in the alarm information; if the fault duration is less than 5 minutes and does not reappear within 1 hour, the fault alarm will be automatically ignored.

7. The method for monitoring a media big data Hadoop cluster according to claim 1, characterized in that: At the data storage layer, adjust the data storage strategy for the RRDTool database and optimize the table structure design for the HBase database.

8. The method for monitoring a media big data Hadoop cluster according to claim 7, characterized in that: The data storage strategy of the RRDTool database is adjusted, the data storage period is set to 2 hours, and the archiving strategy is to archive once a week; the table structure design of the HBase database is optimized, the row key adopts the combination of timestamp and task ID, and the column family is divided into task information column family, data flow column family, and fault information column family; the column qualifiers of the task information column family are task type and task progress, the column qualifiers of the data flow column family are input data flow and output data flow, and the column qualifiers of the fault information column family are fault type and fault occurrence time.

9. The method for monitoring a media big data Hadoop cluster according to claim 1, characterized in that: It also includes using Hive and Pig data analysis tools to conduct in-depth analysis of historical monitoring data stored in HBase, and to explore the resource usage patterns and failure patterns of media big data processing tasks; using Hive to write query statements to count the data traffic distribution of video processing, audio processing, and image processing tasks during the morning peak, evening peak, and late night time periods, as well as the corresponding fault type distribution when the task fails; using Pig to write scripts to analyze the correlation between task failure rate and data traffic fluctuations, so as to derive the failure pattern.