Ambari-based big data platform monitoring method and system
By adopting Ambari-based monitoring method on the Hadoop big data platform, combining activity data and operating status information, a platform performance prediction model is built, which solves the problems and failures that the Hadoop big data platform may encounter during operation, improves the stability and reliability of the platform, and optimizes workflow and efficiency.
Patent Information
- Application Number
- CN202411398397.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2025-05-06
AI Technical Summary
How to timely and accurately discover and deal with various problems and failures that the Hadoop big data platform may encounter during operation to ensure the stable operation of the platform.
A big data platform monitoring method based on Ambari is provided. Through four steps: data collection, data processing, model construction and platform monitoring and early warning, combining activity data and operating status information, a platform performance prediction model is built, performance changes in future time periods are predicted, and early warning analysis and optimization suggestions are carried out.
It improves the stability and reliability of the big data platform, optimizes work processes, improves work efficiency and quality, and promotes the innovation and development of work.
Smart Images

Figure CN119938435A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data management and monitoring technology, and in particular to a big data platform monitoring method and system based on Ambari. Background Art
[0002] With the rapid development of information technology, big data has become an important support for work. As an important platform for big data processing, the stable operation of Hadoop is crucial to the normal development of work. However, Hadoop clusters may encounter various problems and failures during operation. How to timely and accurately discover and handle these problems to ensure the stable operation of the platform is an important challenge facing current work.
[0003] How to integrate work needs and conduct monitoring and early warning on the Hadoop big data platform is a technical problem that needs to be solved. Summary of the invention
[0004] The technical task of the present invention is to address the above shortcomings and provide a big data platform monitoring method and system based on Ambari to solve the technical problem of how to integrate work needs and monitor and warn the Hadoop big data platform.
[0005] In a first aspect, the present invention provides a big data platform monitoring method based on Ambari, which is used to monitor and warn a Hadoop big data platform including a Hadoop cluster and a data platform in combination with activity data, and comprises the following steps:
[0006] Data collection: Based on the activity records and knowledge base, the data generated by historical activities are collected as historical activity data, and the historical operation status information of the Hadoop big data platform is collected through the Ambari API. The operation status information includes the node status and task execution status of the Hadoop cluster;
[0007] Data processing: Use historical activity data and historical operation status information as target data, perform data cleaning and feature extraction on the target data, obtain activity features including activity frequency and participation, and platform performance features including CPU utilization and memory usage, and match activity features and platform performance features in the time dimension to establish common time series features as the sample set;
[0008] Model construction: Build a platform performance prediction model based on machine learning algorithms and train the platform performance prediction model based on sample sets. The performance prediction model uses time series features as input, predicts and outputs performance changes of the Hadoop big data platform in the future time period;
[0009] Platform monitoring and early warning: Collect current personnel data and operating status information, establish common time series features based on the current personnel data and operating status information, use the time series features as input, and use the trained platform performance prediction model to predict the performance changes of the Hadoop big data platform in the future time period, and perform early warning analysis on the Hadoop big data platform based on the prediction results, obtain early warning information and push the early warning information to relevant users, adjust the Hadoop cluster based on the prediction results, and provide optimization suggestions for activities.
[0010] Preferably, when performing data cleaning on the target data, outliers, missing values and duplicates are removed.
[0011] Preferably, when collecting data generated by historical activities based on activity records and knowledge base as historical activity data, the data generated by the work is collected based on the activity records and knowledge base, and the data is sorted into a format suitable for Hadoop processing, and the data is uploaded to HDFS through Hadoop command line tools to obtain activity data, wherein the formats suitable for Hadoop processing include CSV format and JSON format, and the command line tools include hdfs dfs-put; correspondingly, when data cleaning and feature extraction are performed on the target data with historical activity data as the target data, SQL query statements are written through Hive, and the target data is aggregated and analyzed through the SQL query statements, or the target data is batch processed through the data processing frameworks Spark and Flink;
[0012] When collecting historical operating status information of the Hadoop big data platform through the Ambari API, the status of the Hadoop cluster is tracked through the monitoring function provided by Ambari, and the log files of the Hadoop cluster are viewed to analyze the job execution and system performance to obtain the operating status information; correspondingly, when performing data cleaning and feature extraction on the target data with the operating status information as the target data, the target data is queried and analyzed through the tools in the Hadoop big data platform, including Pig and Impala.
[0013] Preferably, the platform prediction model is used to predict performance changes of the Hadoop big data platform based on a linear regression algorithm or a decision tree.
[0014] Preferably, when training the platform performance prediction model based on the sample set, the parameters of the platform performance prediction model are adjusted by a cross-validation method, and the platform performance prediction model is evaluated using accuracy and recall as indicators.
[0015] In a second aspect, the present invention provides an Ambari-based big data platform monitoring system, which is used to monitor and warn a Hadoop big data platform including a Hadoop cluster and a data platform through an Ambari-based big data platform monitoring method as described in any one of the first aspects, wherein the system includes a data acquisition module, a data processing module, a model building module, and a platform monitoring and warning module.
[0016] The data collection module is used to perform the following: collect the data generated by historical activities as historical activity data based on activity records and knowledge base, and collect the historical operation status information of the Hadoop big data platform through the Ambari API. The operation status information includes the node status and task execution status of the Hadoop cluster;
[0017] The data processing module is used to perform the following: taking historical activity data and historical operation status information as target data, respectively, performing data cleaning and feature extraction on the target data, obtaining activity features including activity frequency and participation, and platform performance features including CPU utilization and memory utilization, and matching the activity features and platform performance features in the time dimension, and establishing a common time series feature as a sample set;
[0018] The model building module is used to perform the following: build a platform performance prediction model based on a machine learning algorithm, and train the platform performance prediction model based on a sample set. The performance prediction model uses time series features as input, predicts and outputs performance changes of the Hadoop big data platform in the future time period;
[0019] The platform monitoring and early warning module is used to perform the following: collect current personnel data and operating status information, establish common time series features based on the current personnel data and operating status information, use the time series features as input, and use the trained platform performance prediction model to predict the performance changes of the Hadoop big data platform in the future time period, and perform early warning analysis on the Hadoop big data platform based on the prediction results, obtain early warning information and push the early warning information to relevant users, adjust the Hadoop cluster based on the prediction results, and provide optimization suggestions for activities.
[0020] Preferably, when performing data cleaning on the target data, the data processing module is used to perform the following operations on the target data: removing outliers, missing values and duplicates.
[0021] Preferably, when the data generated by historical activities is collected based on the activity records and the knowledge base as historical activity data, the data processing module is used to perform the following: collect the data generated by the work based on the activity records and the knowledge base, organize the data into a format suitable for Hadoop processing, and upload the data to HDFS through the Hadoop command line tool to obtain activity data, wherein the formats suitable for Hadoop processing include CSV format and JSON format, and the command line tool includes hdfs dfs-put; correspondingly, when data cleaning and feature extraction are performed on the target data with the historical activity data as the target data, write SQL query statements through Hive, perform aggregation analysis on the target data through the SQL query statements, or perform batch processing on the target data through the data processing frameworks Spark and Flink;
[0022] When collecting historical operating status information of the Hadoop big data platform through the Ambari API, the data processing module is used to perform the following: track the status of the Hadoop cluster through the monitoring function provided by Ambari, view the log files of the Hadoop cluster, analyze the job execution and system performance, and obtain the operating status information; correspondingly, when performing data cleaning and feature extraction on the target data with the operating status information as the target data, query and analyze the target data through the tools in the Hadoop big data platform, where the tools include Pig and Impala.
[0023] Preferably, the platform prediction model is used to predict performance changes of the Hadoop big data platform based on a linear regression algorithm or a decision tree.
[0024] Preferably, when the platform performance prediction model is trained based on the sample set, the model building module is used to adjust the parameters of the platform performance prediction model through a cross-validation method, and to evaluate the platform performance prediction model using accuracy and recall as indicators.
[0025] The Ambari-based big data platform monitoring method and system of the present invention have the following advantages:
[0026] 1. Improved the stability and reliability of the big data platform: Through real-time monitoring of the Ambari big data platform, potential failures and anomalies can be discovered and handled in a timely manner to ensure the stable operation of the platform. This can not only reduce work interruptions caused by platform failures, but also avoid significant losses caused by data loss or damage;
[0027] 2. Optimize workflow: Combine data with the operating data of the big data platform, analyze the correlation and impact between the two, provide data support and decision-making basis for work, help optimize personnel activity arrangements, improve work efficiency, improve work methods, etc., and promote scientific and standardized work;
[0028] 3. Improve work efficiency and quality: Through real-time monitoring and analysis, timely discover and deal with problems that may affect work efficiency and quality. For example, when a functional module is frequently used but has obvious performance bottlenecks, it can be optimized or upgraded in time to improve overall work efficiency and quality.
[0029] 4. Promote innovation and development of work: The monitoring and analysis results based on the big data platform can provide new ideas and methods for work. For example, by analyzing the data of personnel activities, it can be found that certain forms of activities are popular and effective, and these forms of activities can be further promoted and optimized; at the same time, the content and form of work can also be adjusted and optimized according to the analysis results. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0031] The present invention is further described below in conjunction with the accompanying drawings.
[0032] Figure 1 This is a flowchart of a big data platform monitoring method based on Ambari in Example 1. DETAILED DESCRIPTION
[0033] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments may be combined with each other.
[0034] The embodiment of the present invention provides a big data platform monitoring method and system based on Ambari, which is used to solve the technical problem of how to integrate work requirements and monitor and warn the Hadoop big data platform.
[0035] Embodiment 1:
[0036] The present invention discloses a big data platform monitoring method based on Ambari, which is used to monitor and warn a Hadoop big data platform including a Hadoop cluster and a data platform in combination with activity data, and includes four steps: data collection, data processing, model construction, and platform monitoring and warning.
[0037] Step S100 Data Collection: Based on the activity records and knowledge base, the data generated by historical activities are collected as historical activity data, and the historical operation status information of the Hadoop big data platform is collected through the Ambari API. The operation status information includes the node status of the Hadoop cluster and the task execution status.
[0038] Ambari is an open source project that has tools to simplify the configuration, management, and monitoring of Apache Hadoop clusters. The Ambari API can be used to build customized management tools to assist in the automation of big data cluster operations, such as rapid deployment, elastic scaling, and automatic backup, making the Hadoop big data platform run more efficiently and easier to maintain.
[0039] Step S200 data processing: taking historical activity data and historical operation status information as target data respectively, performing data cleaning and feature extraction on the target data, obtaining activity features including activity frequency and participation, and platform performance features including CPU utilization and memory usage, and matching the activity features and platform performance features in the time dimension, and establishing common time series features as a sample set.
[0040] In this embodiment, when data generated by historical activities are collected as historical activity data based on activity records and knowledge bases, data generated by the work are collected based on the activity records and knowledge bases, and the data are organized into a format suitable for Hadoop processing, and the data is uploaded to HDFS through Hadoop's command line tool to obtain activity data, where the formats suitable for Hadoop processing include CSV format and JSON format, and the command line tools include hdfs dfs-put; correspondingly, when data cleaning and feature extraction are performed on the target data with historical activity data as the target data, SQL query statements are written through Hive, and aggregation analysis is performed on the target data through the SQL query statements, or batch processing is performed on the target data through the data processing frameworks Spark and Flink.
[0041] When collecting historical operating status information of the Hadoop big data platform through the Ambari API, the status of the Hadoop cluster is tracked through the monitoring function provided by Ambari, and the log files of the Hadoop cluster are viewed to analyze the job execution and system performance to obtain the operating status information; correspondingly, when performing data cleaning and feature extraction on the target data with the operating status information as the target data, the target data is queried and analyzed through the tools in the Hadoop big data platform, including Pig and Impala.
[0042] Among them, when cleaning the target data, outliers, missing values and duplicates are removed.
[0043] When extracting features, this embodiment selects and constructs data features that are helpful for machine learning models to understand. For example, features such as activity frequency and participation are extracted from work data; features such as CPU utilization and memory usage are extracted from big data platform operation data, and work data is matched with big data platform operation data in the time dimension to establish common time series features, and the processed data is stored in Hadoop's HDFS, or a data warehouse is built using tools such as Hive.
[0044] Step S300: Model construction: construct a platform performance prediction model based on a machine learning algorithm, and train the platform performance prediction model based on a sample set. The performance prediction model takes time series features as input, predicts and outputs performance changes of the Hadoop big data platform in future time periods.
[0045] In this embodiment, the platform prediction model is used to predict the performance change of the Hadoop big data platform based on a linear regression algorithm or a decision tree.
[0046] As a specific implementation of model building, select the appropriate algorithm according to the problem type (classification, regression, clustering, etc.). For example, you can use linear regression to predict the performance changes of a Hadoop cluster, or use a decision tree to identify influencing factors.
[0047] As a specific implementation of the platform prediction model, the input (features) include working data and big data platform operation data.
[0048] The working data includes the following:
[0049] Personnel activity frequency: the number of times a person participates in an activity;
[0050] Learning participation: the proportion or activity of people participating in learning;
[0051] Knowledge base usage: how often people access the knowledge base;
[0052] Types of activities: The number of different types of activities.
[0053] The big data platform operation data includes the following:
[0054] CPU utilization: CPU usage of each node in the cluster;
[0055] Memory utilization: memory usage of each node in the cluster;
[0056] Network traffic: the amount of network traffic within the cluster;
[0057] Job execution status: number of successful / failed jobs, average execution time, etc.
[0058] Resource allocation: the ratio of resources allocated to each task.
[0059] The platform prediction model predicts the performance indicators of the platform, including the following:
[0060] Predict future CPU utilization;
[0061] Predict future memory utilization;
[0062] Predict the success rate of operations within a certain period of time in the future;
[0063] Predict the average job execution time within a certain period of time in the future.
[0064] In this embodiment, when model training is performed based on a sample set, the sample set is divided into a training set, a validation set, and a test set to evaluate the generalization ability of the model.
[0065] During model training, the training set data is used to train the selected machine learning model, and the model parameters are adjusted through methods such as cross-validation to obtain the best performance.
[0066] When evaluating the model, the accuracy, recall and other indicators of the model are evaluated on the validation set, and finally a final evaluation is performed on the test set to ensure that the model has good generalization ability.
[0067] Step S400 platform monitoring and early warning: collect current personnel data and operation status information, establish common time series features based on the current personnel data and operation status information, use the time series features as input, and use the trained platform performance prediction model to predict the performance changes of the Hadoop big data platform in the future time period, and perform early warning analysis on the Hadoop big data platform based on the prediction results, obtain early warning information and push the early warning information to relevant users, adjust the Hadoop cluster based on the prediction results, and provide optimization suggestions for the activities.
[0068] In this embodiment, the trained model is used to predict how work data affects the performance of the big data platform, or conversely, how changes in the performance of the big data platform affect the results of the work. Based on the results output by the model, specific suggestions are put forward to improve the work or optimize the performance of the big data platform.
[0069] As the specific implementation of the platform early warning monitoring, when abnormal situations are detected, such as excessive CPU usage, insufficient disk space, etc., alarm information will be generated in a timely manner, and the alarm information will be notified to relevant personnel via SMS, email, etc., and optimization suggestions and support will be provided for the activities based on the monitoring data.
[0070] As an improvement of this embodiment, during the specific implementation process, the data changes, and the model needs to be retrained regularly to adapt to the new situation, and the model is continuously adjusted according to business feedback to optimize its performance.
[0071] Embodiment 2:
[0072] The present invention discloses a big data platform monitoring system based on Ambari, comprising a data acquisition module, a data processing module, a model building module and a platform monitoring and early warning module.
[0073] The data collection module is used to perform the following: based on the activity records and knowledge base, collect the data generated by historical activities as historical activity data, and collect the historical operation status information of the Hadoop big data platform through the Ambari API. The operation status information includes the node status and task execution status of the Hadoop cluster.
[0074] Ambari is an open source project that has tools to simplify the configuration, management, and monitoring of Apache Hadoop clusters. The Ambari API can be used to build customized management tools to assist in the automation of big data cluster operations, such as rapid deployment, elastic scaling, and automatic backup, making the Hadoop big data platform run more efficiently and easier to maintain.
[0075] The data processing module is used to perform the following: taking historical activity data and historical operating status information as target data, respectively, performing data cleaning and feature extraction on the target data, obtaining activity features including activity frequency and participation, and platform performance features including CPU utilization and memory usage, and matching activity features and platform performance features in the time dimension to establish common time series features as a sample set.
[0076] In this embodiment, when the data generated by historical activities is collected as historical activity data based on activity records and knowledge bases, the data processing module is used to perform the following: collect the data generated by the work based on the activity records and knowledge bases, organize the data into a format suitable for Hadoop processing, and upload the data to HDFS through Hadoop's command line tool to obtain activity data, where the formats suitable for Hadoop processing include CSV format and JSON format, and the command line tools include hdfs dfs-put; correspondingly, when data cleaning and feature extraction are performed on the target data with historical activity data as the target data, write SQL query statements through Hive, perform aggregation analysis on the target data through SQL query statements, or perform batch processing on the target data through data processing frameworks Spark and Flink.
[0077] When collecting historical operating status information of the Hadoop big data platform through the Ambari API, the data processing module is used to perform the following: track the status of the Hadoop cluster through the monitoring function provided by Ambari, view the log files of the Hadoop cluster, analyze the job execution and system performance, and obtain the operating status information; correspondingly, when performing data cleaning and feature extraction on the target data with the operating status information as the target data, query and analyze the target data through the tools in the Hadoop big data platform, where the tools include Pig and Impala.
[0078] Among them, when cleaning the target data, outliers, missing values and duplicates are removed.
[0079] When extracting features, this embodiment selects and constructs data features that are helpful for machine learning models to understand. For example, features such as activity frequency and participation are extracted from work data; features such as CPU utilization and memory usage are extracted from big data platform operation data, and work data is matched with big data platform operation data in the time dimension to establish common time series features, and the processed data is stored in Hadoop's HDFS, or a data warehouse is built using tools such as Hive.
[0080] The model building module is used to perform the following: build a platform performance prediction model based on a machine learning algorithm, and train the platform performance prediction model based on a sample set. The performance prediction model takes time series features as input, predicts and outputs the performance changes of the Hadoop big data platform in the future time period.
[0081] In this embodiment, the platform prediction model is used to predict the performance change of the Hadoop big data platform based on a linear regression algorithm or a decision tree.
[0082] As a specific implementation of the model building module, the appropriate algorithm is selected according to the problem type (classification, regression, clustering, etc.). For example, linear regression can be used to predict the performance changes of a Hadoop cluster, or a decision tree can be used to identify influencing factors.
[0083] As a specific implementation of the platform prediction model, the input (features) include working data and big data platform operation data.
[0084] The working data includes the following:
[0085] Personnel activity frequency: the number of times a person participates in an activity;
[0086] Learning participation: the proportion or activity of people participating in learning;
[0087] Knowledge base usage: how often people access the knowledge base;
[0088] Types of activities: The number of different types of activities.
[0089] The big data platform operation data includes the following:
[0090] CPU utilization: CPU usage of each node in the cluster;
[0091] Memory utilization: memory usage of each node in the cluster;
[0092] Network traffic: the amount of network traffic within the cluster;
[0093] Job execution status: number of successful / failed jobs, average execution time, etc.
[0094] Resource allocation: the ratio of resources allocated to each task.
[0095] The platform prediction model predicts the performance indicators of the platform, including the following:
[0096] Predict future CPU utilization;
[0097] Predict future memory utilization;
[0098] Predict the success rate of operations within a certain period of time in the future;
[0099] Predict the average job execution time within a certain period of time in the future.
[0100] In this embodiment, when model training is performed based on a sample set, the sample set is divided into a training set, a validation set, and a test set to evaluate the generalization ability of the model.
[0101] During model training, the training set data is used to train the selected machine learning model, and the model parameters are adjusted through methods such as cross-validation to obtain the best performance.
[0102] When evaluating the model, the accuracy, recall and other indicators of the model are evaluated on the validation set, and finally a final evaluation is performed on the test set to ensure that the model has good generalization ability.
[0103] The platform monitoring and early warning module is used to perform the following: collect current personnel data and operating status information, establish common time series features based on the current personnel data and operating status information, use the time series features as input, and use the trained platform performance prediction model to predict the performance changes of the Hadoop big data platform in the future time period, and perform early warning analysis on the Hadoop big data platform based on the prediction results, obtain early warning information and push the early warning information to relevant users, adjust the Hadoop cluster based on the prediction results, and provide optimization suggestions for activities.
[0104] In this embodiment, the trained model is used to predict how work data affects the performance of the big data platform, or conversely, how changes in the performance of the big data platform affect the results of the work. Based on the results output by the model, specific suggestions are put forward to improve the work or optimize the performance of the big data platform.
[0105] As the specific implementation of the platform early warning monitoring, when abnormal situations are detected, such as excessive CPU usage, insufficient disk space, etc., alarm information will be generated in a timely manner, and the alarm information will be notified to relevant personnel via SMS, email, etc., and optimization suggestions and support will be provided for the activities based on the monitoring data.
[0106] As an improvement of this embodiment, during the specific implementation process, the data changes, and the platform monitoring and early warning module needs to retrain the model regularly to adapt to the new situation, and continuously adjust the model according to business feedback to optimize its performance.
[0107] The system of this embodiment can execute the method disclosed in Example 1 to monitor and warn the Hadoop big data platform including the Hadoop cluster and the data platform.
[0108] The present invention has been shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the means in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.
Claims
1. A big data platform monitoring method based on Ambari, characterized in that: The method is used to monitor and warn the Hadoop big data platform including the Hadoop cluster and the data platform in combination with the activity data, including the following steps: Data collection: Based on the activity records and knowledge base, the data generated by historical activities are collected as historical activity data, and the historical operation status information of the Hadoop big data platform is collected through the Ambari API. The operation status information includes the node status and task execution status of the Hadoop cluster; Data processing: Use historical activity data and historical operation status information as target data, perform data cleaning and feature extraction on the target data, obtain activity features including activity frequency and participation, and platform performance features including CPU utilization and memory usage, and match activity features and platform performance features in the time dimension to establish common time series features as the sample set; Model construction: Build a platform performance prediction model based on machine learning algorithms and train the platform performance prediction model based on sample sets. The performance prediction model uses time series features as input, predicts and outputs performance changes of the Hadoop big data platform in the future time period; Platform monitoring and early warning: Collect current personnel data and operating status information, establish common time series features based on the current personnel data and operating status information, use the time series features as input, and use the trained platform performance prediction model to predict the performance changes of the Hadoop big data platform in the future time period, and perform early warning analysis on the Hadoop big data platform based on the prediction results, obtain early warning information and push the early warning information to relevant users, adjust the Hadoop cluster based on the prediction results, and provide optimization suggestions for activities.
2. The Ambari-based big data platform monitoring method according to claim 1, characterized in that: When cleaning the target data, remove outliers, missing values, and duplicates.
3. The Ambari-based big data platform monitoring method according to claim 1, characterized in that: When the data generated by historical activities is collected based on activity records and knowledge bases as historical activity data, the data generated by the work is collected based on the activity records and knowledge bases, and the data is sorted into a format suitable for Hadoop processing, and the data is uploaded to HDFS through Hadoop command line tools to obtain activity data, where the formats suitable for Hadoop processing include CSV format and JSON format, and the command line tools include hdfs dfs-put; correspondingly, when data cleaning and feature extraction are performed on the target data with historical activity data as the target data, SQL query statements are written through Hive, and the target data is aggregated and analyzed through the SQL query statements, or the target data is batch processed through the data processing frameworks Spark and Flink; When collecting historical operating status information of the Hadoop big data platform through the Ambari API, the status of the Hadoop cluster is tracked through the monitoring function provided by Ambari, and the log files of the Hadoop cluster are viewed to analyze the job execution and system performance to obtain the operating status information; correspondingly, when performing data cleaning and feature extraction on the target data with the operating status information as the target data, the target data is queried and analyzed through the tools in the Hadoop big data platform, including Pig and Impala.
4. The Ambari-based big data platform monitoring method according to claim 1, characterized in that: The platform prediction model is used to predict the performance change of the Hadoop big data platform based on a linear regression algorithm or a decision tree.
5. The Ambari-based big data platform monitoring method according to claim 1, characterized in that: When the platform performance prediction model is trained based on the sample set, the parameters of the platform performance prediction model are adjusted through the cross-validation method, and the platform performance prediction model is evaluated using accuracy and recall as indicators.
6. A big data platform monitoring system based on Ambari, characterized in that: The system is used to monitor and warn a Hadoop big data platform including a Hadoop cluster and a data platform through a big data platform monitoring method based on Ambari as described in any one of claims 1 to 5, wherein the system includes a data acquisition module, a data processing module, a model building module, and a platform monitoring and warning module. The data collection module is used to perform the following: based on the activity records and knowledge base, collect the data generated by historical activities as historical activity data, and collect the historical operation status information of the Hadoop big data platform through the AmbariAPI. The operation status information includes the node status and task execution status of the Hadoop cluster; The data processing module is used to perform the following: taking historical activity data and historical operation status information as target data, respectively, performing data cleaning and feature extraction on the target data, obtaining activity features including activity frequency and participation, and platform performance features including CPU utilization and memory utilization, and matching the activity features and platform performance features in the time dimension, and establishing a common time series feature as a sample set; The model building module is used to perform the following: build a platform performance prediction model based on a machine learning algorithm, and train the platform performance prediction model based on a sample set. The performance prediction model uses time series features as input, predicts and outputs performance changes of the Hadoop big data platform in the future time period; The platform monitoring and early warning module is used to perform the following: collect current personnel data and operating status information, establish common time series features based on the current personnel data and operating status information, use the time series features as input, and use the trained platform performance prediction model to predict the performance changes of the Hadoop big data platform in the future time period, and perform early warning analysis on the Hadoop big data platform based on the prediction results, obtain early warning information and push the early warning information to relevant users, adjust the Hadoop cluster based on the prediction results, and provide optimization suggestions for activities.
7. The Ambari-based big data platform monitoring system according to claim 6, characterized in that: When performing data cleaning on the target data, the data processing module is used to perform the following operations on the target data: removing outliers, missing values, and duplicates.
8. The Ambari-based big data platform monitoring system according to claim 6, characterized in that: When the data generated by historical activities is collected based on the activity records and knowledge base as historical activity data, the data processing module is used to perform the following: collect the data generated by the work based on the activity records and knowledge base, organize the data into a format suitable for Hadoop processing, and upload the data to HDFS through Hadoop's command line tool to obtain activity data, where the formats suitable for Hadoop processing include CSV format and JSON format, and the command line tool includes hdfs dfs-put; correspondingly, when the historical activity data is used as the target data for data cleaning and feature extraction, SQL query statements are written through Hive, and the target data is aggregated and analyzed through the SQL query statements, or the target data is batch processed through the data processing frameworks Spark and Flink; When collecting historical operating status information of the Hadoop big data platform through the Ambari API, the data processing module is used to perform the following: track the status of the Hadoop cluster through the monitoring function provided by Ambari, view the log files of the Hadoop cluster, analyze the job execution and system performance, and obtain the operating status information; correspondingly, when performing data cleaning and feature extraction on the target data with the operating status information as the target data, query and analyze the target data through the tools in the Hadoop big data platform, where the tools include Pig and Impala.
9. The Ambari-based big data platform monitoring system according to claim 6, characterized in that: The platform prediction model is used to predict the performance change of the Hadoop big data platform based on a linear regression algorithm or a decision tree.
10. The Ambari-based big data platform monitoring system according to claim 6, characterized in that: When the platform performance prediction model is trained based on the sample set, the model building module is used to adjust the parameters of the platform performance prediction model through the cross-validation method, and to evaluate the platform performance prediction model using accuracy and recall as indicators.