Big data management platform CDH automatic deployment method, device and system

Through the CDH automated deployment system, the problem of manual deployment of the CDH platform is solved, efficient and stable automated deployment is achieved, and the system reliability and deployment efficiency are improved.

CN120234015AInactive Publication Date: 2025-07-01NANTONG JIUWEI SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510181825.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Manual deployment requires a lot of time and human resources when installing and configuring the big data management platform CDH, which can easily lead to configuration inconsistency and errors, especially when dealing with large clusters, which may lead to system performance degradation or service interruption.

Method used

Provides a CDH automated deployment system, including configuration management module, node management module, deployment execution module and monitoring feedback module. The system automatically completes deployment configuration by defining and storing parameter configuration data, detecting and managing node health assessment coefficients in real time, and monitoring and outputting alarm information in real time.

Benefits of technology

Improves the stability, reliability and deployment efficiency of the system, reduces human errors, ensures rapid response to potential problems, improves system security and stability, and thus improves deployment efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure BDA0005277296230000031
    Figure BDA0005277296230000031
Patent Text Reader

Abstract

The invention discloses a big data management platform CDH automatic deployment method, device and system, and relates to the technical field of monitoring analysis, and the method comprises a configuration management module which is used for defining and storing parameter configuration data corresponding to big data management platform deployment; the node management module is used for detecting each node in the cluster, determining a health assessment coefficient corresponding to each node according to a detection result, and managing each node based on the health assessment coefficient corresponding to each node; the deployment execution module is used for performing deployment configuration on the big data management platform according to the parameter configuration data corresponding to the configuration management module; and the monitoring feedback module is used for monitoring the deployment and configuration process in real time, collecting corresponding state data in the deployment and configuration process in real time, detecting and analyzing the state data, and outputting alarm information based on a detection and analysis result. The method has the effect of improving the deployment efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of monitoring and analysis technologies, and particularly to a method, device, and system for automated deployment of the CDH big data management platform. Background Art

[0002] With the rapid development of big data technologies, enterprises' demand for data processing capabilities has been increasing continuously. In the field of modern information technology, big data has become an important means for enterprises to gain a competitive edge. As a leading big data management platform, CDH is widely used in industries such as finance, healthcare, and retail due to its powerful data processing capabilities and rich ecosystem. CDH integrates multiple components in the Hadoop ecosystem and can support the entire process of data processing from storage to complex analysis.

[0003] In related technologies, during the installation and configuration process of CDH, manual deployment requires a large amount of time and human resources. Especially when dealing with large clusters, the configuration and debugging work may take several days or even weeks, and manual operations are prone to configuration inconsistencies and errors. Especially in the case of a large number of nodes, it may lead to a decline in system performance or service interruption, and there is room for improvement. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, this application provides a method, device, and system for automated deployment of the CDH big data management platform.

[0005] In a first aspect, this application provides an automated deployment system for the CDH big data management platform, including:

[0006] A configuration management module, used to define and store parameter configuration data corresponding to the deployment of the big data management platform;

[0007] A node management module, used to detect each node in the cluster, confirm the corresponding health assessment coefficient for each node based on the detection results, and then manage each node based on the health assessment coefficient corresponding to each node;

[0008] A deployment execution module, used to perform deployment configuration on the big data management platform according to the parameter configuration data corresponding to the configuration management module;

[0009] A monitoring and feedback module, used to monitor the deployment configuration process in real time, collect the corresponding status data during the deployment configuration process in real time, detect and analyze the status data, and output an alarm message based on the detection and analysis results.

[0010] Preferably, the configuration management module includes a template generation unit and a parameter verification unit;

[0011] The template generation unit is used to confirm the deployment requirement information corresponding to the target user, and then generate a standardized deployment template based on the deployment requirement information;

[0012] The parameter verification unit is used to analyze the input parameter configuration data before the deployment operation, obtain the data consistency evaluation coefficient and data integrity evaluation coefficient corresponding to the parameter configuration data, and then manage the parameter configuration data based on the data consistency evaluation coefficient and data integrity evaluation coefficient.

[0013] Preferably, before the deployment operation, analyze the input parameter configuration data, obtain the data consistency evaluation coefficient and data integrity evaluation coefficient corresponding to the parameter configuration data, and then manage the parameter configuration data based on the data consistency evaluation coefficient and data integrity evaluation coefficient, specifically including:

[0014] Before the deployment operation, analyze the input parameter configuration data, then confirm the total number of parameters P corresponding to the parameter configuration data, and respectively perform format detection, dependency detection, and conflict detection on each parameter, and then respectively confirm the number of parameters F with correct format, the number of parameters D with correct dependencies, and the number of parameters N without conflicts;

[0015] Through the formula Confirm the data consistency evaluation coefficient C corresponding to the parameter configuration data, where ω1, ω2, and ω3 represent weight coefficients;

[0016] Perform missingness detection, range verification, and redundancy detection on each parameter, and then respectively confirm the number of parameters M without missing, the number of parameters R with correct range, and the number of parameters E without redundancy;

[0017] Through the formula Confirm the data integrity evaluation coefficient I corresponding to the parameter configuration data, where υ1, υ2, and υ3 represent weight coefficients;

[0018] Compare the data consistency evaluation coefficient C corresponding to the parameter configuration data and the data integrity evaluation coefficient I corresponding to the parameter configuration data with the preset thresholds respectively;

[0019] If both the data consistency evaluation coefficient and the data integrity evaluation coefficient corresponding to the parameter configuration data are equal to the preset thresholds, there is no need to adjust and manage the parameter configuration data;

[0020] If the data consistency evaluation coefficient or the data integrity evaluation coefficient corresponding to the parameter configuration data is lower than the preset threshold, it is necessary to detect and correct the parameter configuration data.

[0021] Preferably, the node management module includes an automatic topology discovery unit and a node health check unit;

[0022] The automatic topology discovery unit is used to identify each node in the cluster through network scanning and node identification technology, and obtain the network topology relationship corresponding to each node;

[0023] The node health check unit is used to regularly detect the status of each node, confirm the health assessment coefficient corresponding to each node according to the detection result, and manage each node based on the health assessment coefficient corresponding to each node.

[0024] Preferably, detecting the status of each node, confirming the health assessment coefficient corresponding to each node according to the detection result, and managing each node based on the health assessment coefficient corresponding to each node specifically includes:

[0025] Construct a health assessment coefficient model corresponding to each node:

[0026]

[0027] The above-mentioned health assessment coefficient model corresponding to each node is obtained by fitting with historical data, where i represents the number corresponding to each node, i = 1, 2, 3......j, Cpui, Nci, Cpi, Ni, Si, Li respectively represent the CPU usage rate score, memory usage rate score, disk usage rate score, network status score, service status score, and log analysis score corresponding to the i-th node, μ1, μ2, μ3, μ4, μ5, μ6 represent weight coefficients, and e is the natural constant;

[0028] Real-time collect the CPU usage rate, memory usage rate, and disk usage rate corresponding to each node, and then confirm the CPU usage rate score, memory usage rate score, and disk usage rate score corresponding to each node, and monitor the network latency, network packet loss rate, and bandwidth usage situation corresponding to each node, and then confirm the network status score corresponding to each node, and check whether the key services corresponding to each node are running normally, and then confirm the service status score corresponding to each node, and analyze the log information corresponding to each node, and then confirm the log analysis score corresponding to each node;

[0029] Input the CPU usage rate score, memory usage rate score, disk usage rate score, network status score, service status score, and log analysis score corresponding to each node into the health assessment coefficient model corresponding to each node, and confirm the health assessment coefficient βi corresponding to each node;

[0030] Compare the health assessment coefficient βi corresponding to each node with the preset health assessment threshold interval [β′, β″], where β′ represents the preset first health assessment threshold and β″ represents the preset second health assessment threshold;

[0031] If there exists a node with a corresponding health assessment coefficient βi < β′, then the node is determined as a faulty node, and a first warning signal is output to the node, and the node is managed based on the first warning signal;

[0032] If there exists a node with a corresponding health assessment coefficient βi within [β′, β″], then the node is determined as a warning node, and a second warning signal is output to the node, and the node is managed based on the second warning signal, where the warning intensity of the first warning signal is greater than that of the second warning signal;

[0033] If there exists a node with a corresponding health assessment coefficient βi > β″, then the node is determined as a healthy node, and no warning signal needs to be output to the node.

[0034] Preferably, the deployment execution module includes a task scheduling unit and a parallel execution unit;

[0035] The task scheduling unit is used to generate a deployment task queue according to the parameter configuration data provided by the configuration management module, and allocate tasks according to the health assessment coefficients of each node;

[0036] The parallel execution unit is used to execute the deployment tasks in parallel with multiple threads.

[0037] Preferably, the monitoring and feedback module includes a log analysis unit and an exception handling unit;

[0038] The log analysis unit is used to perform real-time analysis on the log data generated by each node during the deployment process, identify abnormal information in the log data of each node through preset keywords, and give warnings based on the abnormal information;

[0039] The exception handling unit is used to collect the corresponding status data in real time during the deployment configuration process, detect and analyze the status data, and output an alarm message based on the results of the detection and analysis.

[0040] Preferably, detecting and analyzing the status data and outputting an alarm message based on the results of the detection and analysis specifically includes:

[0041] During a preset time period, the status data corresponding to each warning node in the deployment configuration process is collected in real time, and the health assessment coefficient βn corresponding to each warning node is extracted from the status data, where n represents the number corresponding to each warning node, n = 1, 2, 3......m, and a curve βn(t) of the health assessment coefficient of each warning node changing with time is constructed;

[0042] Through the formula the change coefficient rn of the health assessment coefficient of each warning node is confirmed, where [t1, t2] represents the preset time period;

[0043] Compare the change coefficient rn of the health assessment coefficient of each warning node with a preset change threshold r'.

[0044] If there exists a change coefficient rn of the health assessment coefficient of a warning node such that rn > r', it is determined that the warning node is abnormal, and an abnormal alarm signal is output.

[0045] In a second aspect, the present application provides a method for automated deployment of a big data management platform CDH, including the following steps:

[0046] Define and store parameter configuration data corresponding to the deployment of the big data management platform.

[0047] Detect each node in the cluster, confirm the health assessment coefficient corresponding to each node according to the detection result, and then manage each node based on the health assessment coefficient corresponding to each node.

[0048] Perform deployment configuration on the big data management platform according to the parameter configuration data corresponding to the configuration management module.

[0049] Monitor the deployment configuration process in real time, collect the corresponding status data during the deployment configuration process in real time, detect and analyze the status data, and output an alarm message based on the result of the detection and analysis.

[0050] In a third aspect, the present application provides a device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the program, it implements any one of the big data management platform CDH automated deployment systems.

[0051] In summary, the present application includes at least one of the following beneficial technical effects:

[0052] 1. The present invention provides a big data management platform CDH automated deployment system. By defining and storing parameter configuration data corresponding to the deployment of the big data management platform through the configuration management module, the accuracy of the parameter configuration is effectively ensured. The node management module evaluates and optimizes the node health in real time, the deployment execution module efficiently and automatically completes the platform deployment, and the monitoring and feedback module provides real-time status monitoring and problem alarm, thereby effectively improving the stability, reliability, and deployment efficiency of the system, reducing human errors, and ensuring a quick response to potential problems.

[0053] 2. By constructing the curve of the health assessment coefficient of each warning node changing with time, the change coefficient of the health assessment coefficient of each warning node is confirmed, and the change coefficient of the health assessment coefficient of each warning node is compared with a preset change threshold, and an abnormal warning signal is output for the abnormal warning node based on the comparison result, thereby effectively improving the system security and stability, and thus effectively improving the deployment efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0055] Figure 1 It is a schematic diagram of the system for automatic deployment of the CDH big data management platform in the embodiments of the present application.

[0056] Figure 2 It is a flowchart of the method for automatic deployment of the CDH big data management platform in the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The following will further elaborate on the present application Figure 1-2 in further detail with reference to the accompanying drawings.

[0058] Embodiment 1

[0059] The embodiments of the present application disclose a system for automatic deployment of the CDH big data management platform.

[0060] Referring to Figure 1 , a system for automatic deployment of the CDH big data management platform includes:

[0061] A configuration management module for defining and storing parameter configuration data corresponding to the deployment of the big data management platform;

[0062] A node management module for detecting each node in the cluster, confirming the health assessment coefficient corresponding to each node according to the detection result, and then managing each node based on the health assessment coefficient corresponding to each node;

[0063] A deployment execution module for performing deployment configuration on the big data management platform according to the parameter configuration data corresponding to the configuration management module;

[0064] A monitoring and feedback module for monitoring the deployment configuration process in real time, collecting the corresponding status data during the deployment configuration process in real time, detecting and analyzing the status data, and outputting an alarm message based on the detection and analysis result.

[0065] It should be noted that the configuration management module includes a template generation unit and a parameter verification unit;

[0066] The template generation unit is used to confirm the deployment requirement information corresponding to the target user, and then generate a standardized deployment template based on the deployment requirement information;

[0067] The parameter verification unit is used to analyze the input parameter configuration data before the deployment operation, and obtain the data consistency evaluation coefficient and data integrity evaluation coefficient corresponding to the parameter configuration data, and then manage the parameter configuration data based on the data consistency evaluation coefficient and data integrity evaluation coefficient.

[0068] Specifically, the specific steps for generating a standardized deployment template based on the deployment requirement information are as follows:

[0069] Collect the deployment requirement information corresponding to the target user from the target user. In the embodiment of the present application, the deployment requirement information includes, but is not limited to, cluster scale, node configuration, network topology structure, and security policy, etc. Analyze the common and individual requirements of different application scenarios, and confirm the basic structure and variable parameters of the template;

[0070] Define the variable parameters in the template, including but not limited to the number of nodes, memory allocation, storage path, etc., to ensure the flexibility of the template, and then automatically generate a standardized deployment template based on the basic structure and variable parameters of the template.

[0071] After generating the standardized deployment template, it also includes outputting the generated standardized deployment template to the target user for verification to ensure compliance with the requirements of the application scenario, and optimizing the template structure and parameter settings according to the feedback corresponding to the target user, so as to effectively improve the applicability and deployment efficiency of the template.

[0072] Furthermore, analyzing the input parameter configuration data before the deployment operation, and obtaining the data consistency evaluation coefficient and data integrity evaluation coefficient corresponding to the parameter configuration data, and then managing the parameter configuration data based on the data consistency evaluation coefficient and data integrity evaluation coefficient specifically includes:

[0073] Analyze the input parameter configuration data before the deployment operation, and then confirm the total number of parameters P corresponding to the parameter configuration data, and respectively perform format detection, dependency detection, and conflict detection on each parameter, and then respectively confirm the number of parameters F with correct format, the number of parameters D with correct dependencies, and the number of parameters N without conflicts;

[0074] Through the formula Confirm the data consistency evaluation coefficient C corresponding to the parameter configuration data, where ω1, ω2, and ω3 represent weight coefficients;

[0075] Perform missing detection, range verification, and redundancy detection on each parameter, and then respectively confirm the number of parameters M without missing, the number of parameters R with correct range, and the number of parameters E without redundancy;

[0076] Through the formula Confirm the data integrity evaluation coefficient I corresponding to the parameter configuration data, where υ1, υ2, and υ3 represent weight coefficients;

[0077] Compare the data consistency evaluation coefficient C corresponding to the parameter configuration data and the data integrity evaluation coefficient I corresponding to the parameter configuration data with the preset thresholds respectively;

[0078] If both the data consistency evaluation coefficient and the data integrity evaluation coefficient corresponding to the parameter configuration data are equal to the preset threshold, there is no need to adjust and manage the parameter configuration data;

[0079] If the data consistency evaluation coefficient or the data integrity evaluation coefficient corresponding to the parameter configuration data is lower than the preset threshold, it is necessary to detect and correct the parameter configuration data.

[0080] Specifically, the parameter configuration data includes but is not limited to node information, software version, network configuration, resource allocation, and log management, etc.; the format detection is set to verify whether the format of each parameter is correct (such as data type, data length), the dependency detection is set to confirm that the dependencies between parameters are correct (such as the rationality of the master-slave node configuration), and the conflict detection is set to check whether there are conflicts in the parameters (such as inconsistent values of the same parameter from different sources);

[0081] The missing detection is set to confirm whether there are missing parameters, the range verification is set to check whether the parameter values are within a reasonable range, the redundancy detection is set to ensure that there are no unnecessary redundant parameters. If both the data consistency evaluation coefficient corresponding to the parameter configuration data and the data integrity evaluation coefficient corresponding to the parameter configuration data are equal to 1, there is no need to adjust and manage the parameter configuration data. If it is less than 1, it is necessary to detect and correct the parameter configuration data. Through the above steps, the integrity and consistency of the parameter configuration data can be effectively evaluated, ensuring the reliability and stability of the deployment.

[0082] It should be noted that the node management module includes an automatic topology discovery unit and a node health check unit;

[0083] The automatic topology discovery unit is used to identify each node in the cluster through network scanning and node identification technology, and obtain the network topology relationship corresponding to each node;

[0084] The node health check unit is used to regularly detect the status of each node, confirm the health assessment coefficient corresponding to each node according to the detection results, and manage each node based on the health assessment coefficient corresponding to each node.

[0085] Furthermore, detecting the status of each node, confirming the health assessment coefficient corresponding to each node according to the detection results, and managing each node based on the health assessment coefficient corresponding to each node specifically include:

[0086] Construct a health assessment coefficient model corresponding to each node:

[0087]

[0088] The above-mentioned health assessment coefficient model corresponding to each node is obtained by fitting with historical data, where i represents the number corresponding to each node, Cpui, Nci, Cpi, Ni, Si, and Li respectively represent the CPU usage rate score, memory usage rate score, disk usage rate score, network status score, service status score, and log analysis score corresponding to the i-th node, μ1, μ2, μ3, μ4, μ5, and μ6 represent weight coefficients, and e is the natural constant;

[0089] Real-time collect the CPU usage rate, memory usage rate, and disk usage rate corresponding to each node, and then confirm the CPU usage rate score, memory usage rate score, and disk usage rate score corresponding to each node, and monitor the network latency, network packet loss rate, and bandwidth usage situation corresponding to each node, and then confirm the network status score corresponding to each node, and check whether the key services corresponding to each node are running normally, and then confirm the service status score corresponding to each node, and analyze the log information corresponding to each node, and then confirm the log analysis score corresponding to each node;

[0090] Input the CPU usage rate score, memory usage rate score, disk usage rate score, network status score, service status score, and log analysis score corresponding to each node into the health assessment coefficient model corresponding to each node to confirm the health assessment coefficient βi corresponding to each node;

[0091] Compare the health assessment coefficient βi corresponding to each node with the preset health assessment threshold interval [β′, β″];

[0092] If there is a node whose corresponding health assessment coefficient βi < β′, then determine the node as a faulty node, output a first warning signal to the node, and manage the node based on the first warning signal;

[0093] If the health assessment coefficient βi corresponding to a node is between [β′, β″], then the node is determined to be a warning node, and a second warning signal is output to the node, and the node is managed based on the second warning signal, where the warning intensity of the first warning signal is greater than that of the second warning signal;

[0094] If the health assessment coefficient βi corresponding to a node is > β″, then the node is determined to be a healthy node, and no warning signal needs to be output to the node.

[0095] Specifically, managing the node based on the first warning signal is to take emergency measures. For example, restarting the service, transferring the load or replacing the hardware; managing the node based on the second warning signal is preventive maintenance, such as cleaning the disk and optimizing the network.

[0096] Specifically, in the embodiment of the present application, the process of obtaining the CPU usage rate score is as follows:

[0097] Collect the CPU usage rate corresponding to the node, and compare the CPU usage rate corresponding to the node with a preset threshold range. The preset threshold range is low load 0%-50%, medium load 51%-75%, and high load 76%-100%. If the CPU usage rate corresponding to the node is in the low load, the CPU usage rate score is 1. If the CPU usage rate corresponding to the node is in the medium load, the CPU usage rate score is 0.5. If the CPU usage rate corresponding to the node is in the high load, the CPU usage rate score is 0. Similarly, the memory usage rate score and the disk usage rate score are obtained;

[0098] The process of obtaining the network status score is as follows:

[0099] Respectively confirm the preset threshold ranges in which the network latency, network packet loss rate, and bandwidth usage are located, and then respectively confirm the scores corresponding to the network latency, network packet loss rate, and bandwidth usage. Then, perform a weighted sum of the scores corresponding to the network latency, network packet loss rate, and bandwidth usage to confirm the network status score;

[0100] The process of obtaining the service status score is as follows:

[0101] Respectively confirm the service running status score and the response time score. Among them, the service running status includes normal running, partial abnormality, and serious abnormality, and the response time includes fast response, medium response, and slow response. Then, perform a weighted sum of the service running status score and the response time score to confirm the service status score;

[0102] The process of obtaining the log analysis score is as follows:

[0103] Respectively confirm the error warning score and the anomaly score. The error warnings include no error, few warnings, and frequent errors. The anomalies include no anomaly, occasional anomaly, and persistent anomaly. Then, perform a weighted sum of the error warning score and the anomaly score to confirm the log analysis score.

[0104] It should be noted that the deployment execution module includes a task scheduling unit and a parallel execution unit;

[0105] The task scheduling unit is used to generate a deployment task queue according to the parameter configuration data provided by the configuration management module, and allocate tasks according to the health assessment coefficients of each node;

[0106] The parallel execution unit is used to execute the deployment tasks in parallel with multiple threads.

[0107] Specifically, the allocation principle of allocating tasks according to the health assessment coefficients of each node is from the healthy nodes to the warning nodes in sequence, and the faulty nodes do not receive allocated tasks. Using multi-threaded parallel execution of the deployment tasks can significantly shorten the deployment time. Through the combination of task scheduling and parallel execution, it is possible to greatly improve the deployment efficiency on the premise of ensuring the deployment quality, meeting the rapid deployment requirements of large-scale clusters.

[0108] It should be noted that the monitoring and feedback module includes a log analysis unit and an anomaly handling unit;

[0109] The log analysis unit is used to perform real-time analysis on the log data generated by each node during the deployment process, identify anomaly information in the log data of each node through preset keywords, and give early warnings based on the anomaly information;

[0110] The anomaly handling unit is used to collect the corresponding status data in real time during the deployment configuration process, detect and analyze the status data, and output an alarm message based on the results of the detection and analysis.

[0111] Furthermore, detecting and analyzing the status data and outputting an alarm message based on the results of the detection and analysis specifically include:

[0112] During a preset time period, collect in real time the status data corresponding to each warning node during the deployment configuration process, and extract the health assessment coefficient βn corresponding to each warning node from the status data, where n represents the number corresponding to each warning node, n = 1, 2, 3......m, and construct a curve of the health assessment coefficient of each warning node changing with time βn(t);

[0113] Through the formula Confirm the change coefficient rn of the health assessment coefficient of each warning node, where [t1, t2] represents the preset time period;

[0114] Compare the coefficient of change rn of the health assessment coefficient of each warning node with a preset coefficient of change threshold r'.

[0115] If there exists a coefficient of change rn of the health assessment coefficient of a warning node such that rn > r', it is determined that the warning node is abnormal, and an abnormal warning signal is output.

[0116] Embodiment 2

[0117] The embodiment of the present application also discloses a method for automatic deployment of a big data management platform CDH.

[0118] Refer to Figure 2 , a method for automatic deployment of a big data management platform CDH, including the following steps:

[0119] Define and store parameter configuration data corresponding to the deployment of the big data management platform;

[0120] Detect each node in the cluster, confirm the health assessment coefficient corresponding to each node according to the detection result, and then manage each node based on the health assessment coefficient corresponding to each node;

[0121] Perform deployment configuration on the big data management platform according to the parameter configuration data corresponding to the configuration management module;

[0122] Monitor the deployment configuration process in real time, collect the corresponding status data during the deployment configuration process in real time, detect and analyze the status data, and output an alarm message based on the result of the detection and analysis.

[0123] The above content is only an example and illustration of the concept of the present invention. Those skilled in the art of the present technology can make various modifications or supplements to the described specific embodiments or use similar methods to replace them. As long as they do not deviate from the concept of the invention, they should fall within the protection scope of the present invention.

[0124] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0125] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention.

Claims

1. A big data management platform CDH automatic deployment system, characterized in that: include: Configuration management module, used to define and store parameter configuration data corresponding to the deployment of the big data management platform; The node management module is used to detect each node in the cluster, and determine the health assessment coefficient corresponding to each node according to the detection results, and then manage each node based on the health assessment coefficient corresponding to each node; A deployment execution module, used to deploy and configure the big data management platform according to the parameter configuration data corresponding to the configuration management module; The monitoring feedback module is used to monitor the deployment and configuration process in real time, collect the corresponding status data during the deployment and configuration process in real time, detect and analyze the status data, and output alarm information based on the results of the detection and analysis.

2. According to claim 1, a big data management platform CDH automatic deployment system is characterized in that: The configuration management module includes a template generation unit and a parameter verification unit; The template generating unit is used to confirm the deployment requirement information corresponding to the target user, and then generate a standardized deployment template based on the deployment requirement information; The parameter verification unit is used to analyze the input parameter configuration data before performing the deployment operation, and obtain the data consistency assessment coefficient and the data integrity assessment coefficient corresponding to the parameter configuration data, and then manage the parameter configuration data based on the data consistency assessment coefficient and the data integrity assessment coefficient.

3. The CDH automated deployment system for a big data management platform according to claim 2, characterized in that: Before performing the deployment operation, the input parameter configuration data is analyzed, and the data consistency evaluation coefficient and the data integrity evaluation coefficient corresponding to the parameter configuration data are obtained, and then the parameter configuration data is managed based on the data consistency evaluation coefficient and the data integrity evaluation coefficient, specifically including: Before the deployment operation is performed, the input parameter configuration data is analyzed to confirm the total number of parameters P corresponding to the parameter configuration data, and format detection, dependency detection and conflict detection are performed on each parameter respectively, to confirm the number of parameters F with correct format, the number of parameters D with correct dependency and the number of parameters N without conflict respectively; By formula Confirm the data consistency evaluation coefficient C corresponding to the parameter configuration data, where ω1, ω2, and ω3 are represented as weight coefficients; Perform missingness detection, range verification, and redundancy detection on each parameter, and then confirm the number of parameters M without missingness, the number of parameters R with correct range, and the number of parameters E without redundancy respectively; By formula Confirm the data integrity assessment coefficient I corresponding to the parameter configuration data, where υ1, υ2, and υ3 are weight coefficients; Compare the data consistency assessment coefficient C corresponding to the parameter configuration data and the data integrity assessment coefficient I corresponding to the parameter configuration data with the preset threshold values ​​respectively; If the data consistency assessment coefficient and the data integrity assessment coefficient corresponding to the parameter configuration data are both equal to the preset thresholds, there is no need to adjust and manage the parameter configuration data; If the data consistency assessment coefficient or data integrity assessment coefficient corresponding to the parameter configuration data is lower than the preset threshold, the parameter configuration data needs to be checked and corrected.

4. The CDH automated deployment system for a big data management platform according to claim 1, characterized in that: The node management module includes an automatic topology discovery unit and a node health check unit; The automatic topology discovery unit is used to identify each node in the cluster through network scanning and node identification technology, and obtain the network topology relationship corresponding to each node; The node health check unit is used to regularly perform status detection on each node, confirm the health assessment coefficient corresponding to each node according to the detection result, and manage each node based on the health assessment coefficient corresponding to each node.

5. The CDH automated deployment system for a big data management platform according to claim 4, characterized in that: Perform status detection on each node, determine the health assessment coefficient of each node based on the detection results, and manage each node based on the health assessment coefficient of each node, including: Construct the health assessment coefficient model corresponding to each node: The health assessment coefficient model corresponding to each node mentioned above is obtained by fitting historical data, where i represents the number corresponding to each node, i=1,2,3...j, Cpui, Nci, Cpi, Ni, Si, Li represent the CPU usage score, memory usage score, disk usage score, network status score, service status score, and log analysis score corresponding to the i-th node, μ1, μ2, μ3, μ4, μ5, and μ6 represent weight coefficients, and e is a natural constant; Collect the CPU usage, memory usage and disk usage of each node in real time, and then confirm the CPU usage score, memory usage score and disk usage score of each node, and monitor the network delay, network packet loss rate and bandwidth usage of each node, and then confirm the network status score of each node, and check whether the key services of each node are running normally, and then confirm the service status score of each node, and analyze the log information of each node, and then confirm the log analysis score of each node; Input the CPU usage score, memory usage score, disk usage score, network status score, service status score, and log analysis score corresponding to each node into the health assessment coefficient model corresponding to each node, and confirm the health assessment coefficient βi corresponding to each node; Compare the health assessment coefficient βi corresponding to each node with the preset health assessment threshold interval [β′, β″], where β′ represents the preset first health assessment threshold and β″ represents the preset second health assessment threshold; If there is a health assessment coefficient βi<β′ corresponding to a node, the node is determined to be a faulty node, a first warning signal is output to the node, and the node is managed based on the first warning signal; If there is a node whose corresponding health assessment coefficient βi is between [β′, β″], the node is determined as a warning node, and a second warning signal is output to the node, and the node is managed based on the second warning signal, wherein the warning intensity of the first warning signal is greater than that of the second warning signal; If there is a health assessment coefficient βi>β″ corresponding to a node, the node is determined to be a healthy node, and there is no need to output a warning signal to the node.

6. The CDH automated deployment system for a big data management platform according to claim 5, characterized in that: The deployment execution module includes a task scheduling unit and a parallel execution unit; The task scheduling unit is used to generate a deployment task queue according to the parameter configuration data provided by the configuration management module, and allocate tasks according to the health assessment coefficient of each node; The parallel execution unit is used to execute the deployment task in parallel with multiple threads.

7. The CDH automated deployment system for a big data management platform according to claim 5, characterized in that: The monitoring feedback module includes a log analysis unit and an exception handling unit; The log analysis unit is used to perform real-time analysis on the log data generated by each node during the deployment process, identify abnormal information in the log data of each node through preset keywords, and issue an early warning based on the abnormal information; The exception handling unit is used to collect corresponding status data during the deployment and configuration process in real time, perform detection and analysis on the status data, and output alarm information based on the results of the detection and analysis.

8. The CDH automatic deployment system for a big data management platform according to claim 7, characterized in that: Performing detection and analysis on the status data and outputting alarm information based on the detection and analysis results, specifically including: In a preset time period, the status data corresponding to each warning node in the deployment configuration process is collected in real time, and the health assessment coefficient βn corresponding to each warning node is extracted from the status data, where n represents the number corresponding to each warning node, n=1,2,3...m, and a curve βn(t) showing the change of the health assessment coefficient of each warning node over time is constructed; By formula Determine the change coefficient rn of the health assessment coefficient of each warning node, where [t1, t2] represents the preset time period; Compare the change coefficient rn of the health assessment coefficient of each warning node with the preset change threshold r′; If there is a change coefficient rn>r′ of the health assessment coefficient of the warning node, it is determined that the warning node is abnormal, and an abnormal alarm signal is output.

9. A method for automatic deployment of a big data management platform CDH, applied to a big data management platform CDH automatic deployment system as described in any one of claims 1 to 8, characterized in that: The following steps are involved: Define and store parameter configuration data corresponding to the deployment of the big data management platform; Test each node in the cluster, determine the health assessment coefficient of each node based on the test results, and manage each node based on the health assessment coefficient of each node. Deploy and configure the big data management platform according to the parameter configuration data corresponding to the configuration management module; Monitor the deployment and configuration process in real time, collect the corresponding status data during the deployment and configuration process in real time, perform detection and analysis on the status data, and output alarm information based on the results of the detection and analysis.

10. A device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements a big data management platform CDH automatic deployment system as described in any one of claims 1 to 8.