Big data cluster deployment method, apparatus, device, and medium
By deploying data acquisition, monitoring and alerting, and visualization analysis services for big data clusters, the challenges of cluster management have been solved, and automated monitoring and efficient management have been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BOE TECHNOLOGY GROUP CO LTD
- Filing Date
- 2023-05-31
- Publication Date
- 2026-07-28
AI Technical Summary
Big data cluster management is difficult, especially due to the presence of multiple servers and containers, which leads to low efficiency in monitoring and managing the cluster's operation.
This paper provides a method for deploying big data clusters. By deploying data acquisition services, monitoring and alarm services, and data visualization and analysis services to the cluster, the method enables monitoring and management of the cluster's operation and uses Grafana and Prometheus for visualization.
It improves the service deployment efficiency of big data clusters and enables automated monitoring and management of cluster operation without requiring manual operation by users.
Smart Images

Figure CN117716338B_ABST
Abstract
Description
[0001] This invention claims priority to the earlier application, which has the application number PCT / CN2022 / 106091 and is entitled "Big Data Cluster Deployment Method and Data Processing Method Based on Big Data Cluster", filed on July 15, 2022, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This invention relates to the field of computer technology, and in particular to a method, apparatus, equipment and medium for deploying big data clusters. Background Technology
[0003] With the rapid development of computer and information technology, the scale of industry application systems has expanded rapidly, and the data generated by these applications has grown exponentially. Today, industries or enterprises with data volumes reaching hundreds of terabytes (TB), tens of petabytes (PB), or even hundreds of petabytes have emerged. In order to effectively process big data, research on big data management and application methods has arisen.
[0004] In related technologies, high-speed data processing and storage are primarily achieved by deploying big data clusters and running services within them. However, big data clusters typically consist of multiple servers, each capable of containing multiple containers. Different containers can be used to provide big data cluster services for different service objectives, making the management of big data clusters particularly challenging. Therefore, a method is urgently needed to monitor the operational status of each component within a big data cluster, thereby assisting in its management. Summary of the Invention
[0005] This invention provides a method, apparatus, device, and medium for deploying big data clusters to address the shortcomings of related technologies.
[0006] According to a first aspect of the present invention, a method for deploying a big data cluster is provided, the method comprising: Deploy data acquisition services, monitoring and alarm services, and data visualization and analysis services for big data clusters; The system collects operational data from the big data cluster through a data acquisition service and uploads the operational data to a data visualization and analysis service through a monitoring and alarm service. Through data visualization and analysis services, a monitoring interface is created based on the operational data. This interface is used to graphically display the operational status of the big data cluster.
[0007] According to a second aspect of the present invention, a big data cluster deployment apparatus is provided, the apparatus comprising: The deployment module is used to deploy data acquisition services, monitoring and alarm services, and data visualization and analysis services for big data clusters. The data processing module is used to collect operational data of the big data cluster through the data acquisition service and upload the operational data to the data visualization and analysis service through the monitoring and alarm service. The display module is used to display the running data through the data visualization analysis service and the monitoring interface. The monitoring interface is used to display the running status of the big data cluster in a graphical way.
[0008] According to a third aspect of the present invention, a computing device is provided, the computing device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it performs the operations performed by the big data cluster deployment method provided in the first aspect above.
[0009] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which a program is stored, and when the program is executed by a processor, it performs the operations performed by the big data cluster deployment method provided in the first aspect above.
[0010] As described in the above embodiments, this invention deploys data acquisition services, monitoring and alarm services, and data visualization and analysis services for big data clusters. This allows the data acquisition service to collect operational data from the big data cluster, and the monitoring and alarm service to upload this data to the data visualization and analysis service. The data visualization and analysis service then displays the operational data on a monitoring page, enabling monitoring of the big data cluster's operation and assisting in its management. Furthermore, the entire service deployment process requires no manual user intervention, significantly improving the efficiency of big data cluster service deployment.
[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0013] Figure 1 This is a flowchart illustrating a big data cluster deployment method according to an embodiment of the present invention.
[0014] Figure 2 This is a schematic diagram of a deployment interface according to an embodiment of the present invention.
[0015] Figure 3 This is a schematic diagram of a physical layer monitoring interface according to an embodiment of the present invention.
[0016] Figure 4This is a schematic diagram of another physical layer monitoring interface according to an embodiment of the present invention.
[0017] Figure 5 This is a schematic diagram of another physical layer monitoring interface according to an embodiment of the present invention.
[0018] Figure 6 This is a schematic diagram of another physical layer monitoring interface according to an embodiment of the present invention.
[0019] Figure 7 This is a schematic diagram of an HDFS component monitoring interface according to an embodiment of the present invention.
[0020] Figure 8 This is a schematic diagram of a YARN component monitoring interface according to an embodiment of the present invention.
[0021] Figure 9 This is a schematic diagram of a Clickhouse component monitoring interface according to an embodiment of the present invention.
[0022] Figure 10 This is a schematic diagram of a component layer monitoring interface according to an embodiment of the present invention.
[0023] Figure 11 This is a flowchart illustrating the deployment process of a physical layer monitoring service according to an embodiment of the present invention.
[0024] Figure 12 This is a flowchart illustrating the deployment process of a component-layer monitoring service according to an embodiment of the present invention.
[0025] Figure 13 This is a schematic diagram illustrating the data flow of a monitoring service according to an embodiment of the present invention.
[0026] Figure 14 This is a flowchart illustrating the deployment process of a gateway proxy service according to an embodiment of the present invention.
[0027] Figure 15 This is a schematic diagram of an information viewing interface according to an embodiment of the present invention.
[0028] Figure 16 This is a flowchart illustrating an embodiment of the present invention for obtaining a container network address.
[0029] Figure 17 This is a schematic diagram of a web interface for managing HDFS components, as shown in an embodiment of the present invention.
[0030] Figure 18This is a schematic diagram of a web interface for managing YARN components according to an embodiment of the present invention.
[0031] Figure 19 This is a schematic diagram of a client interface for managing Hive components according to an embodiment of the present invention.
[0032] Figure 20 This is a schematic diagram of a client interface for managing Clickhouse components according to an embodiment of the present invention.
[0033] Figure 21 This is a schematic diagram of another information viewing interface according to an embodiment of the present invention.
[0034] Figure 22 This is a schematic diagram of a deployment interface according to an embodiment of the present invention.
[0035] Figure 23 This is a schematic diagram of another deployment interface according to an embodiment of the present invention.
[0036] Figure 24 This is a schematic diagram of another deployment interface according to an embodiment of the present invention.
[0037] Figure 25 This is a schematic diagram of another deployment interface according to an embodiment of the present invention.
[0038] Figure 26 This is a flowchart illustrating a virtual IP address switching process according to an embodiment of the present invention.
[0039] Figure 27 This is a flowchart illustrating a virtual IP address switching process according to an embodiment of the present invention.
[0040] Figure 28 This is a flowchart illustrating a virtual IP address switching process according to an embodiment of the present invention.
[0041] Figure 29 This is a schematic diagram illustrating a virtual IP address switching result according to an embodiment of the present invention.
[0042] Figure 30 This is a flowchart illustrating a management node migration process according to an embodiment of the present invention.
[0043] Figure 31 This is a schematic diagram illustrating the principle architecture of a big data cluster deployment method according to an embodiment of the present invention.
[0044] Figure 32 This is a block diagram illustrating a big data cluster deployment device according to an embodiment of the present invention.
[0045] Figure 33 This is a schematic diagram of the structure of a computing device according to an embodiment of the present invention. Detailed Implementation
[0046] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.
[0047] To aid in understanding this invention, the technical terms involved in this invention will first be introduced.
[0048] Grafana is an open-source monitoring data analysis and visualization suite used to create monitoring dashboards for visualizing monitoring data. For example, Grafana can be used to visualize and analyze time-series data obtained from infrastructure and application data analysis. Furthermore, it can be applied to other areas requiring data visualization and analysis. Grafana helps query, visualize, alert, and analyze metrics and data, providing a fast and flexible visualization experience that allows users to visualize data in the way they need.
[0049] Prometheus is an open-source monitoring and alerting system based on a time-series database. It's a leading monitoring solution, a one-stop monitoring and alerting platform with few dependencies, comprehensive functionality, and high integration and ease of use. Grafana can connect to Prometheus data sources to display monitoring data graphically. Prometheus supports data collection through various exporters and data reporting via Pushgateway. Prometheus's performance is sufficient to provide monitoring and alerting capabilities for clusters of tens of thousands of machines.
[0050] Exporter: This is a general term for a type of data collection component in Prometheus, primarily used for data acquisition. In addition to collecting data from the target, the Exporter can also convert the collected data into a format supported by Prometheus. Unlike traditional data collection components, it does not send data to the Central Processing Unit (CPU), but instead waits for the central server to actively retrieve it. The Exporter is the object monitored by Prometheus. The Exporter's interface can be exposed to the Prometheus service in the form of a Hypertext Transfer Protocol (HTTP) service. The Prometheus service can obtain the monitoring data to be collected by accessing the interface provided by the Exporter. There are many common Exporters, such as Node-exporter, Mysqld-exporter, and HAproxy-exporter, which support service monitoring of service providers such as HAProxy, StatsD, Graphite, and Redis.
[0051] Node-exporter is used to collect server-level operational metrics. For example, it can be used to collect basic monitoring data such as system load average (Loadavg), file system (Filesystem), and memory management (Meminfo). Additionally, it can be used to monitor the usage of components such as CPU, memory, disk, and I / O on container servers. To collect data using Node-exporter, you first need to deploy Node-exporter on the service from which you want to collect data.
[0052] Apache Knox: The primary goal of the Apache Knox project is to provide access to Apache Hadoop through an HTTP resource broker. The Apache Knox gateway can extend the reach of Apache Hadoop services to users outside the Hadoop cluster without compromising Hadoop security.
[0053] Knox gateways can provide security for multiple Hadoop clusters, offering the following advantages: (1) Simplified access: By encapsulating Kerberos (a computer network authorization protocol used to securely authenticate personal communications in insecure networks) into the cluster, the service provision scope of Hadoop’s Representational State Transfer (REST) service and HTTP service is extended.
[0054] (2) Improve security: Expose Hadoop’s REST / HTTP services without revealing network details, and provide out-of-the-box use via the Secure Socket Layer (SSL) protocol.
[0055] (3) Centralized control: Centralized security guarantees for REST service APIs (Application Programming Interfaces) are implemented, and requests are routed to multiple Hadoop clusters.
[0056] (4) Enterprise integration: Supports multiple authentication systems such as Lightweight Directory Access Protocol (LDAP), Active Directory (AD), Single Sign On (SSO), and Security Assertion Markup Language (SAML).
[0057] The big data cluster deployment method provided by this invention will be described in detail below.
[0058] This invention provides a method for deploying a big data cluster. After deploying servers and containers in a big data cluster via drag-and-drop node deployment, it automatically deploys corresponding monitoring services (including data acquisition services, monitoring and alarm services, and data visualization and analysis services) to the deployed servers and containers. This allows for monitoring of the big data cluster's operational status through the deployed monitoring services, and the monitoring results can be visualized, assisting in the operation and maintenance of the big data cluster and lowering the technical threshold for its operation and maintenance. Furthermore, the solution provided by this invention enables automated deployment of monitoring services, improving the efficiency of service deployment in the big data cluster.
[0059] The above is merely an exemplary description of the application scenarios of the present invention and does not constitute a limitation on the application scenarios of the present invention. In more possible implementations, the big data cluster deployment method provided by the present invention can also be used to deploy monitoring services in other service clusters, and the present invention does not limit this.
[0060] The above-mentioned big data cluster deployment method can be executed by computing devices, which can be servers, such as one server, multiple servers, server clusters, etc. This invention does not limit the type or number of computing devices.
[0061] To make it easier to understand, we will first explain the process of deploying servers and containers in a big data cluster by dragging and dropping nodes.
[0062] In some embodiments, relevant technical personnel can build the basic network environment required for constructing a big data cluster on any service according to actual technical needs. The server with the basic network environment can be used as a management server, and subsequently, servers and / or containers can be deployed on the basis of the management server to realize the deployment of the big data cluster.
[0063] Optionally, relevant technical personnel can add servers to the big data cluster according to their actual needs to build a big data cluster that includes multiple servers.
[0064] In some embodiments, a deployment interface may be provided to offer deployment functionality for the big data cluster. Optionally, a "Add Physical Pool" control can be set within the deployment interface, allowing a server to be added to the big data cluster. For example, a "Add Physical Pool" area can be set within the deployment interface to place the "Add Physical Pool" control within the deployment resource pool area. See also... Figure 2 , Figure 2 This is a schematic diagram of a deployment interface according to an embodiment of the present invention, such as... Figure 2 As shown, the deployment interface is divided into a node creation area, a temporary resource pool area, and a deployment resource pool area. The "Add Physical Pool" button in the deployment resource pool area is the control for adding a new physical pool. By using the "Add Physical Pool" button, the server can be added to the big data cluster.
[0065] By adding a physical pool control in the deployment interface, users can add servers to the big data cluster according to their actual technical needs, so that the created big data cluster can meet the technical requirements and ensure the smooth operation of subsequent data processing.
[0066] In one possible implementation, adding a physical pool can be achieved through the following process, thereby adding a server to the big data cluster: Step 1: In response to the trigger operation of adding a physical pool control, display the Add Physical Pool interface, which includes an ID retrieval control and a password retrieval control.
[0067] See Figure 3 , Figure 3 This is a schematic diagram of an interface for adding a physical pool according to an embodiment of the present invention. After the new physical pool control is triggered, it can be displayed on the visual interface as shown below. Figure 3 The interface for adding a physical pool shown in the image has two input boxes: one for obtaining the IP address and the other for obtaining the password.
[0068] Step 2: Obtain the server identifier corresponding to the physical pool to be added through the identifier retrieval control, and obtain the password to be verified through the password retrieval control.
[0069] In one possible implementation, relevant technical personnel can enter the server identifier of the server to be added to the big data cluster in the identifier acquisition control and enter a pre-set password in the password acquisition control, so that the computing device can obtain the server identifier corresponding to the physical pool to be added through the identifier acquisition control and obtain the password to be verified through the password acquisition control.
[0070] Optionally, after obtaining the server identifier corresponding to the physical pool to be added through the identifier acquisition control and the password to be verified through the password acquisition control, the password to be verified can be verified.
[0071] Step 3: If the password to be verified passes, display the physical pools to be added in the deployment resource pool area.
[0072] By setting an identifier retrieval control, users can input the server identifier of the server to be added to the big data cluster, thus meeting their customization needs. By setting a password retrieval control, users can input a password to be verified in the password retrieval interface. The user's identity is verified based on the password to determine whether the user has the right to add the server to the big data cluster, thereby ensuring the security of the big data cluster deployment process.
[0073] Once the password verification is successful, the server corresponding to the obtained server identifier can be added to the big data cluster.
[0074] The above process completes the construction of the hardware environment for the big data cluster, resulting in a big data cluster including at least one server. Containerization can then be performed on this server to enable the big data cluster to provide big data processing capabilities to users.
[0075] In some embodiments, the deployment interface includes a node creation area, which includes node creation controls and at least one big data component. The big data component includes at least an HDFS component, a YARN component, a Clickhouse component, and a Hive component.
[0076] The HDFS component can be used to provide data storage functionality. In other words, to provide data storage functionality to users, it is necessary to deploy the containers corresponding to the nodes of the HDFS component in the big data cluster so as to provide distributed storage services for users' data through the deployed containers to meet their needs.
[0077] YARN components can be used to provide data analysis functions. That is, to provide data analysis functions to users, it is necessary to deploy the containers corresponding to the nodes of the YARN component in the big data cluster so that the containers corresponding to the nodes of the HDFS component can obtain data from the containers corresponding to the nodes of the HDFS component, and perform data analysis based on the obtained data to meet the user's data analysis needs.
[0078] The Hive component can convert the data stored in the container corresponding to the node of the HDFS component into a queryable data table, so that data querying and processing can be performed based on the data table to meet the user's data processing needs.
[0079] It's important to note that while both YARN and Hive components can provide data analysis capabilities, they differ in that using YARN requires developing a series of codes to perform data processing tasks after the data processing tasks are submitted to the YARN component. In contrast, using Hive simplifies the process entirely by using Structured Query Language (SQL) statements.
[0080] ClickHouse is a columnar storage database that can meet users' needs for storing large amounts of data. Compared to commonly used row-based storage databases, ClickHouse offers faster read speeds. Furthermore, ClickHouse allows for partitioned storage of data, enabling users to retrieve only one or a few partitions for processing based on their specific needs, rather than accessing all data in the database. This reduces the data processing load on computing devices.
[0081] In some embodiments, at least one big data component may be displayed on the deployment interface so that users can select from the big data components displayed on the deployment interface based on actual technical needs. When any big data component is selected, in response to the triggering operation of the node creation control, the node to be deployed corresponding to the selected big data component is displayed in the temporary resource pool area.
[0082] It should be noted that different components include different nodes. Optionally, the HDFS component includes NameNode (nn) node, DataNode (dn) node, and SecondaryNameNode (sn) node; the YARN component includes ResourceManager (rm) node and NodeManager (nm) node; the Hive component includes Hive (hv) node; and the Clickhouse component includes Clickhouse (ch) node.
[0083] Based on the relationships between the components and nodes described above, the computing device can display the corresponding nodes in the temporary resource pool area as nodes to be deployed, depending on the selected big data component. Users can drag and drop the nodes to be deployed from the temporary resource pool area to the physical pool in the deployment resource pool area, enabling the deployment of containers in the big data cluster through node drag-and-drop operations.
[0084] In some embodiments, in response to a drag-and-drop operation on a node to be deployed in the temporary resource pool area, the node to be deployed can be displayed in the physical pool in the deployment resource pool area of the deployment interface.
[0085] The deployment resource pool area may include at least one physical pool. In one possible implementation, when displaying a node to be deployed in a physical pool in the deployment resource pool area of the deployment interface in response to a drag operation on a node to be deployed in the temporary resource pool area, for any node to be deployed, the node to be deployed may be displayed in the physical pool indicated when the drag operation ends in response to the drag operation on the node to be deployed.
[0086] By providing drag-and-drop functionality for the nodes displayed in the deployment interface, users can drag and drop each node to be deployed into the corresponding physical pool according to their actual technical needs, thereby meeting their customized requirements.
[0087] It should be noted that after dragging all the nodes to be deployed from the temporary resource pool area to the deployment resource pool area, you can respond to the start deployment operation in the deployment interface and deploy the container corresponding to the node to be deployed on the server corresponding to the physical pool where the node is located.
[0088] The above embodiments mainly introduce the process of adding physical pools and deploying the corresponding containers of nodes in physical pools. Optionally, the deployment resource pool area can also be equipped with a physical pool deletion control.
[0089] In cases where the deployment of a resource pool includes the deletion of a physical pool control, technical personnel can delete the physical pool by deleting the physical pool control.
[0090] In one possible implementation, a physical pool corresponds to a delete physical pool control. Relevant technicians can trigger the delete physical pool control corresponding to any physical pool, and the computing device can respond to the triggering operation of any delete physical pool control by no longer displaying the physical pool corresponding to the triggered delete physical pool control in the deployment resource pool area.
[0091] Still with Figure 2 Taking the deployment interface shown as an example, as Figure 2 The deployment interface shown has an "×" button in the upper right corner of each physical pool displayed in the deployment resource pool area. This button is the delete physical pool control. Users can delete the corresponding physical pool by triggering any "×" button.
[0092] By setting a delete physical pool control in the deployment interface, users can delete any physical pool according to their actual needs, thereby removing the server corresponding to that physical pool from the big data cluster. This meets users' technical requirements and is easy to operate. Users only need a simple control to trigger the operation to complete the modification of the big data cluster, which greatly improves the efficiency of operation.
[0093] It should be noted that when a service is removed from the big data cluster, the containers deployed on the server will also be removed from the big data cluster.
[0094] After introducing the basic deployment process of the big data platform, the big data cluster deployment method provided by this invention will be described in detail below.
[0095] See Figure 1 , Figure 1 This is a flowchart illustrating a big data cluster deployment method according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes: Step 101: Deploy data acquisition services, monitoring and alarm services, and data visualization and analysis services for the big data cluster.
[0096] A big data cluster can include multiple servers, each of which can deploy at least one container to provide big data services. The containers deployed in the big data cluster are created through drag-and-drop node operations in the deployment interface, which provides deployment functionality for the big data cluster, including but not limited to deploying servers within the cluster and deploying containers on those servers.
[0097] Optionally, the containers deployed on the server may include containers corresponding to the Hadoop Distributed File System (HDFS) component, containers corresponding to the Yet Another Resource Negotiator (YARN) component, containers corresponding to the database tool (Clickhouse) component, containers corresponding to the data warehouse tool (Hive) component, etc. The present invention does not limit the specific type of the deployed containers.
[0098] See Figure 2 , Figure 2 This is a schematic diagram of a deployment interface according to an embodiment of the present invention, such as... Figure 2 As shown, the deployment interface offers HDFS, YARN, Hive, and Clickhouse components. Users can select the appropriate component to deploy its corresponding container on the big data cluster. It's important to note that different types of components correspond to different nodes. For example, the HDFS component includes NameNode (nn), DataNode (dn), and SecondaryNameNode (sn) nodes; the YARN component includes ResourceManager (rm) and NodeManager (nm) nodes; the Hive component includes Hive (hv) nodes; and the Clickhouse component includes Clickhouse (ch) nodes. After selecting a component, the nodes included in that component will be displayed in the temporary resource pool area of the deployment interface. Users can adjust the nodes to be deployed in the temporary resource pool area (including adding and deleting nodes). Users can drag and drop nodes from the temporary resource pool to various physical pools in the deployment resource pool (one physical pool corresponds to one server in a big data cluster) to deploy the containers corresponding to the nodes in the physical pool on the corresponding server, achieving the goal of containerized deployment of the big data cluster through drag-and-drop operations.
[0099] After containerizing the big data cluster through the deployment interface, the deployed big data cluster can provide users with basic data processing functions of a big data processing platform. For example, the deployed big data cluster can be delivered to users in the form of a big data processing platform. Users can configure computing tasks in the big data processing platform, such as setting the data required to execute computing tasks and the computing instructions (such as computing rules) based on the execution of computing tasks. This allows the corresponding data to be obtained through a generally deployed big data cluster, and the data to be processed to obtain processing results that meet the user's needs.
[0100] In addition to deploying basic big data clusters to provide basic data processing functions, data acquisition services, monitoring and alarm services, and data visualization and analysis services can also be deployed for big data clusters. This allows for the monitoring of the operation of big data clusters through the deployed data acquisition services, monitoring and alarm services, and data visualization and analysis services, thus providing strong support for the operation and maintenance of big data clusters.
[0101] Among them, the data acquisition service is used to collect the operational data of service providers in the big data cluster, the monitoring and alarm service is used to transfer data between the data acquisition service and the data visualization and analysis service, and the data visualization and analysis service is used to display the reported operational data graphically.
[0102] It should be noted that the data acquisition service can collect not only the operational data of service providers in the big data cluster, but also production data generated by users during actual production processes, and storage data stored on local servers or containers, including but not limited to various types of production data, multimedia data, tensor data, and structured data. For example, in a production park-production line scenario, the data acquisition service can be used to collect production data generated by various processes, stations, raw materials, and algorithms on the production line; in a residential park-commercial / cultural tourism scenario, the data acquisition service can be used to collect multimedia data, tensor data, and structured data generated based on services such as passenger flow statistics, human recognition, hotspot tracking, and restricted area statistics. Among these, multimedia data can be video frame data collected by cameras deployed in the park, tensor data can be vectorized data generated based on facial recognition, and structured data can be various types of data used for big data computing and analysis. This invention does not limit the specific data types.
[0103] Optionally, the data acquisition service can be provided by deploying the Exporter service (i.e., deploying the Exporter component) in the big data cluster, the monitoring and alarm service can be provided by deploying the Prometheus service (i.e., deploying the Prometheus system) in the big data cluster, and the data visualization and analysis service can be provided by deploying the Granafa service (i.e., deploying the Granafa plugin) in the big data cluster.
[0104] Step 102: Collect the operational data of the big data cluster through the data acquisition service, and upload the operational data to the data visualization and analysis service through the monitoring and alarm service.
[0105] Step 103: Through data visualization analysis services, the monitoring interface is based on the running data display. The monitoring interface is used to display the running status of the big data cluster in a graphical way.
[0106] This invention deploys data acquisition services, monitoring and alarm services, and data visualization and analysis services for big data clusters. The data acquisition service collects operational data from the big data cluster, and the monitoring and alarm service uploads this data to the data visualization and analysis service. The data visualization and analysis service then displays the operational data on a monitoring page, enabling monitoring of the big data cluster's operation and assisting in its management. Furthermore, the entire service deployment process requires no manual user intervention, significantly improving the efficiency of big data cluster service deployment.
[0107] After introducing the basic implementation process of the big data cluster deployment method provided by the present invention, the following describes various optional embodiments of the present invention.
[0108] It should be noted that monitoring and alerting services and data visualization and analysis services can be used as general-purpose underlying services. Once deployed, they can provide corresponding services to the entire big data cluster. Therefore, by deploying Prometheus and Grafana services once in a big data cluster, monitoring and alerting services and data visualization and analysis services can be provided to the entire big data cluster through the deployed Prometheus and Grafana services.
[0109] For different types of service objects, the specific types of data acquisition services required may differ. Optionally, the data acquisition service may include a first data acquisition service and a second data acquisition service. The first data acquisition service can be used to collect server operation data, and the second data acquisition service can be used to collect container operation data. Therefore, the first data acquisition service can also be called a physical layer data acquisition service, and the second data acquisition service can also be called a component layer data acquisition service.
[0110] When deploying data acquisition services, Exporter components matching the type of the service object can be deployed for each service object. If the service object is a server, the Exporter component corresponding to the first data acquisition service can be deployed for it; for example, the Node-exporter component can be deployed for the server. If the service object is a container, the Exporter component corresponding to the second data acquisition service can be deployed for it. However, it should be noted that the type of Exporter component corresponding to different types of containers is different. For example, the Exporter component used to provide data acquisition services for containers corresponding to the HDFS component and the YARN component can be Hadoop-exporter, and the Exporter component used to provide data acquisition services for the Clickhouse component can be Clickhouse-exporter.
[0111] It should be noted that for the first data acquisition service used to collect server runtime data, a corresponding first data acquisition service needs to be deployed on each server, specifically for collecting runtime data from the corresponding server. However, for the second data acquisition service used to collect container runtime data, the data acquisition service for each type of container only needs to be deployed once on one server in the big data cluster. Then, the same type of containers in the entire big data cluster can collect runtime data through the deployed second data acquisition service. For example, if the big data cluster includes 3 servers, and each server has a container corresponding to an HDFS component, then the Hadoop-exporter service (that is, the Hadoop-exporter component) can be deployed on one of the servers. This Hadoop-exporter service can then be used to collect runtime data from the containers corresponding to the HDFS components on these 3 servers.
[0112] After a preliminary introduction to data acquisition services, monitoring and alarm services, and data visualization and analysis services, the following section describes the detailed process of deploying these services in a big data cluster.
[0113] In some embodiments, step 101, when deploying data acquisition services, monitoring and alarm services, and data visualization and analysis services for a big data cluster, can be achieved through the following steps: Step 1011: Based on the service providers currently deployed in the big data cluster, deploy data acquisition services, monitoring and alarm services, and data visualization and analysis services for the current service providers in the big data cluster.
[0114] It should be noted that among the multiple servers included in a big data cluster, there can be a management server. Once the management server is deployed, the foundation of the big data cluster platform is established; that is, the big data infrastructure service platform has been built, allowing for the further deployment of various services. This management server can serve as the core device for providing big data cluster services. Therefore, for step 1011, when deploying data acquisition services, monitoring and alarm services, and data visualization and analysis services for the big data cluster, it can be achieved as follows: Based on the servers and containers currently included in the big data cluster, deploy data acquisition services on each server of the big data cluster, and deploy monitoring and alarm services and data visualization and analysis services on the management server of the big data cluster.
[0115] Optionally, when deploying monitoring and alerting services and data visualization and analysis services on the management server of the big data cluster, the Prometheus service can be deployed on the management server to enable the deployment of monitoring and alerting services on the management server. In addition, the Grafana service can be deployed on the management server of the big data cluster to enable the deployment of data visualization and analysis services on the management server.
[0116] When deploying data acquisition services on each server in a big data cluster, a first data acquisition service and a second data acquisition service can be deployed on the management server, and the first data acquisition server can be deployed on every server other than the management server. Specifically, both the first and second data acquisition services are deployed on the management server. The first data acquisition service is used to collect operational data from the management server, while the second data acquisition service is used to provide data acquisition services to various containers within the big data cluster.
[0117] Optionally, for servers in a big data cluster, runtime data may include disk space usage data, network traffic data, CPU usage data, memory usage data, etc.; for containers in a big data cluster, runtime data may include big data service usage data and container status data, etc. The present invention does not limit the specific type of runtime data.
[0118] In one possible implementation, when deploying the first data acquisition service on the management server, a Node-exporter service can be deployed on the management server to collect runtime data from the management server. When deploying the second data acquisition service on the management server, a Hadoop-exporter service can be deployed on the management server to collect runtime data from containers corresponding to the HDFS component and the YARN component in the big data cluster; alternatively, a Clickhouse-exporter service can be deployed on the management server to collect runtime data from containers corresponding to the Clickhouse component in the big data cluster.
[0119] It should be noted that, in order to ensure that the deployed monitoring services (including monitoring and alarm services, data visualization and analysis services, and data acquisition services) can run smoothly, the deployed services can be debugged after deployment to ensure that the deployed services can be used normally.
[0120] In addition, the Prometheus data source can be pre-configured in the Grafana service to ensure that the deployed Grafana service and the Prometheus service can communicate with each other, so that the Prometheus service can pass the data to the Grafana service after fetching data from various exporter services.
[0121] In addition, Grafana offers a variety of available page templates. Technical personnel can download the template that best suits their specific technical needs, creating monitoring pages tailored to those needs. These customized pages can then be imported into the Grafana service. For example, the URL of a pre-made monitoring page can be imported directly into Grafana. Furthermore, a mirror image of the Grafana service with the imported monitoring pages can be generated. This image image can then be run to start the Grafana service, ensuring that the monitoring pages are displayed correctly based on the imported URLs.
[0122] Step 1012: In response to the update operation of the deployed service provider in the big data cluster, update the data acquisition service, monitoring and alarm service and data visualization and analysis service deployed in the big data cluster. The update operation of the deployed service provider includes deleting the deployed service provider and deploying a new service provider.
[0123] It should be noted that the monitoring and alarm service has a corresponding first configuration file, which records which data collection services the data to be transmitted to the data visualization and analysis service originates from. Since update operations can include deleting an existing service provider and deploying a new one, the service deployment process for these two different types of update operations will be described below.
[0124] When the update operation on an already deployed service provider is to deploy a new service provider, step 1012 can include the following implementation: In one possible implementation, in response to the addition of a new server in the big data cluster, a first data acquisition service is deployed on the new server, and the first configuration file corresponding to the monitoring and alarm service is modified based on the new service provider, so that the monitoring and alarm service and data visualization analysis service already deployed in the big data cluster can provide services to the new server.
[0125] Optionally, to distinguish the first data acquisition services deployed on the management server and the newly added server, the first data acquisition services deployed on the management server and the newly added server can be named separately to differentiate them. For example, the naming rule Node-exporter-ip can be used, and the IP field in the named service name can be used to distinguish different first data acquisition services.
[0126] Additionally, it should be noted that the monitoring and alarm service can correspond to a primary configuration file. For example, the Prometheus service can maintain a primary configuration file. This primary configuration file can be used to configure the service information of the data acquisition service corresponding to the data to be transmitted to the data visualization and analysis service. Therefore, after deploying the primary data acquisition service on a new server, the primary configuration file maintained by the Prometheus service can be modified to update it. This ensures that the updated primary configuration file includes the service configuration information of the newly deployed primary data acquisition service, indicating that the running data collected by the newly deployed primary data acquisition service is also to be transmitted to the data visualization and analysis service.
[0127] The service configuration information of the first data acquisition service may include various types of information such as service name and the server to which the service belongs, and this invention does not limit this information.
[0128] In another possible implementation, in response to the addition of a new container in the big data cluster, the first configuration file corresponding to the monitoring and alarm service is modified based on the new server, so that the second data acquisition service, monitoring and alarm service, and data visualization and analysis service already deployed in the big data cluster can provide services to the new container.
[0129] It should be noted that, as described in the above embodiments, for the same type of component on different servers, the corresponding containers only need to have the second data acquisition service deployed once to collect their runtime data. Furthermore, the first configuration file maintained by the Prometheus service can also record the service configuration information of the second data acquisition service. Optionally, the service configuration information of the second data acquisition service may include various types of information such as the container information (e.g., container name, container type) of the container it operates on; this invention does not limit this information.
[0130] Therefore, when adding a new container in a big data cluster, the second data acquisition service deployed on the management server can be updated directly based on the container information of the new container, so that the second data acquisition service deployed on the management server can provide data acquisition services for the new container. This is equivalent to deploying the second data acquisition service on the new container without having to redeploy the second data acquisition service.
[0131] The above embodiment illustrates the deployment of a second data acquisition service on a management server. In more possible implementations, a server can be randomly selected from the servers included in the big data cluster to deploy the second data acquisition service. It should be noted that different types of second data acquisition services can be deployed on different servers, but for the same type of second data acquisition service, it only needs to be deployed once on one server. However, regardless of which server the second data acquisition service is deployed on, the first configuration file maintained by the monitoring and alarm service deployed on the management server must be updated.
[0132] It should be noted that, due to the difference in nature between the physical layer and the component layer, physical layer monitoring can be deployed during the deployment of the basic service platform, but component layer monitoring can only be deployed after the big data components have been fully deployed. That is, the process of deploying the first data acquisition service for a new server can be performed concurrently with the deployment of the new server, while the process of deploying the second data acquisition service for a new container can be performed after the new container has been deployed.
[0133] Therefore, a status query interface can be provided to query the current deployment status of components in the big data cluster. If the container corresponding to a certain big data component has been deployed, the URL address of the pre-made image file for the container corresponding to that big data component can be returned; otherwise, an empty value can be returned so that the current component deployment can be determined based on the return value.
[0134] In other embodiments, when the update operation on the deployed service provider is to delete the deployed service provider, step 1012 may include the following implementation: The above embodiments mainly introduce the service deployment method when adding a new service provider. In more possible implementation methods, servers and / or containers already deployed in the big data cluster can also be deleted.
[0135] In one possible implementation, in response to the deletion operation of a server already deployed in the big data cluster, the first data acquisition service corresponding to the server is deleted, and the first configuration file corresponding to the monitoring and alarm service is modified based on the deleted server, so that the monitoring and alarm service and data visualization analysis service already deployed in the big data cluster no longer provide services to the deleted server.
[0136] Optionally, in response to the deletion operation of a server in the big data cluster, the first data acquisition service deployed for that server can be deleted, and the first configuration file maintained by the monitoring and alarm service can be modified to delete the service configuration information of the first data acquisition service corresponding to that server from the first configuration file, so that when data is obtained from the deployed data acquisition service based on the updated first configuration file, data will not be obtained from the deleted server.
[0137] In another possible implementation, in response to the deletion of a container already deployed in the big data cluster, the first configuration file corresponding to the monitoring and alarm service is modified based on the deleted server, so that the second data acquisition service, monitoring and alarm service, and data visualization and analysis service already deployed in the big data cluster no longer provide services to the deleted server.
[0138] Optionally, in response to the deletion operation of a container in the big data cluster, the first configuration file maintained by the monitoring and alarm service can be modified to delete the service configuration information of the second data collection service corresponding to the container from the first configuration file, so that when obtaining data collected from the deployed data collection service, data will not be obtained from the deleted container.
[0139] Optionally, different functional interfaces can be provided to implement different types of service deployment processes. For example, an initial physical pool interface can be provided to provide the deployment function of the first data collection service for newly added servers; a delete physical pool interface can be provided to delete the first data collection service corresponding to a server when it is deleted; and a deployment interface can also be provided to deploy the second data collection service for newly added containers, and / or to delete the second data collection service corresponding to a container when it is deleted.
[0140] It should be noted that in order to automatically deploy the data acquisition service, monitoring and alarm service, and data visualization and analysis service on any node, and to ensure that the three can recognize each other and exchange data, the data acquisition service, monitoring and alarm service, and data visualization and analysis service need to be deployed in the same OverLay network. That is, the Prometheus server, Grafana service, and all Exporters need to be set up in the same OverLay network.
[0141] After setting up the Prometheus server, Grafana service, and all Exporters on the same Overlay network, the deployed services can be named according to pre-defined naming rules when starting the container. This allows the named services to be recorded in the first configuration file, so that the Prometheus service can crawl data based on the first configuration file.
[0142] For example, when deploying Node-exporter, each Node-exporter service deployed on each machine can be named according to the naming rule of node-exporter-server name (such as ip1, ip2, ...). The naming result can be recorded in the first configuration file of the Prometheus service, allowing the Prometheus service to access the data from each exporter. For services like Prometheus and Gafana, which only need to be deployed once, they can be directly named Prometheus and Grafana. When configuring the data source in Grafana, you only need to configure "http: / / prometheus:9090", and the Grafana service can then read the data from the Prometheus service.
[0143] Overlay networks can virtualize network connections between multiple hosts and run applications within this virtual network to isolate applications from the underlying network. By deploying data acquisition services, monitoring and alarm services, and data visualization and analysis services in an overlay network, a natural protection mechanism can be formed to provide additional security protection for big data clusters and enhance their security.
[0144] After completing the deployment of data acquisition service, monitoring and alarm service and data visualization and analysis service through the above embodiments, step 102 can be used to collect the operation data of the big data cluster through the data acquisition service and upload the operation data to the data visualization and analysis service through the monitoring and alarm service.
[0145] Optionally, the server's operating data can be collected through the deployed first data acquisition service, and the container's operating data can be collected through the deployed second data acquisition service. The data collected by the first and second data acquisition services can be transmitted to the data visualization and analysis service through the monitoring and alarm service, so that the data visualization and analysis service can display the operating data on the monitoring interface through step 103.
[0146] In some embodiments, step 103, when monitoring the interface based on the running data display, can be implemented in the following way: In one possible implementation, a physical layer monitoring interface is displayed based on the operational data of the servers in the big data cluster. This interface is used to display the operational data of the servers within the big data cluster. Optionally, the physical layer monitoring interface can also be referred to as a physical layer dashboard.
[0147] See Figure 3 , Figure 3 This is a schematic diagram of a physical layer monitoring interface according to an embodiment of the present invention. In a big data cluster where only a management server is deployed, it can display... Figure 3 The monitoring interface shown is as follows. Figure 3 As shown, in a big data cluster that only includes the management server, the number of monitored servers is displayed as 1, and it is possible to monitor... Figure 3 The physical layer monitoring interface shown displays disk space usage data (i.e., disk space utilization and disk read / write capacity) and network traffic data (i.e., network traffic) of the management server.
[0148] In addition, in such Figure 3 The physical layer monitoring interface shown can also be equipped with a page-turning function control, which allows users to view more types of operational data. Figure 3 The physical layer monitoring interface shown triggers this function control, and the computing device can then display the following: Figure 4 The physical layer monitoring interface shown is available in [reference]. Figure 4 , Figure 4 This is a schematic diagram of another physical layer monitoring interface according to an embodiment of the present invention. Figure 4 The physical layer monitoring interface shown displays CPU usage data (i.e., CPU utilization), memory usage data (i.e., memory utilization), and disk space usage data (i.e., disk space of each partition).
[0149] The above Figure 3 and Figure 4 The image shown is only the display format of the physical layer monitoring interface when the big data cluster includes only the management server. If a new server is added to the big data cluster, in addition to the management server, the physical layer monitoring interface will display as shown below. Figure 5 and Figure 6 The style shown.
[0150] Figure 5 This is a schematic diagram of another physical layer monitoring interface according to an embodiment of the present invention, such as... Figure 5As shown, when a new server is added to the big data cluster, the number of monitored servers will change from 1 to 2. Furthermore, two curves will be displayed for disk space utilization, disk read / write capacity, and network traffic, with each curve corresponding to one server, in order to monitor the operation of two servers.
[0151] Figure 6 This is a schematic diagram of another physical layer monitoring interface according to an embodiment of the present invention. Figure 6 It can be based on, for example Figure 5 The page-turning operation shown in the physical layer monitoring interface is obtained, in such... Figure 6 In the physical layer monitoring interface shown, the curves displayed for CPU utilization and memory utilization will both become two lines, with each curve corresponding to one server, in order to monitor the operation of two servers.
[0152] In another possible implementation, a component-level monitoring interface is displayed based on the runtime data of containers within the big data cluster. This interface is used to display the runtime data of the containers in the big data cluster. Optionally, the component-level monitoring interface can also be called a component-level dashboard.
[0153] Optionally, the component layer monitoring interface can include various types to display the running status of containers corresponding to different types of big data components. For example, the component layer monitoring interface can include an HDFS component monitoring interface (or HDFS monitoring dashboard) to display the running status of containers corresponding to HDFS components, a YARN component monitoring interface (or YARN monitoring dashboard) to display the running status of containers corresponding to YARN components, and a Clickhouse component monitoring interface (or Clickhouse monitoring dashboard) to display the running status of containers corresponding to Clickhouse components.
[0154] It should be noted that, due to the relationship between the Hive component and the HDFS and YARN components, the Hive component generally will not have problems running when the HDFS and YARN components are running normally. Therefore, under normal circumstances, the Hive component monitoring interface for the container corresponding to the Hive component can be omitted from the component layer monitoring interface. In other possible implementations, in order to ensure the completeness and comprehensiveness of the monitoring process, a Hive component monitoring interface can also be set up. This invention does not limit this.
[0155] See Figure 7 , Figure 7 This is a schematic diagram of an HDFS component monitoring interface according to an embodiment of the present invention. Figure 7The HDFS component monitoring interface shown displays big data service usage data and container status data. For example, you can see... Figure 7 The HDFS component monitoring interface shown displays HDFS capacity usage (HDFS_Capacity), file descriptor usage (File_descriptor_Usage), HDFS file count (HDFS_File_Number), HDFS block distribution status (HDFS_Block_Status), DataNode node status (DataNode_Status), NameNode node JVM heap memory usage (NameNode_JVM_Heap_Mem), NameNode node JVM heap garbage collection (GC) count (NameNode_JVM_GC_Count), and NameNode node JVM heap GC time (NameNode_JVM_GC_Time).
[0156] See Figure 8 , Figure 8 This is a schematic diagram of a YARN component monitoring interface according to an embodiment of the present invention. Figure 8 The YARN component monitoring interface shown displays big data service usage data and container status data. For example, you can see... Figure 8 The YARN component monitoring interface shown displays YARN CPU usage (YARN_CPU_Usage), YARN memory usage (YARN_Mem_Usage), NodeManager node memory usage (NodeManager_Memory_Usage), application status in YARN (Application_Status), application exceptions in YARN (Application_Exception), ResourceManager node JVM heap GC time (ResourceManager_JVM_GC_Time), NameNode node status (NodeManager_Status), and ResourceManager node JVM heap memory usage (ResourceManager_JVM_Mem).
[0157] See Figure 9 , Figure 9 This is a schematic diagram of a Clickhouse component monitoring interface according to an embodiment of the present invention. Figure 9 The ClickHouse component monitoring interface shown displays big data service usage data and container status data. For example, you can see... Figure 9 The Clickhouse component monitoring interface shown displays the number of Clickhouse queries, the number of Clickhouse merges and rows per minute, Clickhouse read / write activity, Clickhouse compressed read buffer size per minute, current Clickhouse connections, number of Clickhouse tasks, Clickhouse cache rate, and Clickhouse memory usage.
[0158] It's important to note that because the monitoring service for the servers is deployed along with the servers, there won't be any issues monitoring the servers within the big data cluster through the monitoring interface. However, the monitoring service for the containers can only be deployed after the containers for the big data components are deployed. Therefore, when monitoring the big data cluster through the monitoring interface, you might encounter situations where you can't view their operational data because the components haven't been deployed yet. See [link to relevant documentation]. Figure 10 , Figure 10 This is a schematic diagram of a component layer monitoring interface according to an embodiment of the present invention, such as... Figure 10 As shown, for components that have not yet been deployed, a message can be displayed: "Component has not been deployed and cannot be viewed at this time".
[0159] For ease of understanding, the following sections describe the process of deploying physical layer monitoring services to display physical layer monitoring pages and the process of deploying component layer monitoring services to display component layer monitoring pages. Deploying physical layer monitoring services includes deploying the first data acquisition service, monitoring and alarm service, and data visualization and analysis service, while deploying component layer monitoring services includes deploying the second data acquisition service, monitoring and alarm service, and data visualization and analysis service.
[0160] See Figure 11 , Figure 11 This is a flowchart illustrating the deployment process of a physical layer monitoring service according to an embodiment of the present invention, such as... Figure 11As shown, after the basic service platform of the big data cluster is deployed, the Prometheus service, Granafa service, and Node-exporter-ip1 service (ip1 is also the identifier of the management server) can be automatically deployed on the management server for the big data cluster. When a physical pool operation is detected, if it is a delete physical pool operation to delete server ip2, the delete physical pool interface can be called to delete the Node-exporter-ip2 service; if it is a add physical pool operation to add server ip2, the initialize physical pool interface can be called to automatically deploy the Node-exporter-ip2 service on the new physical pool (i.e., the server ip2 newly added to the big data cluster). After the deletion or deployment of the Node-exporter-ip2 service is completed, the first configuration file maintained by the Prometheus service can be modified, and the corresponding command can be executed through the Prometheus service to dynamically load the first configuration file to obtain the latest monitoring list (used to indicate the data source that needs to be captured), so that the physical layer monitoring page can be displayed based on the latest monitoring list.
[0161] The information recorded in the first configuration file before modification may include: static_configs: - targets: [node-exporter-ip1:9100] The information recorded in the modified first configuration file can be: static_configs: - targets: [node-exporter-ip1:9100,node-exporter-ip2:9100] See Figure 12 , Figure 12 This is a flowchart illustrating the deployment process of a component-layer monitoring service according to an embodiment of the present invention, such as... Figure 12As shown, after the basic service platform for the big data cluster is deployed, Prometheus and Granafa services can be automatically deployed on the management server for the big data cluster. When a new type of big data component deployment operation is detected, if it is a big data component deletion operation, the deployment interface can be called to automatically delete its corresponding data acquisition service; if it is a big data component addition operation, the deployment interface can be called to automatically deploy the corresponding data acquisition service for the newly added container. After the deletion or deployment of the data acquisition service is completed, the first configuration file maintained by the Prometheus service can be modified, and the corresponding command can be executed through the Prometheus service to dynamically load the first configuration file, thereby updating the data on the component layer monitoring page.
[0162] Taking adding the HDFS-exporter service as an example, when modifying the Prometheus configuration file, you can add the following information to the first configuration file: - job_name: "hdfs_exporter static_configs: - targets: [hdfs-exporter:9131] According to the above embodiments, the data flow of the monitoring service implementation process provided by the present invention can be found in [reference needed]. Figure 13 , Figure 13 This is a schematic diagram illustrating the data flow of a monitoring service according to an embodiment of the present invention, such as... Figure 13 As shown, runtime data of servers and containers can be collected through the Exporter service. For example, the Node-exporter service (including Node-exporter-ip1, Node-exporter-ip2, etc.) can collect runtime data of servers, while the Hadoop-exporter service and Clickhouse-exporter service can collect runtime data of containers corresponding to big data components. The Prometheus service can then capture the runtime data collected by the Exporter service and pass the captured data to the Grafana service. The Grafana service has pre-configured physical layer dashboards and component layer dashboards (including HDFS monitoring dashboard, YARN monitoring dashboard, and Clickhouse monitoring dashboard), so the data transmitted from the Prometheus service can be used to display physical layer monitoring pages, HDFS monitoring pages, YARN monitoring pages, and Clickhouse monitoring pages.
[0163] Optionally, if the operational data collected through the data acquisition service is abnormal, alarm information can be sent through the monitoring and alarm service so that operation and maintenance personnel can handle the abnormality in a timely manner.
[0164] The above embodiments mainly describe the process of deploying monitoring services for big data clusters. In other possible implementations, a gateway proxy service can also be deployed for the big data cluster. This gateway proxy service provides access functionality to users outside the big data cluster. The gateway proxy service can be a Knox service.
[0165] It should be noted that the deployment of the big data cluster in this invention is based on the OverLay network. The hostname and port provided by the OverLay network are inaccessible from the outside. Communication with the big data cluster can only be achieved by deploying the Knox service inside the OverLay network and sending read and write requests to the Knox service from the outside.
[0166] Taking the process of accessing the container corresponding to the HDFS component as an example, the HDFS component can include the NameNode node and DataNode nodes. The NameNode can be regarded as the manager in the distributed file system, mainly responsible for managing the file system namespace, cluster configuration information, and storage block replication. The NameNode node can store the file system metadata in memory, which mainly includes file information, information about the file blocks corresponding to each file, and information about each file block in the DataNodes. The DataNode is the basic unit of file storage. It stores file blocks in the local file system, saves the metadata of the file blocks, and periodically sends all existing file block information to the NameNode.
[0167] When an external client initiates a file read / write request to the NameNode, the NameNode returns the DataNode information for the portion it manages, based on the file size and file block configuration. The client divides the file into multiple file blocks and, according to the DataNode address information, writes these blocks sequentially to each DataNode block. However, in an Overlay network, the DataNode connection information returned by the NameNode is the hostname and port within the Overlay network, which is inaccessible from the outside. To enable communication between the external network and the DataNode nodes, a Knox service must be deployed within the Overlay network. External read / write requests are sent to the Knox service, which then facilitates communication with the DataNode nodes, thus enabling data read / write operations.
[0168] In some embodiments, when deploying a gateway proxy service for a big data cluster, the gateway proxy service can be deployed on the management server of the big data cluster.
[0169] Optionally, if new servers and / or containers are added to the big data cluster, corresponding gateway proxy services need to be deployed for the new servers and / or containers.
[0170] However, it should be noted that gateway proxy service is also a general-purpose basic service. Therefore, for a big data cluster, the gateway proxy service only needs to be deployed once on the management server. When deploying the gateway proxy service for new servers and / or containers in the future, only the already deployed gateway proxy service needs to be updated.
[0171] Optionally, the gateway proxy service can maintain a second configuration file, which can be used to record the objects to be served by the gateway proxy service. When updating an already deployed gateway proxy service, the second configuration file maintained by the gateway proxy service can be updated to include newly added servers and / or containers as objects to be served by the gateway proxy service, thereby achieving the effect of deploying the gateway proxy service for newly added servers and / or containers.
[0172] It should be noted that since big data services are primarily provided through containers during the actual operation of big data clusters, external access is generally mostly via containers. Therefore, when deploying the gateway proxy service, it is necessary to ensure that the containers corresponding to the big data components are already deployed or will be deployed in the big data cluster. Furthermore, the component parameters of the deployed big data components can be used as configuration parameters for the gateway proxy service deployment.
[0173] See Figure 14 , Figure 14 This is a flowchart illustrating the deployment process of a gateway proxy service according to an embodiment of the present invention, such as... Figure 14 As shown, when a user triggers the deployment start operation in the deployment interface to deploy a big data cluster, it can determine whether the objects to be deployed include containers corresponding to HDFS, YARN, or Hive components. If the objects to be deployed include containers corresponding to HDFS, YARN, or Hive components, the component parameters of the HDFS, YARN, or Hive components to be deployed can be used as configuration parameters. By calling the Knox plugin, the second configuration file can be modified, thereby realizing the deployment of the Knox service.
[0174] Once the gateway proxy service is deployed, users can obtain the network addresses of the deployed containers through the deployed gateway proxy service, so that they can access the big data cluster through the obtained network addresses.
[0175] In some embodiments, an information viewing control can be provided in the deployment interface. The information viewing control can be used to provide the function of viewing the container network address so that users can obtain the container network address through the information viewing control. In response to the triggering operation of the information viewing control, the network address of the container displayed in the deployment interface is displayed.
[0176] Optionally, the information viewing control can be set in the deployment resource pool area, so that when displaying the network address of the container shown in the deployment interface, the network address of the container shown in the deployment resource pool area can be displayed.
[0177] Still as Figure 2 Taking the deployment interface shown as an example, in such a case... Figure 2 In the deployment interface shown, the button labeled "Connection Information" is the information viewing control. Users can obtain the network addresses of the containers already deployed in the big data cluster by triggering the button labeled "Connection Information".
[0178] Optionally, when a user triggers the information viewing control, the computing device can display an information viewing interface to show the user the network addresses of containers deployed in the big data cluster.
[0179] It's important to note that different types of containers correspond to different types of network addresses. If the deployed components are HDFS, YARN, and Hive, the returned address will be the Knox proxy address. HDFS and YARN return the URL of the proxy web page, while Hive returns the URL of the proxy Java Database Connectivity (JDBC). If the deployed component is Clickhouse, the returned address will be the Clickhouse JDBC URL, which will be displayed on the information viewing interface.
[0180] If HDFS, YARN, Hive, and Clickhouse components have already been deployed in the big data cluster, the information viewing interface can be found here. Figure 15 , Figure 15 This is a schematic diagram of an information viewing interface according to an embodiment of the present invention, such as... Figure 15 As shown, the network address of the container corresponding to the NameNode node and ResourceManager node is the URL address of the proxy web page, the network address of the container corresponding to the Hive node is the URL address of the proxy JDBC, and the network address of the container corresponding to the Clickhouse node is the URL address of the proxy JDBC.
[0181] Optionally, after obtaining the container's network address, users can copy the web page's URL into their browser, enter their username and password, and access the management pages of the HDFS and YARN components; alternatively, they can configure the proxy JDBC URL and username / password using a database connection tool to connect to the Hive and Clickhouse components for data analysis.
[0182] See Figure 16 , Figure 16 This is a flowchart illustrating an embodiment of the present invention for obtaining a container network address, such as... Figure 16As shown, if a user requests to view connection information, the computing device can query the network deployment status in the big data cluster. Since the HDFS, YARN, and Hive components need to access the service through the Knox proxy, if the deployed services include those corresponding to the HDFS, YARN, and Hive components, the URL address of the Knox proxy can be returned. If the deployed services are those corresponding to the HDFS and YARN components, the returned URL address is the URL address of the web page for the Knox proxy; if the deployed services are those corresponding to the Hive component, the URL address of the JDBC proxy for Knox can be returned. If the deployed services only include those corresponding to the Clickhouse component, the JDBC URL address of Clickhouse can be returned.
[0183] It should be noted that the URLs corresponding to the HDFS and YARN components are the URLs of the web pages, and therefore can be accessed by copying them into a browser; while the URLs corresponding to the Hive and Clickhouse components are the URLs of the proxy JDBC, and therefore can be accessed through a database connection client.
[0184] A schematic diagram of the interface used to manage different components can be found here. Figures 17 to 20 , Figure 17 This is a schematic diagram of a web interface for managing HDFS components according to an embodiment of the present invention. Figure 18 This is a schematic diagram of a web interface for managing YARN components according to an embodiment of the present invention. Figure 19 This is a schematic diagram of a client interface for managing Hive components according to an embodiment of the present invention. Figure 20 This is a schematic diagram of a client interface for managing Clickhouse components according to an embodiment of the present invention.
[0185] Additionally, if the big data component has not yet been deployed in the big data cluster, it can display something like this: Figure 21 The information viewing interface shown is available in [reference]. Figure 21 , Figure 21 This is a schematic diagram of another information viewing interface according to an embodiment of the present invention. Figure 21 The information viewing interface shown displays a "No data available" message, indicating that the big data component has not yet been deployed in the big data cluster.
[0186] Through the above embodiments, users can achieve the purpose of accessing big data clusters from the outside, thereby enhancing the flexibility of using big data clusters.
[0187] Optionally, the present invention can also provide a function for the overall migration of services deployed on a management server. For example, in the event of an alarm, there may be insufficient storage resources on the management server. As the core of the entire big data cluster, insufficient storage resources on the management server may cause the entire big data cluster to malfunction. To ensure the smooth operation of the big data cluster, a service migration function can be provided, which allows for the overall migration of services deployed on the management server to other servers with more storage resources.
[0188] In some embodiments, the management node corresponding to the management server can be displayed in the deployment interface so that users can trigger the service migration process by dragging and dropping nodes in the deployment interface.
[0189] In one possible implementation, the deployment interface displays multiple physical pools (each physical pool corresponds to a server) in the deployment resource pool. When displaying the management node corresponding to the management server, the management node can be displayed in the physical pool corresponding to the management server.
[0190] Users can drag and drop management nodes, which will appear in the physical pool corresponding to the management server, to other physical pools to migrate services deployed on the management server to other servers.
[0191] In one possible implementation, in response to a drag-and-drop operation on the management node in the deployment interface, the configuration data in the management server is copied to the destination server indicated by the drag-and-drop operation, thereby migrating the services deployed on the management server to the destination server. The destination server is the server corresponding to the physical pool where the drag-and-drop operation ends.
[0192] Optionally, the deployment interface can include a deployment control that users can trigger after moving the management node from the management server to the destination server, thereby initiating a background service migration operation. See also Figure 22 , Figure 22 This is a schematic diagram of a deployment interface according to an embodiment of the present invention. Figure 22 In the deployment interface shown, the button labeled "Start Deployment" is the deployment control.
[0193] The following is as follows Figure 22 The deployment interface shown illustrates the drag-and-drop process for managing nodes. Figure 22As shown, the server with IP address 10.10.239.152 is the management server, and the physical pool corresponding to the management server displays management nodes. If the server with IP address 10.10.177.23 is used as the destination server, after the user drags and drops the management node from the physical pool corresponding to the management server to the physical pool corresponding to the destination server, the deployment interface will update as shown. Figure 23 See the form shown. Figure 23 , Figure 23 This is a schematic diagram of another deployment interface according to an embodiment of the present invention. Figure 23 In the deployment interface shown, the management nodes in the physical pool corresponding to the management server will be displayed as pending relocation. Figure 23 The dotted line indicates that the management node is in a state to be moved, and the management node can be displayed in the physical pool corresponding to the destination server. However, in cases such as... Figure 23 In the state shown, only the nodes have been moved at the interface level; the actual services have not yet been migrated.
[0194] Optionally, users can, for example Figure 23 The deployment interface shown triggers the "Start Deployment" button, which serves as the deployment control, to initiate the service migration process.
[0195] During service migration, prompts can be displayed to remind users to wait for the migration to complete. For example, a prompt can be displayed in response to dragging and dropping the management node in the deployment interface, indicating that the management server and the destination server are being redeployed.
[0196] See Figure 24 , Figure 24 This is a schematic diagram of another deployment interface according to an embodiment of the present invention. Figure 24 The deployment interface shown displays the message "Deploying in progress. This operation may take some time. Please wait patiently."
[0197] It should be noted that the service migration process is essentially the process of copying the configuration data from the management server to the destination server. Once the data copy is complete, the management node migration is considered complete. At this point, the computing device can automatically delete the management node corresponding to the management server from the deployment interface by following the drag-and-drop instructions in the deployment interface, and then display the management node corresponding to the destination server in the deployment interface.
[0198] For example, the deployment interface can be updated to, as follows: Figure 25 As shown in the diagram. See also Figure 25 , Figure 25This is a schematic diagram of another deployment interface according to an embodiment of the present invention. Figure 25 In the deployment interface shown, the physical pool corresponding to the original management server (i.e., the server with IP address 10.10.239.152) no longer has a management node, while the physical pool corresponding to the destination server (i.e., the server with IP address 10.10.177.23) shows a management node. At this time, the server with IP address 10.10.177.23 can be used as the new management server in the big data cluster.
[0199] Optionally, to ensure that users can continue to use the services provided by the big data cluster during the service migration process, and to achieve the goal of service migration without the user's awareness, Internet Protocol (IP) technology can be used.
[0200] A virtual IP address is an IP address that does not correspond to a specific computer or a specific network interface card (NIC). All data packets sent to this IP address eventually reach the destination process on the host machine via the actual NIC. A common use case for virtual IPs is in high availability (HA) systems. Typically, a system may experience downtime due to routine maintenance or unexpected events. To improve the high availability of the system's external services, a primary / standby configuration is used. When the host providing the service, M, fails, the service switches to the standby host, S, to continue providing services. Users are unaware of this process. In this scenario, the IP address providing services to clients is a virtual IP address. When host M fails, the virtual IP address floats onto the standby host and continues providing services. In this case, the virtual IP address is not associated with a specific computer host or a specific physical NIC; it is a virtual or logical concept that can move and float freely. This shields the system's internal details from external access while facilitating maintainability and scalability.
[0201] In other words, a virtual IP address can be pre-configured for the management server. However, since the virtual IP address is not the IP address of any real server in the network environment, in order to ensure that the virtual IP address can be used to locate the corresponding server in the network environment, an Address Resolution Protocol (ARP) cache can be maintained.
[0202] Optionally, each server can maintain an ARP cache to store the mapping between IP addresses and physical addresses (Media Access Control, MAC) within the same network (i.e., the ARP cache table). When a server in the Ethernet sends data, it first looks up the MAC address corresponding to the target IP in this cache table, and then sends the data to the server corresponding to that MAC address.
[0203] Therefore, the mapping between virtual IP addresses and real server MAC addresses can be maintained through ARP caching, so that the real server in the network can be found based on the virtual IP address.
[0204] In one possible implementation, after the service migration is complete, the virtual IP address pre-configured for the management server can be deleted, and the virtual IP address can be configured for the destination server to achieve the switching of the server proxied by the virtual IP address, thereby deleting the configuration data in the management server.
[0205] Optionally, when switching the server proxied by the virtual IP address from one server (referred to as the origin server) to another server (referred to as the destination server), the destination server can send an ARP broadcast to servers within the same network. All servers that receive the broadcast will automatically refresh their maintained ARP cache tables to achieve service switching, so that requests can be sent to the service corresponding to the new host through the virtual IP address in the future.
[0206] The above embodiment is illustrated by taking the switching of the server proxied by the virtual IP address after the service migration is completed as an example. In more possible implementations, before switching the server proxied by the virtual IP address, the service corresponding to the configuration data can be started on the destination server based on the configuration data copied from the management server. If all services on the destination server start normally, the virtual IP address that was previously set for the management server is deleted, and the virtual IP address is configured for the destination server.
[0207] It should be noted that this invention mainly introduces big data services and monitoring services deployed on the management server. In general, database services (such as MySQL services), message queue services (such as RabbitMQ services), and front-end and back-end services of the big data platform are also deployed on the management server. When migrating services in the management server to the destination server, all of the above services need to be migrated.
[0208] Optionally, when starting the various services on the target server, they can be started in the following order: "Database Service → Message Queue Service → Big Data Service → Big Data Platform Front-end and Back-end Services → Monitoring Service". Alternatively, other startup orders can be used. For example, the startup order of the Big Data Platform Front-end and Back-end Services and the Big Data Service can be swapped. However, it is necessary to ensure that basic services such as the Database Service and Message Queue Service are started before the Big Data Platform Front-end and Back-end Services and the Big Data Service, while additional functions such as the Monitoring Service can generally be started last.
[0209] By testing the service startup status, the virtual IP address can be switched only when the migrated service can start normally. This ensures that services can be provided to users normally before and after the virtual IP address switch, so as to achieve the goal of making the user feel no difference.
[0210] The following example illustrates the process of switching virtual IP addresses, using the original server's real IP address as 10.10.177.11, the destination server's real IP address as 10.10.177.12, and the virtual IP address as 10.10.177.15: See Figure 26 , Figure 26 This is a flowchart illustrating a virtual IP address switching process according to an embodiment of the present invention, such as... Figure 26 As shown, the management node is set up on the original server with IP address 10.10.177.11. When providing services externally, the original server uses the virtual IP address 10.10.177.15 to identify itself. When migrating the management node to the destination server with IP address 10.10.177.12, the data on the original server is first copied. After the data copy is complete, the services on the destination server are started in the following order: "MySQL service → RabbitMQ service → Big Data Platform front-end and back-end services → Big Data service → Monitoring service (including Prometheus service and Granafa service)".
[0211] See Figure 27 , Figure 27 This is a flowchart illustrating a virtual IP address switching process according to an embodiment of the present invention, such as... Figure 27 As shown, after starting each service on the target server, you can check whether the service is working properly by accessing the port.
[0212] If the service is working properly, you can enter as follows: Figure 28 The process shown is described in the following document. Figure 28 , Figure 28This is a flowchart illustrating a virtual IP address switching process according to an embodiment of the present invention, such as... Figure 28 As shown, a virtual IP address can be deleted from the original server, and then added to the destination server. The destination server then sends an ARP broadcast to other servers in the same network to update its ARP cache, thus achieving virtual IP switching. After the virtual IP switching is complete, the service deployed on the original server can be deleted.
[0213] See Figure 29 , Figure 29 This is a schematic diagram illustrating a virtual IP address switching result according to an embodiment of the present invention, such as... Figure 29 As shown, after the switch, the management node will be located on the destination server, and the virtual IP address 10.10.177.15 will be used as the IP address of the destination server during communication.
[0214] The management node migration process provided in the above embodiments can be found in [reference needed]. Figure 30 , Figure 30 This is a flowchart illustrating a management node migration process according to an embodiment of the present invention, such as... Figure 30 As shown, if a user drags a management node on the deployment interface (e.g., from the original machine to the new machine), the computing device can copy data from the original machine to the new machine and start services on the new machine sequentially (including MySQL, RabbitMQ, Knox, Prometheus, Granafa, etc.). The device uses a polling approach to determine if the services deployed on the new machine are functioning correctly. If they are, the device switches the server proxied by the virtual IP address and sends ARP broadcasts to machines on the network to update the ARP cache table, completing the virtual IP address switch. After the virtual IP address switch is complete, the services and data on the original machine can be deleted.
[0215] The above embodiments mainly describe the service migration function provided by the present invention from the perspective of responding to management server anomalies through service migration. In more possible implementations, when no anomalies occur in the big data cluster, users can also migrate the services deployed on the management server as a whole according to actual technical needs. The present invention does not limit this.
[0216] See Figure 31 , Figure 31 This is a schematic diagram illustrating the principle architecture of a big data cluster deployment method according to an embodiment of the present invention, such as... Figure 31As shown, users can select the type and parameters of the nodes to be deployed in the deployment interface. The selected nodes will then be displayed in the temporary resource pool of the deployment interface. Nodes in the temporary resource pool can be transferred to the deployment resource pool by dragging and dropping or automatic allocation. Users can perform operations on the physical pool and nodes in the deployment resource pool. Physical pool operations can include adding, deleting, prioritizing, and monitoring physical pools. Node operations can include deploying, moving, deleting, and retiring nodes. Furthermore, this invention can provide physical layer dashboards and component layer dashboards (including HDFS monitoring dashboards, YARN monitoring dashboards, and Clickhouse monitoring dashboards) so that users can view the monitored operational data through the dashboards. Additionally, functions such as viewing connection information and restoring factory settings are also available.
[0217] The solution provided by this invention can provide automated deployment monitoring services and gateway proxy services on the basis of automated deployment of big data components. It provides users with a complete big data infrastructure platform. Without logging into the server, users can automatically deploy big data components through drag-and-drop and click operations on the front-end page. At the same time, the monitoring of machine operation status, the monitoring of component service functions, and the gateway proxy function are automatically completed, which greatly reduces the technical threshold of big data operation and maintenance and can improve service security and customer experience.
[0218] Corresponding to the embodiments of the aforementioned methods, the present invention also provides embodiments of corresponding big data cluster deployment devices and the computing devices used thereon.
[0219] See Figure 32 , Figure 32 This is a block diagram illustrating a big data cluster deployment device according to an embodiment of the present invention, such as... Figure 32 As shown, the device includes: Deployment module 3201 is used to deploy data acquisition services, monitoring and alarm services, and data visualization and analysis services for big data clusters. The data processing module 3202 is used to collect the operation data of the big data cluster through the data acquisition service, and upload the operation data to the data visualization and analysis service through the monitoring and alarm service. Display module 3203 is used to display the running data monitoring interface through data visualization analysis services. The monitoring interface is used to display the running status of the big data cluster in a graphical way.
[0220] In some embodiments of the present invention, the deployment module 3201, when used to deploy data acquisition services, monitoring and alarm services, and data visualization and analysis services for a big data cluster, is used for: Based on the service providers currently deployed in the big data cluster, deploy data acquisition services, monitoring and alarm services, and data visualization and analysis services for the current service providers in the big data cluster; In response to update operations on deployed service providers in the big data cluster, update the data acquisition service, monitoring and alarm service, and data visualization and analysis service deployed in the big data cluster. Update operations on deployed service providers include deleting deployed service providers and deploying new service providers.
[0221] In some embodiments of the present invention, the service provider is a server and / or a container, and the big data cluster includes multiple servers, among which there is a management server; Deployment module 3201, when used to deploy data acquisition services, monitoring and alarm services, and data visualization and analysis services for the current service providers in the big data cluster, is used for: Based on the servers and containers currently included in the big data cluster, deploy data acquisition services on each server of the big data cluster, and deploy monitoring and alarm services and data visualization and analysis services on the management server of the big data cluster.
[0222] In some embodiments of the present invention, the data acquisition service includes a first data acquisition service and a second data acquisition service, wherein the first data acquisition service is used to acquire server operation data and the second data acquisition service is used to acquire container operation data. Deployment module 3201, when used to deploy a data acquisition service on each server in a big data cluster based on the servers and containers currently included in the big data cluster, is used for: Deploy the first data acquisition service and the second data acquisition service on the management server, and deploy the first data acquisition server on each of the multiple servers except the management server.
[0223] In some embodiments of the present invention, the service provider is a server and / or a container, and the monitoring and alarm service corresponds to a first configuration file; When an update operation on an already deployed service provider is to deploy a new service provider, the deployment module 3201, in responding to an update operation on an already deployed service provider in the big data cluster, updates the data acquisition service, monitoring and alarm service, and data visualization and analysis service already deployed in the big data cluster, and performs at least one of the following: In response to the addition of a new server in the big data cluster, the first data acquisition service is deployed on the new server, and the first configuration file corresponding to the monitoring and alarm service is modified based on the new service provider, so that the monitoring and alarm service and data visualization analysis service already deployed in the big data cluster can provide services to the new server. In response to the addition of a new container in the big data cluster, the first configuration file corresponding to the monitoring and alarm service is modified based on the new server, so that the second data acquisition service, monitoring and alarm service, and data visualization and analysis service already deployed in the big data cluster can provide services to the new container.
[0224] In some embodiments of the present invention, when the update operation on the deployed service provider is to delete the deployed service provider, the deployment module 3201, in response to the update operation on the deployed service provider in the big data cluster, updates the data acquisition service, monitoring and alarm service, and data visualization and analysis service deployed in the big data cluster, is used for at least one of the following: In response to the deletion operation of the server already deployed in the big data cluster, the first data acquisition service corresponding to the server is deleted, and the first configuration file corresponding to the monitoring and alarm service is modified based on the deleted server, so that the monitoring and alarm service and data visualization analysis service already deployed in the big data cluster no longer provide services to the deleted server. In response to the deletion of containers already deployed in the big data cluster, the first configuration file corresponding to the monitoring and alarm service is modified based on the deleted server, so that the second data acquisition service, monitoring and alarm service, and data visualization and analysis service already deployed in the big data cluster no longer provide services to the deleted server.
[0225] In some embodiments of the present invention, the data acquisition service, the monitoring and alarm service, and the data visualization and analysis service are located in the same overlay network. The data acquisition service is the Exporter service, the monitoring and alarm service is the Prometheus service, and the data visualization and analysis service is the Grafana service.
[0226] In some embodiments of the present invention, the first data acquisition service is the Node-Exporter service; when the container is a container corresponding to the HDFS component or the YARN component, the second data acquisition service is the Hadoop-Exporter service; when the container is a container corresponding to the Clickhouse component, the second data acquisition service is the Clickhouse-Exporter service.
[0227] In some embodiments of the present invention, the service provider is a server and / or a container; for a server in a big data cluster, the running data includes at least one of disk space usage data, network traffic data, CPU usage data, and memory usage data; for a container in a big data cluster, the running data includes at least one of big data service usage data and container status data.
[0228] In some embodiments of the present invention, the service provider is a server and / or a container; When used for monitoring an interface based on running data display, the display module 3203 is used for at least one of the following: Based on the operational data of the servers in the big data cluster, a physical layer monitoring interface is displayed. The physical layer monitoring interface is used to display the operational data of the servers in the big data cluster. Based on the runtime data of containers in the big data cluster, a component-level monitoring interface is displayed. This interface is used to show the runtime data of containers in the big data cluster.
[0229] In some embodiments of the present invention, the device further includes: The sending module is used to send alarm information through the monitoring and alarm service in response to abnormalities in the operational data collected by the data acquisition service.
[0230] In some embodiments of the present invention, the big data cluster includes multiple servers, and one of the multiple servers is a management server. The display module 3203 is also used to display the deployment interface and display the management node corresponding to the management server in the deployment interface; The device also includes: The copy module is used to respond to drag-and-drop operations on the management node in the deployment interface, and copy the configuration data in the management server to the destination server indicated by the drag-and-drop operation.
[0231] In some embodiments of the present invention, a virtual IP address is pre-configured for the management server; The second deletion module is used to delete the virtual IP address that was pre-set for the management server and to configure the virtual IP address to the destination server; The second deletion module is also used to delete configuration data in the management server.
[0232] In some embodiments of the present invention, the startup module is used to start the service corresponding to the configuration data in the destination server based on the configuration data copied from the management server; The second deletion module is also used to perform the steps of deleting the virtual IP address that was pre-set for the management server and configuring the virtual IP address to the destination server if all services in the destination server are started normally.
[0233] In some embodiments of the present invention, the display module 3203 is also configured to display a prompt message in response to a drag-and-drop operation on the management node in the deployment interface, the prompt message indicating that the management server and the destination server are being redeployed.
[0234] In some embodiments of the present invention, the second deletion module is further configured to delete the management node corresponding to the management server from the deployment interface according to the instructions of the drag operation on the management node in the deployment interface. The display module 3203 is also used to display the management node corresponding to the destination server in the deployment interface.
[0235] In some embodiments of the present invention, the deployment module 3201 is further configured to deploy a gateway proxy service for the big data cluster, the gateway proxy service being used to provide access functionality to users outside the big data cluster.
[0236] In some embodiments of the present invention, the big data cluster includes multiple servers, and one of the multiple servers is a management server. Deployment module 3201, when used to deploy gateway proxy services for big data clusters, is used for: Deploy a gateway proxy service on the management server; In response to the addition of a new service provider in the big data cluster, deploy a gateway proxy service for the new service provider.
[0237] In some embodiments of the present invention, the display module 3203 is also used to display a deployment interface and display an information viewing control in the deployment interface. The information viewing control is used to provide the function of viewing the container network address. The display module 3203 is also used to display the network address of the container shown in the deployment interface in response to a trigger operation of the information viewing control.
[0238] In some embodiments of the present invention, the deployment interface includes a deployment resource pool area, which is used to display the nodes corresponding to the containers deployed on each server of the big data cluster; Display module 3203, when used to display information viewing controls in the deployment interface, is used for: Display information viewing controls in the deployment resource pool area of the deployment interface; The display module 3203, when used to display the network address of the container shown in the deployment interface in response to a trigger operation of the information viewing control, is used for: In response to a trigger action on the information viewing control, display the network addresses of the containers shown in the deployment resource pool area.
[0239] In some embodiments of the present invention, the service providers deployed in the big data cluster are deployed based on node drag-and-drop operations in the deployment interface.
[0240] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0241] The present invention also provides a computing device, see [link to relevant documentation]. Figure 33 , Figure 33 This is a schematic diagram of the structure of a computing device according to an embodiment of the present invention. Figure 33 As shown, the computing device includes a processor 3310, a memory 3320, and a network interface 3330. The memory 3320 stores computer instructions that can run on the processor 3310. The processor 3310 is used to implement the big data cluster deployment method provided in any embodiment of the present invention when executing the computer instructions. The network interface 3330 is used to implement input / output functions. In more possible implementations, the computing device may also include other hardware, which is not limited by the present invention.
[0242] This invention also provides a computer-readable storage medium, which can take many forms, such as RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (e.g., hard disk drives), solid-state drives, any type of storage disk (e.g., optical discs, DVDs), or similar storage media, or combinations thereof. Specifically, the computer-readable medium can also be paper or other suitable media capable of printing programs. A computer program is stored on the computer-readable storage medium, and when executed by a processor, the computer program implements the big data cluster deployment method provided in any embodiment of this invention.
[0243] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the big data cluster deployment method provided in any embodiment of the present invention.
[0244] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, apparatus, computing device, computer-readable storage medium, or computer program product. Therefore, one or more embodiments of this specification can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-readable program code.
[0245] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments corresponding to computing devices are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0246] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of this invention. In some cases, the actions or steps described in this invention may be performed in a different order than those shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0247] The embodiments of the subject matter and functional operation described in this specification can be implemented in the following ways: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by or control of the operation of a big data cluster deployment device. Alternatively or additionally, the program instructions may be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information and transmit it to a suitable receiving device for execution by the big data cluster deployment device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof.
[0248] The processing and logic flow described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform corresponding functions by operating on input data and generating output. The processing and logic flow can also be executed by dedicated logic circuitry—such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the device can also be implemented as dedicated logic circuitry.
[0249] Suitable computers for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, or the computer will be operatively coupled to such mass storage devices to receive data from or transfer data to them, or both. However, a computer is not required to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0250] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.
[0251] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.
[0252] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0253] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the invention. In some cases, the actions described in the invention can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0254] Other embodiments of this specification will readily occur to those skilled in the art upon consideration of the specification and practice of the invention claimed herein. This specification is intended to cover any variations, uses, or adaptations that follow the general principles of this specification and include common knowledge or customary techniques in the art not claimed herein. That is, this specification is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
[0255] The above description is merely an optional embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification shall be included within the scope of protection of this specification.
Claims
1. A method for deploying a big data cluster, characterized in that, The method includes: Deploy data acquisition services, monitoring and alarm services, and data visualization and analysis services for big data clusters; The data acquisition service collects the operational data of the big data cluster, and the monitoring and alarm service uploads the operational data to the data visualization and analysis service. Through the data visualization and analysis service, the monitoring interface is based on the running data display, and the monitoring interface is used to graphically display the running status of the big data cluster; The big data cluster includes multiple servers, among which there is a management server, and the management server is pre-configured with a virtual IP address; The method further includes: Display the deployment interface, and display the management node corresponding to the management server in the deployment interface; In response to a drag-and-drop operation on the management node in the deployment interface, the configuration data in the management server is copied to the destination server indicated by the drag-and-drop operation; Based on the configuration data copied from the management server, start the service corresponding to the configuration data on the destination server; Delete the virtual IP address that was previously set for the management server, and then configure the virtual IP address to the destination server.
2. The method according to claim 1, characterized in that, The aforementioned deployment of data acquisition services, monitoring and alarm services, and data visualization and analysis services for big data clusters includes: Based on the service providers currently deployed in the big data cluster, deploy data acquisition services, monitoring and alarm services, and data visualization and analysis services for the current service providers in the big data cluster; In response to update operations on deployed service providers in the big data cluster, the data acquisition service, monitoring and alarm service, and data visualization and analysis service deployed in the big data cluster are updated. Update operations on deployed service providers include deleting deployed service providers and deploying new service providers.
3. The method according to claim 2, characterized in that, The service provider is a server and / or a container, and the big data cluster includes multiple servers, among which there is a management server. The provision of data acquisition services, monitoring and alarm services, and data visualization and analysis services for the service providers currently included in the big data cluster, within the big data cluster, includes: Based on the servers and containers currently included in the big data cluster, a data acquisition service is deployed on each server of the big data cluster, and a monitoring and alarm service and a data visualization and analysis service are deployed on the management server of the big data cluster.
4. The method according to claim 3, characterized in that, The data acquisition service includes a first data acquisition service and a second data acquisition service. The first data acquisition service is used to collect the server's operating data, and the second data acquisition service is used to collect the container's operating data. The step of deploying a data acquisition service on each server of the big data cluster, based on the servers and containers currently included in the big data cluster, includes: A first data acquisition service and a second data acquisition service are deployed on the management server, and the first data acquisition server is deployed on each of the plurality of servers other than the management server.
5. The method according to claim 4, characterized in that, The service provider is a server and / or a container, and the monitoring and alarm service has a corresponding first configuration file; When an update operation on an already deployed service provider is to deploy a new service provider, the response to the update operation on the already deployed service provider in the big data cluster, updating the data acquisition service, monitoring and alarm service, and data visualization and analysis service already deployed in the big data cluster, includes at least one of the following: In response to the addition of a new server in the big data cluster, a first data acquisition service is deployed on the new server, and the first configuration file corresponding to the monitoring and alarm service is modified based on the new service provider, so that the monitoring and alarm service and data visualization analysis service already deployed in the big data cluster can provide services to the new server. In response to the addition of a new container in the big data cluster, the first configuration file corresponding to the monitoring and alarm service is modified based on the new server, so that the second data acquisition service, monitoring and alarm service and data visualization and analysis service already deployed in the big data cluster can provide services to the new container.
6. The method according to claim 5, characterized in that, When the update operation on a deployed service provider is to delete the deployed service provider, the response to the update operation on the deployed service provider in the big data cluster, updating the data acquisition service, monitoring and alarm service, and data visualization and analysis service deployed in the big data cluster, includes at least one of the following: In response to the deletion operation of the server already deployed in the big data cluster, the first data acquisition service corresponding to the server is deleted, and the first configuration file corresponding to the monitoring and alarm service is modified based on the deleted server, so that the monitoring and alarm service and data visualization analysis service already deployed in the big data cluster no longer provide services to the deleted server. In response to the deletion operation of the container already deployed in the big data cluster, the first configuration file corresponding to the monitoring and alarm service is modified based on the deleted server, so that the second data acquisition service, monitoring and alarm service and data visualization analysis service already deployed in the big data cluster no longer provide services to the deleted server.
7. The method according to claim 6, characterized in that, The data acquisition service, the monitoring and alarm service, and the data visualization and analysis service are located in the same overlay network. The data acquisition service is an Exporter service, the monitoring and alarm service is a Prometheus service, and the data visualization and analysis service is a Grafana service.
8. The method according to claim 7, characterized in that, The first data acquisition service is the Node-Exporter service; when the container is a container corresponding to the HDFS component or the YARN component, the second data acquisition service is the Hadoop-Exporter service; when the container is a container corresponding to the Clickhouse component, the second data acquisition service is the Clickhouse-Exporter service.
9. The method according to claim 1, characterized in that, The service provider is a server and / or a container; for the server in the big data cluster, the operational data includes at least one of disk space usage data, network traffic data, CPU usage data, and memory usage data; for the container in the big data cluster, the operational data includes at least one of big data service usage data and container status data.
10. The method according to claim 1, characterized in that, The big data cluster includes multiple service providers, which include servers and / or containers. The monitoring interface based on the running data display includes at least one of the following: Based on the operating data of the servers in the big data cluster, a physical layer monitoring interface is displayed, which is used to display the operating data of the servers in the big data cluster. Based on the runtime data of the containers in the big data cluster, a component layer monitoring interface is displayed, which is used to display the runtime data of the containers in the big data cluster.
11. The method according to claim 1, characterized in that, The method further includes: In response to an anomaly in the operational data collected through the data acquisition service, an alarm message is sent through the monitoring and alarm service.
12. The method according to claim 1, characterized in that, After copying the configuration data from the management server to the destination server indicated by the drag-and-drop operation on the management node in the deployment interface, the method further includes: Delete the configuration data in the management server.
13. The method according to claim 1, characterized in that, The step of deleting the virtual IP address pre-configured for the management server and configuring the virtual IP address for the destination server includes: If all services on the destination server start normally, then the steps of deleting the virtual IP address that was previously set for the management server and configuring the virtual IP address to the destination server are performed.
14. The method according to claim 1, characterized in that, The method further includes: In response to a drag-and-drop operation on the management node in the deployment interface, a prompt message is displayed indicating that the management server and the destination server are being redeployed.
15. The method according to claim 12, characterized in that, After deleting the configuration data in the management server, the method further includes: Following the drag-and-drop instructions for the management node in the deployment interface, delete the management node corresponding to the management server from the deployment interface, and display the management node corresponding to the destination server in the deployment interface.
16. The method according to claim 1, characterized in that, The method further includes: A gateway proxy service is deployed for the big data cluster, which provides access functionality to users outside the big data cluster.
17. The method according to claim 16, characterized in that, The big data cluster includes multiple servers, among which there is a management server; Deploying a gateway proxy service for the big data cluster includes: Deploy a gateway proxy service on the management server; In response to the addition of a new service provider in the big data cluster, a gateway proxy service is deployed for the new service provider.
18. The method according to claim 16, characterized in that, The method further includes: An information viewing control is displayed in the deployment interface, which provides the function of viewing the container network address; In response to a triggering operation of the information viewing control, the network address of the container displayed in the deployment interface is shown.
19. The method according to claim 18, characterized in that, The deployment interface includes a deployment resource pool area, which is used to display the nodes corresponding to the containers that have been deployed on each server of the big data cluster. The information viewing control displayed on the deployment interface includes: The information viewing control is displayed in the deployment resource pool area of the deployment interface; The step of displaying the network address of the container shown in the deployment interface in response to a trigger operation on the information viewing control includes: In response to a trigger operation on the information viewing control, the network address of the container displayed in the deployment resource pool area is shown.
20. The method according to claim 1, characterized in that, The service providers deployed in the big data cluster are obtained through node drag-and-drop operations in the deployment interface.
21. A big data cluster deployment device, characterized in that, include: The deployment module is used to deploy data acquisition services, monitoring and alarm services, and data visualization and analysis services for big data clusters. The data processing module is used to collect the operation data of the big data cluster through the data acquisition service, and upload the operation data to the data visualization and analysis service through the monitoring and alarm service. The display module is used to display the running data monitoring interface based on the data visualization analysis service. The monitoring interface is used to display the running status of the big data cluster in a graphical manner. The big data cluster includes multiple servers, among which there is a management server, and the management server is pre-configured with a virtual IP address; The device further includes: The display module is used to display the deployment interface and the management node corresponding to the management server in the deployment interface; The copy module is used to copy the configuration data in the management server to the destination server indicated by the drag operation in response to the drag operation on the management node in the deployment interface. The startup module is used to start the service corresponding to the configuration data on the destination server based on the configuration data copied from the management server. The second deletion module is used to delete the virtual IP address that was pre-set for the management server and to configure the virtual IP address to the destination server.
22. A computing device, characterized in that, The computing device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the operations performed by the big data cluster deployment method as described in any one of claims 1 to 20.
23. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, performs the operations performed by the big data cluster deployment method as described in any one of claims 1 to 20.