Method for rapidly deploying Spark cluster

Through a one-click rapid deployment method, the deployment process of Spark cluster is simplified by using integrated batch operation commands and initialization scripts, solving the complex and cost-effective problems of traditional deployment methods, and achieving rapid deployment and efficient replication.

CN119945911APending Publication Date: 2025-05-06XIAMEN MEIYABAIKE INFORMATION SECURITY RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411895822.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The traditional Spark cluster deployment method has complex operation, large dependencies, high personnel requirements and poor replication, resulting in high deployment time and cost.

Method used

Adopt a one-click rapid deployment method, unified management of cluster nodes, and integrated batch operation commands and initialization scripts are used to simplify the deployment process and reduce technical requirements.

Benefits of technology

The rapid deployment and replication of Spark clusters are realized, which reduces labor costs, improves construction efficiency, and simplifies deployment operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945911A_ABST
    Figure CN119945911A_ABST
Patent Text Reader

Abstract

The invention discloses a rapid deployment method for a Spark cluster, and the method specifically comprises the steps: configuring a corresponding node IP through a service planning template, and enabling the service planning template to be jointly deployed by the Spark cluster, a Hadoop cluster and a Zookeeper cluster; the method comprises the following steps of: carrying out initialization operation on an operation system by utilizing an initSystem.sh initialization script; executing the initSystem.sh initialization script to carry out one-key deployment on the Spark cluster, the Hadoop cluster and the Zookeeper cluster, wherein the initSystem.sh initialization script is used for carrying out configuration, distribution and initialization operation on a cluster deployment package by utilizing an integrated ck batch operation command and / etc / cluster. Conf configuration; carrying out one-key deployment on the Spark cluster, the Hadoop cluster and the Zookeeper cluster; the method comprises the following steps of: starting a Spark cluster, a Hadoop cluster and a Zookeeper cluster, integrating the cluster starting into a startCluster.sh script, and executing the startCluster.sh script, so as to finish the starting of the Spark cluster, the Hadoop cluster and the Zookeeper cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and mainly to a method for rapid deployment of a Spark cluster. Background Art

[0002] With the continuous development of AI technology and the introduction of large models, the requirements for fast computing of massive data are getting higher and higher. In order to meet the computing of massive data at the T level or even the PT level, Hadoop provides the offline data computing MapReduce technology, but MapReduce also has disadvantages. The computing based on disk IO occupies a large amount of disk IO. As the computing data continues to increase, the disadvantages are exposed more and more obviously, and the computing becomes slower and slower. With the continuous development of society, it is necessary to compute and mine massive data, and also to respond quickly to get results. Therefore, a technology based on memory massive data computing has emerged, namely Spark. Major enterprises or individuals have more and more demands for Spark based on large data volume and memory computing. How to quickly deploy Spark clusters has become more and more important.

[0003] The traditional deployment method is complex, highly dependent, requires high personnel, and has poor reproducibility, resulting in long deployment time and cost. In order to solve these deployment problems, the present invention is based on the idea of ​​batch operation, highly integrates related operations, and quickly deploys Spark clusters through one-click, thereby achieving rapid deployment and rapid replication, while significantly reducing the technical requirements for personnel and saving labor costs.

[0004] Spark cluster deployment depends on many components. Conventional Spark cluster deployment also requires the deployment of dependent components, which is time-consuming. It is inconvenient to copy when there are many projects. At the same time, the deployment personnel need to have certain basic knowledge of Linux, Spark and Hadoop, which is not conducive to the rapid deployment and rapid replication of Spark clusters. The deployment personnel's technical ability requirements are high, and the corresponding labor costs will also increase. Summary of the invention

[0005] In view of the above shortcomings, the present invention proposes a method for rapid deployment of Spark clusters, which utilizes unified management of cluster nodes, reasonably divides cluster nodes according to the needs of deployment of each component, and utilizes integrated cluster batch operation related commands referred to as ck and initialization scripts. Deployers only need to configure the cluster IP, execute the initialization script on the management node to perform some routine configurations, such as closing the firewall, optimizing the operating system, etc., and then use the integrated ck batch operation to deploy the Spark cluster. The deployment of the Spark cluster can be completed in a few simple steps, which can greatly save deployment time, does not require very high professional skills of the deployment personnel, can also save costs, can be quickly copied in the project, and improve construction efficiency. The specific steps are as follows:

[0006] Use the service planning template to configure the corresponding node IP, and the service planning template is deployed by the Spark cluster, Hadoop cluster and Zookeeper cluster;

[0007] Use the initSystem.sh initialization script to initialize the operating system;

[0008] Execute the initSystem.sh initialization script to perform one-click deployment of the Spark cluster, Hadoop cluster, and Zookeeper cluster. The initSystem.sh initialization script uses the integrated ck batch operation command and the configuration of / etc / cluster.conf to configure, distribute, and initialize the cluster deployment package;

[0009] The cluster startup is integrated into the startCluster.sh script, and the startCluster.sh script is executed to complete the startup of the Spark cluster, Hadoop cluster, and Zookeeper cluster.

[0010] Furthermore, after the service planning template completes the node IP configuration, the configured service planning template is uploaded to the server / etc directory during deployment, and the parsing program parses the uploaded cluster deployment configuration template / etc / cluster.xls file to generate a temporary / etc / cluster.conf file.

[0011] Spark cluster deploys each node service planning template. Operators only need to configure the IP of the corresponding node according to the template, without having to consider the node allocation of each service within the cluster, which greatly simplifies the deployment steps.

[0012] Furthermore, the parsing program parses the / etc / cluster.conf temporary file again, and further utilizes the integrated ck batch operation command to perform cluster deployment operations.

[0013] The parser parses the temporary file / etc / cluster.conf generated by the deployment template, and then uses the integrated ck batch operation command to perform related cluster deployment operations, which can simplify the deployment operation and facilitate deployment.

[0014] Furthermore, the initialization operation includes closing the firewall, eliminating the need for encryption between cluster nodes, and optimizing the operating system for cluster high performance.

[0015] The prepared initialization script initSystem.sh is used to perform related initialization operations on the operating system, such as closing the firewall, eliminating the need for passwords between cluster nodes, and optimizing the operating system for cluster high performance, which simplifies deployment operations.

[0016] Run the installCluster.sh script to deploy Zookeeper, Hadoop, and Spark cluster-related services with one click. The script uses the integrated ck batch operation command and the / etc / cluster.conf configuration to configure, distribute, and initialize each cluster deployment package.

[0017] The initialization operation also includes environment preparation, adjusting the operating system versions of all nodes to be consistent, configuring the internal network, and enabling the nodes to communicate with each other through host names or IP addresses. SSH is used to avoid passwords between cluster nodes.

[0018] The specific steps of configuring the cluster deployment package are as follows: configuring the slaves node of the Hadoop cluster; after the configuration is completed, copying the configured cluster deployment package to each node, and performing the cpush operation on the management node through the integrated ck batch operation.

[0019] The slaves node is the core configuration file, which can list the host names or IP addresses of all Worker nodes.

[0020] The core configuration file also needs to set environment variables and define default parameters of the cluster during configuration.

[0021] Cluster startup is integrated into the startCluster.sh script. Executing this script can complete the startup of the Zookeeper cluster, Hadoop cluster, and Spark cluster.

[0022] According to a second aspect of the present invention, a computer program product is provided, on which one or more computer programs are stored. When the one or more computer programs are executed by a computer processor, the above method is implemented.

[0023] The above one or more technical solutions in the embodiments of the present application have at least one of the following technical effects:

[0024] The present invention uses cluster node configuration files and highly integrated scripts to quickly deploy Hadoop clusters using batch operations in one click, thereby reducing the replicability of cluster deployment, reducing productivity, and being able to respond to project implementation work in a timely and rapid manner, avoiding labor waste; specifically, cluster deployment templates and centralized batch operation ck commands are mainly used for one-click rapid deployment. In order to simplify the Spark deployment process and the relevant technical reserves of operators, relevant operators only need to configure the corresponding node IP according to the deployment template to achieve rapid deployment.

[0025] In addition, the present invention uses an Excel deployment template to allocate node IPs of the main roles of the Spark cluster for deployment, which makes the configuration clearer and simpler for operators.

[0026] The present invention is applicable to people who have little knowledge of Spark technology and need to deploy Spark to calculate data offline. It is also applicable to areas with many projects and great application prospects in areas such as rapid replication during project implementation, improved delivery time, and savings in productivity and labor costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The accompanying drawings illustrate the embodiments and are used together with the description to explain the principles of the present invention. It will be easy to recognize other embodiments and many expected advantages of the embodiments because they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale with each other. The same reference numerals refer to corresponding similar parts.

[0028] Figure 1 A schematic flow chart of a method for rapid deployment of a Spark cluster according to an embodiment of the present invention is shown.

[0029] Figure 2 It is a structural diagram of a computer system suitable for implementing an electronic device of an embodiment of the present application. DETAILED DESCRIPTION

[0030] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It should also be noted that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0031] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0032] Figure 1A schematic diagram of a method for rapid deployment of a Spark cluster according to an embodiment of the present invention is shown. Figure 1 As shown:

[0033] S1. Use the service planning template to configure the corresponding node IP, and the service planning template is jointly deployed by the Spark cluster, Hadoop cluster and Zookeeper cluster;

[0034] Spark cluster deployment depends on Hadoop cluster and Zookeeper cluster. The three clusters can share the same set of cluster nodes or be deployed separately, depending on the scale of the specific business and server conditions. At the same time, there are multiple services running on different nodes, so it is necessary to plan the Zookeeper cluster and Hadoop service nodes in advance.

[0035] The present invention provides a Spark cluster deployment service planning template for each node. The operator only needs to configure the IP address of the corresponding node according to the template, and does not need to consider the node allocation problem of each service within the cluster, which greatly simplifies the deployment steps.

[0036] In this embodiment, there are five nodes with IP addresses 192.168.1.1-5, which are used to deploy the Spark cluster. The dependent Zookeeper cluster and Hadoop cluster are deployed at the same time, and the cluster nodes are shared. The allocation of each cluster node is as follows:

[0037] Table 1 Spark cluster node configuration

[0038]

[0039] Among them, the role service port is the default port and can be filled in and modified. For high availability, the Master role is generally configured with an odd number of nodes, with 3 nodes by default.

[0040] Table 2 Zookeeper cluster node configuration

[0041]

[0042] Among them, the role service port is the default port and can be filled in and modified. Zookeeper is generally configured with an odd number of nodes, and the default is 3 nodes.

[0043] Table 3 Hadoop cluster node configuration

[0044]

[0045]

[0046] Among them, the role port is the default port and can be modified by yourself.

[0047] The above configurations are for each cluster. When deploying, the deployer only needs to upload the configured template to the / etc directory of the server. The parser will parse the uploaded cluster deployment configuration template / etc / cluster.xls file and generate / etc / cluster.conf for easy parsing of ck batch operations. The configuration files are as follows:

[0048]

[0049] The parser parses the temporary file / etc / cluster.conf generated by the deployment template, and then uses the integrated ck batch operation command to perform related cluster deployment operations, which can simplify the deployment operation and facilitate deployment.

[0050] S2. Use the initSystem.sh initialization script to initialize the operating system;

[0051] The prepared initialization script initSystem.sh is used to perform related initialization operations on the operating system, such as closing the firewall, eliminating the need for passwords between cluster nodes, and optimizing the operating system for cluster high performance, which simplifies deployment operations.

[0052] Various environmental preparations are also required during initialization, including setting various environment variables, adjusting the operating system versions of all nodes to be consistent, configuring the internal network, enabling nodes to communicate with each other through host names or IP addresses, setting up SSH password-free login from the master node to all slave nodes, and simplifying subsequent command execution.

[0053] S3. Execute the initSystem.sh initialization script to perform one-click deployment on the Spark cluster, Hadoop cluster, and Zookeeper cluster. The initSystem.sh initialization script uses the integrated ck batch operation command and the configuration of / etc / cluster.conf to configure, distribute, and initialize the cluster deployment package;

[0054] Execute the installCluster.sh script to deploy Zookeeper, Hadoop, and Spark cluster-related services with one click. The script uses the integrated ck batch operation command and the / etc / cluster.conf configuration to configure, distribute, and initialize each cluster deployment package. The relevant configuration operations are as follows:

[0055] Configure the slaves of the Hadoop cluster:

[0056] ck hadoop_datanode:hostname|grep-v"\*\*"|grep-v"\-\-"> / usr / local / hadoop / etc / hadoop / slaves;

[0057] As the core configuration file, slaves can list the host names or IP addresses of all Worker nodes. The core configuration file also includes setting environment variables and defining the default parameters of the cluster, among which environment variables such as SPARK_HOME, JAVA_HOME, etc.; the spark-defaults.conf file defines Spark default parameters such as spark.master, spark.executor.memory, etc.

[0058] After the configuration is completed, the deployment packages of each cluster will be copied to each node and executed on the management node through the integrated ck batch operation command cpush. The specific steps are as follows:

[0059] Synchronize the Spark service directory / usr / local / spark-3.4.0 to the / usr / local / of the installation node: cpush spark_worker: / usr / local / spark-3.4.0 / usr / local / spark / usr / local / ;

[0060] Among them, / usr / local / spark is the soft link directory of / usr / local / spark-3.4.0;

[0061] After the software package distribution is completed, the relevant initialization operations will be performed on each cluster. The specific steps are as follows:

[0062] Write 0 to the / data / zk / myid file of the first IP node of node zookeeper in the / etc / cluster.conf configuration: ck zookeeperk:0 "echo 0> / data / zk / myid";

[0063] The script installCluster.sh configures, distributes, and initializes the relevant clusters according to the configured / etc / cluster.conf, completing the deployment of the configured cluster.

[0064] S4. Cluster startup is integrated into the startCluster.sh script, and the startCluster.sh script is executed to complete the startup of the Spark cluster, Hadoop cluster, and Zookeeper cluster. The specific steps are as follows:

[0065] The Zookeeper component does not have its own batch startup operation. Each node that needs to be deployed needs to be started. This script uses some batch operations of ck that have been integrated in advance to complete the startup of each Zookeeper node: ck zk: / usr / local / zk / bin / zkServer.sh start;

[0066] Start the Hadoop cluster using its own startup script:

[0067] / usr / local / hadoop / sbin / start-dfs.sh;

[0068] Spark cluster starts, using its own startup script:

[0069] / usr / local / spark / sbin / start-all.sh;

[0070] Reference below Figure 2 , which shows a schematic diagram of the structure of a computer system 200 suitable for implementing an electronic device of an embodiment of the present application. Figure 2 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0071] like Figure 2 As shown, the computer system 200 includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 202 or a program loaded from a storage part 208 into a random access memory (RAM) 203. In the RAM 203, various programs and data required for the operation of the system 200 are also stored. The CPU 201, the ROM 202, and the RAM 203 are connected to each other via a bus 204. An input / output (I / O) interface 205 is also connected to the bus 204.

[0072] The following components are connected to the I / O interface 205: an input section 206 including a keyboard, a mouse, etc.; an output section 207 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 208 including a hard disk, etc.; and a communication section 209 including a network interface card such as a LAN card, a modem, etc. The communication section 209 performs communication processing via a network such as the Internet. A drive 210 is also connected to the I / O interface 205 as needed. A removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 210 as needed, so that a computer program read therefrom is installed into the storage section 208 as needed.

[0073] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 209, and / or installed from the removable medium 211. When the computer program is executed by the central processing unit (CPU) 201, the above functions defined in the method of the present application are executed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wireless, wireline, optical cable, RF, etc., or any suitable combination of the foregoing.

[0074] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0075] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0076] The modules involved in the embodiments of the present application may be implemented by software or by hardware.

[0077] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or it may exist independently and not be assembled into the electronic device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device: configures the corresponding node IP using the service planning template, and the service planning template is jointly deployed by the Spark cluster, the Hadoop cluster and the Zookeeper cluster; initializes the operating system using the initSystem.sh initialization script; executes the initSystem.sh initialization script to perform one-click deployment of the Spark cluster, the Hadoop cluster and the Zookeeper cluster, and the initSystem.sh initialization script uses the integrated ck batch operation command and the configuration of / etc / cluster.conf to configure, distribute and initialize the cluster deployment package; the cluster startup is integrated into the startCluster.sh script, and the startCluster.sh script is executed to complete the startup of the Spark cluster, the Hadoop cluster and the Zookeeper cluster.

[0078] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above invention concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.

Claims

1. A method for rapid deployment of a Spark cluster, characterized in that: include: Use the service planning template to configure the corresponding node IP, and the service planning template is deployed by the Spark cluster, Hadoop cluster and Zookeeper cluster; Use the initSystem.sh initialization script to initialize the operating system; Execute the initSystem.sh initialization script to perform one-click deployment of the Spark cluster, Hadoop cluster, and Zookeeper cluster. The initSystem.sh initialization script uses the integrated ck batch operation command and the configuration of / etc / cluster.conf to configure, distribute, and initialize the cluster deployment package; The cluster startup is integrated into the startCluster.sh script, and the startCluster.sh script is executed to complete the startup of the Spark cluster, Hadoop cluster, and Zookeeper cluster.

2. The method according to claim 1, characterized in that After the service planning template completes the node IP configuration, the configured service planning template is uploaded to the server / etc directory during deployment, and the parser parses the uploaded cluster deployment configuration template / etc / cluster.xls file to generate a temporary / etc / cluster.conf file.

3. The method according to claim 2, characterized in that The parsing program parses the / etc / cluster.conf temporary file again, and further uses the integrated ck batch operation command to perform cluster deployment operations.

4. The method according to claim 1, characterized in that: The initialization operation includes closing the firewall, eliminating the need for encryption between cluster nodes, and optimizing the operating system for cluster high performance.

5. The method according to claim 4, characterized in that The initialization operation also includes environment preparation, adjusting the operating system versions of all nodes to be consistent, configuring the internal network, and enabling the nodes to communicate with each other through host names or IP addresses. SSH is used to avoid passwords between cluster nodes.

6. The method according to claim 1, characterized in that The specific steps for configuring the cluster deployment package are as follows: Configure the slaves nodes of the Hadoop cluster; after the configuration is completed, copy the configured cluster deployment package to each node, and perform the cpush operation on the management node through the integrated ck batch operation.

7. The method according to claim 6, characterized in that The slaves node is the core configuration file, which can list the host names or IP addresses of all Worker nodes.

8. The method according to claim 7, characterized in that The core configuration file also needs to set environment variables and define default parameters of the cluster during configuration.

9. A computer program product, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

10. A computing system, characterized in that: The method comprises a processor and a memory, wherein the processor is configured to execute the method according to any one of claims 1 to 8.