A configuration method of a big data cluster and a docking method of a big data platform
By generating cluster configuration information for big data clusters, the problems of high complexity and large manual input when connecting to different big data platforms are solved, thus simplifying the connection process and reducing costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU SHUMEI TECHNOLOGY CO LTD
- Filing Date
- 2022-12-26
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies present challenges in connecting to different big data platforms, including high complexity and significant manual labor costs.
By obtaining the configuration file of the platform to be connected, the cluster configuration information of the big data cluster is generated, and the cluster components are configured based on this information, simplifying the connection process.
Cluster configuration information is generated during the initial connection, and subsequent connections only require this information to complete the connection, which simplifies the complexity of big data platform connection and reduces manual input costs.
Smart Images

Figure CN116010088B_ABST
Abstract
Description
Configuration methods for big data clusters and integration methods with big data platforms Technical Field
[0001] This application relates to the field of big data technology, and in particular to the configuration methods of big data clusters and the docking methods of big data platforms. Background Technology
[0002] With the continuous updating and iteration of big data technology, when upper-layer applications use big data platforms to run computing tasks, different vendors' big data platforms often need to be processed differently. This means that the parameters of multiple configuration files for each component of the big data platform need to be modified and compared. Summary of the Invention
[0003] This application aims to at least partially address one of the technical problems in the related art.
[0004] Therefore, the first objective of this application is to propose a configuration method for a big data cluster. This method can generate cluster configuration information for the big data cluster based on the configuration file corresponding to the platform to be connected, and configure the cluster components included in the big data cluster based on this cluster configuration information. Thus, during the initial connection, the configuration process with the platform to be connected can be saved by generating cluster configuration information. Subsequent connections only require the cluster configuration information to complete the connection, simplifying the complexity of big data platform connection and reducing manual input costs.
[0005] The second objective of this application is to propose a method for connecting to a big data platform.
[0006] The third objective of this application is to propose a configuration device for a big data cluster.
[0007] The fourth objective of this application is to propose a docking device for a big data platform.
[0008] The fifth objective of this application is to provide an electronic device.
[0009] The sixth objective of this application is to provide a computer-readable storage medium.
[0010] The seventh objective of this application is to provide a computer program product.
[0011] To achieve the above objectives, the first aspect of this application proposes a method for configuring a big data cluster, including:
[0012] Obtain the configuration file corresponding to the platform to be integrated;
[0013] The configuration information for the big data cluster is generated based on the configuration file; the big data cluster is used for data processing services.
[0014] Based on the cluster configuration information, the cluster components included in the big data cluster are configured.
[0015] Optionally, as a first possible implementation of the first aspect, generating cluster configuration information for the big data cluster based on the configuration file includes:
[0016] Based on the configuration items in the configuration file corresponding to the platform to be connected, replace the corresponding configuration items in the application's configuration file to obtain the cluster configuration information of the big data cluster.
[0017] Alternatively, as a second possible implementation of the first aspect, the configuration item is in the form of a key-value pair;
[0018] The key in the key-value pair is used to indicate a component in the big data cluster;
[0019] The value in the key-value pair is used to indicate the calling port.
[0020] Alternatively, as a third possible implementation of the first aspect, the method further includes:
[0021] Obtain the authentication information corresponding to the platform to be connected;
[0022] The configuration of the cluster components included in the big data cluster based on the cluster configuration information includes:
[0023] In response to user operation, based on the cluster configuration information, the authentication information and the target resource information of the component to be configured are sent to the big data cluster, so that if the big data cluster passes the authentication information, the target resource of the component is created and configured in the big data cluster according to the target resource information.
[0024] Alternatively, as a fourth possible implementation of the first aspect, the information of the target resource of the component to be configured includes at least one of the following:
[0025] Location of the HDFS component's storage directory;
[0026] The upper limit of storage space for the storage directory of the HDFS component;
[0027] The maximum amount of computing resources that can be used in the computing resource queue of YARN components;
[0028] The upper limit of storage space for the database tables of the HIVE component.
[0029] Optionally, as a fifth possible implementation of the first aspect, obtaining the configuration file corresponding to the platform to be connected includes:
[0030] In response to the selection operation, determine the type and / or version information of the platform to be connected;
[0031] Based on the type and / or version information, determine the configuration file corresponding to the platform to be connected from the pre-stored configuration files.
[0032] To achieve the above objectives, a second aspect of this application proposes a method for connecting to a big data platform, comprising:
[0033] Obtain the cluster configuration information corresponding to the platform to be connected; the cluster configuration information is generated based on the big data cluster configuration method described in the first aspect;
[0034] Based on the cluster configuration information, the system connects to the cluster components included in the big data cluster to complete the connection with the platform to be connected.
[0035] Optionally, as a first possible implementation of the second aspect, the step of connecting to the cluster components included in the big data cluster based on the cluster configuration information to complete the connection with the platform to be connected includes:
[0036] Replace the configuration information in the application used to connect to the cluster components in the big data cluster with the cluster configuration information to complete the connection with the platform to be connected.
[0037] Alternatively, as a second possible implementation of the second aspect, the method further includes:
[0038] In response to the received operation command, the cluster component being connected completes the processing of the data in the platform to be connected.
[0039] Optionally, as a third possible implementation of the second aspect, obtaining the cluster configuration information corresponding to the platform to be connected includes:
[0040] Obtain the cluster configuration information corresponding to the platform to be integrated through the interface; or...
[0041] Download the cluster configuration information corresponding to the platform to be integrated.
[0042] To achieve the above objectives, a third aspect of this application provides a configuration apparatus for a big data cluster, comprising:
[0043] The first acquisition module is used to acquire the configuration file corresponding to the platform to be connected;
[0044] A generation module is used to generate cluster configuration information for a big data cluster based on the configuration file; the big data cluster is used for data processing services.
[0045] The configuration module is used to configure the cluster components included in the big data cluster based on the cluster configuration information.
[0046] Optionally, as a first possible implementation of the third aspect, the generation module is further configured to:
[0047] Based on the configuration items in the configuration file corresponding to the platform to be connected, replace the corresponding configuration items in the application's configuration file to obtain the cluster configuration information of the big data cluster.
[0048] Alternatively, as a second possible implementation of the third aspect, the configuration item is in the form of a key-value pair;
[0049] The key in the key-value pair is used to indicate a component in the big data cluster;
[0050] The value in the key-value pair is used to indicate the calling port.
[0051] Alternatively, as a third possible implementation of the third aspect, the apparatus further includes:
[0052] The second acquisition module is used to acquire the authentication information corresponding to the platform to be connected;
[0053] The configuration module is also used for:
[0054] In response to user operation, based on the cluster configuration information, the authentication information and the target resource information of the component to be configured are sent to the big data cluster, so that if the big data cluster passes the authentication information, the target resource of the component is created and configured in the big data cluster according to the target resource information.
[0055] Alternatively, as a fourth possible implementation of the third aspect, the information of the target resource of the component to be configured includes at least one of the following:
[0056] Location of the HDFS component's storage directory;
[0057] The upper limit of storage space for the storage directory of the HDFS component;
[0058] The maximum amount of computing resources that can be used in the computing resource queue of YARN components;
[0059] The upper limit of storage space for the database tables of the HIVE component.
[0060] Optionally, as a fifth possible implementation of the third aspect, the first acquisition module further includes:
[0061] In response to the selection operation, determine the type and / or version information of the platform to be connected;
[0062] Based on the type and / or version information, determine the configuration file corresponding to the platform to be connected from the pre-stored configuration files.
[0063] To achieve the above objectives, a fourth aspect of this application provides a docking device for a big data platform, comprising:
[0064] The third acquisition module is used to acquire cluster configuration information corresponding to the platform to be connected; the cluster configuration information is generated based on the big data cluster configuration method described in the first aspect;
[0065] The docking module is used to dock with the cluster components contained in the big data cluster based on the cluster configuration information, and to complete the docking with the platform to be docked.
[0066] Alternatively, as a first possible implementation of the fourth aspect, the docking module is further configured to:
[0067] Replace the configuration information in the application used to connect to the cluster components in the big data cluster with the cluster configuration information to complete the connection with the platform to be connected.
[0068] Alternatively, as a second possible implementation of the fourth aspect, the apparatus further includes:
[0069] The processing module is used to respond to the received operation instructions and complete the processing of data in the platform to be connected based on the cluster component being connected.
[0070] Optionally, as a third possible implementation of the fourth aspect, the third acquisition module is further configured to:
[0071] Obtain the cluster configuration information corresponding to the platform to be integrated through the interface; or...
[0072] Download the cluster configuration information corresponding to the platform to be integrated.
[0073] To achieve the above objectives, a fifth aspect of this application provides an electronic device comprising:
[0074] At least one processor; and
[0075] A memory communicatively connected to the at least one processor; wherein,
[0076] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the configuration method of the big data cluster described in the first aspect, or to perform the docking method of the big data platform described in the second aspect.
[0077] To achieve the above objectives, a sixth aspect of this application provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the configuration method of the big data cluster of the first aspect, or to execute the docking method of the big data platform of the second aspect.
[0078] To achieve the above objectives, a seventh aspect of this application provides a computer program product, characterized in that it includes a computer program, which, when executed by a processor, implements either the configuration method for the big data cluster described in the first aspect or the docking method for the big data platform described in the second aspect.
[0079] The technical solutions provided in this application have the following beneficial effects:
[0080] By obtaining the configuration file corresponding to the platform to be connected, the cluster configuration information of the big data cluster is generated based on the configuration file. The big data cluster is used for data processing services, and the cluster components included in the big data cluster are configured based on this configuration information. Since the cluster configuration information of the big data cluster can be generated based on the configuration file corresponding to the platform to be connected, and the cluster components included in the big data cluster can be configured based on this configuration information, the configuration process with the platform to be connected can be saved during the initial connection. Subsequent connections only require the cluster configuration information to complete the connection, simplifying the complexity of big data platform connection and reducing manual input costs.
[0081] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0082] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0083] Figure 1 is a flowchart illustrating a configuration method for a big data cluster provided in an embodiment of this application;
[0084] Figure 2 is a flowchart illustrating another method for configuring a big data cluster provided in an embodiment of this application;
[0085] Figure 3 is a flowchart illustrating a big data platform docking method provided in an embodiment of this application;
[0086] Figure 4 is a flowchart illustrating another big data platform docking method provided in the embodiments of this application;
[0087] Figure 5 is a schematic diagram of the configuration of a big data cluster and the process of connecting to a big data platform in a scenario provided by an embodiment of this application;
[0088] Figure 6 is a schematic diagram of the configuration device for a big data cluster provided in an embodiment of this application;
[0089] Figure 7 is a schematic diagram of another configuration device for a big data cluster provided in an embodiment of this application;
[0090] Figure 8 is a schematic diagram of the structure of a docking device for a big data platform provided in an embodiment of this application;
[0091] Figure 9 is a schematic diagram of the structure of another big data platform docking device provided in an embodiment of this application;
[0092] Figure 10 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0093] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0094] In related technologies, the following approach is often used to connect to different big data platforms: First, collect the configuration files and authentication files of all components in the big data platform to be connected. Then, modify each configuration file to be adapted. Finally, create the resources required for computing via commands. This process requires significant manpower due to subtle differences between different versions of different platforms, increasing the complexity and labor costs of the connection process.
[0095] This application addresses the problems of high complexity and high manual input costs in connecting different big data platforms in related technologies by proposing a configuration method for big data clusters and a connection method for big data platforms.
[0096] The following describes, with reference to the accompanying drawings, a method for configuring a big data cluster and a method for connecting to a big data platform, according to embodiments of this application.
[0097] Figure 1 is a flowchart illustrating a configuration method for a big data cluster provided in an embodiment of this application.
[0098] As shown in Figure 1, the configuration method for this big data cluster includes the following steps:
[0099] Step 101: Obtain the configuration file corresponding to the platform to be connected.
[0100] In this embodiment, the platform to be connected can be any platform, including but not limited to big data platforms.
[0101] It should be noted that the execution entity in this embodiment is the engine management component. In this embodiment, the engine management component can obtain the configuration file corresponding to the platform to be connected through various public, legal and compliant means. For example, the engine management component can collect the configuration file corresponding to the platform to be connected in real time, or it can obtain the configuration file corresponding to the platform to be connected from other devices through network transmission or physical copying, or it can obtain the configuration file corresponding to the platform to be connected through other public, legal and compliant means. This embodiment does not impose any restrictions on this.
[0102] In one possible implementation of this embodiment, the configuration file corresponding to the platform to be connected, obtained by the engine management component, can be a component configuration file within the Hadoop ecosystem. Hadoop is a software framework capable of distributed processing of large amounts of data, providing a reliable, efficient, and scalable data processing solution. It implements a distributed file system with high fault tolerance and offers high throughput access to application data, making it suitable for applications with extremely large datasets. Data in the file system can be accessed in a streaming manner.
[0103] Step 102: Generate cluster configuration information for the big data cluster based on the configuration file, wherein the big data cluster is used for data processing services.
[0104] In this embodiment, after obtaining the configuration file corresponding to the platform to be connected, the engine management component can generate cluster configuration information for the big data cluster based on the configuration file. The big data cluster is used for data processing services. Optionally, the data processing services may include computing, storage, etc.
[0105] In one possible implementation of this embodiment, when the configuration file corresponding to the platform to be connected is a component configuration file within the Hadoop system, the big data cluster generated based on the configuration file is the Hadoop cluster, and the cluster configuration information of the generated big data cluster is the cluster configuration information of the Hadoop cluster.
[0106] Step 103: Configure the cluster components included in the big data cluster based on the cluster configuration information.
[0107] In this embodiment, the engine management component can configure the cluster components included in the big data cluster based on the cluster configuration information of the generated big data cluster. Optionally, the cluster components included in the big data cluster may include HDFS (Hadoop Distributed File System), YARN (Yet Another Resource Negotiator), HIVE, Kafka, etc. Among them, HDFS is characterized by high fault tolerance and is designed to be deployed on low-cost hardware, providing high throughput access to application data, and is suitable for applications with very large datasets. YARN is a new Hadoop resource manager, a general-purpose resource management system that provides unified resource management and scheduling for upper-layer applications. HIVE is a data warehouse tool based on Hadoop, used for data extraction, transformation, and loading.
[0108] The big data cluster configuration method provided in this embodiment obtains the configuration file corresponding to the platform to be connected, and generates cluster configuration information for the big data cluster based on the configuration file. The big data cluster is used for data processing services, and the cluster components included in the big data cluster are configured based on the cluster configuration information. Since the cluster configuration information can be generated based on the configuration file corresponding to the platform to be connected, and the cluster components included in the big data cluster can be configured based on this cluster configuration information, the configuration process with the platform to be connected can be saved during the initial connection. Subsequent connections only require the cluster configuration information to complete the connection, simplifying the complexity of big data platform connection and reducing manual input costs.
[0109] To clearly illustrate the previous embodiment, this embodiment provides another method for configuring a big data cluster. Figure 2 is a flowchart illustrating another method for configuring a big data cluster provided in this embodiment.
[0110] As shown in Figure 2, the configuration method for this big data cluster may include the following steps:
[0111] Step 201: In response to the selection operation, determine the type and / or version information of the platform to be connected.
[0112] In this embodiment, the platform to be connected can be any platform, including but not limited to big data platforms.
[0113] It should be noted that the execution entity in this embodiment is also the engine management component. In this embodiment, the engine management component has pre-collected various types and / or versions of configuration files, so that the configuration file corresponding to the platform to be connected can be determined by the platform type and / or version information selected by the user.
[0114] In this embodiment, the engine management component can determine the type and / or version information of the platform to be connected in response to the user's selection operation.
[0115] Step 202: Based on the type and / or version information, determine the configuration file corresponding to the platform to be connected from the pre-stored configuration files.
[0116] In this embodiment, after the engine management component determines the type and / or version information of the platform to be connected in response to the user's selection operation, it can determine the configuration file corresponding to the platform to be connected from the pre-stored configuration file based on the type and / or version information of the platform to be connected.
[0117] Step 203: Obtain the authentication information corresponding to the platform to be connected.
[0118] In this embodiment, the engine management component can also obtain the authentication information corresponding to the platform to be connected, so as to perform authentication based on the authentication information.
[0119] In one possible implementation of this embodiment, obtaining the authentication information corresponding to the platform to be connected is equivalent to obtaining the authentication file corresponding to the platform to be connected. Optionally, the engine management component can pre-collect various types of authentication files, thereby determining the authentication file corresponding to the platform to be connected based on the authentication type selected by the user. For example, the engine management component can respond to the user's selection operation, determine the type of authentication file, and then, based on the authentication type, determine the authentication file corresponding to the platform to be connected from the pre-stored authentication files.
[0120] Step 204: Replace the corresponding configuration items in the application's configuration file with the configuration items in the configuration file of the platform to be connected, so as to obtain the cluster configuration information of the big data cluster.
[0121] It should be noted that, in this embodiment, the configuration file corresponding to the platform to be connected includes at least one configuration item, so that the corresponding configuration item in the application's configuration file can be replaced according to the configuration item in the configuration file corresponding to the platform to be connected, in order to obtain the cluster configuration information of the big data cluster.
[0122] In one possible implementation of this embodiment, the configuration items in the configuration file corresponding to the platform to be connected can be in key-value pair format. The key in the key-value pair indicates a component in the big data cluster, and the value indicates the calling port. Optionally, the components in the big data cluster may include HDFS, YARN, HIVE, Kafka, etc. For example, configuration items in XML (eXtensible Markup Language) format are in key-value pair format.
[0123] Step 205: In response to the user operation, based on the cluster configuration information, send authentication information and target resource information of the component to be configured to the big data cluster, so that if the big data cluster passes the authentication information, it can create and configure the target resource of the component in the big data cluster according to the target resource information.
[0124] In this embodiment, in response to a user operation, based on the cluster configuration information of the big data cluster, the authentication information corresponding to the platform to be connected and the target resource information of the component to be configured can be sent to the big data cluster. This allows the big data cluster to create and configure the target resource of the component based on the target resource information if the authentication information is approved. Optionally, the target resource information of the component to be configured includes at least one of the following:
[0125] Location of the HDFS component's storage directory;
[0126] The storage space limit of the storage directory of the HDFS component;
[0127] The maximum amount of computing resources that can be used in the computing resource queue of YARN components;
[0128] The upper limit of storage space for the database tables of the HIVE component.
[0129] The big data cluster configuration method provided in this embodiment determines the type and / or version information of the platform to be connected in response to a selection operation. Based on this type and / or version information, it identifies the corresponding configuration file for the platform from a pre-stored configuration file, thereby obtaining the authentication information for the platform. Then, based on the configuration items in the platform's configuration file, it replaces the corresponding configuration items in the application's configuration file to obtain the cluster configuration information for the big data cluster. Subsequently, in response to user operations, based on the cluster configuration information, it sends authentication information and target resource information of the component to be configured to the big data cluster. If the authentication information is approved, the big data cluster creates and configures the target resource of the component based on the target resource information. This method effectively masks the differences between different platforms, enabling unified control over platform configuration, simplifying the complexity of big data platform integration, and reducing manual input costs.
[0130] The above embodiments describe the implementation process of the configuration method of the big data cluster from the perspective of the engine management component. The following describes a possible implementation method of the big data platform docking method from the perspective of the application side, with reference to Figure 3. Figure 3 is a flowchart of a big data platform docking method provided by the embodiment of this application.
[0131] As shown in Figure 3, the connection method of this big data platform may include the following steps:
[0132] Step 301: Obtain the cluster configuration information corresponding to the platform to be connected.
[0133] In this embodiment, the platform to be connected can be any platform, including but not limited to big data platforms.
[0134] It should be noted that the execution entity in this embodiment is the application side, which may be the front end or the back end, and this embodiment does not limit it.
[0135] In this embodiment, the application can obtain cluster configuration information corresponding to the platform to be connected. This cluster configuration information can be generated based on the big data cluster configuration method provided in this application embodiment.
[0136] In one possible implementation of this embodiment, the application can obtain the cluster configuration information corresponding to the platform to be connected through an interface, or directly download the cluster configuration information corresponding to the platform to be connected.
[0137] Step 302: Connect to the cluster components included in the big data cluster based on the cluster configuration information to complete the connection with the platform to be connected.
[0138] In this embodiment, after obtaining the cluster configuration information corresponding to the platform to be connected, the application can connect to the cluster components included in the big data cluster based on the cluster configuration information, thereby completing the connection with the platform to be connected. Optionally, the cluster components may include HDFS, YARN, HIVE, Kafka, etc.
[0139] The big data platform integration method provided in this embodiment obtains the cluster configuration information corresponding to the platform to be integrated, and then integrates the cluster components contained in the big data cluster based on the cluster configuration information to complete the integration with the platform to be integrated. Since only the cluster configuration information corresponding to the platform to be integrated needs to be obtained during the integration process, the complexity of big data platform integration is simplified and the manual input cost is reduced.
[0140] To clearly illustrate the previous embodiment, this embodiment provides another method for connecting to a big data platform. Figure 4 is a flowchart illustrating another method for connecting to a big data platform provided in this application embodiment.
[0141] As shown in Figure 4, the connection method of this big data platform may include the following steps:
[0142] Step 401: Obtain the cluster configuration information corresponding to the platform to be connected.
[0143] It should be noted that the execution process of this step can participate in the execution process of step 301 in the previous embodiment, and the principle is the same, so it will not be repeated here.
[0144] Step 402: Replace the configuration information in the application used to connect to the cluster components in the big data cluster with the cluster configuration information to complete the connection with the platform to be connected.
[0145] In this embodiment, the application can replace the configuration information used in the application to connect to the cluster components in the big data cluster with cluster configuration information to complete the connection with the platform to be connected.
[0146] Step 403: In response to the received operation command, process the data in the platform to be connected based on the connected cluster components.
[0147] In this embodiment, after completing the connection with the platform to be connected, the application can also respond to the received operation instructions and process the data in the platform based on the connected cluster components. For example, the application can upload resource files to HDFS for storage, or submit Spark tasks to the cluster to use the cluster's storage and computing resources for data development, and so on. Spark is a fast and general-purpose computing engine designed for large-scale data processing. It can be used to perform a variety of operations, including SQL queries, text processing, machine learning, etc., and supports interactive computing and complex algorithms.
[0148] The big data platform integration method provided in this embodiment obtains the cluster configuration information corresponding to the platform to be integrated. It then replaces the configuration information used in the application to integrate the cluster components within the big data cluster with this new cluster configuration information, thus completing the integration with the platform. In response to received operation commands, it processes the data from the platform based on the integrated cluster components. Therefore, by replacing the configuration information used in the application to integrate the big data cluster with the new cluster configuration information, the integration with the platform is completed, simplifying the complexity of big data platform integration and reducing manual input costs. Furthermore, after integration, it can process the data from the platform based on the integrated cluster components upon receiving operation commands.
[0149] To clearly illustrate the above-disclosed embodiments, examples are provided below.
[0150] As shown in Figure 5, the first step is to configure the big data cluster.
[0151] Configuring a big data cluster involves selecting the platform type, importing the configuration file for the platform to be integrated (component configuration files within the Hadoop ecosystem), selecting the authentication type, and importing the authentication file. This will then generate a big data cluster configuration set. Specifically, the steps are as follows:
[0152] S1: Collect configuration files for the big data platform, confirm the authentication type, and obtain the authentication file. For example, to connect to the xx big data platform, it is necessary to collect hdfs-site.xml, core-site.xml, yarn-site.xml, with the authentication type being Kerberos, and the corresponding authentication file user.keytab.
[0153] S2: Upload the collected configuration files to the engine management system. Select the corresponding big data platform type and version (including but not limited to FusionInsight, TBDS, HDP, CDH, etc.), and upload the collected configuration files in the order shown on the engine management page. After uploading, a complete cluster configuration will be generated internally by the system.
[0154] S3: Create corresponding computing and storage resources on different components through the system page. For example, create a storage directory on HDFS and set the maximum storage size of the directory; create a computing resource queue on YARN and set the maximum computing resources that the queue can use, etc.
[0155] Next, configure the cluster.
[0156] By selecting a big data platform already entered in the engine management system, you can use the resources associated with that platform. This can be done in two ways:
[0157] Method 1: Officially integrate the application into the system, and the system will automatically push platform configurations. At this time, you can also manage the resources used by the application and control permissions.
[0158] Method 2: Applications that are not connected to the system can download the platform configuration that has already been entered from the system page, and then manually connect using the downloaded configuration file.
[0159] Finally, use cluster resources for development.
[0160] Once an application is connected to a big data platform, it can use the resources configured in S2 to perform computation, storage, and other activities. For example, it can upload resource files to HDFS for storage; submit Spark tasks to the cluster and use the cluster's storage and computing resources for data development, etc.
[0161] In summary, engine management enables applications to adapt to big data platforms without requiring modifications to configuration files or resources via backend commands, simplifying the complexity of big data platform integration and reducing manual input costs.
[0162] Corresponding to the configuration method of the big data cluster provided in the embodiments of Figures 1 and 2 above, this disclosure also provides a configuration device for the big data cluster. Since the configuration device for the big data cluster provided in this application corresponds to the configuration method for the big data cluster provided in the embodiments of Figures 1 and 2 above, the implementation method of the configuration method for the big data cluster is also applicable to the configuration device for the big data cluster provided in this application, and will not be described in detail in this application.
[0163] Figure 6 is a schematic diagram of the configuration device for a big data cluster provided in an embodiment of this application.
[0164] As shown in Figure 6, the configuration device for the big data cluster may include: a first acquisition module 61, a generation module 62, and a configuration module 63.
[0165] The first acquisition module 61 is used to acquire the configuration file corresponding to the platform to be connected.
[0166] Module 62 is used to generate cluster configuration information for the big data cluster based on the configuration file; the big data cluster is used for data processing services.
[0167] Configuration module 63 is used to configure the cluster components contained in the big data cluster based on cluster configuration information.
[0168] Furthermore, in one possible implementation of this application embodiment, the generation module 62 is further configured to:
[0169] Based on the configuration items in the configuration file of the platform to be connected, replace the corresponding configuration items in the application's configuration file to obtain the cluster configuration information of the big data cluster.
[0170] Furthermore, in one possible implementation of this application embodiment, the configuration item is in the form of a key-value pair;
[0171] In this key-value pair, the key is used to indicate a component in the big data cluster;
[0172] The value in the key-value pair is used to indicate the calling port.
[0173] Furthermore, in one possible implementation of this application embodiment, the first acquisition module 61 further includes:
[0174] In response to the selection action, determine the type and / or version information of the platform to be connected;
[0175] Based on type and / or version information, determine the configuration file corresponding to the platform to be integrated from the pre-stored configuration files.
[0176] Based on the above embodiments, this application also provides a possible implementation of a configuration device for a big data cluster. Figure 7 is a schematic diagram of another configuration device for a big data cluster provided by this application. Based on the previous embodiment, the configuration device for the big data cluster further includes: a second acquisition module 64.
[0177] The second acquisition module 64 is used to acquire the authentication information corresponding to the platform to be connected.
[0178] Furthermore, in one possible implementation of this application embodiment, the configuration module 63 is further configured to:
[0179] In response to user actions, based on cluster configuration information, authentication information and target resource information of the component to be configured are sent to the big data cluster. If the big data cluster approves the authentication information, it creates and configures the target resource of the component in the big data cluster according to the target resource information.
[0180] Furthermore, in one possible implementation of this application embodiment, the information of the target resource of the component to be configured includes at least one of the following:
[0181] Location of the HDFS component's storage directory;
[0182] The storage space limit of the storage directory of the HDFS component;
[0183] The maximum amount of computing resources that can be used in the computing resource queue of YARN components;
[0184] The upper limit of storage space for the database tables of the HIVE component.
[0185] The big data cluster configuration device provided in this embodiment obtains the configuration file corresponding to the platform to be connected, and generates cluster configuration information for the big data cluster based on the configuration file. The big data cluster is used for data processing services, and the cluster components included in the big data cluster are configured based on the cluster configuration information. Since the cluster configuration information can be generated based on the configuration file corresponding to the platform to be connected, and the cluster components included in the big data cluster can be configured based on this cluster configuration information, the configuration process with the platform to be connected can be saved during the initial connection. Subsequent connections only require the cluster configuration information to complete the connection, simplifying the complexity of big data platform connection and reducing manual input costs.
[0186] Corresponding to the big data platform docking method provided in the embodiments of Figures 3 and 4 above, this disclosure also provides a big data platform docking device. Since the big data platform docking device provided in this application corresponds to the big data platform docking method provided in the embodiments of Figures 1 and 2 above, the implementation method of the big data platform docking method is also applicable to the big data platform docking device provided in this application, and will not be described in detail in this application.
[0187] Figure 8 is a schematic diagram of the structure of a docking device for a big data platform provided in an embodiment of this application.
[0188] As shown in Figure 8, the docking device for the big data platform may include: a third acquisition module 81 and a docking module 82.
[0189] The third acquisition module 81 is used to acquire cluster configuration information corresponding to the platform to be connected; the cluster configuration information is basically generated by the configuration method of the big data cluster proposed in any of the above embodiments.
[0190] The docking module 82 is used to dock with the cluster components contained in the big data cluster based on the cluster configuration information, and to complete the docking with the platform to be docked.
[0191] Furthermore, in one possible implementation of this application embodiment, the docking module 82 is also used for:
[0192] Replace the configuration information in the application used to connect to the cluster components in the big data cluster with the cluster configuration information to complete the connection with the platform to be connected.
[0193] Furthermore, in one possible implementation of this application embodiment, the third acquisition module 81 is further configured to:
[0194] Obtain the cluster configuration information corresponding to the platform to be integrated through the interface; or...
[0195] Download the cluster configuration information corresponding to the platform to be integrated.
[0196] Based on the above embodiments, this application also provides a possible implementation of a big data platform docking device. Figure 8 is a schematic diagram of another big data platform docking device provided in this application. Based on the previous embodiment, the big data platform docking device further includes a processing module 83.
[0197] The processing module 83 is used to respond to the received operation instructions and complete the processing of data in the platform to be connected based on the connected cluster components.
[0198] The big data platform docking device provided in this embodiment obtains the cluster configuration information corresponding to the platform to be docked, and then docks with the cluster components contained in the big data cluster based on the cluster configuration information to complete the docking with the platform to be docked. Since only the cluster configuration information corresponding to the platform to be docked needs to be obtained to complete the docking process, the complexity of big data platform docking is simplified and the manual input cost is reduced.
[0199] To implement the above embodiments, this application also proposes an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to execute the configuration method of the big data cluster proposed in any of the above embodiments of this application, or to execute the docking method of the big data platform proposed in any of the above embodiments of this application.
[0200] Figure 10 is a schematic diagram of an electronic device provided in an embodiment of this application. It should be noted that the electronic device shown in Figure 10 is merely an example and should not impose any limitations on the function and scope of use of the embodiments of this application.
[0201] As shown in Figure 10, the electronic device may include: a housing 11, a processor 12, a memory 13, a circuit board 14, and a power supply circuit 15. The circuit board 14 is disposed inside the space enclosed by the housing 11, and the processor 12 and the memory 13 are disposed on the circuit board 14. The power supply circuit 15 is used to supply power to various circuits or devices of the electronic device. The memory 13 is used to store executable program code. The processor 12 runs a program corresponding to the executable program code by reading the executable program code stored in the memory 13, which is used to execute the configuration method of the big data cluster proposed in any of the above embodiments of this application, or to execute the docking method of the big data platform proposed in any of the above embodiments of this application.
[0202] The specific execution process of the above steps by the processor 12, as well as the steps further executed by the processor 12 by running executable program code, can be found in the description of the embodiments shown in Figures 1-5 of this application, and will not be repeated here.
[0203] To implement the above embodiments, this application also proposes a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the configuration method of the big data cluster proposed in any of the above embodiments of this application, or to execute the docking method of the big data platform proposed in any of the above embodiments of this application.
[0204] To implement the above embodiments, this application also proposes a computer program product, including a computer program that, when executed by a processor, implements the configuration method for a big data cluster proposed in any of the above embodiments of this application, or implements the docking method for a big data platform proposed in any of the above embodiments of this application.
[0205] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art...
[0206] Those skilled in the art may combine and integrate the different embodiments or examples described in this specification and the features 5 of the different embodiments or examples.
[0207] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, the terms "first" and "second" are limited to this specific context.
[0208] The "second" feature may explicitly or implicitly include at least one of those features. In the description of this application,
[0209] The term "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations, wherein they may be implemented not in the order shown or discussed, including substantially simultaneously depending on the functionality involved.
[0210] Alternatively, the functions may be performed in the reverse order, as will be understood by those skilled in the art to which the embodiments of this application pertain.
[0211] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, including...).
[0212] The processor (or other system that fetches and executes instructions from an instruction execution system, apparatus, or device) is used, or used in conjunction with such instruction execution system, apparatus, or device. For the purposes of this specification,
[0213] "Computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit a program for use in or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, a computer-readable medium can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optical scanning of the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0214] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0215] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0216] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0217] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A configuration method for a big data cluster, characterized in that, include: Obtain the configuration file corresponding to the platform to be connected, wherein the configuration file is a component configuration file within the Hadoop system, and the configuration file is in XML format; generate cluster configuration information for a big data cluster based on the configuration file; the big data cluster is used for data processing services, wherein the big data cluster is a Hadoop cluster, and the cluster configuration information is the cluster configuration information of the Hadoop cluster; generating the cluster configuration information for the big data cluster based on the configuration file includes: replacing the corresponding configuration items in the application's configuration file according to the configuration items in the configuration file corresponding to the platform to be connected, so as to obtain the cluster configuration information of the big data cluster, wherein the configuration items are in key-value pair format, the key in the key-value pair is used to indicate the component in the big data cluster, and the value in the key-value pair is used to indicate the calling port; based on the cluster configuration information, configure the cluster groups included in the big data cluster. The method further includes configuring the cluster components, wherein the cluster components include HDFS, YARN, and HIVE; the method also includes: connecting to the cluster components contained in the big data cluster based on the cluster configuration information to complete the connection with the platform to be connected, including: replacing the configuration information in the application used to connect to the cluster components in the big data cluster with the cluster configuration information to complete the connection with the platform to be connected; the method also includes: obtaining the authentication information corresponding to the platform to be connected; the configuration of the cluster components contained in the big data cluster based on the cluster configuration information includes: responding to user operation, sending the authentication information and the target resource information of the component to be configured to the big data cluster based on the cluster configuration information, so that if the big data cluster passes the authentication information, it creates and configures the target resource of the component in the big data cluster according to the target resource information.
2. The method according to claim 1, characterized in that, The target resource information of the component to be configured includes at least one of the following: the location of the HDFS component's storage directory; the upper limit of the HDFS component's storage directory's storage space; the upper limit of the computing resources available for the YARN component's computing resource queue; and the upper limit of the storage space for the HIVE component's database tables.
3. The method according to any one of claims 1-2, characterized in that, The step of obtaining the configuration file corresponding to the platform to be connected includes: in response to the selection operation, determining the type and / or version information of the platform to be connected; and determining the configuration file corresponding to the platform to be connected from the pre-stored configuration files based on the type and / or version information.
4. A method for connecting to a big data platform, characterized in that, include: Obtain the cluster configuration information corresponding to the platform to be connected; the cluster configuration information is generated based on the configuration method of the big data cluster as described in any one of claims 1-3; Based on the cluster configuration information, the system connects to the cluster components included in the big data cluster to complete the connection with the platform to be connected.
5. The method according to claim 4, characterized in that, The method further includes: in response to a received operation instruction, processing the data in the platform to be connected based on the cluster component being connected.
6. The method according to claim 4, characterized in that, The step of obtaining the cluster configuration information corresponding to the platform to be connected includes: obtaining the cluster configuration information corresponding to the platform to be connected through an interface; or, downloading the cluster configuration information corresponding to the platform to be connected.
7. A configuration device for a big data cluster, characterized in that, The apparatus implements the configuration method of the big data cluster as described in claim 1. The apparatus includes: a first acquisition module, used to acquire a configuration file corresponding to the platform to be connected; a generation module, used to generate cluster configuration information of the big data cluster based on the configuration file; the big data cluster is used for data processing services; and a configuration module, used to configure the cluster components included in the big data cluster based on the cluster configuration information.
8. A docking device for a big data platform, characterized in that, include: The third acquisition module is used to acquire cluster configuration information corresponding to the platform to be connected; the cluster configuration information is generated based on the configuration method of the big data cluster according to any one of claims 1-3; the connection module is used to connect to the cluster components included in the big data cluster based on the cluster configuration information, and complete the connection with the platform to be connected.
9. An electronic device, characterized in that, include: At least one processor; The at least one processor is also connected to a memory. The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the configuration method of the big data cluster according to any one of claims 1-3, or to perform the docking method of the big data platform according to any one of claims 4-6.
10. A computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the configuration method of the big data cluster according to any one of claims 1-3, or to execute the docking method of the big data platform according to any one of claims 4-6.
11. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the configuration method of the big data cluster according to any one of claims 1-3, or implements the docking method of the big data platform according to any one of claims 4-6.
Citation Information
Patent Citations
Configuration updating method and device, server and electronic equipment
CN111580884A