A reasoning platform multi-cluster system deployment method and device, a terminal and a medium
Patent Information
- Application Number
- CN202310784756.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-29
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-06-29
AI Technical Summary
当前,对AIStation推理平台多集群系统的部署一般是用原始命令一步一步将各种驱动、第三方组件及系统模块业务安装部署,这种方式对部署人员的sh或Python语言要求比较高,命令复杂且多,成百上千行,执行起来难度大且容易出错,严重影响推理平台多集群系统的部署效率
[0016]本发明提供的一种推理平台多集群系统部署方法、装置、终端及介质,相对于现有技术,具有以下有益效果:基于自动化容器操作的开源平台实现,构建完整的部署流程,可借助脚本使用简单命令实现多集群系统的自动部署,无需要求部署人员了解sh或Python语言,减轻人员工作负担,提高部署效率,降低管理成本投入。部署过程中,实时记录部署结果,若部署异常发出提示,便于部署人员了解部署状态。
Smart Images

Figure CN116860263B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cluster deployment, and specifically to a method, apparatus, terminal, and medium for deploying a multi-cluster system for an inference platform. Background Technology
[0002] AIStation is an AI development resource platform for AI enterprise inference scenarios. It aims to solve problems such as cumbersome inference service deployment processes, lack of unified management of AI applications, and unreasonable allocation of computing resources. It provides convenient, stable, and efficient one-stop AI application deployment and computing resource management software products for customers such as government, security, finance, and traditional IT companies.
[0003] AIStation inference platform supports lifecycle management and microservice architecture for containerized applications, provides multiple ways to publish inference services and continuous delivery capabilities, simplifies the inference service deployment process, and provides users with a stable, fast and flexible production environment service deployment platform.
[0004] Figure 1 This is a diagram illustrating the cluster architecture of the AIStation inference platform. The AIStation inference platform comprises multiple clusters ( Figure 1 (The example below uses two clusters.) Each cluster contains one master node and at least one slave node. Currently, the deployment of multi-cluster systems for the AIStation inference platform typically involves installing and deploying various drivers, third-party components, and system modules step by step using raw commands. This method requires deployment personnel to have a high level of proficiency in sh or Python languages. The commands are complex and numerous, numbering in the hundreds or thousands, making execution difficult and prone to errors, which seriously affects the deployment efficiency of multi-cluster systems for the inference platform. Summary of the Invention
[0005] To address the aforementioned issues, this invention provides a method, apparatus, terminal, and medium for deploying a multi-cluster system on an inference platform. Based on an open-source platform for automated container operations, it constructs a complete deployment process. Automated deployment of multi-cluster systems can be achieved with the help of scripts, eliminating the need for deployment personnel to understand .sh or Python languages, thus reducing the workload of personnel and improving deployment efficiency.
[0006] In a first aspect, the technical solution of the present invention provides a method for deploying a multi-cluster system of an inference platform, implemented based on an open-source platform for automated container operation, comprising the following steps: Install the operating system on all nodes in each cluster; each cluster consists of one master node and several slave nodes. Upload the platform code script installation package to the master node of each cluster and modify the configuration file in the platform code script installation package; Based on the configuration parameters in the modified configuration file, each cluster is deployed under the master node of each cluster. The process involves uploading the platform code script installation package to the master node of each cluster and modifying the configuration files within the package. Specifically, this includes the following steps: Upload the platform code script installation package to the main directory of the master node of each cluster; Extract the platform code script installation package to the main directory to generate an installation directory; Execute the configuration modification command in the installation directory to modify the configuration file; The process of modifying configuration files by executing configuration modification commands in the installation directory includes the following steps: Execute the first configuration modification command to modify the parameters in the domain name address association file; the parameters in the domain name address association file include the high availability environment status, management node address, compute node address, GPU node address, and CPU node address; Check if the domain address association file has been modified. If the modification is not completed, continue to check whether the domain address association file has been modified. Once the modifications are complete, check whether the modified parameters are consistent with the corresponding standard parameters. If they match, a message will appear indicating that the domain name and address have been successfully associated with the file. If there is a discrepancy, check whether the number of times the first configuration modification command has been executed has reached the threshold. If the threshold is reached, a message indicating that the modification of the domain address associated file failed will be issued. If the threshold is not reached, the first configuration modification command will be executed again to modify the parameters in the domain name address association file.
[0007] In an optional implementation, the configuration modification command is executed in the installation directory to modify the configuration file, which further includes the following steps: Execute the second configuration modification command to modify the parameters in the inference platform variable file; the parameters in the inference platform variable file include the inference platform version number, network segment address, database address required for deployment, port number, username, and password; Check if the variable files of the inference platform have been modified. If the modification is not completed, the inference platform variable file will be continuously checked to see if the modification is complete. Once the modifications are complete, check whether the modified parameters are consistent with the corresponding standard parameters. If they match, a message will be displayed indicating that the inference platform variable file was successfully created. If there is a discrepancy, check whether the number of times the second configuration modification command has been executed has reached the threshold. If the threshold is reached, a message indicating that the modification of the variable file on the inference platform has failed will be issued. If the threshold is not reached, the second configuration modification command will be executed again to modify the parameters in the inference platform variable file.
[0008] In an optional implementation, each cluster is deployed separately under the master node of each cluster, specifically including the following steps: Log in to the master node of the first cluster and deploy the first cluster; Determine if the first cluster has been successfully deployed; If not, continue to determine whether the first cluster has been deployed successfully; If so, record the deployment results of the first cluster, then log in to the master node of the second cluster and deploy the second cluster. Determine if the second cluster has been successfully deployed; If not, continue to determine whether the second cluster has been deployed successfully; If so, record the deployment results of the second cluster, then log in to the master node of the third cluster and deploy the third cluster. This process continues until all clusters are fully deployed. Check all cluster deployment results for any clusters that failed to deploy; If it does not exist, a message will be displayed indicating that the cluster system deployment is complete; If it exists, a message will be displayed indicating that the cluster system deployment failed, and the cluster identifier of the failed deployment will be given.
[0009] In one optional implementation, the cluster is deployed by including the following steps: Configure SSH passwordless login and Network Time Protocol (NTP) service on each node; Configure the domain name system service for the cluster; Install the application container engine and deployment tools on each node; Install the distributed system infrastructure, container image repository open-source project, database management system, lightweight directory access protocol and graphics card driver on each node; Log in to the container image repository open source project, check whether the operating system image has been successfully downloaded, and push it to the project created in the container image repository; If not, a cluster deployment error will be displayed; If so, configure the cluster to make all nodes in the cluster and all the servers deployed on the nodes ready. Create services outside the cluster.
[0010] In an optional implementation, deploying the cluster further includes the following steps: When installing a database management system, check if the first cluster already has a database management system installed. If so, skip installing the database management system; Otherwise, install the database management system normally.
[0011] In an optional implementation, deploying the cluster further includes the following steps: When installing Lightweight Directory Access Protocol (LMAP), check if the first cluster has LMAP already installed. If so, skip installing the Lightweight Directory Access Protocol; Otherwise, install the Lightweight Directory Access Protocol normally.
[0012] In an optional implementation, deploying the cluster further includes the following steps: When creating an external service, check if the first cluster has already created an external service; If so, skip creating the service outside the cluster; Otherwise, create the service outside the cluster normally.
[0013] Secondly, the technical solution of the present invention provides a multi-cluster system deployment device for an inference platform, implemented based on an open-source platform for automated container operation, comprising: Operating system installation module: Installs the operating system on all nodes in each cluster; each cluster consists of one master node and several slave nodes; Configuration modification module: Upload the platform code script installation package to the master node of each cluster and modify the configuration file in the platform code script installation package; Cluster deployment module: Based on the configuration parameters in the modified configuration file, each cluster is deployed under the master node of each cluster.
[0014] Thirdly, the technical solution of the present invention provides a terminal, comprising: Memory, used to store the deployment program for the multi-cluster system of the inference platform; A processor is configured to implement the steps of the inference platform multi-cluster system deployment method as described above when executing the inference platform multi-cluster system deployment program.
[0015] Fourthly, the present invention provides a computer-readable storage medium storing an inference platform multi-cluster system deployment program, wherein the inference platform multi-cluster system deployment program, when executed by a processor, implements the steps of the inference platform multi-cluster system deployment method as described in any of the above claims.
[0016] This invention provides a method, apparatus, terminal, and medium for deploying a multi-cluster system on an inference platform. Compared to existing technologies, it offers the following advantages: Based on an open-source platform for automated container operations, it constructs a complete deployment process. Automated deployment of multi-cluster systems can be achieved using simple commands via scripts, eliminating the need for deployment personnel to understand shell scripts or Python, thus reducing workload, improving deployment efficiency, and lowering management costs. During deployment, the results are recorded in real time, and alerts are issued if deployment anomalies occur, facilitating understanding of the deployment status for deployment personnel. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of the AIStation inference platform cluster architecture.
[0019] Figure 2 This is a schematic diagram of a multi-cluster system deployment method for an inference platform provided in an embodiment of the present invention.
[0020] Figure 3 This is a schematic diagram illustrating the deployment process principle of a specific embodiment of a multi-cluster system deployment method for an inference platform provided by the present invention.
[0021] Figure 4 This is a schematic diagram of the process of modifying the domain name address association file in a specific embodiment of a multi-cluster system deployment method for an inference platform provided by the present invention.
[0022] Figure 5 This is a schematic diagram of the process for modifying variable files on an inference platform, in a specific embodiment of a multi-cluster system deployment method for an inference platform provided by an embodiment of the present invention.
[0023] Figure 6 This is a schematic block diagram of the structure of the multi-cluster system deployment device for the inference platform provided in this embodiment of the invention.
[0024] Figure 7 This is a schematic diagram of the structure of a terminal provided in an embodiment of the present invention. Detailed Implementation
[0025] To enable those skilled in the art to better understand the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.
[0027] The key terms used in this invention will be explained below.
[0028] Kubernetes (K8S) is an open-source platform for automating container operations. It is used to manage containerized applications across multiple hosts in a cloud platform.
[0029] Hosts is a system file without an extension that can be opened with tools such as Notepad. Its function is to create an association "database" between some commonly used website domain names and their corresponding IP addresses. When a user enters a website address that needs to be logged in into a browser, the system will first automatically look for the corresponding IP address in the Hosts file. Once found, the system will immediately open the corresponding webpage. If not found, the system will then submit the website address to the DNS domain name resolution server for IP address resolution.
[0030] NTP stands for Network Time Protocol, which is a protocol used to synchronize computer time.
[0031] DNS: Domain Name System, a mapping system that links domain names to IP addresses, providing users with access to the Internet.
[0032] Kubeadm is a deployment tool for Kubernetes (k8s), and its deployment method is relatively simple. It only requires two commands: `kubeadminit` (initialization) and `kubeadm join` (adding a node to the master). It can quickly deploy a k8s cluster.
[0033] Docker is an open-source application container engine that allows developers to package their applications and dependencies into a portable image, which can then be deployed on machines running any popular Linux or Windows operating system, and also enables virtualization.
[0034] Hadoop is a distributed system infrastructure developed by the Apache Software Foundation.
[0035] Harbor is an open-source container image repository project designed for enterprise users. It includes essential enterprise functions such as permission management (RBAC), LDAP, auditing, security vulnerability scanning, image verification, management interface, self-registration, and high availability (HA). It also features image replication and Chinese language support tailored to the characteristics of Chinese users.
[0036] MariaDB: A database management system that is a fork of MySQL and is mainly maintained by the open-source community. MariaDB is licensed under the GPL to ensure full compatibility with MySQL, including its API and command line, making it an easy alternative to MySQL.
[0037] OpenLDAP is a free and open-source implementation of the Lightweight Directory Access Protocol (LDAP). GPU driver: Graphics card driver.
[0038] Figure 2 This is a schematic diagram of a multi-cluster system deployment method for an inference platform provided by an embodiment of the present invention. This method is implemented based on an open-source platform for automated container operations. Figure 2 The executing entity can be a multi-cluster system deployment device for an inference platform. The multi-cluster system deployment method for an inference platform provided in this embodiment of the invention is executed by a computer device; correspondingly, the multi-cluster system deployment device for the inference platform runs on the computer device. Depending on different requirements, the order of the steps in this flowchart can be changed, and some steps can be omitted.
[0039] like Figure 2 As shown, the method includes the following steps.
[0040] S1 installs the operating system on all nodes in each cluster; each cluster consists of one master node and several slave nodes.
[0041] S2 uploads the platform code script installation package to the master node of each cluster and modifies the configuration file in the platform code script installation package.
[0042] S3, based on the configuration parameters in the modified configuration file, deploys each cluster separately under the master node of each cluster.
[0043] The multi-cluster system deployment method for inference platforms provided in this invention is based on an open-source platform for automated container operation. It constructs a complete deployment process and can automatically deploy multi-cluster systems using simple commands via scripts. This eliminates the need for deployment personnel to understand sh or Python, reducing their workload, improving deployment efficiency, and lowering management costs.
[0044] To further understand the present invention, a specific embodiment is provided below to illustrate the invention in a more detailed manner. Figure 3 This is a schematic diagram of the deployment process of this specific embodiment, which is divided into three stages: before installation, during installation and after installation. First, the operating system is installed and the deployment files are configured. Then, the cluster is initialized, HA is enabled, dependent components and the inference platform are deployed, and after installation, activation and checks are performed.
[0045] Specifically, this particular embodiment includes the following steps.
[0046] S101 installs the operating system on all nodes in each cluster.
[0047] Each cluster consists of one master node and several slave nodes, with CentOS 7.7 GNOME Desktop operating system installed on the master and node nodes within each cluster.
[0048] S102, upload the platform code script installation package to the master node of each cluster, and modify the configuration file in the platform code script installation package.
[0049] S102.1 Upload the platform code script installation package to the main directory of the master node of each cluster.
[0050] S102.2 Extract the platform code script installation package to the main directory to generate an installation directory.
[0051] S102.3, execute the configuration modification command in the installation directory to modify the configuration file.
[0052] Specifically, the installation package ais-install.tar.gz is uploaded to the / home directory of each cluster master1 node. The installation package is then decompressed and the configuration files are modified. First, the installation package ais-install.tar.gz is decompressed to the current directory, creating an ais-install directory. Then, on each cluster master node, the relevant commands are executed in the ais-install directory to modify the configuration.
[0053] The modifications to the configuration files include changes to the domain name address association file (host file) and the inference platform variable file (ais.cfg file).
[0054] Figure 4 This is a diagram illustrating the process of modifying the domain name address association file, such as... Figure 4 As shown, modifying the domain name address association file includes the following steps.
[0055] S102.3.1 Execute the first configuration modification command to modify the parameters in the domain name address association file.
[0056] The parameters in the domain name address association file include the high availability environment status, management node address, compute node address, GPU node address, and CPU node address.
[0057] Specifically, on each cluster master node, execute the command `cp config / hosts.single config / hosts` in the ais-install directory to modify the configuration as follows: #AIStationDB 100.3.14.203 ais-db #AIStationStorage 100.3.14.203 ais-hdfs #AIStationImageRepo 100.3.14.203harbor-infp.com #AIStationMaster 100.3.14.41 ais-master1 #AIStationGPUNode 100.3.14.203 ais-gpu-node1 100.3.14.206 ais-gpu-node2 100.3.14.207 ais-gpu-node3 #AIStationCPUNode 100.3.14.45 ais-cpu-node1 100.3.14.46 ais-cpu-node2 #AIStationEDGENode 100.3.14.208 ais-edge-node1 100.3.14.209 ais-edge-node2 #End S102.3.2, Check whether the domain name address association file has been modified.
[0058] S102.3.3 If the modification is not completed, the domain name address association file will be continuously checked to see if the modification is complete.
[0059] S102.3.4 If the modification is complete, check whether the modified parameters are consistent with the corresponding standard parameters.
[0060] S102.3.5 If they match, a message will be displayed indicating that the domain name address was successfully associated with the file.
[0061] S102.3.6 If there is a discrepancy, check whether the number of times the first configuration modification command has been executed has reached the threshold.
[0062] S102.3.7 If the threshold is reached, a message indicating that the modification of the domain name address association file failed will be issued.
[0063] S102.3.8 If the threshold is not reached, the first configuration modification command will be executed again to modify the parameters in the domain name address association file.
[0064] After making modifications, check the results. Only proceed with subsequent operations if the modifications are correct to avoid deployment errors. If a modification is incorrect the first time, try modifying it multiple times. If the modification still fails after several attempts, provide a timely warning.
[0065] Figure 5 This is a diagram illustrating the process of modifying variable files on the inference platform, such as... Figure 5 As shown, modifying the inference platform variable file includes the following steps.
[0066] S102.3.9 Execute the second configuration modification command to modify the parameters in the inference platform variable file.
[0067] The parameters in the inference platform variable file include the inference platform version number, network address, database address required for deployment, port number, username, and password.
[0068] Specifically, on each cluster master node, execute the command `cp config / ais.cfg.single config / ais.cfg` in the ais-install directory and modify the configuration as follows: # commit ID and version AIS_0_COMMIT_ID=d6ee99e4c045a6716e5c653d7da8e9ae6f5a8b03 AIS_1_AISTATION_VERSION=v2.4 # net AIS_2_MASTER_NETWORK=100.3.14 AIS_8_HARBOR_DOMAIN_NEAT=harbor-infp.com AIS_3_NFS_PATH= / home / aistation / ais AIS_9_LOG_LEVEL=DEBUG #harbor AIS_10_HARBOR_USER=admin AIS_11_HARBOR_PASSWORD=Harbor12345 AIS_12_HARBOR_DOMAIN=harbor-infp.com:14444 AIS_13_HARBOR_INSTALL=100.3.14.44 AIS_14_HARBOR_PROJECT=ais-ci44 AIS_15_HARBOR_IP=100.3.14.44:14444 AIS_16_HARBOR_DOCKERURL=100.3.14.44:2375 #hadoop AIS_17_HADOOP_IP=100.3.14.44 AIS_18_HADOOP_PORT=14000 AIS_19_HADOOP_LINK=http: / / 100.3.14.44:14000 AIS_20_HADOOP_USER=root #database AIS_21_DB_HOST=100.3.14.43 AIS_22_DB_PORT=7306 AIS_23_DB_USER=root AIS_24_DB_PWD=root AIS_25_DB_NAME=ais_ci43 AIS_26_DB_CHARSET=utf8 #ldap AIS_27_LDAP_IP=100.3.14.43 AIS_28_LDAP_PORT=389 AIS_29_LDAP_TYPE=LDAP AIS_30_LDAP_USER=cn=admin,dc=ldap,dc=inspur,dc=com AIS_30_1_LDAP_DEFAULT_POSIX_GROUP_BASEDN=dc=ldap,dc=inspur,dc=com AIS_31_LDAP_PWD=admin AIS_32_LDAP_LINUXBASE=dcldapdcinspurdccom AIS_33_LDAP_DIR=dc=ldap,dc=inspur,dc=com AIS_34_LDAP_DOTDOMAIN=ldap.inspur.com AIS_35_LDAP_DOMAINBASE=ldap #license and system serving AIS_36_LICENSE_MACHINECODE=b4:05:5d:5d:8c:a0 AIS_37_LICENSE_MODE=development #cluster manager AIS_38_CLUSTER_CMSNODE=ais-master1 AIS_39_CLUSTER_OUTCLUSTER=100.3.14.44 AIS_40_CLUSTER_DOMAIN=testcicd.build44 #monitor AIS_41_LOG_NODE=ais-master1 AIS_42_KIBANA_PORT=70006 AIS_43_INFLUXDB_ADDRESS=100.3.14.44:8083 AIS_44_INFLUXDB_USER=admin AIS_45_INFLUXDB_PASSWORD=123456 #knative AIS_46_KNATIVE_SKIPPINGIMG=harbor-infp.com:14444 #dns server AIS_47_DNS_SERVER=100.3.14.44 #ntp server AIS_48_NTP_SERVER=100.3.14.44 #Virtual IP (No configuration required for non-high availability deployment) AIS_49_VIRTUAL_IP=100.3.14.249 #Always or IfNotPresent AIS_50_IMAGE_PULLPOLICY=Always #PATH AIS_51_DOCKER_DATA_INSTALL_PATH= / var / lib AIS_52_HADOOP_INSTALL_PATH= / usr / local AIS_53_HADOOP_DATA_INSTALL_PATH= AIS_54_HARBOR_INSTALL_PATH= / home AIS_55_HARBOR_DATA_INSTALL_PATH= / dataHarbor AIS_56_MARIADB_DATA_INSTALL_PATH= / var / lib / mysql AIS_57_KUBELET_DATA_INSTALL_PATH= / var / lib AIS_58_ETCD_DATA_INSTALL_PATH= / var / lib #Train AIS_60_TRAIN_URL=https: / / 100.7.36.88:72002 ################### The following variables remain unchanged by default #################### #temp path AIS_1001_BASE_PATH= / home / aistation / ais AIS_1005_INSTALL_PATH= / root / ais-install-v2.4 AIS_1006_TAR_PATH= / home / aistation AIS_1007_ADD_PROCESS=1 #file distribution AIS_1002_DISTR_BLOCKSIZE=4194304 AIS_1003_DISTR_RETRYTIMES=5 AIS_1004_DISTR_RETRYCODE=16 # north AIS_61_CLOUD_IP= S102.3.10, Check whether the variable file of the inference platform has been modified.
[0069] S102.3.11 If the modification is not completed, continuously check whether the inference platform variable file has been modified.
[0070] S102.3.12, If the modification is complete, check whether the modified parameters are consistent with the corresponding standard parameters.
[0071] If S102.3.13 matches, then the inference platform variable file will be successfully filed.
[0072] S102.3.14 If inconsistent, check whether the number of times the second configuration modification command has been executed has reached the threshold.
[0073] S102.3.15 If the threshold is reached, a prompt indicating that the modification of the inference platform variable file has failed will be issued.
[0074] S102.3.16 If the threshold is not reached, the second configuration modification command will be executed again to modify the parameters in the inference platform variable file.
[0075] After making modifications, check the results. Only proceed with subsequent operations if the modifications are correct to avoid deployment errors. If a modification is incorrect the first time, try modifying it multiple times. If the modification still fails after several attempts, provide a timely warning.
[0076] S103, based on the configuration parameters in the modified configuration file, deploys each cluster separately under the master node of each cluster.
[0077] This specific embodiment deploys each cluster sequentially and monitors the cluster deployment status in a timely manner.
[0078] S103.1 Log in to the master node of the first cluster and deploy the first cluster.
[0079] S103.2, Determine whether the first cluster has been deployed successfully.
[0080] S103.3 If not, continue to determine whether the first cluster has been deployed.
[0081] S103.4 If so, after recording the deployment result of the first cluster, log in to the master node of the second cluster and deploy the second cluster.
[0082] S103.5, determine whether the second cluster has been deployed successfully.
[0083] S103.6 If not, continue to determine whether the second cluster has been deployed.
[0084] S103.7 If so, then after recording the deployment result of the second cluster, log in to the master node of the third cluster and deploy the third cluster.
[0085] S103.8, and so on, until all clusters are fully deployed.
[0086] S103.9 checks if there are any clusters that failed to deploy in the results of all cluster deployments.
[0087] If S103.10 does not exist, a message will be displayed indicating that the cluster system deployment is complete.
[0088] S103.11, if it exists, will indicate that the cluster system deployment failed and will provide the cluster identifier of the failed deployment.
[0089] This specific embodiment records the deployment status of each cluster. After all clusters are deployed, it detects whether there are any clusters that have failed to deploy and provides timely prompts to facilitate maintenance. After maintenance, the clusters that failed to deploy can be deployed separately, thus improving deployment efficiency.
[0090] The deployment of the cluster includes the following steps.
[0091] Step 1: Configure SSH passwordless login and Network Time Protocol (NTP) service on each node.
[0092] Step 2: Configure the Domain Name System (DNS) service for the cluster.
[0093] Step 3: Install the application container engine and deployment tools on each node.
[0094] Step 4: Install the distributed system infrastructure, container image repository open source project, database management system, lightweight directory access protocol and graphics card driver on each node.
[0095] Step 5: Log in to the container image repository open source project, check whether the operating system image has been successfully downloaded, and push it to the project created in the container image repository.
[0096] Step 6: If not, a cluster deployment error message will be displayed.
[0097] Step 7: If so, configure the cluster so that all nodes in the cluster and all the servers deployed on the nodes are ready.
[0098] Step 8: Create an external service.
[0099] The database management system, lightweight directory access protocol, and external services are shared by all clusters, so they only need to be deployed on one cluster.
[0100] Accordingly, when installing a database management system, the system checks whether the first cluster already has a database management system installed; if so, the installation is skipped; otherwise, the database management system is installed normally. When installing a lightweight directory access protocol, the system checks whether the first cluster already has a lightweight directory access protocol installed; if so, the installation is skipped; otherwise, the lightweight directory access protocol is installed normally. When creating an external service, the system checks whether the first cluster already has an external service created; if so, the creation of the external service is skipped; otherwise, the external service is created normally.
[0101] Specifically, when deploying the cluster, the installation is performed step by step in the / home / ais-install directory of each cluster master node.
[0102] Execute the following commands respectively: bash -x ais.sh [remoteUser] [remotePassword / domain][master_ip] [step_NO].
[0103] remoteUser: Currently, only the root user is supported, which is a unified root user for the cluster; remotePassword / domain: the unified root user password of the cluster; the second parameter in step 2 is the domain of each cluster; master_ip: the IP address of the master1 node; the third parameter in step 2 is the IP address of the master1 node of each cluster; step_NO: installation step.
[0104] 1. cp ais.sh.single ais.sh.
[0105] Copy a non-high-availability script for use in each subsequent deployment command 2. Configure ssh password-free login and ntp service.
[0106] bash -x ais.sh root [remotePassword] [master_ip]1 3. Configure the dns server of the cluster; step 2 needs to be executed for as many times as there are clusters, and the domain of each cluster is passed in the second parameter in turn, and the IP address of the master node of each cluster is passed in the third parameter.
[0107] bash -x ais.sh root [domain] [master_ip]2 4. Copy and install docker and kubeadm software on all nodes.
[0108] bash -x ais.sh root [remotePassword] [master_ip]3 5. Install hadoop.
[0109] bash -x ais.sh root [remotePassword] [master_ip]4 6. Install harbor.
[0110] bash -x ais.sh root [remotePassword] [master_ip]5 7. Install mariadb: when deploying multiple clusters, all clusters share one database. If the database has been installed in the first cluster, skip this step.
[0111] bash -x ais.sh root [remotePassword] [master_ip]6 8. Install OpenLDAP: When deploying multiple clusters, all clusters share a single LDAP. If OpenLDAP is already installed on the first cluster, skip this step.
[0112] bash -x ais.sh root [remotePassword] [master_ip]7 9. Install the GPU driver.
[0113] bash -x ais.sh root [remotePassword] [master_ip]8 10. Load and push images: This step requires ensuring that the images are successfully loaded and pushed to the project created in Harbor. You can log in to Harbor and go to the project to verify and view this.
[0114] bash -x ais.sh root [remotePassword] [master_ip]9 11. setup cluster (non high available).
[0115] Configure a cluster so that all nodes and the various services deployed on those nodes are in a normal, ready state.
[0116] bash -x ais.sh root [remotePassword] [master_ip]10 12. Set up services outside the cluster: Create services outside the cluster. When deploying multiple clusters, all clusters share a set of services outside the cluster. If services outside the cluster have already been created in the first cluster, other clusters can skip this step.
[0117] bash -x ais.sh root [remotePassword] [master_ip]11 The above completes the deployment of the multi-cluster system for the inference platform. After that, the deployment can be activated and checked.
[0118] The foregoing has described in detail an embodiment of a method for deploying a multi-cluster system of an inference platform. Based on the method for deploying a multi-cluster system of an inference platform described in the above embodiment, this invention also provides an apparatus for deploying a multi-cluster system of an inference platform corresponding to the method.
[0119] Figure 6This is a schematic block diagram of the structure of a multi-cluster system deployment device for an inference platform provided in an embodiment of the present invention. In this embodiment, the multi-cluster system deployment device 600 for an inference platform can be divided into multiple functional modules according to the functions it performs, such as... Figure 6 As shown. The functional modules may include: an operating system installation module 610, a configuration modification module 620, and a cluster deployment module 630. The module referred to in this invention is a series of computer program segments that can be executed by at least one processor and perform a fixed function, and which are stored in memory.
[0120] Operating System Installation Module 610: Installs the operating system on all nodes of each cluster; each cluster includes one master node and several slave nodes.
[0121] Configuration Modification Module 620: Uploads the platform code script installation package to the master node of each cluster and modifies the configuration file in the platform code script installation package.
[0122] Cluster Deployment Module 630: Based on the configuration parameters in the modified configuration file, it deploys each cluster under the master node of each cluster.
[0123] In an optional implementation, the configuration modification module 620 uploads the platform code script installation package to the master node of each cluster and modifies the configuration file in the platform code script installation package. Specifically, this includes the following steps: uploading the platform code script installation package to the main directory of the master node of each cluster; decompressing the platform code script installation package to the main directory to generate an installation directory; and executing the configuration modification command in the installation directory to modify the configuration file.
[0124] In an optional implementation, the configuration modification module 620 executes a configuration modification command in the installation directory to modify the configuration file, specifically including the following steps: executing a first configuration modification command to modify the parameters in the domain name address association file; wherein, the parameters in the domain name address association file include high availability environment status, management node address, compute node address, GPU node address, and CPU node address; detecting whether the domain name address association file has been modified; if the modification is not complete, continuously detecting whether the domain name address association file has been modified; if the modification is complete, detecting whether the modified parameters are consistent with the corresponding standard parameters; if consistent, indicating that the domain name address association file is successfully modified; if inconsistent, detecting whether the number of times the first configuration modification command has been executed has reached a threshold; if the threshold has been reached, issuing a domain name address association file modification failure prompt; if the threshold has not been reached, re-executing the first configuration modification command to modify the parameters in the domain name address association file.
[0125] In an optional implementation, the configuration modification module 620 executes a configuration modification command in the installation directory to modify the configuration file. Specifically, this includes the following steps: executing a second configuration modification command to modify the parameters in the inference platform variable file; wherein the parameters in the inference platform variable file include the inference platform version number, network segment address, database address required for deployment, port number, username, and password; checking whether the inference platform variable file modification is complete; if the modification is not complete, continuously checking whether the inference platform variable file modification is complete; if the modification is complete, checking whether the modified parameters are consistent with the corresponding standard parameters; if consistent, indicating that the inference platform variable file modification was successful; if inconsistent, checking whether the number of times the second configuration modification command has been executed has reached a threshold; if the threshold has been reached, issuing a prompt indicating that the inference platform variable file modification failed; if the threshold has not been reached, re-executing the second configuration modification command to modify the parameters in the inference platform variable file.
[0126] In an optional implementation, the cluster deployment module 630 deploys each cluster under the master node of each cluster, specifically including the following steps: logging into the master node of the first cluster and deploying the first cluster; determining whether the deployment of the first cluster is complete; if not, continuing to determine whether the deployment of the first cluster is complete; if yes, recording the deployment result of the first cluster, logging into the master node of the second cluster and deploying the second cluster; determining whether the deployment of the second cluster is complete; if not, continuing to determine whether the deployment of the second cluster is complete; if yes, recording the deployment result of the second cluster, logging into the master node of the third cluster and deploying the third cluster; and so on, until all clusters are deployed; detecting whether there are any clusters that failed to deploy in the deployment results; if not, indicating that the cluster system deployment is complete; if yes, indicating that the cluster system deployment failed, and providing the identifier of the cluster that failed to deploy.
[0127] In an optional implementation, the cluster deployment module 630 deploys the cluster, specifically including the following steps: configuring SSH passwordless login and Network Time Protocol (NTP) services on each node; configuring the cluster's Domain Name System (DNS) service; installing the application container engine and deployment tools on each node; installing the distributed system infrastructure, container image repository open-source project, database management system, Lightweight Directory Access Protocol (HDLP), and graphics card driver on each node; logging into the container image repository open-source project, checking whether the operating system image has been successfully downloaded, and pushing it to the project created in the container image repository; if not, indicating a cluster deployment error; if yes, configuring the cluster to make all nodes on the cluster and the various servers deployed on the nodes ready; and creating external services.
[0128] In an optional implementation, when the cluster deployment module 630 deploys the cluster, it further includes the following steps: When installing the database management system, it checks whether the first cluster has already installed the database management system; if so, it skips the installation of the database management system; otherwise, it installs the database management system normally. When installing the Lightweight Directory Access Protocol (Light DLAP), it checks whether the first cluster has already installed the Lightweight Directory Access Protocol (Light DLAP); if so, it skips the installation of the Lightweight Directory Access Protocol (Light DLAP); otherwise, it installs the Lightweight Directory Access Protocol (Light DLAP) normally. When creating an external service, it checks whether the first cluster has already created an external service; if so, it skips the creation of the external service; otherwise, it creates the external service normally.
[0129] The inference platform multi-cluster system deployment device in this embodiment is used to implement the aforementioned inference platform multi-cluster system deployment method, so its function corresponds to the function of the above method, and will not be repeated here.
[0130] Figure 7 A schematic diagram of a terminal 700 provided in an embodiment of the present invention includes: a processor 710, a memory 720, and a communication unit 730. The processor 710 is used to implement the following steps when executing the multi-cluster system deployment program for the inference platform stored in the memory 720: Install the operating system on all nodes in each cluster; each cluster consists of one master node and several slave nodes. Upload the platform code script installation package to the master node of each cluster and modify the configuration file in the platform code script installation package; Based on the configuration parameters in the modified configuration file, each cluster is deployed under the master node of each cluster.
[0131] This invention is based on an open-source platform for automated container operations, which builds a complete deployment process. It can automatically deploy multi-cluster systems using simple commands via scripts, without requiring deployment personnel to know sh or Python, thus reducing the workload of personnel, improving deployment efficiency, and reducing management costs. The terminal 700 includes a processor 710, a memory 720, and a communication unit 730. These components communicate via one or more buses. Those skilled in the art will understand that the server structure shown in the figures does not constitute a limitation of the present invention. It can be a bus topology or a star topology, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0132] The memory 720 can be used to store the execution instructions of the processor 710. The memory 720 can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. When the execution instructions in the memory 720 are executed by the processor 710, the terminal 700 is able to perform some or all of the steps in the above method embodiments.
[0133] The processor 710 serves as the control center of the storage terminal, connecting various parts of the electronic terminal via various interfaces and lines. It executes software programs and / or modules stored in the memory 720, and calls data stored in the memory to perform various functions of the electronic terminal and / or process data. The processor can be composed of integrated circuits (ICs), such as a single packaged IC or multiple packaged ICs with the same or different functions connected together. For example, the processor 710 may only include a central processing unit (CPU). In this embodiment of the invention, the CPU may have a single processing core or include multiple processing cores.
[0134] The communication unit 730 is used to establish a communication channel, enabling the storage terminal to communicate with other terminals. It can receive user data sent by other terminals or send user data to other terminals.
[0135] The present invention also provides a computer storage medium, which may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0136] The computer storage medium stores a deployment program for a multi-cluster inference platform system. When the deployment program is executed by a processor, it performs the following steps: Install the operating system on all nodes in each cluster; each cluster consists of one master node and several slave nodes. Upload the platform code script installation package to the master node of each cluster and modify the configuration file in the platform code script installation package; Based on the configuration parameters in the modified configuration file, each cluster is deployed under the master node of each cluster.
[0137] This invention is based on an open-source platform for automated container operations, which builds a complete deployment process. It can automatically deploy multi-cluster systems using simple commands via scripts, without requiring deployment personnel to know sh or Python, thus reducing the workload of personnel, improving deployment efficiency, and reducing management costs. Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or other media capable of storing program code. It includes several instructions to cause a computer terminal (which may be a personal computer, server, or a second terminal, network terminal, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0138] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0139] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0140] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0141] The above-disclosed embodiments are merely preferred embodiments of the present invention, but the present invention is not limited thereto. Any non-creative variations that can be conceived by those skilled in the art, as well as any improvements and modifications made without departing from the principles of the present invention, should fall within the protection scope of the present invention.
Claims
1. A method for deploying a multi-cluster system for an inference platform, characterized in that, An open-source platform implementation based on automated container operations includes the following steps: Install the operating system on all nodes in each cluster; each cluster consists of one master node and several slave nodes. Upload the platform code script installation package to the master node of each cluster and modify the configuration file in the platform code script installation package; Based on the configuration parameters in the modified configuration file, each cluster is deployed under the master node of each cluster. The process involves uploading the platform code script installation package to the master node of each cluster and modifying the configuration files within the package. Specifically, this includes the following steps: Upload the platform code script installation package to the main directory of the master node of each cluster; Extract the platform code script installation package to the main directory to generate an installation directory; Execute the configuration modification command in the installation directory to modify the configuration file; The process of modifying configuration files by executing configuration modification commands in the installation directory includes the following steps: Execute the first configuration modification command to modify the parameters in the domain name address association file; the parameters in the domain name address association file include the high availability environment status, management node address, compute node address, GPU node address, and CPU node address; Check if the domain address association file has been modified. If the modification is not completed, continue to check whether the domain address association file has been modified. Once the modifications are complete, check whether the modified parameters are consistent with the corresponding standard parameters. If they match, a message will appear indicating that the domain name and address have been successfully associated with the file. If there is a discrepancy, check whether the number of times the first configuration modification command has been executed has reached the threshold. If the threshold is reached, a message indicating that the modification of the domain address associated file failed will be issued. If the threshold is not reached, the first configuration modification command will be executed again to modify the parameters in the domain name address association file; Execute the second configuration modification command to modify the parameters in the inference platform variable file; the parameters in the inference platform variable file include the inference platform version number, network segment address, database address required for deployment, port number, username, and password; Check if the variable files of the inference platform have been modified. If the modification is not completed, the inference platform variable file will be continuously checked to see if the modification is complete. Once the modifications are complete, check whether the modified parameters are consistent with the corresponding standard parameters. If they match, a message will be displayed indicating that the inference platform variable file was successfully created. If there is a discrepancy, check whether the number of times the second configuration modification command has been executed has reached the threshold. If the threshold is reached, a message indicating that the modification of the variable file on the inference platform has failed will be issued. If the threshold is not reached, the second configuration modification command will be executed again to modify the parameters in the inference platform variable file.
2. The method for deploying a multi-cluster system of an inference platform according to claim 1, characterized in that, Under the master node of each cluster, each cluster is deployed separately, specifically including the following steps: Log in to the master node of the first cluster and deploy the first cluster; Determine if the first cluster has been successfully deployed; If not, continue to determine whether the first cluster has been deployed successfully; If so, record the deployment results of the first cluster, then log in to the master node of the second cluster and deploy the second cluster. Determine if the second cluster has been successfully deployed; If not, continue to determine whether the second cluster has been deployed successfully; If so, record the deployment results of the second cluster, then log in to the master node of the third cluster and deploy the third cluster. This process continues until all clusters are fully deployed. Check all cluster deployment results for any clusters that failed to deploy; If it does not exist, a message will be displayed indicating that the cluster system deployment is complete; If it exists, a message will be displayed indicating that the cluster system deployment failed, and the cluster identifier of the failed deployment will be given.
3. The method for deploying a multi-cluster system of an inference platform according to claim 2, characterized in that, Deploying a cluster involves the following steps: Configure SSH passwordless login and Network Time Protocol (NTP) service on each node; Configure the domain name system service for the cluster; Install the application container engine and deployment tools on each node; Install the distributed system infrastructure, container image repository open-source project, database management system, lightweight directory access protocol and graphics card driver on each node; Log in to the container image repository open source project, check whether the operating system image has been successfully downloaded, and push it to the project created in the container image repository; If not, a cluster deployment error will be displayed; If so, configure the cluster to make all nodes in the cluster and all the servers deployed on the nodes ready. Create services outside the cluster.
4. The method for deploying a multi-cluster system of an inference platform according to claim 3, characterized in that, Deploying a cluster also includes the following steps: When installing a database management system, check if the first cluster already has a database management system installed. If so, skip installing the database management system; Otherwise, install the database management system normally.
5. The method for deploying a multi-cluster system for an inference platform according to claim 4, characterized in that, Deploying a cluster also includes the following steps: When installing Lightweight Directory Access Protocol (LMAP), check if the first cluster has LMAP already installed. If so, skip installing the Lightweight Directory Access Protocol; Otherwise, install the Lightweight Directory Access Protocol normally.
6. The method for deploying a multi-cluster system of an inference platform according to claim 5, characterized in that, Deploying a cluster also includes the following steps: When creating an external service, check if the first cluster has already created an external service; If so, skip creating the service outside the cluster; Otherwise, create the service outside the cluster normally.
7. A deployment device for a multi-cluster system of an inference platform, characterized in that, An open-source platform implementation based on automated container operations, used to perform the method according to any one of claims 1 to 6, comprising: Operating system installation module: Installs the operating system on all nodes in each cluster; each cluster consists of one master node and several slave nodes; Configuration modification module: Upload the platform code script installation package to the master node of each cluster and modify the configuration file in the platform code script installation package; Cluster deployment module: Based on the configuration parameters in the modified configuration file, each cluster is deployed under the master node of each cluster.
8. A terminal, characterized in that, include: Memory, used to store the deployment program for the multi-cluster system of the inference platform; A processor is configured to implement the steps of the inference platform multi-cluster system deployment method as described in any one of claims 1-6 when executing the inference platform multi-cluster system deployment program.
9. A computer-readable storage medium, characterized in that, The readable storage medium stores an inference platform multi-cluster system deployment program, which, when executed by a processor, implements the steps of the inference platform multi-cluster system deployment method as described in any one of claims 1-6.
Citation Information
Patent Citations
Model deployment method and device and model reasoning method and device
CN112329945A
Deployment method of K8S cluster based on hyper-fusion platform and electronic equipment
CN115051846A