Automatic deployment method, device and equipment for large-model all-in-one machine and medium
By using automated BMC and switch configuration, IPMI protocol and other technologies, the automatic deployment of large-scale integrated machines has been achieved, solving the problem of high threshold for large-scale deployment, improving deployment efficiency and reducing costs.
Patent Information
- Application Number
- CN202510866016.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-31
AI Technical Summary
Deploying large-scale pre-trained models has a high barrier to entry, requiring massive computing power and complex training and inference processes, resulting in high deployment difficulty and low efficiency for enterprises.
This paper provides an automated deployment method for large-scale integrated machines. Through automated processes such as BMC and switch configuration, IPMI protocol, DHCP and TFTP services, it realizes the automatic deployment of server network connectivity, system image mounting and large-scale service, simplifying the configuration process.
It has enabled the automated deployment of large-scale integrated machines, reducing deployment difficulty, improving batch deployment efficiency, and reducing manual intervention and costs.
Smart Images

Figure CN120880892A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to an automatic deployment method, apparatus, equipment and medium for a large-scale all-in-one computer. Background Technology
[0002] In recent years, artificial intelligence technology has developed rapidly, especially large-scale pre-trained models (large models) such as GPT and BERT, which have demonstrated powerful capabilities in fields such as natural language processing and computer vision. Large models require massive computing power (such as GPU / TPU clusters), massive amounts of data, and complex training and inference processes, which poses a very high barrier to entry for enterprise deployment.
[0003] Therefore, there is a need for an automated deployment method for large-scale integrated machines to simplify deployment, lower the deployment threshold, and improve batch deployment efficiency. Summary of the Invention
[0004] This invention provides an automatic deployment method, apparatus, equipment, and medium for large-scale integrated machines, which can simplify the deployment difficulty of large-scale integrated machines, lower the deployment threshold, and improve the efficiency of batch deployment.
[0005] According to one aspect of the present invention, an automatic deployment method for a large-scale integrated machine is provided, comprising:
[0006] Configure the BMC of the rack server and the switch according to the network plan, power on the rack, and connect the deployment box to the management network.
[0007] After the deployment service detects network connectivity, it parses the deployment configuration list and notifies the user to begin deployment;
[0008] The temporary system image of ramdisk is automatically mounted via the IPMI protocol. After the server starts, the official system image is pulled and written to the system disk.
[0009] After configuring the network on the client, pull the platform deployment package and deploy the driver, Kubernetes, and large model service in sequence.
[0010] Optionally, the BMC configuration includes setting the IPMI interface address and user password; the switch configuration includes setting management network policies according to network planning.
[0011] The deployment box stores a ramdisk temporary operating system, a formal operating system image, and a platform installation package, and runs DHCP, TFTP, and deployment services to provide automatic deployment support.
[0012] Optionally, the deployment configuration list includes preset information such as the server IPMI interface address, IPMI user password, and platform system drive letter.
[0013] Optionally, the automatic mounting of the ramdisk temporary system image via the IPMI protocol includes:
[0014] The deployment service remotely controls the server to mount a temporary ISO image via the IPMI protocol;
[0015] Configure the server to boot from CD-ROM next time.
[0016] Optionally, after the server starts, it pulls the official system image and writes it to the system disk, including:
[0017] The server obtains a temporary management network IP address via the DHCP protocol;
[0018] The client is deployed by pulling a production operating system image in qcow2 format from the TFTP service.
[0019] Write to the system disk according to the drive letter specified in the deployment configuration list.
[0020] Optionally, the platform deployment package includes drivers, Kubernetes components, and a large model service product, wherein the large model service product is deployed in a container using a Chart template.
[0021] Optionally, the method further includes:
[0022] The client performs basic configuration of the platform, including network parameter verification and service component status detection, and reports the deployment result information to the deployment service.
[0023] Send a deployment completion notification to the user.
[0024] According to another aspect of the present invention, an automatic deployment device for a large-scale integrated model is provided, comprising:
[0025] The configuration unit is used to configure the BMC of the rack server and the switch according to the network plan, power on the rack, and connect the deployment box to the management network.
[0026] The parsing unit is used to parse the deployment configuration list and notify the user to start deployment after the deployment service detects network connectivity.
[0027] The write unit is used to automatically mount the ramdisk temporary system image via the IPMI protocol. After the server starts, it pulls the official system image and writes it to the system disk.
[0028] The pull unit is used to deploy the platform deployment package after the client configures the network, and then deploy the driver, Kubernetes and large model services in sequence.
[0029] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0030] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the automatic deployment method for large-scale integrated machines according to any embodiment of the present invention.
[0031] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the automatic deployment method of the large-scale all-in-one machine according to any embodiment of the present invention.
[0032] This invention provides an automatic deployment method, apparatus, device, and medium for large-scale integrated machines. It enables automatic, unmanned deployment by connecting the box to the large-scale integrated machine and automatically triggering platform deployment. Deployment personnel can simply configure the deployment configuration list in advance and upload it to the deployment box, thereby reducing deployment difficulty and saving deployment costs.
[0033] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 This is a flowchart of an automatic deployment method for a large-scale integrated machine according to an embodiment of the present invention;
[0036] Figure 2 This is a schematic diagram of an automatic deployment method for a large-scale integrated machine according to an embodiment of the present invention;
[0037] Figure 3 This is a schematic diagram of the structure of an automatic deployment device for a large-scale integrated machine according to an embodiment of the present invention;
[0038] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the automatic deployment method of the large-scale integrated machine according to an embodiment of the present invention. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0040] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0041] like Figure 1 As shown in the figure, this embodiment of the invention provides an automatic deployment method for a large-scale integrated machine, which may include the following steps:
[0042] S110. Configure the BMC of the rack server and the switch according to the network plan, power on the rack, and connect the deployment box to the management network.
[0043] Figure 2 This is a schematic diagram of an automatic deployment method for a large-scale integrated machine provided by an embodiment of the present invention. According to the network plan, the BMC (Baseboard Management Controller) address and user password are configured for the server in the rack. Specifically, this includes setting the IPMI (Intelligent Platform Management Interface) IP address, username, and password for subsequent remote server management. Network planning and configuration are performed on the switches, including dividing the management network VLAN and setting port policies (such as link aggregation and port security) to ensure network connectivity between the management network and the service network. After completing the connection between the server and the switch, the entire rack is powered on. The portable deployment box is connected to the in-band management network port of the switch via a network cable. The deployment box contains a built-in ramdisk temporary operating system, a production operating system image, and a platform installation package, and runs DHCP, TFTP, and deployment services.
[0044] S120. After detecting network connectivity, the deployment service parses the deployment configuration list and notifies the user to start deployment.
[0045] The deployment service within the deployment box monitors the management network status in real time, determining network connectivity via ping commands or network protocol stack monitoring. If connectivity is lost, it retryes periodically until it is established. Once connected, the deployment service parses the deployment configuration list uploaded to the deployment box by the user beforehand, obtaining preset information such as the server IPMI interface address, user password, and platform system drive letter. The deployment service then sends a deployment start notification to the user via SMS or email, including the project name and deployment start time.
[0046] S130: Automatically mounts the ramdisk temporary system image via the IPMI protocol. After the server starts, it pulls the official system image and writes it to the system disk.
[0047] The deployment service remotely controls the server via the IPMI protocol, mounts the ramdisk temporary ISO system image to the server's virtual optical drive, configures the server to boot from CDROM next time, and then restarts the server.
[0048] After the server boots from the ramdisk temporary system, it obtains a temporary management network IP address via the DHCP protocol and establishes communication with the deployment box.
[0049] The deployment client (with a built-in temporary system) pulls the production operating system image and deployment configuration list in qcow2 format from the TFTP service. Based on the system drive letter specified in the list (such as / dev / sda), it writes the image to the server's physical hard drive. After completion, it reboots to enter the production operating system.
[0050] S140. After configuring the network on the client, pull the platform deployment package and deploy the driver, K8S, and large model service in sequence.
[0051] Static network configuration: After the production operating system starts, the client obtains a temporary IP address via DHCP, pulls the platform network configuration file from the TFTP service, and configures the server's static IP address, gateway, DNS and other parameters to ensure business network connectivity.
[0052] The client downloads the platform deployment package from the TFTP service, which includes hardware drivers, K8S components, and large model service products.
[0053] First, install the drivers for server hardware such as GPU / network card, then deploy the K8S cluster, including components such as kubelet and kube-controller-manager, to build a containerized runtime environment.
[0054] Based on the K8S platform, large models and intelligent agent service products are deployed in containers using Chart templates, and the configuration and startup of the model inference engine and service interfaces are completed.
[0055] In this embodiment of the invention, BMC configuration includes setting the IPMI interface address and user password; switch configuration includes setting management network policies according to network planning.
[0056] The deployment box stores a ramdisk temporary operating system, a production operating system image, and a platform installation package, and runs DHCP, TFTP, and deployment services to provide automatic deployment support.
[0057] The BMC is an independent management unit of the server, which is remotely managed via the IPMI protocol. Configuring the IPMI interface address requires accessing the server's BMC management interface, typically via buttons on the server's front panel or by pressing a specific hotkey during startup. Assign a fixed IP address, subnet mask, and gateway to the IPMI network interface, ensuring it is on the same network segment as the switch's management network so that the deployment box can access the server's BMC over the network.
[0058] In the BMC management interface, create or modify the login user account and password for the IPMI interface. This information needs to be recorded in the deployment configuration list for authentication when the deployment service remotely controls the server via the IPMI protocol, ensuring that only authorized devices can access the server's underlying management functions.
[0059] Based on the network architecture design of the large-scale integrated machine, such as the separation of management network, service network and storage network, the management network VLAN is divided on the switch, and the corresponding port is added to the VLAN to ensure that the management traffic of the deployment box and the server BMC is isolated from the service traffic.
[0060] Restrict MAC address access permissions on the switch's management network ports, allowing only the MAC addresses of the deployment box and server BMC to pass through, preventing unauthorized devices from accessing the network; if the switch supports this, bind multiple physical ports as logical aggregation ports to improve management network bandwidth and redundancy; prioritize management network traffic to ensure that data packets for services such as DHCP and TFTP are transmitted first during deployment, avoiding network congestion that could affect deployment efficiency.
[0061] The ramdisk temporary operating system is a memory-virtual temporary system image containing a minimal system kernel and drivers. It is used to temporarily load during server startup and can execute and deploy client programs without relying on the local hard drive. The pre-installed operating system image is in qcow2 format and supports rapid deployment to the server's physical hard drive. It also includes installation packages for drivers, K8S components, large model service products, and intelligent agent services required for the large model all-in-one machine, stored in compressed packages or chart templates.
[0062] The DHCP service temporarily assigns a management network IP address to the server, ensuring that the server can communicate with the deployment box via the network after startup to obtain the images and configuration files required for deployment; the TFTP service provides simple file transfer functionality, used by the server to pull the ramdisk temporary system image, the production operating system image, the deployment configuration list, and the platform installation package from the deployment box; the core control service is responsible for detecting network connectivity, parsing the deployment configuration list, calling the IPMI protocol to control the server startup process, coordinating the deployment client to execute image writing and platform service deployment, and sending deployment status notifications to users.
[0063] In this embodiment of the invention, the deployment configuration list includes preset information such as the server IPMI interface address, IPMI user password, and platform system drive letter.
[0064] The deployment configuration manifest is a predefined configuration file that users fill in before deployment based on the hardware plan and platform requirements of the large-scale integrated machine, and stores it in the deployment box. Its core function is to provide key parameters for the automated deployment process, avoiding manual intervention during deployment and achieving unattended deployment.
[0065] The network IP address of the IPMI interface is used by the deployment service to remotely access the server's BMC via the IPMI protocol. The deployment service establishes a connection with the server's BMC through this address to perform low-level operations such as remotely mounting a ramdisk temporary system image, setting the boot order, and restarting the server. For example, the list might record 192.168.1.101 (corresponding to server 1 in rack) or 192.168.1.102 (corresponding to server 2), etc. The IPMI address of each server must be consistent with the switch's management network segment.
[0066] The IPMI user password is the authentication credential for accessing the server's IPMI interface, and includes a username and password. When deploying services to operate the server via the IPMI protocol, this password is required for authentication, ensuring that only authorized deployment processes can control the server. For example, the manifest might default to the username "admin" and the password "Password123," and this information must be exactly the same as the login information configured in the server's BMC.
[0067] The platform system drive letter is the target hard drive letter for the installation of the official operating system and the large model platform.
[0068] Before deployment, users can fill in the IPMI address, password, and drive letter information for each server in the rack by using the UI provided by the deployment service or by directly editing the configuration file, and then upload it to the deployment box. During deployment, the deployment service parses the list and automatically executes the deployment process according to preset parameters.
[0069] In this embodiment of the invention, automatically mounting the ramdisk temporary system image via the IPMI protocol includes:
[0070] The deployment service remotely controls the server to mount a temporary ISO image via the IPMI protocol;
[0071] Configure the server to boot from CD-ROM next time.
[0072] IPMI is a server hardware-level management protocol that allows deployed services to remotely access the server's BMC over a network, enabling control over the server's underlying hardware without relying on an operating system.
[0073] After detecting the management network connectivity and parsing the deployment configuration list, the deployment service sends a remote mount command to the server BMC via the IPMI protocol based on the server IPMI interface address and user password in the list. The deployment service specifies the path of the temporary ISO system image of ramdisk stored in the deployment box and maps the image as a virtual optical drive of the server via the IPMI protocol. After receiving the mount request, the BMC verifies the operation permission of the deployment service through the IPMI user password. After successful verification, the virtual mount of the ISO image is completed.
[0074] The server boots from the local hard drive by default. The boot order needs to be modified via BMC to ensure that the server loads the ramdisk temporary system from the virtual optical drive first after restarting, rather than the local hard drive system that is not deployed.
[0075] Remote access to BMC boot configuration: The deployment service connects to the server BMC management interface via the IPMI protocol to access the boot order configuration module; sets the CDROM boot priority to the highest, overriding the default hard disk boot order; the deployment service sends a boot order modification command, the BMC saves the configuration, and the deployment service sends a server restart command via the IPMI protocol or waits for the server to restart naturally next time, so that the new boot order takes effect.
[0076] In this embodiment of the invention, after the server starts, it pulls the official system image and writes it to the system disk, including:
[0077] The server obtains a temporary management network IP address via the DHCP protocol;
[0078] The client is deployed by pulling a production operating system image in qcow2 format from the TFTP service.
[0079] Write to the system disk according to the drive letter specified in the deployment configuration list.
[0080] DHCP is used to automatically assign temporary network parameters to the server, avoiding manual configuration of IP addresses and ensuring that the server and the deployment box can communicate within the management network.
[0081] The deployment box provides DHCP service: The deployment box runs DHCP service and presets an IP address pool in the management network segment; after the server boots from the ramdisk temporary system, it automatically sends a DHCP Discover broadcast packet to find a DHCP server; after the deployment box's DHCP service receives the request, it assigns a temporary IP address to the server, along with parameters such as subnet mask and gateway. The IP lease period is usually set to the duration of the deployment process to ensure that the IP does not expire during the deployment process.
[0082] TFTP is used for lightweight file transfer. The deployment box provides the operating system image and configuration files to the server through the TFTP service, without the need for complicated authentication mechanisms, making it suitable for efficient file transfer in deployment scenarios.
[0083] After the server obtains a temporary IP address, the deployment client connects to the TFTP server of the deployment box via the TFTP protocol. Following the deployment service instructions, the client requests to pull the qcow2 format image stored in the deployment box. The TFTP service transfers the image file in chunks to the server's memory. Upon receiving the chunks, the deployment client performs an integrity check to ensure the image is undamaged. qcow2 is the image format for QEMU virtual machines, supporting sparse file storage and snapshot functionality. It can be quickly deployed to the server's hard drive, and the image size is smaller than a full system installation package, reducing transfer time.
[0084] Specifying drive letters can avoid system installation errors caused by drive letter conflicts in traditional USB flash drive deployments. By specifying the target drive letter in advance through the configuration list, it can ensure that the system is accurately written to the local hard drive of the server.
[0085] The deployment client obtains the deployment configuration list from the TFTP service and parses out the preset system drive letter; the deployment client calls system tools to partition and format the specified drive letter; and uses a disk cloning tool to write the qcow2 image to the target drive letter, ensuring that the system image completely covers the target disk.
[0086] If the disk letter is already in use during the write process, the deployment client will perform the operation according to the fault tolerance policy in the configuration list to avoid manual intervention; after the write is completed, the deployment client will record the disk letter mapping relationship and verify the disk mount status when the production system starts to ensure deployment reliability.
[0087] In this embodiment of the invention, the platform deployment package includes drivers, K8S components, and a large model service product, wherein the large model service product is deployed in a container using a Chart template.
[0088] Driver software that adapts to server hardware devices, including drivers for hardware such as GPUs, network cards, and storage controllers, ensures that the operating system correctly identifies and calls hardware resources. For example, GPU drivers support parallel computing acceleration for large models, network card drivers ensure network communication efficiency, and storage drivers optimize disk read and write performance.
[0089] Kubernetes (K8S) components are core components of the container orchestration platform, used to manage and deploy containerized applications. These components include: a control plane component responsible for global cluster management and decision-making; and node components running on each server node, responsible for container lifecycle management and network proxying. K8S components are used to build containerized runtime environments, enabling automated deployment, scaling, and fault recovery for large-scale services.
[0090] The large-scale model service product comprises inference services and intelligent agent applications developed based on large-scale models, such as chatbots and text generation engines. It is packaged as a Docker container image and includes the model inference engine, service interfaces, configuration files, etc.
[0091] Chart is the application packaging format for Helm (Kubernetes package manager). It contains a set of YAML configuration files that describe the resource objects required to deploy large model services in a Kubernetes cluster. The Chart template for the large model service product is stored in the platform deployment package. The deployment client parses the Chart template using Helm commands and generates a Kubernetes resource manifest based on the parameters in values.yaml. Helm submits the resource manifest to the Kubernetes API Server, where the Kubernetes cluster creates corresponding Deployment (container group), Service (service exposure), and other objects, thus achieving automated deployment of the large model service.
[0092] Chart templates define service deployment parameters through a unified Chart format, ensuring consistency across multiple environments. They support dynamic adjustment of service resource configurations by modifying values.yaml, adapting to clusters with different computing power scales without code modification. Charts can define dependencies between services, and Helm will deploy dependent components in sequence, reducing deployment complexity.
[0093] After the client pulls the platform deployment package from the TFTP service, it performs the deployment in the order of "driver → K8S component → large model service product". First, the driver is installed to ensure that the hardware is ready; then the K8S cluster is deployed to build a container running platform; finally, the large model service is deployed through the Chart template to realize the business function online.
[0094] During deployment, the deployment client monitors the service startup status through the K8S API. If a component fails to deploy, it will automatically retry or log the error and notify the user. After the large model service is deployed, K8S will automatically expose the service interface, making it easy for external parties to call the inference service.
[0095] In embodiments of the present invention, the method may further include:
[0096] The client performs basic configuration of the platform, including network parameter verification and service component status detection, and reports the deployment result information to the deployment service.
[0097] Send a deployment completion notification to the user.
[0098] The client deployment automatically verifies the server's network configuration parameters, including whether the static IP address, subnet mask, gateway, and DNS server match the deployment configuration list, ensuring connectivity between the service network and the management network. For example, it verifies whether the server IP belongs to the planned service network segment and whether the gateway points to the correct switch port.
[0099] Verify that the hardware drivers for GPU, network card, etc., are installed successfully; check the running status of the K8S control plane and node components to ensure the health of the cluster;
[0100] Check whether the containers for the large model inference engine and agent service are starting normally, whether the service ports are listening, and verify the status code returned by the service interface through HTTP requests.
[0101] If a component is in an abnormal state, the deployment client will log the error and attempt to restart the service. If multiple retries fail, the deployment will be terminated and the fault type will be marked.
[0102] The deployment client organizes the platform configuration results into structured data, including: hardware information such as server IP and BMC address; operating system version and disk partitioning; deployment status and error codes of each service component; and deployment time and resource usage statistics. This information is then reported to the deployment service via HTTP API or message queue, and the deployment service updates the deployment status records in its database upon receipt.
[0103] When the client deployment completes the basic platform configuration and all service components are in normal status, a "deployment complete" signal is sent to the deployment service, triggering the notification process.
[0104] The notification content and format include the project name and deployment task ID; deployment start and end times, total time taken; a summary of deployment results; access entry information; and operational suggestions. The notification can be sent to the user's registered mobile phone number; a detailed report can be sent to the user's email address, along with a download link for the deployment logs and a status statistics table; and a prominent icon can be displayed on the deployment management interface, which can be clicked to view deployment details.
[0105] like Figure 3 As shown, this embodiment of the invention provides an automatic deployment device for a large-scale integrated machine, which may include:
[0106] Configuration unit 310 is used to configure the BMC of the rack server and the switch according to the network plan, power on the rack, and connect the deployment box to the management network.
[0107] The parsing unit 320 is used to parse the deployment configuration list and notify the user to start deployment after the deployment service detects network connectivity.
[0108] The write unit 330 is used to automatically mount the ramdisk temporary system image via the IPMI protocol. After the server starts, it pulls the official system image and writes it to the system disk.
[0109] Pull unit 340 is used to deploy the platform deployment package after the client configures the network, and then deploy the driver, K8S and large model service in sequence.
[0110] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the automatic deployment device for large-scale all-in-one machines. In other embodiments of the present invention, the automatic deployment device for large-scale all-in-one machines may include more or fewer components than illustrated, or combine some components, or split some components, or arrange different components. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0111] The information interaction and execution process between the various units in the above-mentioned device are based on the same concept as the method embodiment of the present invention, and the specific details can be found in the description of the method embodiment of the present invention, and will not be repeated here.
[0112] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0113] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0114] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0115] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the automated deployment method for large-scale integrated machines.
[0116] In some embodiments, the large-scale all-in-one machine automatic deployment method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the large-scale all-in-one machine automatic deployment method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the large-scale all-in-one machine automatic deployment method by any other suitable means (e.g., by means of firmware).
[0117] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0118] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0119] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0120] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0121] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0122] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0123] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0124] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for automatically deploying a large-scale integrated machine, characterized in that, include: Configure the BMC of the rack server and the switch according to the network plan, power on the rack, and connect the deployment box to the management network. After the deployment service detects network connectivity, it parses the deployment configuration list and notifies the user to begin deployment; The temporary system image of ramdisk is automatically mounted via the IPMI protocol. After the server starts, the official system image is pulled and written to the system disk. After configuring the network on the client, pull the platform deployment package and deploy the driver, Kubernetes, and large model service in sequence.
2. The method according to claim 1, characterized in that, The BMC configuration includes setting the IPMI interface address and user password; the switch configuration includes setting management network policies according to network planning. The deployment box stores a ramdisk temporary operating system, a formal operating system image, and a platform installation package, and runs DHCP, TFTP, and deployment services to provide automatic deployment support.
3. The method according to claim 1, characterized in that, The deployment configuration list includes preset information such as the server IPMI interface address, IPMI user password, and platform system drive letter.
4. The method according to claim 1, characterized in that, The automatic mounting of the ramdisk temporary system image via the IPMI protocol includes: The deployment service remotely controls the server to mount a temporary ISO image via the IPMI protocol; Configure the server to boot from CD-ROM next time.
5. The method according to claim 1, characterized in that, After the server starts, it pulls the official system image and writes it to the system disk, including: The server obtains a temporary management network IP address via the DHCP protocol; The client is deployed by pulling a production operating system image in qcow2 format from the TFTP service. Write to the system disk according to the drive letter specified in the deployment configuration list.
6. The method according to claim 1, characterized in that, The platform deployment package includes drivers, Kubernetes components, and a large model service product, wherein the large model service product is deployed in a container using a Chart template.
7. The method according to claim 1, characterized in that, The method also includes: The client performs basic configuration of the platform, including network parameter verification and service component status detection, and reports the deployment result information to the deployment service. Send a deployment completion notification to the user.
8. An automatic deployment device for a large-scale integrated model machine, characterized in that, include: The configuration unit is used to configure the BMC of the rack server and the switch according to the network plan, power on the rack, and connect the deployment box to the management network. The parsing unit is used to parse the deployment configuration list and notify the user to start deployment after the deployment service detects network connectivity. The write unit is used to automatically mount the ramdisk temporary system image via the IPMI protocol. After the server starts, it pulls the official system image and writes it to the system disk. The pull unit is used to deploy the platform deployment package after the client configures the network, and then deploy the driver, Kubernetes and large model services in sequence.
9. An electronic device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the automatic deployment method for the large model all-in-one machine according to any one of claims 1-7.
10. A computer-readable medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the automatic deployment method of the large model all-in-one machine according to any one of claims 1-7.