Fast method for creating multi-node clusters

By using a bootable OS to restrict the copying and installation of the boot program OS during the multi-node cluster creation process, combined with PXE service and high-speed data switch, the problem of low target OS transfer efficiency under USB boot is solved, and cluster creation is made more efficient and faster.

CN121887808APending Publication Date: 2026-04-17DELL PROD LP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DELL PROD LP
Filing Date
2024-10-15
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

During the creation of a multi-node cluster, data transfer efficiency is low when using USB boot to install the target OS, especially under low-bandwidth network conditions where it takes too long, affecting the efficiency of cluster creation.

Method used

By using a bootable OS to restrict the copying and installation of the boot program OS, combined with PXE services and high-speed data switches, the target OS can be quickly transferred and cluster nodes can be efficiently re-imagined.

Benefits of technology

It shortens cluster creation time, improves the installation efficiency of multi-node clusters, reduces data transmission time under low-bandwidth network conditions, and improves the speed and efficiency of cluster creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121887808A_ABST
    Figure CN121887808A_ABST
Patent Text Reader

Abstract

The disclosed systems and methods may launch a first node of a multi-node cluster from an attached storage medium, such as a USB driver, into a launcher OS. The launcher OS may initiate PXE and file sharing services on the first node. The disclosed features may copy one or more software packages for a target OS from the attached storage medium to a persistent storage device of the first node. The method may also include a PXE booting one or more other nodes of the multi-node cluster into the bootstrap OS and connecting one or more of the other nodes to the first node via a high speed switch. One or more software packages including software for the target OS are copied from the persistent storage device to at least one of the one or more other nodes via the high speed switch.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of enterprise-scale systems management, and more specifically to operating system (OS) deployment. Background Technology

[0002] As the value and use of information continue to grow, individuals and businesses seek additional ways to process and store information. One option available to users is an information processing system. Information processing systems typically process, compile, store, and / or communicate information or data for business, personal, or other purposes, allowing users to leverage the value of information. Because technologies and information processing needs and requirements vary between different users or applications, information processing systems may also vary regarding: what information is processed, how it is processed, how much information is processed, stored, or communicated, and how quickly and efficiently it can be processed, stored, or communicated. Variations in information processing systems allow them to be general-purpose or configured for specific users or purposes (such as financial transaction processing, airline ticketing, enterprise data storage, or global communications). Furthermore, information processing systems can include various hardware and software components that can be configured to process, store, and communicate information, and may include one or more computer systems, data storage systems, and networking systems.

[0003] Typically, server-level information processing can boot an OS from various sources, including local storage media such as optical discs (CDs), hard disk drives (HDDs), or solid-state drives (SSDs); locally attached external media such as universal serial bus (USB) drives; or remote, network-attached media such as preboot execution environment (PXE) boot storage devices.

[0004] The server can load and execute two or more OS types of varying sizes and capabilities. OS types mentioned in this disclosure include bootable OS, bootloader OS, and target OS. A bootable OS can be loaded directly into and executed from the server's system memory in response to a cold reset and / or hard reset. A bootloader OS can refer to an intermediate memory footprint OS suitable for providing services to facilitate the installation of a target runtime OS (such as a Windows family of OSes from Microsoft or a Linux-based OS distribution).

[0005] In at least some deployments, the resources of multiple servers can be configured into a multi-node cluster, where each server can be referred to as a node. Such clusters can be deployed in hyperconverged infrastructure (HCI) appliances, characterized by tightly integrated and centrally managed compute, storage, and networking resources. Dell Technologies' VxRail series of HCI appliances is an example of an HCI appliance.

[0006] The number of nodes included in a single cluster can range from two to 64 or more. During cluster creation, each node may need to be re-imaged using one or more OS types.

[0007] During server re-imaging, the server can boot into a bootable OS, copy the bootloader OS and the target OS, install the bootloader OS, and boot into the bootloader OS. All of these operations are performed within the bootable OS. The target OS is very large. For example, the Openshift NIM package size exceeds 30GB and the MSAZ NIM package size is >25GB. The time required to install such a large remote shared file can be unacceptable, especially when data is transferred via management port switches that often have low data transfer capacity. Therefore, downloading the target HCIOS using a BMC-restricted network (1Gb / s) can take a very long time (e.g., more than 1 hour). Using a USB image file improves data transfer performance, but it must be performed on each node. Summary of the Invention

[0008] The disclosed features enable the rapid creation of multi-node clusters by minimizing or otherwise reducing the operations performed within the bootable OS by restricting the bootable OS to copying the bootloader OS to the system and then installing the bootloader OS. When a node boots into the bootloader OS, the node can copy the target OS and provide a shared folder on the data port network (25G) for sharing between the bootloader OS and the target OS for replication by other re-imaging nodes.

[0009] The disclosed methods and systems for creating multi-node clusters address common problems associated with this process, allowing for re-imaging of each node in the cluster. However, instead of re-imaging each node using USB boot and transmitting images for the target OS and other software via a relatively slow network switch, the disclosed features allow for re-imaging in a conventional manner (e.g., via USB boot) to launch a bootloader OS that can initiate a PXE service and support file sharing to transmit the target OS image to each of the other nodes in the cluster via a high-speed data switch.

[0010] In at least one aspect, the disclosed system and method can boot a first node of the multi-node cluster from an attached storage medium, such as a USB drive, into a bootloader OS. The bootloader OS can initiate PXE and file-sharing services on the first node. The disclosed features can copy one or more software packages for a target OS from the attached storage medium to a persistent storage device on the first node. The method may further include booting one or more other nodes of the multi-node cluster into the bootloader OS via PXE and connecting one or more of the other nodes to the first node via a high-speed switch. One or more software packages, including software for the target OS, are copied from the persistent storage device to at least one of the one or more other nodes via the high-speed switch.

[0011] Starting the first node may include triggering operations, such as inserting a USB drive into a USB port of the information processing system. Each node may include a data port connected to the high-speed switch and each node may include a management port connected to the management network switch.

[0012] In at least some implementations, PXE booting of one or more other nodes can be triggered by simultaneously or substantially simultaneously powering on each of the one or more other nodes. In some implementations, booting into the bootloader OS includes loading an initial OS (e.g., a bootable OS) into system memory in response to power-on and running the initial OS to load and install the bootloader OS.

[0013] The technical advantages of this disclosure will be readily apparent to those skilled in the art based on the figures, descriptions, and claims included herein. The objectives and advantages of the embodiments will be realized and achieved, at least by means of the elements, features, and combinations specifically pointed out in the claims.

[0014] It should be understood that the foregoing general description and the following detailed description are both illustrative and explanatory, and not limiting of the claims set forth in this disclosure. Attached Figure Description

[0015] A more complete understanding of this embodiment and its advantages can be obtained by referring to the following description taken in conjunction with the accompanying drawings, wherein like reference numerals indicate like features, and in the drawings:

[0016] Figure 1 The traditional process for creating a multi-node cluster based on existing technologies is described.

[0017] Figure 2 The creation of a multi-node cluster is described;

[0018] Figure 3 A flowchart illustrating the cluster creation method is shown; and

[0019] Figure 4 It shows that it is suitable for use with Figures 1 to 3 The subject matter shown in the figure and described in the appended description is a representative information processing system used in combination. Detailed Implementation

[0020] By reference Figures 1 to 4 To best understand the exemplary embodiments and their advantages, the same numbering is used to indicate the same and corresponding parts, unless otherwise explicitly indicated.

[0021] For the purposes of this disclosure, an information processing system may include any tool or set of tools operable to compute, classify, process, transmit, receive, retrieve, create, transform, store, display, represent, detect, record, reproduce, dispose of, or utilize information, intelligence, or data of any form for commercial, scientific, control, entertainment, or other purposes. For example, an information processing system may be a personal computer, a personal digital assistant (PDA), a consumer electronics device, a network storage device, or any other suitable device, and may vary in size, shape, performance, functionality, and price. An information processing system may include memory, one or more processing resources such as a central processing unit (“CPU”), a microcontroller, or hardware or software control logic. Additional components of the information processing system may include one or more storage devices, one or more communication ports for communicating with external devices, and various input / output (“I / O”) devices (such as a keyboard, mouse, and video display). The information processing system may also include one or more buses operable to transmit communication between various hardware components.

[0022] Additionally, the information processing system may include firmware for controlling and / or communicating with, for example, hard disk drives, network circuitry, memory devices, I / O devices, and other peripheral devices. For example, a management program and / or other components may include firmware. As used in this disclosure, firmware includes software embedded in an information processing system component for performing predefined tasks. Firmware is typically stored in non-volatile memory or memory that does not lose stored data upon power failure. In some embodiments, firmware associated with an information processing system component is stored in non-volatile memory accessible to one or more information processing system components. In similar or alternative embodiments, firmware associated with an information processing system component is stored in non-volatile memory dedicated to and including as a part of that component.

[0023] For the purposes of this disclosure, computer-readable media may include any tool or set of tools that can retain data and / or instructions for a period of time. Computer-readable media may include, but is not limited to: storage media, such as direct access storage devices (e.g., hard disk drives or floppy disks), sequential access storage devices (e.g., magnetic tape drives), compact optical discs, CD-ROMs, DVDs, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and / or flash memory; and communication media, such as wires, optical fibers, microwaves, radio waves, and other electromagnetic and / or optical carrier waves; and / or any combination of the foregoing.

[0024] For the purposes of this disclosure, information processing resources may be broadly defined as any component system, apparatus or device of an information processing system, including but not limited to processors, service processors, basic input / output systems (BIOS), buses, memory, I / O devices and / or interfaces, storage resources, network interfaces, motherboards and / or any other components and / or elements of the information processing system.

[0025] In the following description, details are illustrated by way of example to facilitate discussion of the disclosed subject matter. However, it will be apparent to those skilled in the art that the disclosed embodiments are exemplary and not an exhaustive list of all possible embodiments.

[0026] Throughout this disclosure, hyphenated reference numerals denote specific instances of elements, while unhyphenated reference numerals denote elements generally. Thus, for example, "device 12-1" refers to an instance of a device category, which may be collectively referred to as "device 12," and any of the device categories may be generally referred to as "device 12."

[0027] As used herein, when two or more elements are referred to as being “coupled” to each other, such terms indicate, where applicable, that the two or more elements are in electronic communication, mechanical connection, including thermal and fluid communication, thermal or mechanical connection, whether indirect or direct, with or without an intermediate element.

[0028] Now refer to the diagram, Figure 1 The traditional creation of a multi-node cluster 10 is depicted, in which USB drives 12 are used. Figure 1Each of the four nodes 14-0 to 14-3 shown is re-imagined. The USB drive contains an installable image file, referred to herein as an all-in-one image file, because the all-in-one image file may include a bootable OS, a bootloader OS, and software packages for one or more target OSes. After the USB drive 12 has successfully booted one of the nodes 14, the USB drive 12 can be physically delivered to a USB port on the next node 14 and inserted into the USB port to repeat the re-imaging process. While the data transfer rate of the USB drive may be acceptable, the conventional process is inefficient for repeating potentially time-consuming re-imaging operations serially.

[0029] Turn now Figure 2 The characteristics for creating a multi-node cluster 100 based on the disclosed topic are described. This can be achieved by efficiently re-imaging each of the nodes 114. Figure 2 A multi-node cluster of 100. It can be used with... Figure 1 The first node 10 is re-imaged using the USB drive 112 in essentially the same manner. The re-image process boots the first node 114-0 into a bootloader OS. The bootloader OS enables PXE services and file sharing, which allow the remaining nodes 114-1 to 114-3 to significantly reduce the data transfer time required to copy OS image files to and from the first node 114-0 by utilizing the high-speed data switch 120 to which each node 114 is connected. (Comparison) Figure 1 and Figure 2 You can see Figure 1 The remaining process described requires a time interval of approximately N*T, where N is the number of nodes and T is the time required to transfer the required data via the USB drive. In contrast, Figure 2 The file transfer interval is approximately 2*T and depends heavily on the number of nodes. More specifically, such as Figure 2 As described in the text, re-imaging of the first node 114-0 requires... Figure 1 The amount of time required to re-image the first node 114-0 is the same or substantially the same, but the amount of time required to re-image all the remaining nodes 114-1 to 114-3 is reduced by re-imagering them substantially in parallel via a high-speed data switch.

[0030] Now for reference Figure 3The flowchart illustrates an exemplary method 300 for efficiently re-imaging multiple nodes to create a multi-node cluster. The illustrated method 300 begins by booting (operation 302) a first node of the multi-node cluster into a bootloader OS from an attached storage medium. The bootloader OS may initiate (operation 304) PXE and file-sharing services on the first node. Next, one or more software packages for a target OS may be copied (operation 306) from the attached storage medium to a persistent storage device in the first node. The illustrated method 300 also includes booting other nodes of the multi-node cluster into the bootloader OS via PXE (operation 310) and connecting one or more of the other nodes to the first node via a high-speed switch (operation 312). Next, the illustrated method 300 copies (operation 314) one or more software packages from the persistent storage device to at least one of the one or more other nodes via the high-speed switch.

[0031] Now for reference Figure 4 , Figures 1 to 2 Any or more of the elements shown can be implemented as by Figure 4 The illustrated information processing system 400 is an example of an information processing system or implemented within said information processing system. The illustrated information processing system includes one or more general-purpose processors or central processing units (CPUs) 401, which are communicatively coupled to memory resources 410 and input / output hubs 420, with various I / O resources and / or components communicatively coupled to said input / output hubs. Figure 4 The I / O resources explicitly depicted include a network interface 440 (commonly referred to as a NIC (Network Interface Card)), storage resources 430, and additional I / O devices, components, or resources 450. As a non-limiting example, these additional I / O devices, components, or resources include a keyboard, mouse, display, printer, speaker, microphone, etc. The illustrated information processing system 400 includes a baseboard management controller (BMC) 460, which provides out-of-band management resources, as well as other features and services, that can be coupled to a management server (not depicted). In at least some embodiments, the BMC 460 can still manage the information processing system 400 even when it is powered off or in a standby state. The BMC 460 may include a processor, memory, and out-of-band network interfaces that are separate from and physically isolated from the in-band network interfaces of the information processing system 400 and / or other embedded information processing resources. In some implementations, the BMC 460 may include a remote access controller (e.g., a Dell remote access controller or an integrated Dell remote access controller) or a chassis management controller, or may be a component of said remote access controller or chassis management controller.

[0032] This disclosure covers all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments herein that will be understood by those skilled in the art. Similarly, where appropriate, the appended claims cover all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments herein that will be understood by those skilled in the art. Furthermore, references in the appended claims to a device or system or a component of a device or system adapted to, arranged to, capable of, configured to, enabled to, operable to, or effectively perform a particular function cover that device, system, or component, whether or not it or the particular function is activated, turned on, or unlocked, provided that the device, system, or component is so adapted, arranged, capable of, configured, enabled, operable, or effective.

[0033] All examples and conditional language described herein are intended to assist the reader in understanding the present disclosure and the educational objectives provided by the inventors to advance the concepts in the art, and should be construed as not being limited to such specific examples and conditions. While embodiments of the present disclosure have been described in detail, it should be understood that various changes, substitutions, and modifications may be made to embodiments of the present disclosure without departing from the spirit and scope thereof.

Claims

1. A method for creating a multi-node cluster, comprising: The first node of the multi-node cluster is booted into the bootloader OS from the attached storage medium; Initiate a pre-execution environment (PXE) and file sharing service on the first node; as well as Copy one or more software packages for the target OS from the attached storage medium to a permanent storage device in the first node; PXE boots one or more other nodes of the multi-node cluster into the bootloader OS; One or more of the other nodes are connected to the first node via a high-speed switch; The one or more software packages are copied from the persistent storage device to at least one of the one or more other nodes via the high-speed switch.

2. The method of claim 1, wherein starting the first node includes inserting a USB drive into a USB port of the information processing system.

3. The method of claim 1, wherein each node includes a data port connected to the high-speed switch.

4. The method of claim 3, wherein each node includes a management port connected to a management network switch.

5. The method of claim 1, wherein PXE startup of the one or more other nodes includes powering on each of the one or more other nodes.

6. The method of claim 5, wherein energizing each of the one or more other nodes comprises energizing each of the one or more other nodes substantially simultaneously.

7. The method of claim 1, wherein booting into the bootloader OS includes loading an initial OS into system memory in response to power-on and running the initial OS to load and install the bootloader OS.

8. An information processing system, comprising: Central Processing Unit (CPU); System memory, including processor-executable instructions, which, when executed by the CPU, cause the system to perform a multi-node cluster creation operation, the multi-node cluster creation operation including: The first node of the multi-node cluster is booted into the bootloader OS from the attached storage medium; Initiate a Pre-Execution Environment (PXE) and file sharing service on the first node; and Copy one or more software packages for the target OS from the attached storage medium to a permanent storage device in the first node; PXE boots one or more other nodes of the multi-node cluster into the bootloader OS; One or more of the other nodes are connected to the first node via a high-speed switch; The one or more software packages are copied from the persistent storage device to at least one of the one or more other nodes via the high-speed switch.

9. The information processing system of claim 8, wherein starting the first node includes inserting a USB drive into the USB port of the information processing system.

10. The information processing system of claim 8, wherein each node includes a data port connected to the high-speed switch.

11. The information processing system of claim 10, wherein each node includes a management port connected to a management network switch.

12. The information processing system of claim 8, wherein PXE activation of the one or more other nodes includes powering on each of the one or more other nodes.

13. The information processing system of claim 12, wherein energizing each of the one or more other nodes comprises energizing each of the one or more other nodes substantially simultaneously.

14. The information processing system of claim 8, wherein booting into the boot program OS includes loading an initial OS into system memory in response to power-on and running the initial OS to load and install the boot program OS.