GPU hot plug control method and GPU BOX
Through the automated control of the GPU BOX and control unit, hot-swapping of GPUs is possible without shutting down the server, solving the time-consuming and risky issues of hot-swapping in existing technologies, improving system availability and resource utilization, and reducing operating costs and the risk of hardware damage.
Patent Information
- Application Number
- CN202510853632.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-26
AI Technical Summary
In the existing technology, the GPU hot-swap process requires task interruption, time-consuming and professional operation, resulting in low system availability, resource waste and high operating costs, and the risk of hardware damage and data loss.
The GPU BOX and control unit monitor the hot-swap button status and automatically control the GPU plug-in and plug-out operations, enabling hot-swap without shutting down the server. Combined with load balancing scheduling and security processing logic, the safety and efficiency of the plug-in and plug-out process are ensured.
It improves system availability and flexibility, reduces the risk of hardware damage and data loss, reduces manual intervention, improves resource utilization and management efficiency, and ensures business continuity and system stability.
Smart Images

Figure CN120705101A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a GPU hot-plug control method and a GPU BOX. Background Art
[0002] In today's rapidly developing AI data centers and machine learning fields, the demand for hot-swappable GPUs is becoming increasingly prominent. However, during large-model pre-training and fine-tuning, the GPUs involved in the task only perform a portion of the entire task. If a GPU experiences an alarm or anomaly, the current solution is to replace the GPU, which requires interrupting part or all of the task and requiring power removal and unpacking to replace the GPU. This not only causes the task to fail and need to be restarted, but also results in a significant waste of time and resources. For example, a long-term large-model pre-training task may be completely wasted due to a GPU failure, requiring a significant amount of time and energy to retrain.
[0003] Furthermore, inserting and removing a GPU card from a chassis involves opening the chassis, preparing the GPU power cables, checking the slots, powering it on, testing it, and then reinstalling it in the rack. This process is not only time-consuming but also requires professional personnel to ensure that the insertion and removal process does not cause hardware damage or data loss due to static electricity, power mismatches, and other factors. Furthermore, traditional insertion and removal methods require system downtime for each GPU replacement, severely impacting system availability and efficiency. For large data centers or high-performance computing environments, frequent hardware replacements can significantly increase operating costs. Summary of the Invention
[0004] In view of this, the embodiments of the present disclosure provide a GPU BOX, a GPU hot-swap control system and method, which can solve the problem that GPUs in the prior art cannot achieve efficient and fast hot-swap.
[0005] In a first aspect, an embodiment of the present disclosure provides a GPU hot-plug control method, the method comprising:
[0006] By acquiring the actual register value corresponding to the status information of the hot-swap button of the GPU BOX from the control unit; the status information is any one of being continuously pressed, briefly pressed, and not pressed; the GPU BOX is used to install the GPU card;
[0007] Based on a preset mapping relationship between the register value and different state information of the hot-swap button, obtaining the state type corresponding to the actual register value;
[0008] Get the slot information where the GPU BOX is connected to the target server;
[0009] Determining the installation status of the GPU BOX according to the slot information;
[0010] Determine a target request of the GPU BOX according to the installation status and the status type;
[0011] The operating state of the GPU BOX is controlled based on the target request.
[0012] In a second aspect, the present application discloses a GPU BOX, comprising:
[0013] The box body has a receiving slot for installing a GPU module, and the box body is provided with a hot-swap button; the GPU module includes a GPU card;
[0014] The box body is equipped with a control mainboard and a network module. The control mainboard is internally integrated with a slave control unit, and the slave control unit is used to obtain the actual register value corresponding to the status information of the hot-swap button;
[0015] The control mainboard is provided with a connector that matches the plug-in slot of the target server, and the network module is connected to the target server network via the connector.
[0016] The GPU hot-swap control method disclosed in the present application obtains the actual register value corresponding to the status information of the hot-swap button of the GPU BOX from a control unit; obtains the status type corresponding to the actual register value based on a preset mapping relationship between the register value and different status information of the hot-swap button; obtains the slot information of the GPU BOX connected to the target server; determines the installation status of the GPU BOX based on the slot information; determines the target request of the GPU BOX based on the installation status and status type, and controls the operating status of the GPU BOX based on the target request. The operating status of the GPU BOX is controlled by the different states of the hot-swap button and the installation status of the GPU BOX. The GPU BOX can be plugged or unplugged at any time according to actual needs without shutting down the server, greatly improving the availability and flexibility of the system. By monitoring and judging the status of the hot-swap button and properly handling the target request, insertion or removal operations at inappropriate times are avoided, reducing the risk of hardware damage and system failure. The automated hot-swap control process reduces manual intervention, reduces the workload of administrators, and improves system management efficiency. At the same time, the system can respond to hot-swap operations in real time, improving resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 A flowchart of a GPU hot-plug control method provided in an embodiment of the present disclosure.
[0019] Figure 2 This is a flowchart of a first embodiment of the method for controlling the operating state of a GPU BOX based on a target request disclosed in this application.
[0020] Figure 3 This is a flow chart of a second embodiment of the method for controlling the operating state of a GPU BOX based on a target request disclosed in this application.
[0021] Figure 4 A schematic diagram of the hot plug event processing process provided by an embodiment of the present disclosure.
[0022] Figure 5 This is a schematic diagram of the application of the GPU hot-swap control system provided in this embodiment.
[0023] Figure 6 A three-dimensional schematic diagram of a GPU hot-swap control system provided in an embodiment of the present disclosure.
[0024] Figure 7 for Figure 6 A partial enlarged view of middle A.
[0025] Figure 8 A three-dimensional schematic diagram of a GPU BOX provided in an embodiment of the present disclosure.
[0026] Figure 9 for Figure 8 Schematic diagram of the second angle.
[0027] Figure 10 for Figure 8 A three-dimensional schematic diagram of the push-pull limit device in FIG.
[0028] Figure 11 Schematic diagram of the intelligent computing architecture of the GPU hot-swap control system provided in an embodiment of the present disclosure.
[0029] Explanation of the reference numerals: 100, extended backplane box; 111, second mounting position; 112, limiting hole; 200, GPU BOX; 210, box body; 211, locking protrusion; 212, through hole; 213, first ventilation hole; 214, second ventilation hole; 220, GPU module; 230, connecting piece; 240, power module; 250, push-pull limiting device; 251, first rod section; 252, second rod section; 253, screwing piece; 254, pin shaft. DETAILED DESCRIPTION
[0030] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0031] Reference Figure 1 The present application discloses a GPU hot-swap control method for hot-swap control between a GPU BOX and a target server, the method comprising:
[0032] S100 , obtaining an actual register value corresponding to status information of a hot-swap button of the GPU BOX from a control unit.
[0033] The status information is any one of continuously pressed, briefly pressed, and not pressed. Specifically, continuously pressed means that the button is pressed for no less than 6 seconds, and briefly pressed means that the button is pressed for less than 6 seconds.
[0034] The GPU BOX is used to install a GPU card. In this embodiment, the GPU BOX can install all types of GPU cards on the market to meet the needs of different types of products.
[0035] Furthermore, the slave control unit can be a microcontroller (MCU), which is connected to the hot swap button and can detect the button's status. The MCU has internal registers for storing button status information. Software programs can read the actual register values from the MCU registers via a specific communication protocol (such as I2C or SPI). For example, when the hot swap button is pressed continuously, the register value may be set to 0x01; when pressed briefly, to 0x02; and when not pressed, to 0x03.
[0036] Using register values to represent the status information of the hot-swap button can convert the physical status into a digital signal, facilitating subsequent processing and judgment.
[0037] S200 , based on a preset mapping relationship between a register value and different status information of a hot-swap button, obtaining a status type corresponding to an actual register value.
[0038] The status type is inserted or not inserted.
[0039] The establishment of the preset mapping relationship enables the main control unit to quickly and accurately convert register values into meaningful state types. This method simplifies the logic of state judgment, improves the processing efficiency of the system, and avoids complex calculation and judgment processes.
[0040] S300: Obtain the slot information where the GPU BOX is connected to the target server.
[0041] Obtaining slot information can help users clearly identify which GPU BOX to operate on, thus avoiding misoperation.
[0042] Each slot has a unique identifier, and the corresponding slot status can be monitored in real time through a slot information monitoring chip. Alternatively, the slot information to which the GPU BOX is connected can be obtained by querying the server's hardware management interface (such as the BMC-Baseboard Management Controller). For example, the BMC maintains a device connection table that records the device information connected to each slot. The software program can read the slot number where the GPU BOX is located from this table. Slot information is the key to determining the specific location and connection status of the GPU BOX. By obtaining slot information, the system can accurately identify different GPU BOXes and provide accurate positioning for subsequent operations.
[0043] S400: Determine the installation status of the GPU BOX according to the slot information.
[0044] The installation status indicates whether the GPU BOX is plugged into the target server or whether the GPU BOX is not plugged into the target server.
[0045] Specifically, the GPU Box installation status can be determined based on the slot information and the server's hardware status. For example, if the signal connection corresponding to the slot is stable and the server can detect the presence of the GPU Box (e.g., through a hardware detection protocol), the GPU Box can be determined to be properly installed. However, if the slot signal connection is abnormal, this may indicate a problem with the GPU Box installation. Accurately determining the GPU Box installation status is a crucial prerequisite for ensuring the normal operation of the system. By determining the installation status, the system can promptly identify any problems that may arise during installation and avoid system failures caused by improper installation.
[0046] S500: Determine the target request of the GPU BOX according to the installation status and status type.
[0047] The target request includes an insertion power-on request or a power-off unplug request.
[0048] Specifically, when the installation status is that the GPU BOX and the target server are plugged in and the status type is expected to be plugged in, the target request of the corresponding GPU BOX is determined to be an insertion and power-on request; when the installation status is that the GPU BOX and the target server are plugged in and the status type is expected to be unplugged, the target request of the corresponding GPU BOX is determined to be a power-off and unplugging request; when the installation status is that the GPU BOX and the target server are not plugged in and the status type is expected to be plugged in, the target request of the corresponding GPU BOX is determined to be an insertion and power-on request; when the installation status is that the GPU BOX and the target server are not plugged in and the status type is expected to be unplugged, the target request of the corresponding GPU BOX is determined to be a power-off and unplugging request.
[0049] The "power off" in the power off unplug request means that the GPU BOX wants to be disconnected from the server and no longer requires additional power. During the entire process, the target server does not need to be powered off. The plugging and unplugging of a single GPU BOX does not affect the normal operation of other plugged-in GPU BOXes.
[0050] Determining target requests based on installation status and status type enables the system to make reasonable responses based on actual conditions, improving the system's intelligence and operational safety.
[0051] S600: Control the operating state of the GPU BOX based on the target request.
[0052] The GPU Box's operating status is controlled based on target requests, enabling automated hot-swap control of the GPU Box. Specifically, power checks and initialization are performed during insertion and power-up, ensuring the GPU Box's proper operation. Resource release and power-off are performed during removal, improving system stability and hardware lifespan.
[0053] According to the solution disclosed in this embodiment, when inserting or removing a GPU card, that is, inserting or removing a GPU BOX in this embodiment, the GPU BOX can be directly inserted or removed to connect or disconnect with the target server without shutting down the corresponding server.
[0054] Furthermore, in this embodiment, when the target request is a power-off removal request, the system initiates a corresponding protection program upon receiving a signal indicating that the hardware is about to be removed. The operating system suspends all data transmission and processing tasks related to the GPU and triggers the migration of tasks handled by the unplugged GPU to another suitable GPU. After completing data / task processing on the unplugged GPU, the GPU is powered off to ensure that data flow stops safely. For example, ongoing graphics rendering tasks will be temporarily interrupted to prevent data loss or corruption.
[0055] The GPU hot-swap control method disclosed in the present application obtains the actual register value corresponding to the status information of the hot-swap button of the GPU BOX from a control unit; obtains the status type corresponding to the actual register value based on a preset mapping relationship between the register value and different status information of the hot-swap button; obtains the slot information of the GPU BOX connected to the target server; determines the installation status of the GPU BOX based on the slot information; determines the target request of the GPU BOX based on the installation status and status type, and controls the operating status of the GPU BOX based on the target request. The operating status of the GPU BOX is controlled by the different states of the hot-swap button and the installation status of the GPU BOX. The GPU BOX can be plugged or unplugged at any time according to actual needs without shutting down the server, greatly improving the availability and flexibility of the system. By monitoring and judging the status of the hot-swap button and properly handling the target request, insertion or removal operations at inappropriate times are avoided, reducing the risk of hardware damage and system failure. The automated hot-swap control process reduces manual intervention, reduces the workload of administrators, and improves system management efficiency. At the same time, the system can respond to hot-swap operations in real time, improving resource utilization.
[0056] Furthermore, the GPU Box can be flexibly operated according to actual needs, such as installation, removal, and self-test. At the same time, the system can promptly detect installation problems and issue alarms, facilitating maintenance and management for administrators. During hot-swap operations, the system will first perform necessary resource release and data preservation operations to avoid data loss and system failures caused by sudden removal of the GPU Box, thereby improving system security. The system can reasonably control the GPU Box based on different status information, for example, shutting it down when not in use to save energy, and starting it up promptly when needed to improve resource utilization efficiency.
[0057] Furthermore, if the status is a request to unplug, the system indicator light of the GPU BOX is set to flash quickly (e.g., 5 times per second) to indicate that a request has been received. If the application-level hot-plug task processing software is not registered, the control slot is powered off, the current status of each slot is saved, and all indicators are turned off.
[0058] The GPU hot-swap control method disclosed in this embodiment also includes: configuring hot-swap service software, the hot-swap service software establishing a network connection with different target servers, for obtaining GPU operating status information of each target server, and triggering an early warning signal when the GPU operating status information is abnormal; the GPU operating status information includes the GPU core processor occupancy rate and the GPU memory occupancy rate.
[0059] Specifically, software registration and deregistration functions can be added to the hot-swap service software, and the application-level hot-swap task processing software can be deployed on a server with an interconnected network to uniformly manage and process the hot-swap-related AI tasks and hot-swap events of each GPU server. That is, after the application-level hot-swap task processing software is deployed on a server with an interconnected network, SHPS establishes a network connection with each GPU server to achieve unified management.
[0060] The hot-swap service software monitors the GPU core processor utilization and GPU memory utilization of each GPU server in real time. When it detects that the utilization exceeds a threshold, it controls the corresponding GPU Box's system indicator to flash. A GPU load balancing scheduling algorithm is developed based on the performance, power consumption, and domain expertise of each GPU card. The hot-swap service software also supports manually specifying a server's GPU Box for GPU task migration and power-off. It also supports hibernation and wakeup operations for specific GPU Boxes on specific servers according to policies and set time periods. It continuously monitors GPU Box failures and GPU and memory utilization across multiple servers, displaying these statuses in real time on the software interface and GPU Box indicator lights.
[0061] Reference Figure 2 In a first embodiment, the method for controlling the operating state of a GPU BOX based on a target request includes:
[0062] A100, in response to the power-off unplug request, removes the GPU in the corresponding GPU BOX from the task scheduling list through the hot-plug service software.
[0063] When a power-off unplug request is received, if the GPU is directly powered off without any processing, the tasks running on the GPU may be suddenly interrupted, resulting in data loss or incorrect processing results. By removing the corresponding GPU from the task scheduling list through the hot-plug service software, the system can orderly stop the ongoing tasks on the GPU, avoid the adverse effects of abnormal task interruptions, and ensure the stability of system operation. During the task scheduling process, the system will assign tasks according to the status of each GPU. If the GPU to be unplugged is not removed from the task scheduling list, the system will continue to assign new tasks to it. At this time, the GPU is about to be unplugged, which will cause resource allocation conflicts and affect the normal operation of other GPUs. Removing it from the task scheduling list can avoid the occurrence of such conflicts.
[0064] For users, they hope that the system will not be damaged when performing GPU hot-plug operations. Through the processing of hot-plug service software, the system can respond to power-off unplugging requests in a safe manner, allowing users to perform operations with confidence, thereby improving users' trust in the system and usage experience.
[0065] A200, based on the GPU load balancing scheduling algorithm, determines the target GPU from other GPUs, and has the target GPU execute the tasks to be processed by the removed GPU.
[0066] Different GPUs may have differences in performance and load. The GPU load balancing scheduling algorithm can select a target GPU with a lighter load and appropriate performance from other GPUs based on the operating status information of each GPU (such as GPU core processor occupancy and GPU memory occupancy) to execute the tasks pending on the removed GPU. This can avoid situations where some GPUs are overloaded while others are idle, achieve balanced resource allocation, and improve overall resource utilization. Reasonable task allocation allows each GPU to operate at its optimal performance, reducing the waiting time for task processing and improving the overall processing efficiency of the system. For example, in a multi-GPU deep learning training environment, the load balancing scheduling algorithm can speed up model training.
[0067] In many critical business scenarios, such as big data analysis and scientific computing, task continuity is crucial. When a GPU needs to be removed, its pending tasks are promptly assigned to other target GPUs for execution. This ensures that business is not interrupted by the removal of a single GPU, ensuring business continuity and minimizing the impact of hardware maintenance or replacement.
[0068] Reference Figure 3 In the second embodiment, “controlling the operating state of the GPU BOX based on the target request” includes:
[0069] B100, in response to the insertion power-on request, powers on the GPU in the newly inserted GPU BOX through the hot-swap service software;
[0070] B200 , in response to the power-on completion instruction, adds the GPU in the newly inserted GPU BOX to the task scheduling list.
[0071] The hot-swap service software features specialized control logic. When responding to a power-on request, it will power on the newly inserted GPU in the GPU BOX according to specific procedures and specifications. This prevents damage to the GPU hardware caused by current surges caused by sudden power-ups. For example, the software may perform preliminary checks to ensure the GPU's connection is normal and the voltage is stable before gradually applying power to ensure GPU safety. A standardized power-on process helps reduce the probability of failures caused by improper power-up. The hot-swap service software monitors and manages the power-on process. If an anomaly is detected (such as excessive current or abnormal voltage), timely measures can be taken, such as interrupting the power-on operation, to prevent the failure from further escalating, thereby improving system reliability.
[0072] Hot-swap service software automatically responds to insertion requests and performs power-on operations, eliminating the need for manual intervention at each power-on step and significantly improving operational efficiency. This automated operation can save significant time and labor costs in large-scale data centers or scenarios where frequent GPU replacement is required. The software-executed power-on operations adhere to unified standards and procedures, ensuring consistent operation every time. This ensures that different GPUs receive identical power-on treatment, avoiding issues caused by operational variability and improving system stability and maintainability.
[0073] After powering up, the newly inserted GPU in the GPU BOX is added to the task scheduling list. The system can immediately incorporate the new computing resources into the task allocation system. This allows the new GPU to share the load of other GPUs in subsequent task processing, improving the system's overall computing power and processing efficiency. For example, when performing large-scale image rendering or deep learning training, the newly added GPU can accelerate task completion. As business development and demand increase, the system's computing resources may need to be continuously expanded. Promptly adding new GPUs to the task scheduling list allows the system to quickly adapt to business changes, better meet growing business needs, and avoid business bottlenecks caused by insufficient resources.
[0074] The task scheduling list can reasonably allocate tasks based on factors such as the performance and load of each GPU. After adding a new GPU to the list, the system can re-evaluate resource allocation and achieve more optimized task scheduling. For example, when the load on a GPU is too high, the system can allocate some tasks to the newly added GPU, thereby achieving dynamic resource balance and improving resource utilization. This dynamic resource management method enables the system to flexibly respond to various changes, such as hardware failures and business peaks. When a problem occurs on a GPU, the system can adjust task allocation in a timely manner and use other GPUs to continue completing tasks; during business peaks, new GPUs can be inserted at any time and added to the task scheduling list to enhance the system's processing capabilities.
[0075] Furthermore, when multiple actual register values sent from the control unit are received, corresponding target requests are executed in a preset priority order, so that more important tasks can be processed first.
[0076] Specifically, when receiving actual register values from multiple slave control units—that is, when multiple buttons are pressed simultaneously—each corresponding slave control unit generates a corresponding signal. An algorithm (such as a state machine or priority queue) analyzes the aggregated signals to determine the priority and validity of the current operation. This method efficiently handles the parallel operation of multiple buttons, ensuring the accuracy and reliability of the hot-swap process.
[0077] Furthermore, the GPU hot-swap control method disclosed in the present application also includes: in response to a request to introduce a new GPU BOX, scanning the corresponding PCIe slots through ROM software on the target server, identifying vacant slots dedicated to GPUs, and marking the vacant slots as independent slot resources; configuring corresponding slot information for the new GPU BOX based on the independent slot resources; the slot information includes one or both of the physical address and priority of the corresponding slot.
[0078] By scanning the corresponding PCIe slots using the ROM (read-only memory) software on the target server, available slots dedicated to GPUs can be accurately identified. In servers, there are numerous PCIe slots with diverse uses. Manually identifying available GPU-dedicated slots is not only time-consuming and labor-intensive, but also prone to errors. Using ROM software for automated scanning can quickly and accurately identify available slots, improving the efficiency and accuracy of resource identification. Marking the identified available slots as independent slot resources facilitates clear management of GPU slot resources. This marking method allows the system to intuitively distinguish which slots are available and which are occupied, facilitating subsequent resource allocation and scheduling. In large-scale data centers, where there are numerous servers and complex PCIe slot resources, marking independent slot resources can greatly simplify resource management and improve management efficiency.
[0079] Based on independent slot resources, the new GPU Box is configured with corresponding slot information, including the physical address and priority of the corresponding slot. Different GPU Boxes may have different performance and usage requirements. By configuring the physical address and priority, the GPU Box can be matched with the most suitable slot. For example, a GPU Box with higher performance requirements can be assigned a higher priority slot to ensure better resource support and improve operational efficiency.
[0080] Slot configuration enables the system to dynamically adjust resources. When system load changes or new business demands arise, GPU Box slot information can be reconfigured based on actual conditions, enabling flexible resource allocation. For example, during peak business hours, some GPU Boxes can be relocated to higher-priority slots to meet higher computing demands. During low business hours, the slot priorities of some GPU Boxes can be lowered to conserve resources.
[0081] By configuring accurate slot information for a new GPU Box, you can avoid resource conflicts. In a multi-GPU system, improper slot configuration can cause multiple GPU Boxes to compete for the same resources, impacting system stability and performance. Proper slot configuration ensures that each GPU Box has independent resource space, reducing the possibility of resource conflicts and improving system stability. Slot information configuration can also improve system compatibility with different types of GPU Boxes. GPU Boxes produced by different manufacturers may have different interface standards, performance characteristics, and other aspects. By configuring appropriate slot information, the system can better adapt to these differences, ensuring that various types of GPU Boxes can function properly in the system, thereby improving system compatibility and versatility.
[0082] Specifically, the GPU hot-swap control method disclosed in this application can achieve hardware-level hot-swap, effectively solving the problem of the server needing to shut down, power off, open the chassis, and other operations when not performing important tasks.
[0083] Specifically, it includes: re-adjusting and implementing the ROM software on the target server motherboard to allocate specific independent slot resources to the GPU. Specifically, during the server production or initialization phase, technicians use specific programming tools (such as a programmer) to write configuration instructions to the motherboard ROM software. These instructions contain detailed information about the GPU slot resource allocation, such as the physical address and priority corresponding to each GPU slot. After receiving the instructions, the ROM software updates its own configuration data. It will rescan the PCIe slots on the motherboard, identify the slots specifically used for GPU BOX, and mark these slots as independent resources. At the same time, the ROM software will establish a slot resource mapping table to record the status and related information of each GPU BOX slot.
[0084] Change the monitoring and processing behavior of BIOS when the computer is powered on, and detect all important hardware disconnection, connection and other events. Specifically, technicians enter instructions for changing monitoring and processing behaviors through the BIOS setup interface (usually entered by pressing a specific button when the server is powered on). These instructions include setting the type of hardware to be monitored (such as PCIe devices, memory modules, etc.), event triggering conditions (such as time thresholds for hardware connection and disconnection), etc. After receiving the instructions, the BIOS will start a background monitoring program, which will regularly scan the status of the hardware devices and monitor the connection and disconnection events of important hardware in real time by detecting signal changes on the PCIe bus, plug-in detection pins of the device, etc. When an event is detected, the BIOS will perform corresponding processing according to the preset rules.
[0085] Furthermore, the application also includes: judging whether the monitored event is GPU-related based on the slot resources (if the slot where the event occurs is marked as a GPU slot in the slot resource mapping table, it is judged to be a GPU-related event). After detecting the hardware event, the monitoring program obtains information from the slot resource mapping table of the ROM software to make a judgment; if it is a GPU addition or deletion event, a GPU hardware addition or deletion notification is triggered to the CPU through the motherboard bus; if it is not a GPU addition or deletion event, a major hardware fault notification is triggered to the CPU through the motherboard bus and a machine shutdown or restart protection mechanism is triggered.
[0086] By using the slot resource mapping table to determine whether a monitored event is GPU-related, hardware events can be accurately classified. In complex server systems, hardware events are numerous, and different types of events require different handling methods. Using the slot resource mapping table for judgment allows the system to quickly and accurately distinguish GPU-related events from other hardware events, providing a foundation for subsequent targeted processing.
[0087] After detecting a hardware event, the monitoring program directly retrieves information from the ROM software's slot resource mapping table for judgment, avoiding complex analysis and detection processes and significantly improving event handling efficiency. This rapid response mechanism enables the system to react to events promptly, reducing potential risks caused by untimely event handling.
[0088] When a GPU addition or removal event is detected, a GPU hardware addition or removal notification is sent to the CPU via the motherboard bus. This allows the CPU to promptly monitor GPU hardware changes and adjust system resource allocation and scheduling accordingly. For example, when a new GPU is added, the CPU can allocate more computing tasks to the new GPU, improving overall system performance. When a GPU is removed, the CPU can rebalance the load across other GPUs to ensure stable system operation. For events other than GPU addition or removal, a major hardware failure is detected and a machine shutdown or reboot is triggered. This timely protective measure prevents further escalation of the hardware failure and damage to other hardware components.
[0089] Accurate event identification and targeted handling mechanisms help reduce the risk of system crashes. By promptly handling GPU-related events and major hardware failures, the system can quickly adjust and protect itself when problems arise, maintaining stable operation. This stability is particularly important for mission-critical systems, minimizing business interruptions and data loss caused by system failures.
[0090] This approach standardizes and automates the system's hardware fault handling. Whether it's a GPU-related incident or a major hardware failure, there are clear processes and mechanisms for handling it, making it easier for system administrators to troubleshoot and fix the problem. Furthermore, automatically triggered protection mechanisms reduce the need for manual intervention, improving system maintainability.
[0091] Furthermore, the GPU hot-plug control method disclosed in this application also includes configuring security processing logic. Specifically, configuring the security processing logic includes: sending a modification protection mechanism instruction to the BIOS through the BIOS setup interface; in response to the modification protection mechanism instruction, triggering the BIOS to execute a preset modification of the internal initial security processing logic. The preset modification includes: when the BIOS detects an addition or deletion of PCIe hardware, recording the event information and transmitting it to the target operating system, rather than controlling the server to restart or shut down as traditionally done, thus effectively handling server power outages.
[0092] By modifying the protection mechanism in the BIOS setup interface, the BIOS can record event information when it detects the addition or removal of PCIe hardware. This provides the system with real-time monitoring of hardware changes. In security-critical environments, such as financial institutions' data centers or military systems, any unauthorized hardware insertion or removal can pose a security risk. Recording these events can help administrators promptly identify unusual hardware changes and take appropriate security measures to prevent potential security threats, such as data leaks or system damage caused by malicious hardware.
[0093] Recorded hardware event information serves as an important basis for security audits. Administrators can regularly review these records to understand system hardware usage and changes. For example, during compliance checks, this can demonstrate that system hardware operations comply with regulations and procedures. If a security incident occurs, these records can also be used for post-incident investigations to determine key information such as the time of the incident and the hardware involved, helping to quickly locate the problem and identify responsibility.
[0094] When a system malfunctions, recorded PCIe hardware addition or removal events can help administrators quickly determine whether the issue is related to hardware changes. For example, if the system begins to behave abnormally after a certain point in time, and there happens to be a record of PCIe hardware plugging and unplugging operations at that time, this hardware can be prioritized for inspection and troubleshooting, significantly reducing troubleshooting time. This event information is also very helpful for hardware management and maintenance. Administrators can use these records to understand hardware replacement frequency, usage duration, and other information, allowing them to rationally plan hardware maintenance and update cycles. Furthermore, these records can be used as a reference when performing hardware upgrades or replacements to ensure the correctness and safety of the operation.
[0095] With technological advancements, the types and performance of PCIe hardware are constantly being updated. By configuring security processing logic, the system can better adapt to these hardware changes. When a new PCIe device is inserted or an old device is removed, the BIOS can promptly record and pass this information to the target operating system. The operating system can then adjust its configuration and driver based on this information to ensure the proper functioning of the new hardware, improving the system's compatibility with different hardware. By sending instructions to the BIOS to modify the protection mechanism through the BIOS setup interface, users can flexibly adjust the system's security protection policy based on actual needs. Different application scenarios may have different requirements for hardware plug-in and unplug management. Users can select the appropriate protection mechanism based on their own security needs and business characteristics to achieve the optimal balance between system security and flexibility.
[0096] At the same time, the GPU hot-plug control method disclosed in this application can achieve system reliability-level hot-plugging, effectively solving the problems of hardware-level hot-plugging in the non-power-off state, where the live hardware plug-in operation is prone to sudden changes in current and voltage, resulting in plug-in sparks, and live electricity easily causing electrostatic breakdown. This protects the motherboard and hardware chips, and extends the service life of the GPU and the entire machine.
[0097] Specifically, when the status is request to unplug, the system indicator light of the GPU BOX is first set to flash quickly (such as 5 times per second) to indicate that the request has been received. If the system driver and hot-swap service software determine that the current status is in use and request to unplug, it will first send a control signal to the system indicator light of the GPU BOX, requesting that the indicator light be set to flash quickly (such as 5 times per second), indicating that the system has received the unplug request. After receiving the control signal, the system indicator light of the GPU BOX begins to flash rapidly at a frequency of 5 times per second. If the application-level hot-swap task processing software is not registered, the system driver and hot-swap service software will send a power-off command to the corresponding controller to control the power off of the slot. The system saves the current status information of each slot, and finally turns off all the indicators, indicating that the hot-swap operation is completed.
[0098] In addition, through the GPU hot-swap control method disclosed in this application, by adding software registration and deregistration functions and hot-swap service software that can be deployed on a server with an interoperable network, application security-level hot-swap can also be achieved, which can realize hot-swap control in scenarios such as training and inference.
[0099] Further references Figure 4 During hot plug event processing, BIOS upgrades and settings specifically include: The BIOS reserves PCIE resources in the PCI Bus Driver to allocate memory and I / O resources for all PCIe RPs that support Hot Plug. For example, the BIOS reserves 4KB of I / O, 16MB of non-prefetchable memory, and 16MB of prefetchable memory for the PCIe RP. A Hot Plug Control option has been added to the BIOS Setup Menu. This control option has multiple bits, each corresponding to a subnode. For example, Bit 0 controls the Hot Plug of PCIE Root Port D4F0 on Subnode 0. When reporting a Hot Plug Event, LID#Pin is used as the SCI input signal source to be transmitted to the CPU. Therefore, during BIOS initialization, it is necessary to configure the registers related to the LID#Pin SCI interrupt reporting, such as LID#Pin SCI Enable and the interrupt triggering method.
[0100] When a PCIe Root Port is hot-removed, an Uncorrectable Error is generated. This error can cause an MCE, causing the Host OS to stop functioning. Because Surprise Down Errors are inevitable during the Hot Plug process, the BIOS shields Surprise Down Errors for PCIe RPs that support Hot Plug, preventing them from being reported to the Machine Check Architecture (MCA).
[0101] OS configuration specifically includes: using the ACPI Hotplug HotPlug solution based on SCI Interrupt, the OS must support ACPI Hotplug event processing, and when compiling the kernel, check and configure kernel support for ACPI Hotplug.
[0102] Reference Figure 5 The present application also discloses a GPU hot-swap control system, which specifically includes several target servers. A PCIE Switch expansion card can be installed in each slot of each target server, and each expansion slot of the PCIE Switch expansion card can be plugged into a GPU BOX.
[0103] Taking the target server supporting 4 PCIE slots as an example, PCIE Switch technology expansion technology can be used. If the Switch used supports 1 to 5, through expansion, a single server hot-swappable slot can be connected to 4*5 GPU BOXes, that is, 4*5 GPUs can be connected. In specific operations, multiple GPU slots can be grouped according to the Switch, which can facilitate the grouped design and implementation of power supply, control, computing power, bus, etc.
[0104] The hot-plug operation of each GPU BOX adopts the GPU hot-plug control method disclosed in the first aspect of this application.
[0105] Reference Figure 6 The present application discloses a GPU hot-swap control system, and what is intended to be protected is a PCIe-based GPU hot-swap control system, in which the hot-swap control of each GPU BOX in the system adopts the GPU hot-swap control method disclosed in the first aspect of the present application.
[0106] The system specifically includes a GPU BOX 200, a target server, a power supply unit, an expansion backplane box 100, a master control unit, and a slave control unit. The GPU BOX 200 is a PCIE-based GPU BOX that can accommodate a variety of commercially available GPU cards without changing the existing GPU card structure. GPU cards can be swapped in and out, and powered on, without shutting down or opening the target server. This pioneering GPU service architecture innovation enables hot swapping from outside the server. The expansion backplane box 100 is located on the back or side of the target server for easy access and to avoid interference with other components.
[0107] The interior of the extended backplane box 100 includes a first accommodating area and a second accommodating area. The first accommodating area includes N first mounting positions for correspondingly installing N target servers; the second accommodating area includes M second mounting positions 111, each first mounting position is corresponding to P second mounting positions 111, and each second mounting position 111 is used to install a GPU BOX 200; wherein, 0<N<M, P>1.
[0108] Specific reference Figure 7 In this embodiment, the number of GPUs connected to the server is expanded, so M must be greater than N, and each server must be connected to at least two GPUs. In the prior art, server CPUs and motherboards generally directly support a limited number of PCIE slots, typically 4-8.
[0109] In this embodiment, the second installation position 111 is an independent channel. A limiting hole 112 is defined in the second installation position 111 to assist in limiting the position of the GPU BOX after a single GPU BOX is installed in place.
[0110] In this embodiment, M power supply groups are provided. Each power supply group is matched with a power module 240 of a GPU BOX and is used to independently control the power supply of a single GPU BOX.
[0111] Due to their product structure and status, existing GPUs can only be inserted into server PICE slots. Since PCIE power supply is limited and GPU power can reach hundreds of watts, additional power supply lines are required from the server's power supply. In the solution disclosed in this embodiment, if a power supply group fails, such as a short circuit or overcurrent, it will only affect the single GPU box that matches it, without affecting other GPU boxes. This ensures that other GPU boxes in the entire server system can continue to operate normally, greatly improving system reliability. For example, in a large data center, hundreds of GPU boxes work together. Without independent power supply groups, a single power failure could cause all GPUs to stop working, resulting in significant losses.
[0112] When different GPUs run different tasks, their power demands fluctuate. Independent power supply groups can more precisely adjust the power supply to a single GPU Box based on power fluctuations, reducing the impact of power fluctuations on other GPU Boxes. For example, when a GPU Box experiences a sudden and significant power surge during complex deep learning calculations, the independent power supply group can quickly respond and provide sufficient power, preventing power fluctuations from affecting the stable operation of other GPU Boxes.
[0113] Furthermore, the power supply to individual GPU Boxes can be independently turned on or off based on actual needs. For example, during system maintenance, upgrades, or task scheduling, the power to a particular GPU Box can be conveniently disconnected without affecting the normal operation of other GPU Boxes, effectively improving the flexibility and efficiency of system management. If a problem occurs with a GPU Box, the matching power supply group can be checked to quickly determine if the fault is a power issue. Compared to multiple GPUs sharing a single power supply, this independent setup makes troubleshooting simpler and more accurate, shortening repair time.
[0114] Depending on the workload of different GPU BOXes, independent power supply groups can provide power on demand. When a GPU BOX is idle, its power supply can be reduced to reduce unnecessary power consumption. For example, at night when the server load is low, for some GPU BOXes that do not need to work temporarily, energy can be saved by turning off their corresponding power supply groups. Different GPUs may have different power requirements. Independent power supply groups can be configured according to the specific power requirements of each GPU BOX, avoiding the problem of wasted or insufficient power resources due to unified power supply. For example, GPU BOXes with lower power requirements can be equipped with smaller power supply groups, thereby optimizing the power resource allocation of the entire server system.
[0115] In this embodiment, a voltage monitoring chip, such as MAX6393, can be used to monitor the power supply voltage of the PCIE interface in real time. When a voltage abnormality is detected, timely measures can be taken to control power on and off to avoid damage to the device. At the same time, a current sensor, such as INA219, can be used to monitor the current consumption of the PCIE device. The working status of the device can be determined based on the current size. When the device is in standby or idle state, the power can be turned off in time to save energy.
[0116] You can also use hot-swap controller chips like the LTC4261 to safely control the power supply of the PCIE interface when the GPU is inserted or removed, preventing voltage spikes and current shocks during the insertion and removal process, and protecting the server and PCIe devices from damage. When the GPU BOX is detected, the hot-swap controller will gradually power on the PCIE interface according to the preset timing; and when the GPU BOX is detected to be removed, it will smoothly power off the interface.
[0117] In this application, when the GPU BOX is about to be unplugged, the detection circuit will sense the plugging and unplugging action and quickly send a signal to the system. For example, when the GPU BOX begins to be unplugged, the detection circuit will detect the voltage or current change at the interface and then immediately notify the system that the hardware removal operation is about to occur. After receiving the signal that the hardware is about to be removed, the system will start the corresponding protection program, and the operating system will suspend all data transmission and processing tasks related to the GPU in the GPU BOX to ensure that the data stops flowing in a safe state. For example, the ongoing graphics rendering task will be temporarily interrupted to avoid data loss or damage.
[0118] There is usually a data buffer between the GPU and the system to temporarily store data being transferred. When a hot-plug operation is detected, the buffer will save the data that has not been processed. Just like water flowing through a reservoir, when the pipe is about to be disconnected, the reservoir will retain some water and wait for appropriate processing. Before the GPU BOX is unplugged, the system will ensure that the data in the buffer is synchronized with other storage devices or system components. This can ensure the integrity of the data and will not lose important information even if a hot-plug operation occurs. For example, data being transferred from the GPU to the memory will be completed before being unplugged, or at least be safely saved in the buffer.
[0119] In this embodiment, the GPU BOX is connected to the computer system via a PCIe interface. The PCIe standard defines a complete hot-plug protocol that specifies the behavior of hardware and software during hot-plug events, ensuring that the system can handle hot-plug events in an orderly manner. For example, the PCIe hot-plug protocol requires that when the system detects that the GPU BOX is unplugged, it releases related resources in a specific order to avoid system crashes or data corruption.
[0120] The master control unit is configured on the motherboard of the corresponding target server and is used to connect to the slave control units. Each slave control unit is connected to the hot-swap button of a GPU BOX to obtain the status information of the corresponding hot-swap button in real time and send it to the CPU of the corresponding target server through the master control unit. The CPU of the target server controls the hot-swap operation of the corresponding GPU BOX based on the status information signal.
[0121] Furthermore, the main control unit is preferably I 2 C main control chip, I 2 The input and output pins and interrupt pins of the C main control chip are respectively connected to the I 2 C control pin and interrupt pin connection.
[0122] The slave control unit is preferably I 2 C from the controller, I 2 Different range register values of the slave controller correspond to different status information of the hot plug button. 2 C main control chip connects multiple I 2 C bus, each I 2 C bus to connect multiple I 2 C slave controller.
[0123] Among them, "I 2 C From the different range register values of the controller, the corresponding hot swap button status information" can be understood as: the hot swap button will have different states under different operations, such as not pressed, pressed, long pressed, etc. We set I 2 C reflects these states in the register values of the slave controller. When the button state changes, the register value changes accordingly and triggers an interrupt to notify the main controller.
[0124] Suppose we use an 8-bit register to store the status information of the hot swap button. The following are register values in different ranges and their corresponding button states: 1) Not pressed state: When the button is not pressed, the register value remains at 0x00. 2) Short press state: When a button press signal is detected, timing starts. If a button release signal is detected within the set short press time threshold (such as 200ms), the register value is set to a value in the range of 0x01-0x0F, such as 0x01. 3) Long press state: If the button press time exceeds the set long press time threshold (such as 1000ms), the register value is set to a value in the range of 0x10-0x1F, such as 0x10. 4) Abnormal state: If a hardware failure or other abnormal conditions occur during the detection process, the register value is set to a value in the range of 0x20-0xFF, such as 0x20. When the button state changes and the register value changes, I2 The slave controller will trigger an interrupt signal to the master controller. After receiving the interrupt signal, the master controller will send an interrupt signal to the master controller through I 2 C bus to read the register value, thereby obtaining the current status information of the button. Through the above scheme, we can achieve different ranges of register values corresponding to different status information of the hot plug button. When the button state changes, the register value changes accordingly and triggers an interrupt signal. The main controller can use I 2 C bus to read the register value to obtain the current state of the button.
[0125] Reference Figure 8 and Figure 9 The GPU BOX 200 includes a box body 210 , and the box body 210 has an accommodating slot for installing a GPU module 220 , and the GPU module includes a GPU card.
[0126] A control mainboard and a network module are installed in the box body 210. A slave control unit is integrated inside the control mainboard, which is used to obtain the actual register value corresponding to the status information of the hot-swap button; a connector 230 is installed on the control mainboard that matches the plug-in slot of the target server, and the network module is connected to the target server network through the connector.
[0127] The control motherboard is provided with a PCIe slot for inserting the GPU module, which can fix the GPU card according to the original fixing mode.
[0128] The box also includes a power module 240 connected to the GPU module 220;
[0129] The control mainboard also integrates a control module, a storage module and a PCIe interconnection module. The GPU module 220, the control module, the storage module and the network module are all connected to the PCIe interconnection module.
[0130] Furthermore, the ends of the connector 230 and the power module 240 are cantilevered out of the box body 210 , and the ends of the connector 230 and the power module 240 are matched with the plug-in slots of the target server.
[0131] A push-pull limit device 250 is provided on the outside of the box body 210. The push-pull limit device 250 is used to assist in pushing the box body 210 to engage and fix with the target server in the first state, assist in pulling out the box body 210 and disconnecting it from the target server in the second state, and lock the relative position of the box body 210 and the target server in the third state.
[0132] Specifically, the first state is the process of plugging the GPU BOX 200 into the corresponding slot of the target server. The push-pull limit device 250 applies force to the box body 210 under the action of external force, pushing the box body 210 along the corresponding channel to approach the target server until it is engaged and fixed with the target server.
[0133] The second state is the process of disconnecting the GPU BOX 200 from the corresponding slot of the target server. The push-pull limit device 250 applies force to the box body 210 under the action of external force, pulling the box body 210 outward along the corresponding channel and disconnecting it from the target server.
[0134] In the third state, when the GPU BOX 200 is engaged and fixed with the target server, the push-pull limiter 250 is in a vertical state, which is used to lock the relative position of the box body 210 and the target server to ensure the stability of the engagement and fixation with the target server.
[0135] A plurality of first ventilation holes 213 are provided on the first side, and a plurality of second ventilation holes 214 are provided on the side opposite to the first side. The plurality of first ventilation holes 213 and the plurality of second ventilation holes 214 are arranged opposite to each other, forming a good air convection channel inside the box body 210. Cold air can be discharged from the ventilation holes on one side, pass through the interior of the box body 210, and be discharged from the ventilation holes on the other side carrying heat, thereby effectively cooling the internal components of the box body 210 and accelerating the air circulation speed. This convection ventilation method is much more effective than natural heat dissipation or single-directional ventilation, and can promptly reduce the temperature around the GPU.
[0136] The ventilation area formed by the multiple first ventilation holes 213 and the ventilation area formed by the multiple second ventilation holes 214 are both set to match the GPU module 220, that is, the ventilation area is determined based on the heat generation and heat dissipation requirements of the GPU module 220. This ensures that there is sufficient air flow through the GPU module 220 to provide it with just the right heat dissipation capacity, avoiding energy waste due to an excessively large ventilation area or insufficient heat dissipation due to an excessively small ventilation area.
[0137] In this embodiment, through air duct and heat dissipation simulation tests, the maximum contact surface of the air duct inside the GPU BOX 200 is designed to pass through the GPU module 220, so that the cooling air can directly and fully contact the heat source, maximize the heat exchange efficiency between the air and the GPU module 220, quickly conduct away the heat generated by the GPU card, and further enhance the overall ventilation and heat dissipation effect.
[0138] The size of the receiving slot is not less than that of the GPU module 220, and the receiving slot is adapted to different types of GPU cards; the GPU BOX 200 provided in this application has receiving slots adapted to different types of GPU cards, and is compatible with GPU cards of various specifications, allowing users to select suitable GPU cards according to actual needs without having to worry about the problem of being unable to install due to mismatched receiving slots; in addition, the well-adaptable receiving slot can facilitate future GPU card upgrades, ensuring that the equipment can keep up with the pace of technological development and extending the service life of the equipment.
[0139] When users need to improve the system's graphics processing capabilities, they simply remove the existing GPU card from the slot and replace it with a more powerful one. This convenient upgrade method allows the equipment to maintain high performance without undergoing large-scale hardware replacement. In the event of a GPU failure, the slots that accommodate different types of GPU cards make replacement even easier. Maintenance personnel can quickly find and install a suitable replacement GPU, reducing equipment downtime and improving maintenance efficiency.
[0140] The end of the power module 240 is a hole-shaped structure, which facilitates the connection and disconnection of the power line at one time when plugging and unplugging the server.
[0141] The end of connector 230 features multiple rows of holes, effectively simplifying plugging and unplugging operations. Operators no longer need to spend significant time and effort manually aligning pins and holes, improving plugging and unplugging efficiency. This is particularly important during large-scale server deployment or maintenance, where frequent plugging and unplugging operations become easier and faster, saving both labor and time. Connector 230, used to route the network module's network signal to the OSFP port on the target server's CPU control board, effectively reduces contact resistance, helping to minimize loss and interference during signal transmission. This ensures stable and efficient network signal transmission, reduces data packet loss, and improves the quality and reliability of network communications.
[0142] In this embodiment, the housing is made of durable materials and features self-alignment. The ends of the power module and connectors utilize elastic connectors for easy insertion and removal, preventing damage during insertion and removal, and thus extending the device's service life. The GPU module and the storage module each have an integrated DMA controller. A P2PDMA-enabled channel is established between the GPU and storage modules, enabling direct access between the storage device and the GPU module's video memory, supporting direct read and write access to the disk.
[0143] Specifically, the channel supporting P2P DMA is a transmission channel between the DMA controller inside the GPU module and the DMA controller inside the storage module. In this embodiment, the transmission rate of the channel supporting P2P DMA is V1, 100Gb / s≤V1≤400Gb / s. The channel supporting P2P DMA allows direct memory data transmission between devices without CPU intervention. The higher transmission rate (100Gb / s-400Gb / s) can allow large amounts of data to move quickly between different devices (such as GPUs, storage devices, network interface cards, etc.), effectively speeding up data transmission between GPUs and reducing waiting time, thereby giving full play to the parallel computing capabilities of multiple GPUs. The transmission rate of 100Gb / s-400Gb / s provides sufficient bandwidth expansion space for the system, allowing the system to adapt to the growth of data volume and computing needs in the next few years or even longer. In this way, when upgrading and expanding the system, there is no need to immediately replace the entire data transmission architecture, reducing the upgrade cost and complexity.
[0144] A GPUDirect P2P connection channel is established between the GPU module and other GPU modules in the same target server; a GPUDirect RDMA connection channel is established between the GPU module and other GPU modules in different target servers.
[0145] The GPUDirect RDMA connection channel has a transmission rate of V2; 2Tb / s ≤ V2 ≤ 8TbGb / s. This high transmission rate enables lightning-fast data movement between GPUs, between GPUs and storage devices, or between other compute nodes, enabling rapid data sharing and synchronization of computational results, enhancing the cluster's overall computing power and scalability. For example, in a supercomputer cluster, multiple GPU nodes working together via high-speed channels can solve more complex scientific computing and engineering simulation problems.
[0146] In this embodiment, the storage module is preferably an NVMe SSD, and the interface of the NVMe SSD is preferably an Edge 1.S interface, which can meet the data transmission requirements of high-performance storage devices, thereby improving the overall performance and efficiency of the GPU BOX 200.
[0147] The PCIe interconnect module is a PCIe switch chip that relays and strengthens signals from PCIe, control circuits, and other devices that are attenuated by length. The network module is preferably an RDMA network card, which enables high-speed, low-latency network data transmission. In a distributed computing environment, RDMA network cards can significantly improve the efficiency of data exchange between nodes.
[0148] Compared to the traditional data access method in which data is first transferred from the disk to the system memory and then processed and scheduled by the CPU to be moved to the GPU video memory, the GPU BOX disclosed in this embodiment provides a direct access method that can significantly reduce data transmission latency and improve data processing efficiency. This solution bypasses the CPU and allows direct communication between the GPU and storage device, greatly improving the efficiency and speed of data transmission.
[0149] Specifically, during deep learning training, the GPU needs to frequently read large amounts of training data from storage devices. Through the P2P DMA-enabled channel established in this application, the GPU can quickly acquire data, speed up training, and shorten model training time. In high-performance computing fields such as scientific computing and simulation, large amounts of data need to be processed. The established P2P DMA-enabled channel can improve data transmission efficiency, enabling the GPU to perform calculations more efficiently and improve computing performance. For data-intensive applications such as video processing and image analysis, the established P2P DMA-enabled channel can speed up data reading and writing, improving the application's processing power and response speed.
[0150] Among them, the internal PCIe interconnection module, GPU module control module, storage module, and network module in each GPU BOX are all integrated and packaged, preferably packaged in a GPU BOX. The GPU BOX is installed on the server through an extended backplane box. The installed GPU BOX and the server constitute an intelligent computing server, which effectively shortens the physical distance between the components, effectively reduces the path length of signal transmission, reduces signal delay and interference, improves the stability and speed of data transmission, and ensures high-speed and accurate data interaction between modules, thereby improving the computing performance of the entire GPU unit; after the modules are integrated and packaged, the communication between modules is more efficient, reducing unnecessary energy consumption; at the same time, unified power management is also more convenient, and can intelligently adjust power consumption according to the working status of the GPU unit, thereby reducing overall energy consumption and improving energy utilization efficiency.
[0151] In this embodiment, each GPU BOX is independently configured. These independently configured GPU BOXes act like standardized components. During server or computing system deployment, they simply need to be installed into the corresponding interface, eliminating the need for complex wiring and debugging. This plug-and-play nature significantly shortens system deployment time and improves operational efficiency. When computing demands increase, independent GPU BOXes can be easily added to boost the system's computing power. Whether increasing the number of GPU units within a single server or expanding across a cluster of multiple servers, this can be easily accomplished without requiring large-scale modifications to the existing system. Each GPU BOX is independently configured, so if one unit fails, it does not affect the normal operation of other units. The system can quickly identify and isolate the failed unit, allowing computing tasks to continue using the remaining functioning GPU units, thereby improving the reliability and availability of the entire system and reducing downtime caused by hardware failures. Independently configured GPU BOXes offer greater operability during maintenance. If a unit experiences a problem, it can be directly removed from the system for repair or replacement without affecting the normal operation of other units or the entire system. This makes maintenance simpler and more efficient, reducing maintenance costs and complexity. Since each GPU module is independent, its performance can be monitored and managed individually. Through the monitoring software, the working status, performance indicators and other information of each unit can be obtained in real time, so that potential problems can be discovered in time and adjustments and optimizations can be made to ensure that the entire system is always in the best operating state.
[0152] Reference Figure 8 and Figure 10 The push-pull limiting device 250 includes a first rod segment 251, a second rod segment 252, and a screw member 253. The first rod segment 251 is connected to the outer side of the box body 210 via a pin 254 and has the freedom to rotate about the pin 254, facilitating effective control of the box body 210. A locking hole is defined on a first side of the box body 210, and a through hole 212 is defined on the side opposite the first side. The through hole 212 is used to extend the ends of the connector 230 and the power module 240.
[0153] The first rod segment 251 has a locking portion, specifically a C-shaped groove, located at the end of the first rod segment 251. The second rod segment 252 is located at the end of the first rod segment 251 away from the pin 254 and has a mounting portion for mounting the screw member 253.
[0154] In the first state and the second state, the second rod segment 252 is away from the box body 210 and has no contact with the box body 210; in the third state, the second rod segment 252 or the first rod segment 251 rotates around the pin shaft 254 under the action of external force until it is in contact with the side wall of the box body 210, the locking portion presses against the preset position of the target server and the screw member 253 is fixed to the locking hole to lock the relative position of the box body 210 and the target server.
[0155] In the third state, the engaging portion is located within the engaging hole and abuts against the inner wall of the engaging hole, providing resistance to further inward rotation of the first rod segment 251, thereby notifying the operator that the rotation is in place. The screw member 253 can then be manipulated to lock the box body 210, which is simple and efficient. The screw member 253 is preferably a set screw to facilitate the insertion, removal, and fixing of the GPU BOX 200.
[0156] In specific operations, the first rod segment 251 can be rotated to a horizontal position and then pushed inward until it cannot be pushed any further, indicating that it is plugged into the target server. Then the first rod segment 251 can be rotated to a vertical state. At this time, the engaging portion on the first rod segment 251 is located in the limiting hole 112 and is pressed against the inner wall side of the limiting hole 112, that is, the first rod segment 251 can no longer be controlled to rotate inward. Then, the fixing screws are tightened by hand to lock the second rod segment 252 with the box body 210.
[0157] Reference Figure 11 The following describes a detailed description of the GPU hot-swap control system in conjunction with specific embodiments. Within a single GPU BOX, the GPU module, control module, storage module, and network module are all connected to the PCIe interconnect module via the first-class PCIe protocol. Each GPU BOX on the same node is connected to the PCIe switch via the second-class PCIe protocol. Each GPU BOX on a different node is connected to the RDMA switch module via the third-class PCIe protocol. The second-class PCIe protocol version is lower than the first-class PCIe protocol version, the third-class PCIe protocol version is consistent with the first-class PCIe protocol version, and the fourth-class PCIe protocol version is no higher than the second-class PCIe protocol version.
[0158] Among them, the first type of PCIe protocol is preferably the PCIe-Gen5 / 6 protocol, the second type of PCIe protocol is preferably the PCIe-Gen4 / 5 protocol, and the fourth type of PCIe protocol is preferably the PCIe-GEN3 / 4 / 5 protocol.
[0159] In this embodiment, the CPU unit may include several servers, each of which is connected to one or more GPU BOXes; each GPU BOX is an independent hot-swappable unit; when at least two GPU BOXes are installed on the same server, a single-server, multi-BOX model is formed. Through the solution provided by this application, users can flexibly adjust GPU resources based on actual business needs. For example, when conducting large-scale deep learning training, more GPU units can be added to the server to enhance parallel computing capabilities; when the business volume is small, the number of connected GPU units can be reduced to reduce energy consumption and costs.
[0160] The network module in the GPU BOX is connected to the corresponding OSFP interface in a single server. Different GPU BOXes are connected to different OSFP interfaces. Each network module in the GPU BOX has its own unique OSFP interface in the corresponding server.
[0161] The OSFP interface features high bandwidth and low latency, meeting the needs of high-speed transmission of large amounts of data between GPU units and servers. Connecting different GPU BOXes to different OSFP interfaces effectively avoids network conflicts and bandwidth competition. Each GPU unit can independently use the bandwidth resources of an OSFP interface, ensuring the stability and reliability of data transmission and improving the overall performance of the system. Because each GPU unit is independently connected to a different OSFP interface on the server, each GPU unit can be independently monitored and managed. Operations and maintenance personnel can obtain real-time information on the network usage of each GPU unit and promptly identify and resolve potential network issues such as bandwidth bottlenecks and network congestion. When a network failure occurs on a GPU unit, the independence of the interface allows the faulty unit to be quickly located and isolated from other normal units, reducing the impact of the failure on the entire system. This also facilitates troubleshooting and repair, improving operation and maintenance efficiency.
[0162] Each GPU Box on the same node is connected to the PCIe Switch via the PCIe Type 2 protocol. The version of the PCIe Type 2 protocol is lower than that of the PCIe Type 1 protocol. A GPUDirect P2P connection channel is established between the GPU modules in each GPU Box on the same node, enabling P2P DMA communication between GPUs directly using the PCIe bus.
[0163] Specifically, the GPU module integrates a DMA controller connected to the GPU card to control data read and write to the GPU card and control data exchange with the PCIe switch chip. The storage module integrates a DMA controller connected to the storage chip to control data read and write to the storage chip and control data exchange with the PCIe switch chip. The channel that supports P2PDMA is the transmission channel between the DMA controller within the GPU module and the DMA controller within the storage module.
[0164] Furthermore, the control module is preferably an integrated system-on-chip, i.e., SoC, which is responsible for the initialization configuration of the DMA controller inside the GPU module and the DMA controller inside the storage module, and can support P2PDMA (i.e., point-to-point direct memory access) without the participation of the server CPU, that is, data can be read directly from the storage module without being transferred through the server CPU and memory.
[0165] Each GPU Box on a different node is connected to the RDMA switch module via the PCIe Type 3 protocol. The version of the PCIe Type 3 protocol is consistent with the version of the PCIe Type 1 protocol. A GPUDirect RDMA (GDR) connection channel is established between the GPU modules in each GPU Box on different nodes, and cross-node GPU interconnection is achieved through the RDMA switch module.
[0166] The CPU unit (i.e., server) is not co-located with the GPU BOX, and is connected to the GPU BOX via the PCIe Type 4 protocol, which is no later than the PCIe Type 2 protocol. Specifically, the CPU unit manages the GPU module, network module, and storage module via the PCIe bus. This non-co-located design eliminates reliance on components such as PCIe interface switches and retimers.
[0167] A hot-swappable connection channel is established between the GPU BOX and the CPU unit. The GPU BOX is integrated into the GPU BOX and connected to the server. A hot-swappable connection channel is established between the two. This channel is a dynamically pluggable interface link based on a specific high-speed data transmission protocol. This channel has high-speed, low-latency data transmission capabilities, enabling real-time data exchange and command transmission between the GPU BOX and the CPU unit. Furthermore, when the server is operating normally, the GPU BOX can be unplugged and plugged in at any time. During hot-swappable operations, there is no electrical damage to the server system or the GPU BOX, ensuring system stability and data integrity. This channel uses a reliable interface design to ensure stable data transmission and prevent signal interference and electrical problems that may occur during the plug-in and plug-out process.
[0168] In this embodiment, the CPU unit is solely responsible for initializing the issuance of configuration instructions, transmission control, and monitoring. In this architecture, tasks previously handled by the server CPU are now handled by the SoC chip within the GPU Box, completely eliminating the need for the server CPU or memory. The internal PCIe interconnect module primarily performs data routing and exchange, ensuring accurate data transmission between the storage module and the GPU module. Therefore, in this simple one-to-one connection scenario, the DMA controllers of the storage module and GPU module are primarily involved in the P2P DMA transmission process. Therefore, through the internal configuration of the GPU Box, users and applications can directly transfer data from storage devices to GPU memory. The preferred fourth-class PCIe protocol is PCIe-GEN3 / 4 / 5.
[0169] In this embodiment, the P2P DMA-supported channels, GPUDirect P2P connection channels, and GPUDirect RDMA connection channels do not occupy CPU lanes. That is, they do not occupy the physical channels for data transmission between the CPU and other devices, namely, the channels of the PCI-Express (PCIe) bus. The CPU exchanges data with these devices via PCIe channels. Each PCIe channel has a certain bandwidth for data transmission. The more channels there are, the greater the total data transmission bandwidth. However, the number of PCIe channels available to the CPU is limited. By setting up channels that do not occupy CPU lanes in this application, CPU interference is effectively reduced, significantly improving data transmission speeds between local devices, and effectively reducing data transmission latency.
[0170] In order to maximize the performance of GPU cards and disks, the three-stage intelligent computing architecture system prefers to use a higher version of PCIe-GEN5 / 6 (i.e., the first type of PCIe protocol) for data transmission in this functional area inside the GPU BOX. Without occupying CPU Lane, it supports each GPU card to establish a link binding with the high-speed network chip and realize direct read and write access to the disk.
[0171] The application discloses a three-stage intelligent computing architecture system based on the PCIe protocol, which realizes the interaction between GPU cards in different GPU BOXes without occupying the CPU Lane. Specifically, it includes two situations: the first situation is the interaction between different GPUs in the same node (i.e., the same server), and the second situation is the interaction between different GPUs in different nodes (i.e., the same server).
[0172] For scenario 1: single-server multi-GPU BOX mode, the constructed GPUDirect P2P connection channel enables communication between different GPUs on the same node (i.e., the same server) (i.e., GPU-to-GPU communication). The following describes the data transmission process in detail, using deep learning training as an example. Deep learning training programs (such as TensorFlow and PyTorch) initiate collaborative computations between multiple GPUs. For example, in model parallelism, different GPUs are responsible for different parts of the model and need to exchange intermediate computation results to complete the forward and backward propagation of the entire model. For example, when GPU BOX A needs to copy data to GPU B, the specific transmission process includes: 1) Responding to a request instruction (i.e., a request instruction for GPU BOX A to copy data to GPU B), obtaining request information for the request instruction; the request information includes the source (i.e., GPU BOX A), the destination (i.e., GPU BOX B), the transfer length, and the transfer mode. 2) Responding to an initialization configuration instruction, invoking the control module (SoC chip) within the source (i.e., GPU BOX A) to initialize and configure the DMA controllers in the source GPU module and the source network module. 3) In response to the initialization configuration instruction, the control module (SoC chip) of the target end (i.e., GPU BOX B) is called to initialize the configuration of the DMA controller in the target end GPU module and the DMA controller in the target end network module. 4) Through the PCIe switch (external PCIe switch) between the source and target ends, a GPUDirect P2P connection channel is established between the initialized DMA controller in the source end GPU module and the DMA controller in the target end GPU module. 5) In response to the GPUDirect P2P connection channel establishment completion signal, the source end GPU is triggered to execute the target data sending instruction. 6) Read the target data from the source GPU memory based on the transmission length and transmission mode (i.e., read the data from the GPU memory via the DMA controller of the GPU in GPU Box A). Transmit the target data to the GPU module in GPU Box B via the GPUDirect P2P connection channel (i.e., transmit the target data via the PCIe interconnect module in GPU Box A, the PCIe switch, and the PCIe interconnect module in GPU Box B). Then, write the target data to the specified location in the GPU memory via the DMA controller in the GPU module in GPU Box B. The PCIe switch is responsible for correctly routing the data to GPU Box B.
[0173] Each GPU uses the data and model parameters in the video memory to perform forward and backward propagation calculations, and sends the calculated gradient data directly to the video memory of other GPUs through the corresponding PCIe switch. At the same time, it also obtains the gradient data sent by other GPUs from the PCIe switch. In this process, there is no need for the CPU to transfer and process data, achieving efficient gradient synchronization.
[0174] In the single-server multi-GPU BOX mode, GPU-to-GPU communication preferably uses the medium version of the PCIe-GEN4 / 5 protocol (i.e., the second type of PCIe protocol), which can support 20 GPU cards for horizontal topology networking communication.
[0175] For the second scenario: multi-server, multi-GPU BOX mode, the established GPUDirect RDMA connection channel enables communication between different GPUs on different nodes (i.e., different servers) (i.e., host-to-host cross-node communication). For example, deep learning training programs (such as TensorFlow and PyTorch) initiate collaborative computations across multiple GPUs across nodes. In model parallelism, different GPUs are responsible for different parts of the model and need to exchange intermediate computation results to complete the forward and backward propagation of the entire model. Taking the example of multi-level, multi-GPU distributed computing for deep learning training, specifically when GPU BOX A needs to copy data to GPU BOX, the specific transmission process includes: 1) Responding to a request instruction (i.e., a request instruction for GPU BOX A to copy data to GPU BOX), obtaining the request instruction's request information; the request information includes the source, destination, transfer length, and transfer mode. 2) Responding to an initialization configuration instruction, calling the control module on the source side (i.e., GPU BOX A) to initialize and configure the DMA controllers in the source GPU module and the source network module. 3) In response to the initialization configuration instruction, the control module of the target end (i.e., GPU BOX B) is called to initialize the configuration of the DMA controller in the target end GPU module and the DMA controller in the target end network module. 4) Through the RDMA switch (i.e., RDMA switch module) between the source and target ends, an RDMA connection channel, i.e., a GPUDirect RDMA connection channel, is established between the initialized DMA controller in the source end network module and the DMA controller in the target end network module. 5) In response to the RDMA connection channel establishment completion signal, the source end GPU is triggered to execute the instruction to send the target data. 6) Based on the transmission length and transmission mode, the target data is read from the GPU memory of the source end (i.e., the GPU DMA controller in GPU BOX A reads the data from the GPU memory), transmits the target data to the GPU module of the target end via the GPUDirect RDMA connection channel, and writes the target data to the specified location of the GPU memory of the target end GPU module via the DMA controller in the target end GPU module.
[0176] Specifically, the data is transmitted to the network module of the source end through the PCIe interconnection module of the source end, the first server corresponding to the source end is determined through the OSFP interface connected to the network module, and the data is transmitted to the RDMA switch through the first server; the second server corresponding to the RDMA switch and located at a different node from the source end is determined, and the target data is transmitted to the network module of the target end through the second server; and then the target data is written to the specified location of the video memory of the GPU module of the target end through the PCIe interconnection module of the target end.
[0177] In terms of host-to-host network interconnection technology for multiple intelligent computing servers, it is preferred to adopt a higher version of the PCIe-GEN5 / 6 protocol to achieve 2Tb-8Tb high-speed node network switching interconnection for a single intelligent computing server.
[0178] Furthermore, the CPU unit in this application is compatible with the products of all domestic trusted computing CPU manufacturers, while ensuring that the performance of the corresponding intelligent computing server is not affected.
[0179] In this application, the interconnection pattern of different modules within the GPU BOX constitutes the first segment of the intelligent computing architecture; the interaction pattern between different GPU BOXes constitutes the second segment of the intelligent computing architecture; and the interaction pattern between the GPU BOX and the CPU unit constitutes the third segment of the intelligent computing architecture. Among them, the second segment of the intelligent computing architecture includes the interaction between different GPU BOXes on the same node, as well as the interaction between different GPU BOXes on different nodes. Different intelligent computing architecture segments allow the system to integrate hardware resources of different types and performances. This heterogeneous computing approach can fully utilize the advantages of different hardware and improve the overall computing efficiency of the system. Furthermore, the three-stage intelligent computing architecture system can effectively improve computer performance. Specifically, the first stage of the intelligent computing architecture optimizes the interconnection of modules within the GPU unit, which can reduce the data transmission delay between modules and increase the speed of data processing; the interaction mode between GPU units of the same node and different nodes in the second stage of the intelligent computing architecture can give full play to the parallel computing advantages of the GPU. Multiple GPU units can process different data subsets at the same time, and then merge data and summarize results through an efficient interaction mode, effectively enhancing the parallel computing capability; the third stage of the intelligent computing architecture enables the GPU unit and the CPU unit to work together. The CPU can be responsible for task scheduling, data preprocessing and other tasks, while the GPU focuses on large-scale parallel computing. This collaborative computing method can make full use of the control capabilities of the CPU and the computing power of the GPU to improve the overall computing performance of the system.
[0180] The application adopts a three-stage architecture, modularizing the GPU BOX, PCIe Switch, RDMA switch module and CPU unit. Each module is relatively independent, which facilitates flexible configuration and expansion according to different application scenarios and needs. A hot-swappable connection channel is established between the GPU BOX and the CPU unit. During system operation, related components can be easily plugged in and out without powering off the target server, effectively improving the maintainability and availability of the system.
[0181] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A GPU hot-swap control method, characterized in that: include: By acquiring the actual register value corresponding to the status information of the hot-swap button of the GPU BOX from the control unit; the status information is any one of being continuously pressed, briefly pressed, and not pressed; the GPU BOX is used to install the GPU card; Based on a preset mapping relationship between the register value and different state information of the hot-swap button, obtaining the state type corresponding to the actual register value; Get the slot information where the GPU BOX is connected to the target server; Determining the installation status of the GPU BOX according to the slot information; Determine a target request of the GPU BOX according to the installation status and the status type; The operating state of the GPU BOX is controlled based on the target request.
2. The GPU hot-swap control method according to claim 1, wherein: The determining, according to the installation status and the status type, a target request of the corresponding GPU BOX includes: When the installation state is that the GPU BOX is plugged in with the target server and the state type is that it is expected to be unplugged, determining that the target request corresponding to the GPU BOX is a power-off unplugging request; When the installation status is that the GPU BOX is not plugged in to the target server and the status type is expected to be unplugged, it is determined that the target request corresponding to the GPU BOX is a power-off unplug request.
3. The GPU hot-swap control method according to claim 2, wherein: The controlling the operating state of the GPU BOX based on the target request includes: In response to the power-off unplug request, the GPU in the corresponding GPU BOX is removed from the task scheduling list through the configured hot-plug service software; According to the GPU load balancing scheduling algorithm, a target GPU is determined from other GPUs, and the target GPU executes the tasks to be processed by the removed GPU.
4. The GPU hot-swap control method according to claim 1, wherein: The determining, according to the installation status and the status type, a target request of the corresponding GPU BOX includes: When the installation status is that the GPU BOX is plugged into the target server and the status type is expected insertion, determining that the target request corresponding to the GPU BOX is an insertion power-on request; When the installation status is that the GPU BOX is not plugged in to the target server and the status type is expected to be plugged in, it is determined that the target request corresponding to the GPU BOX is an insertion power-on request.
5. The GPU hot-swap control method according to claim 4, wherein: The controlling the operating state of the GPU BOX based on the target request includes: In response to the insertion power-on request, powering on the GPU in the newly inserted GPU BOX through configured hot-plug service software; In response to the power-on completion instruction, the GPU in the newly inserted GPU BOX is added to the task scheduling list.
6. The GPU hot-swap control method according to claim 1, wherein: When multiple actual register values sent from the control unit are received, corresponding target requests are executed in order of preset priorities.
7. The GPU hot-swap control method according to claim 1, wherein: Also includes: In response to a new GPUBOX introduction request, the ROM software on the target server scans the corresponding PCIe slots, identifies an empty slot dedicated to the GPU, and marks the empty slot as an independent slot resource; The corresponding slot information is configured for the new GPU BOX according to the independent slot resource; the slot information includes one or both of a physical address and a priority of the corresponding slot.
8. The GPU hot-swap control method according to claim 1, wherein: It also includes configuring security processing logic; The configuration security processing logic includes: sending a modification protection mechanism instruction to the BIOS through the BIOS setup interface; In response to the modification protection mechanism instruction, triggering the BIOS to perform a preset modification of the internal initial security processing logic; The preset modification includes: when the BIOS detects that an event of adding or deleting PCIe hardware occurs, recording the event information and transmitting it to the target operating system.
9. The GPU hot-swap control method according to claim 1, wherein: Also includes: Monitor the GPU operating status of each target server in real time, and trigger an early warning signal and perform task migration when the GPU operating status information is abnormal; The GPU operation status information abnormality includes: the GPU core processor occupancy rate exceeds a first preset threshold, or the GPU memory occupancy rate exceeds a second preset threshold.
10. A GPU BOX, characterized in that: include: The box body has a receiving slot for installing a GPU module, and the box body is provided with a hot-swap button; the GPU module includes a GPU card; The box body is equipped with a control mainboard and a network module. The control mainboard is internally integrated with a slave control unit, and the slave control unit is used to obtain the actual register value corresponding to the status information of the hot-swap button; The control mainboard is provided with a connector that matches the plug-in slot of the target server, and the network module is connected to the target server network via the connector.
11. The GPU BOX according to claim 10, wherein: The control mainboard is provided with a PCIe slot for plugging in the GPU module; The box body also includes a power supply module connected to the GPU module; The control mainboard also integrates a control module, a storage module and a PCIe interconnection module, and the GPU module, the control module, the storage module and the network module are all connected to the PCIe interconnection module; The end of the power module is matched with the plug-in slot of the target server.