Distributed resource management method, apparatus and system, and device and storage medium
The switches in the distributed resource management system connect the target devices and computing units to realize resource pooling and coordinated scheduling, solving the efficiency problems caused by resource isolation in the server, and improving the system operation efficiency and flexibility.
Patent Information
- Application Number
- PCT/CN2024/095399
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-05-27
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, various types of resources in the server architecture are in an isolated state, resulting in a single linear resource scheduling method, which cannot work together, affecting the overall operational efficiency.
Through the distributed resource management system, the switch is used to connect the target device and the target computing unit based on a high-speed serial cache coherence bus to realize resource pooling, and synchronous power-on, reset operations and resource scheduling, including resource reset and allocation when power-on or reset instructions are received.
It improves the efficiency and flexibility of resource management, realizes efficient coordination and life cycle management of resources, and improves the consistency and reliability of system operation.
Smart Images

Figure CN2024095399_03072025_PF_FP_ABST
Abstract
Description
Distributed resource management method, device, system, equipment and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 27, 2023, with application number 202311824720.5, and application name “Distributed Resource Management Method, Apparatus, System, Equipment and Storage Medium”, all contents of which are incorporated by reference into this application. Technical Field
[0003] The present application belongs to the field of computer technology, and in particular relates to a distributed resource management method, apparatus, system, device, and non-volatile readable storage medium. Background Art
[0004] Scenarios such as artificial intelligence and machine learning, high-performance computing, and cloud and edge computing environments are complex and diverse. To meet resource demands, server hardware architectures must be optimized and restructured to improve resource utilization and reduce maintenance costs. In related technologies, various resource types within server architectures are isolated, and correspondingly, the resource scheduling methods for each type of resource are linear, meaning that different types of resources cannot collaborate, which impacts the overall operational efficiency of the server.
[0005] Summary of the Invention
[0006] The present application provides a distributed resource management method, apparatus, system, device, and non-volatile readable storage medium.
[0007] In a first aspect, the present application provides a distributed resource management method, which is applied to a distributed resource management system deployed in a server. The distributed resource management system includes a switch and multiple resource pools. The multiple resource pools are obtained by connecting the switch to a first resource corresponding to a target device in the server or a second resource corresponding to a target computing unit in the server based on a high-speed serial cache coherence bus. The method includes:
[0008] When receiving a power-on command, the switch, the target device, and the target computing unit are controlled to be powered on synchronously;
[0009] When a reset instruction is received, a reset operation is performed on the device to be reset indicated by the reset instruction; the device to be reset includes at least one of a target device, a target computing unit, and a switch;
[0010] Based on the resource scheduling request, target resources in multiple resource pools are scheduled; resource scheduling includes resource reset and resource allocation.
[0011] In some embodiments, the method further comprises:
[0012] When target fault information corresponding to any resource pool is collected, the fault location, fault type, and fault recovery strategy are determined based on the target fault information.
[0013] In some embodiments, the switch includes a core switch and multiple access switches, the core switch is connected to the multiple access switches, and each access switch is used to connect to multiple target devices of the same type or to multiple target computing units of the same type based on a high-speed serial cache consistency bus.
[0014] In some embodiments, the method further comprises:
[0015] Obtain resource usage information corresponding to the target device and target computing unit;
[0016] Based on the resource usage information, resource analysis is performed on the target device and the target computing unit to obtain resource monitoring data.
[0017] In some embodiments, a switch is deployed in a switch chassis, a target device is deployed in a device chassis, and a target computing unit is deployed in a host chassis; upon receiving a power-on instruction, controlling the switch, the target device, and the target computing unit to be powered on synchronously includes:
[0018] When a baseboard controller in the switch chassis receives a power-on signal and power-on success signals sent from the device chassis and the host chassis, the baseboard controller controls the logic unit in the switch chassis to supply power to the switch based on the first enable signal; the power-on signal is generated based on the power-on instruction;
[0019] The baseboard controller in the control switch chassis sends a power-on signal to the baseboard controller in the device chassis and the baseboard controller in the host chassis to supply power to the target device and the target computing unit.
[0020] In some embodiments, the power-on signal sent to the logic unit in the device chassis is used for the logic unit in the device chassis to power the target device based on the second enable signal;
[0021] The power-on signal sent to the logic unit in the host chassis is used for the logic unit in the host chassis to power the target computing unit based on the third enable signal.
[0022] In some embodiments, the power-on success signal includes a first power-on success signal; and the method includes:
[0023] A power supply unit of the control device chassis supplies power to various components in the device chassis based on the standby voltage;
[0024] After each component in the device chassis receives the standby voltage, a power-on success signal is generated and sent to the logic unit and baseboard controller of the device chassis;
[0025] The baseboard controller of the control device chassis sends a first power-on success signal to the baseboard controller in the switch chassis.
[0026] In some embodiments, the power-on success signal includes a second power-on success signal; and the method further includes:
[0027] The power supply unit of the host chassis controls the power supply to each component in the host chassis based on the standby voltage;
[0028] After each component in the host chassis receives the standby voltage, a power-on success signal is generated and sent to the logic unit and baseboard controller of the host chassis;
[0029] The baseboard controller of the control host chassis sends a second power-on success signal to the baseboard controller in the switch chassis.
[0030] In some embodiments, the method further comprises:
[0031] Controlling a baseboard controller in the host chassis to scan first interfaces corresponding to the host chassis and the switch chassis to obtain a first topology corresponding to the host chassis and the switch chassis;
[0032] The baseboard controller in the switch chassis is controlled to scan the second interface corresponding to the device chassis and the switch chassis to obtain a second topology corresponding to the switch chassis and the device chassis.
[0033] In some embodiments, the resource scheduling request includes a resource release request; and resource scheduling is performed on target resources in multiple resource pools based on the resource scheduling request, including:
[0034] In the case of receiving a resource release request, removing the device to be adjusted corresponding to the resource to be adjusted indicated by the resource release request from the second topology map and determining device information corresponding to the device to be adjusted; the target resources in the multiple resource pools include the resource to be adjusted;
[0035] Reset the device to be adjusted based on the device information.
[0036] In some embodiments, resetting the device to be adjusted includes:
[0037] sending a first reset signal to a logic unit in the switch chassis based on a baseboard controller in the switch chassis;
[0038] The first reset signal is sent via the target interface to the logic unit in the device chassis corresponding to the device to be adjusted by the logic unit in the switch chassis; the logic unit is used to forward the first reset signal to the device to be adjusted to achieve reset.
[0039] In some embodiments, the resource scheduling request further includes a resource acquisition request; and the method further includes:
[0040] Based on the resource acquisition request, the resources to be adjusted are allocated to the designated computing unit indicated by the resource acquisition request.
[0041] In some embodiments, the reset instruction includes a system reset instruction, and the device to be reset includes a switch, a target computing unit, and a target device; when the reset instruction is received, performing a reset operation on the device to be reset indicated by the reset instruction includes:
[0042] Based on the system reset instruction, the target computing unit generates a system reset signal; the target computing unit implements a reset operation on the target computing unit based on the system reset signal;
[0043] The target computing unit is controlled to send a system reset signal to the logic unit in the switch chassis and the logic unit in the device chassis to implement the reset operation on the switch and the target device.
[0044] In some embodiments, sending a system reset signal to a logic unit in a switch chassis and a logic unit in a device chassis includes:
[0045] Controlling the target computing unit to send a system reset signal to the baseboard controller in the switch chassis via the baseboard controller in the host chassis;
[0046] Controlling the baseboard controller in the switch chassis to send a system reset signal to the logic unit and the switch in the switch chassis to perform a reset operation on the switch;
[0047] The baseboard controller in the control switch chassis sends a system reset signal from the target interface to the logic unit of the device chassis; the system reset signal is used for the logic unit of the device chassis to perform a reset operation based on the system reset signal.
[0048] In some embodiments, the reset instruction includes a device reset instruction, and the device to be reset includes a target reset device; when the reset instruction is received, performing a reset operation on the device to be reset indicated by the reset instruction includes:
[0049] Based on the device reset instruction, the baseboard controller in the switch chassis generates a device reset signal and sends the device reset signal to the logic unit in the switch chassis;
[0050] The logic unit in the control switch chassis sends the device reset signal from the target interface to the logic unit in the target device chassis corresponding to the device reset instruction, so as to perform a reset operation on the target reset device indicated by the device reset instruction.
[0051] In some embodiments, the distributed resource management system further includes a general management controller; and the method further includes:
[0052] The device asset information and interface connection status information corresponding to the target device in the distributed resource management system are obtained through the general management controller.
[0053] In a second aspect, the present application provides a distributed resource management device, which is applied to a distributed resource management system deployed in a server. The distributed resource management system includes a switch and multiple resource pools. The multiple resource pools are obtained by connecting the switch to a first resource corresponding to a target device in the server or a second resource corresponding to a target computing unit in the server based on a high-speed serial cache coherence bus. The device includes:
[0054] A first control module is configured to control the switch, the target device, and the target computing unit to be powered on synchronously upon receiving a power-on instruction;
[0055] A first reset module is configured to, upon receiving a reset instruction, perform a reset operation on a device to be reset indicated by the reset instruction; the device to be reset includes at least one of a target device, a target computing unit, and a switch;
[0056] The first scheduling module is used to perform resource scheduling on target resources in multiple resource pools based on resource scheduling requests; resource scheduling includes resource resetting and resource allocation.
[0057] In a third aspect, the present application provides a distributed resource management system, wherein the distributed resource management system is used to execute any of the distributed resource management methods in the first aspect.
[0058] In a fourth aspect, the present application provides an electronic device comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the distributed resource management method of any one of the above-mentioned first aspects is implemented.
[0059] In a fifth aspect, the present application provides a non-volatile readable storage medium. When the instructions in the non-volatile readable storage medium are executed by a processor of an electronic device, the electronic device can execute the steps in the distributed resource management method in some embodiments of the first aspect mentioned above.
[0060] In some embodiments of the present application, a switch in a distributed resource management system is used to connect the first resource corresponding to the target device and the second resource corresponding to the target computing unit based on a high-speed serial cache consistency bus to form a resource pool, thereby realizing hardware decoupling of the first resource and the second resource in the server. Through the switch, the first resource and the second resource in the distributed resource management system can be efficiently coordinated, thereby improving the operating efficiency of the distributed resource management system. At the same time, when a power-on instruction is received, the units in the distributed resource management system can be controlled to be powered on in a centralized manner, thereby improving the operating consistency of the distributed resource management system. When a reset instruction is received, a reset operation can be performed on the device to be reset indicated by the reset instruction, and resource scheduling (including resource reset and resource allocation) can be performed on the target resources in multiple resource pools based on a resource scheduling request. In this way, the distributed resource management system in some embodiments of the present application can realize the reset of the entire system or device, and support resource reset and resource reallocation during the resource scheduling process, providing a more efficient and flexible resource management architecture, realizing life cycle management of pooled resources, improving the practicality and flexibility of resource management to a certain extent, and the overall resource management solution of the distributed resource management system also further improves the system operating efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0062] FIG1 is a flowchart of a distributed resource management method according to an embodiment of the present application;
[0063] FIG2 is a schematic diagram of a server deployed with a distributed resource management system provided in an embodiment of the present application;
[0064] FIG3 is a schematic diagram of a connection architecture between a switch and a resource pool provided in an embodiment of the present application;
[0065] FIG4 is a schematic diagram of a distributed pooled whole-machine solution provided in an embodiment of the present application;
[0066] FIG5 is an overall block diagram of a distributed pooling management software provided in an embodiment of the present application;
[0067] 6 is a topological diagram of a distributed resource management system with coordinated power-on and power-off functions provided by an embodiment of the present application;
[0068] FIG7 is a flowchart of a resource scheduling step provided by an embodiment of the present application;
[0069] FIG8 is a topological diagram of a distributed resource management system reset function provided in an embodiment of the present application;
[0070] FIG9 is a structural diagram of a distributed resource management device provided in an embodiment of the present application;
[0071] FIG10 is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0072] The following will be combined with the accompanying drawings of some embodiments of the present application to clearly and completely describe the technical solutions of some embodiments of the present application. Obviously, some of the embodiments described are part of the embodiments of the present application, but not all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0073] Figure 1 is a step flow chart of a resource management method provided in some embodiments of the present application, which is applied to a distributed resource management system deployed in a server. The distributed resource management system includes a switch and multiple resource pools. The multiple resource pools are obtained by connecting the switch to the first resource corresponding to the target device in the server or the second resource corresponding to the target computing unit in the server based on a high-speed serial cache consistency bus.
[0074] In some embodiments of the present application, the server may be a distributed resource pooling server, in which a distributed resource management system is deployed. The distributed resource management system may be a resource management system with a switch as the core, which may specifically decouple and pool different resources within the server based on the switch, and realize collaborative scheduling of resources on this basis. The distributed resource management system includes a general management controller, a management engine, a switch, and multiple resource pools. Among them, the general management controller can be used to manage each resource pool and switch, for example: asset management, power on and off management, etc. The general management controller may be a pooled system management controller (PSMC), the management engine may be a pooled management engine, and the switch may be a CXL Switch, that is, a switch based on a high-speed serial cache consistency bus. The target devices in the server may include memory (DDR, Double Data Rate SDRAM), hard disks may include NVME (Non Volatile Memory Host Controller Interface Specification), SSD (Solid State Disk), E3.S (Enterprise and Data Center Standard Form Factor E3.S), acceleration devices may include FPGA (Field Programmable Gate Array), GPU (graphics processing unit), and other devices. The target computing unit may be a processor (CPU, Central Processing Unit). Accordingly, the first resource may include processor resources, the second resource may include memory resources, storage resources, and acceleration resources, and the multiple resource pools may include a processor resource pool, a memory resource pool, a storage resource pool, and an acceleration resource pool.
[0075] In a possible implementation, FIG2 shows a schematic diagram of a server deployed with a distributed resource management system. As shown in FIG2 , a server can be divided into multiple distributed resource management systems according to demand, and multiple distributed resource management systems can be uniformly managed by a data center monitoring and management platform. The number of resource pools in different distributed resource management systems can be set and divided according to demand. For example, for a distributed resource management system, it can include a general management controller (such as the resource pool whole machine management system in FIG2 ), a management engine (such as the resource pool management engine in FIG2 ), a switch (such as the high-performance switching unit in FIG2 ) and multiple resource pools (including a general computing unit resource pool, a heterogeneous computing unit resource pool, a memory resource pool and a storage resource pool). The resource pool whole-machine management system and the resource pooling management engine can manage, monitor and deploy and control switches (such as the high-performance switching unit in Figure 2), general computing units (including CPUs), heterogeneous computing units including GPUs, FGPAs, ASICs (Application Specific Integrated Circuit, integrated circuit chip technology), DPUs (Data Processing Units, data processors), memories (including DRAM, Dynamic Random Access Memory, dynamic random access memory), hard disks (including SSDs) and various resource pools based on Ethernet (Eth, EtherNet). Among them, the resource pool whole-machine management system can include the functions of the remote monitoring management interface (Application Programming Interface, API), the whole system reset management function, the whole system power on and off management function and the centralized asset management function. The resource pooling management engine can include the functions of topology identification, topology display and dynamic resource allocation.
[0076] In some embodiments of the present application, the switch is connected to the target device and the target computing unit based on a high-speed serial cache coherent bus (CXL, Compute Express Link). In order to enable the switch to fully meet the usage requirements, the switch may include a core switch and an access switch. Specifically, the core switch may be connected to multiple access switches. Exemplarily, the core switch and multiple access switches may be connected in a star topology, and the core switch and the access switch may constitute a high-performance switching unit. The core switch and the access switch may include one or more switching chips, and each access switch is connected to a resource pool. For any access switch, it can be connected to the same type of target device or the same type of target computing unit. That is, an access switch can be connected to multiple memories to build a memory resource pool; an access switch can be connected to multiple hard disks to build a storage resource pool; an access switch can be connected to multiple acceleration devices to build an acceleration resource pool; an access switch can be connected to multiple processors to build a processor resource pool. It is connected to the core switch through multiple access switches. In this way, memory resources, processor resources, heterogeneous acceleration resources, storage resources and other resources are distributed and decoupled into pools through a high-speed serial cache consistency bus, and the switch is used to allocate computing power and resources within the server.
[0077] For example, Figure 3 shows a schematic diagram of a switch-resource pool connection architecture. As shown in Figure 3, a core switch is connected to five access switches, which are respectively connected to a processor resource pool, a storage resource pool, an acceleration resource pool, and a memory resource pool. It is understood that for target devices or computing units of the same type, multiple resource pools can be divided as needed and connected to access switches. For example, the two access switches shown in Figure 3 are connected to two memory resource pools.
[0078] Exemplarily, FIG4 shows a schematic diagram of a distributed pooled whole-machine solution, and FIG4 includes the whole-machine architecture corresponding to the distributed resource management system, including: Ethernet switches, processors, switches (such as CXL switches), memory resource pools, acceleration resource pools, storage resource pools, and infrastructure. Among them, the contents pointed to by the two ends of the bidirectional arrows in FIG4 can be used as master devices or slave devices during the wiring process, and the contents pointed to by the unidirectional arrows can be used as slave devices during the wiring process. For example, if bidirectional arrow 1 points to the switch and the acceleration resource pool, it can indicate that the acceleration resource pool can be used as a master device to read resource data of other devices through the switch. If the unidirectional arrow 1 points to the switch and the memory resource pool, it can indicate that the external device can read data in the memory resource pool based on the switch.
[0079] Exemplarily, a distributed pooled resource system can be implemented through distributed pooled management software. FIG5 shows an overall block diagram corresponding to a distributed pooled management software. As shown in FIG5 , the upper-layer management software may include a boot loader, an operating system kernel, and a management software layer. The boot loader and the operating system kernel include a management unit and various hardware drivers, such as an inter-IC bus (I2C) driver, a universal asynchronous receiver / transmitter (UART) driver, and a serial peripheral interface (SPI) driver, and provide a unified upper-layer interface and management services for architecture platforms such as x86 (microprocessor) and ARM (Advanced RISC Machine, processor); the management software layer may include commonly used applications such as firmware management, energy saving and heat dissipation, log management, fault diagnosis, and remote control. The upper-layer management software can communicate with the unified management module (UMM) based on a standard software interface (Restful API), thereby implementing various functions of the distributed pooled resource system. Furthermore, users can directly manage the distributed pooled resource system by operating a web interface (such as buttons, selection boxes, etc.). The Unified Management Module (UMM) provides standard RESTful APIs to upper-level management software. It also communicates with system hardware (memory, storage devices, I / O devices, switch modules, network modules, cooling modules, power supply modules, etc.) through the Unified Management Module Interface (UMMI), enabling upper-level software to control and manage system hardware through the UMM. The UMMI can include interfaces such as the Power Management Bus (PMBus), System Management Bus (SMBus), PCI Express (PCIe), Cache Coherent XL (CXL), Universal Asynchronous Receiver / Transmitter (UART), Inter-IC Bus (I2C), and Serial Peripheral Interface (SPI). Specifically, device management can be performed using the PMBus and SMBus; in-band management can be performed using the PCIe and CXL; out-of-band management can be performed using the Universal Asynchronous Receiver / Transmitter (UART) and Inter-IC Bus (I2C); and security management can be performed using the SPI.
[0080] As shown in FIG1 , the method may include:
[0081] Step 101: When a power-on instruction is received, the switch, the target device, and the target computing unit are controlled to be powered on synchronously.
[0082] In some embodiments of the present application, when the distributed resource management system receives a power-on instruction, it can control the switch, target device, and target computing unit to be powered on synchronously based on the power-on instruction, and correspondingly, it can also control the switch, target device, and target computing unit to be powered off synchronously. The power-on instruction can be triggered based on a preset action, such as pressing the power button of the switch chassis corresponding to the switch. Based on the power-on instruction, a power-on signal is generated and given to the baseboard controller and the logic unit in the switch chassis, and the logic unit controls the power-on of the switch, and based on the power-on signal transmission between the baseboard controller in the switch chassis and the baseboard controllers in the host chassis and the device chassis, the power-on of the target computing unit corresponding to the host chassis and the target device corresponding to the device chassis is controlled. The communication between the baseboard controller in the switch chassis and the baseboard controllers in the host chassis and the device chassis can be achieved through an Ethernet network.
[0083] Step 102: When a reset instruction is received, a reset operation is performed on the device to be reset indicated by the reset instruction; the device to be reset includes at least one of a target device, a target computing unit, and a switch.
[0084] In some embodiments of the present application, when the distributed resource management system receives a reset instruction, a reset operation is performed on the device to be reset indicated by the reset instruction. The reset instruction may carry an indication of the device to be reset that needs to be reset, and the device with reset is reset based on the reset instruction. The reset instruction may be issued by a processor, and the device to be reset may include at least one of a target device, a target computing unit, and a switch. That is, the distributed resource management system may reset a single target device, a target computing unit, and a switch, or may reset multiple target devices, target computing units, and / or switches. It is understandable that the device to be reset may include a target device, a target computing unit, and a switch, that is, the distributed resource management system may support the function of resetting the entire system.
[0085] In actual application scenarios, when a system error occurs or a device cannot be recognized, a reset instruction may be triggered and the system and / or hardware (device) may be restarted based on the reset instruction.
[0086] Step 103: Based on the resource scheduling request, resource scheduling is performed on target resources in multiple resource pools; resource scheduling includes resource resetting and resource allocation.
[0087] In some embodiments of the present application, based on a resource scheduling request, resource scheduling can be performed on the target resource indicated by the resource scheduling request based on the content indicated in the resource scheduling request. The resource scheduling request may include a resource adjustment request, a resource release request, and a resource acquisition request. Exemplarily, a resource adjustment request may be used to indicate that the target resource corresponding to the first task is released and the target resource is allocated to the second task. A resource release request may be used to indicate that the target resource corresponding to the first task is released, and a resource acquisition request may be used to indicate that the target resource is allocated to the second task. The target resource is any one or more resources in a plurality of resource pools, for example, the target resource may be a memory resource, an acceleration resource, a processor resource, and the like. Resource scheduling may include resource reset (resource release) and resource matching. Resource reset may be achieved by performing a reset operation on the target device or target computing unit corresponding to the resource, for example, when the target resource is a memory resource, the memory resource may be reset by resetting the memory device. Resource matching may be matching resources to corresponding tasks on demand.
[0088] In summary, in some embodiments of the present application, a switch in a distributed resource management system is used to connect the first resource corresponding to the target device and the second resource corresponding to the target computing unit based on a high-speed serial cache consistency bus to form a resource pool, thereby realizing hardware decoupling of the first resource and the second resource in the server. Through the switch, the first resource and the second resource in the distributed resource management system can be efficiently coordinated, thereby improving the operating efficiency of the distributed resource management system. At the same time, when a power-on instruction is received, the units in the distributed resource management system can be controlled to be powered on in a centralized manner, thereby improving the operating consistency of the distributed resource management system. When a reset instruction is received, a reset operation can be performed on the device to be reset indicated by the reset instruction, and resource scheduling (including resource reset and resource allocation) can be performed on the target resources in multiple resource pools based on a resource scheduling request. In this way, the distributed resource management system in some embodiments of the present application can realize the reset of the entire system or device, and in the resource scheduling process, support resource reset and resource reallocation, provide a more efficient and flexible resource management architecture, realize the life cycle management of pooled resources, improve the practicality and flexibility of resource management to a certain extent, and the overall resource management solution of the distributed resource management system also further improves the system operation efficiency.
[0089] Optionally, some embodiments of the present application further include the following steps:
[0090] Step 201: When target fault information corresponding to any resource pool is collected, the fault location, fault type, and fault recovery strategy are determined based on the target fault information.
[0091] In some embodiments of the present application, the distributed resource management system also includes a general management controller and a node management controller. The node management controller is used to perform power on and off management and asset management of the resources and equipment corresponding to each resource pool. The general management controller serves as a central management node, and the node management controllers are distributed nodes and are uniformly managed by the general management controller. The general management controller is used to perform remote monitoring management, whole system reset management, whole system power on and off management, and centralized asset management.
[0092] When the node management controller corresponding to any resource pool monitors, views and analyzes the health status and fault information of the resource pool, if the server in the node management controller collects the target fault information corresponding to the resource pool, the target fault information can be transmitted to the client in the master management controller through the network based on the intelligent platform management interface (IPMI) or the Redfish protocol, and the client in the master management controller determines the location of the fault, the type of fault and the fault recovery strategy based on the target fault information. Among them, the server can include a protocol layer, a parsing layer and a driver layer. The protocol layer includes the intelligent platform management interface (IPMI) and the Redfish protocol (The Redfish Scalable Platforms Management API, the specification of the scalable platform management API). The parsing layer is used to extract parameters, and the driver layer may include a JTAG interface and general purpose input and output (GPIO). The client is used to simulate various resource faults in different topologies, and may include an application layer, a functional layer and a protocol layer. The functional layer includes an error annotation script, and the protocol layer includes the intelligent platform management interface (IPMI) and the Redfish protocol. Faults can occur in areas such as fans, CPUs, memory, GPUs, storage devices, network equipment, and PCIe (PCIe Express) plug-in devices. Fault types can be categorized as downtime or non-downtime. Downtime primarily occurs during startup and during runtime. Non-downtime can include abnormal power supply temperature indicators, fan anomalies, device failures, and other non-fatal faults. Fault recovery strategies can be strategies that can repair the fault based on the target fault information.
[0093] Upon receiving target fault information, the client in the master management controller can analyze it using a fault analysis model to determine the fault location, fault type, and fault recovery strategy. This fault analysis model can be developed by continuously training a large amount of annotated fault information data until the model parameters converge, followed by fine-tuning and correction of the parameters. Specifically, the compressed and protocolized target fault information can be input into the fault analysis model, which then outputs the fault location, fault type, and fault recovery strategy.
[0094] In some embodiments of the present application, the distributed resource management system can be provided with a system fault management mechanism, which collects target fault information corresponding to each resource pool, and intelligently locates the fault location and fault type based on the target fault information, and determines the fault recovery strategy, so as to facilitate timely monitoring and discovery of faults in the distributed resource management system when faults occur, and implement rapid response strategies, thereby improving the stability and reliability of the system.
[0095] Optionally, some embodiments of the present application further include the following steps:
[0096] Step 301: Obtain resource usage information corresponding to a target device and a target computing unit.
[0097] In some embodiments of the present application, resource usage information corresponding to a target device or target computing unit can be obtained through a node management controller corresponding to each resource pool. Specifically, the node management controller can monitor the resource usage of the target device or target computing unit in real time to obtain resource usage information. The resource usage information may include processor utilization, memory utilization, network bandwidth utilization, etc.
[0098] Step 302: Based on the resource usage information, perform resource analysis on the target device and the target computing unit to obtain resource monitoring data.
[0099] In some embodiments of the present application, resource analysis can be performed on target devices and target computing units based on the resource usage represented by the resource usage information to obtain resource monitoring data. This resource monitoring data may include resource utilization, task execution time, task completion status, and the like. By obtaining this resource monitoring data, subsequent resource planning and decision-making can be performed based on the resource monitoring data, such as adding or reducing computing units, adjusting resource allocation strategies, and so on, to ensure system performance and efficiency.
[0100] Optionally, the switch is deployed in a switch chassis, the target device is deployed in a device chassis, and the target computing unit is deployed in a host chassis.
[0101] In some embodiments of the present application, the switch can be deployed in a switch chassis (SW chassis), the target device can be deployed in a device chassis (Device chassis), and the target computing unit can be deployed in a host chassis (Host chassis). The switch chassis may include a switch, a baseboard management controller (BMC), a complex programmable logic device (CPLD), a power supply unit, an onboard DC-DC power supply (Voltage regulator, VR), and other components (such as a network card). The device chassis may include a target device, a baseboard controller, a logic unit, a power supply unit, an onboard DC-DC power supply, and other components (such as a network card). The host chassis may include a target computing unit, a baseboard controller, a logic unit, a power supply unit, an onboard DC-DC power supply, and other components (such as a network card).
[0102] Step 401: When a baseboard controller in a switch chassis receives a power-on signal and power-on success signals sent from a device chassis and a host chassis, the baseboard controller controls a logic unit in the switch chassis to supply power to the switch based on a first enable signal; the power-on signal is generated based on a power-on instruction.
[0103] In some embodiments of the present application, when a power-on command is received, a power-on signal is generated and sent to the logic unit and the baseboard controller in the switch chassis. When the baseboard controller in the switch chassis receives the power-on signal and the power-on success signals sent from the device chassis and the host chassis respectively, it indicates that the switch, the target device and the target computing unit can be powered on synchronously based on the power-on signal. The logic unit in the switch chassis then sends a first enable signal to the main power supply (main DC-DC switching regulator) in the switch chassis, so that the main power supply can supply power to the switch and other components in the switch chassis.
[0104] Step 402: Control the baseboard controller in the switch chassis to send a power-on signal to the baseboard controller in the device chassis and the baseboard controller in the host chassis to supply power to the target device and the target computing unit.
[0105] In some embodiments of the present application, a baseboard controller in a switch chassis sends a power-on signal to baseboard controllers in a device chassis and a host chassis to power on a target device and a target computing unit. The power-on signal sent to the logic unit in the device chassis is used to enable the logic unit in the device chassis to power the target device based on a second enable signal; and the power-on signal sent to the logic unit in the host chassis is used to enable the logic unit in the host chassis to power the target computing unit based on a third enable signal.
[0106] Specifically, after the baseboard controller in the switch chassis sends a power-on signal to the baseboard controller in the device chassis, the baseboard controller in the device chassis sends the power-on signal to the logic unit in the device chassis via the Inter-IC Bus (I2C) / Universal Asynchronous Receiver / Transmitter (UART) interface. The logic unit then sends a second enable signal to the main power supply in the device chassis, allowing the main power supply to supply power to the target device and other components in the device chassis. After the target device and other components are powered on, a power-on completion signal can be sent again to the baseboard controller in the switch chassis via the network based on the baseboard controller in the device chassis to notify the baseboard controller in the switch chassis that the device chassis has been powered on. Correspondingly, the power-on method of the host chassis is similar to that of the device chassis. Specifically, after the baseboard controller in the switch chassis sends a power-on signal to the baseboard controller in the host chassis, the baseboard controller in the host chassis sends the power-on signal to the logic unit in the host chassis via the Inter-IC Bus (I2C) / Universal Asynchronous Receiver / Transmitter (UART) interface. The logic unit then sends a third enable signal to the main power supply in the host chassis, allowing the main power supply to supply power to the target computing unit (CPU) and other components in the host chassis. After the target computing unit and other components are powered on, a power-on completion signal can be sent again to the baseboard controller in the switch chassis via the network based on the baseboard controller in the host chassis to notify the baseboard controller in the switch chassis that the host chassis has been powered on.
[0107] In some embodiments of the present application, when the baseboard controller in the switch chassis receives a power-on signal and a power-on success signal sent from the device chassis and the host chassis, it is characterized that the switch, the target device and the target computing unit can be powered on synchronously. Therefore, the switch is powered on based on the logic unit in the switch chassis respectively, and the target device and the target computing unit can be powered on based on the interaction between the baseboard controller in the switch chassis and the baseboard controllers in the device chassis and the host chassis. In this way, the distributed resource management system can still support centralized power-on and power-off control on the basis of resource decoupling pooling, thereby achieving power-on consistency.
[0108] Optionally, the power-on success signal includes a first power-on success signal. The first power-on success signal is a power-on success signal sent by the baseboard controller in the device chassis to the baseboard controller in the switch chassis.
[0109] Some embodiments of the present application may include the following steps:
[0110] Step 501: Control a power supply unit of a device chassis to supply power to various components in the device chassis based on a standby voltage.
[0111] In some embodiments of the present application, a power supply unit (PSU) controlling the device chassis generates a standby voltage (the standby voltage can be obtained by conversion based on an on-board DC-DC power supply) and sends the standby voltage to each component in the device chassis (including a baseboard controller, a logic unit, etc.). The standby voltage is used to wake up each component in the device chassis so that each component can operate normally.
[0112] Step 502: After each component in the device chassis receives the standby voltage, a power-on success signal is generated and sent to the logic unit and baseboard controller of the device chassis.
[0113] In some embodiments of the present application, after each component in the device chassis receives a standby voltage, a power good signal (PG) is generated and sent to a logic unit in the device chassis. The logic unit then sends the power good signal to a baseboard controller in the device chassis via I2C / UART. Specifically, after the last standby voltage is sent to the corresponding component, a power good signal is generated and sent to the logic unit in the device chassis.
[0114] Step 503: The baseboard controller of the control device chassis sends a first power-on success signal to the baseboard controller in the switch chassis.
[0115] In some embodiments of the present application, when a baseboard controller in a device chassis receives a power-on success signal, the baseboard controller controlling the device chassis sends a first power-on success signal to the switch chassis. The first power-on success signal indicates that the device chassis has entered standby mode and can be powered on.
[0116] In some embodiments of the present application, power is first supplied to each component in the device chassis through the power supply unit of the device chassis, and then a power-on success signal is sent to the baseboard controller in the switch chassis. When the baseboard controller in the switch chassis receives the first power-on success signal, it can perform subsequent synchronous power-on based on the power-on signal and the second power-on success signal sent by the host chassis.
[0117] Optionally, the power-on success signal includes a second power-on success signal. The second power-on success signal is a power-on success signal sent by the baseboard controller in the host chassis to the baseboard controller in the switch chassis.
[0118] Some embodiments of the present application may include the following steps:
[0119] Step 601: Control the power supply unit of the host chassis to supply power to various components in the host chassis based on the standby voltage.
[0120] In some embodiments of the present application, a power supply unit (PSU) that controls the host chassis generates a standby voltage (the standby voltage can be obtained by converting an on-board DC-DC power supply) and sends the standby voltage to each component in the host chassis (including a baseboard controller, a logic unit, etc.). The standby voltage is used to wake up each component in the host chassis so that each component can work normally.
[0121] Step 602: After each component in the host chassis receives the standby voltage, a power-on success signal is generated and sent to the logic unit and baseboard controller of the host chassis.
[0122] In some embodiments of the present application, after each component in the host chassis receives a standby voltage, a power-on success signal (power good, PG) is generated and sent to a logic unit in the host chassis. The logic unit then sends the power-on success signal to the host chassis' baseboard controller via I2C / UART. Specifically, after the last standby voltage is sent to the corresponding component, a power-on success signal is generated and sent to the logic unit in the host chassis.
[0123] Step 603: Control the baseboard controller of the host chassis to send a second power-on success signal to the baseboard controller in the switch chassis.
[0124] In some embodiments of the present application, when the baseboard controller in the host chassis receives a power-on success signal, the baseboard controller controlling the host chassis sends a first power-on success signal to the switch chassis. The first power-on success signal is used to indicate that the host chassis has entered standby mode and can be powered on.
[0125] In some embodiments of the present application, power is first supplied to each component in the host chassis through the power supply unit of the host chassis, and then a power-on success signal is sent to the baseboard controller in the switch chassis. When the baseboard controller in the switch chassis receives the second power-on success signal, it can perform subsequent synchronous power-on based on the power-on signal and the first power-on success signal sent by the device chassis.
[0126] As an example, Figure 6 shows a topology diagram for coordinated power-on and power-off in a distributed resource management system. The steps for synchronously powering on a switch, target device, and target computing unit are described based on Figure 6: 1. After power is supplied by the power supply units in the device chassis and host chassis, the onboard DC-DC power supplies generate a standby voltage to the components in the device chassis and host chassis. 2. After receiving the standby voltage, the components in the device chassis and host chassis each generate a power-on success signal and send it to the baseboard controller in the switch chassis, which is then sent by the logic unit and baseboard controller. 3. Upon receiving both the power-on success signal and the power-on signal, the baseboard controller in the switch chassis sends the power-on signal to the logic unit in the switch chassis. The logic unit in the switch chassis sends a first enable signal to the main power supply (main DC-DC switching regulator) in the switch chassis, enabling the main power supply to supply power to the switch and other components in the switch chassis. 4. Simultaneously, the power-on signal is sent by the baseboard controller in the switch chassis to the baseboard controllers in the device chassis and host chassis, respectively. 5. The baseboard controllers in the device chassis and the host chassis send a power-on signal to the logic units in their respective chassis. The logic units then send an enable signal to the main power supply in each chassis to supply power to the target device and target computing unit.
[0127] Optionally, some embodiments of the present application may include the following steps:
[0128] Step 701: Control a baseboard controller in a host chassis to scan first interfaces corresponding to the host chassis and the switch chassis, and obtain a first topology corresponding to the host chassis and the switch chassis.
[0129] In some embodiments of the present application, a baseboard controller in a host chassis scans the first interface corresponding to the host chassis and the switch chassis. The first interface may include an interface in which the host chassis and the switch chassis are connected, such as an interface connecting the host chassis to the switch chassis and an interface connecting the switch chassis to the host chassis, to obtain local interface information corresponding to the host chassis and first interface information corresponding to the switch chassis. The local interface information corresponding to the host chassis may include identification information (ID) of the interface connecting the host chassis side to the switch chassis, and the first interface information includes identification information of the interface connecting the switch chassis side to the host chassis. Based on the local interface information corresponding to the host chassis and the first interface information corresponding to the switch chassis, a first topology map corresponding to the host chassis and the switch chassis is constructed. The first topology map may represent the connection relationship between the host chassis and the switch chassis and the corresponding relationship between the interfaces.
[0130] Step 702: Control the baseboard controller in the switch chassis to scan the second interface corresponding to the device chassis and the switch chassis to obtain a second topology corresponding to the switch chassis and the device chassis.
[0131] In some embodiments of the present application, a baseboard controller in a device chassis scans the second interface corresponding to the device chassis and the switch chassis. The second interface may include an interface in which the device chassis and the switch chassis have a connection relationship, such as an interface connecting the device chassis to the switch chassis and an interface connecting the switch chassis to the device chassis, to obtain local interface information corresponding to the device chassis and second interface information corresponding to the switch chassis. The local interface information corresponding to the device chassis may include identification information (ID) of the interface connecting the device chassis side to the switch chassis, and the second interface information includes identification information of the interface connecting the switch chassis side to the device chassis. Based on the local interface information corresponding to the device chassis and the second interface information corresponding to the switch chassis, a second topology map corresponding to the device chassis and the switch chassis is constructed. The second topology map may represent the connection relationship between the device chassis and the switch chassis and the corresponding relationship between the interfaces.
[0132] Based on the first topology map and the second topology map, it is convenient to view the information of all resources in the distributed resource management system, such as resource node type, power-on status, overall health status, management IP and other functions; it supports viewing interface topology interconnection information, and can view the connection status of each interface and the information of the target device or target computing unit corresponding to the connected resource through the Web (webpage) / Redfish page.
[0133] It is understandable that in order to improve the system operation efficiency, the first topology map and the second topology map can be obtained in advance and stored in a designated location. In this way, when operations need to be performed based on the first topology map and the second topology map, the first topology map and the second topology map can be obtained directly based on the designated location.
[0134] In some embodiments of the present application, the distributed resource management system can realize automatic discovery of resource topology and construction of topology map, support system topology view viewing function, and improve the convenience and uniformity of centralized resource management.
[0135] Optionally, the resource scheduling request includes a resource release request.
[0136] Accordingly, step 103 may include the following steps:
[0137] Step 801: upon receiving a resource release request, remove the device to be adjusted corresponding to the resource to be adjusted indicated in the resource release request from the second topology map and determine device information corresponding to the device to be adjusted; target resources in multiple resource pools include the resource to be adjusted.
[0138] In some embodiments of the present application, upon receiving a resource release request, the management engine can determine the resource to be adjusted indicated by the resource release request, and hot-remove the device to be adjusted corresponding to the resource to be adjusted from the second topology map, indicating that the device to be adjusted is no longer exchanging data. Based on the second topology map, the device information corresponding to the device to be adjusted is determined, wherein the resource to be adjusted may include a first resource and a second resource, and the device to be adjusted may include a target computing unit and a target device. The device information may include the physical location of the device. Exemplarily, the device to be adjusted and information related to the device to be adjusted can be removed from the second topology map.
[0139] It is understandable that before receiving a resource release request, it is necessary to ensure that the application layer process related to the resource to be adjusted has ended to avoid abnormal access by the application layer to the device to be adjusted corresponding to the resource to be adjusted, resulting in program abnormality.
[0140] Step 802: Reset the device to be adjusted based on the device information.
[0141] In some embodiments of the present application, a reset operation is performed on the device to be adjusted based on the device information. After the reset is completed, the device to be adjusted can be restarted to restore the corresponding operating state of the device to be adjusted to a default value.
[0142] In some embodiments of the present application, upon receiving a resource release request, the device to be adjusted corresponding to the resource to be adjusted may be reset to release the resource to be adjusted, thereby enabling dynamic scaling and release of resources through the management engine.
[0143] Optionally, step 802 may include:
[0144] Step 8021: Send a reset signal to a logic unit in the switch chassis based on a baseboard controller in the switch chassis.
[0145] Step 8022: The reset signal is sent via the target interface to the logic unit in the device chassis corresponding to the device to be adjusted by the logic unit in the switch chassis; the logic unit is used to forward the reset signal to the device to be adjusted to achieve reset.
[0146] In some embodiments of the present application, resetting the device to be adjusted may include generating a first reset signal based on a baseboard controller in a switch chassis, sending the first reset signal to a logic unit in the switch chassis, and transmitting the first reset signal via a target interface to a logic unit in a device chassis corresponding to the device to be adjusted by the logic unit in the switch chassis. Upon receiving the first reset signal, the logic unit in the device chassis forwards the first reset signal to the device to be adjusted, thereby resetting the device to be adjusted.
[0147] In some embodiments of the present application, the baseboard controller and logic unit in the switch chassis transmit signals to the device chassis corresponding to the device to be adjusted, thereby resetting the device to be adjusted. This releases the resources to be adjusted, improving the flexibility of resource allocation.
[0148] Optionally, the resource scheduling request also includes a resource acquisition request. The embodiment of the present application further includes the following steps:
[0149] Step 901: Based on a resource acquisition request, allocate the resources to be adjusted to the designated computing unit indicated by the resource acquisition request.
[0150] In some embodiments of the present application, after the resources to be adjusted are released, the resources to be adjusted may be allocated to a designated computing unit indicated by the resource acquisition request based on the resource acquisition request, for use by the designated computing unit.
[0151] For example, Figure 7 illustrates a flowchart of resource scheduling steps. As shown in Figure 7, Figure 7 includes a general computing unit resource pool, namely, a target computing unit resource pool, including a resource pool based on general computing units 1-n; and a heterogeneous computing unit resource pool, namely, an accelerator device resource pool, including a resource pool based on heterogeneous computing units (such as GPUs and FPGAs) 1-n. A high-performance switching unit is a switch, which can be composed of multiple switching chips. In actual application scenarios, when an application stops running, the corresponding resources can be released back to the corresponding resource pool to facilitate efficient resource flow and full utilization. For example, to release a heterogeneous accelerator card device (target device) and assign it to a general computing unit (target computing unit), first ensure that the application layer process associated with the heterogeneous accelerator card device has terminated. The user then triggers a resource scheduling request, including a resource release request and a resource acquisition request. Specifically, upon receiving the resource scheduling instruction sent by the user, the target computing unit (general computing unit) initiates a resource scheduling request to the management engine. Based on the resource release request indicated by the management engine, the device to be adjusted (heterogeneous computing unit) corresponding to the resource to be adjusted (heterogeneous acceleration card resource) is hot removed, and a request is sent to the switch to obtain the physical location of the device to be adjusted (heterogeneous computing unit) corresponding to the resource to be adjusted. Based on the physical location of the device, the heterogeneous computing power device resources are reset and the device to be adjusted is restarted to restore the operating status to the default value. When the reset is completed, the management engine reallocates the heterogeneous computing power device resources to the designated computing unit (the general computing unit indicated by the resource acquisition request) based on the resource acquisition request. In this way, the designated computing unit will see the newly added device (heterogeneous acceleration card device) without the business being aware of it, and the dynamic switching of heterogeneous computing power resources can be completed.
[0152] In some embodiments of the present application, resource acquisition requests can be used to achieve on-demand allocation of resources and dynamic resource allocation, thereby improving the flexibility of resource allocation.
[0153] Optionally, the reset instruction includes a system reset instruction, and the devices to be reset include a switch, a target computing unit, and a target device.
[0154] In some embodiments of the present application, the reset instruction may include a system reset instruction, which is used to instruct to reset the entire system. Accordingly, the device to be reset may include a switch, a target computing unit, and a target device.
[0155] Accordingly, step 102 may include the following steps:
[0156] Step 1001: Based on a system reset instruction, a target computing unit generates a system reset signal; the target computing unit implements a reset operation on the target computing unit based on the system reset signal.
[0157] Step 1002: Control the target computing unit to send a system reset signal to the logic unit in the switch chassis and the logic units of each device chassis to implement a reset operation on the switch and each target device.
[0158] In some embodiments of the present application, based on a system reset instruction, the target computing unit generates a system reset signal and sends the system reset signal to other components (such as a baseboard controller, logic unit, network card, etc.) in the host chassis corresponding to the target computing unit, thereby resetting the target computing unit and related devices. The target computing unit then sends the system reset signal to the logic unit in the switch chassis and the logic unit in the device chassis, so that the logic unit in the switch chassis and the logic unit in the device chassis perform a reset operation on the switch and the target device based on the system reset signal. For example, as shown in FIG8 , the target computing unit can send the system reset signal to the logic unit in the host chassis, which in turn sends the system reset signal to the baseboard controller in the host chassis and other devices in the host chassis. The baseboard controller in the host chassis sends the system reset signal to the baseboard controller in the switch chassis based on the Ethernet switch, and the switch is reset based on the system reset signal sent by the baseboard controller in the switch chassis via the logic unit in the switch chassis. The baseboard controller in the switch chassis sends the system reset signal to the logic unit in the device chassis based on the target interface in the switch chassis, and the target device is reset based on the system reset signal sent by the logic unit in the device chassis.
[0159] Optionally, step 1002 may include the following steps:
[0160] Step 1101: Control the target computing unit to send a system reset signal to the baseboard controller in the switch chassis via the baseboard controller in the host chassis.
[0161] In some embodiments of the present application, when the baseboard controller in the host chassis receives a system reset signal sent by the target computing unit, the system reset signal is sent to the baseboard controller in the switch chassis based on Ethernet.
[0162] Step 1102: Control the baseboard controller in the switch chassis to send a system reset signal to the logic unit and the switch in the switch chassis to perform a reset operation on the switch.
[0163] In some embodiments of the present application, a baseboard controller in a switch chassis sends a system reset signal to a logic unit in the switch chassis, and the logic unit in the switch chassis sends the system reset signal to the switch to reset the switch.
[0164] Step 1103: Control the baseboard controller in the switch chassis to send a system reset signal from the target interface to the logic unit of the device chassis; the system reset signal is used for the logic unit of the device chassis to perform a reset operation based on the system reset signal.
[0165] In some embodiments of the present application, a baseboard controller in a switch chassis sends a system reset signal to a logic unit in the switch chassis. The logic unit in the switch chassis then sends the system reset signal to a logic unit in a device chassis via a target interface. The logic unit in the device chassis then sends the system reset signal to a target device to reset the target device.
[0166] In some embodiments of the present application, a system reset signal is generated by a target computing unit, and through signal transmission between a host chassis, a switch chassis, and a device chassis, the entire system can be reset based on an automatically detected system reset signal.
[0167] Optionally, the reset instruction includes a device reset instruction, and the device to be reset includes a target reset device. Some embodiments of the present application further include the following steps:
[0168] Step 1201: Based on a device reset instruction, a baseboard controller in a switch chassis generates a device reset signal and sends the device reset signal to a logic unit in the switch chassis.
[0169] Step 1202: Control the logic unit in the switch chassis to send a device reset signal from the target interface to the logic unit in the target device chassis corresponding to the device reset instruction, so as to perform a reset operation on the target reset device indicated by the device reset instruction.
[0170] In some embodiments of the present application, upon receiving a device reset instruction, a baseboard controller in a switch chassis generates a device reset signal and sends the device reset signal to a logic unit in the switch chassis. The logic unit in the switch chassis, based on a target interface, sends the device reset signal to a logic unit in a target device chassis corresponding to the device reset instruction. The logic unit in the target device chassis then sends the device reset signal to a target reset device in the target device chassis, thereby resetting the target reset device. The target device chassis is the device chassis where the target reset device indicated by the device reset instruction is located.
[0171] In some embodiments of the present application, a device reset signal is generated by a baseboard controller in a switch chassis, and through signal transmission between the switch chassis and the device chassis, the corresponding target reset device can be reset based on the automatically detected device reset signal.
[0172] Optionally, some embodiments of the present application may include the following steps:
[0173] Step 1301: Obtain device asset information and interface connection status information corresponding to a target device in a distributed resource management system through a general management controller.
[0174] In some embodiments of the present application, the node management controller can collect asset information corresponding to the corresponding resource pool (including device asset information corresponding to the target device and computing asset information corresponding to the target computing unit) and interface connection status information, and the node management controller can send the asset information corresponding to the resource pool and the interface connection status information to the general management controller so that the general management controller can perform asset monitoring and interface connection status monitoring in the distributed resource management system.
[0175] Figure 9 is a structural diagram of a distributed resource management device provided in some embodiments of the present application, which is applied to a distributed resource management system deployed in a server. The distributed resource management system includes a switch and multiple resource pools. The multiple resource pools are obtained by connecting the switch to the first resource corresponding to the target device in the server or the second resource corresponding to the target computing unit in the server based on a high-speed serial cache consistency bus.
[0176] As shown in FIG9 , the device may specifically include:
[0177] The first control module 1401 is configured to control the switch, the target device, and the target computing unit to be powered on synchronously upon receiving a power-on instruction;
[0178] The first reset module 1402 is configured to, upon receiving a reset instruction, perform a reset operation on a device to be reset indicated by the reset instruction; the device to be reset includes at least one of a target device, a target computing unit, and a switch;
[0179] The first scheduling module 1403 is configured to perform resource scheduling on target resources in multiple resource pools based on a resource scheduling request; resource scheduling includes resource resetting and resource allocation.
[0180] Optionally, the device further comprises:
[0181] The first determination module is configured to determine a fault location, a fault type, and a fault recovery strategy based on the target fault information when target fault information corresponding to any resource pool is collected.
[0182] Optionally, the device further comprises:
[0183] A first acquisition module is used to obtain resource usage information corresponding to the target device and the target computing unit;
[0184] The first analysis module is used to perform resource analysis on the target device and the target computing unit based on the resource usage information to obtain resource monitoring data.
[0185] Optionally, the switch is deployed in a switch chassis, the target device is deployed in a device chassis, and the target computing unit is deployed in a host chassis; the first control module 1401 includes:
[0186] a first control submodule, configured to control a logic unit in the switch chassis to supply power to the switch based on a first enable signal when a baseboard controller in the switch chassis receives a power-on signal and power-on success signals sent from the device chassis and the host chassis; the power-on signal is generated based on the power-on instruction;
[0187] The second control submodule is used to control the baseboard controller in the switch chassis to send a power-on signal to the baseboard controller in the device chassis and the baseboard controller in the host chassis to supply power to the target device and the target computing unit.
[0188] Optionally, the power-on success signal includes a first power-on success signal; and the device further includes:
[0189] a second control module, configured to control a power supply unit of the device chassis to supply power to various components in the device chassis based on a standby voltage;
[0190] A first sending module is used to generate a power-on success signal and send it to the logic unit and baseboard controller of the device chassis after each component in the device chassis receives the standby voltage;
[0191] The third control module is configured to control the baseboard controller of the device chassis to send a first power-on success signal to the baseboard controller in the switch chassis.
[0192] Optionally, the power-on success signal includes a second power-on success signal; and the device further includes:
[0193] a fourth control module, configured to control a power supply unit of the host chassis to supply power to various components in the host chassis based on a standby voltage;
[0194] A second sending module is used to generate a power-on success signal and send it to the logic unit and baseboard controller of the host chassis after each component in the host chassis receives the standby voltage;
[0195] The fifth control module is configured to control the baseboard controller of the host chassis to send a second power-on success signal to the baseboard controller in the switch chassis.
[0196] Optionally, the device further comprises:
[0197] A second acquisition module is used to control the baseboard controller in the host chassis to scan the first interface corresponding to the host chassis and the switch chassis, and obtain a first topology corresponding to the host chassis and the switch chassis;
[0198] The third acquisition module is configured to control the baseboard controller in the switch chassis to scan the second interface corresponding to the device chassis and the switch chassis, and obtain a second topology corresponding to the switch chassis and the device chassis.
[0199] Optionally, the resource scheduling request includes a resource release request; the first scheduling module 1403 includes:
[0200] A second determining module is configured to, upon receiving a resource release request, remove the device to be adjusted corresponding to the resource to be adjusted indicated by the resource release request from the second topology map and determine device information corresponding to the device to be adjusted; the target resources in the plurality of resource pools include the resource to be adjusted;
[0201] The second reset module is configured to reset the device to be adjusted based on the device information.
[0202] Optionally, the second reset module includes:
[0203] a third sending module, configured to send a first reset signal to a logic unit in the switch chassis based on a baseboard controller in the switch chassis;
[0204] The fourth sending module is used to send the first reset signal to the logic unit in the device chassis corresponding to the device to be adjusted through the target interface through the logic unit in the switch chassis; the logic unit is used to forward the first reset signal to the device to be adjusted to achieve reset.
[0205] Optionally, the resource scheduling request further includes a resource acquisition request; and the apparatus further includes:
[0206] The first allocation module is configured to allocate the resources to be adjusted to the designated computing unit indicated by the resource acquisition request based on the resource acquisition request.
[0207] Optionally, the reset instruction includes a system reset instruction, and the devices to be reset include a switch, a target computing unit, and a target device; the first reset module 1402 includes:
[0208] The first generating module is configured to generate a system reset signal by the target computing unit based on the system reset instruction; the target computing unit implements a reset operation on the target computing unit based on the system reset signal;
[0209] The sixth control module is used to control the target computing unit to send a system reset signal to the logic unit in the switch chassis and the logic unit in the device chassis to implement a reset operation on the switch and the target device.
[0210] Optionally, the sixth control module includes:
[0211] A first control submodule, configured to control the target computing unit to send a system reset signal to the baseboard controller in the switch chassis via the baseboard controller in the host chassis;
[0212] A second control submodule is configured to control a baseboard controller in the switch chassis to send a system reset signal to a logic unit and a switch in the switch chassis, so as to perform a reset operation on the switch;
[0213] The third control submodule is used to control the baseboard controller in the switch chassis to send a system reset signal from the target interface to the logic unit of the device chassis; the system reset signal is used for the logic unit of the device chassis to perform a reset operation based on the system reset signal.
[0214] Optionally, the reset instruction includes a device reset instruction, and the device to be reset includes a target reset device; the first reset module 1402 includes:
[0215] a fifth sending module, configured to generate a device reset signal by a baseboard controller in the switch chassis based on the device reset instruction, and send the device reset signal to a logic unit in the switch chassis;
[0216] The seventh control module is used to control the logic unit in the switch chassis to send the device reset signal from the target interface to the logic unit in the target device chassis corresponding to the device reset instruction, so as to perform a reset operation on the target reset device indicated by the device reset instruction.
[0217] Optionally, the distributed resource management system further includes a general management controller; the device further includes:
[0218] The fourth acquisition module is used to acquire device asset information and interface connection status information corresponding to the target device in the distributed resource management system through the general management controller.
[0219] The present application also provides a distributed resource management system for executing the distributed resource management method of some embodiments of the present application.
[0220] The present application also provides an electronic device, see Figure 10, including: a processor 1501, a memory 1502, and a computer program 15021 stored in the memory and executable on the processor, and when the processor executes the program, it implements the distributed resource management method of some embodiments of the present application.
[0221] The present application also provides a non-volatile readable storage medium. When the instructions in the non-volatile readable storage medium are executed by the processor of an electronic device, the electronic device can execute the distributed resource management method of some embodiments of the present application.
[0222] For some embodiments of the device, since they are basically similar to some embodiments of the method, the description is relatively simple, and the relevant parts can be referred to the partial description of some embodiments of the method.
[0223] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems may also be used together with the teachings herein. Based on the above description, it is apparent that the structure required for constructing such systems is suitable. In addition, the present application is not directed to any specific programming language. It should be understood that various programming languages may be utilized to implement the present application described herein, and the description of the specific languages above is provided for the purpose of disclosing the preferred embodiment of the present application.
[0224] In the description provided herein, a large number of specific details are described. However, it is understood that some embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.
[0225] Similarly, it should be understood that in order to streamline the present application and aid in understanding one or more of the various inventive aspects, in the above description of some embodiments of the present application, various features of the present application are sometimes grouped together into some embodiments, figures, or descriptions thereof. However, this disclosed method should not be interpreted as reflecting the following intention: that the claimed application requires more features than the features explicitly recited in each claim. More precisely, as reflected in the claims below, inventive aspects lie in less than all the features of some of the previously disclosed embodiments. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim itself serving as some embodiments of the present application.
[0226] It will be appreciated by those skilled in the art that the modules in the devices of some embodiments may be adaptively changed and arranged in one or more devices that are different from some embodiments. The modules or units or components in some embodiments may be combined into one module or unit or component, and furthermore may be divided into a plurality of submodules or subunits or subcomponents. All features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device so disclosed may be combined in any combination, except that at least some of such features and / or processes or units are mutually exclusive. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.
[0227] Some embodiments of the various components of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the sorting device according to the present application. The present application can also be implemented as a device or apparatus program for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0228] It should be noted that some embodiments illustrate rather than limit the present application, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of appropriately programmed computers. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.
[0229] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in some embodiments of the aforementioned method and will not be repeated here.
[0230] It should be noted that all actions of acquiring signals, information or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0231] The above are only some preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.
[0232] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A distributed resource management method, characterized in that, Applied to a distributed resource management system deployed in a server, the distributed resource management system includes a switch and multiple resource pools. The multiple resource pools are obtained by the switch connecting to the first resources corresponding to target devices in the server or the second resources corresponding to target computing units in the server based on a high-speed serial cache coherence bus. The method includes: When receiving a power-on instruction, controlling the switch, the target device, and the target computing unit to power on synchronously; When receiving a reset instruction, performing a reset operation on the device to be reset indicated by the reset instruction; the device to be reset includes at least one of the target device, the target computing unit, and the switch; Based on a resource scheduling request, performing resource scheduling on the target resources in the multiple resource pools; the resource scheduling includes resource reset and resource allocation.
2. The method according to claim 1, wherein The method further includes: When acquiring target fault information corresponding to any resource pool, based on the target fault information, determining the fault location, fault type, and fault recovery strategy.
3. The method according to claim 1, characterized in that, The switch includes a core switch and multiple access switches. The core switch is connected to the multiple access switches, and each access switch is used to connect to multiple target devices of the same type or multiple target computing units of the same type based on the high-speed serial cache coherence bus.
4. The method according to claim 1, characterized in that, The method further includes: Obtaining the resource usage information corresponding to the target device and the target computing unit; Based on the resource usage information, performing resource analysis on the target device and the target computing unit to obtain resource monitoring data.
5. The method according to claim 1, wherein The switch is deployed in a switch chassis, the target device is deployed in a device chassis, and the target computing unit is deployed in a host chassis; When receiving a power-on instruction, controlling the switch, the target device, and the target computing unit to power on synchronously includes: When the baseboard controller in the switch chassis receives a power-on signal and a power-on success signal sent from the device chassis and the host chassis, controlling the logic unit in the switch chassis to supply power to the switch based on a first enable signal; the power-on signal is generated based on the power-on instruction; Controlling the baseboard controller in the switch chassis to send the power-on signal to the baseboard controller in the device chassis and the baseboard controller in the host chassis to supply power to the target device and the target computing unit.
6. The method according to claim 5, characterized in that The power-on signal sent to the logic unit in the device chassis is used for the logic unit in the device chassis to supply power to the target device based on a second enable signal; The power-on signal sent to the logic unit in the host chassis is used for the logic unit in the host chassis to supply power to the target computing unit based on a third enable signal.
7. The method according to claim 5, characterized in that, The power-on success signal includes a first power-on success signal; the method includes: Controlling the power supply unit of the device chassis to supply power to each component in the device chassis based on a standby voltage; After each component in the device chassis receives the standby voltage, it generates a power-on success signal and sends it to the logic unit and the baseboard controller of the device chassis; Control the baseboard controller of the device chassis to send the first power-on success signal to the baseboard controller in the switch chassis.
8. The method according to claim 5, wherein The power-on success signal includes a second power-on success signal; the method further includes: Control the power supply unit of the host chassis to supply power to each component in the host chassis based on the standby voltage; After each component in the host chassis receives the standby voltage, it generates a power-on success signal and sends it to the logic unit and the baseboard controller of the host chassis; Control the baseboard controller of the host chassis to send the second power-on success signal to the baseboard controller in the switch chassis.
9. The method according to claim 5, wherein The method further includes: Control the baseboard controller in the host chassis to scan the first interface corresponding to the host chassis and the switch chassis, and obtain the first topology map corresponding to the host chassis and the switch chassis; Control the baseboard controller in the switch chassis to scan the second interface corresponding to the device chassis and the switch chassis, and obtain the second topology map corresponding to the switch chassis and the device chassis.
10. The method according to claim 9, wherein The resource scheduling request includes a resource release request; based on the resource scheduling request, performing resource scheduling on the target resources in the multiple resource pools includes: In the case of receiving the resource release request, remove the device to be adjusted corresponding to the resource to be adjusted indicated by the resource release request from the second topology map and determine the device information corresponding to the device to be adjusted; the target resources in the multiple resource pools include the resource to be adjusted; Based on the device information, reset the device to be adjusted.
11. The method according to claim 10, characterized in that, The resetting of the device to be adjusted includes: Based on the baseboard controller in the switch chassis, send a first reset signal to the logic unit in the switch chassis; Through the logic unit in the switch chassis, send the first reset signal to the logic unit in the device chassis corresponding to the device to be adjusted through the target interface; the logic unit is used to forward the first reset signal to the device to be adjusted to implement the reset.
12. The method according to claim 10, wherein The resource scheduling request further includes a resource acquisition request; the method further includes: Based on the resource acquisition request, allocate the resource to be adjusted to the specified computing unit indicated by the resource acquisition request.
13. The method according to claim 1, characterized in that, The reset instruction includes a system reset instruction, and the device to be reset includes the switch, the target computing unit, and the target device; In the case of receiving the reset instruction, performing a reset operation on the device to be reset indicated by the reset instruction includes: Based on the system reset instruction, the target computing unit generates a system reset signal; the target computing unit implements a reset operation on the target computing unit based on the system reset signal; Control the target computing unit to send the system reset signal to the logic unit in the switch chassis and the logic unit in the device chassis to implement a reset operation on the switch and the target device.
14. The method according to claim 13, wherein Controlling the target computing unit to send the system reset signal to the logic unit in the switch chassis and the logic unit of the device chassis includes: Controlling the target computing unit to send the system reset signal to the baseboard controller in the switch chassis via the baseboard controller in the host chassis; Controlling the baseboard controller in the switch chassis to send the system reset signal to the logic unit and the switch in the switch chassis to perform a reset operation on the switch; Controlling the baseboard controller in the switch chassis to send the system reset signal from the target interface to the logic unit of the device chassis; the system reset signal is used for the logic unit of the device chassis to perform a reset operation based on the system reset signal.
15. The method according to claim 1, characterized in that, The reset instruction includes a device reset instruction, and the device to be reset includes a target reset device; when receiving the reset instruction, performing a reset operation on the device to be reset indicated by the reset instruction includes: Based on the device reset instruction, generating a device reset signal by the baseboard controller in the switch chassis and sending the device reset signal to the logic unit in the switch chassis; Controlling the logic unit in the switch chassis to send the device reset signal from the target interface to the logic unit in the target device chassis corresponding to the device reset instruction to perform a reset operation on the target reset device indicated by the device reset instruction.
16. The method according to claim 1, characterized in that, The distributed resource management system further includes a general management controller; the method further includes: Obtaining the device asset information and interface connection status information corresponding to the target device in the distributed resource management system through the general management controller.
17. A distributed resource management device, characterized in that, Applied to a distributed resource management system deployed in a server, the distributed resource management system includes a switch and multiple resource pools, and the multiple resource pools are obtained by the switch connecting to the first resource corresponding to the target device in the server or the second resource corresponding to the target computing unit in the server based on a high-speed serial cache coherence bus; the device includes: A first control module, configured to control the switch, the target device, and the target computing unit to power on synchronously when receiving a power-on instruction; A first reset module, configured to perform a reset operation on the device to be reset indicated by the reset instruction when receiving the reset instruction; the device to be reset includes at least one of the target device, the target computing unit, and the switch; A first scheduling module, configured to perform resource scheduling on the target resources in the multiple resource pools based on a resource scheduling request; the resource scheduling includes resource reset and resource allocation. The resource scheduling includes resource reset and resource allocation.
18. A distributed resource management system, characterized in that, The distributed resource management system is used to execute the distributed resource management method according to any one of claims 1-16.
19. An electronic device, characterized in that, Including: A processor, a memory, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the distributed resource management method according to any one of claims 1-16.
20. A non-volatile readable storage medium, characterized in that, When the instructions in the non-volatile readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute one or more of the distributed resource management methods recited in claims 1-16.
Citation Information
Patent Citations
Fault memory bank processing method and device, electronic equipment and storage medium
CN116483613A
IO expansion architecture, IO switch and PCIe device
CN117041184A
Equipment reset method and device, storage medium and electronic equipment
CN117251039A
Distributed resource management method, device, system and equipment and storage medium
CN117472596A
Methods and apparatus to manage workload domains in virtual server racks
US20180157532A1
Cited By
MCU-based IPMB implementation system
CN120578612A