Resource pool and reset control method thereof, and timing control method of main system
Through resource pool design and management of the in-band and out-band management chip integrated network management of the board, the problem that traditional PCIe architecture cannot integrate external networks and internal networks is solved, and efficient PCIe resource utilization and system security improvement are achieved.
Patent Information
- Application Number
- CN202510897108.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-06-30
AI Technical Summary
Traditional PCIe architectures cannot integrate external and internal networks together, resulting in insufficient PCIe resources, affecting the bandwidth requirements of multiple networks, and increasing hardware costs.
The resource pool design is adopted, including switching boards and management boards, and network management is integrated through in-band and out-of-band management chips, communication bandwidth and equipment number are dynamically adjusted, and the switching boards are used to connect the coprocessor and the main resource pool, and management is combined with management modules.
Hardware isolation for in-band and out-band network management is realized, system security is improved, communication resources are dynamically adjusted, and PCIe resource utilization and system stability are improved.
Smart Images

Figure CN120406701B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of PCIe topology technology, and in particular to a resource pool and a reset control method thereof, and a timing control method of a main system. Background Art
[0002] In the traditional PCIe (peripheral component interconnect express) architecture, each CPU (central processing unit) in a single computer is connected to a switch (SW) via a fixed PCIe link. Each SW is then connected to each GPU via a fixed PCIe link, with each GPU having a dedicated fixed channel bandwidth (e.g., x16). Multi-computer networking requires adding an InfiniBand network adapter (used in high-performance computing and data centers) or an RDMA network adapter (remote direct memory access) to a single computer, combined with a switch, to achieve multi-computer networking.
[0003] In related technologies, the PCIe architecture can only manage the internal network and cannot integrate the external network and the internal network. Summary of the Invention
[0004] The present application provides a resource pool and a reset control method thereof, and a timing control method for a main system, to at least solve the problem of how to integrate out-of-band network management and in-band network management to achieve hardware isolation of in-band and out-of-band network management.
[0005] The present application provides a resource pool, including: at least one switch board and a management board, the management board including: a management module, an in-band management chip and an out-of-band management chip, wherein:
[0006] A switch card is connected to at least one coprocessor resource pool and a main resource pool, and is used to implement data interaction between the coprocessor resource pool and the main resource pool, and to connect to an external network based on an external network interface;
[0007] In-band management chip, connected to the switch board;
[0008] Out-of-band management chip, connected to the switch board;
[0009] The management module is connected to the in-band management chip and the out-of-band management chip respectively, and is used to perform in-band management of the switch board based on the in-band management chip, and is also used to manage data transmitted through the external network interface based on the out-of-band management chip.
[0010] The present application also provides a switching device, which includes the resource pool as described above, including a box, and a first area and a second area are provided in the box along a first direction, wherein:
[0011] The first area is provided with a switching board, and the second area is provided with a management board. The switching board is detachably connected to the management board.
[0012] The present application also provides a timing control method for a main system, wherein the main system includes: a main resource pool, a coprocessor resource pool, and the above resource pool; or the main system includes: a main resource pool, a coprocessor resource pool, and the above switching device; the timing control method for the main system includes:
[0013] After the main resource pool, the coprocessor resource pool, and the resource pool are powered, the baseboard management controller of the coprocessor resource pool and the baseboard management controller of the main resource pool send a preparation completion signal to the baseboard management controller of the resource pool;
[0014] When receiving the power-on signal, the baseboard management controller of the resource pool controls the power-on based on the programmable logic device, and sends the power-on signal to the baseboard management controller of the coprocessor resource pool and the baseboard management controller of the resource pool through the network;
[0015] After receiving the power-on signal and completing the power-on process, the baseboard management controller of the coprocessor resource pool and the baseboard management controller of the resource pool respectively feed back a power-on completion signal to the baseboard management controller of the resource pool;
[0016] Upon receiving the shutdown signal, the resource pool controls the shutdown and outputs a relation signal to the baseboard management controller of the coprocessor resource pool and the baseboard management controller of the resource pool respectively;
[0017] The baseboard management controller of the coprocessor resource pool and the baseboard management controller of the resource pool receive the shutdown signal and shut down respectively.
[0018] The present application also provides a reset timing control method for a main system, wherein the main system includes: a main resource pool, a coprocessor resource pool, and the above resource pool; or the main system includes: a main resource pool, a coprocessor resource pool, and the above switching device; the reset timing control method for the main system includes:
[0019] After the main system is powered on, the resource pool management module scans the device identifiers of the coprocessor resource pool and the main resource pool and confirms the logical topology connection;
[0020] When the main resource pool receives the reset signal, the programmable logic device of the main resource pool sends a reset signal to the switch card of the resource pool;
[0021] After the switch card of the resource pool receives the reset signal and resets, the switch card of the resource pool sends a reset signal to the programmable logic device of the coprocessor resource pool;
[0022] The programmable logic device of the coprocessor resource pool receives the reset signal and resets.
[0023] The present application also provides a method for controlling the reset timing of a resource pool, which is applied to the above resource pool, or the above switching device; the method for controlling the reset timing of a resource pool includes:
[0024] The management module of the control resource pool sends a reset instruction to the programmable logic device of the resource pool;
[0025] The programmable logic device of the control resource pool resets the switching chip in the switching board corresponding to the reset instruction, or sends a reset instruction to the coprocessor resource pool corresponding to the reset instruction.
[0026] Through the present application, since at least one switch board and a management board are integrated into a resource pool, the management board manages at least one switch board, and is connected to the coprocessor resource pool and the main resource pool respectively through the switch board, thereby providing a variety of coprocessor resource pool resources for the upstream main resources, and dynamically adjusting the communication bandwidth and the number of communication devices between the upstream and downstream. At the same time, the management module performs in-band management on at least one switch board through the in-band management chip, and manages the data transmitted by the external network interface set on the switch board through the out-of-band management chip. Thus, the out-of-band network management and in-band network management are integrated to realize the hardware isolation of in-band and out-of-band network management, which greatly improves the security of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0028] Figure 1 A structural diagram of a resource pool provided in an embodiment of the present application;
[0029] Figure 2 A structural diagram of another resource pool provided in an embodiment of the present application;
[0030] Figure 3 An application diagram of a resource pool provided in an embodiment of the present application;
[0031] Figure 4 A structural diagram of another resource pool provided in an embodiment of the present application;
[0032] Figure 5 A structural diagram of another resource pool provided in an embodiment of the present application;
[0033] Figure 6 A topological diagram of a management board in a resource pool provided in an embodiment of the present application;
[0034] Figure 7 A topological diagram of a management module in a resource pool provided in an embodiment of the present application;
[0035] Figure 8 A diagram of the topology of a switching chip in a resource pool provided in an embodiment of the present application;
[0036] Figure 9 A physical structure diagram of a switch board in a resource pool provided in an embodiment of the present application;
[0037] Figure 10 A physical structure diagram of a management board in a resource pool provided in an embodiment of the present application;
[0038] Figure 11 A physical diagram of the front and back structures of a resource pool provided in an embodiment of the present application;
[0039] Figure 12 A physical diagram of a resource pool structure provided in an embodiment of the present application;
[0040] Figure 13 A physical diagram of another resource pool structure provided in an embodiment of the present application;
[0041] Figure 14 A flow chart of a timing control method for a main system provided in an embodiment of the present application;
[0042] Figure 15 This is a flow chart of a reset timing control method for a main system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0044] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0045] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0046] In the traditional PCIe (peripheral component interconnect express) architecture, each CPU (central processing unit) in a single computer is connected to a switch (SW) via a fixed PCIe link. Each SW is then connected to each GPU via a fixed PCIe link, with each GPU having a dedicated fixed channel bandwidth (e.g., x16). Multi-computer networking requires adding an InfiniBand network adapter (used in high-performance computing and data centers) or an RDMA network adapter (remote direct memory access) to a single computer, combined with a switch, to achieve multi-computer networking.
[0047] In related technologies, the PCIe architecture can only manage internal networks and cannot integrate external and internal networks. Furthermore, when networking multiple machines, each machine must have its own PCIe resource for the IB network or RDMA network card. To meet the bandwidth requirements of the multi-machine network, these network cards require nearly 50% of the PCIe resources available for expansion on a single machine, resulting in a shortage of PCIe resources.
[0048] The rapid development of complex computing scenarios such as artificial intelligence, machine learning, and high-performance computing has placed new demands on data center architectures. Traditional computing architectures face three core contradictions in large-scale AI model training: the conflict between isolated hardware resources and scalable computing needs, the conflict between fixed bandwidth allocation and dynamic load fluctuations, and the conflict between physical device scalability limitations and the elasticity requirements of computing clusters. In GPU (Graphics Processing Unit) training and large-scale AI (Artificial Intelligence) model applications, PCIe resource pools have become a key solution. PCIe resource pools use virtualization technology to split physical channels into logical resource pools and, combined with intelligent scheduling algorithms (such as QoS (Quality of Service)-based prioritization), enable on-demand bandwidth allocation. PCIe tree topologies inherently limit the number of IDs (identity documents) and the depth of the hierarchy (traditional architectures only support 256 device IDs). However, large-scale AI model training often requires computing clusters composed of hundreds of GPUs. Resource pools, by incorporating the latest PCIe technology and switch chips, restructure the physical topology into a mesh structure, significantly increasing the upper limit of GPU cluster size. Therefore, this paper proposes a PCIe resource pool architecture design scheme.
[0049] In the traditional PCIe architecture, each CPU (Central Processing Unit) in a single machine is connected to the switch via a fixed PCIe link, and each switch is connected to each GPU via a fixed PCIe link. Each GPU has a dedicated fixed channel bandwidth (e.g., x16). Multi-machine networking requires adding an InfiniBand (IB) or RDMA (remote direct memory access) network card to a single machine, along with a switch, to achieve multi-machine networking. Each machine in a multi-machine network requires dedicated PCIe resources for the IB or RDMA network card. To meet the bandwidth requirements of the multi-machine network, these network cards consume nearly 50% of the PCIe resources available for expansion on a single machine. Furthermore, the need to add network cards, optical modules, and switches to the system further increases the price of the entire machine.
[0050] The embodiment of the present application provides a resource pool, such as Figure 1 As shown, it includes: at least one switching board A and a management board B, and the management board B includes: a management module, an in-band management chip and an out-of-band management chip, wherein,
[0051] Switch board A is connected to at least one coprocessor resource pool and a main resource pool, and is used to implement data exchange between the coprocessor resource pool and the main resource pool, and to connect to the external network based on the external network interface;
[0052] Specifically, the coprocessor resource pool can be a device, specifically a GPU (Graphics Processing Unit), SSD (Solid State Disk or Solid State Drive), FPGA (Field Programmable Gate Array), network card, memory, and other resources. The main resource pool is specifically the host resource pool. The interface exposed by switch board A connects to at least one coprocessor resource pool and the main resource pool. The external network interface is specifically an RJ45, a type of information outlet (i.e., communication outlet) connector in a cabling system. The external network interface can specifically connect to external network switching equipment or fans, for example. Multiple switch boards A are connected via PCIe links.
[0053] In-band management chip, connected to the switch board;
[0054] Out-of-band management chip, connected to the switch board;
[0055] Specifically, the specific models of the in-band management chip and the out-of-band management chip are 88E6190X.
[0056] The management module is connected to the in-band management chip and the out-of-band management chip respectively, and is used to perform in-band management of the switch board based on the in-band management chip, and is also used to manage data transmitted through the external network interface based on the out-of-band management chip.
[0057] Specifically, the management module uses an in-band management chip to manage the switch cards in-band. It also uses an out-of-band management chip and an extended network management chip to manage data transmitted through the external network interface. The external network interface allows access to the management module and any device. Simultaneously, the switch cards access the management module through the out-of-band management chip, serving as a shared management network port for the management module. This fusion of out-of-band and in-band network management achieves hardware isolation between in-band and out-of-band network management, significantly improving system security.
[0058] Specifically, since at least one switching board A and a management board B are integrated into a resource pool, the management board B manages at least one switching board A, and is connected to the coprocessor resource pool and the main resource pool through the switching board A respectively, thereby providing a variety of coprocessor resource pool resources for the upstream main resources, and dynamically adjusting the communication bandwidth and the number of communication devices between the upstream and downstream.
[0059] It should be noted that during on-site maintenance, the network port of the out-of-band management chip on the right can be used when necessary to improve the convenience of operation and maintenance.
[0060] Through the present application, since at least one switch board A and management board B are integrated into a resource pool, management board B manages at least one switch board A and is connected to the coprocessor resource pool and the main resource pool through switch board A, thereby providing a variety of coprocessor resource pool resources for the upstream main resources, and dynamically adjusting the communication bandwidth and number of communication devices between the upstream and downstream. At the same time, the management module performs in-band management on at least one switch board A through the in-band management chip, and manages the data transmitted by the external network interface set on the switch board A through the out-of-band management chip. Thus, the out-of-band network management and in-band network management are integrated to achieve hardware isolation of in-band and out-of-band network management, greatly improving the security of the system.
[0061] In some optional embodiments, reference Figure 2 , the switch board includes: a first channel interface, wherein,
[0062] A first end of the first channel interface is connected to the in-band management chip, and a second end of the first channel interface is selectively connected to the in-band management chip or the out-of-band management chip.
[0063] Specifically, the channel interface is MCIO (Mini Cool Edge IO, a multi-channel input / output technology) to achieve PCIe x16 bandwidth. Each switch board includes at least one first channel interface.
[0064] refer to Figure 2 Taking four switch boards A as an example, the first end of the first channel interface of each switch board A is connected to the in-band management chip, specifically to pins P2, P3, P4, P5, P6, P7, P9, and P10 of the in-band management chip. The second end of the first channel interface of each switch board A is connected to the out-of-band management chip, specifically to pins P5, P6, P7, and P10 of the out-of-band management chip. For example, the second end of the first channel interface of switch board A can be selectively connected to pins P10 of the in-band management chip (indicated by the dotted line) or P0 of the in-band management chip (indicated by the solid line).
[0065] In some optional embodiments, reference Figure 2 , the management board also includes: a first network management chip, wherein,
[0066] The first network management chip is connected to the second end of the first channel interface and the out-of-band management chip respectively, and is used to enable the external network interface to communicate directly with the in-band network.
[0067] Specifically, the model of the network management chip may be 88E1512. The external network interface is directly connected to the in-band management chip via the first channel interface and the first network management chip, thereby directly establishing a connection between the external network interface and the in-band network.
[0068] In some optional implementations, the management module, the in-band management chip, and the out-of-band management chip are provided on the first baseboard. Figure 2 The management module includes: at least one management module module, a second network management chip and a baseboard management controller, wherein,
[0069] At least one management module, connected to the in-band management chip, for controlling the in-band management chip;
[0070] a second network management chip connected to the baseboard management controller and the out-of-band management chip respectively;
[0071] The baseboard management controller is connected to the second network management chip and is used to manage the external network interface based on the out-of-band management chip.
[0072] Specifically, the management module is a CPU, and the management module is connected to the switch board A based on an in-band management chip, and controls the in-band management chip and the switch board A.
[0073] For example, Figure 2 As shown, there can be two management modules, and the management modules are connected to the in-band management chip at the same time. The information of the two management modules can be synchronized through the network, which speeds up the transmission of redundant information and improves the stability of the system.
[0074] Specifically, the baseboard management controller is a BMC, which is connected to the in-band management chip via the second network management chip, thereby managing the external network interface, specifically transmitting data to the external network interface, or receiving data transmitted from the external network.
[0075] It should be noted that the management module can specifically adopt a low-power system management module. To achieve efficient management of the switch chip, the management board is designed with two independent low-power system management modules. These two management modules are attached to the management board in the form of plug-in cards. When the system is turned on and running, they serve as redundant backups for each other. By switching the signals (USB, network, I2C) of the management switch chip, uninterrupted management of the switch card can be achieved, providing stability for the PCIe resource pool. In order to achieve redundant monitoring, by switching the communication signals of the baseboard management controller and the user interface signals, users can achieve uninterrupted management and always understand the operating status of the system.
[0076] In some optional embodiments, reference Figure 2 The switch board includes: a second baseboard (not shown), a plurality of switch chips, a plurality of third network management chips, a plurality of fourth network management chips and a second channel interface, wherein,
[0077] The second channel interface is used to connect the switch chip in the switch board to achieve communication between the switch chips in the switch board;
[0078] Multiple switching chips are integrated on the second baseboard, the multiple switching chips are communicatively connected to multiple switching chips on other switching boards via first channel interfaces, and are respectively connected to at least one coprocessor resource pool and a main resource pool;
[0079] Multiple switching chips for implementing data exchange between the coprocessor resource pool and the main resource pool;
[0080] a plurality of third network management chips, each of which is connected between the external network interface and the first channel interface, and is used to enable the external network interface to communicate with the first channel interface;
[0081] A plurality of fourth network management chips are each connected between the first channel interface and the switching chip, and are used to enable the switching chip to communicate with the first channel interface.
[0082] Specifically, all switch chips are connected through the first and second channel interfaces, creating a meshed PCIe topology that works with the CPUs in the primary resource pool. Multipath interconnection achieves redundancy, allowing multiple physical links between any two nodes. If a PCIe link fails, data automatically switches to other available paths, preventing a single point of failure from paralyzing the entire system. Each PCIe endpoint (such as a GPU) can establish independent channels with multiple targets simultaneously, significantly improving effective bandwidth utilization. Modularity and dual-management redundancy are also integrated into the system, providing robust management of the PCIe resource pool.
[0083] Specifically, the switching chip is a PCIe SW. A PCIe Switch is a hardware device that provides expansion or aggregation capabilities and allows more devices to be connected to a PCIe port. Multiple switching chips within the same switching board are connected via the second channel interface, and switching chips on different switching boards A are connected via the first channel interface. The number of first channel interfaces corresponds to the number of switching chips within switching board A, and the number of second channel interfaces corresponds to the number of switching boards A. In addition, the first channel interface and the second channel interface can be reserved to achieve changes in the SW topology, not just the interconnection of two SWs using PCIe x16 bandwidth. Furthermore, by changing the connection method of the supplementary MCIO cable, different PCIe resource pool topologies can be constructed within a single machine.
[0084] For example, when there are two switching chips in a switching board A, and when there are four switching boards A, each switching chip includes a second channel interface for connecting to another switching chip in the same switching board, and also includes three second channel interfaces for connecting to three different switching boards A to realize the interconnection of 4-layer switching boards.
[0085] Specifically, the baseboard management controller manages the external network interface based on the out-of-band management chip, the first channel interface, and the third network management chip. Simultaneously, at least one management module manages the coprocessor resource pool through the in-band management chip, the first channel interface, and the fourth network management chip.
[0086] In some optional embodiments, reference Figure 4 , the switch board includes: a plurality of first selection paths, wherein,
[0087] A plurality of first selection paths, each of which is respectively connected to the second ends of the plurality of first channel interfaces and the in-band management chip, are used for the in-band management chip to communicate with a plurality of switch boards at the same time.
[0088] Specifically, the selection path is a MUX (ie, a multiplexer), and the second end of each first channel interface can be connected to the in-band management chip through the first selection path.
[0089] In some optional embodiments, reference Figure 2 and Figure 3 , the baseboard management controller is used to configure the in-band management chip, the out-of-band management chip, the second network management chip and each third network management chip; or,
[0090] The in-band management chip is used to configure the first network management chip and configure each third network management chip based on the first selection path; or,
[0091] Each switching chip is configured for the connected fourth network management chip.
[0092] Specifically, the baseboard management controller performs initial configuration on the in-band management chip, the out-of-band management chip, the second network management chip, and each third network management chip. The in-band management chip may also perform initial configuration on the first network management chip and each third network management chip. Each switch chip may also perform initial configuration on the connected fourth network management chip.
[0093] In some optional implementations, when the resource pool includes multiple switch boards A, the switch boards A are stacked.
[0094] Specifically, the switch boards A are stacked to save space.
[0095] In some optional implementations, the switch card further includes: a plurality of data center front-end interfaces, wherein:
[0096] Multiple data center front-end interfaces are connected to a switching chip and to a coprocessor resource pool or a main resource pool, so as to enable the switching chip to communicate with the coprocessor resource pool and the main resource pool respectively.
[0097] Specifically, the data center front-end interface is a CDFP (Data Center Front Panel) connector. The CDFP connector is an external X16 PCIe connector. Each layer of switch board A has multiple data center front-end interfaces.
[0098] Specifically, the resource pool contains four identical layers of switch boards A, among which the single-layer switch board A integrates two high-performance switch chips, 10 CDFP connectors, 8 X16 PCIe MCIO connectors (MCIO uses MCIO cables inside the chassis to interconnect the eight switch chips), 1 external management network port, a board power connector, and 1 board management connector.
[0099] By stacking four switch cards A, you can arrange eight switch chips within a 2U space and still provide 40 external PCIe ports. After stacking switch cards A, the eight switch chips are connected via MCIO cables within the chassis. Different connection methods can be used to construct different resource pool topologies, including but not limited to one 2x1 plus 2x3, two 2x2s, one 2x4, or all eight switch chips interconnected. This supports resource pools of varying sizes, enabling hardware isolation of PCIe resource pools and achieving higher resource pool security.
[0100] In some optional implementations, the management board further includes: a second selection path, the first baseboard includes: a baseboard interface,
[0101] A second selection path is connected to the baseboard management controller and the baseboard interface respectively;
[0102] The backplane interface is connected to the out-of-band management chip and at least one management module respectively, and is used to extend the signal path so that the backplane management controller is connected to the out-of-band management chip and at least one management module respectively for backup communication.
[0103] Specifically, the baseboard management controller is connected to the second selection path based on the NCSI (Network Controller – Sideband Interface technology), the second selection path is connected to the baseboard interface based on the NCSI communication method, the baseboard interface is connected to at least one management module module based on the PCIe method, and the baseboard interface is connected to the in-band management chip based on the MDI (Ethernet port using twisted pair cable) method.
[0104] In some optional implementations, the management board further includes: a high-speed connector, wherein:
[0105] High-speed connector, connected to the switch board, used to connect to external management devices or expansion interfaces.
[0106] Specifically, the high-speed connector is a Slimline connector. The management board controls the switching board through the Slimline connector, which includes the error interrupt signal INT of the switching chip, the out-of-band management signal I2C, the network management signal SERDES (short for SERializer (serializer) / DESerializer (deserializer)) and the power-on control signals PWREN and PWRGD, as well as the external SERDES signal from the Slimline connector, which is converted into an MDI signal through the network management chip to provide an external network management interface for the management board.
[0107] In some optional implementations, the management board further includes: a programmable logic device, wherein:
[0108] The programmable logic device is connected to the baseboard management controller and is used to control the power-on and power-off timing and reset timing, and send the power-on status and reset status of the entire machine to the baseboard management controller.
[0109] Specifically, the programmable logic device is a CPLD (ie, complex programmable logic device), which is used to control the power-off timing and reset timing of the resource pool, and send the power-on status and reset status of the resource pool to the baseboard management controller.
[0110] Specifically, the CPLD monitors the operating status of the low-power system management module through the GPIO of the low-power system management module. In addition to monitoring the operating status of the low-power system management module through the GPIO of the low-power system management module, the BMC chip also monitors the system information of the low-power system management module through the UART serial port. The two low-power system management modules also monitor each other's operating status through GPIO. Monitoring the operating status of the low-power system management module through multiple links can achieve rapid monitoring and rapid switching after failures. The BMC supports monitoring and management of out-of-band information of the entire machine, monitoring PSU status information, Power VR status information, and overall temperature information, and implementing fan control and system display.
[0111] In some optional implementations, the third network management chip and the fourth network management chip have the same address to facilitate management.
[0112] Specifically, to facilitate maintenance, the third and fourth network management chips on a Layer 4 switch card are assigned the same address, 0x00. Each high-performance switch chip acts as the network management chip master for configuration and management. Each of the four network ports on a Layer 4 switch card is configured by an in-band management chip, an out-of-band management chip, and a baseboard management (BMC). The in-band management chip requires two identically addressed network management chips, the first and third, necessitating the addition of a multiplexing (MUX) for time-sharing management. After configuring the first network management chip, with address 0x00, the MUX must be switched, and the next network management chip, with address 0x00, must be configured via MDIO. As the primary master, the BMC can configure the in-band management chip, the out-of-band management chip, the third network management chip corresponding to a network port, and the BMC's own second network management chip. Their slave addresses are 0x02, 0x03, 0x00, and 0x01, respectively. Based on this MDIO topology design, initial network configuration management can be implemented within the PCIe resource pool, enabling faster in-band management of the network for the switch chip and out-of-band management of the entire machine after startup.
[0113] refer to Figure 8As an example, a single switch card supports up to two Broadcom PEX89144 series switch chips. Each switch chip has nine PCIe GEN5 x16 lanes, with a single full-duplex speed of up to 128 GB / s. The PCIe resources of each switch chip on the second backplane are connected to five x16 CDFP connectors and four x16 MCIO connectors. The dense layout of the ten CDFP connectors on the front panel allows the PCIe resources of both switches to be fully accessible. Two MCIOs are located next to the switch chips on the switch card, interconnecting the switch chips within the card. Reserving these MCIOs allows for flexible switch chip topologies, extending beyond the PCIe x16 bandwidth connection between two switch chips. The card also supports six PCIe x16 MCIOs for Layer 4 switch card interconnection. By adjusting the rear MCIO cable connections, different PCIe resource pool topologies can be constructed within a single machine.
[0114] refer to Figure 8 In addition to PCIe signals, the CDFP connector also carries 100M CLK signals from the clock buffer, I2C signals from the management board, and reset signals for the management board's low-power system management module, baseboard management controller, and PCA9555 controlled by the switch chip. The 100M CLK provides downstream devices with a clock source identical to the switch chip, improving system stability. The management board's I2C signals enable automatic detection of upstream and downstream devices and out-of-band information acquisition from downstream devices. Through reset signal linkage, the low-power system management module, baseboard management controller, and switch chip can receive reset signals from the upstream host and, based on internal policy control, reset downstream devices controlled by the low-power system management module, baseboard management controller, and switch chip.
[0115] To maximize PCB (baseboard) utilization and minimize PCB waste, the management board layout utilizes a cable riser for the AIC card (graphics card) of the low-power system management module, reducing PCB waste in the fan area. To accommodate the size of the rear window, the PSU area is hollowed out, resulting in an overall rectangular design. Furthermore, to dissipate heat from the low-power system management module and the switching chip, the fan module faces the low-power system management module and the switching chip. To power and manage the switching board, the power connector and signal connector are placed in the upper left corner of the board. To reduce wiring from the low-power system management module to the baseboard management controller, the baseboard management controller is placed between the two low-power system management modules, reducing the PCIe link length of the AIC card. The AIC card is placed to the side of the low-power system management module.
[0116] refer to Figure 3 , can connect to multiple main resource pools, such as H1 to Hn, and can connect to NVMES, GPU, FPGA, NIC, etc. Figure 6 and Figure 7 , Figure 7 This is a topological connection diagram of the management module. COM Express, the form factor for a computer-on-module (COM), is a highly integrated and compact PC that can be used in design applications like integrated circuit components. Each COM Express module integrates core CPU and memory functions, general-purpose I / O, USB, audio, graphics (PEG), and Ethernet. The low-power system management module (SMM) uses an Intel ATOM C3000 processor with a base frequency of 2.4GHz. Measuring 95mm by 125mm, it is primarily responsible for data processing, memory control, and terminal processing, and is the core module of the server system. The system supports the PEX8780 high-performance switch chip and peripheral I / O device expansion. It uses two 64GB SoDIMMs (SoDIMMs) for memory and a low-power System Management Module (SMM) Type 7 standard interface (USB interface standard). It supports two NVME (NVM Express) and SATA flash memory modules. It also supports four USB 3.0 & 2.0 interfaces, VGA (video transmission standard) and Gigabit Ethernet (GE, 1000Mbps Ethernet). A PCI Express x8 port is reserved for expansion, including AIC. It also has 20 HSIO ports shared by PCIe, SATA, and USB 3.0. It also supports eMMC (Embedded Multi Media Card) 5.0 and eMMC 4.5. The system uses a low-power Intel CPU for low-power management. Its modular design, integrating the Intel CPU into a single module, facilitates platform upgrades and switching. Moreover, when in use, if the requirements for redundant management are not high, a low-power system management module can be used in conjunction with a baseboard management controller to reduce the cost of the entire machine.
[0117] The embodiment of the present application provides a switching device, such as Figure 9 As shown, the switching device includes the resource pool as above, including a box, along a first direction, a first area and a second area are set in the box, wherein,
[0118] The first area is provided with a switching board, and the second area is provided with a management board. The switching board is detachably connected to the management board.
[0119] refer to Figure 9 and Figure 10 , Figure 9 For multi-layer switch boards, the first area is equipped with at least one tray. Inside the chassis, at least one set of slides is provided. Each tray is retractable and mounted on each set of slides. Each tray contains at least one layer of switch chip boards. A locking hook structure is provided inside the chassis to lock the trays in place. To facilitate maintenance of the switch boards and prevent the four layers of switch boards from needing to be disassembled and assembled sequentially before removal, the four layers of switch boards are arranged on two trays. Each tray can be individually pulled out using handles on both sides. During installation, once the trays are in place, the locking mechanism is engaged to secure the trays to the chassis, allowing cable assembly to proceed. Figure 10 For the management board. The lock hook structure can refer to Figure 13 .
[0120] In some optional implementations, the management module includes: a first management module module, a second management module module and a baseboard management controller;
[0121] The first management module module and the second management module module are spaced apart, and the baseboard management controller is arranged between the first management module module and the second management module module; the first management module module and the second management module module face opposite directions.
[0122] Specifically, refer to Figure 10 Because the two management modules have 1+1 redundancy and share a single baseboard management controller (BMC), the BMC is placed between the two low-power system management modules (LPSMs) to minimize routing distances for the PCIe, USB (Universal Serial Bus), and LPC (Linear Predictive Coding) signals from the two management modules to the BMC. The two LPSMs are symmetrically positioned. This facilitates signal switching between the two LPSMs. Specifically, by placing the BMC between the first and second management modules, with the first and second modules facing opposite directions, routing the BMC to the two management modules is minimized, significantly reducing routing distance.
[0123] refer to Figure 12 , Figure 12 For the structure of the PCIe resource pool, in order to occupy fewer Us within the 42U space of the cabinet to arrange the PCIe resource pool and support a larger number of interfaces, the size of the entire machine is defined as 2U height to maximize support for 40 CDFP interfaces, realizing a high-density PCIe resource pool.
[0124] refer to Figure 11The management board also supports power input and distribution for the entire system, cooling control for the entire system, and internal or external resource expansion. It supports 1+1 PSU (power supply) input (compatible with HVDC PSU (power supply) input), where the PSUs are stacked one on top of the other and blindly plugged into the management board via the PSU backplane. If an HVDC PSU is used, since the HVDC PSU does not have a built-in fan, a fan connector is reserved on the management board to dissipate heat specifically for the HVDC PSU via the chassis' built-in fan. It supports N+1 rotor-level fan cooling control, with fans blindly plugged directly into the management board.
[0125] refer to Figure 5 , supporting resource expansion for the low-power system management module, including support for one PCIe 3.0 x8 network card, one PCIe 3.0 x2 NVMe M.2 hard drive or SATA 3.0 M.2 hard drive (two hard drives connected to the management board via an adapter board), and one PCIe 3.0 x4 debug port (for software). The 1+1 PSU redundancy design improves system redundancy and stability, meeting system power requirements while reducing power costs.
[0126] refer to Figure 11 , Figure 11 This shows the front and back windows of the resource pool. The left and right mounting ears are the user interface modules for the system, including a power button, indicator light, user identifier (UID), Type-C debug port (standard USB form factor), VGA (Video Graphics Array), and a USB port for the low-power system management module. In the center of the front window are four layers of switch cards, each layer containing a system management network port and 10 CDFP ports. The rear panel contains four 6056 fans, two half-height AIC (Nvidia graphics card manufacturers call them Add-In Cards) network cards, and two 1+1 redundant PSUs.
[0127] The embodiment of the present application provides a timing control method of a main system, such as Figure 13 As shown, the main system includes: a main resource pool, a coprocessor resource pool and the above resource pool; or the main system includes: a main resource pool, a coprocessor resource pool and the above switching device; the timing control method of the main system includes:
[0128] Step 1: After the main resource pool, the coprocessor resource pool, and the resource pool are powered on, the baseboard management controller of the coprocessor resource pool and the baseboard management controller of the main resource pool send a preparation completion signal to the baseboard management controller of the resource pool;
[0129] Step 2: upon receiving the power-on signal, the baseboard management controller of the resource pool controls the power-on based on the programmable logic device, and sends the power-on signal to the baseboard management controller of the coprocessor resource pool and the baseboard management controller of the resource pool via the network;
[0130] Step 3: After receiving the power-on signal and completing the power-on process, the baseboard management controller of the coprocessor resource pool and the baseboard management controller of the resource pool each send a power-on completion signal back to the baseboard management controller of the resource pool;
[0131] Step 4: upon receiving the shutdown signal, the resource pool controls the shutdown and outputs a relationship signal to the baseboard management controller of the coprocessor resource pool and the baseboard management controller of the resource pool respectively;
[0132] Step 5: The baseboard management controller of the coprocessor resource pool and the baseboard management controller of the resource pool receive the shutdown signal and shut down respectively.
[0133] Specifically, refer to Figure 14 The main resource pool includes: management modules, devices, power-on power conversion chips, standby power conversion chips, programmable logic devices, baseboard managers, and power supplies. The resource pool includes: switching chips, management modules, devices, power-on power conversion chips, standby power conversion chips, programmable logic devices, baseboard managers, and power supplies. The coprocessor resource pool includes: coprocessors, power-on power conversion chips, standby power conversion chips, programmable logic devices, baseboard managers, and power supplies.
[0134] Specifically, the power-on and power-off timing control method of the main system can be carried out according to the following steps:
[0135] 1. After the power supplies in each resource pool are powered, the standby power conversion chip switches out the standby power and sends the last standby power (Power Good, a key signal in computer hardware used to indicate stable power output) to the programmable logic device.
[0136] 2. When the standby power supply in each chassis is complete and the PLD and BMC are functioning normally, the PLDs in the main resource pool and coprocessor resource pool send the standby power-on PG signal to the BMC via I2C (two-wire serial bus) / UART (universal asynchronous receiver / transmitter).
[0137] 3. The baseboard management controller of each resource pool sends a signal indicating that the power-on status has been completed to the baseboard management controller of the PCIe resource pool via the network. Each resource pool waits for the power-on signal from the PCIe resource pool to proceed to the next step.
[0138] 4. After pressing the PCIe resource pool power button, the power-on signal is given to the programmable logic device, which controls the EN (enable) signal of the main power supply.
[0139] 5. At the same time, a power-on signal is sent to the baseboard management controller, which then sends the power-on signal to the baseboard management controllers of each resource pool via the network.
[0140] 6. The baseboard management controller of each resource pool sends a power-on signal to the programmable logic device of each resource pool through I2C / UART. The programmable logic device controls the EN of each VR (voltage regulator) and sends a power-on completion signal to the baseboard management controller and low-power system management module of the PCIe resource pool through the network and the baseboard management controller after receiving the last power PG.
[0141] 7. After the entire system is powered off through the baseboard management controller or by long-pressing the power button, a shutdown signal is sent to the programmable logic device. The programmable logic device controls the output of the main power EN signal, thereby powering off the main power supply of the resource pool.
[0142] 8. The baseboard management controller of the PCIe resource pool sends a shutdown signal to the baseboard management controllers of each resource pool. The baseboard management controllers of each resource pool use programmable logic devices to shut down the main power supply of the resource pool.
[0143] It's important to note that the above-described coordinated power-on and power-off process enables collaboration among multiple resource pools, allowing the upstream resource pool to properly identify downstream PCIe devices in accordance with the PCIe protocol specification. In other words, the resource pools can coordinate the power-on and power-off timing of both the main resource pool and the coprocessor resource pool.
[0144] The embodiment of the present application provides a reset timing control method for a main system, such as Figure 15 As shown, the main system includes: a main resource pool, a coprocessor resource pool and the above resource pool; or the main system includes: a main resource pool, a coprocessor resource pool and the above switching device; the reset timing control method of the main system includes:
[0145] Step 1: After the main system is powered on, the resource pool management module scans the device identifiers of the coprocessor resource pool and the main resource pool, and confirms the logical topology connection;
[0146] Step 2: When the main resource pool receives the reset signal, the programmable logic device of the main resource pool sends a reset signal to the switch card of the resource pool;
[0147] Step 3: After the switch card of the resource pool receives the reset signal and resets, the switch card of the resource pool sends a reset signal to the programmable logic device of the coprocessor resource pool;
[0148] Step 4: The programmable logic device of the coprocessor resource pool receives the reset signal and resets.
[0149] Specifically, the reset timing control method of the main system can be performed according to the following steps:
[0150] 1. After power is applied, each resource pool independently switches to standby mode. After power is applied, the low-power system management module automatically powers on, and the programmable logic device identifies the presence of the switch card and reports this to the low-power system management module.
[0151] 2. The low-power system management module scans the peer coprocessor ID (device identifier) using SMB_HOST (Server Message Block, a network file sharing protocol) to confirm the topology connection of the data center front-end interface.
[0152] 3. After pressing the resource pool power button, the switch chip powers on, identifies its own ID, and reports it to the low-power system management module. The low-power system management module creates a topology table and matches it with the switch chip ports, and configures the switch chip ports as uplink or downlink.
[0153] 4. After the button is pressed, the baseboard management controller notifies each resource pool through the switch to turn on the main power supply.
[0154] 5. After powering on each resource pool, wait for a reset signal. The master resource pool's management module and programmable logic device send a reset signal to each module. The master resource pool's baseboard management controller sends the reset signal to the PCIe resource pool's baseboard management controller via the switch.
[0155] 6. The reset of other devices in the PCIe resource pool, except the switch chip, is triggered by the reset of the low-power system management module. The reset of the switch chip can be forwarded via the two-level data center front-end interface through the upstream reset of the main resource pool, or the baseboard management controller can notify the programmable logic device through I2C, and the programmable logic device sends the reset to the switch chip and the data center front-end interface.
[0156] 7. The coprocessor resource pool is reset by the front-end interface of the data center and sent to the programmable logic device, which is then sent to each device or directly to the GPU or SSD.
[0157] It should be noted that the above-described resource pool coordinated reset process enables collaborative operation between multiple resource pools, allowing the upstream resource pool to properly identify downstream PCIe devices in accordance with the PCIe protocol specification. In other words, the resource pools can coordinate the reset timing of both the main resource pool and the coprocessor resource pool.
[0158] In some optional implementations, the reset timing control method of the main system further includes:
[0159] Step (1): When the main resource pool receives a reset signal, the programmable logic device of the main resource pool sends a reset signal to the programmable logic device of the resource pool;
[0160] Step (2): After the programmable logic device of the resource pool receives the reset signal and resets, the programmable logic device of the resource pool sends a reset signal to the programmable logic device of the coprocessor resource pool, or the programmable logic device of the resource pool sends a reset signal to the switch board of the resource pool;
[0161] Step (3): The programmable logic device of the coprocessor resource pool receives the reset signal and resets.
[0162] Specifically, the reset timing control method of the main system can also be implemented by programmable logic devices of each resource pool. Of course, there is another reset timing control method in which the programmable logic device and the switching chip can be reset together.
[0163] An embodiment of the present application provides a method for controlling the reset timing of a resource pool. The method is applied to the above resource pool, or the method is applied to the above switching device. The method for controlling the reset timing of a resource pool includes:
[0164] Step 1: The management module of the control resource pool sends a reset instruction to the programmable logic device of the resource pool;
[0165] Step 2: The programmable logic device in the control resource pool resets the switch chip in the switch board corresponding to the reset instruction, or sends a reset instruction to the coprocessor resource pool corresponding to the reset instruction.
[0166] Specifically, this resource pool reset sequence control method uses the resource pool's management module to control resets. First, when the management module needs to reset a switch chip or device, it sends a command to the PCIe resource pool's data center front-end interface via I2C / UART. Then, the SW resource pool's data center front-end interface resets the switch chip or sends a reset signal to the coprocessor resource pool's data center front-end interface or GPU / SSD. This enables the resource pool's self-reset process through the management module.
[0167] In some optional implementations, the method for controlling the reset timing of the resource pool further includes:
[0168] Step (1): When the switch chip of the resource pool receives the reset signal of the main resource pool, it sends a reset signal to the coprocessor resource pool;
[0169] Step (2): The switching chip module of the resource pool sends a reset completion signal to the management module of the resource pool.
[0170] Specifically, this resource pool reset timing control method uses the switch chip in the resource pool to control the reset. For example, in certain situations, when the switch chip needs to independently reset a downstream device, after receiving the reset signal from the master resource pool, the switch chip proactively sends a reset signal through the hardware link based on its own upstream and downstream configurations. After sending the reset signal, the switch chip notifies the baseboard management controller and low-power system management module via the INT interrupt pin that the downstream device has been reset.
[0171] The above is a detailed introduction to a resource pool and its reset control method, and a timing control method of the main system provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A resource pool, characterized in that: include: At least one switch board and a management board, wherein the management board includes: a management module, an in-band management chip and an out-of-band management chip, wherein: A switch card is connected to at least one coprocessor resource pool and a main resource pool, and is used to implement data interaction between the coprocessor resource pool and the main resource pool, and to connect to an external network based on an external network interface; An in-band management chip connected to the switch board; An out-of-band management chip connected to the switch board; A management module is connected to the in-band management chip and the out-of-band management chip respectively, and is used to perform in-band management of the switch board based on the in-band management chip, and is also used to manage data transmitted through the external network interface based on the out-of-band management chip.
2. The resource pool according to claim 1, characterized in that: The switch board includes: a first channel interface, wherein: A first end of the first channel interface is connected to the in-band management chip, and a second end of the first channel interface is selectively connected to the in-band management chip or to an out-of-band management chip.
3. The resource pool according to claim 2, characterized in that: The management board also includes: a first network management chip, wherein: The first network management chip is connected to the second end of the first channel interface and the out-of-band management chip respectively, and is used to enable the external network interface to communicate directly with the in-band network.
4. The resource pool according to claim 2, characterized in that: The management module, the in-band management chip and the out-of-band management chip are arranged on the first baseboard. The management module includes: at least one management module module, a second network management chip and a baseboard management controller, wherein: At least one management module, connected to the in-band management chip, for controlling the in-band management chip; A second network management chip is connected to the baseboard management controller and the out-of-band management chip respectively; A baseboard management controller is connected to the second network management chip and is used to manage the external network interface based on the out-of-band management chip.
5. The resource pool according to claim 4, characterized in that: The switch board includes: a second baseboard, a plurality of switch chips, a plurality of third network management chips, a plurality of fourth network management chips and a second channel interface, wherein: The second channel interface is used to connect the switch chip in the switch board to achieve communication between the switch chips in the switch board; a plurality of switch chips integrated on the second baseboard, wherein the plurality of switch chips are communicatively connected to the plurality of switch chips on other switch boards via the first channel interface, and are respectively connected to at least one coprocessor resource pool and a main resource pool; Multiple switching chips for implementing data exchange between the coprocessor resource pool and the main resource pool; a plurality of third network management chips, each of which is connected between the external network interface and the first channel interface, and is used to enable the external network interface to communicate with the first channel interface; A plurality of fourth network management chips are each connected between the first channel interface and the switching chip, and are used to enable the switching chip to communicate with the first channel interface.
6. The resource pool according to claim 5, characterized in that: The switch board includes: a plurality of first selection paths, wherein: A plurality of first selection paths, each of which is respectively connected to the second ends of the plurality of first channel interfaces and the in-band management chip, are used for the in-band management chip to communicate with a plurality of switch boards at the same time.
7. The resource pool according to claim 6, characterized in that: The baseboard management controller is used to configure the in-band management chip, the out-of-band management chip, the second network management chip and each third network management chip; or, The in-band management chip is used to configure the first network management chip, and configure each third network management chip based on the first selection path; or, Each of the switching chips is configured for the connected fourth network management chip.
8. The resource pool according to claim 3, characterized in that: When the resource pool includes multiple switch boards, the switch boards are stacked.
9. The resource pool according to claim 4, characterized in that: The switch card also includes: a plurality of data center front-end interfaces, wherein: A plurality of data center front-end interfaces are connected to a switching chip and to a coprocessor resource pool or a main resource pool, so as to enable the switching chip to communicate with the coprocessor resource pool and the main resource pool respectively.
10. The resource pool according to claim 4, characterized in that: The management board also includes: a second selection path, the first baseboard includes: a baseboard interface, a second selection path connected to the baseboard management controller and the baseboard interface respectively; The backplane interface is connected to the out-of-band management chip and at least one management module respectively, and is used to extend the signal path so that the backplane management controller is connected to the out-of-band management chip and at least one management module respectively for backup communication.
11. The resource pool according to claim 10, characterized in that: The management board also includes: a high-speed connector, wherein: A high-speed connector is connected to the switch board and is used to connect to an external management device or an expansion interface.
12. The resource pool according to claim 11, characterized in that: The management board also includes: a programmable logic device, wherein: The programmable logic device is connected to the baseboard management controller and is used to control the power-on and power-off timings and the reset timings, and send the whole machine power-on status and reset status to the baseboard management controller.
13. The resource pool according to claim 12, characterized in that: The third network management chip and the fourth network management chip have the same address for easy management.
14. A switching device, characterized in that: The switching device includes the resource pool according to any one of claims 12 or 13, comprising a box, and a first area and a second area are provided in the box along a first direction, wherein: The first area is provided with a switch board, and the second area is provided with a management board. The switch board is detachably connected to the management board.
15. The switching device according to claim 14, wherein: The management module includes: a first management module module, a second management module module and a baseboard management controller; The first management module module and the second management module module are spaced apart, and the baseboard management controller is arranged between the first management module module and the second management module module; the first management module module and the second management module module face opposite directions.
16. A timing control method for a main system, characterized in that: The main system includes: A main resource pool, a coprocessor resource pool, and a resource pool according to any one of claims 12 or 13; Alternatively, the main system includes: a main resource pool, a coprocessor resource pool, and a switching device according to any one of claims 14 or 15; and the timing control method of the main system includes: After the main resource pool, the coprocessor resource pool and the resource pool are powered on, the baseboard management controller of the coprocessor resource pool and the baseboard management controller of the main resource pool send a preparation completion signal to the baseboard management controller of the resource pool; When receiving the power-on signal, the baseboard management controller of the resource pool controls the power-on based on the programmable logic device, and sends the power-on signal to the baseboard management controller of the coprocessor resource pool and the baseboard management controller of the resource pool respectively through the network; After receiving the power-on signal and completing the power-on process, the baseboard management controller of the coprocessor resource pool and the baseboard management controller of the resource pool respectively feed back a power-on completion signal to the baseboard management controller of the resource pool; Upon receiving the shutdown signal, the resource pool controls the shutdown and outputs a relation signal to the baseboard management controller of the coprocessor resource pool and the baseboard management controller of the resource pool respectively; The baseboard management controller of the coprocessor resource pool and the baseboard management controller of the resource pool receive the shutdown signal and shut down respectively.
17. A reset timing control method for a main system, characterized in that: The main system includes: A main resource pool, a coprocessor resource pool, and a resource pool according to any one of claims 12 or 13; Alternatively, the main system includes: a main resource pool, a coprocessor resource pool, and a switching device according to any one of claims 14 or 15; and the reset timing control method of the main system includes: After the main system is powered on, the resource pool management module scans the device identifiers of the coprocessor resource pool and the main resource pool and confirms the logical topology connection; When the main resource pool receives the reset signal, the programmable logic device of the main resource pool sends a reset signal to the switch card of the resource pool; After the switch card of the resource pool receives the reset signal and resets, the switch card of the resource pool sends a reset signal to the programmable logic device of the coprocessor resource pool; The programmable logic device of the coprocessor resource pool receives the reset signal and resets.
18. The reset timing control method of the main system according to claim 17, characterized in that: Also includes: When the main resource pool receives the reset signal, the programmable logic device of the main resource pool sends a reset signal to the programmable logic device of the resource pool; After the programmable logic device of the resource pool receives the reset signal and resets, the programmable logic device of the resource pool sends a reset signal to the programmable logic device of the coprocessor resource pool, or the programmable logic device of the resource pool sends a reset signal to the switch board of the resource pool; The programmable logic device of the coprocessor resource pool receives the reset signal and resets.
19. A method for controlling reset timing of a resource pool, characterized in that: The method is applied to the resource pool according to any one of claims 12 or 13, or the method is applied to the switching device according to any one of claims 14 or 15; the reset timing control method of the resource pool includes: The management module of the control resource pool sends a reset instruction to the programmable logic device of the resource pool; The programmable logic device of the control resource pool resets the switching chip in the switching board corresponding to the reset instruction, or sends a reset instruction to the coprocessor resource pool corresponding to the reset instruction.
20. The method for controlling the reset timing of a resource pool according to claim 19, wherein: Also includes: When the switch chip of the resource pool receives the reset signal of the main resource pool, it sends a reset signal to the coprocessor resource pool; The switching chip module of the resource pool sends a reset completion signal to the management module of the resource pool.
Citation Information
Patent Citations
Equipment reset method and device, storage medium and electronic equipment
CN117251039A
Server system, resource scheduling method of server system, chip and chip grain
CN118210634A