Management of network devices in servers
By configuring a supplementary network device in the server to replace the faulty primary network device, the performance issues caused by network failures were resolved, ensuring the stable and efficient operation of the server.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUPER MICRO COMPUTER INC(US)
- Filing Date
- 2025-11-18
- Publication Date
- 2026-05-22
AI Technical Summary
Failure of network devices in the server can lead to packet loss, poor data transmission quality, connection interruption, bottlenecks or latency spikes, affecting server performance and response time. Furthermore, outdated firmware may cause compatibility issues and system instability.
Configure one or more supplementary network devices in the server to replace the faulty main network device, and achieve seamless replacement through the Baseboard Management Controller (BMC) and Intelligent Platform Management Interface (IPMI) to ensure the continuity of processor operation.
It enables automatic replacement of the main network device without interrupting processor operation when an error is detected, ensuring the stability and efficient operation of the server and avoiding performance degradation caused by network failures.
Smart Images

Figure CN122073554A_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to computer technology, including but not limited to methods, apparatus, structures, devices, and systems for managing computer devices or systems (e.g., network devices housed in server racks). Background Technology
[0002] Servers play a central role in driving big data and artificial intelligence (AI) applications by providing the processing, storage, and networking capabilities needed to manage and analyze the massive amounts of data generated from diverse sources, including Internet of Things (IoT) devices, social media, and enterprise systems. Servers heavily rely on network devices such as network device cards (NICs), routers, and switches to communicate with other servers, devices, and the Internet. These network devices work closely with the server's processor to manage data delivery, routing, and traffic control, ensuring seamless communication. However, potential problems with these network devices can lead to severe disruptions. For example, a faulty NIC can cause packet loss, resulting in poor data transmission quality or even connection interruptions. Routers or switches experiencing high traffic or misconfiguration can cause bottlenecks or latency spikes, impacting server performance and response times. Furthermore, outdated firmware on network devices can cause compatibility issues with newer processors, leading to unexpected crashes or system instability. Summary of the Invention
[0003] According to some embodiments of the present application disclosed herein, at least regular monitoring, firmware updates, and maintenance of network devices used in servers are essential to ensure efficient server operation and maintain stable connections. Various embodiments of the present application relate to methods, apparatuses, structures, devices, and systems for managing network devices of computer devices or systems (e.g., server computers housed in server racks). In addition to a set of main network devices already coupled and configured to work with the server's processor, the server also includes one or more supplementary network devices. Upon detecting a fault in one of the main network devices, the server configures one or more supplementary network devices to replace the faulty main network device, for example, without interrupting the operation of the associated processor coupled to the main network device.
[0004] In some embodiments, the server is used to perform artificial intelligence operations (e.g., model training, data inference). When one of a plurality of primary network devices mounted on a substrate (e.g., a printed circuit board (PCB)) fails to operate, a supplementary network device replaces the faulty primary network device, for example, by applying simple commands via an Intelligent Platform Management Interface (IPMI) associated with the server's Baseboard Management Controller (BMC).
[0005] In one aspect, some embodiments include a computer system that further includes a plurality of network devices and a first processor device coupled to the plurality of network devices. The plurality of network devices are configured to receive input signals and provide output signals, for example, according to a plurality of network protocols, and include a first group of primary network devices and a group of one or more supplementary network devices. The first processor device is configured to monitor the operation of the first group of primary network devices and, based on a determination that the first primary network device in the first group of primary network devices has an error, configure a first supplementary network device in the group of one or more supplementary network devices to replace the first primary network device.
[0006] In some embodiments, the first processor device is further configured to pair a plurality of second processor devices with the first group of main network devices by pairing each second processor device with at least one different main network device in the first group of main network devices. Furthermore, in some embodiments, the first processor device includes a central processing unit (CPU), and each second processor device includes a graphics processing unit (GPU). Furthermore, in some embodiments, the computer system further includes the plurality of second processor devices.
[0007] In some embodiments, the first processor device is configured to execute firmware to determine that the first primary network device has the error and enable the System Management Mode (SMM) where the first supplementary network device replaces the first primary network device. Alternatively, in some embodiments, the first processor device is configured to execute an operating system including an error handler to determine that the first primary network device has the error, release the first primary network device, and retrain and employ the first supplementary network device.
[0008] In some embodiments, the error includes one of the following: hardware failure of the first master network device, driver or firmware problem, resource exhaustion or overload, signal integrity problem, and link layer protocol error.
[0009] In another embodiment, some implementations include a method implemented at a computer system comprising a plurality of network devices and a first processor device coupled to the plurality of network devices. The plurality of network devices includes a first group of primary network devices and a group of supplementary network devices. The method includes monitoring the operation of the first group of primary network devices. The method further includes configuring a first supplementary network device in the group of one or more supplementary network devices to replace the first primary network device based on a determination that the first primary network device in the first group of primary network devices has an error.
[0010] In another aspect, some embodiments include a computer system. The computer system includes multiple network devices for receiving input signals and providing output signals, for example, according to multiple network protocols. The multiple network devices include a first group of primary network devices and a group of supplementary network devices. The computer system further includes a first processor device coupled to the multiple network devices and a memory having instructions stored therein, which, when executed by the one or more processors, cause the processor to perform operations including monitoring the first group of primary network devices and, based on a determination that the first primary network device in the first group has an error, configuring a first supplementary network device in the group of one or more supplementary network devices to replace the first primary network device.
[0011] In another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more programs that, when executed by a first processor device of a computer system, cause the first processor to perform operations including monitoring the operation of a first group of primary network devices. The first processor device is coupled to a plurality of network devices, and the plurality of network devices includes the first group of primary network devices and a group of supplementary network devices. The one or more programs further include instructions for monitoring the operation of the first group of primary network devices and, based on a determination that the first primary network device in the first group of primary network devices has an error, configuring a first supplementary network device in the group of one or more supplementary network devices to replace the first primary network device.
[0012] These illustrative embodiments and implementations are mentioned not to limit or define this disclosure, but to provide examples to aid in understanding it. Additional embodiments are discussed in the detailed description, and further description is provided therein. Attached Figure Description
[0013] To better understand the various described implementation schemes, the following detailed embodiments should be referred to in conjunction with the following figures, in which similar element symbols are used throughout the figures to refer to corresponding parts.
[0014] Figure 1 This is a front view of an instance server rack supporting one or more servers according to some embodiments.
[0015] Figure 2 According to some embodiments, it can be used as Figure 1 A block diagram of an example system module in a typical computer device used for server applications.
[0016] Figure 3A , 3B 3C and 3C are, respectively, perspective view, front view and rear view of an instance server according to some embodiments.
[0017] Figure 4 This is a block diagram of an example computer system including a first processor device and multiple network devices according to some embodiments.
[0018] Figure 5 It is a block diagram of an example computer system according to some embodiments, wherein one or more supplementary network devices are used in place of one or more corresponding main network devices.
[0019] Figure 6 This is a block diagram of an example computer system comprising two processor devices and one or more supplementary network devices according to some embodiments.
[0020] Figure 7A This is a block diagram of an example processor system of a computer system according to some embodiments, the computer system comprising two processor devices, each of which is coupled to a corresponding data switch.
[0021] Figure 7B This is a block diagram of an example processor system of a computer system according to some embodiments, the computer system comprising two processor devices, each of which is coupled to two corresponding data switches.
[0022] Figure 7C This is a block diagram of another example processor system according to some embodiments, the processor system comprising two processor devices, each of which is coupled to a corresponding data switch.
[0023] Figure 8 This is a schematic diagram of a computer system for managing network devices at the firmware level, according to some embodiments.
[0024] Figure 9 This is a schematic diagram of a computer system for managing network devices at the software level, according to some embodiments.
[0025] Figure 10 This is a flowchart of an example method for a network device for managing a computer system according to some embodiments.
[0026] Similar reference figures are used throughout several views in the accompanying drawings to refer to the corresponding parts. Detailed Implementation
[0027] Reference will now be made in detail to specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternatives may be used without departing from the scope of the claims, and that the subject matter may be practiced without these specific details.
[0028] Various embodiments of this application relate to methods, apparatus, structures, devices, and systems for managing network devices of computer devices or systems (e.g., server computers housed in server racks). In addition to a set of main network devices already coupled and configured to work with the server's processor, the server also includes one or more supplementary network devices. Upon detecting a fault in one of the main network devices, the server configures one or more supplementary network devices to replace the faulty main network device, for example, without interrupting the operation of the associated processor coupled to the main network device.
[0029] Figure 1 This is a front view of an example server rack 100 (also referred to as a rack mount, rack cabinet, or simply rack) capable of supporting one or more servers 120 according to some embodiments. The server rack 100 includes a frame 102 and a plurality of slots 104 and can be used in a data center, server room, or network cabinet to support, organize, and manage multiple computing equipment modules 106 (e.g., servers 120, storage devices 116S and 116N, networking equipment, and other types of hardware). Each of the plurality of slots 104 of the server rack 100 is configured to receive and support a corresponding computing equipment module 106. In some embodiments, the plurality of slots 104 includes at least one empty slot 104B, which is not used to provide mechanical support to any equipment module 106 and can receive an equipment module 106 (if desired). In some embodiments, the server rack 100 has a predefined width of 19 or 23 inches, a height of up to 84 inches or more, and a depth selected from 24, 32, 40, or 48 inches. A rack unit (1U) is the standard size of a server 120 and other equipment modules 106 installed in a server rack 100. The server rack 100 provides space for the server 120 and other equipment modules 106, and is 19 inches wide and has a height expressed in rack units (e.g., 1U, 2U, 4U).
[0030] Examples of computing equipment modules 106 supported by multiple slots 104 of server rack 100 include, but are not limited to, firewall module 108, switch box 110, server 120, display device 112, keyboard 114, solid-state drive (SSD) 116S, network attached storage device 116N, and uninterruptible power supply (UPS) 118. Each computing equipment module 106 plays a corresponding role in maintaining the network and computing environment. In some embodiments, firewall module 108 is a network security device that monitors and controls incoming and outgoing network traffic based on predetermined security rules, thereby establishing a barrier between a trusted internal network and an untrusted external network. Firewall module 108 may be placed near the network entry point to protect server rack 100 from unauthorized access, malware, and network attacks. In some embodiments, firewall module 108 includes packet filtering, stateful inspection, VPN support, and intrusion prevention system (IPS). In some embodiments, the switch box 110 is placed near the network entry point along with the firewall module 108 and is configured to receive incoming signals and forward them (e.g., converted to electrical signals) to different servers 120 mounted on the server rack 100. The switch box 110 is used in the server rack 100 to minimize cable length and ensure efficient network service management. The switch box 110 may support different speeds (e.g., 800 Gbps, 1.6 Tbs, 3.2 Tbs), has multiple ports (24, 48, etc.), and provides features such as Virtual LAN (VLAN) support, PoE (Power over Ethernet), and managed or unmanaged capabilities.
[0031] Multiple computing equipment modules 106 of server rack 100 may include multiple servers 120, each configured to provide data, resources, services, or programs to other client devices via one or more wired or wireless communication networks. Each server 120 is mounted in a slot 104 of server rack 100 and configured to provide one or more services (e.g., network hosting, database management, and application support). Servers 120 mounted on server rack 100 can provide higher processing power, larger memory capacity, redundant power supplies, and hot-swappable components to achieve higher availability and reliability compared to individual client devices. In some embodiments, one or more rack servers 120 include multiple graphics processing units (GPUs) configured to perform machine learning operations, for example, in a data center associated with machine learning tasks. In some embodiments, server 120 includes one or more processors, memory storing one or more programs for execution by one or more processors, and a system enclosure for enclosing one or more processors, memory, and power supply components.
[0032] SSD 116S and Network Attached Storage (NAS) device 116N are configured to provide storage space for server 120, which is mounted in server rack 100. SSDs use flash memory to store data and, compared to hard disk drive (HDD) devices, exhibit high speed, low latency, durability, low power consumption, and a variety of capacity and form factor. Conversely, the Network Attached Storage (NAS) device 116N is a dedicated file storage device that provides data access to a network and allows a large number of different types of client devices to retrieve data from a centralized disk capacity. In some embodiments, the NAS device 116N may have high capacity, Redundant Array of Independent Disks (RAID), support for multiple file sharing protocols (NFS, SMB / CIFS, FTP), user management, and backup features. In some embodiments, the SSD 116S is a storage drive for speed and is used, for example, within server 120 housed in the same server rack 100, while the NAS 116N is configured for file sharing, data backup, and remote access.
[0033] In some implementations, the UPS 118 is used to provide emergency power to other computing equipment modules 106 in the event of a power outage, allowing them to remain operational for a sufficient period of time to safely disconnect or switch to an alternative power source. In examples, the UPS 118 is mounted in a server rack 100 or placed in a bottom slot to support weight, thereby providing backup power to other computing equipment modules 106. The UPS 118 provides one or more of the following: battery backup, surge protection, voltage regulation, real-time monitoring, management software, and / or varying runtimes based on capacity and load.
[0034] Server rack 100 further includes multiple mechanical structures configured to provide mechanical support to or facilitate access to multiple computing equipment modules 106. The multiple mechanical structures include one or more of the following: open-frame racks (e.g., without doors or side panels), mounting rails, cable management features (e.g., arms, hooks, and trays), power boards, shelves, drawers, and empty panels. In some embodiments, the multiple mechanical structures also include rack housings (e.g., cabinets), lockable doors, and side panels to protect the computing equipment modules 106 from unauthorized access. In an example, server rack 100 includes or is coupled to multiple panels configured to convert server rack 100 into a server rack. In some embodiments, server rack 100 further includes a cooling or ventilation system to facilitate heat dissipation. Using server rack 100 helps optimize space, improve cooling efficiency, simplify maintenance, and enhance the overall organization and management of information technology (IT) infrastructure.
[0035] Figure 2 It is possible to use according to some embodiments Figure 1The diagram illustrates an example system module 200 in a typical electronic device used in server 120. System module 200 in this electronic device includes at least a processor module 202, a memory module 204 for storing programs, instructions, and data, an input / output (I / O) controller 206, one or more communication interfaces (e.g., a network device 208), and one or more communication buses 240 for interconnecting these components. In some embodiments, the I / O controller 206 allows the processor module 202 to communicate with I / O devices (e.g., a keyboard, mouse, or touchpad) via a universal serial bus interface. In some embodiments, the network device 208 includes one or more interfaces (e.g., for Wi-Fi, Ethernet, and Bluetooth networks), each allowing the electronic device to exchange data with another external source (e.g., a server or other electronic device). In some embodiments, the communication bus 240 includes a circuitry (sometimes referred to as a chipset) that interconnects the various system components included in system module 200 and controls communication between the system components.
[0036] In some embodiments, processor module 202 includes one or more central processing units (CPUs). In some embodiments, processor module 202 includes one or more graphics processing units (GPUs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), tensor processing units (TPUs), microcontrollers (MCUs), neural processing units (NPUs), or combinations thereof. In some embodiments, system module 200 further includes a baseboard management controller (BMC) 224 mounted on a motherboard for remote management (e.g., IPMI, Redfish standards). BMC 224 is configured to provide an interface allowing administrators to monitor, troubleshoot, and update server 120 without physical access. In some embodiments, system module 200 further includes BIOS / UEFI firmware 226 (e.g., included on the motherboard), which is configured to initialize and test hardware components during startup and provide an interface for configuring hardware settings.
[0037] More specifically, in some embodiments, network devices 208 applied to server 120 are configured to manage, route, or facilitate network traffic, thereby enabling communication within a network or the Internet. Examples of network devices 208 include, but are not limited to, NICs (e.g., Ethernet or Wi-Fi adapters), network switches, network routers, load balancers, firewalls, wireless access point (WAP) devices, modems, repeater nodes, network hubs, bridges, gateways, intrusion detection and prevention systems, and virtual private network (VPN) facilities. In some embodiments, a subset of network devices 208 is configured to exchange data with another external source of one or more CPUs. Alternatively or additionally, in some embodiments, a subset of network devices 208 is configured to exchange data with external sources that are not CPU processors (e.g., GPUs). In some implementations, multiple network devices 208 are applied in the network infrastructure of server 120, for example, in a data center or enterprise environment.
[0038] In some embodiments, memory module 204 includes high-speed random access memory, such as DRAM, static random access memory (SRAM), double data rate (DDR) dynamic random access memory (RAM), or other random access solid-state memory devices. In some embodiments, memory module 204 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. In some embodiments, memory module 204, or alternatively, the non-volatile memory devices within memory module 204, include non-transitory computer-readable storage media. In some embodiments, a memory slot is reserved on system module 200 for receiving memory module 204. Once inserted into the memory slot, memory module 204 is integrated into system module 200.
[0039] In some embodiments, system module 200 further includes one or more components selected from memory controller 210, solid-state drive (SSD) 212, hard disk drive (HDD) 214, power supply unit (PSU) 216, power management integrated circuit (PMIC) 218, graphics module 220, and audio module 222. Memory controller 210 is configured to control communication between processor module 202 and memory components in the electronic device that include memory module 204. SSD 212 is configured to apply integrated circuit assemblies to store data in the electronic device and, in many embodiments, is based on NAND or NOR memory. HDD 214 is a conventional data storage device used for storing and retrieving digital information based on electromechanical disks. PSU 216 is configured to receive multiple power signals 260 and provide multiple DC power supplies 250 (e.g., 12 V, 54 V). PMIC 218 is configured to modulate multiple DC power supplies 250 to other desired DC voltage levels (e.g., 5 V, 3.3 V, or 1.8 V) as required by various components or circuits within the electronic device (e.g., processor module 202). Graphics module 220 is configured to generate output images according to image / video formats desired by one or more display devices and feed them to one or more display devices. Audio module 222 is configured to facilitate audio signals to and from the input and output of the electronic device under the control of a computer program.
[0040] It should be noted that the communication bus 240 also interconnects and controls communication among various system components, including components 210 to 224.
[0041] Figure 3A , 3B 3C and 3C are perspective, front, and rear views of an example server 120 according to some embodiments, respectively. Server 120 includes two CPUs 302 and is configured to implement applications in virtualization, AI inference, machine learning, enterprise servers, software-defined storage, or cloud computing. In some embodiments, server 120 further includes a non-volatile memory fast (NVMe) drive 304 or a Serial Advanced Technology Attachment (SATA) drive 306 for accessing mass storage devices (e.g., hard disk drives 214, optical disk drives, and SSDs 212) and handling data workloads. References Figure 3BIn one example, server 120 may include 12 drive bays 308, each configured to receive a corresponding NVMe drive 304 or SATA drive 306. Furthermore, in some embodiments, server 120 further includes multiple memory slots 310 for receiving one or more memory modules 204 (e.g., Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM) dual in-line memory modules (DIMMs)). In this example, server 120 is housed in a compact 1U chassis and is used as a node in a data center.
[0042] In some embodiments, server 120 includes multiple data transmission interfaces ( Figure 3A (Not shown in the image), thus allowing for high-speed connectivity and scalability options. In this example, the data transfer interface includes a set of Peripheral Component Interconnect Rapid (PCIe) slots. A GPU or network device 208 can be integrated into server 120 via a PCIe slot. Additionally, see reference... Figure 3C In some embodiments, a set of data transfer interfaces 312 includes four 16-channel PCIe 5.0 slots and is exposed on the back side of server 120.
[0043] refer to Figure 3B In some embodiments, the front side of server 120 further includes one or more of the following: a power button 314, a Universal Serial Bus (USB) 316, one or more status LEDs 318, and a Unique Identifier (UID) button 320. (See reference) Figure 3C In some embodiments, the rear side of server 120 further includes one or more of the following: USB port 322, local area network (LAN) port 324, display port 326, and access interface 328 to PSU 216.
[0044] Figure 4 This is an example computer system 400 according to some embodiments, comprising a first processor device 402 (e.g., CPU 302, BMC 224) and a plurality of network devices 404 (e.g., Figure 1A block diagram of server 120 in the computer system 400. Multiple network devices 404 are configured to receive input signals and provide output signals to the computer system 400, for example, according to multiple network protocols. The multiple network devices 404 include a first group of primary network devices 404A and a group of supplementary network devices 404S. A first processor device 402 is coupled to the multiple network devices 404. The first processor device 402 is configured to monitor the operation of the first group of primary network devices 404A. Based on a determination that the first primary network device 404A-1 in the first group of primary network devices 404A has an error, the first processor device 402 configures a first supplementary network device 404S-1 in one or more supplementary network devices 404S to replace the first primary network device 404A-1. In some embodiments, after the first supplementary network device 404S-1 replaces the first primary network device 404A-1, the first supplementary network device 404S-1 is regarded as one of the primary network devices 404A and is monitored by the first processor device 402.
[0045] In some embodiments, each network device 404 is configured to manage, route, or facilitate network traffic to enable communication within a network or the Internet. Examples of network devices 404 include, but are not limited to, NICs (e.g., Ethernet or Wi-Fi adapters), network switches, network routers, load balancers, firewalls, WAP devices, modems, repeater nodes, network hubs, bridges, gateways, intrusion detection and prevention systems, and VPN facilities. For example, a NIC includes a physical card for connecting to a network and may be an Ethernet or Wi-Fi adapter. Network switches are configured to manage traffic between different servers 102 within a data center. Routers are configured to route data between different networks, and server 120 may be used, for example, to route traffic in an enterprise network or cloud environment. Load balancers are configured to distribute incoming network traffic. Firewalls are configured to filter traffic to protect server 120 from unauthorized access and potential threats. Modems are configured to modulate and demodulate signals communicated over telephone lines or cables. Repeater nodes are configured to amplify or regenerate signals.
[0046] In some embodiments, the computer system 400 further includes a plurality of second processor devices 406 (e.g., GPUs). The first processor device 402 pairs the plurality of second processor devices 406 with the first group of main network devices 404A by pairing each of the second processor devices 406 with at least one different main network device 404A in the first group of main network devices 404A. See also Figure 4In one example, a first processor device 402 is coupled to four main network devices 404A and four second processor devices 406, and each second processor device 406 is paired with a different main network device 404A. In another example (not shown), one of the second processor devices 406 is paired with two or more main network devices 404A.
[0047] In some embodiments, the computer system 400 further includes a plurality of processor-side data interfaces 408 coupled to a first processor device 402 and a plurality of device-side data interfaces 410 coupled to a plurality of network devices 404. Both the plurality of processor-side data interfaces 408 and the plurality of device-side data interfaces 410 are configured to operate based on a predefined data transfer protocol, and each processor-side data interface 408 and its corresponding device-side data interface 410 are uniquely associated with each other and have a predefined number of channels associated with the predefined data transfer protocol. For example, the predefined data transfer protocol is PCIe, and the predefined number is an integer in the range of 1 to 16 (inclusive). In some embodiments, the first processor device 402 monitors the operation of a first group of main network devices 404A by monitoring the data communication status associated with each of the plurality of processor-side data interfaces 408 or by receiving messages from the processor-side data interfaces 408 indicating whether a corresponding main network device 404A coupled to a data interface 408 is operating correctly.
[0048] In some embodiments, the computer system 400 further includes a data switch 412 coupled between a first processor device 402 and a first group of main network devices 404A. In other words, each network device 404 is indirectly coupled to the first processor device 402 at least via the data switch 412. The data switch 412 is configured to select the first group of main network devices 404A (e.g., multiple network devices 404) to exchange data with the first processor device 402. Furthermore, in some embodiments, the data switch 412 is coupled to multiple network devices 404 via multiple device-side data interfaces 410 and multiple processor-side data interfaces 408.
[0049] In some embodiments, the computer system 400 includes a first processor substrate 414 configured to support a first processor device 402 and an I / O device substrate 416 further configured to support a plurality of network devices 404. In an example, the first processor substrate 414 includes a motherboard of server 120, and a second processor device 406 is also mounted on the motherboard. Alternatively, in some embodiments, the first processor device 402 pairs a plurality of second processor devices 406 with a first group of main network devices 404A. The computer system 400 includes a second processor substrate 418 for supporting the plurality of second processor devices 406, and the second processor substrate 418 is separate from the first processor substrate 414 and the I / O device substrate 416. In some embodiments, each of the first processor substrate 414, the I / O device substrate 416, and the second processor substrate 418 (if any) has a corresponding power supply.
[0050] Figure 5 This is an example computer system 500 according to some embodiments (e.g., Figure 1 The diagram illustrates a server 120, where one or more supplementary network devices 404S are used to replace one or more corresponding primary network devices 404A. The computer system 500 includes a first processor device 402 (e.g., CPU 302, BMC 224) and multiple network devices 404 for receiving input signals and providing output signals to the computer system 500. The multiple network devices 404 include a first group of primary network devices 404A and a group of supplementary network devices 404S. The first processor device 402 monitors the operation of the first group of primary network devices 404A. Based on a determination that the first primary network device 404A-1 in the first group of primary network devices 404A has an error, the first processor device 402 configures a first supplementary network device 404S-1 in the group of one or more supplementary network devices 404S to replace the first primary network device 404A-1.
[0051] In some embodiments, the first supplementary network device 404S-1 is coupled (502) to the first processor device 402 via a corresponding device-side data interface 410-1 and a corresponding processor-side data interface 408-1, for example, without involving a data switch 412. Alternatively, in some embodiments, the first supplementary network device 404S-1 is coupled (504) to the first processor device 402 via the device-side data interface 410-1, the processor-side data interface 408-2, and the data switch 412.
[0052] In some embodiments, the first processor device 402 monitors the operation of the first main network device 404A-1 by at least determining that the first main network device 404A-1 has an error and identifying a corresponding second processor device 406-1 paired with the first main network device 404A-1. The first processor device 402 replaces the first main network device 404A-1 with the first supplementary network device 404S-1 by at least pairing the first supplementary network device 404S-1 with the corresponding second processor device 406-1 to replace the first main network device 404A-1.
[0053] In some embodiments, the first processor device 402 monitors the operation of the first main network device 404A-1 by at least determining that the first main network device 404A-1 has an error and cannot be corrected using multiple error handling operations. A first supplementary network device 404S-1 is configured to replace the first main network device 404A-1 based on the determination that the error cannot be corrected using multiple error handling operations. In some embodiments, the multiple error handling operations are predefined and implemented by the first processor device 402 to correct the error detected in the first main network device 404A-1. If all multiple error handling operations fail to correct the error, then the first main network device 404A-1 is replaced.
[0054] In some embodiments, an error includes one of the following: hardware failure, driver or firmware problem, resource exhaustion or overload, signal integrity problem, and link layer protocol error. Examples of hardware failure include, but are not limited to, physically damaged components, memory corruption, circuit system failure, cable or connector problems, and transceiver problems. Examples of driver or firmware problems include, but are not limited to, outdated, damaged, or incompatible drivers and firmware defects or intermittent failures. Examples of resource exhaustion or overload include, but are not limited to, buffer overflows and high traffic or network congestion. Examples of signal integrity problems include, but are not limited to, electromagnetic interference (EMI) and signal loss or jitter. Examples of link layer protocol errors include, but are not limited to, cyclic redundancy check (CRC) failures and loss of synchronization.
[0055] In some embodiments, based on the determination that each of one or more second primary network devices 404A-2 in a plurality of network devices 404 has a corresponding error, the first processor device 402 configures a corresponding second supplementary network device 404S-2 in a set of one or more supplementary network devices 404S to replace the corresponding second primary network device 404A-2. Furthermore, in some embodiments, the corresponding second supplementary network device 404S-2 is directly coupled (506) to the first processor device 402 via corresponding data interfaces 408-3 and 410-2. Alternatively, in some embodiments, the corresponding second supplementary network device 404S-2 is indirectly coupled (508) to the first processor device 402 via a data switch 412 and corresponding data interfaces 408-4 and 410-2.
[0056] In some embodiments, the CPU and GPU of server 120 are used to perform artificial intelligence operations (e.g., model training, data inference). When the first main network device 404A-1, mounted on substrate 416 (e.g., PCB), fails to operate, the second processor device 406-1 (e.g., GPU) paired with the first main network device 404A-1 cannot transmit data via the first main network device 404A-1. The first processor device 402 includes the BMC 224 of server 120 (…). Figure 2 ), and execute the Intelligent Platform Management Interface (IPMI). In the IPMI (e.g., at the firmware level), commands are executed to replace the faulty main network device 404A-1 with the first supplementary network device 404S-1.
[0057] Figure 6 This is an example computer system 600 according to some embodiments, comprising two processor devices 402 and 602 and one or more supplementary network devices 404S (e.g., Figure 1 The diagram shows a server 120. The computer system 600 includes a first processor device 402 (e.g., CPU 302, BMC 224), a third processor device 602, and a plurality of network devices 404 for receiving input signals and providing output signals to the computer system 600. The plurality of network devices 404 includes a first group of main network devices 404A, a second group of main network devices 404B, and a group of one or more supplementary network devices 404S. The first processor device 402 monitors the operation of the first group of main network devices 404A. Based on a determination that the first main network device 404A-1 in the first group of main network devices 404A has an error, the first processor device 402 configures a first supplementary network device 404S-1 in the group of one or more supplementary network devices 404S to replace the first main network device 404A-1.
[0058] In some embodiments, the plurality of network devices 404 includes a second set of main network devices 404B. A third processor device 602 is coupled to the plurality of network devices and monitors the operation of the second set of main network devices 404B. Based on the determination that each of one or more third main network devices 404B-3 in the second set of main network devices has a corresponding error, the third processor device 602 configures a corresponding third supplementary network device 404S-3 in a set of one or more supplementary network devices 404S to replace the corresponding third main network device 404B-3.
[0059] In some embodiments, the computer system 600 includes a first processor substrate 414 configured to support a first processor device 402 and a third processor device 602, and an I / O device substrate 416 configured to support a plurality of network devices 404. The first processor device 402 pairs a plurality of second processor devices 406 with a first group of main network devices 404A. The computer system 600 includes a second processor substrate 418 for supporting the plurality of second processor devices 406, and the second processor substrate 418 is separate from the first processor substrate 414 and the I / O device substrate 416. In some embodiments, each of the first processor substrate 414, the I / O device substrate 416, and the second processor substrate 418 (if any) has a corresponding power supply.
[0060] In some embodiments, the computer system 600 includes a first processor substrate 414 configured to support a first processor device and a third processor device 602, and an I / O device substrate 416 configured to support a plurality of network devices 404. Furthermore, in some embodiments, the computer system 600 further includes a plurality of second processor devices 406 and a second processor substrate 418 for supporting the plurality of second processor devices 406. The second processor devices 406 are coupled to both the first processor device 402 and the third processor device 602. The first processor device 402 and the third processor device 602 are further configured to pair two distinct subsets 406A and 406B of the plurality of second processor devices 406 with a first group of main network devices 404A and a second group of main network devices 404B, respectively. More specifically, a first subset 406A of the second processor devices 406 is paired with the first group of main network devices 404A, and a second subset 406B of the second processor devices 406 is paired with the second group of main network devices 404B.
[0061] Furthermore, in some embodiments, based on the determination that the third main network device 404B-3 in the second group of main network devices 404B has a corresponding error, the first processor device 402 configures a corresponding one 404S-3 in one or more supplementary network devices 404S to replace the third main network device 404B-3. Furthermore, in some embodiments, the corresponding third supplementary network device 404S-3 is directly coupled (606) to the third processor device 602 via corresponding data interfaces 608-1 and 410-3. Alternatively, in some embodiments, the corresponding third supplementary network device 404S-3 is indirectly coupled (610) to the third processor device 602 via a data switch 604 and corresponding data interfaces 608-2 and 410-3.
[0062] Figure 7A It is a computer system according to some embodiments (e.g., Figure 6 A block diagram of an instance processor system 700 of a computer system 600, the computer system comprising two processor devices 402 and 602, each of which is coupled to a corresponding data switch. Figure 7B This is a block diagram of an example processor system 720 of a computer system according to some embodiments, the computer system including two processor devices 402 and 602, each of which is coupled to two respective data switches. Figure 7C This is a block diagram of another example processor system 740 according to some embodiments, the processor system including two processor devices 402 and 602, each of which is coupled to a corresponding data switch. Each of the processor systems 700, 720, and 740 is formed on a first processor substrate 414 and includes a first processor device 402 (e.g., a first CPU), a third processor device 602 (e.g., a second CPU), a BMC 224, and a plurality of data interfaces 408 and 608. In some embodiments, for each processor device 402 or 602, the plurality of data interfaces 408 or 608 includes one or more data interfaces directly coupled to the corresponding processor device 402 or 602. Alternatively or additionally, for each processor device 402 or 602, the plurality of data interfaces 408 or 608 includes (e.g., via...) Figure 7A Data switches 412 or 604 in the middle, via Figure 7B and 7C The data switches 412A, 412B, 604A or 604B in the middle are indirectly coupled to a subset of the data interfaces of the corresponding processor device 402 or 602.
[0063] In some embodiments, the first primary network switch 404A-1 (not shown) is at least connected via a data switch (e.g., Figure 7A The switch 412 in Figure 7B and7C The switch 412A or 412B in the first processor device 402 is coupled to the first processor device 402. When the first main network device 404A-1 is replaced by the first supplementary network device 404S-1 ( Figure 5 When the first supplementary network device 404S-1 is coupled to the first processor device 402, either via a data switch or without a data switch. Furthermore, in some embodiments, when the two main network switches 404A-1 and 404A-2 are replaced with two corresponding supplementary network devices 404S-1 and 404S-2... Figure 5 When the two supplementary network devices 404S-1 and 404S-2 are coupled to the first processor device 402 independently of each other, either via a data switch or without a data switch. In some embodiments, when the third main network switch 404B-3 is replaced by the third supplementary network device 404S-3, Figure 6 When the third supplementary network device 404S-3 is connected via a data switch (e.g., Figure 7A The 604 switch in the middle Figure 7B and 7C The switch 604A or 604B in the middle is coupled to the third processor device 602 without a data switch.
[0064] In some embodiments, each data interface 408 or 608 includes 16 data channels. See also Figure 7A In some embodiments, the first processor device 402 is coupled to a data switch 412 having 144 PCIe switches, and the data switch 412 is further coupled to eight data interfaces 408 having a total of 144 data channels, thus utilizing all 144 of the 144 PCIe switches in the data switch 412. See also Figure 7B In some embodiments, the first processor device 402 is coupled to two data switches 412A and 412B, each having 104 PCIe switches, and each data switch 412A or 412B is further coupled to four data interfaces 408 having a total of 64 data channels. In other words, for each data switch 412A or 412B, 64 of the 104 PCIe switches are used to control 64 data channels, and 44 idle data switches remain idle for controlling the 44 data channels. (See reference...) Figure 7C In some embodiments, the first processor device 402 is coupled to a data switch 412 having 180 PCIe switches, and the data switch 412 is further coupled to 10 data interfaces 408 having a total of 160 data channels, thus utilizing 160 of the 180 PCIe switches of the data switch 412.
[0065] In some embodiments, based on the determination of whether data switches 412, 412A, or 412B have unused switch components (e.g., unused PCIe switches), the computer system determines whether a supplementary network device 404S replacing the primary network device 404A is directly coupled to the first processor device 402 or indirectly coupled to the first processor device 402 via data switches 412, 412A, or 412B. For example, in some cases (e.g., with...) Figure 7A In the associated case, all switch components of data switch 412 are used to couple the main network device 404A or the second processor device 406. Replacing a supplementary network device 404S with a faulty main network device 404A may require direct coupling to the first processor device 402. Alternatively, in some cases (e.g., with...) Figure 7B (or 7C associated) Under these conditions, a set of switch components of data switch 412 has not yet been used to couple the main network device 404A or the second processor device 406. The supplementary network device 404S that replaces the faulty main network device 404A may be directly coupled to the first processor device 402 or indirectly coupled to the first processor device 402 via data switches 412, 412A or 412B.
[0066] Figure 8 This is a schematic diagram of a computer system 800 that manages a network device 404 at the firmware level, according to some embodiments. Figure 9 This is a schematic diagram of a computer system 900 in a software-level management network device 404 according to some embodiments. Each of the computer systems 800 or 900 includes a hardware layer 802, an operating system layer 804, a system software layer 806, and an application software layer 808. The hardware layer 802 includes a processor module 202 (e.g., Figures 4 to 7C CPU 810, processor device 402 and / or 602, memory module 204, storage device (e.g., SSD 212, hard disk drive 214), network device 208 (e.g., Figures 4 to 6The network device 208 and other peripheral devices 812 are included. The network device 208 and peripheral devices 812 may be coupled to the CPU 810 via a data interface 814 (e.g., PCIe). The operating system layer 804 sits above the hardware, acting as an intermediary between the hardware and system software, and is configured to manage hardware resources and provide a stable, consistent way for software applications to interact with the hardware without needing to know the specific details of the hardware. In some embodiments, the operating system 816 includes an error handler 818, implemented on the operating system layer 804. The system software layer 806 is used for system maintenance, performance enhancement, and bridging the gap between the operating system 816 and application software 822. Instances of firmware include, but are not limited to, utility software, device drivers 820, and compilers. The application software layer 808 includes software applications 822 (e.g., web browsers, video games) that utilize the functionality provided by the underlying layers to provide a wide range of features to improve productivity, task management, and enhance the user experience.
[0067] In some embodiments, computer system 800 or 900 includes firmware stored in non-volatile memory (such as read-only memory (ROM) or flash memory). The firmware comprises low-level software directly embedded in the hardware components and provides a basic control layer bridging the hardware layer 802 and the operating system layer 804. In an example, the firmware includes a Basic Input / Output System (BIOS) or a Unified Extensible Firmware Interface (UEFI) 824, which initializes and configures the hardware at startup and provides an interface between the hardware and the operating system 816.
[0068] refer to Figure 8 In some embodiments, a first processor device 402 (e.g., CPU 810) is configured to execute firmware (e.g., UEFI 824) to determine that the first primary network device 404A-1 has an error and to enable a System Management Mode (SMM) in which a first supplementary network device 404S-1 replaces the first primary network device 404A-1. Execution of the operating system 816 may be suspended in the SMM, thereby allowing the firmware to execute preferentially. The firmware maintains the device driver 820 associated with the faulty first primary network device 404A-1 and relinks the device driver 820 to the first supplementary network device 404S-1.
[0069] More specifically, in some embodiments, the first master network device 404A-1 is detected by the root port. Figure 5Uncorrected errors (Operation 832) are detected. For example, Transaction Layer Packets (TLPs) facilitate data transfer between PCIe devices via request and complete, and uncorrected errors can be detected based on malformed TLPs. The Enhanced Downstream Port Containment (EDPC) status and error source identifier (ID) are recorded. The root port programmed I / O (RP PIO) status is deregistered (if applicable). The root port sends a system control interrupt (Operation 834) to UEFI 824. UEFI 824 detects the EDPC status, reads the Advanced Error Report (AER) and EDPC registers, creates a system event log, and updates the common platform error log table. UEFI 824 clears the Uncorrected Error (UCE) status and removes the link from the Downstream Port Containment (DPC). An interrupt (Operation 836) is delivered to the operating system 816, which notifies (Operation 840) the driver 820 of the uncorrected error. The error handler 818 returns relevant information (Operations 838 and 842) to UEFI 824 and CPU 810. UEFI 824 enables SMM and replaces the first primary network device 404A-1 with the first supplementary network device 404S-1. An unexpected hot-plug interrupt can be delivered to the operating system 816.
[0070] refer to Figure 9 In some embodiments, the first processor device 402 (e.g., CPU 810) is configured to execute an operating system 816 including an error handler 818 to determine the first main network device 404A-1 ( Figure 5 If an error is detected in the first primary network device 404A-1, the first primary network device 404A-1 is released, and the first supplementary network device 404S-1 is retrained and adopted. In some embodiments, when an error is detected in the first primary network device 404A-1, the computer device continues to execute the operating system 816 on the first processor device 402, and the device driver 820 cooperates with the operating system 816 to relink the device driver 820 to the first supplementary network device 404S-1.
[0071] More specifically, in some embodiments, the first main network device 404A-1 is detected by the root port, for example, based on a malformed TLP (operation 832). Figure 5Uncorrected errors (UCE) are detected. Software-triggered DPCs can be used for verification purposes. The Enhanced Downstream Port Containment (EDPC) status and error source identifier (ID) are logged. The root port programmed I / O (RP PIO) status is deregistered (if applicable). The root port sends a system control interrupt to UEFI 824. UEFI 824 detects the EDPC status, reads the Advanced Error Report (AER) and EDPC registers, creates a system event log, and updates the common platform or log table. UEFI 824 clears the uncorrected error (UCE) status and removes the link from the downstream port containment (DPC). An interrupt (operation 902) is delivered to the operating system 816, which notifies the device driver 820 (operation 906) of the uncorrectable error. UEFI 824 enables SMM and replaces the first primary network device 404A-1 with the first supplementary network device 404S-1. An unexpected hot-plug interrupt can be delivered to the operating system 816. The root port sends a message signaling interrupt (MSI) to the error handler 818 of the operating system 816. Error handler 818 detects DPC events. It records the DPC status and error source ID, and also records the RP PIO status (if applicable). If the root port has DPC capability, the operating system 816 attempts to recover by releasing the link from the DPC, retraining and reactivating the link, or restoring the sub-device to its operational state. Error handler 818 returns the relevant information (operation 904) to CPU 810.
[0072] Figure 10 It is according to some embodiments for managing computer systems (e.g., Figure 1 A flowchart of an example method 1000 of a network device 404 (server 120) is shown. In some embodiments, method 1000 is performed by a computer system (e.g., a computer device 404) that stores data in a non-transitory computer-readable storage medium. Figures 4 to 6 Instruction control executed by one or more processors (e.g., BMC 224, CPU) in a computer system. Figure 10 Each of the operations shown may correspond to instructions stored in computer memory or computer-readable storage media of server 120. The computer-readable storage media may include disk or optical disk storage devices, solid-state storage devices (e.g., flash memory), or other non-volatile storage devices. Computer-readable instructions stored on the computer-readable storage media may include one or more of the following: source code, assembly language code, object code, or other instruction formats interpreted by one or more processors. Some operations in method 1000 may be combined and / or the order of some operations may be changed.
[0073] Method 1000 is implemented in a computer system comprising a plurality of network devices 404 and a first processor device 402 coupled to the plurality of network devices 404 (operation 1002). The plurality of network devices 404 includes a first group of main network devices 404A and a group of one or more supplementary network devices 404S. The computer system monitors (operation 1004) the operation of the first group of main network devices 404A. Based on the determination that the first main network device 404A-1 in the first group of main network devices 404A has an error, the computer system configures (operation 1006) the first supplementary network device 404S-1 in the group of one or more supplementary network devices 404S to replace the first main network device 404A-1.
[0074] In some embodiments, the first processor device 402 pairs a plurality of second processor devices 406 with the first group of main network devices 404A by pairing each second processor device 406 with at least one different main network device 404A in the first group of main network devices 404A (operation 1010) (operation 1008). Furthermore, in some embodiments, the first processor device 402 includes a central processing unit (CPU), and each second processor device 406 includes a graphics processing unit (GPU). In some embodiments, the first processor device 402 at least determines that the first main network device 404A-1 has an error and identifies a corresponding second processor device 406-1 paired with the first main network device 404A-1 (operation 1008). Figure 5 The first processor device 402 monitors the operation of the first main network device 404A-1. The first processor device 402 replaces the first main network device 404A-1 with the first supplementary network device 404S-1 by at least pairing the first supplementary network device 404S-1 with the corresponding second processor device 406-1 to replace the first main network device 404A-1.
[0075] In some embodiments, the first processor device 402 monitors the operation of the first main network device 404A-1 by at least determining that the first main network device 404A-1 has an error and that the error cannot be corrected using multiple error handling operations. The first supplementary network device 404S-1 replaces the first main network device 404A-1 based on the determination that the error cannot be corrected using multiple error handling operations.
[0076] In some embodiments, the computer system includes a plurality of processor-side data interfaces 408 coupled to a first processor device 402 and a plurality of device-side data interfaces 410 coupled to a plurality of network devices 404. Both the plurality of processor-side data interfaces 408 and the plurality of device-side data interfaces 410 operate based on a predefined data transfer protocol, and each processor-side data interface 408 and its corresponding device-side data interface 410 are uniquely associated with each other and have a predefined number of channels associated with the predefined data transfer protocol. Furthermore, in some embodiments, the predefined data transfer protocol is Peripheral Component Rapid Interconnect (PCIe), and the predefined number is an integer in the range of 1 to 16 (inclusive).
[0077] In some embodiments, the computer system includes a data switch 412 coupled between a first processor device 402 and a first set of main network devices 404A. Figures 4 to 6 Data switch 412 selects a first group of primary network devices 404A to exchange data with the first processor device 402. Furthermore, in some embodiments, the computer system further includes a plurality of processor-side data interfaces 408 coupled to the first processor device 402 and a plurality of device-side data interfaces 410 coupled to a plurality of network devices 404. Data switch 412 is coupled to the plurality of network devices 404 via the plurality of device-side data interfaces 410 and the plurality of processor-side data interfaces 408.
[0078] In some embodiments (e.g., with) Figure 5 (Associated), the first supplementary network device 408S-1 is coupled to the first processor device 402 via the corresponding device-side data interface 410-1 and processor-side data interface 408-1.
[0079] In some embodiments, the computer system includes a first processor substrate 414 configured to support a first processor device 402 and an input / output (I / O) device substrate 416 configured to support a plurality of network devices 404. Furthermore, in some embodiments, the computer system further includes a plurality of second processor devices 406 coupled to the first processor device 402 and a second processor substrate 418 for supporting the plurality of second processor devices 406. The first processor device 402 pairs the plurality of second processor devices 406 with a first set of main network devices 404A.
[0080] In some embodiments, according to one or more second main network devices 404A-2 among a plurality of network devices 404 Figure 5 Each of the following has a corresponding error determination, and the first processor device 402 configures a corresponding second supplementary network device 404S-2 in a group of one or more supplementary network devices 404S to replace the corresponding second main network device 404A-2.
[0081] In some embodiments, the plurality of network devices 404 includes a second set of main network devices 404B, and the computer system further includes (operation 1012) a third processor device 602 coupled to the plurality of network devices 404. The third processor device 602 monitors (operation 1014) the operation of the second set of main network devices 404B. Based on the determination that each of one or more third main network devices 404B-3 in the second set of main network devices 404B has a corresponding error, the third processor device 602 configures (operation 1016) a corresponding third supplementary network device 404S-3 in a set of one or more supplementary network devices 404S to replace the corresponding third main network device 404B-3. Furthermore, in some embodiments, the computer system further includes a first processor substrate 414 configured to support the first processor device 402 and the third processor device 602, and an input / output (I / O) device substrate 416 configured to support the plurality of network devices 404. In some embodiments, the computer system further includes a plurality of second processor devices 406 coupled to both the first processor device 402 and the third processor device 602, and a second processor substrate 418 for supporting the plurality of second processor devices 406. The first processor device 402 and the third processor device 602 connect two different subsets 406A and 406B of the plurality of second processor devices 406. Figure 6 They are paired with the first group of main network devices 404A and the second group of main network devices 404B, respectively.
[0082] In some embodiments, the first processor device 402 is configured to execute firmware to determine that the first primary network device 404A-1 has an error and enable the system management mode (SMM) in which the first supplementary network device 404S-1 replaces the first primary network device 404A-1.
[0083] In some embodiments, the first processor device 402 is configured to execute an operating system including an error handler to determine that the first primary network device 404A-1 has an error, release the first primary network device 404A-1, and retrain and employ the first supplementary network device 404S-1.
[0084] In some embodiments, each of the plurality of network devices 404 includes one of the following: a network interface card, a switch, a router, a load balancer, a firewall, a wireless access point, a modem, a repeater, a hub, a bridge, and a gateway device.
[0085] It should be understood that Figure 10The specific order of operations described herein is merely illustrative and is not intended to indicate that the described order is the only possible order of operations. Those skilled in the art will recognize the various ways in which signal timing on a serial data interface can be managed as described herein. Furthermore, it should be noted that other diagrams (e.g., Figures 1 to 9 The details of other processes described can also be similar to those described above. Figure 10 The method described applies to method 1000. For the sake of brevity, these details will not be repeated here.
[0086] The terminology used in the description of the various embodiments described herein is for the purpose of describing a particular embodiment only and is not intended to be restrictive. As used in the description of the various described embodiments and the appended claims, the singular forms “a / an” and “described” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will be further understood that when the terms “includes / including” and / or “comprises / comprising” are used in this specification, they specify the presence of the stated feature, integer, step, operation, element, and / or component, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, it will be understood that although the terms “first,” “second,” etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another.
[0087] As used herein, depending on the context, the term "if" may optionally be interpreted as meaning "when" or "at the time of" or "in response to determination" or "in response to detection" or "according to determination". Similarly, depending on the context, the phrase "if determination" or "if [the condition or event] is detected" may optionally be interpreted as meaning "when determination" or "in response to determination" or "when [the condition or event] is detected" or "according to the determination that [the condition or event] is detected".
[0088] For purposes of explanation, the foregoing description has been illustrated with reference to specific embodiments. However, the illustrative discussion above is not intended to be exhaustive or to limit the claims to the precise form disclosed. Many modifications and variations are possible in light of the foregoing teachings. Embodiments were chosen and described to best explain the principles of operation and practical application, thereby enabling understanding by others skilled in the art.
[0089] Although the various figures illustrate several logical stages in a specific order, stages that are not dependent on the order can be reordered, and other stages can be combined or decomposed. While some reorderings or other groupings are explicitly mentioned, others will be obvious to those skilled in the art, and therefore the orderings and groupings presented herein are not an exhaustive list of alternatives. Furthermore, it should be recognized that stages can be implemented in hardware, firmware, software, or any combination thereof.
Claims
1. A computer system comprising: Multiple network devices for receiving input signals and providing output signals, the multiple network devices including a first group of main network devices and a group of one or more supplementary network devices; A first processor device, coupled to the plurality of network devices, is configured to: Monitor the operation of the first group of main network devices; and Based on the determination that the first primary network device in the first group of primary network devices has an error, the first supplementary network device in the group of one or more supplementary network devices is configured to replace the first primary network device.
2. The computer system of claim 1, wherein the first processor device is further configured to: The plurality of second processor devices are paired with the first group of main network devices by pairing each second processor device with at least one different main network device in the first group of main network devices.
3. The computer system of claim 2, wherein the first processor device includes a central processing unit (CPU) and each second processor device includes a graphics processing unit (GPU).
4. The computer system of claim 2, wherein the first processor device is configured to monitor the operation of the first main network device at least in the following ways: It was determined that the first main network device had the error; Identify the corresponding second processor device paired with the first main network device; The first processor device is configured to replace the first main network device with the first supplementary network device by at least pairing the first supplementary network device with the corresponding second processor device to replace the first main network device.
5. The computer system of claim 1, wherein the first processor device is configured to monitor the operation of the first main network device at least in the following ways: It is determined that the first main network device has the error; and It was determined that the error could not be corrected using multiple error handling operations; The first supplementary network device is configured to replace the first primary network device based on a determination that the error cannot be corrected using the plurality of error handling operations.
6. The computer system according to claim 1, further comprising: Multiple processor-side data interfaces are coupled to the first processor device; and Multiple device-side data interfaces are coupled to the multiple network devices; The plurality of processor-side data interfaces and the plurality of device-side data interfaces are configured to operate based on a predefined data transmission protocol, and each processor-side data interface and its corresponding device-side data interface are uniquely associated with each other and have a predefined number of channels associated with the predefined data transmission protocol.
7. The computer system of claim 6, wherein the predefined data transfer protocol is Peripheral Component Interconnect (PCIe), and the predefined number is an integer in the range of 1 to 16 (inclusive).
8. The computer system according to claim 1, further comprising: A data switch is coupled between the first processor device and the first group of main network devices, the data switch being configured to select the first group of main network devices to exchange data with the first processor device.
9. The computer system according to claim 8, further comprising: Multiple processor-side data interfaces are coupled to the first processor device; and Multiple device-side data interfaces are coupled to the multiple network devices; The data switch is coupled to the multiple network devices via the multiple device-side data interfaces and the multiple processor-side data interfaces.
10. The computer system of claim 1, wherein the first supplementary network device is coupled to the first processor device via a corresponding device-side data interface and a processor-side data interface.
11. The computer system according to claim 1, further comprising: A first processor substrate, configured to support the first processor device; and Input / output (I / O) device substrate configured to support the plurality of network devices.
12. The computer system according to claim 1, further comprising: A plurality of second processor devices coupled to the first processor device, wherein the first processor device is further configured to pair the plurality of second processor devices with the first group of main network devices; and A second processor substrate is used to support the plurality of second processor devices.
13. A method comprising: In a computer system comprising a plurality of network devices and a first processor device coupled to the plurality of network devices, wherein the plurality of network devices includes a first set of main network devices and a set of supplementary network devices: Monitor the operation of the first group of main network devices; and Based on the determination that the first primary network device in the first group of primary network devices has an error, the first supplementary network device in the group of one or more supplementary network devices is configured to replace the first primary network device.
14. The method of claim 13, further comprising: Based on the determination that each of the one or more second primary network devices in the plurality of network devices has a corresponding error, the corresponding second supplementary network device in the group of one or more supplementary network devices is configured to replace the corresponding second primary network device.
15. The method of claim 13, wherein the plurality of network devices includes a second set of main network devices, and the computer system further includes a third processor device coupled to the plurality of network devices, the method further comprising: Monitor the operation of the second group of main network devices; and Based on the determination that each of the one or more third main network devices in the second group of main network devices has a corresponding error, the corresponding third supplementary network device in the one or more supplementary network devices in the group is configured to replace the corresponding third main network device.
16. The method of claim 15, wherein the computer system further comprises a first processor substrate configured to support the first processor device and the third processor device, and an input / output I / O device substrate configured to support the plurality of network devices.
17. The method of claim 15, wherein the computer system further comprises: A plurality of second processor devices coupled to both the first processor device and the third processor device, wherein the first processor device and the third processor device are further configured to pair two different subsets of the plurality of second processor devices with the first group of main network devices and the second group of main network devices, respectively. and A second processor substrate is used to support the plurality of second processor devices.
18. A non-transitory computer-readable storage medium having instructions stored thereon, which, when executed by a first processor device of a computer system, cause the first processor device to perform operations including: Monitoring the operation of a first group of main network devices, wherein the first processor device is coupled to a plurality of network devices, the plurality of network devices including the first group of main network devices and a group of supplementary network devices; and Based on the determination that the first primary network device in the first group of primary network devices has an error, the first supplementary network device in the group of one or more supplementary network devices is configured to replace the first primary network device.
19. The non-transitory computer-readable storage medium of claim 18, wherein the first processor device is configured to execute firmware to (1) determine that the first primary network device has the error, and (2) enable the system management mode (SMM) in which the first supplementary network device replaces the first primary network device.
20. The non-transitory computer-readable storage medium of claim 18, wherein the first processor device is configured to execute an operating system including an error handler to (1) determine that the first primary network device has the error, (2) release the first primary network device, and (3) retrain and employ the first supplementary network device.