System and method for managing device failover and data routing in network system

By introducing controllers, schedulers, and routing units into the switching equipment, the device status and data routing paths are dynamically managed, solving the problem of system performance degradation caused by device failures and achieving high availability and efficient data transmission.

CN121967311APending Publication Date: 2026-05-01AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
Filing Date
2025-10-20
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In complex systems, device failures can lead to a chain reaction, affecting the overall system performance and efficiency. Existing technologies struggle to effectively manage device failures and maintain system reliability.

Method used

By introducing controllers, schedulers, and routing units into the switching equipment, the device status is dynamically monitored and active and passive states are reallocated when a fault occurs, and data routing paths are dynamically switched to ensure the continuity of data traffic.

Benefits of technology

It achieves high availability and efficient data transmission in the event of device failure, reduces downtime, and avoids underutilization of expensive resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967311A_ABST
    Figure CN121967311A_ABST
Patent Text Reader

Abstract

The present technology relates to systems and methods for managing device failover and data routing in a network system. The apparatus includes an input configured to receive an input voltage and an output configured to provide an output voltage. The apparatus includes a first circuit configured to generate a first signal associated with the output voltage. The apparatus further includes a first comparator configured to compare the first signal to a first reference voltage and generate a second signal based on the comparison. The apparatus further includes a switch configured to receive the second signal and adjust a first resistance in a current path between the input and the output based on the second signal. The device implements multi-stage surge current control, allowing dynamic adjustment of the surge current at different stages of the power-on phase.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to computer networks. Background Technology

[0002] In modern computing and networking environments, reliable and efficient communication between devices is crucial for maintaining system performance and uptime. Many systems involve multiple devices—such as network interface cards (NICs), storage devices, and processing units—working together to handle large volumes of data traffic. These devices can be interconnected via switches that manage data routing between the devices and external systems, including host systems and other endpoints.

[0003] Some methods for data transfer between devices rely on Direct Memory Access (DMA), which allows devices to directly access memory without burdening the central processing unit (CPU). This improves overall efficiency by reducing processing overhead and enabling faster data transfers. For example, Peripheral Component Interconnect High Speed ​​(PCIe) is a standard that supports high-speed communication between devices such as network interface cards, processing units, and memory controllers. PCIe enables direct connections between devices via a bus architecture, facilitating efficient data flow between multiple endpoints through switches.

[0004] As systems become more complex, especially those with high-performance workloads such as artificial intelligence (AI) and machine learning (ML), the likelihood of device failure increases. These workloads typically rely on multiple devices working collaboratively, and the failure of one device can have a cascading effect on the entire system. For example, when a NIC that transmits data to one or more processing units fails, associated processing units may become idle, resulting in a loss of processing power and reduced overall system efficiency.

[0005] Various methods have been explored to resolve device failures in complex systems, but they have proven to be insufficient. It is important to recognize the need for new and improved systems and methods. Summary of the Invention

[0006] On one hand, this disclosure relates to a switching device comprising: a first port coupled to a first device; a second port coupled to a second device; a controller coupled to the first port, the controller being configured to assign an active state to the first device and a passive state to the second device; a scheduler coupled to the controller, the scheduler being configured to monitor the operational state of the first device by detecting a first fault associated with the first device; and a routing unit coupled to the controller, the routing unit being configured to determine a first routing path between the first device and the first port for managing data traffic, the routing unit including a routing table configured to store the first routing path; wherein, in response to the scheduler detecting the first fault, the controller is configured to reassign the active state from the first device to the second device and the passive state from the second device to the first device, and the routing unit is configured to determine a second routing path between the second device and the second port, and update the routing table to store the second routing path.

[0007] On the other hand, this disclosure relates to a switching device comprising: a first port coupled to a first device; a second port coupled to a second device; a controller coupled to the first port, the controller being configured to assign an active state to the first device and a passive state to the second device; a scheduler coupled to the controller, the scheduler being configured to monitor the operational state of the first device by detecting a first fault associated with the first device; and a routing unit coupled to the controller, the routing unit being configured to determine a first routing path between the first device and the first port for managing data traffic; wherein, in response to the scheduler detecting the first fault, the controller is configured to reassign the active state from the first device to the second device and the passive state from the second device to the first device, and the routing unit is configured to determine a second routing path between the second device and the second port.

[0008] On the other hand, this disclosure relates to a method comprising: assigning an active state to a first device coupled to a first port and a passive state to a second device coupled to a second port by a controller; monitoring the operational state of the first device by a scheduler; determining a first routing path between the first device and the first port for managing data traffic by a routing unit; reassigning the active state to the second device and the passive state to the first device in response to detecting a first fault associated with the first device; and determining a second routing path between the second device and the second port for managing data traffic by the routing unit. Attached Figure Description

[0009] A further understanding of the nature and advantages of particular embodiments can be achieved by referring to the remainder of the specification and the accompanying drawings, in which the same reference numerals are used to refer to similar components. In some instances, sublabels are associated with reference numerals to indicate one of a plurality of similar components. When reference numerals are cited without specifying sublabels, they are intended to refer to all such plurality of similar components.

[0010] Figure 1 This is a block diagram illustrating the architecture of a computing system according to various embodiments of the present technology.

[0011] Figure 2 This is a block diagram illustrating the architecture of a computing system according to various embodiments of the present technology.

[0012] Figure 3 This is a block diagram illustrating various embodiments of a switching device according to the present technology. Detailed Implementation

[0013] This technology relates to a switching device for managing device status and data traffic among multiple devices. In an embodiment, the switching device includes a first port coupled to a first device and a second port coupled to a second device. The device further includes a controller configured to assign an active state to the first device and a passive state to the second device. The device further includes a scheduler configured to monitor the operational status of the first device and detect faults. Upon detecting a fault, the controller reassigns the active state to the second device and the passive state to the first device, ensuring continuous data traffic flow and reducing downtime through dynamic switching. Other embodiments also exist.

[0014] One general aspect includes a switching device comprising: a first port coupled to a first device; a second port coupled to a second device; a controller coupled to the first port, the controller being configured to assign an active state to the first device and a passive state to the second device; a scheduler coupled to the controller, the scheduler being configured to monitor the operational state of the first device by detecting a first fault associated with the first device; and a routing unit coupled to the controller, the routing unit being configured to determine a first routing path between the first device and the first port for managing data traffic, the routing unit including a routing table configured to store the first routing path. In response to the scheduler detecting the first fault, the controller is configured to reassign the active state from the first device to the second device and the passive state from the second device to the first device, and the routing unit is configured to determine a second routing path between the second device and the second port, and update the routing table to store the second routing path.

[0015] The implementation may include one or more of the following features: The first device includes a first network interface card (NIC) and the second device includes a second NIC. The scheduler is configured to monitor the operational status of the first device based on predefined time intervals. The first fault is detected based on the loss of electrical connectivity between the first device and the first port. The first fault is detected based on errors in the configuration space of the first device. The first fault is detected based on the success rate of data transactions between the first device and the first port. The switching device further includes a third port coupled to a third device. The third device includes a graphics processing unit (GPU). The third device includes a storage device. The first device is coupled to the second device via a Peripheral Component Interconnect High Speed ​​(PCIe) interface. The switching device further includes a fourth port coupled to a host, and the controller is configured to communicate the active status of the first device to the host.

[0016] According to another embodiment, the present technology provides a switching device comprising: a first port coupled to a first device; a second port coupled to a second device; a controller coupled to the first port, the controller being configured to assign an active state to the first device and a passive state to the second device; a scheduler coupled to the controller, the scheduler being configured to monitor the operational state of the first device by detecting a first fault associated with the first device; and a routing unit coupled to the controller, the routing unit being configured to determine a first routing path between the first device and the first port for managing data traffic. In response to the scheduler detecting the first fault, the controller is configured to reassign the active state from the first device to the second device and the passive state from the second device to the first device, and the routing unit is configured to determine a second routing path between the second device and the second port.

[0017] The implementation scheme may include one or more of the following features: The first device includes a first network interface card (NIC). The first fault is detected based on the loss of electrical connectivity between the first device and the first port. The first fault is detected based on errors in the configuration space of the first device. The first fault is detected based on the success rate of data transactions between the first device and the first port.

[0018] According to yet another embodiment, the present technology provides a method comprising: assigning an active state to a first device coupled to a first port and a passive state to a second device coupled to a second port by a controller; monitoring the operational state of the first device by a scheduler; determining a first routing path between the first device and the first port for managing data traffic by a routing unit; reassigning the active state to the second device and the passive state to the first device in response to detecting a first fault associated with the first device; and determining a second routing path between the second device and the second port for managing data traffic by the routing unit. In various embodiments, the first device includes a first network interface card (NIC) and the second device includes a second NIC. The first fault is detected based on the loss of electrical connectivity between the first device and the first port.

[0019] The following description is presented to enable those skilled in the art to make and use the invention in the context of a particular application. Those skilled in the art will readily understand the various modifications and uses in different applications, and the general principles defined herein are applicable to a wide range of embodiments. Therefore, the art is not intended to be limited to the presented embodiments, but should be accorded the broadest scope consistent with the principles and novel features disclosed herein.

[0020] In the following detailed description, numerous specific details are set forth to provide a more thorough understanding of the art. However, those skilled in the art will understand that the art can be practiced without being limited to these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the art.

[0021] Readers should pay attention to all papers and documents filed concurrently with and published for public examination with this specification, the contents of which are incorporated herein by reference. Unless otherwise expressly stated, all features disclosed in this specification (including any appended claims, abstracts, and figures) may be replaced by alternative features used for the same, equivalent, or similar purposes. Therefore, unless otherwise expressly stated, each disclosed feature is merely one example of a general series of equivalent or similar features.

[0022] Furthermore, any element not expressly stated in the claims as a “component” for performing the specified function or a “step” for performing a particular function shall not be construed as a “component” or “step” as specified in paragraph 6 of section 112 of 35 U.S.C. Specifically, the use of “step of…” or “action of…” in the claims herein is not intended to invoke the provisions of paragraph 6 of 35 U.S.C. 112.

[0023] When an element is referred to herein as “connected” or “coupled” to another element, it should be understood that the element may be directly connected to the other element or have an intermediary element present between the elements. Conversely, when an element is referred to as “directly connected” or “directly coupled” to another element, it should be understood that there is no intermediary element in the “direct” connection between the elements. However, the presence of a direct connection does not preclude the existence of other connections in which intermediary elements may be present.

[0024] Furthermore, the terms left, right, front, back, top, bottom, forward, backward, clockwise, and counterclockwise are used for interpretive purposes only and are not limited to any fixed direction or orientation. Rather, they are used only to indicate the relative position and / or orientation between the various parts of an object and / or component.

[0025] Furthermore, for ease of description, the methods and processes described herein may be described in a specific order. However, it should be understood that unless the context otherwise indicates, intermediate processes may occur before and / or after any part of the described process, and various other procedures may be reordered, added, and / or omitted according to various embodiments.

[0026] Unless otherwise indicated, all figures used herein to express quantities, dimensions, etc., should be understood to be modified by the term “about” in all instances. In this application, unless otherwise specifically stated, the use of the singular includes the plural, and the use of the terms “and” and “or” means “and / or”, unless otherwise indicated. Furthermore, the use of the terms “comprising” and “having”, as well as other forms (e.g., “includes,” “included,” “has,” “have,” and “had”), should be considered non-exclusive. Additionally, terms such as “element” or “component” cover both elements and components comprising one unit and elements and components comprising more than one unit, unless otherwise specifically stated.

[0027] As used herein, the phrase “at least one of…” preceding a series of items separated by the terms “and” or “or” modifies the entire list, not each member of the list (i.e., each item). The phrase “at least one of…” does not require selection of at least one of every listed item; rather, it allows the meaning to include at least one of any item and / or at least one of any combination of items. For example, the phrases “at least one of A, B, and C” or “at least one of A, B, or C” each refer to only A, only B, or only C; and / or any combination of A, B, and C. In examples where the intention is to select “at least one of each of A, B, and C,” or alternatively, “at least one of A, at least one of B, and at least one of C,” this is explicitly stated.

[0028] Figure 1 This is a block diagram illustrating the architecture of a computing system 100 according to various embodiments of the present technology. This diagram is merely illustrative and should not unduly limit the scope of the claims. Many variations, alternatives, and modifications will be recognized by those skilled in the art.

[0029] In various implementations, System 100 represents a distributed computing architecture that interconnects multiple hardware components to facilitate seamless communication and high-speed data transfer. For example, System 100 is designed to support high-speed communication between multiple devices, such as network interface cards (NICs), graphics processing units (GPUs), and storage controllers. These devices are interconnected via switches, which facilitates data routing between devices and external systems, such as host systems and other endpoints. System 100 can be applied to various computing environments, such as data centers, AI / ML workloads, cloud computing, high-performance computing systems, and / or the like.

[0030] Depending on the implementation, system 100 may utilize direct memory access (DMA) to transfer data between components. For example, the term "direct memory access" can refer to a process where a device can directly transfer data between its own memory and system memory without intervention from the central processing unit (CPU). This mechanism reduces CPU overhead and accelerates data transfer rates, which is beneficial in high-performance computing environments where multiple devices frequently exchange large amounts of data. For example, in AI / ML workloads, the NIC can transfer data directly to the GPU for processing without requiring the CPU to handle each transaction.

[0031] In various implementations, PCI Express (PCIe) is used to facilitate high-speed communication between components. PCIe is a high-speed serial bus interface that allows low-latency, high-bandwidth data exchange between connected devices such as CPUs, memory, NICs, GPUs, and memory controllers. It supports chip-to-chip and board-to-board interconnects via cards and connectors, allowing multiple devices to communicate through a shared data path. PCIe is useful in high-performance computing environments where large amounts of data need to be efficiently transferred between processing units and memory.

[0032] As shown, system 100 includes a memory management unit (MMU) 101, which can be configured to manage memory access across devices within system 100. It translates virtual addresses (e.g., used by software) into physical addresses (e.g., used by hardware) to ensure that devices connected to system 100 can access appropriate memory locations. Instances of memory management units may include, but are not limited to, I / O memory management units (IOMMUs), CPU MMUs, GPU MMUs, virtual MMUs, and / or the like. Depending on the implementation, MMU 101 may be implemented as a separate dedicated hardware unit or integrated directly within the CPU as part of a system-on-a-chip (SoC) architecture.

[0033] In various embodiments, system 100 includes a root complex 102. The term "root complex" can refer to a central component in a PCIe hierarchy that connects a host system (e.g., CPU and / or system memory) to PCIe devices or endpoints. Root complex 102 acts as a communication bridge between the PCIe architecture and the host system, managing communication between upstream and downstream devices. During system initialization, root complex 102 can perform device enumeration, identify PCIe devices connected to the system, and assign an address to each device.

[0034] In some embodiments, system 100 further includes switch 103. For example, the term "switch" may refer to a hardware component that facilitates communication between multiple devices by managing the flow of data across a shared communication path. Examples of switches may include, but are not limited to, PCIe switches, Ethernet switches, InfiniBand switches, Fibre Channel switches, and / or the like. In some instances, switch 103 includes a PCIe switch designed to connect various PCIe-compatible devices, such as NICs, GPUs, storage devices, and other peripherals. The PCIe switch acts as an intermediary between these devices and root complex 102, facilitating high-speed data transfer between devices on the PCIe bus.

[0035] In various embodiments, system 100 may include one or more endpoint devices (e.g., devices 104, 105, 106). For example, the term "endpoint" or "endpoint device" may refer to any device connected to a shared bus that communicates with other components in the system via a switch or root complex. Examples of endpoints may include, but are not limited to, NICs, GPUs, storage devices, and / or other peripheral devices. For example, device 104 may include a first NIC responsible for handling network communications and data transfers to and from external networks. In systems where large amounts of data need to be ingested or distributed, such as in cloud computing or high-performance data centers, NICs are beneficial for efficiently moving data across systems.

[0036] In some instances, device 105 may include a first GPU, and device 106 may include a second GPU. GPUs can be used to handle computationally intensive tasks, such as AI model training, parallel data processing, or high-speed rendering. In various AI / ML workloads, multiple GPUs can be used to handle large datasets, increasing computational throughput and reducing the time required to complete large-scale computations.

[0037] In various implementations, these endpoint devices work together to enable high-speed data transfer across the system. For example, in an AI / ML workload, data from an external network (e.g., network 107) may be delivered to device 104 (e.g., a first NIC), which then transfers the data to device 105 (e.g., a first GPU) for processing. The processed data may then be shared with device 106 (e.g., a second GPU) for additional computation or stored in external storage, all facilitated by switch 103.

[0038] However, as the complexity of System 100 increases—especially in high-performance environments such as data centers, AI / ML applications, and cloud computing—the risk of device failure also increases. Devices such as NICs or GPUs may experience failures due to hardware malfunctions, network problems, or other factors, potentially leaving expensive resources like GPUs underutilized or idle. For example, if the NIC (e.g., device 104) responsible for transmitting data to the GPU (e.g., device 105) fails, the GPU may be unable to receive the necessary data for processing, resulting in a loss of processing power and reduced overall system efficiency. Therefore, System 100 is expected to implement a high-availability configuration that ensures continuous performance even in the event of hardware failure.

[0039] Figure 2 This is a block diagram illustrating the architecture of a computing system 200 according to various embodiments of the present technology. This diagram is merely illustrative and should not unduly limit the scope of the claims. Those skilled in the art will recognize many variations, alternatives, and modifications.

[0040] In various implementations, system 200 represents a distributed computing architecture that interconnects multiple hardware components to facilitate seamless communication and high-speed data transfer. For example, system 200 may include at least one of an MMU 201, a root complex 202, a switch 203, and / or one or more endpoint devices. In various instances, one or more endpoint devices may include a first NIC 204, a second NIC 207, a first processor 205, and a second processor 206. These endpoint devices are connected to switch 203, which manages data flow between the endpoints and an external network 208. MMU 201 manages access to shared memory resources, while root complex 202 acts as a bridge between switch 203 and the host system, facilitating communication between the CPU, memory, and various endpoint devices.

[0041] In some embodiments, one or more NICs (e.g., first NIC 204 and / or second NIC 207) may be responsible for receiving data from network 208 and performing DMA transfers to one or more processors (e.g., first processor 205 and / or second processor 206) for computation. In various instances, first processor 205 and second processor 206 may include one or more GPUs configured to handle computationally intensive tasks, such as AI model training, parallel data processing, or high-speed rendering. DMA allows data to be transferred directly from the NIC to system memory and / or peer devices (e.g., GPUs), bypassing the CPU, which reduces overhead and increases overall data transfer efficiency.

[0042] However, like any component in a system, a NIC can encounter errors, such as hardware failures, network problems, or other factors. When a NIC fails, it may lose its ability to transmit data, and in some cases, this can leave multiple GPUs without the data they need to process. Since GPUs are typically much more expensive than NICs, a failure in a NIC can lead to severe underutilization of expensive computing resources, resulting in inefficient system operation.

[0043] To address this issue, System 200 implements a failover mechanism to ensure uninterrupted operation in the event of a NIC failure. This mechanism allows the system to dynamically switch from a failed NIC to a backup NIC, ensuring system availability and continued full utilization of GPU resources. By automatically detecting errors and rerouting data traffic to the operational NIC, System 200 maintains high availability and minimizes downtime, providing an efficient and reliable computing environment.

[0044] According to some embodiments, operation of system 200 begins with an enumeration process during system initialization. During enumeration, root complex 202 identifies all devices connected via switch 203 and assigns a unique address to each device for communication. This process ensures that each endpoint device (e.g., NIC and processor) is recognized by system 200 and ready to communicate with root complex 202 and other components.

[0045] In some instances, the enumeration process may involve determining the operational status of the connected NICs. These statuses determine the role each NIC will play within the system. For example, the term "operational status" may refer to the current state or operating mode assigned to a particular device, such as whether the device is active, passive, or in standby mode. Operational status can be determined by monitoring various metrics, such as device activity, data transmission success rate, network connectivity, error detection, and / or the like.

[0046] In some embodiments, the first NIC 204 may initially be assigned an active state. For example, the term "active state" may refer to a state in which a device (e.g., the NIC) is responsible for handling active data transmission between the system and an external network (e.g., network 208). In this state, the NIC operates as the primary network interface, actively participating in sending and receiving data. On the other hand, during normal operation, the second NIC 207 may be placed in a passive state. For example, the term "passive state" may refer to a standby state in which a device (e.g., the NIC) remains idle but is ready to take over in the event of a failure in the active device. The device in the passive state does not handle active data transmission but monitors potential failover scenarios for the system. In the passive state, the second NIC 207 hides the system's operational flow to prevent conflicts within the system's device hierarchy.

[0047] During normal operation, the active NIC (e.g., the first NIC 204) manages all data transmissions between system 200 and external devices, including communication with network 208 and other internal system components such as the processor. Meanwhile, the passive NIC (e.g., the second NIC 207) remains inactive but remains ready to take over in the event of a failure. Throughout this process, switch 203 monitors the operational status of the active NICs to detect any potential problems or malfunctions.

[0048] Depending on the implementation, switch 203 continuously monitors the operational status of the first NIC 204 using various methods. For example, fault detection may be based on the loss of electrical connectivity between the first NIC 204 and switch 203, firmware errors, or by monitoring the error rate during data transmission. If switch 203 detects a loss of network connectivity or a high transmission failure rate, this can trigger a failover mechanism. Other fault detection mechanisms may include checking the health status reported by the NIC's internal diagnostics, or receiving an error signal when the NIC fails to respond to a regular data request.

[0049] When a fault is detected in the first NIC 204, switch 203 immediately triggers a failover process. Switch 203 can reassign the active state to the second NIC 207, making it the new primary network interface, while placing the first NIC 204 in a passive state for further investigation or repair. In the passive state, the first NIC 204 becomes hidden from the host system, meaning the host system no longer sees it in the device hierarchy, preventing the host from attempting to communicate with the failed device. The second NIC 207 now takes over all network traffic responsibility, seamlessly replacing the failed NIC without requiring a system reboot or manual intervention. This failover process ensures minimal disruption to system operation, allows for continuous network connectivity, and prevents expensive computing resources (e.g., GPUs) from being underutilized.

[0050] Figure 3 This is a block diagram illustrating various embodiments of a switching device 300 according to the present technology. This diagram is merely illustrative and should not unduly limit the scope of the claims. Those skilled in the art will recognize many variations, alternatives, and modifications.

[0051] In various implementations, the switch device 300 can be used for larger distributed systems (e.g., Figure 2 As part of system 200. Switch 300 can be configured to manage data routing between multiple endpoint devices (e.g., NICs, processors, or other peripherals) and external networks, ensuring seamless communication and high-speed data transmission. In some embodiments, switch device 300 plays a central role in implementing a failover mechanism that ensures continuous system operation even when some devices fail. This is achieved by monitoring the operational status of connected devices (e.g., NICs) and dynamically reconfiguring data paths when a failure is detected.

[0052] As shown, the switch device 300 may include one or more ports (e.g., a first port 301a, a second port 301b, a third port 301c, and a fourth port 301d). For example, the term "port" may refer to a physical or logical interface on the switch through which data is transmitted and received. A port acts as a connection point between endpoint devices (e.g., a NIC, a processor) and an external network, allowing data flow between these components. Examples of ports may include, but are not limited to, PCIe ports, Ethernet ports, InfiniBand ports, or other communication interfaces. Depending on the implementation, a port may be used as an upstream port or a downstream port. An upstream port connects the switch to an upstream component (e.g., a host system or a higher-level network), while a downstream port connects the switch to a downstream component (e.g., an endpoint device).

[0053] In various implementations, the switch device 300 may be implemented as a PCIe switch and may be coupled to one or more endpoint devices. The one or more endpoint devices may be connected via a PCIe interface. For example, the term "PCIe interface" may refer to a physical or logical connection that allows devices to communicate via the PCIe standard. In some embodiments, a first port 301a may be configured to couple to a first device 313a. A second port 301b may be configured to couple to a second device 313b. In some instances, the first device 313a may include a first NIC and the second device 313b may include a second NIC.

[0054] In some cases, a third port 301c may be configured to couple to a third device 313c, which may include a GPU or a storage device. For example, the term "storage device" may refer to a hardware component used for storing and retrieving data. Depending on the application, the storage device may be volatile or non-volatile and is responsible for temporarily or permanently retaining data. Examples of storage devices may include, but are not limited to, hard disk drives (HDDs), solid-state drives (SSDs), and / or the like.

[0055] In some instances, the fourth port 301d may be coupled to host 314. For example, the term "host" or "host system" may refer to a central computing system that manages and coordinates the operation of connected devices. Host 314 may be responsible for initiating data transfers to and from endpoint devices, allocating device addresses, or managing memory allocation. Host 314 may include, but is not limited to, a CPU, memory, I / O subsystems, and / or the like.

[0056] In some embodiments, the switching device 300 further includes one or more processing layers responsible for various stages of data processing, error detection, and protocol management when data flows through the switch. The one or more processing layers may include, but are not limited to, a SerDes layer 302, a physical layer 303, a multiplexer / demultiplexer layer 304, a data link layer 305, a transaction layer 306, and / or similar layers.

[0057] In some implementations, the SerDes layer 302 may include a serializer-deserializer circuit that converts parallel data into serial data for transmission over a high-speed communication link, and then converts the serial data back into parallel data for further processing. The SerDes layer 302 achieves high-speed data transmission by reducing the number of data lines required for communication, which helps maintain high data transfer rates between devices.

[0058] Following the SerDes conversion, data can move through physical layer 303, which handles the physical transmission of data across the communication medium, ensuring that signals are properly synchronized and transmitted with minimal loss. Multiplexer / demultiplexer layer 304 manages data flow by combining multiple data signals into a single stream (e.g., multiplexing) or separating a single data stream into multiple signals (e.g., demultiplexing). These processing layers achieve efficient use of the communication channel by dynamically managing available bandwidth and ensuring that data is transmitted to the appropriate endpoints.

[0059] In various embodiments, the data link layer 305 and transaction layer 306 handle higher-level communication protocols, ensuring that data packets are properly formatted, validated, and transmitted across the switching device 300. For example, the data link layer 305 provides error detection and correction mechanisms to ensure that data transmitted between devices is reliable and error-free. The transaction layer 306 manages the actual data transmission transactions between devices, determining how data is sent, received, and processed at each endpoint.

[0060] According to some embodiments, the switch device 300 may include a switch core 312. For example, the term "switch core" refers to the central processing unit of the switch that manages the overall data flow and controls how data is routed and processed within the switch. In various instances, the switch core 312 plays a crucial role in ensuring high availability and continuity of system operation when endpoint devices (e.g., NICs) fail. By continuously monitoring the status of connected devices and dynamically reassigning their roles (e.g., switching between active and passive states), the switch core 312 ensures uninterrupted data transmission even in the event of hardware or network problems. This process is beneficial in high-performance computing environments where the failure of a single component could lead to severe disruptions or underutilization of resources. For example, the switch core 312 may include at least one of a buffer 307, a routing unit 308, an arbitration unit 309, a scheduler 310, a controller 311, and / or the like.

[0061] In various implementations, switch core 312 includes controller 311. For example, the term "controller" can refer to a hardware component that manages the state of devices and data flow within the switch. Depending on the implementation, controller 308 may be implemented as a dedicated hardware module or as part of a software-defined network system.

[0062] In some instances, controller 311 can be configured to determine and assign the status of devices connected to the switch (e.g., first device 313a and second device 313b) and monitor the overall operation of the switch. For example, controller 311 can assign an active status to first device 313a (e.g., first NIC) and a passive status to second device 313b (e.g., second NIC). Under normal operating conditions, the active device handles all data transmissions, while the passive device remains idle but is ready to take over in case of a failure.

[0063] In various embodiments, scheduler 310 may be coupled to controller 311. For example, the term "scheduler" may refer to a component responsible for managing the timing and coordination of tasks within the system. Instances of schedulers may include, but are not limited to, polling schedulers, priority-based schedulers, credit-based schedulers, and / or the like. Scheduler 310 may be configured to manage the execution and sequencing of data transmission tasks, ensuring that resources are allocated effectively and devices operate synchronously. Depending on the implementation, scheduler 310 may be configured to coordinate data flow, manage task timing, and / or detect the operational status of endpoint devices.

[0064] In various instances, scheduler 310 may work in conjunction with controller 311 to monitor the operational status of a connected device (e.g., first device 313a). For example, scheduler 310 may be configured to monitor the operational status of first device 313a by detecting a first fault associated with the first device. For example, the term "fault" may refer to any event or condition in which a device fails to perform its intended function or suffers degraded performance. Faults may include, but are not limited to, hardware failures, network communication errors, configuration errors, loss of connectivity, and / or similar issues.

[0065] Depending on the implementation, fault detection can be implemented in various ways. For example, the scheduler 310 may detect a fault based on the loss of electrical connectivity between the device and the ports to which it is connected (e.g., first port 301a and first device 313a). This may involve detecting a sudden drop in signal strength or a complete loss of signal. In some instances, faults may be detected based on configuration space errors, where the configuration register of the first device 313a returns an invalid or corrupted value, indicating a fault.

[0066] In some cases, scheduler 310 can monitor the success rate of data transactions between the first device 313a and the rest of the system (e.g., the first port 301a). If the success rate drops below a predefined threshold, it can indicate that the first device 313a has encountered a problem. For example, frequent transmission errors, aborted transactions, or dropped packets can be signs of device failure. In various embodiments, scheduler 310 can be configured to monitor the operational status of the first device 313a based on predefined time intervals. For example, scheduler 310 can perform routine health checks on the first device 313a, such as querying the device's status updates, verifying data integrity, or testing communication responsiveness.

[0067] In various implementations, switch core 312 also includes routing unit 308. For example, the term "routing unit" can refer to a component responsible for determining the path taken by data within the switch, ensuring that data is directed to the appropriate device or network destination. Routing unit 308 manages data flow by allocating and updating routing paths between ports and connected devices based on the current state of the network and / or the operational state of the devices.

[0068] In some instances, routing unit 308 may be configured to determine a first routing path between first device 313a and first port 301a for managing data traffic. For example, the term "routing path" may refer to the communication route followed by data packets traveling between devices within a system. The routing path may be determined based on various factors, such as network topology, bandwidth availability, the operating state of the device (e.g., active or passive), and / or the like. For example, when first device 313a is in an active state, routing unit 308 facilitates data transactions by directing data traffic along the first routing path between first device 313a and first port 301a.

[0069] In some instances, the routing unit may utilize a routing table that stores information about available routes and the status of connected devices. For example, the term "routing table" may refer to a database or data structure that maintains a record of possible route paths for data transmission between devices and ports. In various embodiments, the routing table contains an entry for each connected device, specifying which port it is associated with, its current status (e.g., active or passive), and the route path for data to reach its destination. For example, the routing table may contain an entry storing a first route path between the first device 313a and the first port 301a, ensuring that data sent from the system is properly routed to the active NIC. The routing unit 308 may dynamically update the routing table in response to changes in network conditions, such as device failure or recovery.

[0070] When a fault is detected by scheduler 310, multiple components within switch device 300 work together to maintain system operation. For example, in response to scheduler 310 detecting a first fault, controller 311 is configured to reassign active state from first device 313a to second device 313b and passive state from second device 313b to first device 313a. Routing unit 308 can be configured to determine a second routing path between second device 313b and second port 301b and update the routing table to store the second routing path. This dynamic reallocation ensures uninterrupted continuous data flow, minimizes downtime, and maintains system reliability even in the event of device failure.

[0071] According to various embodiments, switch core 312 further includes buffer 307. For example, the term "buffer" may refer to a memory element or storage area used to temporarily store data. Buffer 307 is used to smooth data flow by accommodating differences in data transmission rates between different components or devices. For example, data arriving from a NIC or external network may arrive at a rate higher than the system can handle, so buffer 307 temporarily stores this data until the system is ready to process it or transmit it to its final destination. In some cases, buffer 307 temporarily stores data packets while routing unit 308 determines the routing path for forwarding the data to its destination. Buffer 307 also plays an important role in failover scenarios, where it stores data when switch core 312 reassigns states (e.g., from an active device to a passive device) and updates the routing table.

[0072] In some embodiments, the switch core 312 further includes an arbitration unit 309. For example, the term "arbitration unit" may refer to a component responsible for managing access to shared resources (e.g., data paths or communication channels). In various instances, when multiple devices connected to the switch device 300 simultaneously request access to the same resource, the arbitration unit 309 determines which device receives priority based on predefined rules or scheduling algorithms. This process ensures efficient data flow between devices and prevents resource contention or traffic bottlenecks. Examples of arbitration mechanisms include priority-based arbitration, round-robin arbitration, and weighted fair queuing.

[0073] While the foregoing is a complete description of specific embodiments, various modifications, alternative constructions, and equivalents may be used. Therefore, the foregoing description and illustrations should not be construed as limiting the scope of the technology as defined by the appended claims.

Claims

1. A switching device, comprising: The first port is coupled to the first device; The second port is coupled to the second device; A controller coupled to the first port, the controller being configured to assign an active state to the first device and a passive state to the second device; A scheduler, coupled to the controller, is configured to monitor the operational status of the first device by detecting a first fault associated with the first device; and A routing unit coupled to the controller, the routing unit being configured to determine a first routing path between the first device and the first port for managing data traffic, the routing unit including a routing table configured to store the first routing path; In response to the scheduler detecting the first fault, the controller is configured to reassign the active state from the first device to the second device and the passive state from the second device to the first device, and the routing unit is configured to determine a second routing path between the second device and the second port and update the routing table to store the second routing path.

2. The switching device according to claim 1, wherein the first device includes a first network interface card (NIC) and the second device includes a second NIC.

3. The switching device of claim 1, wherein the scheduler is configured to monitor the operational status of the first device based on a predefined time interval.

4. The switching device of claim 1, wherein the first fault is detected based on the loss of electrical connectivity between the first device and the first port.

5. The switching device of claim 1, wherein the first fault is detected based on an error in the configuration space of the first device.

6. The switching device of claim 1, wherein the first fault is detected based on the success rate of data transactions between the first device and the first port.

7. The switching device according to claim 1, further comprising a third port coupled to a third device.

8. The switching device of claim 7, wherein the first device is configured to perform a direct memory access (DMA) transfer to the third device.

9. The switching device of claim 7, wherein the third device includes a graphics processing unit (GPU).

10. The switching device according to claim 7, wherein the third device includes a storage device.

11. The switching device of claim 1, wherein the first device is coupled to the second device via a peripheral component interconnect high-speed PCIe interface.

12. The switching device of claim 1, further comprising a fourth port coupled to a host, and the controller being configured to communicate the active state of the first device to the host.

13. A switching device, comprising: The first port is coupled to the first device; The second port is coupled to the second device; A controller coupled to the first port, the controller being configured to assign an active state to the first device and a passive state to the second device; A scheduler, coupled to the controller, is configured to monitor the operational status of the first device by detecting a first fault associated with the first device; and A routing unit coupled to the controller, the routing unit being configured to determine a first routing path between the first device and the first port for managing data traffic; In response to the scheduler detecting the first fault, the controller is configured to reassign the active state from the first device to the second device and the passive state from the second device to the first device, and the routing unit is configured to determine a second routing path between the second device and the second port.

14. The switching device according to claim 13, wherein the first device includes a first network interface card (NIC).

15. The switching device of claim 13, wherein the first fault is detected based on the loss of electrical connectivity between the first device and the first port.

16. The switching device of claim 13, wherein the first fault is detected based on an error in the configuration space of the first device.

17. The switching device of claim 13, wherein the first fault is detected based on the success rate of data transactions between the first device and the first port.

18. A method comprising: The controller assigns an active state to a first device coupled to a first port and a passive state to a second device coupled to a second port. The operating status of the first device is monitored by the scheduler; The routing unit determines a first routing path for managing data traffic between the first device and the first port; In response to detecting a first fault associated with the first device, the active state is reassigned to the second device, and the passive state is reassigned to the first device; and The routing unit determines a second routing path between the second device and the second port for managing data traffic.

19. The method of claim 18, wherein the first device includes a first network interface card (NIC) and the second device includes a second NIC.

20. The method of claim 18, wherein the first fault is detected based on the loss of electrical connectivity between the first device and the first port.