Apparatus and method for rapid recovery from failures of some devices in optical data switching or computing device cluster

By employing N:1 backup protection optical switch equipment in optical data exchange or computing equipment clusters, and utilizing optical couplers and 1×N optical switches, rapid switching to redundant equipment is achieved in the event of equipment failure. This solves the problem of time-consuming and labor-intensive fault recovery in existing technologies, and improves the utilization efficiency and recovery speed of the equipment.

WO2026017154A1PCT designated stage Publication Date: 2026-01-22ACCELINK TECHNOLOGIES CO LTD

Patent Information

Application Number
PCT/CN2025/109329
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-18
Filing Date
2025-07-18
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

In existing technologies, when optical data exchange or computing equipment clusters fail, it is necessary to manually replace the fiber optic connection, which results in huge manpower and time costs for fault recovery. In addition, existing monitoring methods are highly complex and difficult to restore equipment functionality quickly.

Method used

The protective optical switch equipment adopts N:1 backup, which enables rapid switching to redundant equipment in case of equipment failure through optical couplers and 1×N optical switches. The management equipment centrally controls the optical switches and optical couplers, enabling rapid replication of optical signals and replacement of redundant equipment.

Benefits of technology

It enables rapid recovery of optical data exchange or computing equipment clusters in the event of equipment failure, saving recovery time, conserving cluster equipment resources, and improving equipment utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025109329_22012026_PF_FP_ABST
    Figure CN2025109329_22012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data communications, and provides an apparatus and method for rapid recovery from failures of some devices in an optical data switching or computing device cluster. A management device confirms the identifier m of a second-layer device that sends an alarm message; on the basis of a pre-established mapping relationship between the identifier m of the second-layer device and an optical path port Pm in each protection optical switch device, the management device controls an optical switch in each protection optical switch device in the protection optical switch array, such that a common port thereof is connected to the optical path port Pm, thereby replicating an optical signal, imported from a first-layer device library into the second-layer device m, to a second-layer redundant device by means of the optical switch in each protection optical switch device and an optical coupler m. In the present invention, when some devices fail, functions of the devices can be rapidly switched to a redundant device, thereby reducing failure recovery time and saving cluster device resources.
Need to check novelty before this filing date? Find Prior Art

Description

A device and method for rapid recovery from partial equipment failure in an optical data switching or computing equipment cluster.

[0001] Cross-reference of related applications

[0002] This application claims priority to the following patent application:

[0003] (1) A Chinese patent application filed on July 18, 2024, with application number 202410963185.X, entitled “A device and method for rapid recovery of partial equipment failure in an optical data exchange or computing equipment cluster”. Technical Field

[0004] This invention relates to the field of data communication technology, and in particular to a device and method for rapid recovery of partial equipment failure in an optical data exchange or computing equipment cluster. Background Technology

[0005] Optical data switching or computing equipment has several optical ports. The optical signal received from the optical fiber at each port is converted into an electrical signal by a photoelectric converter, processed internally as a digital signal, and then forwarded to the corresponding target port. The electrical signal is then converted back into an optical signal by an electro-optical converter and transmitted out through the optical fiber. This entire process enables high-speed, reliable data transmission and is widely used in data centers, supercomputers, and other systems.

[0006] In data centers, optical data switching or computing equipment often appears in the form of clusters, with several devices divided into multiple functional levels. The characteristic is that devices at the same level have the same functions, number of ports, and similar connection methods.

[0007] Taking the Spine-Leaf architecture as an example, this data center network topology consists of two data exchange layers: Spine and Leaf. The Leaf layer comprises access switches that aggregate traffic from servers and connect directly to the Spine or network core. Spine switches interconnect all Leaf switches in the full mesh topology. Each layer has several functionally identical data exchange devices with the same model, function, and number of ports, connected according to a specific pattern, exhibiting high symmetry. Because the data exchange devices have a large number of ports, a failure in the device itself or some of its ports can reduce the availability of the entire cluster, causing data exchange congestion.

[0008] In emerging artificial intelligence (AI) supercomputer clusters, computing devices are often connected to data exchange equipment in a clustered manner. Taking a 100-node DGX SuperPOD as an example, each superPOD includes 20 A100 GPUs, and each A100 GPU includes 8 IB Compute Connectors, connected to the SuperPOD's 8 Leaf computer switches. All GPUs within the same SuperPOD, or even larger-scale GPUs across multiple SuperPODs, complete computational tasks in parallel. The failure of one GPU can cause the entire cluster's computational tasks to be interrupted and rolled back, wasting a significant amount of computing power.

[0009] Therefore, overcoming the shortcomings of the existing technology is an urgent problem to be solved in this technical field.

[0010] Application content

[0011] The technical problem to be solved by this invention is how to reduce the structural complexity of array optical switch monitoring and how to monitor the state of array optical switches without customer signal light input.

[0012] The present invention adopts the following technical solution:

[0013] In a first aspect, the present invention provides a rapid recovery device for partial equipment failure in an optical data exchange or computing equipment cluster, the rapid recovery device comprising one or more protective optical switching devices, including:

[0014] The protective optical switch device has N optical path ports on one end, namely optical path port Q1, optical path port Q2, ..., optical path port QN, and N+1 ports on the other side, namely optical path port P1, optical path port P2, ..., optical path port PN, and optical path port PC.

[0015] Optical port Qi and optical port Pi are connected in series to the common port and the first port of optical coupler i, respectively. The second port of optical coupler i is connected to optical port PC via a 1×N optical switch. The i-th port of the 1×N optical switch is coupled to the second port of optical coupler i, and the common port of the 1×N optical switch is coupled to optical port PC, i∈[1,N].

[0016] When the 1×N optical switch switches to its own k-th port and its own common port, the optical paths of optical path port Pk and optical path port PC are both connected to optical path port Qk.

[0017] Preferably, an optical amplifier is connected in series between the common port of the 1×N optical switch and the optical path port PC.

[0018] Preferably, the rapid recovery device includes one or more protective optical switch devices, specifically including protective optical switch device 1, protective optical switch device 2, ..., protective optical switch device M, wherein the optical switch in each protective optical switch device is uniformly managed by a management device;

[0019] The output of the first-level device library includes at least M sets of transmission ports, and each set of transmission ports includes at most N transmission ports;

[0020] Each transmission port is carried by optical path ports Q1-QN of a protection optical switch device, while the optical path ports P1-PN of each protection optical switch device are respectively assigned to Layer 2 device 1-Layer 2 device N;

[0021] Each protective optical switch device has an optical path port PC connected to a redundant device, which has the functional attributes of each second-layer device; each second-layer device and the redundant device also have a data control signal interaction link with the management device.

[0022] Preferably, the management device is used to acquire alarm messages originating from any one of the second-layer devices 1 to N, and further includes:

[0023] The management device confirms the identifier m of the second-layer device that sent the alarm message;

[0024] Based on the pre-established mapping relationship between the identifier m of the second-layer device and the optical path port Pm in each protective optical switch device, the management device controls the optical switch in each protective optical switch device to make its common port and optical path port Pm conduct, thereby replicating the optical signal imported from the first-layer device library into the second-layer device m to the second-layer redundant device through the optical switch and optical coupler m in each protective optical switch device.

[0025] The second-layer redundant device replaces the function of the second-layer device m and continues to interact with the first-layer device library.

[0026] Preferred options also include:

[0027] When any of the second-layer devices x from second-layer device 1 to second-layer device N confirms that its communication with the first-layer device library is in a temporarily followable state, the second-layer device x sends an optical path detection request to the management device.

[0028] The management device confirms whether the current second-layer redundant device is in a replacement state. If the second-layer redundant device is in a replacement state, it returns a response message indicating that the second-layer redundant device is in a replacement state to the second-layer device x.

[0029] If the second-layer redundant device is in an idle state, a response message for the second-layer redundant device to enter temporary follow is returned. At the same time as sending the response message, the common port in each protection optical switch device is controlled to be turned on with the optical path port Qx assigned to the second-layer device x in the optical switch device. Thus, by comparing the optical signals interacted between the second-layer redundant device and the second-layer device x, the management device confirms the stability of the state of the second-layer device x. Here, x belongs to [1, N].

[0030] Preferably, the temporary followable state specifically includes:

[0031] The first-layer device library and the second-layer devices are in a heartbeat-linked connection state; or...

[0032] The first-layer device library and the second-layer devices are in a state of continuous data transmission in an unencrypted state; or,

[0033] The first-layer device library and the second-layer devices are in an inherent script execution state, wherein the execution instructions of the script are known in advance by the management device.

[0034] Preferably, the step of confirming the state stability of the second-layer device x by comparing the optical signals interacted between the second-layer redundant device and the second-layer device x further includes:

[0035] The management equipment controls the operating power of the optical amplifiers in each protective optical switch device. With the goal of having the same optical power received by the second-layer device x and the second-layer redundant device, it confirms whether the adjustment of the operating power of the optical amplifiers in each protective optical switch device is within the preset range, thereby confirming the stability of the working state of each optical path port.

[0036] Preferably, the fast recovery device includes one or more protective optical switch devices, specifically including a north-facing protective optical switch array composed of U protective optical switch devices and a south-facing protective optical switch array composed of V protective optical switch devices; N third-layer devices and one third-layer redundant device are connected in series between the north-facing protective optical switch array and the south-facing protective optical switch array, wherein each third-layer device has V optical interfaces facing the south-facing protective optical switch array and U optical interfaces facing the north-facing protective optical switch array;

[0037] The optical interfaces on the other side of the northbound protection optical switch array and the southbound protection optical switch array, in addition to the optical interface coupled to the third-layer device, are used to connect to the northbound device library and the southbound device library, respectively.

[0038] In this system, the optical switches in each protective optical switch device are managed uniformly by the management device, and the corresponding third-layer devices and third-layer redundant devices also establish data interaction channels with the management device.

[0039] Preferably, the second-layer equipment specifically comprises computer rack 1, computer rack 2, ..., computer rack 8; the second-layer redundant equipment specifically comprises redundant computer racks; the first-layer equipment library specifically comprises 8 leaf computing switches; coupled with 32 backup optical switches, wherein N is 8, and the fast recovery device includes:

[0040] The j'th port of the k'th H100 of the redundant computer rack is connected to the PC port of the 4k'+j'-4 backup optical switch; wherein, each H100 includes a set of 8 optical interfaces, numbered from 1 to 8;

[0041] The j' port of the k'th H100 of the i'th computer rack is connected to the Pi' port of the 4k'+j'-4th backup optical switch;

[0042] The Qi' port of the 4k'+j'-4th backup optical switch is connected to the 4i'+k-4th port of the j'th leaf calculation switch; where k' and j' are natural numbers.

[0043] Secondly, the present invention also provides a method for rapid recovery of partial equipment failure in an optical data exchange or computing equipment cluster. The rapid recovery device includes protective optical switch device 1, protective optical switch device 2, ..., protective optical switch device M, wherein the optical switches in each protective optical switch device are uniformly managed by a management device; the output end of the first-layer device library includes at least M sets of transmission ports, and each set of transmission ports includes at most N transmission ports; each set of transmission ports is carried by optical path ports Q1 to QN of one protective optical switch device, while the corresponding optical path ports P1 to PN of each protective optical switch device are respectively assigned to second-layer devices 1 to N; the method includes:

[0044] The management device is used to acquire alarm messages originating from any of the second-layer devices, from second-layer device 1 to second-layer device N.

[0045] The management device confirms the identifier m of the second-layer device that sent the alarm message;

[0046] Based on the pre-established mapping relationship between the identification number m of the second-layer device and the optical path port Pm in each protective optical switch device, the management device controls the optical switch in each protective optical switch device to make its common port and optical path port Pm conduct, thereby replicating the optical signal imported from the first-layer device library into the second-layer device m to the second-layer redundant device through the optical switch and optical coupler m in each protective optical switch device.

[0047] The second-layer redundant device replaces the function of the second-layer device m and continues to interact with the first-layer device library.

[0048] Preferably, the method further includes:

[0049] When any of the second-layer devices x from second-layer device 1 to second-layer device N confirms that its communication with the first-layer device library is in a temporarily followable state, the second-layer device x sends an optical path detection request to the management device.

[0050] The management device confirms whether the current second-layer redundant device is in a replacement state. If the second-layer redundant device is in a replacement state, it returns a response message indicating that the second-layer redundant device is in a replacement state to the second-layer device x.

[0051] If the second-layer redundant device is in an idle state, a response message is returned to the second-layer redundant device to enter temporary follow-up. At the same time as sending the response message, the common port in each protection optical switch device is controlled to be turned on with the optical path port Qx assigned to the second-layer device x in the optical switch device. In this way, by comparing the optical signals interacted between the second-layer redundant device and the second-layer device x, the management device confirms the stability of the state of the second-layer device x.

[0052] Preferably, the temporary followable state specifically includes:

[0053] The first-layer device library and the second-layer devices are in a heartbeat-linked connection state; or...

[0054] The first-layer device library and the second-layer devices are in a state of continuous data transmission in an unencrypted state; or,

[0055] The first-layer device library and the second-layer devices are in an inherent script execution state, wherein the execution instructions of the script are known in advance by the management device.

[0056] Preferably, the step of confirming the state stability of the second-layer device x by comparing the interactive optical signals between the second-layer redundant device and the second-layer device x further includes:

[0057] The management equipment controls the operating power of the optical amplifiers in each protective optical switch device. With the goal of having the same optical power received by the second-layer device x and the second-layer redundant device, it confirms whether the adjustment of the operating power of the optical amplifiers in each protective optical switch device is within the preset range, thereby confirming the stability of the working state of each optical path port.

[0058] Compared with the prior art, the beneficial effects of the embodiments of the present invention are as follows:

[0059] This invention implements N:1 backup for devices at the same level in optical data switching or computing clusters. This allows for rapid switching of functionality to redundant devices when some devices fail, saving recovery time and conserving cluster resources. In contrast, existing technologies typically require manual replacement of fiber optic connections to replace faulty devices or ports when optical data switching or computing clusters fail. This not only reduces the utilization efficiency of the equipment but also incurs significant manpower and time costs for recovery. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0061] Figure 1 is a schematic diagram of the backup optical switching device in a rapid recovery device for partial equipment failure in an optical data exchange or computing equipment cluster provided in an embodiment of the present invention.

[0062] Figure 2 is a schematic diagram of the backup optical switch with an optical amplifier in a rapid recovery device for partial equipment failure in an optical data exchange or computing equipment cluster provided in an embodiment of the present invention.

[0063] Figure 3 is a schematic diagram of the architecture of a rapid recovery device for partial equipment failure in an optical data exchange or computing equipment cluster provided in an embodiment of the present invention;

[0064] Figure 4 is a schematic diagram of the architecture of another rapid recovery device for partial equipment failure in an optical data exchange or computing equipment cluster provided by an embodiment of the present invention;

[0065] Figure 5 is a schematic diagram of the architecture of a rapid recovery device for partial equipment failure in an optical data exchange or computing equipment cluster provided in an embodiment of the present invention.

[0066] Figure 6 is a schematic diagram of the architecture of a rapid recovery device for partial equipment failure in an optical data exchange or computing equipment cluster provided in an embodiment of the present invention.

[0067] Figure 7 is a schematic diagram of the backup optical switch with multiple inputs and multiple outputs in a rapid recovery device for partial equipment failure in an optical data exchange or computing equipment cluster provided in an embodiment of the present invention.

[0068] Figure 8 is a flowchart illustrating a method for rapid recovery of partial equipment failure in an optical data exchange or computing equipment cluster provided by an embodiment of the present invention.

[0069] Figure 9 is a flowchart illustrating another method for rapid recovery of partial equipment failure in an optical data exchange or computing equipment cluster provided by an embodiment of the present invention. Detailed Implementation

[0070] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0071] In the description of this invention, the terms "inner", "outer", "longitudinal", "lateral", "upper", "lower", "top", "bottom", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and do not require that this invention must be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.

[0072] In this invention, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0073] In various embodiments of the present invention, the applicable "first layer" and "second layer" can represent the concepts of "upper layer" and "lower layer" respectively in a specific architecture, or the concepts of "lower layer" and "upper layer" respectively (the meaning here is that the location of the redundant device in various embodiments of the present invention can be set to the opposite layer in a mirror-like manner); and in other specific architectures, they can also represent the concepts of "northbound" and "southbound" respectively, or the concepts of "southbound" and "northbound" respectively. Based on the core technical concept of the present invention, other architectural description methods can also be applied, and no further limitations are made here.

[0074] In this application, unless otherwise expressly specified and limited, the term "connection" should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral part; it can be a direct connection or an indirect connection through an intermediate medium. Furthermore, the term "coupled" can refer to an electrical connection that enables signal transmission.

[0075] This invention proposes an N:1 backup for optical data switching or computing clusters at the same level. Specifically, it involves adding a redundant device of the same model to N parallel optical data switching or computing devices. The optical paths connected to ports at the same location on each device are merged with the backup path via a splitter. The backup path connecting a port at the same location on the N parallel devices is connected to the corresponding port on the redundant device via a 1:N optical switch. When one of the N parallel optical data switching or computing devices, such as the Nth device, fails, the management device detects the failure and transfers the status and data of that device to the redundant device. The failed device's port is disabled, and the status of the backup optical switch connected to all ports is switched to port N. Thus, the functions that would normally require the Nth optical data switching or computing device are transferred to the redundant device.

[0076] Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0077] Example 1:

[0078] This invention provides a rapid recovery device for partial equipment failure in an optical data exchange or computing equipment cluster. The rapid recovery device includes one or more protective optical switching devices, as shown in Figure 1, including:

[0079] The protective optical switch device has N optical path ports on one end, namely optical path port Q1, optical path port Q2, ..., optical path port QN, and N+1 ports on the other side, namely optical path port P1, optical path port P2, ..., optical path port PN, and optical path port PC.

[0080] Optical port Qi and optical port Pi are connected in series to the common port and the first port of optical coupler i, respectively. The second port of optical coupler i is connected to optical port PC via a 1×N optical switch. The i-th port of the 1×N optical switch is coupled to the second port of optical coupler i, and the common port of the 1×N optical switch is coupled to optical port PC, i∈[1,N].

[0081] The splitting ratio of the optical coupler can be arbitrary. Since an additional optical switch introduces insertion loss on the optical path from PC to QN, to ensure that the insertion loss on the optical path from PN to QN is similar to that on the optical path from PC to QN, in one embodiment, the optical coupler employs an asymmetric splitting ratio. For example, the PC path splits 60%, and the PN path splits 40%, so the insertion loss on the optical path from PC to QN, after adding the optical switch, is similar to that on the optical path from PN to QN.

[0082] When the 1×N optical switch switches to its own k-th port and its own common port, the optical paths of optical path port Pk and optical path port PC are both connected to optical path port Qk.

[0083] This patent embodiment proposes a protective optical switch device that can perform N:1 backup of devices at the same level in optical data exchange or computing clusters. This allows the function of the device to be quickly switched to the redundant device when some devices fail, saving fault recovery time and conserving cluster device resources.

[0084] In contrast, when optical data exchange or computing clusters fail, existing technologies typically require manual replacement of fiber optic connections to replace faulty equipment or ports. This not only reduces the utilization efficiency of the equipment itself, but also consumes significant human resources and time for fault recovery.

[0085] The insertion of an optical coupler into the optical path introduces at least 3dB of insertion loss, potentially exceeding the power budget of the existing optical module. This can be addressed by increasing the splitting ratio of the PN to QN optical path and adding an optical amplifier at the PC port to compensate for the power loss from the PC to the QN port. For example, the splitting ratio of the optical coupler in the PN to QN optical path can be asymmetrically distributed, minimizing the loss between PN and QN. The larger insertion loss between PC and QN is compensated by the optical amplifier, while adding an optical amplifier at the PC port compensates for the additional losses caused by the low splitting ratio of the optical coupler and the additional losses from the optical switch. This ensures that the insertion of a backup optical switch does not significantly impact the power budget of the optical modules interconnecting different levels of optical data exchange and computing devices.

[0086] As shown in Figure 2, in conjunction with an embodiment of the present invention, another improved backup optical switch device is provided, wherein an optical amplifier is connected in series between the common port of the 1×N optical switch and the optical path port PC. The function of this optical amplifier is not simply to amplify one of the optical signals originating from optical path ports Q1-QN split by the optical coupler. The key usage of the optical amplifier in the improved structure shown in Figure 2 will be further elaborated in the subsequent extended embodiments of the present invention.

[0087] As shown in Figure 3, the rapid recovery device includes one or more protective optical switch devices, specifically including protective optical switch device 1, protective optical switch device 2, ..., protective optical switch device M, wherein the optical switch in each protective optical switch device is uniformly managed by the management device.

[0088] The output of the first-level device library includes at least M sets of transmission ports, and each set of transmission ports includes at most N transmission ports.

[0089] Each transmission port is carried by optical path ports Q1-QN of a protection optical switch, while the corresponding optical path ports P1-PN of each protection optical switch are respectively assigned to Layer 2 devices 1-N.

[0090] Each protective optical switch device has an optical path port PC connected to a redundant device, which has the functional attributes of each second-layer device; each second-layer device and the redundant device also have a data control signal interaction link with the management device.

[0091] For a network topology consisting of N identical optical data switching or computing devices, an additional redundant optical data switching or computing device can be added, using a set of N:1 backup optical switches for redundancy. Assuming this type of device (i.e., the second-layer device in Figure 3) has M optical ports, then M N-port protection optical switches or a backup optical switch ratio of no less than N:1 are required. The m-th optical port of the redundant device among the N parallel devices is connected to the PC port of the m-th N:1 backup optical switch; the M-th optical port (where m ∈ natural numbers 1 to N) of the Nth device (where N ∈ natural numbers 1 to M) is connected to the PN port of the m-th N:1 backup optical switch; the QN port of the m-th N:1 backup optical switch is connected to the k-th optical port (where k ∈ natural numbers 1 to K, K = M x N) of another layer device that should originally be connected to the m-th port of the Nth device. It should be emphasized that the K optical ports of the other layer are not necessarily located on the same device. Here, it only refers to the total number of optical ports of other layer devices connected to the N devices to be protected, which together have K = M x N optical ports. The specific allocation method is unrelated to the N devices in this layer and the M N:1 backup optical switches.

[0092] When all N Layer 2 devices are functioning normally, the redundant device ports are closed, and all backup optical switches are in any state. When one of the N Layer 2 devices fails, the management device transfers the data and status of the failed device to the redundant device. Let the failed device be the Nth device. All M N:1 backup optical switches are switched to state N, and the optical port of the Nth failed device is closed, thus replacing the Nth failed device with the redundant device.

[0093] In a basic usage scenario, referring to the fast recovery device architecture shown in Figure 2 (it should be emphasized that the following methods are not limited to the fast recovery device architecture shown in Figure 2, but are also applicable to the fast recovery device architectures shown in Figures 4 and 6 below in feasible scenarios), the management device is used to obtain alarm messages originating from any of the second-layer devices 1 to N, and further includes:

[0094] The management device confirms the identifier m of the second-layer device that sent the alarm message; based on the pre-established mapping relationship between the identifier m of the second-layer device and the optical path port Pm in each protective optical switch, the management device controls the optical switches in each protective optical switch to make their common port and optical path port Pm connected, thereby replicating the optical signal imported from the first-layer device library into the second-layer device m to the second-layer redundant device via the optical switches and optical couplers m in each protective optical switch; the second-layer redundant device replaces the function of the second-layer device m and continues to interact with the first-layer device library.

[0095] The following process, based on the protective optical switch device with an improved optical amplifier structure as shown in Figure 2, can be further illustrated when applied to a fast recovery device architecture similar to that described in Figure 3. Further examples of this process include:

[0096] Any Layer 2 device x among Layer 2 devices 1 to N, upon confirming that its communication with the Layer 1 device library is in a temporarily followable state, sends an optical path detection request to the management device. Specifically, the temporarily followable state includes: a heartbeat message connection between the Layer 1 device library and the Layer 2 devices; or, continuous data transmission in an unencrypted state between the Layer 1 device library and the Layer 2 devices; or, an inherent script execution state between the Layer 1 device library and the Layer 2 devices, wherein the script execution instructions are known in advance by the management device. The temporarily followable state signifies that the communication content between the Layer 1 device library and the Layer 2 devices is predictable, allowing the corresponding redundant Layer 2 devices to match and verify with Layer 2 device x through the parsed content even when temporarily accessing the Layer 1 device library.

[0097] The management device confirms whether the current Layer 2 redundant device is in a replacement state. If the Layer 2 redundant device is in a replacement state, it returns a response message indicating that the Layer 2 redundant device is in a replacement state to the Layer 2 device x. The replacement state indicates that a Layer 2 device in the current system architecture has failed, or that the optical path connection between the corresponding Layer 2 device and the Layer 1 device library has encountered a problem. Furthermore, the management device has controlled the associated protective optical switch to switch the optical path link between the problematic Layer 2 device and the Layer 1 device library to the optical path link between the Layer 2 redundant device and the Layer 1 device library. Therefore, the replacement state here can also be understood as the Layer 2 redundant device replacing a failed Layer 2 device.

[0098] If the second-layer redundant device is in an idle state, a response message for the second-layer redundant device to enter temporary follow is returned. At the same time as sending the response message, the common port in each protection optical switch device is controlled to be turned on with the optical path port Qx assigned to the second-layer device x in the optical switch device. Thus, by comparing the optical signals interacted between the second-layer redundant device and the second-layer device x, the management device confirms the stability of the state of the second-layer device x. Here, x belongs to [1, N].

[0099] The aforementioned method of comparing the optical signals of the second-layer redundant device with those of the second-layer device x, allowing the management device to verify the stability of the inherent optical path of the second-layer device x, can be implemented in the following way:

[0100] The management equipment controls the operating power of the optical amplifiers in each protective optical switch. With the goal of ensuring the same optical power received by the second-layer device x and the second-layer redundant device, it verifies whether the adjustment of the operating power of the optical amplifiers in each protective optical switch is within a preset range, thereby confirming the stability of the operating status of each optical path port. In practical implementation, the above method is not solely for confirming the stability of the operating status of each optical path port; it also includes handling various faults, such as those requiring a switch to backup. As an optional solution, the implementation can also incorporate error detection and alarm mechanisms, going beyond just detecting operating power.

[0101] As shown in Figure 3 above, the fast recovery device architecture of this invention also provides an example of the fast recovery device architecture as shown in Figure 4. The fast recovery device includes one or more protective optical switch devices, specifically a north-facing protective optical switch array composed of U protective optical switch devices and a south-facing protective optical switch array composed of V protective optical switch devices. N third-layer devices and one third-layer redundant device are connected in series between the north-facing and south-facing protective optical switch arrays. Each third-layer device has V optical interfaces facing the south-facing protective optical switch array and U optical interfaces facing the north-facing protective optical switch array. Here, the first-layer device library of this invention is also described as the north-facing device of the third-layer device, and the corresponding second-layer device is also described as the south-facing device of the third-layer device. The protective optical switch array of the first-layer device library is also described as N:1 north-facing protective switch devices × U of the third-layer device; the protective optical switch array of the second-layer device is also described as N:1 south-facing protective switch devices × V of the third-layer device.

[0102] The northbound and southbound protection optical switch arrays each have an optical interface on one side, in addition to the optical interface coupled to the Layer 3 devices, used to connect to the northbound and southbound device libraries, respectively. The optical switches in each protection optical switch device are managed uniformly by the Layer 3 device management equipment, and each Layer 3 device and its redundant components also establish data exchange channels with the management equipment. In practical implementation, the terms "northbound" and "southbound" are more often used to refer to the northbound or southbound direction of a specific Layer 3 device.

[0103] In an optical data switching or computing equipment cluster, devices at a certain layer often have both northbound and southbound connection ports. When introducing redundant equipment through backup optical switches, backup optical switches for both the northbound and southbound connection ports need to be configured simultaneously. As shown in Figure 4, this layer has N optical data switching or computing devices, each with U northbound optical ports and V southbound optical ports. A redundant device is configured with U backup optical switches for the northbound ports (at least N:1 ratio) and V backup optical switches for the southbound ports (at least N:1 ratio). When the Nth device fails, the N:1 backup optical switches for both the northbound and southbound ports simultaneously switch to the Nth port to replace the failed device.

[0104] Besides the fast recovery device architecture example shown in Figure 4, in an optical data switching or computing equipment cluster, devices at different layers can independently choose whether to configure redundant devices and their backup optical switches. Figure 5 shows that the first layer has N devices, each with V southbound optical ports, and the second layer has L devices, each with U northbound ports, where N×V=L×U. The N devices in the first layer are configured with one redundant device and V southbound backup optical switches with a port ratio of at least N:1, providing a fault protection mechanism for the N devices in the first layer. The L devices in the second layer are configured with one redundant device and U northbound backup optical switches with a port ratio of at least L:1, providing a fault protection mechanism for the L devices in the second layer. The fault protection mechanisms of the two adjacent layers can coexist independently or exist separately as needed.

[0105] As shown in Figure 3 above, the embodiment of the present invention also provides an example of a fast recovery device architecture as shown in Figure 6 (i.e., a 256-card DGX H100 Super POD 8:1 backup solution). The second-layer devices specifically include computer rack 1, computer rack 2, ..., computer rack 8; the second-layer redundant devices specifically include redundant computer racks; the first-layer device library specifically includes 8 leaf computing switches; and is equipped with 32 backup optical switches, where N is 8. The fast recovery device includes:

[0106] The j'th port of the k'th H100 of the redundant computer rack is connected to the PC port of the 4k'+j'-4 backup optical switch; wherein, each H100 includes a set of 8 optical interfaces, numbered from 1 to 8;

[0107] The j' port of the k'th H100 of the i'th computer rack is connected to the Pi' port of the 4k'+j'-4th backup optical switch;

[0108] The Qi' port of the 4k'+j'-4th backup optical switch is connected to the 4i'+k-4th port of the j'th leaf calculation switch; where k' and j' are natural numbers.

[0109] Figure 7 shows a possible improvement scheme proposed by this invention based on the 1:N backup optical switch device of embodiment 1, namely, a functional diagram of the N:S backup optical switch with multiple backups.

[0110] In some clusters, an N:1 backup may not meet the failure rate requirements, so multi-path backup can be used, such as S backup ports backing up N ports. As shown in Figure 7, an NxS optical switch can be used to achieve S optical ports backing up N optical ports, where the S ports of the optical switch can be independently configured to connect to any N ports.

[0111] Example 2:

[0112] This invention also provides a method for rapid recovery of partial equipment failure in an optical data exchange or computing equipment cluster. The rapid recovery device includes protective optical switch device 1, protective optical switch device 2, ..., protective optical switch device M, wherein the optical switches in each protective optical switch device are uniformly managed by a management device; the output end of the first-layer device library includes at least M sets of transmission ports, and each set of transmission ports includes at most N transmission ports; each set of transmission ports is carried by optical path ports Q1-QN of one protective optical switch device, while the corresponding optical path ports P1-PN of each protective optical switch device are respectively assigned to second-layer device 1-second-layer device N; as shown in Figure 8, the method includes:

[0113] In step 201, the management device is used to acquire alarm messages originating from any of the second-layer devices, from second-layer device 1 to second-layer device N.

[0114] In step 202, the management device confirms the identifier m of the second-layer device that sent the alarm message. The second-layer device corresponding to the identifier m is referred to as the second-layer device m.

[0115] In step 203, based on the pre-established mapping relationship between the identifier m of the second-layer device and the optical path port Pm in each protective optical switch device, the management device controls the optical switch in each protective optical switch device to make its common port and optical path port Pm conduct, thereby replicating the optical signal imported from the first-layer device library into the second-layer device m into the second-layer redundant device through the optical switch and optical coupler m in each protective optical switch device.

[0116] In step 204, the second-layer redundant device replaces the function of the second-layer device m and continues to interact with the first-layer device library.

[0117] This patent implements N:1 backup for devices at the same level in optical data exchange or computing clusters. This allows for rapid switching of functionality to redundant devices when some devices fail, saving recovery time and conserving cluster resources. In contrast, existing technologies typically require manual replacement of fiber optic connections to replace faulty devices or ports when optical data exchange or computing clusters fail. This not only reduces the utilization efficiency of the equipment but also incurs significant manpower and time costs for recovery.

[0118] As shown in Figure 9, in conjunction with the embodiments of the present invention, there is also a preferred extended implementation scheme, the method of which includes:

[0119] In step 301, any second-layer device x among second-layer device 1 to second-layer device N, upon confirming that its communication with the first-layer device library is in a temporarily followable state, sends an optical path detection request to the management device.

[0120] The temporary followable state here means that the communication content between the first-layer device library and the second-layer device can be predicted in advance. Therefore, even when the corresponding second-layer redundant device is temporarily connected to the first-layer device library, it can still be matched and verified with the second-layer device x through the parsed content.

[0121] In step 302, the management device confirms whether the current second-layer redundant device is in a replacement state. If the second-layer redundant device is in a replacement state, it returns a response message indicating that the second-layer redundant device is in a replacement state to the second-layer device x.

[0122] The "alternate state" indicates that a certain Layer 2 device in the current system architecture has failed, or that the optical path connection between the corresponding Layer 2 device and the Layer 1 device library has encountered a problem. Furthermore, the management device has controlled the corresponding protective optical switch to switch the optical path link between the faulty Layer 2 device and the Layer 1 device library to the optical path link between the redundant Layer 2 device and the Layer 1 device library. Therefore, the "alternate state" here can also be understood as the Layer 2 redundant device being in a state of replacing a faulty Layer 2 device.

[0123] Specifically, the temporary followable state includes: the first layer device library and the second layer device are in a heartbeat message connection state; or, the first layer device library and the second layer device are in a continuous data transmission state in an unencrypted state; or, the first layer device library and the second layer device are in an inherent script execution state, wherein the execution instructions of the script are known in advance by the management device.

[0124] In step 303, if the second-layer redundant device is in an idle state, a response message for the second-layer redundant device to enter temporary follow is returned. At the same time as sending the response message, the common port in each protection optical switch device is controlled to be turned on with the optical path port Qx assigned to the second-layer device x in the optical switch device. Thus, by comparing the optical signals interacted between the second-layer redundant device and the second-layer device x, the management device confirms the stability of the state of the second-layer device x.

[0125] The step 303 above, which involves comparing the optical signals interacting between the second-layer redundant device and the second-layer device x to confirm the state stability of the second-layer device x, is further refined in the following specific implementations of this invention:

[0126] The management device controls the operating power of the optical amplifiers in each protective optical switch. With the objective of ensuring that the optical power received by the second-layer device x and the second-layer redundant device are the same, it confirms whether the adjustment of the operating power of the optical amplifiers in each protective optical switch is within a preset range, thereby confirming the stability of the operating state of each optical path port. This forms a method-level closed loop with the protective optical switch containing optical amplifiers as shown in Figure 2 of Embodiment 1. In specific implementation, the above method is not necessarily solely for confirming the stability of the operating state of each optical path port; it also includes handling various faults, such as those requiring switching to backup. As an optional solution, the implementation process can also be combined with error detection, alarms, etc., not limited to the detection of operating power.

[0127] It is worth noting that the information interaction and execution process between the modules and units in the above-mentioned device and system are based on the same concept as the processing method embodiment of the present invention. For details, please refer to the description in the method embodiment of the present invention, and will not be repeated here.

[0128] Those skilled in the art will understand that all or part of the steps in the various methods of the embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0129] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A rapid recovery device for partial equipment failure in an optical data exchange or computing equipment cluster, characterized in that, The rapid recovery device comprises one or more protection optical switch devices, which comprise: The protection optical switch device has N optical path ports at one end, which are optical path port Q1, optical path port Q2, …, and optical path port QN, respectively, and has N+1 ports at the other end, which are optical path port P1, optical path port P2, …, and optical path port PN, and optical path port PC. The optical path port Qi and the optical path port Pi are connected in series at the common port of the optical coupler i and the first port of the optical coupler i, respectively, and the second port of the optical coupler i is connected to the optical path port PC through a 1×N optical switch; wherein the i-th port of the 1×N optical switch is coupled to the second port of the optical coupler i, the common port of the 1×N optical switch is coupled to the optical path port PC, and i∈[1, N]. When the 1×N optical switch switches to the k-th port of itself and the common port of itself, the optical path of the optical path port Pk and the optical path of the optical path port PC are both connected to the optical path port Qk.

2. The device for fast recovery from partial failure of a cluster of optical data switching or computing devices according to claim 1, wherein, An optical amplifier is further connected in series between the common port of the 1×N optical switch and the optical path port PC.

3. The device for fast recovery of partial device failure in a cluster of optical data switching or computing devices according to claim 1, wherein, The rapid recovery device comprises one or more protection optical switch devices, which comprise a protection optical switch array composed of protection optical switch device 1, protection optical switch device 2, …, and protection optical switch device M, wherein the optical switches in each protection optical switch device are uniformly managed by a management device. The output end of the first layer device library comprises at least M sets of transmission ports, and each set of transmission ports comprises at most N transmission ports. Each set of transmission ports is carried by the optical path ports Q1-QN of one protection optical switch device, and the optical path ports P1-PN of the corresponding protection optical switch devices are respectively allocated to the second layer devices 1-2. Wherein, the optical path port PC of each protection optical switch device is connected to a redundant device, and the redundant device has the functional attributes of each second layer device; each second layer device and the redundant device further establish a data control signal interaction link with the management device.

4. The apparatus for fast recovery from partial equipment failure in a cluster of optical data switching or computing equipment according to claim 3, wherein, For the same network topology level of N identical optical data exchange or calculation devices, a redundant optical data exchange or calculation device is added, and a set of N:1 backup optical switches are used for redundant backup.

5. The apparatus for fast recovery from partial equipment failure in a cluster of optical data switching or computing equipment according to claim 4, wherein, When the N second layer devices are working normally, the redundant device port is closed; When one of the N second layer devices fails, the management device transfers the data and state of the failed second layer device to the redundant device.

6. The apparatus for fast recovery from partial equipment failure in a cluster of optical data switching or computing equipment according to claim 3, wherein, The management device is used for acquiring an alarm message originating from any one of the second layer devices 1-2, and further comprises: The management device confirms the identification number m of the second layer device sending the alarm message; According to the pre-established mapping relationship between the identification number m of the second layer device and the optical path port Pm in each protection optical switch device, the management device controls the optical switch in each protection optical switch device to make the common port and the optical path port Pm conductive, so that the optical signal of the first layer device library is copied to the second layer redundant device through the optical switch and the optical coupler m in each protection optical switch device. The second layer redundant device replaces the role of the second layer device m and continues to interact with the first layer device bank.

7. The apparatus for fast recovery from partial equipment failure in a cluster of optical data switching or computing equipment according to claim 3, wherein, Further comprising: Any second layer device x in the second layer device bank 1-second layer device N, when confirming that the communication between itself and the first layer device bank is in a temporarily followable state, sends a light path detection request to the management device; The management device confirms whether the current second layer redundant device is in a replacement state, and if the second layer redundant device is in a replacement state, returns a response message to the second layer device x that the second layer redundant device is in a replacement state; If the second layer redundant device is in an idle state, a response message is returned that the second layer redundant device enters a temporary follow state, wherein at the same time of sending the response message, the common port in each protection optical switch device is controlled to be conductive with the corresponding optical path port Qx allocated to the second layer device x in the optical switch device, so that the state stability of the second layer device x is confirmed by the management device by comparing the optical signals exchanged between the second layer redundant device and the second layer device x; wherein x belongs to [1, N].

8. The apparatus for fast recovery from partial equipment failure in a cluster of optical data switching or computing devices according to claim 7, wherein, The replacement state indicates that there is a second layer device failure in the current system architecture, or the optical path connection between the corresponding second layer device and the first layer device bank has a problem, and the optical path link between the second layer device and the first layer device bank has been switched to the optical path link between the second layer redundant device and the first layer device bank by the management device controlling the matching protection optical switch device.

9. The apparatus for fast recovery from partial equipment failure in a cluster of optical data switching or computing devices of claim 7, wherein, The temporarily followable state specifically includes: The first layer device bank and the second layer device are in a heartbeat message maintained link state; or, The first layer device bank and the second layer device are in a continuous data sending state in a non-encrypted state; or, The first layer device bank and the second layer device are in an inherent script execution state, wherein the execution instructions of the script are known in advance by the management device.

10. The apparatus for fast recovery from partial equipment failure in a cluster of optical data switching or computing equipment according to claim 9, wherein, The meaning expressed by the temporarily followable state is that the communication content between the current first layer device bank and the second layer device can be predicted in advance, and the corresponding second layer redundant device can match and verify with the second layer device x through the corresponding parsed content even when it is temporarily connected to the first layer device bank.

11. The apparatus for fast recovery from partial equipment failure in a cluster of optical data switching or computing equipment according to claim 7, wherein, The management device controls the working power of the optical amplifier in each protection optical switch device, and in the case of taking the same optical power accepted by the second layer device x and the second layer redundant device as the target, confirms whether the adjustment of the working power of the optical amplifier in each protection optical switch device is within the preset range, thereby confirming the stability of the working state of each optical path port. ​ 12. The apparatus for fast recovery from partial equipment failure in a cluster of optical data switching or computing equipment according to claim 1, wherein, The rapid recovery device includes one or more protection optical switch devices, specifically a northward protection optical switch array composed of U protection optical switch devices and a southward protection optical switch array composed of V protection optical switch devices; N third layer devices and a third layer redundant device are connected in series between the northward protection optical switch array and the southward protection optical switch array, wherein each third layer device has V optical interfaces on the side facing the southward protection optical switch array and U optical interfaces on the side facing the northward protection optical switch array; The other side optical interfaces of the northward protection optical switch array and the southward protection optical switch array, except the optical interfaces coupled with the third layer devices, are respectively used for connecting a northward device bank and a southward device bank; The optical switch in each protection optical switch device is uniformly managed by a management device, and each third layer device and the third layer redundant device also establish a data interaction channel with the management device.

13. The apparatus for fast recovery from partial equipment failure in a cluster of optical data switching or computing devices of claim 12, wherein, The first layer device bank is also referred to as the northward device of the third layer device, and the corresponding second layer device is also referred to as the southward device of the third layer device; The protection optical switch array of the first layer device bank is also referred to as the northward N:1 protection switch device×U of the third layer device; The protection optical switch array of the second layer device is also referred to as the southward N:1 protection switch device×V of the third layer device.

14. The apparatus for fast recovery from partial equipment failure in a cluster of optical data switching or computing devices of claim 12, wherein, When introducing a redundant device through a backup optical switch, a backup optical switch with a northward connection port and a southward connection port needs to be configured at the same time.

15. The apparatus for fast recovery from partial equipment failure in a cluster of optical data switching or computing equipment according to claim 1, wherein, The second layer device is specifically computer rack 1, computer rack 2, …, computer rack 8; the second layer redundant device is specifically a redundant computer rack; the first layer device bank is specifically 8 leaf computing switches; 32 backup optical switches are matched, wherein N is 8, and the rapid recovery device includes: The j'th port of the k'th H100 of the redundant computer rack is connected to the PC port of the 4k'+j'-4'th backup optical switch; wherein each H100 includes a group of 8 optical interfaces, and the numbers 1 to 8 are formed; The j'th port of the k'th H100 of the i'th computer rack is connected to the Pi' port of the 4k'+j'-4'th backup optical switch; The Qi' port of the 4k'+j'-4'th backup optical switch is connected to the 4i'+k-4'th port of the j'th leaf computing switch; wherein k' and j' are natural numbers.

16. The apparatus for fast recovery from partial equipment failure in a cluster of optical data switching or computing equipment of claim 1, wherein, The optical coupler adopts an asymmetric splitting ratio.

17. A method for fast recovery from partial equipment failure in a cluster of optical data switching or computing devices, characterized in that, The rapid recovery device includes protection optical switch device 1, protection optical switch device 2, …, protection optical switch device M, wherein the optical switch in each protection optical switch device is uniformly managed by a management device; the output end of the first layer device bank includes at least M sets of transmission ports, each set of transmission ports includes at most N transmission ports; each set of transmission ports is carried by the optical path port Q1-optical path port QN of a protection optical switch device, and the optical path port P1-optical path port PN of the corresponding each protection optical switch device is respectively allocated to the second layer device 1-second layer device N; the method includes: The management device is used to acquire an alarm message originating from any one of the second layer device 1-second layer device N; The management device confirms the identification number m of the second layer device sending the alarm message; According to the pre-established mapping relationship between the identification number m of the second layer device and the optical path port Pm in each protection optical switch device, the management device controls the optical amplifier in each protection optical switch device, so that the common port is connected with the optical path port Pm, thereby the optical signal of the first layer device bank is copied to the second layer redundant device through the optical amplifier in each protection optical switch device and the optical coupler m. The second layer redundant device replaces the role of the second layer device m and continues to interact with the first layer device bank.

18. The method for fast recovery from partial equipment failure in a cluster of optical data switching or computing equipment according to claim 17, wherein, The method further comprises: Any second layer device x in the second layer devices 1-second layer devices N, when confirming that the communication between itself and the first layer device bank is in a temporary followable state, the second layer device x sends an optical path detection request to the management device; The management device confirms whether the current second layer redundant device is in a replacement state, if the second layer redundant device is in a replacement state, returns a response message to the second layer device x that the second layer redundant device is in a replacement state; If the second layer redundant device is in an idle state, returns a response message that the second layer redundant device enters a temporary followable state, wherein at the same time of sending the response message, the common port in each protection optical switch device is controlled to be connected with the corresponding optical path port Qx in the optical switch device allocated to the second layer device x, thereby the state stability of the second layer device x is confirmed by the management device by comparing the optical signals exchanged between the second layer redundant device and the second layer device x.

19. The method for fast recovery from partial equipment failure in a cluster of optical data switching or computing equipment according to claim 18, wherein, The temporary followable state specifically comprises: The first layer device bank and the second layer device are in a heartbeat packet maintaining link state; or, The first layer device bank and the second layer device are in a continuous data sending state in a non-encrypted state; or, The first layer device bank and the second layer device are in an inherent script execution state, wherein the execution instruction of the script is pre-known by the management device.

20. The method for fast recovery from partial equipment failure in a cluster of optical data switching or computing equipment according to claim 18, wherein, The state stability of the second layer device x is confirmed by the management device by comparing the optical signals exchanged between the second layer redundant device and the second layer device x, and specifically further comprises: The management device controls the working power of the optical amplifier in each protection optical switch device, in the case of taking the same optical power accepted by the second layer device x and the second layer redundant device as a target, confirms whether the adjustment of the working power of the optical amplifier in each protection optical switch device is within a preset range, thereby confirming the stability of the working state of each optical path port.

Citation Information

Patent Citations

  • Multi-channel optical fiber automatic backup device for broadcasting network

    CN103441792A

  • Self-diagnostic method for PON protection system, and PON protection system

    CN103959684A

  • Multi-backup OTDR optical amplification device with shared light source and control method

    CN106452569A

  • On-line monitoring optical communication link monitoring protection device and method

    CN118118088A

  • Method and device for realizing business pretection by adopting tunable light source

    CN1503495A

Cited By

  • High-availability network access authentication method and system based on multiple optical modules

    CN121967085A