Optical fiber link fault monitoring method and system of AI data center, and optical module

By detecting the number of symbol errors in data frames and the performance parameters of optical modules in the fiber optic links of AI data centers, and combining this with signal monitoring according to the IEEE 802.3ck standard, the problem of concealed faults in fiber optic links has been solved, enabling rapid early warning and fault location, and improving the stability and operating efficiency of fiber optic links.

CN121619019APending Publication Date: 2026-03-06HUAXIA XINZHIZHI PHOTONICS TECH (BEIJING) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202610139814.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-02
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Fiber optic link failures in AI data centers are highly concealed. In particular, signal transmission instability and increased bit error rate caused by dirt on the optical transmission link end face affect computing efficiency and economic losses. Existing technologies make it difficult to quickly locate and warn of such failures.

Method used

By periodically detecting the number of symbol errors in optical transmission link data frames and combining this with the reporting of optical module performance parameters, rapid early warning of link faults and fault source location can be achieved. The PAM4 code signal of the KP4 FEC frame structure of the IEEE 802.3ck standard is used for monitoring to promptly detect fiber optic contamination and report alarm information.

Benefits of technology

It enables rapid location and early warning of fiber optic link failures, reduces repetitive training and inference interruptions, lowers maintenance costs, and improves the operational efficiency and reliability of AI data centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121619019A_ABST
    Figure CN121619019A_ABST
Patent Text Reader

Abstract

The invention relates to an optical fiber link fault monitoring method and system of an AI data center and an optical module, and belongs to the technical field of optical communication, the method comprises the following steps: an upstream optical module is instructed to send a standard test signal to a downstream optical module, and the downstream optical module receives the standard test signal to obtain a data frame error symbol number which is less than or equal to a first early warning value; if so, judging whether the working parameters of the upstream and downstream optical modules are normal, if so, returning to the beginning, and if not, reporting a first monitoring alarm signal to a network management server; if the data frame error symbol number is gt; judging whether the working parameters of the upstream and downstream optical modules are normal or not according to the first early warning value; and the downstream optical module reports the first monitoring alarm signal to the network management server if the first monitoring alarm signal is abnormal, and reports the second monitoring alarm signal to the network management server if the second monitoring alarm signal is normal. According to the invention, the maintenance cost of the AI data center optical fiber network is reduced, the reliability of the optical network is improved, and the improvement of the working efficiency of the AI data center network is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of optical communication technology. Specifically, this invention relates to a method, system, and optical module for monitoring optical fiber link faults in AI data centers. Background Technology

[0002] The explosion of artificial intelligence technologies, represented by generative AI and large-scale language models, has led to an explosive growth in the demand for bandwidth, efficiency, and stability of data transmission in computing networks. Traditional low- and medium-speed optical modules can no longer support various high-load tasks related to AI. This has directly driven 400G / 800G / 1.6T optical modules from technology reserves to large-scale applications, and they are increasingly becoming key components of optical transmission links. Through photoelectric conversion, they have realized high-bandwidth, low-latency optical interconnect technology for xPU accelerated computing units. Tens of thousands of optical interconnect links are the "neural network" for the operation of AI data centers, and pluggable optical modules greatly improve the maintainability of optical networks in AI data centers.

[0003] When tens of thousands of xPU clusters are training in an AI data center, the efficiency of data communication between a large number of parallel nodes directly determines the computing power. If there are faults such as high latency or high packet loss rate in the interconnected optical transmission links, it will inevitably lead to a significant decrease in parallel computing efficiency, extended AI training time, slower inference response, and in severe cases, even interruption of computing services.

[0004] AI data centers have dense optical interconnects, and failures in optical transmission links are often difficult to detect. In particular, dirt or moisture on the end face of the optical interconnect transmission link caused by the data center environment can lead to a sudden increase in the bit error rate of optical transmission, resulting in interruption of system training or inference. After the fault is cleared, training or inference needs to be restarted, which not only has a high time cost, but also reduces economic benefits, making it difficult to guarantee the efficiency of AI cluster operation.

[0005] AI data center on-site commissioning engineers are often not optical communication professionals or professional optical module commissioning and maintenance personnel. On-site operators may easily cause system errors by improperly operating the optical module's optoelectronic interface or optical transmission link interface without taking measures such as electrostatic protection. The phenomenon of arbitrarily bending the optical fiber transmission link is common. Operators often misjudge the failure of the optical transmission link as a failure of the optical module itself, thereby prolonging the fault location and handling time and causing economic losses.

[0006] The 400G / 800G / 1.6T optical modules for AI data centers are short-range optical communication products, belonging to intensity modulation, direct detection (IM-DD) optical module products. They employ PAM4 modulation technology. The PAM4 modulation format is significantly more sensitive to end-face reflections in optical transmission links than the traditional NRZ modulation format. This is essentially due to its multi-level signal characteristics, which reduce anti-interference redundancy. End-face reflections in optical transmission links can cause signal distortion, eye diagram degradation, and other problems. The noise margin of PAM4 modulated signals is only about 1 / 3 of that of NRZ modulated signals (theoretical value). At the same data rate, the OSNR margin degradation at the receiver is 9.5dB. Moreover, end-face reflections in optical transmission links are "systematic interference," not random noise, and will continuously affect the stability of signal transmission. When the end-face reflection coefficient of the optical transmission link exceeds a certain threshold, the system noise margin will exceed the critical value, and the signal transmission bit error rate will rise sharply.

[0007] AI data center 400G / 800G / 1.6T optical modules are typically short-distance, multimode fiber transmission products. Dirt on the multimode fiber endface is a common problem in fiber optic communication, usually caused by dust, oil, grease, moisture, liquid residue, fingerprints, etc. (See attached image for details.) Figures 1 to 4 As shown, contamination can severely impact signal transmission. 1) Signal attenuation: Contamination hinders optical signal transmission, reducing the signal power at the receiving end and affecting the quality of optical signal transmission; 2) Increased bit error rate: Reflected light interference and signal attenuation can cause data transmission errors, reducing transmission accuracy; 3) Link instability or even interruption: It can cause "flashover" of optical modules, and severe contamination may prevent the fiber optic link from functioning properly, affecting service continuity. According to statistics, approximately 65% ​​of optical module failures in AI data centers are due to "contamination" at the optical connector end face in the optical transmission link; 4) It can cause optical multipath interference, changing the amplitude and phase of the signal, thus affecting signal quality. For example, when the delay is large enough, the delayed modulated optical signal cannot be eliminated by the FFE filter and will act as residual inter-symbol interference (ISI), reducing BER performance.

[0008] As AI data centers scale towards million-card-level optical interconnects, the number of optical connection nodes increases exponentially, significantly increasing the complexity of optical transmission links in AI data center networks. Seemingly minor link failures can occur repeatedly, severely impacting the training and inference efficiency of high-performance servers, leading to incalculable economic losses. There is an urgent need in this field for a technology that can provide early warning and rapid location of fiber optic communication failures. Summary of the Invention

[0009] To address the aforementioned issues, this invention employs a technique that periodically detects the number of symbol errors in optical transmission link data frames and combines this with the reporting of optical module performance parameters. This enables rapid early warning of link faults and fault source location, allowing for timely cleaning and maintenance of the contaminated end face of the optical fiber, ensuring the stability of the optical fiber communication link, and improving troubleshooting efficiency.

[0010] In a first aspect, the present invention provides a method for monitoring fiber optic link faults in AI data centers, comprising the following steps:

[0011] Step S1: Start the subsequent monitoring process based on the link idle time or the period T;

[0012] Step S2: Instruct the upstream optical module to send a standard test signal to the downstream optical module;

[0013] Step S3: After receiving the standard test signal sent by the upstream optical module, the downstream optical module performs data frame statistics to obtain the number of data frame error symbols;

[0014] Step S4: Compare the number of data frame error symbols with the first warning value to determine whether it exceeds the first warning value;

[0015] Step S5: If the number of data frame error symbols is less than or equal to the first warning value, extract the operating parameters of the upstream and downstream optical modules, compare them with the operating parameter threshold of the optical modules, and determine whether the operating parameters of the upstream and downstream optical modules are within the operating parameter threshold range of the optical modules. If the operating parameters of the upstream and downstream optical modules are within the operating parameter threshold range of the optical modules, return to step S1.

[0016] Step S6: If the number of data frame error symbols > the first warning value, extract the operating parameters of the upstream and downstream optical modules, compare them with the operating parameter threshold of the optical modules, and determine whether the operating parameters of the upstream and downstream optical modules are within the operating parameter threshold range of the optical modules.

[0017] Step S7: If the operating parameters of the upstream and downstream optical modules exceed the threshold range of the optical module operating parameters, the downstream optical module will report the first monitoring alarm signal to the network management server, and notify the system maintenance personnel to troubleshoot the optical module fault based on the alarm information issued by the network management server.

[0018] Step S8: If the operating parameters of the upstream and downstream optical modules are within the threshold range of the optical module operating parameters, the downstream optical module will report the second monitoring alarm signal to the network management server, and notify the system maintenance personnel to troubleshoot the fiber optic link fault based on the alarm information issued by the network management server.

[0019] Furthermore, the first monitoring alarm signal includes: upstream and downstream optical module address information and optical module fault alarm information;

[0020] The second monitoring alarm signal includes: upstream and downstream optical module address information, and fiber optic fault alarm information.

[0021] Furthermore,

[0022] The period T can range from 1 to 7 days, or from 2 to 4 weeks.

[0023] The first warning value ranges from 8 to 12.

[0024] Furthermore, the test signal is a PAM4 code signal conforming to the KP4 FEC frame structure of the IEEE 802.3ck standard;

[0025] The first warning value is 10.

[0026] In a second aspect, the present invention provides an optical module for monitoring fiber optic link faults, the optical module integrating an MCU chip and a fiber optic access port; the optical module provides photoelectric data conversion function and is used to be installed in a computing server, leaf layer switch, spine layer switch or core switch;

[0027] The MCU chip executes by calling built-in programs or instructions:

[0028] The subsequent monitoring process is initiated based on the link idle time or on a cycle T.

[0029] The upstream optical module sends a standard test signal to the downstream optical module;

[0030] After receiving the standard test signal sent by the upstream optical module, the downstream optical module performs data frame statistics to obtain the number of data frame error symbols.

[0031] Compare the number of data frame error symbols with the first warning value to determine whether it exceeds the first warning value;

[0032] If the number of data frame error symbols is less than or equal to the first warning value, the operating parameters of the upstream and downstream optical modules are extracted and compared with the operating parameter threshold of the optical modules to determine whether the operating parameters of the upstream and downstream optical modules are within the operating parameter threshold range. If the operating parameters of the upstream and downstream optical modules are within the operating parameter threshold range, the process returns to the next round of monitoring.

[0033] If the number of data frame error symbols is greater than the first warning value, the operating parameters of the upstream and downstream optical modules are extracted and compared with the operating parameter threshold of the optical modules to determine whether the operating parameters of the upstream and downstream optical modules are within the operating parameter threshold range of the optical modules.

[0034] If the operating parameters of the upstream and downstream optical modules exceed the threshold range of the optical module operating parameters, the downstream optical module will report the first monitoring alarm signal to the network management server.

[0035] If the operating parameters of the upstream and downstream optical modules are within the threshold range of the optical module operating parameters, the downstream optical module will report the second monitoring alarm signal to the network management server.

[0036] Furthermore, the first monitoring alarm signal includes: upstream and downstream optical module address information and optical module fault alarm information; the second monitoring alarm signal includes: upstream and downstream optical module address information and optical fiber fault alarm information.

[0037] Furthermore, the range of values ​​for the period T includes: 1 to 7 days, or 2 to 4 weeks;

[0038] The first warning value ranges from 8 to 12;

[0039] The test signal is a PAM4 code signal conforming to the KP4 FEC frame structure of the IEEE 802.3ck standard.

[0040] A third aspect of the present invention provides a fiber optic link fault monitoring system for AI data centers, comprising:

[0041] An optical module, as described in the second aspect of the present invention, is used for optical fiber link fault monitoring; a plurality of the optical modules are respectively installed in a computing server, a leaf layer switch, a spine layer switch, and a core switch; the optical modules in the computing server are connected to the optical modules in the leaf layer switches via optical fibers, the optical modules in the leaf switches are connected to the optical modules in the spine layer switches via optical fibers, and the optical modules in the spine switches are connected to the optical modules in the core switch via optical fibers.

[0042] The network management server is connected to the computing server and core switch in the system via network cable; it is used to monitor the network working status of the system, and to receive upstream and downstream optical module address information, optical module fault alarm information or fiber optic fault alarm information reported by any optical module in the system, and to notify system maintenance personnel to troubleshoot optical module faults or fiber optic faults.

[0043] The optical module works in conjunction with the network management server to execute the fiber optic link fault monitoring method for an AI data center as described in the first aspect of this invention.

[0044] Furthermore, the computing server is used to provide AI computing power resources.

[0045] Furthermore, the leaf layer switch is connected to several of its subordinate computing server nodes, and is responsible for aggregating the traffic of the computing servers, while establishing interconnection links with the spine layer switch above it.

[0046] The spine layer switch is interconnected with all leaf layer switches through multiple equal bandwidth links, enabling non-blocking communication between any leaf layer switch nodes.

[0047] The core switch is used for traffic scheduling across multiple spine layer switches in a cluster, expanding the network scale of the AI ​​data center.

[0048] The advantages of this invention compared to the prior art are:

[0049] During the deployment and commissioning phase of fiber optic transmission systems, especially during the commissioning phase of optical interconnects in AI data centers, this invention enables rapid location of fiber optic link faults. Faults can be efficiently eliminated by replacing passive optical components or cleaning the optical interface end face.

[0050] During the operation of an AI data center, the fiber optic transmission link uses idle intervals to monitor the data transmission effect of the fiber optic link. When problems are detected, early warning information is promptly output to users. This allows for early detection of fiber optic link failures, preventing business training failures or interruptions caused by fiber optic link failures when computing units or systems directly perform training or inference, thus avoiding significant economic losses and improving the long-term security and reliability of the AI ​​data center.

[0051] This invention enables regular maintenance of the optical interconnect transmission links in AI data centers, completely eliminating "flashover" faults during training or inference caused by dirt on the fiber end face of the optical transmission link, improving the operating efficiency of AI data centers, and avoiding significant economic losses caused by repeated training and inference interruptions.

[0052] This invention patent reduces the maintenance cost of optical fiber networks in AI data centers, improves the reliability of optical networks, saves manpower and time costs in troubleshooting optical link faults, and promotes the improvement of the working efficiency of AI data center networks. Attached Figure Description

[0053] Figure 1 This is a photograph of the fiber optic end face under normal conditions.

[0054] Figure 2 This is a photo of the fiber optic end face when it is dusty.

[0055] Figure 3 This is a photograph of the fiber optic end face when liquid residue is present.

[0056] Figure 4 This is a photograph of the fiber optic end face when there is grease residue.

[0057] Figure 5 This is a schematic diagram of a method for monitoring fiber optic link faults in an AI data center, provided as an embodiment of the present invention.

[0058] Figure 6 This is a schematic diagram of the structure of an optical fiber link fault monitoring system for an AI data center, provided as an embodiment of the present invention. Detailed Implementation

[0059] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this patent application are within the scope of protection of this application.

[0060] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0061] This invention achieves rapid early warning and fault source location of link faults by periodically detecting the number of symbol errors in optical transmission link data frames and combining this with the reporting of optical module performance parameters. This enables timely detection of fiber optic faults and timely cleaning and maintenance of dirty end faces, ensuring the stability of optical communication links and improving troubleshooting efficiency.

[0062] In a first aspect, the present invention provides a method for monitoring fiber optic link faults in AI data centers, the flowchart of which is attached. Figure 5 As shown, the specific steps include:

[0063] Step S1: Periodically start and execute subsequent monitoring processes according to link idle time or period T.

[0064] The purpose of this step is to monitor the fiber optic link based on a certain period or link idle time, which can both ensure timely detection of faults and save monitoring costs.

[0065] For example, the period T can range from 1 to 7 days, or from 2 to 4 weeks.

[0066] Step S2: Instruct the upstream optical module to send a specific standard test signal to the downstream optical module.

[0067] The specific test signal is a PAM4 code signal conforming to the KP4 FEC frame structure of the IEEE 802.3ck standard.

[0068] The purpose of this step is to have the upstream optical module emit a specific standard signal to detect the data transmission between the upstream and downstream optical modules.

[0069] The upstream optical module refers to the optical module built into the upstream device, and the downstream optical module refers to the optical module built into the downstream device. The upstream and downstream devices are connected by an optical fiber transmission link, forming an optical link. The upstream and downstream devices include computing servers and switches.

[0070] Step S3: After receiving the specific standard test signal sent by the upstream optical module, the downstream optical module performs data frame statistics to obtain the number of data frame error symbols.

[0071] Step S4: Compare the number of data frame error symbols with the first warning value to determine whether it exceeds the first warning value.

[0072] For example, the first warning value ranges from 8 to 12.

[0073] Ideally, to avoid frequent false alarms, the first warning value should be set to 10.

[0074] In this step, the number of error symbols in the data frames transmitted by the optical fiber link is obtained and compared with the first warning value to characterize the multipath interference warning threshold of the optical fiber transmission link, thereby obtaining a warning of "fault" in the optical fiber link transmission.

[0075] Step S5: If the number of data frame error symbols is less than or equal to the first warning value, extract the operating parameters of the upstream and downstream optical modules and compare them with the threshold values ​​of the optical module operating parameters to determine whether the operating parameters of the upstream and downstream optical modules are within the threshold value range. If the operating parameters of the upstream and downstream optical modules are within the threshold value range, it indicates that the optical fiber transmission link is normal, and return to step S1. If they are not within the threshold value range, the downstream optical module will report the first monitoring alarm signal to the network management server. The first monitoring alarm signal includes the address information of the upstream and downstream optical modules, optical module fault alarm information, etc., and notify the system maintenance personnel to troubleshoot the optical module fault according to the alarm information issued by the network management server.

[0076] Step S6: If the number of data frame error symbols is greater than the first warning value, extract the operating parameters of the upstream and downstream optical modules, compare them with the threshold of the optical module operating parameters, and determine whether the operating parameters of the upstream and downstream optical modules are within the threshold range of the optical module operating parameters.

[0077] For example, the 400G / 800G / 1.6T optical modules used in AI data centers all employ PAM4 modulation technology. The IEEE 802.3 standard uniformly adopts the KP4 FEC error correction code for this modulation pattern. KP4 FEC technology can correct a maximum of 15 errors per data frame, meaning that 15 or fewer errors can restore the original information through the error correction mechanism. As the "dirtiness" of the fiber optic transmission link endface gradually increases, the number of error symbols detected by the receiver also gradually increases. If the "dirtiness" of the fiber optic transmission endface is severely degraded, when the number of error symbols in the data frames received by the transmission link receiver exceeds 15, the optical link transmission system will have uncorrectable bit errors. When the transmission error rate reaches a certain value, the fiber optic transmission performance cannot meet the data service transmission requirements, indicating a potential fault in the optical module or optical transmission link of the relevant link. Typically, the number of error symbols in the data frames of optical modules used in AI data centers does not exceed 3 at the time of manufacture. When the number of data frame error symbols exceeds the first warning value, if the operating parameters of the upstream and downstream optical modules are normal, the present invention requires further verification and troubleshooting of the optical fiber link.

[0078] Step S7: If the operating parameters of the upstream and downstream optical modules exceed the threshold range of the optical module operating parameters, the downstream optical module will report the first monitoring alarm signal to the network management server. The first monitoring alarm signal includes the address information of the upstream and downstream optical modules, optical module fault alarm information, etc., and notifies the system maintenance personnel to troubleshoot the optical module fault according to the alarm information issued by the network management server.

[0079] Furthermore, maintenance personnel locate the corresponding optical modules based on the upstream and downstream optical module address information provided by the network management server, and eliminate optical module faults.

[0080] Step S8: If the operating parameters of the upstream and downstream optical modules are within the threshold range of the optical module operating parameters, the downstream optical module will report the second monitoring alarm signal to the network management server. The second monitoring alarm signal includes the address information of the upstream and downstream optical modules, optical fiber fault alarm information, etc., and notifies the system maintenance personnel to troubleshoot the optical fiber link fault according to the alarm information issued by the network management server.

[0081] In this step, if the operating parameters of the upstream and downstream optical modules are within the threshold range of the optical module operating parameters, it proves that the upstream and downstream optical modules are working normally. However, if it is determined in steps S3 to S6 that the number of erroneous symbols received by the downstream optical module has exceeded the set threshold, then it is determined that there is a problem with the optical fiber transmission link between the upstream and downstream optical modules, and an early warning and maintenance information needs to be issued and reported to the network management server.

[0082] Furthermore, maintenance personnel locate the corresponding fiber optic links based on the upstream and downstream optical module address information provided by the network management server, clean the fiber end faces, and eliminate fiber optic link faults.

[0083] In a second aspect, the present invention provides an optical module for monitoring fiber optic link faults, the optical module integrating an MCU chip and a fiber optic access port; the optical module provides photoelectric data conversion function and is used to be installed in a computing server, leaf layer switch, spine layer switch or core switch.

[0084] The MCU chip executes by calling built-in programs or instructions:

[0085] The subsequent monitoring process is initiated based on the link idle time or on a cycle T.

[0086] The upstream optical module sends a standard test signal to the downstream optical module;

[0087] After receiving the standard test signal sent by the upstream optical module, the downstream optical module performs data frame statistics to obtain the number of data frame error symbols.

[0088] Compare the number of data frame error symbols with the first warning value to determine whether it exceeds the first warning value;

[0089] If the number of data frame error symbols is less than or equal to the first warning value, the operating parameters of the upstream and downstream optical modules are extracted and compared with the threshold of the optical module operating parameters to determine whether the operating parameters of the upstream and downstream optical modules are within the threshold range. If the operating parameters of the upstream and downstream optical modules are within the threshold range, the next round of monitoring process is returned. If they are not within the threshold range, the downstream optical module reports the first monitoring alarm signal to the network management server, notifying the system maintenance personnel to troubleshoot the optical module fault based on the alarm information issued by the network management server.

[0090] If the number of data frame error symbols is greater than the first warning value, the operating parameters of the upstream and downstream optical modules are extracted and compared with the operating parameter threshold of the optical modules to determine whether the operating parameters of the upstream and downstream optical modules are within the operating parameter threshold range of the optical modules.

[0091] If the operating parameters of the upstream and downstream optical modules exceed the threshold range of the optical module operating parameters, the downstream optical module will report the first monitoring alarm signal to the network management server.

[0092] If the operating parameters of the upstream and downstream optical modules are within the threshold range of the optical module operating parameters, the downstream optical module will report the second monitoring alarm signal to the network management server.

[0093] Furthermore, the first monitoring alarm signal includes: upstream and downstream optical module address information and optical module fault alarm information; the second monitoring alarm signal includes: upstream and downstream optical module address information and optical fiber fault alarm information.

[0094] Furthermore, the range of values ​​for the period T includes: 1 to 7 days, or 2 to 4 weeks;

[0095] The first warning value ranges from 8 to 12;

[0096] The test signal is a PAM4 code signal conforming to the KP4 FEC frame structure of the IEEE 802.3ck standard.

[0097] A third aspect of the present invention provides a fiber optic link fault monitoring system for AI data centers, used to implement the fiber optic link fault monitoring method for AI data centers described in the first aspect of the present invention. Specifically, as... Figure 6 As shown, the system includes:

[0098] An optical module, as described in the second aspect of this invention, is used for optical fiber link fault monitoring. Several optical modules are respectively installed in a computing server, a leaf layer switch, a spine layer switch, and a core switch. The optical modules in the computing server are connected to the optical modules in the leaf layer switches via optical fibers, the optical modules in the leaf layer switches are connected to the optical modules in the spine layer switches via optical fibers, and the optical modules in the spine layer switches are connected to the optical modules in the core switch via optical fibers.

[0099] The network management server is connected to the computing server and core switch within the system via network cables. It is used to monitor the network operating status of the system and to receive upstream and downstream optical module address information, optical module fault alarm information, or fiber optic fault alarm information reported by any optical module within the system, and to notify system maintenance personnel to troubleshoot optical module faults or fiber optic faults.

[0100] The optical module works in conjunction with the network management server to execute the fiber optic link fault monitoring method for an AI data center as described in the first aspect of this invention.

[0101] The computing server is used to provide AI computing power resources;

[0102] The leaf layer switch is directly connected to several of its subordinate computing server nodes, and is responsible for aggregating the traffic of the computing servers, while establishing interconnection links with the spine layer switch above it.

[0103] The spine layer switch is interconnected with all leaf layer switches through multiple equal bandwidth links, enabling non-blocking communication between any leaf layer switch nodes.

[0104] The core switch is used for traffic scheduling across multiple spine layer switches in a cluster, further expanding the network scale of the AI ​​data center.

[0105] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for monitoring optical fiber link faults in an AI data center, the method comprising: Comprising the following steps: ​ Step S1: start the subsequent monitoring process according to the link idle time or according to the period T; Step S2: instruct the upstream optical module to send a standard test signal to the downstream optical module; Step S3: after receiving the standard test signal sent by the upstream optical module, the downstream optical module performs data frame statistics to obtain the number of data frame symbol errors; Step S4: compare the number of data frame symbol errors with the first warning value to determine whether it exceeds the first warning value; Step S5: if the number of data frame symbol errors is less than or equal to the first warning value, extract the working parameters of the upstream and downstream optical modules, compare them with the optical module working parameter threshold value, and determine whether the working parameters of the upstream and downstream optical modules are within the optical module working parameter threshold value range. If the working parameters of the upstream and downstream optical modules are within the optical module working parameter threshold value range, return to step S1. If they are not within the optical module working parameter threshold value range, the downstream optical module reports a first monitoring alarm signal to the network management server, and notifies the system maintenance personnel to troubleshoot the optical module failure according to the alarm information issued by the network management server; Step S6: if the number of data frame symbol errors is greater than the first warning value, extract the working parameters of the upstream and downstream optical modules, and compare them with the optical module working parameter threshold value to determine whether the working parameters of the upstream and downstream optical modules are within the optical module working parameter threshold value range; Step S7: if the working parameters of the upstream and downstream optical modules exceed the optical module working parameter threshold value range, the downstream optical module reports a first monitoring alarm signal to the network management server, and notifies the system maintenance personnel to troubleshoot the optical module failure according to the alarm information issued by the network management server; Step S8: if the working parameters of the upstream and downstream optical modules are within the optical module working parameter threshold value range, the downstream optical module reports a second monitoring alarm signal to the network management server, and notifies the system maintenance personnel to troubleshoot the optical fiber link failure according to the alarm information issued by the network management server.

2. The optical fiber link failure monitoring method of an AI data center according to claim 1, characterized in that: the first monitoring alarm signal includes upstream and downstream optical module address information and optical module failure alarm information; the second monitoring alarm signal includes upstream and downstream optical module address information and optical fiber failure alarm information.

3. The optical fiber link failure monitoring method of an AI data center according to claim 1, characterized in that: the period T has a value range of 1 day to 7 days, or 2 weeks to 4 weeks; the first warning value has a value range of 8 to 12.

4. The optical fiber link failure monitoring method of an AI data center according to claim 1, characterized in that: the test signal is a PAM4 code type signal with a KP4 FEC frame structure in accordance with the IEEE 802.3ck standard; the first warning value has a value of 10.

5. An optical module for optical fiber link failure monitoring, characterized in that: the optical module is integrated with an MCU chip and a fiber access port; the optical module provides optical-electrical data conversion function and is used for installation in a computing server, a leaf layer switch, a spine layer switch, or a core switch; the MCU chip is used for executing the following functions by calling the built-in program or instruction: starting the subsequent monitoring process according to the link idle time or according to the period T; The upstream optical module sends a standard test signal to the downstream optical module; The downstream optical module receives the standard test signal sent by the upstream optical module and performs data frame statistics to obtain the number of data frame symbol errors; The number of data frame symbol errors is compared with the first early warning value to determine whether it exceeds the first early warning value; If the number of data frame symbol errors is less than or equal to the first early warning value, the working parameters of the upstream and downstream optical modules are extracted, and the working parameters of the optical module are compared with the threshold value of the working parameters of the optical module to determine whether the working parameters of the upstream and downstream optical modules are within the threshold value range of the working parameters of the optical module. If the working parameters of the upstream and downstream optical modules are within the threshold value range of the working parameters of the optical module, the next round of monitoring process is returned. If the number of data frame symbol errors is greater than the first early warning value, the working parameters of the upstream and downstream optical modules are extracted, and the working parameters of the optical module are compared with the threshold value of the working parameters of the optical module to determine whether the working parameters of the upstream and downstream optical modules are within the threshold value range of the working parameters of the optical module. If the working parameters of the upstream and downstream optical modules exceed the threshold value range of the working parameters of the optical module, the downstream optical module reports the first monitoring alarm signal to the network management server. If the working parameters of the upstream and downstream optical modules are within the threshold value range of the working parameters of the optical module, the downstream optical module reports the second monitoring alarm signal to the network management server.

6. The optical module for optical fiber link fault monitoring according to claim 5, wherein: The first monitoring alarm signal includes upstream and downstream optical module address information and optical module fault alarm information; The second monitoring alarm signal includes upstream and downstream optical module address information, optical fiber fault alarm information.

7. The optical module for optical fiber link fault monitoring according to claim 5, wherein: The value range of the period T includes 1 day to 7 days, or 2 weeks to 4 weeks; The value range of the first early warning value is 8 to 12; The test signal is a PAM4 code type signal in the KP4 FEC frame structure conforming to the IEEE 802.3ck standard.

8. An optical fiber link failure monitoring system of an AI data center, characterized by, It comprises: An optical module, which is the optical module for optical fiber link fault monitoring according to any one of claims 5 to 7; a plurality of optical modules are respectively installed in a computing server, a leaf layer switch, a spine layer switch and a core switch; the optical module in the computing server is connected with the optical module in the leaf layer switch through an optical fiber, the optical module in the leaf layer switch is connected with the optical module in the spine layer switch through an optical fiber, and the optical module in the spine layer switch is connected with the optical module in the core switch through an optical fiber; A network management server connected with the computing server and the core switch in the system through a network cable; For monitoring the network working state of the system, and receiving the upstream and downstream optical module address information, optical module fault alarm information or optical fiber fault alarm information reported by any optical module in the system, and notifying the system maintenance personnel to troubleshoot the optical module fault or optical fiber fault; The optical module and the network management server cooperate to perform the AI data center optical fiber link fault monitoring method according to any one of claims 1 to 4.

9. The optical fiber link failure monitoring system of an AI data center of claim 8, wherein: The computing server is used to provide AI computing power resources.

10. The AI data center optical fiber link fault monitoring system according to claim 8, wherein: The leaf layer switch is connected with a plurality of computing server nodes, is responsible for converging the traffic of the computing server, and establishes an interconnection link with the spine layer switch of the upper layer; The spine layer switch is interconnected with all the leaf layer switches through a plurality of equal-bandwidth links, and realizes non-blocking communication between any leaf layer switch nodes. The core switch is used for traffic scheduling of the plurality of spine layer switches across the cluster, and expands the network scale of the AI data center.

Citation Information

Patent Citations

  • Optical module physical link detection method and device, optical module and optical transmission system

    CN113315572A

  • Early warning method and device of optical fiber link and optical transmission system

    CN114759979A

  • Automatic testing method for FTTR optical link

    CN120639173A

  • PON network fault detection method, apparatus and device, and storage medium

    CN120751298A

  • Fault root cause location method and apparatus for fronthaul link

    WO2024066771A1