Optical Module Health Status Detection Method, Device, Storage Medium and Program Product

By collecting the timing data and system log data of the optical module, building a hybrid model for multi-dimensional feature fusion, and dynamically adjusting the threshold, the problem of difficult detection of the sub-health status of the optical module in the existing technology is solved, and the accurate determination of the healthy status of the optical module and the reduction of operation and maintenance costs are achieved.

CN119995710BActive Publication Date: 2025-07-11BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510451101.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing optical module monitoring methods rely on single-dimensional parameters and cannot effectively deal with the sub-health status of optical modules, resulting in reduced availability or abnormal performance of the Wanka cluster.

Method used

By collecting the timing data of the optical module and the system log data, extracting multi-dimensional features, building a hybrid model for fusion, generating fusion features, combining historical health scores and environmental parameters, dynamically adjusting the threshold for health status determination.

Benefits of technology

It realizes accurate determination of the health status of optical modules, and can discover sub-health status in advance, reduce operation and maintenance costs, and ensure business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119995710B_ABST
    Figure CN119995710B_ABST
Patent Text Reader

Abstract

An embodiment of the present disclosure discloses a method, device, storage medium, and program product for detecting the health status of an optical module. The method includes: extracting timing features from the timing data of the optical module and extracting key features from the system log data of the optical module; fusing the timing features and the key features based on a constructed hybrid model to generate fused features to obtain the current health score of the optical module; collecting the historical health scores of the optical module within a preset time range; detecting at least one environmental parameter of the environment where the optical module is located; obtaining a dynamic threshold of the optical module based on the historical health scores and the environmental parameters; and determining the health status of the optical module based on the relationship between the current health score and the dynamic threshold. This method can comprehensively, accurately, and efficiently evaluate the health status of the optical module based on multi-dimensional data features, provide strong support for the operation and maintenance management of the optical module, and is of great significance for ensuring service continuity and reducing operation and maintenance costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of device security detection, and particularly to a method, device, storage medium and program product for detecting the health status of an optical module. Background Art

[0002] In a computing center, an optical module, as a key device for optoelectronic conversion, is widely used for high-speed and long-distance data transmission between servers, switches and storage devices, and is a core component for realizing the efficient operation of a ten-thousand-card cluster. A ten-thousand-card cluster usually contains tens of thousands of optical modules, and their stability and health status are directly related to the performance and reliability of the entire system.

[0003] However, existing optical module monitoring methods mostly rely on single-dimensional parameters, such as optical power and temperature, etc., and monitor and manage the status of optical modules through a fixed threshold mechanism and a static model. Although this method can ensure the efficient operation and stability of the system to a certain extent, it cannot effectively cope with the sub-healthy state of optical modules. In a training scenario, an abnormality of an optical module may cause the entire ten-thousand-card cluster to be unavailable, while the sub-healthy state may reduce the availability of the cluster or cause abnormal performance output. Summary of the Invention

[0004] In view of this, the embodiments of the present disclosure provide a method, device, storage medium and program product for detecting the health status of an optical module, which can comprehensively, accurately and efficiently evaluate the health status of an optical module based on multi-dimensional data features, provide strong support for the operation and maintenance management of optical modules, and are of great significance for ensuring business continuity and reducing operation and maintenance costs.

[0005] In a first aspect, the embodiments of the present disclosure provide a method for detecting the health status of an optical module, adopting the following technical solutions:

[0006] Collect the timing data and system log data of the optical module;

[0007] Extract timing features from the timing data, and extract key features from the system log data;

[0008] Construct a hybrid model, and fuse the timing features and the key features based on the hybrid model to generate fused features;

[0009] Obtain the current health score of the optical module based on the fused features;

[0010] Collect the historical health scores of the optical module within a preset time range;

[0011] Detect at least one environmental parameter of the environment where the optical module is located;

[0012] Based on the historical health score and the environmental parameters, obtain the dynamic threshold of the optical module;

[0013] Based on the relationship between the current health score and the dynamic threshold, determine the health status of the optical module.

[0014] Optionally, the extracting the time series features from the time series data includes:

[0015] Obtain the mean time between failures (MTBF) of the optical module;

[0016] Based on the mean time between failures and a preset time, determine the size of the time window;

[0017] Perform a sliding operation on the time series data according to the size of the time window, and extract time series features from the data within the window.

[0018] Optionally, the extracting the key features from the system log data includes:

[0019] Traverse each log record in the system log data, and use the log_feature_extraction function to extract the key features from each log record;

[0020] Store the key features in a feature dictionary in the form of key-value pairs.

[0021] Optionally, the fusing the time series features and the key features based on the hybrid model to generate fused features includes:

[0022] Capture the long-term dependencies in the time series features through multi-layer dilated convolution operations to generate global features;

[0023] Perform a pooling operation on the global features to generate an aggregated feature vector;

[0024] Perform semantic encoding on the key features to generate an encoded feature sequence;

[0025] Perform semantic information mining on the encoded feature sequence to generate a comprehensive semantic feature vector;

[0026] Use a bidirectional cross-attention mechanism to fuse the aggregated feature vector and the comprehensive semantic feature vector to obtain the fused features.

[0027] Optionally, the obtaining the current health score of the optical module based on the fused features includes:

[0028] Map the fused features to a scalar value through a linear transformation;

[0029] The scalar value is converted into a current health score in percentage through a Sigmoid activation function.

[0030] Optionally, the optical module health status detection method further includes:

[0031] Deploy the initial hybrid model to the cloud server and each edge device as a global model and a local model respectively;

[0032] Offline training is performed on the global model based on historical data, the trained global model parameters are synchronized to each edge device, and the local model is updated;

[0033] When the edge device detects a new abnormality in the optical module, it collects abnormal data, uses the abnormal data to fine-tune the updated local model, and obtains an update gradient of the local model parameters;

[0034] Upload the updated gradients of each edge device to the cloud server for aggregation to generate aggregated gradients;

[0035] The global model is updated based on the aggregated gradient, the updated global model parameters are synchronized to each edge device, and the local model is updated again.

[0036] Optionally, the optical module health status detection method further includes:

[0037] Score the historical health and construct a historical scoring data set in order of size;

[0038] According to a preset percentile, a corresponding historical health score is selected from the historical score data set as a benchmark threshold, where the benchmark threshold is an initial dynamic threshold;

[0039] If there is an environmental parameter that deviates from the corresponding reference parameter value, an adjustment coefficient is calculated based on the deviating environmental parameter;

[0040] A new dynamic threshold is obtained based on the reference threshold and the adjustment coefficient.

[0041] Optionally, the calculation formula of the adjustment coefficient is:

[0042] ;

[0043] in, is the adjustment coefficient; Number the categories that deviate from environmental parameters; is the total number of deviations from environmental parameters; For the The weight of the deviation from the environmental parameters; For the An adjustment factor for deviations from environmental parameters; is the current measured value of the n-th deviation from the environmental parameter; is the reference parameter value of the n-th deviation from the environmental parameter; is the scaling factor of the n-th deviation from the environmental parameter.

[0044] In a second aspect, an embodiment of the present disclosure also provides an optical module health status detection system, which adopts the following technical solutions:

[0045] A data collection module, configured to collect timing data and system log data of the optical module;

[0046] A feature extraction module, configured to extract timing features from the timing data and key features from the system log data;

[0047] A feature fusion module, configured to build a hybrid model and fuse the timing features and the key features based on the hybrid model to generate fused features;

[0048] A score acquisition module, configured to obtain the current health score of the optical module based on the fused features;

[0049] A score collection module, configured to collect the historical health scores of the optical module within a preset time range;

[0050] An environment detection module, configured to detect at least one environmental parameter of the environment where the optical module is located;

[0051] A threshold acquisition module, configured to obtain the dynamic threshold of the optical module based on the historical health scores and the environmental parameters;

[0052] A health judgment module, configured to determine the health status of the optical module based on the relationship between the current health score and the dynamic threshold.

[0053] In a third aspect, an embodiment of the present disclosure also provides a computer device, which adopts the following technical solutions:

[0054] The computer device includes:

[0055] At least one processor; and,

[0056] A memory communicatively connected to the at least one processor; wherein,

[0057] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the optical module health status detection method described above.

[0058] Fourthly, embodiments of the present disclosure further provide a computer-readable storage medium storing computer instructions for causing a computer to execute the optical module health status detection method described in any one of the above.

[0059] Fifthly, embodiments of the present disclosure further provide a computer program product including computer programs / instructions that, when executed by a processor, implement the steps of the method described in any one of the above.

[0060] The optical module health status detection method provided by the embodiments of the present disclosure simultaneously collects timing data and system log data of the optical module, and can obtain the operation information of the optical module from multiple dimensions. This multi-source data collection method ensures comprehensive monitoring of the operating state of the optical module and avoids misjudgment caused by insufficient information from a single data source. Extracting timing features from the timing data and key features from the system log data can deeply mine the useful information in the data. This process of feature extraction is actually a dimensionality reduction process and statistical process on the original data, removing a large amount of redundant information and giving statistical features by combining partial information, making subsequent analysis and processing more efficient. This not only reduces the consumption of computing resources, but also speeds up the detection speed and improves the real-time performance of the system. Constructing a hybrid model to fuse the timing features and key features can give full play to the advantages of different types of features. Obtaining the current health score of the optical module based on the fused features quantifies the health status of the optical module. At the same time, a dynamic threshold determination mechanism is adopted. The dynamic threshold is specifically set for each optical module in combination with the historical health score of the optical module and environmental parameters. Using the dynamic threshold to determine the health status of the optical module can be flexibly adjusted according to the actual operating conditions and environmental changes of the optical module. The accurate determination of the health status of the optical module by this method can detect the sub-healthy state of the optical module in advance, and timely take preventive maintenance measures before the optical module fails severely, ensuring business continuity while reducing operation and maintenance costs.

[0061] The above description is only an overview of the technical solutions of the present disclosure. In order to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the description. In order to make the above and other purposes, features, and advantages of the present disclosure more obvious and understandable, the following preferred embodiments are specifically given and described in detail in conjunction with the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0063] Figure 1 Flow diagram of the optical module health status detection method provided by the embodiments of the present disclosure;

[0064] Figure 2 Flow diagram of the timing feature acquisition method provided by the embodiments of the present disclosure;

[0065] Figure 3 Flow diagram of the method for fusing timing features and key features provided by the embodiments of the present disclosure;

[0066] Figure 4 Flow diagram of the method for obtaining the current health score provided by the embodiments of the present disclosure;

[0067] Figure 5 Flow diagram of the hybrid model update method provided by the embodiments of the present disclosure;

[0068] Figure 6 Flow diagram of the method for obtaining the dynamic threshold provided by the embodiments of the present disclosure;

[0069] Figure 7 Principle block diagram of the optical module health status detection system provided by the embodiments of the present disclosure;

[0070] Figure 8 Structure diagram of a computer device provided by the embodiments of the present disclosure. Detailed implementation manners

[0071] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0072] It should be clear that the embodiments of the present disclosure are specifically illustrated by the following specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0073] It should be noted that the following description relates to various aspects of embodiments within the scope of the appended claims. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of the aspects set forth herein can be used to implement a device and / or practice a method. Additionally, this device can be implemented and this method can be practiced using other structures and / or functionality in addition to one or more of the aspects set forth herein.

[0074] It should also be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present disclosure. Only the components related to the present disclosure are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0075] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the aspects described can be practiced without these specific details.

[0076] Referring to Figure 1 , the present disclosure provides a method for detecting the health status of an optical module, including the following steps:

[0077] S1: Collect the timing data and system log data of the optical module;

[0078] S2: Extract timing features from the timing data and key features from the system log data;

[0079] S3: Construct a hybrid model, and fuse the timing features and key features based on the hybrid model to generate fused features;

[0080] S4: Obtain the current health score of the optical module based on the fused features;

[0081] S5: Collect the historical health scores of the optical module within a preset time range;

[0082] S6: Detect at least one environmental parameter of the environment where the optical module is located;

[0083] S7: Obtain the dynamic threshold of the optical module based on the historical health scores and environmental parameters;

[0084] S8: Judge the health status of the optical module based on the relationship between the current health score and the dynamic threshold.

[0085] The optical module health status detection method provided by the present disclosure collects the timing data and system log data of the optical module simultaneously, and can obtain the operation information of the optical module from multiple dimensions. The timing data reflects the dynamic changes of the optical module over a period of time, such as the fluctuation of optical power over time, the real-time change of temperature, etc.; the system log data records various events and abnormal information during the operation of the optical module, including error messages, configuration changes, etc. This multi-source data collection method ensures the comprehensive monitoring of the operation status of the optical module and avoids misjudgment caused by insufficient information from a single data source. Compared with the traditional method based on experience or simple parameter judgment, this method based on actual operation data can more accurately reflect the actual state of the optical module and reduce the interference of human factors and subjective judgment.

[0086] Extracting timing features from the timing data and key features from the system log data can deeply mine the useful information in the data. Timing features can capture the dynamic characteristics such as the periodicity and trend of the optical module operation, while key features focus on the important information in the system log, such as specific error codes, abnormal events, etc. These information play a key guiding role in judging the health status of the optical module. The process of feature extraction is actually a dimensionality reduction process and statistical process for the original data, removing a large amount of redundant information and giving statistical features by combining partial information, making the subsequent analysis and processing more efficient. This not only reduces the consumption of computing resources, but also speeds up the detection speed and improves the real-time performance of the system.

[0087] Constructing a hybrid model to fuse the timing features and key features can give full play to the advantages of different types of features. The timing features reflect the dynamic changes of the optical module, while the key features reflect the abnormal conditions during the system operation. Fusing the two can more comprehensively and accurately describe the health status of the optical module. For example, combining the timing feature of optical power and the key feature of optical signal loss in the system log can more accurately judge whether there is a potential fault in the optical module. The fused features have stronger expressive power than single timing features or key features. It can capture the associations and interactions between different features, thus providing richer and more valuable information for obtaining the current health score subsequently and making the score result more accurate and reliable.

[0088] Based on the fusion features, the current health score of the optical module is obtained, and the health status of the optical module is quantitatively represented. This quantitative evaluation method makes the health status of the optical module more intuitive and clear, facilitating the management and decision-making of the operation and maintenance personnel. At the same time, a dynamic threshold determination mechanism is adopted. This dynamic threshold is specifically set for each optical module by combining the historical health score of the optical module and the environmental parameters. Using the dynamic threshold to determine the health status of the optical module can be flexibly adjusted according to the actual operating conditions and environmental changes of the optical module. In different application scenarios, the normal operating standards of different optical modules may vary. The dynamic threshold can better adapt to this difference, avoiding misjudgments that may be caused by the traditional fixed threshold method and improving the accuracy and reliability of the status determination.

[0089] In summary, through the accurate determination of the health status of the optical module, the sub-health status of the optical module can be detected in advance. Before the optical module has a serious failure, preventive maintenance measures can be taken in a timely manner, such as replacing components and performing debugging, to avoid the impact of the optical module failure on the business and ensure the continuity of the business. At the same time, preventive maintenance can reduce the operation and maintenance costs and reduce the downtime and repair costs caused by the optical module failure.

[0090] In S1, the timing data (also known as time series data) of the optical module is collected in real time through the Digital Optical Monitoring (DOM) interface. The timing data refers to the operating parameters of the optical module recorded in chronological order, usually sampled and recorded at a fixed time interval (such as 1 second), reflecting the real-time status of the optical module at different time points. These data mainly include the Digital Diagnostic Monitoring (DDM) parameters of the optical module, such as optical power, operating temperature, and supply voltage. Among them, the optical power includes the transmit power (Tx Power) and receive power (Rx Power) of the optical module, with an accuracy of ±0.1 dBm, that is, the error range between the measured value and the true value does not exceed ±0.1 decibel milliwatt (dBm); the operating temperature refers to the internal temperature of the optical module, with an accuracy of ±0.5 °C, that is, the error range between the measured value and the true temperature does not exceed ±0.5 degrees Celsius; the supply voltage refers to the supply voltage of the optical module, with an accuracy of ±1 millivolt (mV), that is, the error range between the measured value and the true voltage does not exceed ±1 mV.

[0091] At the same time, log collection tools such as Flume or Rsyslog or the tail_logs function are used to collect system log data. The system log data refers to the event records and related content generated during the operation of the optical module or its affiliated equipment, usually generated by the device management system or monitoring software. These data not only record the event information such as the operating status, fault alarms, configuration changes, and operation logs of the optical module and its related equipment, but also cover multi-dimensional information such as performance indicators, resource utilization, user behavior, and security events.

[0092] Time series data is mainly used for performance monitoring and fault prediction, and potential problems can be detected in advance by analyzing the changing trends of parameters. System log data is mainly used for fault troubleshooting and operation and maintenance management. By recording events and performance monitoring information, it provides detailed clues to problems and operation history, helping to quickly locate the root cause of problems. By combining time series data and system log data, key parameters such as voltage fluctuations and changes in bit error rate gradients can be integrated, and quantitative modeling can be performed on events such as link oscillations and sudden increases in CRC errors in system logs, so as to achieve comprehensive monitoring from the underlying hardware to the upper-layer applications. It can not only monitor the operating status of optical modules in real time, but also supplement other key information of optical modules through the event information and performance metrics recorded in the logs, thus avoiding difficulties in problem troubleshooting and delays in fault response caused by a single monitoring dimension.

[0093] In S2, referring to Figure 2 the flowchart of the time series feature acquisition method shown, "extracting time series features from time series data" includes the following steps:

[0094] S21: Obtain the mean time between failures (MTBF) of the optical module;

[0095] S22: Determine the time window size based on the mean time between failures and a preset time;

[0096] S23: Perform a sliding window operation on the time series data according to the time window size, and extract time series features from the data within the window.

[0097] In S21, the mean time between failures (MTBF) of the optical module can be obtained through the following methods: For general types of optical modules, the existing average data in the industry or the MTBF value of similar products can be directly referred to; Optical module manufacturers usually indicate the MTBF value of this type of optical module in the product technical manual or specification sheet, which can be directly consulted and obtained; If there are a large number of actual usage records of this optical module, the MTBF can be calculated by counting the number of failures of the optical module within a certain period of time. The specific method is to divide the total running time by the number of failures, thereby obtaining the mean time between failures.

[0098] In S22, by comparing the MTBF of the optical module with the preset time, the maximum value of the two is selected as the time window size. This method comprehensively considers the requirements of long-term reliability and short-term anomaly detection. When the MTBF is large, it indicates that the optical module itself has high reliability, and using a larger time window can more comprehensively capture its characteristics of long-term stable operation; while when the MTBF is small, the preset time can prevent the time window from being too small and losing key trend information, thus ensuring that the performance of the optical module can be effectively analyzed in different situations.

[0099] In S23, set the sliding step size according to actual requirements. Then, based on the determined time window size and sliding step size, perform a sliding operation on the timing data of the optical module. Each time, extract the data within the window for feature extraction to obtain timing features. The timing features include statistical features and frequency-domain features. The statistical features include the mean, standard deviation (or variance), slope, skewness, and kurtosis of various types of data within the window. Among them, the mean is used to measure the absolute level of the data during this time period; the standard deviation or variance is used to measure the fluctuation range of the data during this time period; the slope is calculated by fitting the linear trend of the data within the window, and the positive and negative slopes respectively reflect the upward or downward trend of the data, reflecting the change trend of the recent data; skewness and kurtosis are used to reflect the distribution form of the data, skewness measures the asymmetry degree of the data distribution, and kurtosis measures the sharp or flat degree of the data distribution.

[0100] The frequency-domain features refer to the frequency-domain representation of the key performance data of the optical module, covering the frequency-domain features of parameters such as the optical power, operating temperature, bias current, and voltage fluctuation of the optical module. Specifically, they include the main frequency component and its amplitude, the proportion of signal energy in different frequency bands, and the frequency peak characteristics. By performing a fast Fourier transform (FFT) on the data within the window, the main frequency component and its corresponding amplitude of the signal are obtained, which are used to capture the periodic changes in the data; analyze the distribution of signal energy in different frequency bands after the FFT transformation to understand the frequency component distribution of the data, that is, the proportion of signal energy in different frequency bands; focus on the significant peaks on the spectrum to identify the periodic oscillations of the optical module parameters. For example, if there are significant peaks in the spectrum of the optical power of the optical module, it may mean that there is a stable oscillation source.

[0101] The extracted timing features can capture different change patterns of the optical module parameters, including short-term anomalies (such as sudden spikes), medium-term drifts (such as gradual attenuation), and long-term cycles (such as day-night environmental cycles), etc., providing a basis for the performance analysis and fault prediction of the optical module.

[0102] The system log data includes multiple log records, and the log_feature_extraction function (a function for extracting key features) is used to extract key features from each log record. The key features include the CRC (Cyclic Redundancy Check) error rate, the LOS (Loss of Signal) duration, and the reset trend. Among them, the CRC error rate refers to the frequency of errors caused by failed verification during data transmission; the LOS duration refers to the length of time the signal loss state persists; and the reset trend refers to the frequency and pattern of device reset operations. The extracted key features are stored in the feature dictionary in the form of key-value pairs, where the key is the feature name (such as 'crc_error_rate', 'los_duration','reset_trend'), and the value is the corresponding feature value.

[0103] In S3, the hybrid model includes a TCN time series network module, a Transformer encoding module, a cross-modal attention fusion module, and a health score module. Through the effective combination of these modules, not only can the effective fusion of time series features and key features be achieved, but also the health of the optical module can be evaluated to generate the corresponding health score. And during the inference process, TensorRT can also be used to accelerate and optimize the hybrid model, reducing the inference time and improving the inference efficiency, so that the current health score can be obtained more quickly.

[0104] Refer to Figure 3 Referring to the flowchart of the method for fusing time series features and key features shown, "fusing time series features and key features based on the hybrid model to generate fused features" includes the following steps:

[0105] S31: Capture the long-term dependencies in the time series features through multi-layer dilated convolution operations to generate global features;

[0106] S32: Perform a pooling operation on the global features to generate an aggregated feature vector;

[0107] S33: Perform semantic encoding on the key features to generate an encoded feature sequence;

[0108] S34: Mine semantic information from the encoded feature sequence to generate a comprehensive semantic feature vector;

[0109] S35: Use a bidirectional cross-attention mechanism to fuse the aggregated feature vector and the comprehensive semantic feature vector to obtain the fused features.

[0110] In S31 and S32, the TCN (Temporal Convolutional Network) temporal network module is used to process temporal features. This module includes a first output layer, at least three dilated convolutions with different dilation rates, and a global pooling layer. The first output layer is used to receive temporal features as input and transmit them to the first layer of dilated convolution. When processing temporal data, the dilated convolution adopts the form of one-dimensional convolution (Conv1D). By introducing the dilation rate, without increasing the number of parameters and computational complexity, the receptive field of the convolutional kernel is expanded, thereby effectively capturing the dependencies in temporal data. Among them, the dilated convolution with a low dilation rate is used to capture short-term fluctuations in temporal features (such as second-level changes), extract local temporal features, and reflect the rapid changes in data within a short period of time (such as within a few seconds); the dilated convolution with a medium dilation rate is used to extract medium-term trends in temporal features (such as minute-level changes), and can capture the dynamic changes in data on a medium time scale; the dilated convolution with a high dilation rate is used to capture long-term dependencies in temporal features (such as hourly periodicity), and can identify the periodicity and long-term trends in data over a long time span. For example, if four layers of dilated convolutions are set, the first layer of dilated convolution uses a convolutional kernel with a dilation rate of 1 (i.e., dilation = 1), which extracts local temporal features and outputs a preliminary local feature representation; the second layer of dilated convolution uses a convolutional kernel with a dilation rate of 2 (i.e., dilation = 2), expands the receptive field, captures features on a longer time scale, and outputs a feature representation on a medium time scale; the third layer of dilated convolution uses a convolutional kernel with a dilation rate of 4 (i.e., dilation = 1), further expands the receptive field, extracts a wider range of temporal dependencies, and outputs a long temporal feature representation; the fourth dilated convolution uses a convolutional kernel with a dilation rate of 8 (i.e., dilation = 1), captures global temporal features through a larger dilation rate, and outputs a global temporal feature representation, simply referred to as global features. The global pooling layer performs a global pooling operation on the global features, compresses the feature dimension, and generates a compact aggregated feature vector.

[0111] Through multiple dilated convolution operations, the TCN temporal network module can not only effectively capture multi-scale dependencies in temporal data, enhance the model's perception ability of complex temporal patterns, but also significantly expand the receptive field by gradually increasing the dilation rate (e.g., growing exponentially by 2). This structure enables the TCN to capture long-term dependencies without significantly increasing the computational cost, thereby better handling long-distance temporal dependencies. In addition, the gradual increase of the dilated convolution also allows the model to achieve a large receptive field in a relatively shallow network structure, thus avoiding the problem of gradient disappearance caused by too many layers in traditional convolutional networks while maintaining efficient computation.

[0112] The global pooling operation further compresses the global features into an aggregated feature vector, reducing the feature dimension and improving the computational efficiency of the model. This compact feature representation not only provides high-quality temporal features for subsequent feature fusion but also enhances the model's ability to analyze the temporal data of optical modules, contributing to a more accurate assessment of the health status of optical modules. In addition, this design of TCN also has the advantage of parallel computing, which can fully utilize the acceleration capabilities of modern hardware and is suitable for large-scale data processing.

[0113] In S33 and S34, the Transformer encoding module is used to process the key features. This module includes a second input layer, a Token embedding layer, a positional encoding layer, a multi-head attention layer, a feed-forward network layer, and a feature vector layer. The second input layer receives the key features as input and transmits them to the Token embedding layer. The Token embedding layer uses a pre-trained BERT model to map each Token (unit) in the key features to a feature vector of a preset dimension (e.g., 768 dimensions). The positional encoding layer adds positional information to the Token embedding, preserving the order of the sequence and generating a feature vector with positional information, i.e., the encoded feature sequence. The multi-head attention layer captures the dynamic associations between Tokens in the encoded feature sequence, extracting the complex relationships between features. The feed-forward network layer introduces a non-linear transformation to further extract the abstract representation of the features, increasing the non-linear representation ability of the model. The feature vector layer integrates the feature sequence processed by the multi-head attention and the feed-forward network to generate a comprehensive semantic feature vector.

[0114] In S35, the cross-modal attention fusion module is used to fuse the aggregated feature vector and the comprehensive semantic feature vector. Its structure includes an attention calculation layer, a residual connection layer, and a layer normalization layer. The attention calculation layer contains at least one set of bidirectional cross-attention sub-layers. Each set takes the aggregated feature vector and the comprehensive semantic feature vector as the query vector (Query) and the key-value vector (Key / Value), respectively. Specifically, on the one hand, the aggregated feature vector is used as the query vector, and the comprehensive semantic feature vector is used as the key-value vector to calculate the "temporal attention log features". On the other hand, the comprehensive semantic feature vector is used as the query vector, and the aggregated feature vector is used as the key-value vector to calculate the "log attention temporal features". This bidirectional interaction mechanism is implemented through the multi-head attention mechanism, which can dynamically align and fuse the information of the two modalities in two directions. Finally, the results of multiple sets of bidirectional interactions are concatenated or weighted and summed to generate the concatenated features.

[0115] The splicing feature is further subjected to a residual connection with the original input (the aggregated feature vector and the comprehensive semantic feature vector) through a residual connection layer to ensure the direct transmission of information. Subsequently, the result after the residual connection process enters the layer normalization layer for normalization operations to stabilize the training process. The finally obtained fused feature representation is a comprehensive vector containing temporal and log information, which is used for subsequent health assessment.

[0116] This fusion method utilizes a bidirectional cross-attention mechanism, which can dynamically align the information of the two modalities to achieve information complementarity. At the same time, it combines the multi-head attention mechanism, which can enhance the richness and flexibility of the feature representation. The residual connection and layer normalization improve the training stability of the model. These designs enable the fused feature vector to provide a more comprehensive and accurate input for health assessment, thereby enhancing the model's ability to evaluate the health status of optical modules.

[0117] In S4, referring to Figure 4 the flowchart of the current health score acquisition method shown, "obtaining the current health score of the optical module based on the fused feature" includes the following steps:

[0118] S41: Map the fused feature to a scalar value through a linear transformation;

[0119] S42: Convert the scalar value to a current health score in percentage system through the Sigmoid activation function.

[0120] Specifically, the health score module is used to process the fused feature and generate the current health score (Health Index) of the optical module. This module includes a regression calculation layer and an activation mapping layer. Among them, the regression calculation layer can select a fully connected layer (such as nn.Linear), and map the input fused feature to a scalar value through a series of linear transformations. This scalar value reflects the "degree of deviation" of the optical module from the ideal health state, which is a continuous value. The higher the score, the more likely there are performance degradations or potential faults in the optical module, and the lower the score, the more normal the module works. To make the score more intuitive, the activation mapping layer uses the Sigmoid activation function to map the output of the regression calculation layer to the range of 0 - 1, and then multiplies the mapped result by 100 to obtain the current health score in percentage system.

[0121] Considering that the operating environment and data distribution of optical modules may evolve over time, and static models are difficult to adapt to the parameter differences of multi-vendor modules, in order to keep the hybrid model effective in the long term, the hybrid model needs to be dynamically updated. Referring to Figure 5 the flowchart of the hybrid model update method shown, the method for training and dynamically updating the hybrid model includes the following steps:

[0122] S43: Deploy the initial hybrid model to the cloud server and each edge device as the global model and local model respectively;

[0123] S44: Conduct offline training on the global model based on historical data, synchronize the parameters of the trained global model to each edge device, and update the local model;

[0124] S45: When an edge device detects a new anomaly in the optical module, collect the anomaly data, and use the anomaly data to fine-tune the updated local model to obtain the update gradient of the local model parameters;

[0125] S46: Upload the update gradients of each edge device to the cloud server for aggregation to generate an aggregated gradient;

[0126] S47: Update the global model based on the aggregated gradient, synchronize the parameters of the updated global model to each edge device, and update the local model again.

[0127] In S43, the code and parameters of the initial hybrid model are transmitted to the cloud server and each edge device through the network respectively, and the corresponding running environments are configured on the cloud server and edge devices. A global model is built on the cloud server for centralized optimization and management; local models are built on each edge device for real-time inference and local fine-tuning to ensure that the models can run properly in their respective environments.

[0128] In S44, the cloud server collects historical data (such as a large amount of historical data of 50,000 optical modules from 10 data centers), and uses this data to conduct offline training on the global model to optimize the model parameters. After the training is completed, the cloud server synchronizes the parameters of the global model to each edge device in the form of a binary file or through network transmission. After the edge device receives and loads these parameters, it updates the local model.

[0129] In S45 - S47, an online incremental update technology is adopted. When an edge device detects a new anomaly, relevant data is collected for the new anomaly, and this data is used to perform a small amount of fine-tuning on the local model to improve the sensitivity of the local model to similar situations. Through this continuous learning, the local model can "keep up with the times" and continuously improve its generalization performance and accuracy.

[0130] Meanwhile, during the fine-tuning process, the updated gradients are calculated and uploaded to the cloud through a secure network channel. After the cloud server receives the updated gradients uploaded by each edge device, it aggregates the gradients using a federated learning aggregation algorithm (such as the FedAvg algorithm) to generate aggregated gradients. The cloud server uses the aggregated gradients to update the global model and adjusts the parameters of the global model to improve performance. Subsequently, the updated global model parameters are synchronized to each edge device, and the edge device updates the local model after receiving the parameters, completing one round of model update process.

[0131] By combining federated learning and incremental learning, the model can continuously learn new data and abnormal patterns, gradually adapt to new module batches, new environmental conditions, and potential new failure modes, thereby significantly improving the generalization performance of the model and its adaptability to different scenarios. This dynamic update mechanism not only improves the real-time performance and accuracy of the model but also optimizes the performance of the global model through the federated average algorithm (FedAvg), ensuring the stability and reliability of the system when facing new challenges.

[0132] Furthermore, when introducing a new module model or cross-manufacturer device, based on the existing hybrid model and with the help of a small amount of new data for fine-tuning, a new hybrid model adapted to the new scenario can be quickly derived. This transfer learning method does not require training from scratch, significantly reducing the training time and computational resource requirements. It not only ensures that the solution can support device heterogeneity but also enhances the flexibility in practical applications.

[0133] Meanwhile, the model update process is automatically performed in the background without interfering with the real-time monitoring in the foreground. Through version management and sufficient testing, it can be ensured that the performance of the updated model does not degrade, and it can be rolled back to the previous version if necessary. As time goes by, the knowledge accumulated by the system becomes increasingly rich, and the discrimination ability of the hybrid model for the sub-healthy state of the optical module will gradually become more perfect.

[0134] In S5 - S8, when the current health score is greater than the dynamic threshold, the optical module is determined to be sub-healthy; when the current health score is not greater than the dynamic threshold, the optical module is determined to be healthy.

[0135] Traditional fixed-threshold mechanisms are insensitive to the detection of progressive degradation and lack timely and effective detection means. Therefore, the optical module health state detection method of the present disclosure introduces an adaptive threshold mechanism to adjust the alarm threshold of the current health score according to historical statistics and environmental changes. Referring to Figure 6 the flow schematic diagram of the dynamic threshold acquisition method shown, the method of "acquiring the dynamic threshold of the optical module based on historical health scores and environmental parameters" includes the following steps:

[0136] S71: Construct a historical score data set by arranging the historical health scores in ascending order;

[0137] S72: Select the corresponding historical health score from the historical score dataset as the baseline threshold according to the preset percentile. The baseline threshold is the initial dynamic threshold.

[0138] S73: If there is an environmental parameter deviating from the corresponding baseline parameter value, calculate the adjustment coefficient based on the deviating environmental parameter.

[0139] S74: Obtain a new dynamic threshold based on the baseline threshold and the adjustment coefficient.

[0140] In S71, the system continuously maintains the historical distribution of health scores within a preset time range (such as 7 days), that is, sorts the collected historical health scores in ascending or descending order, and combines the sorted historical health scores into a historical score dataset.

[0141] In S72, if the collected historical health scores are sorted in ascending order, the percentile is the Nth percentile; if the collected historical health scores are sorted in descending order, the percentile is the (100 - N)th percentile. In this way, it means that under normal circumstances, only N% of the time, the historical health score will exceed this value. N can take the value of 90. Through this method, the initially obtained baseline threshold is the initial dynamic threshold (Threshold).

[0142] In S73 and S74, through a sensor network deployed in the environment where the optical module is located, at least one environmental parameter is collected in real time. These parameters include at least one of the device load rate, power supply voltage stability value, computer room temperature, computer room humidity, fan speed, and power supply status value. Among them, the power supply voltage stability value is the result of quantitatively evaluating the power supply voltage stability, and the power supply status value is the result of quantitatively evaluating the power supply status. The baseline parameter value of each environmental parameter is set in advance. For example, for the rack temperature, its baseline parameter value is set to 25°C, and for the computer room humidity, its baseline parameter value is set to 50%. After the environmental parameter is collected, it is compared with the corresponding baseline parameter value. If the two are inconsistent, it indicates that the environmental parameter has deviated; if they are consistent, it indicates that the environmental parameter has not deviated.

[0143] When it is detected that the environmental parameter deviates from its baseline value, a task to update the dynamic threshold is triggered. At this time, the adjustment coefficient is calculated according to the deviating environmental parameter, and the adjustment coefficient is multiplied by the baseline threshold. The resulting product is the new dynamic threshold. In the next round of environmental detection, if it is found that the new environmental parameter deviates and the historical score dataset has been updated, a new baseline threshold is selected from the new historical score dataset, and the dynamic threshold is recalculated accordingly. Among them, the calculation formula for the adjustment coefficient is:

[0144] ;

[0145] In the formula, is the adjustment coefficient; is the type number of the environmental parameter deviation; is the total quantity of the environmental parameter deviation; is the weight of the th type of environmental parameter deviation. The sum of the weights of all environmental parameters (whether there is deviation or not) should be equal to 1. When a certain environmental parameter does not deviate, the value after subtracting the corresponding reference parameter value is 0 and it can be excluded from the adjustment coefficient calculation; is the adjustment factor of the th type of environmental parameter deviation. This adjustment factor is used to control the sensitivity of the environmental parameter deviation to the dynamic threshold. For example, for the rack temperature, its adjustment factor is 0.05, and for the computer room humidity, its adjustment factor is 0.03; is the current measured value of the th type of environmental parameter deviation; is the reference parameter value of the th type of environmental parameter deviation; is the scaling factor of the

[0146] By introducing a threshold dynamic update mechanism, the system can flexibly adjust the threshold according to different operating conditions. For example, in a high-temperature and high-load environment, the degree of deviation of the optical module's health from the normal range may be slightly higher. At this time, the threshold increase can avoid false alarms; while in a well-cooled and business-idle condition, even a slight anomaly is worthy of attention. At this time, the threshold is lowered to improve sensitivity. The dynamic threshold will be recalculated and updated regularly based on new data to ensure that the alarm criteria are always consistent with the latest historical distribution of health scores and environmental conditions. This mechanism overcomes the defect that the fixed threshold is insensitive to gradual degradation, can detect the early signs of health decline earlier, and at the same time reduces the false alarm rate caused by environmental fluctuations.

[0147] Binary alarms can be achieved through the dynamic threshold and the current health score. When it is determined that the optical module is in a sub-healthy state, a maintenance work order is generated based on the detection results, which is used to arrange professional technicians to conduct targeted inspections on the optical module, check for potential fault hazards, and evaluate whether maintenance operations such as component replacement, cleaning, or performance optimization are required to ensure that the optical module can quickly return to a healthy and stable working state and reduce the risk of system failures caused by the performance degradation of the optical module.

[0148] When it is determined that the optical module is in a healthy state, new timing data and system log data are collected for a new round of detection. At the same time, the dynamic threshold is maintained to ensure its timeliness, accuracy, and adaptability. That is, the dynamic threshold is continuously updated and adjusted according to the newly collected data so that it can accurately reflect the normal performance range of the optical module under different operating conditions and working stages, enabling more sensitive and accurate identification of changes in the optical module state and improving the reliability and effectiveness of alarms.

[0149] Optionally, the health scores of all monitored optical modules are sorted from high to low and displayed, which is convenient for operation and maintenance personnel to prioritize attention to high-risk modules. This is the horizontal display. At the same time, for each optical module, the change trend of its health score is displayed in chronological order, which is convenient for observing the speed of health deterioration of a single module. This is the vertical display. The horizontal display method and the vertical display method can be organically combined and presented in the same chart, which is convenient for operation and maintenance personnel to comprehensively and intuitively compare the differences and connections of data in the horizontal and vertical dimensions.

[0150] In summary, this solution monitors multi-dimensional data of the optical module and uses corresponding methods for feature mining for different types of data, thereby maximizing the data value. At the same time, this also facilitates the hybrid model to more efficiently process and analyze data. The hybrid model can be dynamically updated, accurately identify potential problems and failure trends of the optical module, enable operation and maintenance personnel to take preventive maintenance measures in advance, and greatly reduce the operation and maintenance cost. Through preventive maintenance measures, the continuity of business can be effectively guaranteed, and economic losses caused by optical module failures can be avoided. From data monitoring to visual display, the entire process forms a complete and efficient closed loop, providing a solid guarantee for the stable operation of the optical module and the continuous development of the business.

[0151] Refer to Figure 7 , the present disclosure provides an optical module health status detection system, including:

[0152] A data collection module 101 for collecting timing data and system log data of the optical module;

[0153] A feature extraction module 102 for extracting timing features from the timing data and key features from the system log data;

[0154] A feature fusion module 103 for constructing a hybrid model and fusing the timing features and key features based on the hybrid model to generate fusion features;

[0155] A score acquisition module 104 for obtaining the current health score of the optical module based on the fusion features;

[0156] A score collection module 105 for collecting the historical health scores of the optical module within a preset time range;

[0157] An environmental detection module 106 for detecting at least one environmental parameter of the environment where the optical module is located;

[0158] A threshold acquisition module 107 for acquiring a dynamic threshold of the optical module based on historical health scores and environmental parameters;

[0159] A health judgment module 108 for determining the health status of the optical module based on the relationship between the current health score and the dynamic threshold.

[0160] The various variation methods and specific examples in the above-provided optical module health status detection method are equally applicable to the optical module health status detection system provided in this disclosure. Through the foregoing detailed description of the optical module health status detection method, those skilled in the art can clearly know the implementation method of the optical module health status detection system. For the sake of simplicity of the specification, it will not be elaborated here.

[0161] A computer device according to an embodiment of the present disclosure includes a memory and a processor. The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0162] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In an embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the computer device executes all or part of the steps of the optical module health status detection method in the foregoing embodiments of the present disclosure.

[0163] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain a good user experience effect, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included in the protection scope of the present disclosure.

[0164] As Figure 8 FIG. is a schematic structural diagram of a computer device provided in an embodiment of the present disclosure. It shows a schematic structural diagram of a computer device suitable for implementing the computer device in the embodiment of the present disclosure. Figure 8 The shown computer device is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.

[0165] AsFigure 8 As shown, a computer device may include a processor (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the computer device are also stored. The processor, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0166] Generally, the following devices may be connected to the I / O interface: an input device including, for example, a sensor or a visual information acquisition device, etc.; an output device including, for example, a display screen, etc.; a storage device including, for example, a magnetic tape, a hard disk, etc.; and a communication device. The communication device may allow the computer device to communicate with other devices (such as edge computing devices) wirelessly or wiredly to exchange data. Although Figure 8 a computer device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0167] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via the communication device, or installed from a storage device, or installed from the ROM. When the computer program is executed by the processor, all or part of the steps of the optical module health status detection method according to the embodiments of the present disclosure are executed.

[0168] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details are not repeated here.

[0169] A computer-readable storage medium according to an embodiment of the present disclosure stores non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are run by a processor, all or part of the steps of the optical module health status detection method according to the foregoing embodiments of the present disclosure are executed.

[0170] The above-mentioned computer-readable storage medium includes but is not limited to: optical storage media (such as CD-ROMs and DVDs), magneto-optical storage media (such as MOs), magnetic storage media (such as magnetic tapes or external hard drives), media with built-in rewritable non-volatile memories (such as memory cards), and media with built-in ROMs (such as ROM cartridges).

[0171] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0172] The basic principles of the present disclosure have been described above in connection with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. Additionally, the specific details disclosed above are only for illustrative and facilitating understanding purposes, and not for limitation. These details do not limit the present disclosure to necessarily implementing with the above specific details.

[0173] In the present disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.

[0174] In addition, as used herein, the "or" used in the listing of items starting with "at least one" indicates a disjunctive listing, so that for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the term "exemplary" does not mean that the described examples are preferred or better than other examples.

[0175] It should also be noted that in the systems and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.

[0176] Various changes, substitutions, and alterations to the technology described herein may be made without departing from the teachings defined by the appended claims. Additionally, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Processes, machines, manufactures, compositions of events, means, methods, or acts that are currently existing or later to be developed that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein may be utilized. Accordingly, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.

[0177] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the disclosure. Thus, the disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0178] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A method for detecting the health status of an optical module, characterized in that, Including: Collecting the timing data and system log data of the optical module; Extracting timing features from the timing data and key features from the system log data; Constructing a hybrid model, fusing the timing features and the key features based on the hybrid model to generate fused features; Obtaining the current health score of the optical module based on the fused features; Collecting the historical health scores of the optical module within a preset time range; Detecting at least one environmental parameter of the environment where the optical module is located; Obtaining the dynamic threshold of the optical module based on the historical health scores and the environmental parameters; Wherein, obtaining the dynamic threshold of the optical module based on the historical health scores and the environmental parameters includes: Constructing a historical score dataset by arranging the historical health scores in ascending order; Selecting the corresponding historical health score from the historical score dataset as the benchmark threshold according to a preset percentile, and the benchmark threshold is the initial dynamic threshold; If there is an environmental parameter deviating from the corresponding benchmark parameter value, calculating an adjustment coefficient based on the deviating environmental parameter; Obtaining a new dynamic threshold based on the benchmark threshold and the adjustment coefficient; Wherein, the calculation formula of the adjustment coefficient is: ; wherein, is the adjustment coefficient; is the type number of the environmental parameter deviation; is the total number of environmental parameter deviations; is the weight of the th environmental parameter deviation; is the adjustment factor of the th environmental parameter deviation; is the current measured value of the th environmental parameter deviation; is the reference parameter value of the th environmental parameter deviation; is the scale factor of the Judging the health status of the optical module based on the relationship between the current health score and the dynamic threshold.

2. The method for detecting the health status of an optical module according to claim 1, wherein The extracting timing features from the timing data includes: Obtaining the mean time between failures of the optical module; Determining the time window size based on the mean time between failures and a preset time; Performing a sliding operation on the timing data according to the time window size, and extracting timing features from the data within the window.

3. The optical module health status detection method according to claim 1, wherein The extracting key features from the system log data includes: Traversing each log record in the system log data, and using the log_feature_extraction function to extract the key features in each log record; Storing the key features in a feature dictionary in the form of key-value pairs.

4. The optical module health status detection method according to any one of claims 1-3, wherein fusing the timing features and the key features based on the hybrid model to generate fused features includes: Capturing the long-term dependence relationship in the timing features through multi-layer dilated convolution operations to generate global features; Performing a pooling operation on the global features to generate an aggregated feature vector; Performing semantic encoding on the key features to generate an encoded feature sequence; Performing semantic information mining on the encoded feature sequence to generate a comprehensive semantic feature vector; Using a bidirectional cross-attention mechanism to fuse the aggregated feature vector and the comprehensive semantic feature vector to obtain the fused features.

5. The optical module health status detection method according to claim 4, wherein obtaining the current health score of the optical module based on the fused features includes: Mapping the fused features to a scalar value through a linear transformation; Converting the scalar value into a current health score in percentage system through a Sigmoid activation function.

6. The optical module health status detection method according to claim 1, further including: Deploy the initial hybrid model to the cloud server and each edge device as the global model and the local model respectively; Perform offline training on the global model based on historical data, synchronize the trained global model parameters to each edge device, and update the local model; When the edge device detects a new abnormality in the optical module, collect the abnormal data, and use the abnormal data to fine-tune the updated local model to obtain the update gradient of the local model parameters; Upload the update gradients of each edge device to the cloud server for aggregation to generate an aggregated gradient; Update the global model based on the aggregated gradient, synchronize the updated global model parameters to each edge device, and update the local model again.

7. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the optical module health status detection method according to any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, This computer-readable storage medium stores computer instructions for causing a computer to execute the optical module health status detection method according to any one of claims 1-6.

9. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Strong-robustness smart city edge computing data security system and method

    CN116204925A

  • Sensitive personal information identification method based on multi-scale feature dynamic fusion

    CN118551048A

  • AFC equipment health management method and system and intelligent operation and maintenance system

    CN118642928A

  • Power transmission and transformation equipment fault early warning system based on online monitoring

    CN119323003A