Multi-source heterogeneous data fusion driven optoelectronic device fault detection method and system

CN122824291APending Publication Date: 2026-09-25BEIJING XIANGHENG TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611110907.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]有鉴于此,本申请实施例提供了一种多源异构数据融合驱动的光电子器件故障检测方法及系统,以解决现有技术存在的多源数据时序错位、故障传播关系缺失、故障根因定位可信度不足的问题

Benefits of technology

通过获取目标光电子器件的多源运行数据及器件结构信息,基于功能单元间的物理传递关系构建多域物理拓扑,并生成各数据源的数据质量表征;根据多域物理拓扑中的响应传播关系及控制事件,对多源运行数据进行因果时序校正,生成因果对齐运行序列;将因果对齐运行序列与目标光电子器件的健康基线及多物理数字孪生模型的预测结果进行比较,生成各功能单元的多层故障残差;将多层故障残差映射至多域物理拓扑,结合数据质量表征及响应传播方向进行故障证据融合,生成表征异常起点、传播路径及可信度的故障证据图;将故障证据图与故障原型进行匹配,并根据匹配结果、不可解释残差及预测不确定度确定故障候选集合;当故障候选集合存在根因歧义时,利用多物理数字孪生模型对满足安全约束的候选探测动作进行反事实响应模拟,选择探测动作,并根据实际探测响应更新故障候选集合,输出故障检测结果。本申请能够提高多源数据对齐准确性、增强故障传播识别能力、提升故障根因定位可信度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824291A_ABST
    Figure CN122824291A_ABST
Patent Text Reader

Abstract

The application provides a multi-source heterogeneous data fusion driven optoelectronic device fault detection method and system. The method comprises: acquiring multi-source operation data and device structure information, constructing multi-domain physical topology, and forming data quality characterization; correcting operation data according to response propagation relationship and control event, forming causal alignment sequence; comparing the causal alignment sequence with the health baseline and the digital twin prediction result to generate multi-layer fault residual; mapping the residual to the multi-domain physical topology, combining the data quality and the propagation direction to fuse the evidence, and generating the fault evidence graph; matching the fault evidence graph with the fault prototype, determining the candidate set according to the matching result, the uninterpretable residual and the prediction uncertainty; when the root cause is ambiguous, simulating the counterfactual response of the candidate detection action, selecting the action, updating the candidate set according to the actual response, and outputting the fault detection result. The application can improve the multi-source data alignment accuracy, enhance the fault propagation identification ability, and improve the fault root cause positioning credibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of optoelectronic device condition monitoring and fault diagnosis technology, and in particular to a method and system for fault detection of optoelectronic devices driven by multi-source heterogeneous data fusion. Background Technology

[0002] With the development of high-speed optical communication, optoelectronic sensing, and photonic integration technologies, the structural complexity and integration density of optoelectronic devices continue to increase. These devices typically involve multi-domain coupling processes, including optical signal transmission, electrical drive, thermal regulation, and control feedback. To ensure device reliability, existing technologies usually collect operational data such as optical power, current, voltage, temperature, bit error rate status, and control logs, and employ threshold judgment, feature classification, or fault identification models trained on historical samples for status monitoring and fault diagnosis. Some solutions further introduce digital twins or multi-source data fusion to enhance anomaly detection capabilities.

[0003] However, different data sources exhibit significant differences in sampling frequency, response latency, and data quality. Synchronization based solely on timestamps can easily disrupt the sequential relationship of fault responses. Furthermore, device operating mode switching and actual faults may exhibit similar fluctuations, leading to false alarms. Existing fusion methods typically lack constraints on the physical transmission relationships between device functional units and the direction of fault propagation, making it difficult to distinguish between the source of the anomaly and the downstream affected units. For faults with similar symptoms, scarce samples, or unknown types, existing models are prone to forced classification and lack mechanisms to verify the root cause of the fault through controlled probing actions, resulting in insufficient reliability of fault location and inadequate interpretability of results. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method and system for fault detection of optoelectronic devices driven by multi-source heterogeneous data fusion, in order to solve the problems of multi-source data timing misalignment, lack of fault propagation relationship, and insufficient reliability of fault root cause localization in the prior art.

[0005] A first aspect of this application provides a fault detection method for optoelectronic devices driven by multi-source heterogeneous data fusion, comprising: acquiring multi-source operating data and device structure information of a target optoelectronic device; constructing a multi-domain physical topology based on the physical transmission relationship between functional units and generating data quality characterizations of each data source; performing causal timing correction on the multi-source operating data according to the response propagation relationship and control events in the multi-domain physical topology to generate a causal-aligned operating sequence; comparing the causal-aligned operating sequence with the health baseline of the target optoelectronic device and the prediction results of a multi-physics digital twin model to generate multi-layer fault residuals for each functional unit; mapping the multi-layer fault residuals to the multi-domain physical topology, and performing fault evidence fusion by combining the data quality characterization and response propagation direction to generate a fault evidence map characterizing the anomaly starting point, propagation path, and credibility; matching the fault evidence map with the fault prototype, and determining a fault candidate set based on the matching results, unexplainable residuals, and prediction uncertainty; when there is root cause ambiguity in the fault candidate set, using a multi-physics digital twin model to perform counterfactual response simulation on candidate detection actions that meet safety constraints, selecting detection actions, updating the fault candidate set according to the actual detection response, and outputting fault detection results.

[0006] A second aspect of this application provides a fault detection system for optoelectronic devices driven by multi-source heterogeneous data fusion, comprising: a construction module for acquiring multi-source operating data and device structure information of a target optoelectronic device, constructing a multi-domain physical topology based on the physical transfer relationships between functional units, and generating data quality characterizations for each data source; a correction module for performing causal timing correction on the multi-source operating data according to the response propagation relationships and control events in the multi-domain physical topology, and generating a causal-aligned operating sequence; and a comparison module for comparing the causal-aligned operating sequence with the health baseline of the target optoelectronic device and the prediction results of a multi-physics digital twin model, and generating data for each functional unit. The system consists of: a multi-layer fault residual; a fusion module, which maps the multi-layer fault residual to a multi-domain physical topology, combines data quality characterization and response propagation direction to perform fault evidence fusion, and generates a fault evidence map representing the anomaly origin, propagation path, and credibility; a matching module, which matches the fault evidence map with the fault prototype, and determines the fault candidate set based on the matching results, unexplainable residuals, and prediction uncertainty; and an output module, which, when there is root cause ambiguity in the fault candidate set, uses a multi-physics digital twin model to perform counterfactual response simulation on candidate detection actions that meet safety constraints, selects detection actions, updates the fault candidate set based on the actual detection response, and outputs the fault detection results.

[0007] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: By acquiring multi-source operational data and device structure information of the target optoelectronic device, a multi-domain physical topology is constructed based on the physical transmission relationship between functional units, and data quality characterizations of each data source are generated. Causal timing correction is performed on the multi-source operational data according to the response propagation relationship and control events in the multi-domain physical topology, generating a causal-aligned operational sequence. The causal-aligned operational sequence is compared with the health baseline of the target optoelectronic device and the prediction results of the multi-physics digital twin model to generate multi-layer fault residuals for each functional unit. The multi-layer fault residuals are mapped to the multi-domain physical topology, and fault evidence fusion is performed by combining data quality characterizations and response propagation directions to generate a fault evidence map representing the anomaly origin, propagation path, and credibility. The fault evidence map is matched with the fault prototype, and a fault candidate set is determined based on the matching results, unexplainable residuals, and prediction uncertainty. When there is root cause ambiguity in the fault candidate set, the multi-physics digital twin model is used to perform counterfactual response simulations on candidate detection actions that meet safety constraints, select detection actions, update the fault candidate set according to the actual detection response, and output the fault detection results. This application can improve the accuracy of multi-source data alignment, enhance fault propagation identification capabilities, and improve the credibility of fault root cause localization. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a flowchart illustrating the fault detection method for optoelectronic devices driven by multi-source heterogeneous data fusion provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the optoelectronic device fault detection system driven by multi-source heterogeneous data fusion provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0010] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0011] In existing technologies, fault detection of optoelectronic devices typically involves collecting operational data such as optical power, current, voltage, temperature, bit error rate, and control logs, and then using threshold judgment, feature classification, or fault identification models trained on historical samples for anomaly detection. Some solutions also introduce multi-source data fusion or digital twin models to improve device status identification capabilities by integrating operational characteristics from different monitoring dimensions.

[0012] However, different data sources vary in sampling frequency, timestamps, response delays, and data quality, making synchronization based solely on acquisition time prone to causal response misalignment. Furthermore, existing methods typically lack constraints on the physical transmission relationships between functional units within a device and the direction of response propagation, making it difficult to distinguish between the root cause unit of the fault and downstream units affected by fault propagation. For faults with similar symptoms, insufficient fault samples, or unknown types, existing methods are also prone to misclassification or root cause ambiguity, leading to insufficient reliability in fault location.

[0013] To address the aforementioned issues, this application acquires multi-source operational data and device structure information of the target optoelectronic device, constructs a multi-domain physical topology based on the physical transfer relationships between functional units, and generates data quality characterizations for each data source; performs causal timing correction on the multi-source operational data according to the response propagation relationships and control events in the multi-domain physical topology, generating a causal aligned operational sequence; further, it combines a health baseline and a multi-physical digital twin model to generate multi-layer fault residuals for each functional unit, maps the multi-layer fault residuals to the multi-domain physical topology, and fuses fault evidence through data quality and response propagation direction constraints to form a fault evidence map.

[0014] Based on this, this application matches the fault evidence map with the fault prototype and generates a fault candidate set by integrating unexplainable residuals and prediction uncertainties. When there is root cause ambiguity among multiple fault candidates, a multi-physics digital twin model is used to simulate the counterfactual response of candidate detection actions that meet safety constraints. The detection action is selected based on the separability of responses among candidate faults, and the fault candidate set is updated using actual detection responses. Thus, this application can improve the accuracy of causal alignment of multi-source data, enhance the ability to identify fault propagation paths and abnormal starting points, reduce the probability of fault misclassification, and improve the credibility and interpretability of root cause localization of optoelectronic device faults.

[0015] The technical solution of this application will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0016] Figure 1 This is a flowchart illustrating the fault detection method for optoelectronic devices driven by multi-source heterogeneous data fusion provided in this application embodiment. Figure 1 As shown, the method may specifically include: S101: Obtain multi-source operating data and device structure information of the target optoelectronic device, construct a multi-domain physical topology based on the physical transfer relationship between functional units, and generate data quality characterization of each data source. S102, based on the response propagation relationship and control events in the multi-domain physical topology, perform causal timing correction on the multi-source operation data to generate a causal aligned operation sequence; S103 compares the causal alignment operation sequence with the health baseline of the target optoelectronic device and the prediction results of the multi-physics digital twin model to generate multi-layer fault residuals for each functional unit; S104 maps multi-layer fault residuals to multi-domain physical topology, combines data quality characterization and response propagation direction to perform fault evidence fusion, and generates a fault evidence map that represents the anomaly starting point, propagation path and credibility. S105, Match the fault evidence map with the fault prototype, and determine the fault candidate set based on the matching results, unexplained residuals and prediction uncertainty; S106 When there is root cause ambiguity in the fault candidate set, the multi-physical digital twin model is used to simulate the counterfactual response of the candidate detection actions that meet the safety constraints, select the detection action, update the fault candidate set according to the actual detection response, and output the fault detection result.

[0017] In some embodiments, the operation monitoring data, control status data and historical correlation data of the target optoelectronic device are collected according to the data source identifier, and time stamping and format uniformity are applied to each data source; Functional units are divided according to the device structure information, each functional unit is determined as a topology node, and topology edges with propagation direction and response constraints are established based on the signal transmission, energy coupling and control dependence between functional units to generate a multi-domain physical topology. Quality features representing temporal reliability, sampling completeness, signal validity, and sensing stability are extracted from various data sources. These quality features are then normalized and fused based on the anomaly level and availability of the corresponding data sources to generate a data quality representation.

[0018] Specifically, taking an optoelectronic device comprising a laser emitting unit, a modulation driving unit, an optical receiving unit, a temperature control unit, a power management unit, and a main control unit as an example, the data acquisition cycle is set to 100ms. The system configures data source identifiers for optical power monitoring, bias current monitoring, temperature monitoring, power supply monitoring, bit error statistics, control commands, and maintenance records, and establishes a data access table based on these source identifiers. Operational monitoring data includes transmitted optical power, received optical power, laser bias current, modulation current, junction temperature, case temperature, power supply voltage, and bit error count; control status data includes transmit enable, temperature control target value, drive gain, power adjustment command, and protection status; historical correlation data includes device batch, cumulative runtime, historical alarms, maintenance results, and calibration records. After receiving each data point, the system writes it into a unified data frame. The data frame must at least include the device identifier, source identifier, acquisition time, reporting time, parameter name, parameter value, unit, and valid status.

[0019] To address clock deviations from different acquisition devices, the system selects the main control unit clock as the reference clock. It calculates the offset of each acquisition channel relative to the reference clock based on periodic heartbeat messages and estimates clock drift using the offset changes over multiple consecutive heartbeat cycles. For example, with the temperature monitoring channel lagging by 38ms and the optical power monitoring channel leading by 12ms, the corresponding time markers are corrected by shifting them forward and backward, respectively. For data carrying only the reporting time, the system uses transmission delay statistics to infer the acquisition time and marks the inferred result as the estimated time. When the format is standardized, different units are converted to a preset standard unit, enumerated states are mapped to a unified status code, and missing, out-of-bounds, and duplicate values ​​are marked separately, forming multi-source operational data sorted by the reference time.

[0020] The system divides functional units according to the device structure diagram and interface connection relationships, and establishes each functional unit as a topology node. An optical signal transmission edge is established between the laser emitting unit and the optical receiving unit; a power supply edge is established between the power management unit and the modulation / driving unit and the laser emitting unit; a thermal regulation edge is established between the temperature control unit and the laser emitting unit; and a control dependency edge is established between the main control unit and each controlled unit. Each topology edge records the propagation direction, allowable delay range, gain range, normal fluctuation range, and associated parameters. For example, after the driving gain is increased, the modulation current should change within 20ms to 60ms, and the emitted optical power should respond within 40ms to 120ms; after the junction temperature rises, the allowable change directions of the laser bias current and emitted optical power are limited by the corresponding thermal coupling edge. By aggregating all topology nodes and topology edges, a multi-domain physical topology capable of characterizing the relationships between the optical, electrical, thermal, and control domains is generated.

[0021] Data quality characterization is calculated separately for each data source. Time reliability is determined based on the presence of hardware timestamps, clock correction residuals, and transmission delay fluctuations; sampling completeness is determined based on the ratio of actual samples to theoretical samples within the target time window; signal validity is determined based on out-of-bounds ratio, mutation ratio, noise level, and compliance with physical constraints; and sensor stability is determined based on zero-point drift, repeated measurement differences, and short-time variance. All quality characteristics are first standardized to the 0-1 range and then converted into unidirectional indicators where higher values ​​represent higher quality.

[0022] For example, if a temperature channel has an integrity score of 0.92, a time reliability score of 0.88, a signal validity score of 0.95, and a sensing stability score of 0.90, then it is weighted according to a weight of 0.25, 0.25, 0.30, and 0.20 respectively to obtain a quality score of 0.92. Channels with scores not lower than 0.85 are marked as high-reliability channels. When a channel experiences continuous packet loss or drift exceeding limits, the anomaly level is increased and the availability level is decreased, causing the quality score to decrease accordingly. The system ultimately associates the quality score, quality level, and anomaly cause with the data source identifier, providing a reliable basis for subsequent causal time-series correction, residual calculation, and fault evidence fusion.

[0023] Through the above implementation methods, a unified temporal and physical correlation basis can be established while preserving the differences between multi-source data, thereby reducing the interference of low-quality data and error propagation relationships on fault detection results.

[0024] In some embodiments, causal timing correction is performed on multi-source operational data based on response propagation relationships and control events in a multi-domain physical topology to generate a causal-aligned operational sequence, including: Extract causal anchor points that can trigger state changes of functional units from control events, and determine the candidate response delay intervals of each data source relative to the causal anchor points based on the response propagation relationship; Within the candidate response latency interval, calculate the response correlation and propagation order conformity between the corresponding state changes of each data source and the causal anchor point to determine the actual response latency of each data source. Correct the time position of the corresponding multi-source running data according to the actual response delay, and establish a cross-data source association window according to the same causal anchor point; Based on the runtime configuration, the cross-data source association window is segmented into states, and a causal aligned runtime sequence carrying a runtime mode identifier is generated.

[0025] Specifically, continuing with the aforementioned optoelectronic device as an example, the system uses the emit enable, drive gain adjustment, temperature control target value change, and protection reset events generated by the main control unit as candidate control events, and reads the event type, trigger time, affected object, and parameter change from the event log. Only when a control event can cause a state change in at least one functional unit based on the multi-domain physical topology is the corresponding trigger time determined as the causal anchor point. For example, at 10:15:20.500 milliseconds, the main control unit adjusts the drive gain from 0.72 to 0.80. This event can sequentially affect the modulation drive unit, laser emitter unit, and optical receiver unit; therefore, the 500-millisecond moment is determined as the causal anchor point. Based on the allowable delays recorded by the control-dependent edge, power supply edge, and optical signal transmission edge, candidate response delay intervals of 10ms to 60ms, 20ms to 90ms, 40ms to 140ms, and 70ms to 220ms are set for the modulation current, laser bias current, emitted optical power, and received optical power, respectively.

[0026] Within each candidate response delay interval, the system calculates the change characteristics of the data sequence relative to the causal anchor point. First, the baseline mean and fluctuation range of each parameter are determined using a pre-event stability window. Then, slope abrupt changes, mean shifts, and persistent boundary crossings are detected in the post-event sequence, and the response correlation between parameter changes and control parameter changes is calculated. The response correlation can be obtained by weighting the normalized cross-correlation peak value, consistency of change direction, and consistency of change amplitude. Simultaneously, the system determines the theoretical propagation order based on the multi-domain physical topology; for example, the modulation current should change before the transmitted optical power, and the transmitted optical power should change before the received optical power. The propagation order consistency is calculated based on the detected actual abrupt change times.

[0027] For example, if the modulation current rises continuously 32ms after the event, the transmitted optical power rises 86ms after the event, and the received optical power rises 145ms after the event, and all three are within their respective candidate intervals and their order is consistent with the topology, then 32ms, 86ms, and 145ms are determined as the actual response delays of the corresponding data sources. If a channel has multiple candidate mutation points, the mutation point with the highest combined value of response correlation and propagation order conformity is selected; if the combined value is lower than a preset threshold, the theoretical delay is retained and the alignment confidence of the channel is reduced.

[0028] After determining the actual response delay, the system uses the causal anchor point as a unified zero moment and shifts the response data from each data source forward according to the corresponding delay, aligning the state changes triggered by the same control event under a unified causal coordinate system. For the aforementioned example, the detection times of modulation current, transmitted optical power, and received optical power are subtracted by 32ms, 86ms, and 145ms, respectively, and the corrected data is written to the same event window. The first part of the event window preserves the baseline before triggering, while the second part covers the maximum allowable response delay and stable recovery time, for example, set to 300ms before the anchor point to 800ms after the anchor point. The system establishes cross-data source relationships according to device identifier, causal anchor point identifier, and functional unit identifier, and records the original time, correction time, actual response delay, and alignment reliability to avoid losing the original acquisition data after time correction.

[0029] After the cross-data source association window is formed, the system reads the operating configuration such as transmit power level, temperature control strategy, modulation rate, protection status, and load level, and segments the window into states. When the operating configuration remains unchanged within the window, the corresponding data is divided into a single operating state segment; when the configuration changes, it is split into multiple state segments with the time the configuration takes effect as the boundary.

[0030] For example, if a device first enters a low-power preheating mode and then enters a rated emission mode, it is assigned a preheating mode identifier and a rated emission mode identifier, respectively. For brief transitional data during configuration switching, a transitional state identifier is set to prevent normal switching fluctuations from being misjudged as faults. Finally, the system aggregates each state segment according to the correction time sequence, generating a causal-aligned operating sequence that includes causal anchor points, functional unit responses, actual response delays, alignment confidence, and operating mode identifiers. This provides a unified timing input for subsequent health baseline matching, multi-physics prediction, and fault residual calculation.

[0031] In some embodiments, the causal alignment run sequence is compared with the health baseline of the target optoelectronic device and the prediction results of the multi-physics digital twin model to generate multi-layer fault residuals for each functional unit, including: According to the operation mode identifier, the baseline state interval matching the current operation state is extracted from the health baseline, and the causal aligned operation sequence is input into the multi-physics digital twin model to generate the expected response and prediction uncertainty of each functional unit. The actual response of each functional unit is compared with the baseline state range and the expected response to generate the basic residuals that characterize the individual state drift and physical response deviation. Based on the response constraints between associated functional units in the multi-domain physical topology, the consistency deviation of the actual response in terms of transmission direction, response amplitude and response delay is calculated, and cross-unit associated residuals are generated. Extract the degree of change and persistence characteristics of the basic residuals and cross-unit associated residuals, and perform association aggregation according to functional units to generate multi-layer fault residuals.

[0032] Specifically, continuing with the aforementioned optoelectronic device as an example, the system first reads the operating mode identifier, ambient temperature, power supply conditions, modulation rate, and power level from the causal alignment operating sequence, and then retrieves a reference state range consistent with the current configuration from the health baseline library. The health baseline library can be established by the target device in historical operating windows with no alarms, no maintenance anomalies, and data quality meeting requirements. It stores the mean range, rate of change range, and normal response delay of each functional unit parameter according to preheating mode, rated emission mode, low power mode, and high load mode, respectively. Taking the rated emission mode as an example, the reference range for the emitted optical power of the laser emission unit is 9.6 mW to 10.4 mW, the reference range for the bias current is 41 mA to 45 mA, and the reference range for the junction temperature is 46°C to 50°C. If the current ambient temperature is higher than when the baseline was established, the reference range is slightly shifted according to the temperature compensation relationship to avoid including normal environmental differences in the fault residual.

[0033] The system incorporates the control inputs, environmental states, and initial states of each functional unit into a multi-physics digital twin model, aligning the causal sequence of operations. Internally, the model calculates the current response of the modulation drive unit, the optical power and junction temperature response of the laser emitter unit, and the received power and bit error rate changes of the optical receiver unit, sequentially according to the relationships of electric drive, optical output, heat conduction, and control feedback. For example, after adjusting the drive gain from 0.72 to 0.80, the model predicts that the modulation current should increase to 18.5 mA in approximately 30 milliseconds, the emitted optical power should increase to 10.1 mW in approximately 80 milliseconds, and the received optical power should stabilize at 9.3 mW in approximately 140 milliseconds. The system also generates upper and lower limit ranges for each prediction result based on the model parameter reliability, input data completeness, and the dispersion of repeated simulation results. If the predicted emitted optical power is 10.1 mW and the prediction uncertainty is ±0.25 mW, the corresponding expected response range is 9.85 mW to 10.35 mW.

[0034] For each functional unit, the system compares the actual response with the baseline state range and the expected response range. When the actual value deviates from the healthy baseline but is still close to the expected response, it mainly forms an individual state drift residual; when the actual value deviates from the model's expected response, it forms a physical response deviation. Taking the aforementioned laser emission unit as an example, the stable value of the actual emitted optical power is 9.1 milliwatts, which is 0.5 milliwatts lower than the lower limit of the healthy baseline and 0.75 milliwatts lower than the lower limit of the expected response. Therefore, baseline deviation residuals and model prediction residuals are generated respectively. The system also retains the direction of the deviation, recording the value below the reference value as a negative residual and the value above the reference value as a positive residual, and performs normalization processing according to the width of the reference range so that residuals of different dimensions can be compared uniformly.

[0035] Subsequently, the system verifies whether the actual responses between associated functional units conform to propagation constraints based on the multi-domain physical topology. If the modulation current increases while the transmitted optical power decreases, the signal transmission direction is determined to be inconsistent; if the modulation current increases by 12% but the transmitted optical power only increases by 1%, which is lower than the normal gain range recorded by the topology edge, a response amplitude residual is generated; if the transmitted optical power responds 190 milliseconds after the drive change, exceeding the allowable delay of 40 to 140 milliseconds, a response delay residual is generated. For cases where a decrease in power supply from the power management unit causes synchronization anomalies between the modulation drive unit and the laser emission unit, the system also compares the order of changes in multiple downstream nodes to avoid misjudging common upstream influences as multiple independent faults.

[0036] The system further extracts the peak value, average value, slope of change, cumulative deviation, and duration of each residual within a continuous event window. Residuals that recover rapidly after a short-term boundary violation are marked as transient residuals, while residuals that maintain a continuous increase in the same direction for multiple consecutive windows are marked as persistent degradation residuals. For example, if the negative residual of emitted optical power increases for five consecutive windows, with a cumulative duration of 12 minutes, while the positive residual of junction temperature and the residual of response delay increase simultaneously, then the three types of residuals are associated with the laser emitting unit, and the persistent degradation weight is increased. Finally, the system aggregates baseline deviation residuals, model prediction residuals, directional consistency residuals, amplitude residuals, and time delay residuals according to functional unit identifiers, and associates their respective degree of change, persistence characteristics, and prediction uncertainty to generate multi-layer fault residuals, providing structured input for subsequent topology mapping and fault evidence fusion.

[0037] In some embodiments, the causal alignment run sequence is input into a multi-physics digital twin model to generate the expected response and prediction uncertainty of each functional unit, including: The current boundary conditions of the target optoelectronic device are determined based on the operating mode identifier, and the state input of each functional unit is extracted from the causal aligned operating sequence. The state input is input into the multi-physics digital twin model according to the response propagation relationship in the multi-domain physical topology. Based on the physical transfer constraints between functional units, the state changes are predicted step by step to generate the predicted response of each functional unit. The unmodeled bias in the predicted response is corrected by using a data-driven compensation unit, and the correction result is constrained by physical consistency constraints to generate the desired response; Based on the reliability of model parameters, the completeness of state input, and the dispersion of multiple prediction results, the prediction interval corresponding to the expected response is determined, and the prediction uncertainty is generated.

[0038] Specifically, continuing with the aforementioned optoelectronic device as an example, the multi-physics digital twin model consists of an electric drive computing unit, an optical output computing unit, a heat conduction computing unit, and a control feedback computing unit. Each computing unit establishes input-output relationships through topological edges in the multi-domain physical topology. The system first reads the operating mode identifier carried by the causal-aligned operating sequence, such as the rated emission mode, and simultaneously acquires the ambient temperature, supply voltage, modulation rate, target optical power, heat dissipation status, and cumulative device operating time. This information is used to determine the current boundary conditions, where the ambient temperature limits the heat conduction boundary, the supply voltage limits the upper limit of the electric drive, the target optical power and modulation rate limit the control input, and the heat dissipation status limits the heat exchange coefficient. Subsequently, the system extracts the output voltage of the power management unit, the initial current of the modulation drive unit, the junction temperature and initial optical power of the laser emitting unit, and the initial received power of the optical receiving unit from the causal-aligned operating sequence, and organizes them into state inputs according to a unified causal time.

[0039] During model calculations, the electric drive calculation unit first calculates the modulation current change based on the supply voltage, drive gain, and modulation load. Then, the modulation current is used as input to the laser emitting unit, and the optical output calculation unit predicts the emitted optical power by combining the bias current, junction temperature, and device aging parameters. The heat conduction calculation unit calculates the junction temperature evolution based on power consumption, ambient temperature, heat dissipation coefficient, and heat capacity parameters, and feeds the junction temperature change back to the optical output calculation unit to correct the luminous efficiency. The emitted optical power continues to be input to the optical receiving unit along the optical signal transmission edge, and the received optical power and bit error rate changes are predicted by combining link attenuation and receiver sensitivity. Taking an adjustment of the drive gain from 0.72 to 0.80 as an example, the model predicts that the modulation current will rise to 18.6 mA after 32 ms, the emitted optical power will rise to 10.2 mW after 84 ms, the junction temperature will rise by 1.4 °C after 310 ms, and the received optical power will stabilize at 9.4 mW after 146 ms, thus forming the predicted response sequence for each functional unit.

[0040] Considering that packaging stress, individual device differences, and long-term aging are difficult to fully express through mechanistic equations, the system incorporates a data-driven compensation unit. This unit reads model predictions, actual responses, operating modes, and environmental conditions from recent healthy operating samples, learning systematic deviations under different operating conditions. For example, if the target device exhibits a historical deviation of 0.18mW over-predicted emitted optical power in a high-temperature environment, the compensation unit performs a negative correction on the current prediction. The correction amount cannot directly cover the mechanistic prediction but must be verified through physical consistency constraints, including that the emitted optical power must not exceed the upper limit of electrical power conversion, the direction of optical efficiency correction must not violate the thermal decay law with increasing temperature, and the downstream response time must not be earlier than the upstream response time. When the compensation result violates any of these constraints, the system reduces the correction amount until the propagation direction, energy boundary, and response delay requirements are met, yielding the final desired response.

[0041] Prediction uncertainty is determined by a combination of model parameter reliability, state input completeness, and repeated prediction dispersion. Model parameter reliability is calculated based on parameter calibration time, number of calibration samples, and historical validation errors; state input completeness is determined based on the missing proportion, the proportion of estimated values, and data quality score; repeated prediction dispersion is obtained by running the digital twin model multiple times within permissible limits, perturbing thermal resistance, photoelectric conversion efficiency, link attenuation, and initial state.

[0042] For example, if multiple predictions of the transmitted optical power are concentrated between 9.96mW and 10.34mW, the model parameter confidence level is 0.91, and the state input integrity is 0.94, then the system will use 10.15mW as the expected value and generate a prediction range of 9.90mW to 10.40mW. If key temperature inputs are missing or parameters have not been calibrated for a long time, the prediction range will be expanded accordingly. Finally, the expected response, prediction range, main sources of uncertainty, and response delay of each functional unit are written into the prediction results, providing a basis for basic residual calculation and fault confidence determination.

[0043] Through the above-mentioned hierarchical calculation and constrained compensation, it is possible to maintain the consistency between the model output and the optical, electrical and thermal coupling mechanism of the device, reflect the individual deviation of the target device, and distinguish between model uncertainty and real anomalies by prediction interval, so as to provide a stable reference for subsequent fault residual hierarchical calculation.

[0044] In some embodiments, multi-layer fault residuals are mapped to multi-domain physical topology, and fault evidence is fused by combining data quality characterization and response propagation direction to generate a fault evidence graph representing the anomaly initiation point, propagation path, and credibility, including: The multi-layer fault residuals are written into the corresponding topology nodes of the multi-domain physical topology according to the functional unit identifier, and the initial evidence weight of each topology node is determined according to the data quality characterization. Based on the response propagation direction, response constraints, and the correlation strength between topology nodes, the initial evidence weights are propagated in a directed manner to generate fault interpretation relationships between each topology node; When the fault evidence of the upstream topology node can explain the anomaly of the downstream topology node, the root cause weight of the downstream topology node is suppressed; when the anomaly does not conform to the response propagation direction, the credibility of the corresponding fault explanation relationship is reduced. Based on the root cause weights, propagation correlation results, and evidence consistency of each topology node, the anomaly starting point and corresponding propagation path are determined, and the corresponding credibility is written into the multi-domain physical topology to generate a fault evidence graph.

[0045] Specifically, continuing with the aforementioned optoelectronic device as an example, the system writes the baseline deviation residual, model prediction residual, direction consistency residual, amplitude residual, and time delay residual into the topology nodes corresponding to the power management unit, modulation drive unit, laser emission unit, temperature control unit, and optical receiving unit, respectively, according to the functional unit identifier. Each node stores the residual type, residual direction, normalized amplitude, duration, and occurrence time.

[0046] For example, taking a single event of decreased transmitted optical power as an example, the voltage residual of the power management unit is 0.08, the current residual of the modulation drive unit is 0.16, the optical power residual of the laser emitting unit is 0.74, the junction temperature residual is 0.51, and the received power residual of the optical receiving unit is 0.62. The system determines the initial evidence weight based on the data quality score of the data source corresponding to each residual, and performs a weighted fusion of the residual amplitude, persistence, and quality score. If the quality score of the laser emitting unit's optical power channel is 0.94 and the quality score of its temperature channel is 0.90, then the corresponding anomalous evidence receives a higher initial weight; if there is packet loss in the received power channel and the quality score drops to 0.63, then the evidence weight of the corresponding node is reduced to avoid low-quality data dominating the root cause judgment.

[0047] Furthermore, after the initial evidence is written, the system performs fault evidence propagation along the directed edges of the multi-domain physical topology. Each topological edge pre-records its propagation direction, normal response delay, gain range, and correlation strength. The propagation amount is jointly determined by the evidence weight of the upstream node, the correlation strength of the topological edge, and the actual response compliance. For example, the correlation strength from the modulation driving unit to the laser emitting unit is 0.86, with a normal response delay of 20ms to 100ms; the correlation strength from the laser emitting unit to the optical receiving unit is 0.91, with a normal response delay of 30ms to 120ms.

[0048] When the transmitted optical power drops 65ms after the modulation current anomaly, and the received optical power drops 78ms thereafter, the actual propagation sequence and delay both satisfy the constraints. The system establishes a fault interpretation relationship between the modulation drive unit and the laser transmitting unit, and between the laser transmitting unit and the optical receiving unit, and calculates the downstream evidence contribution according to the propagation attenuation coefficient. During propagation, the system only allows evidence to propagate along the topological edge direction, and truncates the corresponding propagation when it exceeds the allowable delay range or the normal gain limit, preventing abnormal evidence from spreading across physically unrelated functional units.

[0049] For downstream anomalies that can be fully explained by upstream anomalies, the system reduces the weight of downstream nodes as independent root causes. Taking the aforementioned event as an example, if the amplitude of the laser emitting unit anomaly and the decrease in modulation current are within the topology edge gain range, and the optical receiving unit anomaly can be explained by the decrease in transmitted optical power, then the independent root cause weights of the laser emitting unit and the optical receiving unit are suppressed respectively. Conversely, if the received optical power decreases before the transmitted optical power, or if the downstream anomaly direction is opposite to the direction defined by the topology edge, then a conflict in the propagation relationship is determined, the credibility of the corresponding fault explanation relationship is reduced, and the possibility of downstream nodes as independent anomaly starting points is retained. When multiple upstream nodes can explain the same anomaly simultaneously, the system compares the explanation coverage, temporal consistency, and evidence quality to avoid repeatedly accumulating evidence of the same anomaly.

[0050] The system further summarizes the original anomaly evidence, upstream explanatory contributions, downstream supporting evidence, and conflicting evidence from each node, and calculates the root cause weight. If the root cause weight of the modulation driving unit is 0.81, the root cause weight of the laser emitting unit after suppression is 0.46, and the root cause weight of the optical receiving unit is 0.22, then the modulation driving unit is identified as the anomaly starting point, and a propagation path is formed along the topological edges that conform to the propagation constraints, connecting the modulation driving unit, the laser emitting unit, and the optical receiving unit. The path confidence is obtained by aggregating the data quality of each node, edge temporal compliance, amplitude compliance, and consistency of multi-source evidence, for example, calculated to be 0.84.

[0051] Finally, the system writes the anomaly origin, propagation order, affected nodes, root cause weight, path credibility, and conflict markers into the multi-domain physical topology to generate a fault evidence graph. The fault evidence graph retains the original residuals of each node and records the explanatory relationships between anomalies, providing a structured basis for subsequent fault prototype matching and root cause ambiguity resolution.

[0052] In some embodiments, a fault evidence map is matched with a fault prototype, and a fault candidate set is determined based on the matching results, unexplained residuals, and prediction uncertainty, including: Extract the anomaly origin, propagation path, and node evidence distribution from the fault evidence graph to generate the current fault representation; Obtain fault prototypes that match the device type and operating mode of the target optoelectronic device, and calculate the matching degree between the current fault characterization and each fault prototype in terms of abnormal location, propagation relationship and evidence change trend. Based on the interpretation results of each fault prototype on the multi-layer fault residual, the unexplainable residual is determined, and the comprehensive credibility of the corresponding fault prototype is generated by combining the prediction uncertainty. Fault prototypes that meet the candidate criteria in terms of overall credibility are selected, and the candidate priority is determined according to the overall credibility. When the overall credibility of all fault prototypes does not meet the candidate criteria, unknown fault candidates are generated based on the unexplained residuals and prediction uncertainties, and a fault candidate set is formed.

[0053] Specifically, continuing with the aforementioned optoelectronic device as an example, the system first reads the root cause weight, abnormal residual type, propagation direction, occurrence time, and path reliability of each topological node from the fault evidence graph, and forms the current fault characterization according to a unified feature order. Assuming that in a single detection, the root cause weight of the modulation driving unit is 0.78, the root cause weight of the laser emitting unit is 0.49, and the root cause weight of the optical receiving unit is 0.21, and the anomaly propagates sequentially along the modulation driving unit, laser emitting unit, and optical receiving unit, and the decrease in modulation current, decrease in transmitted optical power, and decrease in received optical power are the main node evidences, then the system encodes the anomaly starting position, path node sequence, evidence strength of each node, residual change slope, and duration as the current fault characterization.

[0054] The fault prototype library is managed hierarchically according to device type, packaging structure, operating band, and operating mode. Each fault prototype can be built from historical maintenance confirmation samples, accelerated aging test results, and digital twin fault injection results. Taking the rated transmission mode as an example, the prototype library includes fault prototypes such as abnormal drive gain, laser aging, temperature control failure, power supply fluctuation, and receiver link attenuation. Each fault prototype records the typical anomaly start point, allowed propagation path, distribution of node evidence, residual direction, trend of change, and normal difference range. The system filters out mismatched prototypes based on the target device model and rated transmission mode, retaining only fault prototypes with applicable structural relationships and operating boundaries.

[0055] The system calculates the matching degree between the current fault characterization and each fault prototype. The anomaly location matching degree compares whether the current anomaly origin is consistent with the prototype root cause node; the propagation relationship matching degree compares the degree of conformity between the actual path and the prototype path in terms of node order, propagation direction, and time delay range; and the evidence trend matching degree compares the rising and falling direction, growth slope, and continuous change pattern of the residuals at each node. For example, the matching degrees between the current fault characterization and the driving gain anomaly prototype in terms of anomaly origin, propagation order, and evidence trend are 0.92, 0.88, and 0.83, respectively, and the corresponding matching degrees with the laser aging prototype are 0.61, 0.79, and 0.86, respectively. After aggregation according to preset weights, the system obtains the initial matching results for the two prototypes.

[0056] During the prototype interpretation and verification phase, the system utilizes the residual interpretation range corresponding to each fault prototype to cover the current multi-layer fault residuals, and identifies the uncovered or directionally conflicting parts as uninterpretable residuals. For example, a prototype with abnormal drive gain can explain the decrease in modulation current, transmit optical power, and receive optical power, but cannot fully explain the continuous rise in junction temperature; therefore, the junction temperature residual is included in the uninterpretable residuals. The system calculates the degree of uninterpretability based on the amplitude, duration, and number of nodes involved in the uninterpretable residuals, and corrects the prototype reliability by combining this with the prediction uncertainty of each functional unit. When the prediction interval is wide, residuals near the boundary receive a lower penalty; when the prediction interval is narrow and the residuals significantly exceed the boundary, the degree of uninterpretability is increased.

[0057] The system unifies and weights prototype matching degree, explainable residual ratio, unexplainable degree, and prediction uncertainty to generate a comprehensive credibility score. Taking the aforementioned example, the comprehensive credibility score for the prototype with abnormal drive gain is 0.81, for the laser aging prototype it is 0.68, and for the temperature control failure prototype it is 0.43. If the candidate threshold is set to 0.60, the first two prototypes are added to the fault candidate set, and their priority is determined according to the comprehensive credibility score. Simultaneously, supporting and unexplained evidence is associated with each candidate to provide a basis for subsequent root cause ambiguity determination.

[0058] When the overall confidence level of all fault prototypes is lower than the candidate threshold, the system does not force the current anomaly into an existing category. Instead, it generates unknown fault candidates based on the concentrated location of unexplained residuals in the topology, propagation continuity, and prediction uncertainty. If the unexplained residuals are concentrated in the laser emitting unit and continue to spread along the thermally coupled edge, and multiple high-quality data sources provide consistent evidence, then unknown fault candidates with the laser emitting unit as the suspected starting point are generated, and unknown confidence levels and temporary fault labels are assigned.

[0059] The final set of candidate faults includes candidate type, anomaly origin, propagation path, overall credibility, candidate priority, and unknown marker. It can simultaneously support the identification of known faults and the retention of novel faults, avoiding misclassification due to insufficient prototypes.

[0060] In some embodiments, a counterfactual response simulation is performed on candidate detection actions that meet safety constraints using a multi-physical digital twin model. A detection action is selected, and the fault candidate set is updated based on the actual detection response. The fault detection result is then output, including: Based on the abnormal functional units and response sensitivity relationships corresponding to each fault candidate in the fault candidate set, candidate detection actions are generated, and candidate detection actions that do not meet the safety constraints are screened out according to the operating boundary of the target optoelectronic device. While keeping the current operating state unchanged, the response changes of each fault candidate under different candidate detection actions are simulated using a multi-physics digital twin model to generate counterfactual response fingerprints. Based on the degree of separability between the counterfactual response fingerprints corresponding to different fault candidates, and in combination with the safety margin and operational impact of the candidate detection actions, the target detection action is selected. The system controls the target optoelectronic device to perform target detection actions, obtains the actual detection response, matches the actual detection response with each counterfactual response fingerprint, updates the comprehensive confidence of each fault candidate, and outputs the fault detection result containing the fault location, fault type and detection confidence based on the update result.

[0061] Specifically, continuing with the aforementioned optoelectronic device as an example, when the fault candidate set contains both abnormal drive gain and laser aging as candidates, and their combined confidence levels are 0.81 and 0.68 respectively, the system reads the suspected abnormal functional unit, sensitive input parameters, and observable response parameters corresponding to each candidate. Abnormal drive gain is mainly sensitive to the short-time response amplitude of modulation current and emitted optical power, while laser aging is mainly sensitive to the conversion efficiency between bias current and emitted optical power, and junction temperature changes. Based on this, the system generates candidate detection actions such as fine-tuning of drive gain, fine-tuning of bias current, short-time adjustment of temperature control target value, and fine-tuning of emitted power setpoint, and configures the adjustment direction, change amplitude, hold time, and recovery conditions for each action.

[0062] The system performs safety screening based on the device's current operating mode, rated current, allowable junction temperature, optical power fluctuation range, and service continuity requirements. For example, if the current bias current is 43mA and the safety limit is 48mA, increasing the bias current by 1mA and 2mA will be retained as executable actions, while increasing it by 6mA will be screened out. If the current junction temperature is 49℃ and the safety limit is 55℃, actions that increase the temperature control target value will be excluded due to insufficient thermal margin. Actions that may significantly increase the bit error rate, exceed optical power limits, or trigger protection mechanisms will also be screened out. Screened candidate detection actions must also meet the following requirements: a single detection duration of no more than 2 seconds, an interval of at least 5 seconds between adjacent actions, and the ability to restore the original operating configuration after the action is completed.

[0063] During the counterfactual simulation phase, the system maintains the current ambient temperature, supply voltage, modulation rate, load state, and initial device state unchanged, injecting only the candidate fault parameters and candidate detection actions. Taking a 2% increase in drive gain as an example, under the candidate of abnormal drive gain, the model predicts that the modulation current will only increase by 0.3mA, the emitted optical power will increase by 0.08mW, and the response slope will be low; under the candidate of laser aging, the modulation current will increase by 0.9mA, but the emitted optical power will only increase by 0.05mW, and the junction temperature will rise by 0.4℃. The system combines the response start time, peak change, stable change, response slope, recovery time, and cross-cell transfer ratio under each candidate to form the corresponding counterfactual response fingerprint.

[0064] The system further compares the response fingerprint differences of different fault candidates under the same detection action to calculate the degree of separability. The degree of separability can be determined by weighting the normalized distance, response direction difference, and timing difference between key response parameters. If the difference in emitted optical power change between the two candidates is only 0.03mW when the bias current increases by 1mA, the degree of separability is low; if the difference in modulation current increase, emitted optical power increase, and junction temperature change between the two candidates is significant when the drive gain increases by 2%, the degree of separability is high. The system also calculates the safety margin based on the remaining proportion of the action distance from the safety boundary and calculates the operational impact based on the action duration, expected bit error rate change, and recovery time. After comprehensive comparison, the action of increasing the drive gain by 2% with a separability of 0.84, a safety margin of 0.78, and low operational impact is selected as the target detection action.

[0065] During execution, the main control unit adjusts the drive gain from 0.80 to 0.816 within a specified time, holds it for 1.5 seconds, and then restores it to its original value. The system synchronously collects modulation current, transmitted optical power, received optical power, junction temperature, and bit error count, aligning them according to causal propagation delay to form the actual detection response. The detection results show that the modulation current increased by 0.32mA, the transmitted optical power increased by 0.09mW, the junction temperature change was less than 0.1℃, the counterfactual response fingerprint distance between the actual response and the candidate for drive gain anomaly was 0.12, and the fingerprint distance between the actual response and the candidate for laser aging was 0.47. The system updates the candidate confidence based on the distance, increasing the overall confidence of the drive gain anomaly from 0.81 to 0.93 and decreasing the overall confidence of the laser aging from 0.68 to 0.31.

[0066] When the updated highest overall confidence level reaches the confirmation threshold and the difference between its confidence level and that of the second-highest candidate meets the distinction criteria, the system determines the fault location as the modulation drive unit, the fault type as abnormal drive gain, and the detection confidence level as 0.93. It then correlates the actual detection response, matching distance, supporting evidence, and propagation path to generate a fault detection result. If the matching degree between all actual responses and counterfactual response fingerprints is low, unknown fault candidates are retained, existing fault types are not forcibly output, and the detection data is used as feedback samples for subsequent fault prototype expansion and digital twin model correction.

[0067] In some embodiments, after outputting the fault detection result, the method further includes: Acquire subsequent operational data of the target optoelectronic device and maintenance verification data corresponding to the fault detection results, and generate detection feedback samples; Based on the data quality characterization, multi-layer fault residuals, and changes in fault evidence of subsequent operational data, operational windows that simultaneously meet the data credibility conditions and health judgment conditions are selected, and the health baseline is updated using the selection results. The actual detection response and fault evidence map before and after maintenance are compared, and the maintenance verification results are determined based on the degree of fading of the anomaly origin and propagation path. When the maintenance verification results indicate that the fault has been eliminated, the evidence distribution and response characteristics of the corresponding fault prototype are corrected using the detection feedback samples; when the maintenance verification results indicate that the fault has not been eliminated, the corresponding fault candidates are retained and the unmodeled bias of the multi-physical digital twin model is corrected for subsequent fault detection.

[0068] Specifically, continuing with the aforementioned optoelectronic device as an example, after the system outputs a fault detection result indicating abnormal driving gain in the modulation drive unit with a detection confidence level of 0.93, the system continuously collects operational data of the device before, during, and after maintenance, and reads maintenance work orders, parameter calibration records, component replacement records, and maintenance personnel confirmation information. Subsequent operational data includes modulation current, transmitted optical power, received optical power, junction temperature, supply voltage, bit error count, and control commands. Maintenance verification data includes the restored values ​​of drive parameters, deviations before and after calibration, maintenance completion time, and trial operation results.

[0069] The system establishes associations based on device identifiers, fault event identifiers, and maintenance task identifiers. It aggregates evidence maps before fault detection, actual detection responses, maintenance actions, and continuous operation sequences after maintenance into detection feedback samples, and records the source, time range, and quality level of each data item. For samples where maintenance tasks are not yet completed or trial run duration is insufficient, a "to be verified" flag is set, and they are not immediately used for baseline or model updates.

[0070] The system uses a 5-minute candidate operation window to re-evaluate the time reliability, sampling integrity, signal validity, and sensor stability of subsequent operation data after maintenance, and generates corresponding data quality characterizations. Data reliability is only considered met when the quality scores of key data sources within the window are all above 0.85, the missing rate is no higher than 3%, and there is no clock drift exceeding limits. Subsequently, the system calculates the multi-layer fault residuals of each functional unit and checks whether the residuals have fallen back to the allowable range of the healthy baseline. For example, if the modulation current model prediction residual is below 0.12, the transmit optical power baseline deviation residual is below 0.10, and the cross-unit delay residual is below 0.08, and no stable abnormal propagation path is formed for three consecutive windows, the healthy baseline is considered met. For windows that simultaneously meet two types of conditions, the mean value of parameters, fluctuation range, and response delay are extracted according to the operation mode identifier, and the healthy baseline is corrected using an incremental method with limited update amplitude. The single update amount does not exceed 5% of the original baseline interval width to avoid short-term fluctuations or potential residual faults after maintenance contaminating the healthy baseline.

[0071] During maintenance verification, the system performs the same target detection actions as before maintenance under the same operating mode and similar boundary conditions, such as increasing the drive gain by 2% and holding it for 1.5 seconds, to obtain the actual detection response after maintenance. The system compares the modulation current increment, transmit optical power increment, junction temperature change, and recovery time before and after maintenance, and also compares the root cause weight and propagation path confidence of the abnormal starting point in the fault evidence graph. If the root cause weight of the modulation drive unit was 0.81 and the propagation path confidence was 0.84 before maintenance, and decreased to 0.18 and 0.16 respectively after maintenance, and the abnormal evidence of the downstream laser emitting unit and optical receiving unit faded below the threshold, then the fault is determined to be eliminated. If the abnormal starting point is still located in the modulation drive unit, or the original propagation path only partially fades, then the fault is determined to be not completely eliminated, and the corresponding fault candidate and unfaded evidence are retained.

[0072] When the fault has been eliminated, the system writes the distribution of anomalous evidence before maintenance, propagation delay, residual change trend, and post-maintenance fading results into the driving gain anomaly prototype, correcting the evidence weights, allowable variation ranges, and response fingerprints of each node in the prototype, so that subsequent similar faults can obtain more accurate matching results. When the fault has not been eliminated, the system analyzes the persistent systematic differences between the actual response and the digital twin prediction. For example, if the model overestimates the recovery amplitude of emitted optical power under high temperature conditions by 0.15mW, this deviation is treated as an unmodeled deviation sample, correcting the data-driven compensation unit and the credibility of the corresponding parameters, and expanding the relevant prediction range.

[0073] Meanwhile, the system retains the original fault candidates and regenerates fault evidence maps based on new subsequent operational data for later review. Each update generates a version record containing the sample range, parameter changes, verification results, and effective time; if the prediction error on historical healthy samples increases after the update, the update is revoked and the previous valid version is restored. Through this closed-loop update, abnormal samples can be prevented from entering the healthy baseline while continuously correcting fault prototypes and digital twin models, improving the stability and adaptability of subsequent fault detection.

[0074] The following are system embodiments of this application, which can be used to execute the method embodiments of this application. For details not disclosed in the system embodiments of this application, please refer to the method embodiments of this application.

[0075] Figure 2 This is a schematic diagram of the structure of a multi-source heterogeneous data fusion-driven optoelectronic device fault detection system provided in an embodiment of this application. Figure 2 As shown, the system includes: Module 201 is used to acquire multi-source operating data and device structure information of the target optoelectronic device, construct a multi-domain physical topology based on the physical transfer relationship between functional units, and generate data quality characterization of each data source. The correction module 202 is used to perform causal timing correction on multi-source operating data based on the response propagation relationship and control events in the multi-domain physical topology, and generate a causal aligned operating sequence. Comparison module 203 is used to compare the causal aligned running sequence with the health baseline of the target optoelectronic device and the prediction results of the multi-physics digital twin model to generate multi-layer fault residuals for each functional unit; The fusion module 204 is used to map multi-layer fault residuals to multi-domain physical topology, combine data quality characterization and response propagation direction to perform fault evidence fusion, and generate a fault evidence map that represents the anomaly starting point, propagation path and credibility. The matching module 205 is used to match the fault evidence map with the fault prototype, and determine the fault candidate set based on the matching results, unexplained residuals and prediction uncertainty. The output module 206 is used to simulate the counterfactual response of candidate detection actions that meet safety constraints using a multi-physical digital twin model when there is root cause ambiguity in the fault candidate set, select detection actions, update the fault candidate set according to the actual detection response, and output the fault detection result.

[0076] In some embodiments, Figure 2The construction module 201 collects the operation monitoring data, control status data, and historical correlation data of the target optoelectronic device according to the data source identifier, and performs time stamping and format unification on each data source; it divides the functional units according to the device structure information, determines each functional unit as a topology node, and establishes topology edges with propagation direction and response constraints based on the signal transmission, energy coupling, and control dependency relationships between functional units to generate a multi-domain physical topology; it extracts quality features representing time reliability, sampling integrity, signal validity, and sensing stability from each data source, and normalizes and fuses the quality features according to the anomaly degree and availability of the corresponding data source to generate a data quality characterization.

[0077] In some embodiments, Figure 2 The correction module 202 extracts causal anchor points that can trigger state changes of functional units from control events, and determines the candidate response delay intervals of each data source relative to the causal anchor points based on the response propagation relationship; within the candidate response delay intervals, it calculates the response correlation degree and propagation order conformity degree between the corresponding state changes of each data source and the causal anchor points, and determines the actual response delay of each data source; it corrects the time position of the corresponding multi-source running data according to the actual response delay, and establishes a cross-data source association window according to the same causal anchor point; it segments the cross-data source association window according to the running configuration, and generates a causal aligned running sequence carrying the running mode identifier.

[0078] In some embodiments, Figure 2 The comparison module 203 extracts a baseline state interval matching the current operating state from the health baseline according to the operating mode identifier, and inputs the causal aligned operating sequence into the multi-physics digital twin model to generate the expected response and prediction uncertainty of each functional unit; it compares the actual response of each functional unit with the baseline state interval and the expected response respectively to generate the basic residuals characterizing individual state drift and physical response deviation; based on the response constraints between related functional units in the multi-domain physical topology, it calculates the consistency deviation of the actual response in the transmission direction, response amplitude and response delay to generate cross-unit related residuals; it extracts the degree of change and persistence characteristics of the basic residuals and cross-unit related residuals, and performs correlation aggregation according to functional units to generate multi-layer fault residuals.

[0079] In some embodiments, Figure 2The comparison module 203 determines the current boundary conditions of the target optoelectronic device based on the operating mode identifier and extracts the state input of each functional unit from the causal aligned operating sequence. The state input is input into the multi-physics digital twin model according to the response propagation relationship in the multi-domain physical topology. Based on the physical transfer constraints between functional units, the state changes are predicted step by step to generate the predicted response of each functional unit. The unmodeled bias in the predicted response is corrected by the data-driven compensation unit, and the correction result is limited by the physical consistency constraint to generate the expected response. Based on the model parameter confidence, state input completeness and the dispersion of multiple prediction results, the prediction interval corresponding to the expected response is determined, and the prediction uncertainty is generated.

[0080] In some embodiments, Figure 2 The fusion module 204 writes the multi-layer fault residuals into the corresponding topology nodes of the multi-domain physical topology according to the functional unit identifier, and determines the initial evidence weight of each topology node based on the data quality characterization; based on the response propagation direction, response constraints and the correlation strength between topology nodes, the initial evidence weights are propagated in a directed manner to generate fault interpretation relationships between topology nodes; when the fault evidence of the upstream topology node can explain the anomaly of the downstream topology node, the root cause weight of the downstream topology node is suppressed; when the anomaly does not conform to the response propagation direction, the credibility of the corresponding fault interpretation relationship is reduced; based on the root cause weight of each topology node, the propagation correlation result and the evidence consistency, the anomaly starting point and the corresponding propagation path are determined, and the corresponding credibility is written into the multi-domain physical topology to generate a fault evidence graph.

[0081] In some embodiments, Figure 2 The matching module 205 extracts the anomaly origin, propagation path, and node evidence distribution from the fault evidence map to generate the current fault characterization; it acquires fault prototypes that match the device type and operating mode of the target optoelectronic device, and calculates the matching degree between the current fault characterization and each fault prototype in terms of anomaly location, propagation relationship, and evidence change trend; it determines the unexplainable residuals based on the interpretation results of each fault prototype on the multi-layer fault residuals, and generates the comprehensive credibility of the corresponding fault prototypes in combination with the prediction uncertainty; it screens fault prototypes whose comprehensive credibility meets the candidate conditions, and determines the candidate priority according to the comprehensive credibility; when the comprehensive credibility of all fault prototypes does not meet the candidate conditions, it generates unknown fault candidates based on the unexplainable residuals and prediction uncertainty, and forms a fault candidate set.

[0082] In some embodiments, Figure 2The output module 206 generates candidate detection actions based on the abnormal functional units and response sensitivity relationships corresponding to each fault candidate in the fault candidate set, and filters out candidate detection actions that do not meet safety constraints based on the operating boundary of the target optoelectronic device; while keeping the current operating state unchanged, it uses a multi-physics digital twin model to simulate the response changes of each fault candidate under different candidate detection actions, and generates counterfactual response fingerprints; based on the separability between the counterfactual response fingerprints corresponding to different fault candidates, and combined with the safety margin and operational impact of the candidate detection actions, it selects the target detection action; it controls the target optoelectronic device to execute the target detection action, obtains the actual detection response, matches the actual detection response with each counterfactual response fingerprint, updates the comprehensive credibility of each fault candidate, and outputs the fault detection result containing the fault location, fault type and detection credibility based on the update result.

[0083] In some embodiments, Figure 2 The output module 206 acquires subsequent operating data of the target optoelectronic device and maintenance verification data corresponding to the fault detection results, and generates detection feedback samples. Based on the data quality characterization, multi-layer fault residuals, and changes in fault evidence of the subsequent operating data, it filters operating windows that simultaneously meet the data credibility conditions and health judgment conditions, and updates the health baseline using the filtering results. It compares the actual detection response and fault evidence map before and after maintenance, and determines the maintenance verification results based on the degree of fading of the anomaly starting point and propagation path. When the maintenance verification results indicate that the fault has been eliminated, it uses the detection feedback samples to correct the evidence distribution and response characteristics of the corresponding fault prototype. When the maintenance verification results indicate that the fault has not been eliminated, it retains the corresponding fault candidates and corrects the unmodeled bias of the multi-physics digital twin model for subsequent fault detection.

[0084] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0085] Figure 3 This is a schematic diagram of the electronic device 3 provided in an embodiment of this application. Figure 3 As shown, the electronic device 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, it implements the steps in the various method embodiments described above. Alternatively, when the processor 301 executes the computer program 303, it implements the functions of each module / unit in the various system embodiments described above.

[0086] Electronic device 3 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 3 may include, but is not limited to, processor 301 and memory 302. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or different components.

[0087] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0088] The memory 302 can be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. The memory 302 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 3. The memory 302 can also include both internal and external storage units of the electronic device 3. The memory 302 is used to store computer programs and other programs and data required by the electronic device.

[0089] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0090] If integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a readable storage medium (e.g., a computer-readable storage medium). Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which may be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0091] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A fault detection method for optoelectronic devices driven by multi-source heterogeneous data fusion, characterized in that, include: Acquire multi-source operational data and device structure information of the target optoelectronic device, construct a multi-domain physical topology based on the physical transfer relationship between functional units, and generate data quality characterizations for each data source; Based on the response propagation relationship and control events in the multi-domain physical topology, causal timing correction is performed on the multi-source operational data to generate a causal aligned operational sequence; The causal alignment running sequence is compared with the health baseline of the target optoelectronic device and the prediction results of the multi-physics digital twin model to generate multi-layer fault residuals for each functional unit. The multi-layer fault residuals are mapped to the multi-domain physical topology, and the fault evidence is fused by combining the data quality characterization and response propagation direction to generate a fault evidence map that represents the anomaly starting point, propagation path and credibility. The fault evidence map is matched with the fault prototype, and the fault candidate set is determined based on the matching results, unexplained residuals, and prediction uncertainty. When the fault candidate set has root cause ambiguity, the multi-physical digital twin model is used to perform counterfactual response simulation on the candidate detection actions that meet the safety constraints, select the detection action, update the fault candidate set according to the actual detection response, and output the fault detection result.

2. The method according to claim 1, characterized in that, The process involves acquiring multi-source operational data and device structure information of the target optoelectronic device, constructing a multi-domain physical topology based on the physical transfer relationships between functional units, and generating data quality characterizations for each data source, including: The operation monitoring data, control status data and historical correlation data of the target optoelectronic devices are collected according to the data source identifier, and time stamping and format are standardized for each data source. The device structure information is divided into functional units, each functional unit is determined as a topology node, and a topology edge with propagation direction and response constraint is established based on the signal transmission, energy coupling and control dependency relationship between functional units to generate the multi-domain physical topology. Quality features representing time reliability, sampling integrity, signal validity, and sensing stability are extracted from various data sources. These quality features are then normalized and fused based on the anomaly level and availability of the corresponding data sources to generate the data quality representation.

3. The method according to claim 1, characterized in that, The step of performing causal timing correction on the multi-source operational data based on the response propagation relationship and control events in the multi-domain physical topology to generate a causal aligned operational sequence includes: Extract causal anchor points that can trigger state changes of functional units from the control events, and determine the candidate response delay intervals of each data source relative to the causal anchor points based on the response propagation relationship; Within the candidate response delay interval, calculate the response correlation degree and propagation order conformity degree between the corresponding state change of each data source and the causal anchor point to determine the actual response delay of each data source. The time position of the corresponding multi-source running data is corrected according to the actual response delay, and a cross-data source association window is established according to the same causal anchor point; The cross-data source association window is segmented into states based on the running configuration, and the causal aligned running sequence carrying the running mode identifier is generated.

4. The method according to claim 1, characterized in that, The step of comparing the causal aligned running sequence with the health baseline of the target optoelectronic device and the prediction results of the multi-physics digital twin model to generate multi-layer fault residuals for each functional unit includes: According to the operating mode identifier, a baseline state interval matching the current operating state is extracted from the health baseline, and the causal aligned operating sequence is input into the multi-physics digital twin model to generate the expected response and prediction uncertainty of each functional unit. The actual response of each functional unit is compared with the baseline state range and the expected response to generate a basic residual characterizing the individual state drift and physical response deviation. Based on the response constraints between associated functional units in the multi-domain physical topology, calculate the consistency deviation of the actual response in terms of transmission direction, response amplitude and response delay, and generate cross-unit associated residuals. Extract the degree of change and persistence characteristics of the basic residuals and cross-unit associated residuals, and perform association aggregation according to functional units to generate the multi-layer fault residuals.

5. The method according to claim 4, characterized in that, The step of inputting the causal aligned running sequence into the multi-physics digital twin model to generate the expected response and prediction uncertainty of each functional unit includes: The current boundary conditions of the target optoelectronic device are determined based on the operating mode identifier, and the state input of each functional unit is extracted from the causal aligned operating sequence. The state input is input into the multi-physics digital twin model according to the response propagation relationship in the multi-domain physical topology. Based on the physical transfer constraints between functional units, the state changes are predicted step by step to generate the predicted response of each functional unit. The unmodeled bias in the predicted response is corrected using a data-driven compensation unit, and the correction result is constrained by physical consistency constraints to generate the desired response; Based on the reliability of model parameters, the completeness of state input, and the dispersion of multiple prediction results, the prediction interval corresponding to the expected response is determined, and the prediction uncertainty is generated.

6. The method according to claim 1, characterized in that, The process of mapping the multi-layer fault residuals to the multi-domain physical topology, combining the data quality representation and response propagation direction to perform fault evidence fusion, and generating a fault evidence graph representing the anomaly origin, propagation path, and credibility includes: The multi-layer fault residuals are written into the corresponding topology nodes of the multi-domain physical topology according to the functional unit identifier, and the initial evidence weight of each topology node is determined according to the data quality characterization. Based on the response propagation direction, response constraints, and the correlation strength between topology nodes, the initial evidence weights are propagated in a directed manner to generate fault interpretation relationships between each topology node; When the fault evidence of the upstream topology node can explain the anomaly of the downstream topology node, the root cause weight of the downstream topology node is suppressed; when the anomaly does not conform to the response propagation direction, the credibility of the corresponding fault explanation relationship is reduced. Based on the root cause weights, propagation correlation results, and evidence consistency of each topology node, the anomaly starting point and corresponding propagation path are determined, and the corresponding credibility is written into the multi-domain physical topology to generate the fault evidence graph.

7. The method according to claim 1, characterized in that, The step of matching the fault evidence map with the fault prototype and determining the fault candidate set based on the matching results, unexplained residuals, and prediction uncertainties includes: Extract the anomaly origin, propagation path, and node evidence distribution from the fault evidence graph to generate the current fault representation; Obtain fault prototypes that match the device type and operating mode of the target optoelectronic device, and calculate the matching degree between the current fault characterization and each fault prototype in terms of abnormal location, propagation relationship and evidence change trend. Based on the interpretation results of each fault prototype on the multi-layer fault residual, the unexplainable residual is determined, and the comprehensive credibility of the corresponding fault prototype is generated by combining the prediction uncertainty. Fault prototypes that meet the candidate criteria in terms of overall credibility are selected, and the candidate priority is determined according to the overall credibility. When the overall credibility of all fault prototypes does not meet the candidate criteria, unknown fault candidates are generated based on the unexplained residuals and prediction uncertainties, and the fault candidate set is formed.

8. The method according to claim 1, characterized in that, The process of using the multi-physics digital twin model to simulate counterfactual responses of candidate detection actions that meet safety constraints, selecting detection actions, updating the fault candidate set based on actual detection responses, and outputting fault detection results includes: Based on the abnormal functional units and response sensitivity relationships corresponding to each fault candidate in the fault candidate set, candidate detection actions are generated, and candidate detection actions that do not meet the safety constraints are screened out according to the operating boundary of the target optoelectronic device. While keeping the current operating state unchanged, the multi-physics digital twin model is used to simulate the response changes of each fault candidate under different candidate detection actions, and counterfactual response fingerprints are generated. Based on the degree of separability between the counterfactual response fingerprints corresponding to different fault candidates, and in combination with the safety margin and operational impact of the candidate detection actions, the target detection action is selected. The target optoelectronic device is controlled to perform the target detection action, the actual detection response is obtained, the actual detection response is matched with each counterfactual response fingerprint, the comprehensive confidence of each fault candidate is updated, and the fault detection result including fault location, fault type and detection confidence is output according to the update result.

9. The method according to claim 1, characterized in that, After outputting the fault detection results, the following is also included: Acquire subsequent operating data of the target optoelectronic device and maintenance verification data corresponding to the fault detection results, and generate detection feedback samples; Based on the data quality characterization, multi-layer fault residuals, and changes in fault evidence of the subsequent operating data, operating windows that simultaneously meet the data credibility conditions and health judgment conditions are selected, and the health baseline is updated using the selection results. The actual detection response and fault evidence map before and after maintenance are compared, and the maintenance verification results are determined based on the degree of fading of the anomaly origin and propagation path. When the maintenance verification result indicates that the fault has been eliminated, the evidence distribution and response characteristics of the corresponding fault prototype are corrected using the detection feedback sample; when the maintenance verification result indicates that the fault has not been eliminated, the corresponding fault candidate is retained and the unmodeled bias of the multi-physical digital twin model is corrected for subsequent fault detection.

10. A fault detection system for optoelectronic devices driven by multi-source heterogeneous data fusion, characterized in that, include: The module is used to acquire multi-source operating data and device structure information of the target optoelectronic device, construct a multi-domain physical topology based on the physical transfer relationship between functional units, and generate data quality characterizations of each data source. The correction module is used to perform causal timing correction on the multi-source running data based on the response propagation relationship and control events in the multi-domain physical topology, and generate a causal aligned running sequence. The comparison module is used to compare the causal aligned running sequence with the health baseline of the target optoelectronic device and the prediction results of the multi-physics digital twin model to generate multi-layer fault residuals for each functional unit. The fusion module is used to map the multi-layer fault residuals to the multi-domain physical topology, and combine the data quality characterization and response propagation direction to perform fault evidence fusion, generating a fault evidence map that characterizes the anomaly starting point, propagation path and credibility. The matching module is used to match the fault evidence map with the fault prototype, and determine the fault candidate set based on the matching results, unexplained residuals and prediction uncertainty. The output module is used to simulate the counterfactual response of the candidate detection actions that meet the safety constraints using the multi-physical digital twin model when there is root cause ambiguity in the candidate fault set, select the detection action, update the candidate fault set according to the actual detection response, and output the fault detection result.