An optical module adaptive parameter configuration method, medium and system

CN122528643APending Publication Date: 2026-08-07WUHAN JIUHONG PHOTOELECTRIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN JIUHONG PHOTOELECTRIC TECH CO LTD
Filing Date
2026-05-19
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]有鉴于此,本发明提供一种光模块自适应参数配置方法、介质及系统,能够解决现有技术中存在光模块在温度漂移、器件老化与制造工艺偏差三维耦合扰动下参数配置精度持续下降的技术问题

Benefits of technology

[0029]This invention solves the technical problem of continuously decreasing parameter configuration accuracy of optical modules under three-dimensional coupled disturbances of temperature drift, device aging, and manufacturing process deviations. It quantifies device state by constructing an aging-temperature coupled bias index, generates parameter recommendation vectors by using a configuration optimization model that integrates physical constraints with a weighted heterogeneous graph, stabilizes equalizer tap coefficients by using a variable step size minimum mean square-affine projection joint algorithm, and achieves online continuous learning by consolidating regularization combined with an aging physical prior model with elastic weights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122528643A_ABST
    Figure CN122528643A_ABST
Patent Text Reader

Abstract

The application provides a kind of optical module adaptive parameter configuration method, medium and system, belong to optical module adaptive parameter configuration technical field, the application is determined by real-time acquisition optical module physical parameter calculation aging-temperature coupling bias index to update intensity grade online learning, with weighted heterogeneous graph as input, utilize physical constraint embedded configuration optimization model to generate initial parameter configuration recommendation vector, solve graph Laplace potential field by optical module parameter disturbance propagation algorithm and obtain optimized parameter configuration candidate set by fusing random walk greedy search, adopt variable step size least mean square-affine projection joint algorithm to carry out stable convergence update to equalizer tap coefficient, finally, through edge-center two-level collaborative architecture, parameter configuration is issued, solve the technical problem that optical module is in temperature drift, device aging and manufacturing process deviation three-dimensional coupling disturbance under parameter configuration precision continues to decline.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of optical module adaptive parameter configuration technology, specifically, it relates to an optical module adaptive parameter configuration method, medium and system. Background Technology

[0002] Optical modules are core components of fiber optic communication systems, responsible for electro-optic and photoelectric conversion and signal modulation. In engineering practice, parameters such as the bias current, extinction ratio, and equalizer tap coefficients of optical modules need to be dynamically configured according to link status to maintain system bit error rate and optical signal-to-noise ratio. Existing technologies typically use rule-based static lookup table methods or simple regression models to recommend optical module parameters. Some solutions introduce convolutional neural networks or long short-term memory networks to model historical monitoring data to improve the automation of configuration recommendations. In digital coherent optical communication systems, equalizer parameter updates commonly use a fixed-step least mean square algorithm, and graph neural networks are increasingly being introduced into optical network topology modeling to capture the physical relationships between nodes.

[0003] However, the above methods have significant limitations in dealing with three-dimensional coupled disturbances such as temperature drift, device aging, and manufacturing process deviations: static lookup tables cannot track non-stationary drift, fixed step size minimum mean square algorithm has insufficient convergence stability under high-order modulation format, traditional neural networks suffer configuration accuracy degradation due to catastrophic forgetting in distributed drift environment, and existing graph neural network schemes lack explicit constraints on physical transmission laws, making it difficult to guarantee the physical rationality of parameter configuration recommendation vectors.

[0004] In existing technologies, due to the lack of a systematic approach that integrates aging physics priors, nonlinear transmission constraints, non-stationary sequence modeling, and online continuous learning, existing configuration optimization methods cannot simultaneously ensure compliance with physical constraints, non-stationary long-range memory, and catastrophic forgetting suppression when optical modules face three-dimensional coupled perturbations of temperature, aging, and manufacturing process deviations during long-term service. This leads to a continuous decline in parameter configuration accuracy. In other words, existing technologies suffer from the technical problem of continuously decreasing parameter configuration accuracy of optical modules under three-dimensional coupled perturbations of temperature drift, device aging, and manufacturing process deviations. Summary of the Invention

[0005] In view of this, the present invention provides an adaptive parameter configuration method, medium and system for optical modules, which can solve the technical problem in the prior art where the parameter configuration accuracy of optical modules continues to decline under the three-dimensional coupling disturbance of temperature drift, device aging and manufacturing process deviation.

[0006] The present invention is implemented as follows: The first aspect of the present invention provides a method for adaptive parameter configuration of an optical module, comprising the following steps:

[0007] The bias current, extinction ratio, optical signal-to-noise ratio, temperature and service life of the optical module are collected in real time. The bias current, extinction ratio, optical signal-to-noise ratio, temperature and service life are input into the aging-temperature coupled bias index calculation formula to obtain the aging-temperature coupled bias index. The online learning update intensity level is determined according to the interval to which the aging-temperature coupled bias index belongs.

[0008] The collected bias current, extinction ratio, optical signal-to-noise ratio, and temperature are used as graph node features, and the pre-collected fiber dispersion coefficient and link distance are used as graph edge features to construct a weighted heterogeneous graph. The weighted heterogeneous graph is then input into a physical constraint embedded configuration optimization model to obtain an initial parameter configuration recommendation vector.

[0009] Using the current configuration state corresponding to the collected bias current, extinction ratio, optical signal-to-noise ratio and temperature as the excitation source, the optical module parameter perturbation propagation algorithm is executed on the constructed weighted heterogeneous graph to solve the graph Laplace linear equation system to obtain the global parameter perturbation propagation potential field. Greedy search is performed along the gradient direction of the global parameter perturbation propagation potential field and random walk exploration is fused to obtain the optimized parameter configuration candidate set.

[0010] Using the collected bias current, extinction ratio, optical signal-to-noise ratio and temperature as inputs, the equalizer tap coefficients are updated using a variable step size minimum mean square-affine projection joint algorithm. The update step size is adaptively adjusted according to the moving average of error power, and tap coefficient regularization constraints are applied to obtain a stable and convergent tap coefficient configuration.

[0011] Based on the obtained online learning update intensity level, elastic weight consolidation regularization online update is performed on the physical constraint embedded configuration optimization model. The aging physical prior model is used to constrain the parameter update direction, thus completing the gradual adaptation of the physical constraint embedded configuration optimization model to the current distribution.

[0012] The optimized parameter configuration candidate set and the stable converged tap coefficient configuration are distributed to the edge-center two-level collaborative architecture. High-frequency small-amplitude disturbances with bias current changes less than 5% of the rated bias current and temperature changes less than 3℃ are processed by the optical module's local adaptive rule engine. The cross-module collaborative optimization task is routed to the central server, which calls the physical constraint embedded configuration optimization model and the optical module parameter disturbance propagation algorithm to solve the problem collaboratively, and completes the final parameter configuration distribution.

[0013] The aging-temperature coupling bias index is calculated using the formula... The weighting coefficients , , Data was obtained by performing multiple linear regression analysis on no fewer than 200 sets of device experimental data.

[0014] The interval division rule for the online learning update intensity level is as follows: when The time level is low; when The time level is medium; when The time level is high; the interval boundary value is determined by statistical analysis of the critical state where the configuration prediction error exceeds 5%.

[0015] The online learning update intensity level is determined by the online update intensity index. Further refinement, online updates of the intensity index The calculation uses the formula Weighting coefficients , Obtained through Pareto front analysis.

[0016] The core of the physical constraint embedded configuration optimization model is a 3-layer graph attention network. The attention weight calculation of each graph attention network layer incorporates a physical constraint regularization factor, which is formed by normalizing the product of the Kerr nonlinear coefficient and the optical power of the current node.

[0017] The physical constraint embedded configuration optimization model sets up a Mamba selective state space module after the 3-layer graph attention network layer. It dynamically generates the diagonal elements of the state transition matrix by linear projection of the 64-dimensional latent space vector, and establishes an input-dependent selective historical physical state retention mechanism.

[0018] The physical constraint embedded configuration optimization model sets up a mode skipping bypass mechanism. The node state monitoring module continuously calculates the Mahalanobis distance between the input feature vectors of nodes at adjacent time points. When the Mahalanobis distance exceeds the mutation detection threshold, the bypass channel is triggered to directly input the original node input feature vectors to the output configuration head, bypassing the 3-layer graph attention network layer.

[0019] The mutation detection threshold is determined by statistically analyzing the distribution of Mahalanobis distances when at least 100 measured mode-hopping events occur, and taking the 99th percentile as the threshold.

[0020] The optical module parameter perturbation propagation algorithm constructs a weighted undirected graph from the discrete parameter configuration state in the parameter space. The edge weight between two nodes is defined by the reciprocal of the Euclidean distance between the normalized difference of the fiber dispersion coefficient and the normalized difference of the Kerr nonlinear coefficient. The solution of the graph Laplace linear equations adopts the sparse conjugate gradient method.

[0021] The variable step size minimum mean square-affine projection joint algorithm uses the error vectors of the most recent frames to form an affine subspace during each update, projects the update direction of the equalizer tap coefficients to the affine subspace, and applies tap coefficient regularization constraints by superimposing a mean square norm penalty term onto the cost function.

[0022] The rule for adaptively adjusting the update step size based on the moving average of error power is as follows: when the moving average of error power increases, the update step size increases to accelerate convergence; when the moving average of error power decreases, the update step size decreases to fine-tune the update step size. The upper and lower limits of the update step size are determined by grid search.

[0023] The aging physical prior model is a physical equation describing the monotonic degradation of the bias current and extinction ratio of optical module devices with service time. The degradation parameters are obtained by curve fitting of time-series data collected from no less than 50 optical modules in accelerated aging experiments.

[0024] The elastic weight consolidation regularization estimates parameter importance through a diagonal approximation of the Fisher information matrix, and applies a quadratic penalty term proportional to the parameter importance to important parameters in historical tasks, so that the physical constraint embedded configuration optimization model does not deviate excessively from the parameter region that is effective for historical distributions when learning new distributions.

[0025] The training dataset for the physical constraint embedded configuration optimization model is constructed by deploying no less than 50 optical modules of different batches and service durations on a test bench, and no less than 100,000 sets of labeled samples are constructed. The optimal parameter configuration label vector is determined by an expert system in combination with exhaustive search.

[0026] The output configuration head of the physical constraint embedded configuration optimization model is subject to a physical knowledge soft constraint loss term. The physical knowledge soft constraint loss term penalizes the parameter configuration recommendation vector that violates the steady-state condition of the laser rate equation. The loss function is a weighted sum of the mean square error loss and the physical knowledge soft constraint loss term.

[0027] A second aspect of the present invention provides a computer-readable storage medium storing program instructions that, when executed in a computer, are used to perform the aforementioned optical module adaptive parameter configuration method.

[0028] A third aspect of the present invention provides an optical module adaptive parameter configuration system, comprising the aforementioned computer-readable storage medium, wherein the system is a computer, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.

[0029] This invention solves the technical problem of continuously decreasing parameter configuration accuracy of optical modules under three-dimensional coupled disturbances of temperature drift, device aging, and manufacturing process deviations. It quantifies device state by constructing an aging-temperature coupled bias index, generates parameter recommendation vectors by using a configuration optimization model that integrates physical constraints with a weighted heterogeneous graph, stabilizes equalizer tap coefficients by using a variable step size minimum mean square-affine projection joint algorithm, and achieves online continuous learning by consolidating regularization combined with an aging physical prior model with elastic weights.

[0030] This invention incorporates Kerr nonlinear constraint factors into the graph attention network layer of the physically constrained embedded configuration optimization model, enabling the aggregation process to automatically suppress the influence of physically unreasonable neighbors. The input-dependent dynamic selection mechanism of the Mamba selective state space module establishes long-range memory for non-stationary drift caused by aging, compensating for memory decay in traditional sequence models under non-stationary scenarios. The variable step size minimum mean square-affine projection joint algorithm uses multi-frame error information through affine subspace projection to suppress error propagation under high-order modulation formats, while the adaptive step size ensures parallel fast tracking and fine convergence. Elastic weight consolidation regularization protects historically important parameters with diagonal approximation of the Fisher information matrix, and combines the aging physical prior model to constrain the parameter update direction, enabling the model to achieve a gradual balance between old and new knowledge in a distributed drift environment, thereby systematically suppressing the destruction of configuration accuracy by three-dimensional coupling perturbations.

[0031] In summary, this invention solves the technical problem mentioned in the background art of the continuous decline in parameter configuration accuracy of optical modules under three-dimensional coupling disturbances of temperature drift, device aging, and manufacturing process deviations. Attached Figure Description

[0032] Figure 1 This is a flowchart of the method of the present invention.

[0033] Figure 2 A schematic diagram of the topology and edge weight distribution of a weighted heterogeneous graph for 32 optical module nodes.

[0034] Figure 3 The weight distribution diagram for selectively retaining the historical physical state of nodes with different service durations for the Mamba selective state space module. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below.

[0036] like Figure 1 The diagram shown is a flowchart of an adaptive parameter configuration method for an optical module provided by the first aspect of the present invention. This method includes the following steps:

[0037] S01. Real-time acquisition of the bias current, extinction ratio, optical signal-to-noise ratio, temperature and service life of the optical module. Input the bias current, extinction ratio, optical signal-to-noise ratio, temperature and service life into the aging-temperature coupled bias index calculation formula to obtain the aging-temperature coupled bias index. Determine the online learning update intensity level according to the interval to which the aging-temperature coupled bias index belongs.

[0038] S02. Using the bias current, extinction ratio, optical signal-to-noise ratio and temperature collected in S01 as graph node features, and the pre-collected fiber dispersion coefficient and link distance as graph edge features, a weighted heterogeneous graph is constructed. The weighted heterogeneous graph is then input into the physical constraint embedded configuration optimization model to obtain the initial parameter configuration recommendation vector.

[0039] S03. Using the current configuration state corresponding to the bias current, extinction ratio, optical signal-to-noise ratio and temperature collected in S01 as the excitation source, the optical module parameter perturbation propagation algorithm is executed on the weighted heterogeneous graph constructed in S02. The graph Laplace linear equation system is solved to obtain the global parameter perturbation propagation potential field. Greedy search is performed along the gradient direction of the global parameter perturbation propagation potential field and random walk exploration is fused to obtain the optimized parameter configuration candidate set.

[0040] S04. Using the bias current, extinction ratio, optical signal-to-noise ratio and temperature collected in S01 as inputs, the equalizer tap coefficients are updated using a variable step size minimum mean square-affine projection joint algorithm. The update step size is adaptively adjusted according to the moving average of error power, and tap coefficient regularization constraints are applied to obtain a stable and convergent tap coefficient configuration.

[0041] S05. Based on the online learning update intensity level obtained in S01, perform elastic weight consolidation regularization online update on the physical constraint embedded configuration optimization model, and use the aging physical prior model to constrain the parameter update direction to complete the gradual adaptation of the physical constraint embedded configuration optimization model to the current distribution.

[0042] S06. The optimized parameter configuration candidate set obtained in S03 and the stable converged tap coefficient configuration obtained in S04 are sent to the edge-center two-level collaborative architecture. High-frequency small-amplitude disturbances with bias current changes less than 5% of the rated bias current and temperature changes less than 3℃ are processed by the optical module's local adaptive rule engine. The cross-module collaborative optimization task is routed to the central server, which calls the physical constraint embedded configuration optimization model and the optical module parameter disturbance propagation algorithm to solve the problem collaboratively and complete the final parameter configuration distribution.

[0043] The formula for calculating the aging-temperature coupling bias index is as follows:

[0044] ;

[0045] in The aging-temperature coupling bias index has a dimension of 1. This represents the drift of the bias current relative to the rated bias current, expressed in mA. This is the rated bias current, in mA. This refers to the temperature deviation from the standard operating temperature, expressed in °C. Standard operating temperature, unit: °C Service duration, in hours (h) Device design lifetime, in hours (h) , , The weighting coefficients, all with a dimension of 1, were obtained through multiple linear regression analysis of experimental data from no fewer than 200 sets of devices within a temperature range of 0–70℃ and a service life range of 0–50,000 hours. Typical values ​​are approximately 0.4, 0.35, and 0.25, respectively. The interval division rule for the online learning update intensity level is as follows: when… At that time, the online learning update intensity level was low; when At that time, the online learning update intensity level was medium; when At that time, the online learning update intensity level was high; the above interval boundary values ​​were determined by statistical analysis of the critical state in the experimental data where the prediction error exceeded 5%.

[0046] The adjustment of the online learning update intensity level is based on the online update intensity index. Further refinement, online updates of the intensity index The calculation formula is expressed as follows:

[0047] ;

[0048] in The intensity index is updated online, with a dimension of 1. The aging-temperature coupling bias index has a dimension of 1. The mean square error between the recommended parameter configuration vector and the measured optimal parameter configuration vector, with units consistent with the dimensions of the recommended parameter configuration vector. To validate the baseline mean square error on the validation set, units are... same , These are weighting coefficients, all with a dimension of 1. They were obtained through Pareto front analysis of forgetting rate versus tracking speed in at least 50 online deployment iterations, with typical values ​​of approximately 0.5. When the elastic weight consolidation regularization coefficient is set to a high value, the learning rate is set to a low value; when At that time, the elastic weight consolidation regularization coefficient and the learning rate are both set to a medium value; when At that time, the elastic weight consolidation regularization coefficient is set to a low value, and the learning rate is set to a high value. The specific values ​​of the elastic weight consolidation regularization coefficient and the learning rate for each level are determined by conducting no less than 30 iterative experiments on different drift rate scenarios on a simulation platform, and the optimal combination of levels is determined by configuring a comprehensive score of prediction error convergence speed and catastrophic forgetting degree.

[0049] The specific structure of the physical constraint embedded configuration optimization model is as follows: The model takes a weighted heterogeneous graph as input, where each optical module corresponds to a node. The node input feature vector includes four normalized physical parameters: bias current, extinction ratio, optical signal-to-noise ratio, and temperature. The edge features include two normalized physical quantities: fiber dispersion coefficient and link distance. The first layer of the model is a node feature embedding layer, which uses a two-layer fully connected network to map the node input feature vector to a 64-dimensional latent space vector. The edge features are mapped to a 64-dimensional edge latent space vector via an independent fully connected network. The core of the model is a three-layer graph attention network layer, where the physical constraint regularization factor is incorporated into the attention weight calculation of each graph attention network layer. The physical constraint regularization factor is formed by normalizing the product of the Kerr nonlinear coefficient and the current node's optical power. The Kerr nonlinear coefficient is obtained by pre-calibration of the fiber material parameters, and the current node's optical power is obtained by real-time monitoring by the optical module. The physical constraint regularization factor is added to the attention score as a physical penalty term for the edge weights, so that the graph convolution aggregation process follows the physical law of nonlinear transmission. After the three-layer graph attention network, the temporal sequence of the 64-dimensional latent space vector of each node is sent to the Mamba selective state space module. The Mamba selective state space module contains an input-dependent selection matrix. By dynamically generating the diagonal elements of the state transition matrix through linear projection of the 64-dimensional latent space vector, it determines which node... Some historical physical states are preserved and passed on; the model internally sets up a mode-hopping bypass mechanism: the node state monitoring module continuously calculates the Mahalanobis distance of the node input feature vector at adjacent time points. When the Mahalanobis distance exceeds the mutation detection threshold, the bypass channel is triggered to directly input the original node input feature vector to the output configuration head, bypassing the 3-layer graph attention network layer; the mutation detection threshold is obtained by statistically analyzing the distribution of Mahalanobis distances when at least 100 measured mode-hopping events occur, and taking the 99th percentile as the threshold; the output configuration head is a two-layer fully connected network that outputs the parameter configuration recommendation vectors for each optical module and applies a physical knowledge soft constraint loss term. The physical knowledge soft constraint loss term is used to address violations of the laser rate equation stability. The parameter configuration recommendation vector under the state condition is penalized; the steps for establishing the training dataset of the physical constraint embedded configuration optimization model specifically include: deploying no less than 50 optical modules of different batches and different service durations on the test bench, covering the temperature range of 0 to 70℃ and no less than 5 batches with manufacturing process deviations, collecting the bias current, extinction ratio, optical signal-to-noise ratio, temperature and corresponding optimal parameter configuration annotation vector of each optical module under different temperatures and service durations, the optimal parameter configuration annotation vector is determined by the expert system combined with exhaustive search, constructing no less than 100,000 sets of labeled samples, and dividing the samples into layers according to temperature range, service duration range and process batch, ensuring that the proportion of each layer in the training set and the validation set is consistent;The specific steps for training the physically constrained embedded configuration optimization model include: supervised training using the Adam optimizer with an initial learning rate of 0.001, for a total of no less than 200 training epochs; the loss function is a weighted sum of the mean squared error loss and the physical knowledge soft constraint loss term, the weights of which are determined using a grid search on the validation set; after training, the configuration prediction error across temperature and service duration intervals is verified on an independent test set; if the configuration prediction error exceeds a preset threshold, the training data is expanded and retraining is performed; in the online deployment phase, incremental updates of the physically constrained embedded configuration optimization model are performed according to the online learning update intensity level, using elastic weight consolidation regularization; the elastic weight consolidation regularization coefficient is determined through forgetting rate experiments on historical datasets.

[0050] The physically constrained embedded configuration optimization model models the optical network topology as a weighted heterogeneous graph, enabling the model to naturally capture the physical relationships between optical modules. The attention weights of the graph attention network layer incorporate physical constraint regularization factors, allowing the aggregation process to automatically suppress the influence of physically unreasonable neighbors. The Mamba selective state space module establishes long-range memory for non-stationary drift caused by aging through an input-dependent dynamic selection mechanism, compensating for the memory decay deficiency of traditional sequence models in non-stationary scenarios. The mode-hopping bypass mechanism ensures that the parameter configuration recommendation vector under abnormal states is not affected by the graph smoothing effect. The physical knowledge soft constraint loss term constrains the physical rationality of the parameter configuration recommendation vector from the output end. The synergy of multiple mechanisms enables the physically constrained embedded configuration optimization model to maintain configuration accuracy under three-dimensional coupled perturbations of temperature, aging, and manufacturing process.

[0051] The specific implementation of the optical module parameter perturbation propagation algorithm is as follows: Discrete parameter configuration states in the optical module parameter space are constructed as a weighted undirected graph. Each node in the graph corresponds to a parameter configuration state. The edge weight between two nodes is defined as the reciprocal of the Euclidean distance between the normalized difference of the fiber dispersion coefficient and the normalized difference of the Kerr nonlinear coefficient. A larger edge weight indicates a higher physical similarity between the two parameter configuration states. The weighted undirected graph is equivalent to a resistor network. A unit voltage excitation is applied to the node with the current parameter configuration state, while the other nodes are grounded. The potential of each node is obtained by solving the graph Laplace linear equations. The potential of each node is the parameter perturbation propagation response of each parameter configuration state relative to the current parameter configuration state. The effective resistance value between nodes corresponds to the parameter configuration jump cost. The potential gradient direction is used as the greedy search direction, and a random walk with a probability proportional to the edge weight is performed several times before executing the greedy search step to obtain the optimized parameter configuration candidate set. The solution of the graph Laplace linear equations adopts the sparse conjugate gradient method, with a computational complexity of O(log n). Quantity, of which This represents the number of nodes in a weighted undirected graph.

[0052] The optical module parameter perturbation propagation algorithm transforms the parameter search problem into a potential solution problem for a resistor network. It utilizes the spectral properties of the graph Laplacian operator to globally capture the connectivity structure of the parameter space, overcoming the defect of local gradient search being prone to getting trapped in local optima. The fusion of random walk and greedy search steps enables the search process to dynamically balance between exploring new parameter configuration states and converging to a high-quality parameter configuration state. The effective resistance value measures the cost of parameter configuration jumps, making the search path tend to jump between parameter configuration states with high physical similarity, reducing the risk of link performance fluctuations caused by sudden parameter configuration changes.

[0053] The specific implementation of the variable step size minimum mean square-affine projection joint algorithm is as follows: the update step size of the equalizer tap coefficients is adaptively adjusted according to the moving average of the error power. When the moving average of the error power increases, the update step size increases to accelerate convergence; when the moving average of the error power decreases, the update step size decreases for fine-grained adjustment. The affine projection algorithm uses the error vectors of the most recent frames to construct an affine subspace during each update, and projects the update direction of the equalizer tap coefficients onto this affine subspace to reduce random walks caused by error propagation under higher-order pulse amplitude modulation formats. The tap coefficient regularization constraint is superimposed on the cost function with a mean square norm penalty term to suppress equalizer tap coefficient oscillations. The upper and lower limits of the update step size and the tap coefficient regularization constraint coefficients are determined by grid search under experiments with no fewer than 20 combinations of optical signal-to-noise ratio and fiber dispersion coefficient, using the oscillation amplitude and convergence speed of the mean square error convergence curve as evaluation indicators.

[0054] The variable step size minimum mean square-affine projection joint algorithm uses the affine subspace projection mechanism to construct a high-dimensional constraint surface using multi-frame error vector information. This allows the equalizer tap coefficient update direction to find the most consistent descent direction in the stochastic gradient noise caused by error propagation, fundamentally suppressing the error propagation effect of the decision feedback equalizer under high-order pulse amplitude modulation format. The adaptive adjustment of the update step size ensures fast tracking when the fiber dispersion changes rapidly and fine convergence when approaching the optimal equalizer tap coefficient, resolving the contradiction between convergence speed and stability in the fixed step size algorithm.

[0055] The aging physical prior model is a physical equation describing the monotonic degradation of bias current and extinction ratio of optical module devices during service. The degradation parameters are obtained by curve fitting of the time-series data of bias current and extinction ratio collected from no less than 50 optical modules in accelerated aging experiments. The aging physical prior model serves as an elastic weight to consolidate the parameter update direction constraint during regularized online updates, so that the parameter update of the physical constraint embedded configuration optimization model is biased in the direction that conforms to the aging physical degradation law.

[0056] The elastic weight consolidation regularization is a parameter update constraint method to prevent catastrophic forgetting in online continuous learning. Its principle is as follows: when updating the parameters of the physically constrained embedded configuration optimization model, a quadratic penalty term proportional to the importance of the parameters in the historical task is applied to the important parameters. The importance of the parameters is estimated by the diagonal approximation of the Fisher information matrix on the historical task data, so that the physically constrained embedded configuration optimization model does not deviate excessively from the parameter region that is effective for the historical distribution when learning a new distribution, thereby achieving a gradual balance between old and new knowledge in the distribution drift environment.

[0057] The Mamba selective state-space module is a sequence modeling structure based on a state-space model. Its core is that the diagonal elements of the state transition matrix are dynamically generated by linear projection of the current input signal, which enables the model to selectively retain the input-dependent historical physical state at different times. This is different from the fixed state transition mechanism in traditional recurrent neural networks and is suitable for modeling non-stationary long-range temporal dependencies caused by the aging of optical modules.

[0058] The Mahalanobis distance is a generalized distance metric that considers the variance and covariance of each dimension. When used for node state mutation detection, it can automatically eliminate the influence of the differences in the dimensions and variances of parameters such as bias current, extinction ratio, optical signal-to-noise ratio, and temperature compared to Euclidean distance, thus improving the robustness of mutation detection to different parameter scales.

[0059] The graph Laplace linear equations are a set of linear equations composed of the Laplace matrix of a weighted undirected graph. Each diagonal element of the Laplace matrix is ​​the sum of all edge weights of the corresponding node, and the off-diagonal elements are the negative values ​​of the corresponding edge weights. Solving the graph Laplace linear equations is equivalent to solving the potential distribution of each node in a resistive network. The calculation results reflect the potential response of each parameter configuration state node relative to the excitation source node, i.e., the global parameter perturbation propagation potential field.

[0060] The affine projection algorithm is an algorithm that uses the affine subspace formed by multiple frame input vectors and error vectors to constrain the direction of adaptive filter coefficient update. Its principle is: each time the equalizer tap coefficient update is constrained to find the minimum norm solution in the affine subspace spanned by the input vectors of the most recent several frames. Compared with the minimum mean square algorithm, it uses more historical error information to reduce the update variance and improves the convergence stability of equalizer tap coefficients in non-stationary channels caused by fiber dispersion.

[0061] Pareto front analysis is a method used in multi-objective optimization to identify trade-offs between different objectives, specifically in determining weighting coefficients. and Using forgetting rate and tracking speed as two optimization objectives, a Pareto optimal solution set is identified from the results of no less than 50 online deployment iterations. The solution with the best combined score for forgetting rate and tracking speed is then selected from this Pareto optimal solution set. and Values.

[0062] The specific implementation of step S01 is as follows: Five physical parameters—bias current, extinction ratio, optical signal-to-noise ratio, temperature, and service life—are collected in real time through the built-in monitoring circuit of the optical module. The sampling period can be set to seconds or minutes depending on the link status. These five parameters are then substituted into the aging-temperature coupled bias index calculation formula to calculate the drift and normalized value of the bias current relative to the rated bias current, the deviation and normalized value of the temperature relative to the standard operating temperature, and the ratio of service life to the design life. The three normalized values ​​are then multiplied by weighting coefficients. , , Summing yields the dimensionless aging-temperature coupled bias exponent. The weighting coefficients were obtained through multiple linear regression analysis of experimental data from no fewer than 200 devices within a temperature range of 0–70℃ and a service life range of 0–50,000 hours. Typical values ​​were approximately 0.4, 0.35, and 0.25, respectively. Based on... The range of values ​​for , when The online learning update intensity level was determined to be low. At the middle of the hour, when The interval boundary value is determined by statistical analysis of the critical state where the configuration prediction error exceeds 5%. The intensity index is updated online. exist Based on this, the mean square error normalization of the current parameter configuration recommended vector and the measured optimal parameter configuration vector is further incorporated, and the weighting coefficients are determined through Pareto front analysis. and The typical value is about 0.5, thus unifying the quantification of both the degree of device degradation and the current configuration error, providing a refined basis for updating intensity in subsequent online learning steps.

[0063] The specific implementation of step S02 is as follows: The four physical parameters collected in step S01—bias current, extinction ratio, optical signal-to-noise ratio, and temperature—are normalized and used as graph node feature vectors; the fiber dispersion coefficient, measured in advance or obtained from the link database, and the link distance are normalized and used as graph edge features. A weighted heterogeneous graph is constructed with optical modules as nodes and links as edges, where the edge weights between nodes comprehensively reflect the physical attributes of the link. The weighted heterogeneous graph is input into a physical constraint embedded configuration optimization model. The first layer of the model, the node feature embedding layer, uses a two-layer fully connected network to map the node input feature vectors to a 64-dimensional latent space vector, and the edge features are mapped to a 64-dimensional edge latent space vector via an independent fully connected network. When calculating the attention weights, the core three-layer graph attention network layer superimposes the physical constraint regularization factor—composed of the normalized product of the Kerr nonlinear coefficient and the current node's optical power—onto the attention score, ensuring that the aggregation process follows the physical laws of nonlinear transmission. After the output of the 3-layer graph attention network, the temporal sequence of the 64-dimensional latent space vector is fed into the Mamba selective state space module. A dynamic selection mechanism dependent on input is used to establish long-range memory for non-stationary drift caused by aging. The node state monitoring module continuously calculates the Mahalanobis distance between the input feature vectors of adjacent nodes. When this distance exceeds the mutation detection threshold (99th percentile), a mode-hopping bypass mechanism is triggered, allowing the original node input feature vectors to bypass the graph attention network layer and directly input into the output configuration head. This ensures that the configuration recommendation under mode-hopping events is not affected by graph smoothing effects. The output configuration head is a two-layer fully connected network that outputs the initial parameter configuration recommendation vectors for each optical module.

[0064] The specific implementation of step S03 is as follows: The discrete parameter configuration states in the optical module parameter space are constructed as a weighted undirected graph. Each node in the graph corresponds to a parameter configuration state. The edge weight between two nodes is defined as the reciprocal of the Euclidean distance formed by the normalized difference of the fiber dispersion coefficient and the normalized difference of the Kerr nonlinear coefficient. A larger edge weight indicates a higher physical similarity between the two parameter configuration states. The weighted undirected graph is then equivalent to a resistive network. A unit voltage excitation is applied to the node in the current configuration state, while the remaining nodes are grounded. A graph Laplace matrix is ​​constructed, where the diagonal elements are the sum of all edge weights for the corresponding node, and the off-diagonal elements are the negative values ​​of the corresponding edge weights. The sparse conjugate gradient method is used to solve the graph Laplace linear equations, with a computational complexity of O(log n). The potential of each node is obtained, which is the global parameter perturbation propagation potential field. The potential gradient direction is used as the greedy search direction, and each step moves to the neighboring node with the largest potential gradient. At the same time, a random walk is performed with a probability proportional to the edge weight for several steps before the greedy search step is performed. Through the dynamic balance between exploration and convergence, the optimized parameter configuration candidate set is finally obtained. The effective resistance value between nodes corresponds to the parameter configuration jump cost, so that the search path tends to jump between parameter configuration states with high physical similarity.

[0065] The specific implementation of step S04 is as follows: Using the bias current, extinction ratio, optical signal-to-noise ratio, and temperature collected in step S01 as inputs, the equalizer tap coefficients are updated using a variable step-size minimum mean square-affine projection joint algorithm. The moving average of the error power is calculated in real time. When the moving average increases, the update step size is adaptively increased to accelerate convergence; when the moving average decreases, the update step size is adaptively decreased for fine-grained adjustment. The upper and lower limits of the update step size are determined through grid search experiments under no fewer than 20 combinations of optical signal-to-noise ratio and fiber dispersion coefficient. In each update, the affine projection algorithm uses the input vectors and error vectors of the most recent frames to construct an affine subspace. The equalizer tap coefficient update amount is constrained within this affine subspace to find the minimum norm solution, utilizing multi-frame error information to reduce random walks caused by error propagation under higher-order pulse amplitude modulation formats. Simultaneously, a mean square norm penalty term is superimposed on the cost function to apply tap coefficient regularization constraints, suppressing tap coefficient oscillations, ultimately obtaining a stable and convergent tap coefficient configuration.

[0066] The specific implementation of step S05 is as follows: Based on the online learning update intensity level obtained in step S01, elastic weight consolidation regularization is performed on the physical constraint embedded configuration optimization model for online updates. Elastic weight consolidation regularization estimates the importance of model parameters to historical tasks through a diagonal approximation of the Fisher information matrix, and applies a quadratic penalty term proportional to the importance of important parameters to protect historical knowledge. The aging physics prior model describes the monotonic degradation law of bias current and extinction ratio with service time. The degradation law parameters are obtained by curve fitting of time-series data from no less than 50 optical modules in accelerated aging experiments. During online updates, the aging physics prior model is used as a constraint on the parameter update direction, causing the model parameter updates to be biased in a direction consistent with the aging physics degradation law. When the online update intensity index... When the elastic weight consolidation regularization coefficient is set to a high value and the learning rate to a low value, the elastic weight consolidation regularization coefficient is set to a high value; when When both are at the middle value, when At that time, the elastic weight consolidation regularization coefficient is set to a low value and the learning rate is set to a high value. The optimal combination of each level is determined through no less than 30 iteration experiments on the simulation platform to complete the gradual adaptation of the physical constraint embedded configuration optimization model to the current distribution.

[0067] The specific implementation of step S06 is as follows: The optimized parameter configuration candidate set obtained in step S03 and the stable convergent tap coefficient configuration obtained in step S04 are distributed to the edge-center two-level collaborative architecture. The local adaptive rule engine of the optical module in the edge layer determines in real time whether the change in bias current is less than 5% of the rated bias current and whether the change in temperature is less than 3°C. High-frequency small-amplitude disturbances that meet the conditions are processed directly locally to reduce communication overhead. Cross-module collaborative optimization tasks that do not meet the conditions are routed to the central server. The central server calls the physical constraint embedded configuration optimization model and the optical module parameter disturbance propagation algorithm to solve the problem collaboratively. By integrating the network topology information and the physical constraints of multiple modules, the final parameter configuration is generated and distributed to each optical module, completing the entire adaptive parameter configuration process.

[0068] It should be noted that the key technologies of this invention include: the physically constrained embedded configuration optimization model directly couples the attention weights of the graph attention network layer with the Kerr nonlinear physical constraint regularization factor, so that the graph convolution aggregation internalizes the nonlinear transmission physical law in the mathematical structure, rather than just a soft constraint at the loss function level, thus fundamentally preventing physically unreasonable parameter recommendation vectors from being amplified and propagated during the aggregation process; the Mamba selective state space module realizes adaptive modeling of non-stationary long-range temporal dependencies caused by aging through the dynamic generation mechanism of the input-dependent state transition matrix, overcoming the inherent defect of the fixed state transition mechanism of traditional recurrent neural networks in which memory decay leads to configuration accuracy degradation in non-stationary scenarios; the variable step size minimum mean square-affine projection joint algorithm unifies and fuses multi-frame error information through affine subspace constraints, eliminating the interference of random gradient noise on the tap coefficient update direction under high-order pulse amplitude modulation format from an algebraic level, while the adaptive step size achieves Pareto balance between convergence speed and stability. The synergistic effect of the three key technologies is reflected in the following aspects: the physical constraint embedded configuration optimization model provides a physically reasonable initial configuration recommendation; the optical module parameter perturbation propagation algorithm efficiently searches for optimization candidates in the physical similarity parameter space; the variable step size minimum mean square-affine projection joint algorithm finely converges the equalizer parameters; and the elastic weight consolidation regularization ensures that online learning does not destroy the accumulated physical knowledge. The four technologies form a complete closed loop of physical constraints, global search, fine equalization, and continuous learning, which systematically guarantees the configuration accuracy of the system under three-dimensional coupled perturbations.

[0069] It should be noted that this invention also solves the following technical problems: In the process of searching the parameter space of optical modules, traditional local gradient search methods are prone to getting trapped in local optima because they cannot perceive the global connectivity structure of the parameter space, resulting in low quality of the parameter configuration candidate set. This invention transforms the parameter search problem into a problem of solving the potential of a resistive network through an optical module parameter perturbation propagation algorithm. It uses the spectral properties of the graph Laplacian operator to globally capture the connectivity structure of the parameter space, and uses the node potential distribution as the global parameter perturbation propagation potential field to guide the search direction from a global perspective. The fusion of random walk and greedy search steps enables the search process to dynamically balance between exploring new parameter configuration regions and converging to a high-quality parameter configuration. The effective resistance value measures the cost of parameter configuration jumps, making the search path tend to jump between parameter configuration states with high physical similarity, reducing the risk of link performance fluctuations caused by sudden parameter configuration changes. The sparse conjugate gradient method controls the computational complexity of the solution to within a certain range. The scale ensures the engineering feasibility of the algorithm in a large-scale parameter space, thus solving the technical problem that traditional local gradient search is prone to getting trapped in local optima in the parameter space of optical modules.

[0070] A second aspect of the present invention provides a computer-readable storage medium storing program instructions that, when executed in a computer, are used to perform the aforementioned optical module adaptive parameter configuration method.

[0071] A third aspect of the present invention provides an optical module adaptive parameter configuration system, comprising the aforementioned computer-readable storage medium, wherein the system is any one of a computer, a server, or a microcontroller, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes program instructions stored in the computer-readable storage medium.

[0072] Specifically, the principle of this invention is as follows: The fundamental reason why this invention can solve the above-mentioned technical problems is that it explicitly embeds physical constraints into the entire chain of model structure, sequence modeling, and online learning, rather than imposing constraints only at the input or output end. The attention weights of the graph attention network layer incorporate a physical regularization factor composed of the product of Kerr nonlinear coefficients and optical power, making the graph convolution aggregation process mathematically equivalent to an approximate solution of the physical laws of nonlinear transmission, ensuring the rationality of the recommended parameters in the physical space. The Mamba selective state space module, through a dynamic generation mechanism of input-dependent state transition matrices, mathematically realizes adaptive long-range memory for non-stationary time series, ensuring that the model's ability to track aging drift does not decay over time. The affine projection algorithm constrains the update amount of the equalizer tap coefficients to find the minimum norm solution within the affine subspace spanned by multiple frame input vectors, algebraically reducing the interference of stochastic gradient noise on the update direction and resolving the inherent contradiction between convergence speed and stability in fixed-step-size algorithms. Elastic weight consolidation regularization quantifies the importance of parameters to historical tasks using the Fisher information matrix. By delineating a historical knowledge protection region in the parameter space through a quadratic penalty term, the online update achieves a Pareto optimal balance between adapting to the new distribution and preserving historical knowledge. The above mechanisms work together to form a logically closed-loop three-dimensional coupled perturbation response system.

[0073] The following provides a specific embodiment 1 of the present invention, and the specific implementation of each step in this embodiment 1 is described in detail below.

[0074] The specific implementation method of step S01 is as follows.

[0075] Real-time acquisition of the bias current of the optical module (Unit: mA), Extinction Ratio (Unit: dB) Optical signal-to-noise ratio (Unit: dB) Temperature (Unit: °C) and service duration (Unit: h), calculate the bias current drift. and temperature deviation ,in Rated bias current (unit: mA). The standard operating temperature (unit: °C) is used. Substituting the above parameters into the aging-temperature coupling bias index calculation formula, the formula is as follows:

[0076] ;

[0077] In the formula, The aging-temperature coupled bias index is dimensionless, and the three terms normalize current drift, temperature deviation, and service life to dimensionless values, respectively. Device design lifetime (unit: hours) , , These are weighting coefficients, all with a dimension of 1. They were obtained through multiple linear regression analysis of experimental data from no fewer than 200 sets of devices within a temperature range of 0–70℃ and a service life range of 0–50,000 hours. Typical values ​​are approximately 0.4, 0.35, and 0.25, respectively. The online learning update intensity level is determined according to the following rules within the specified interval: When When, the level is low; when At that time, the level is medium; when At that time, the level was high, and the boundary values ​​of the above intervals were determined by statistical analysis of the critical states in the experimental data where the prediction error exceeded 5%. Furthermore, the intensity index was updated online. The calculation formula is expressed as follows:

[0078] ;

[0079] In the formula, To update the strength index online, with a dimension of 1, the two terms are weighted sums of the aging-temperature coupling bias index and the mean square error ratio. The mean square error between the recommended parameter configuration vector and the measured optimal parameter configuration vector, with units consistent with the dimensions of the recommended parameter configuration vector. To validate the baseline mean square error on the validation set, units are... Same, make Normalized to dimensionless 1 , The weighting coefficients, all with a dimension of 1, were obtained through Pareto front analysis of the forgetting rate and tracking speed in at least 50 online deployment iterations, with typical values ​​around 0.5. Pareto front analysis uses forgetting rate and tracking speed as two optimization objectives, identifying the Pareto optimal solution set from the iterative experimental results, and selecting the solution with the best overall score from this set. and Values ​​can be obtained. When When the elastic weight consolidation regularization coefficient is set to a high value, the learning rate is set to a low value; when When both are at the middle value, when At that time, the elastic weight consolidation regularization coefficient is set to a low value, and the learning rate is set to a high value. The specific values ​​of each level are determined by comprehensive scoring through no less than 30 iterative experiments on different drift rate scenarios on the simulation platform.

[0080] The specific implementation method of step S02 is as follows.

[0081] Collected , , , Four normalized physical parameters construct nodes The input feature vector is expressed by the following formula:

[0082] ;

[0083] In the formula, For nodes The input feature vector has a dimension of 1. For nodes The bias current of the corresponding optical module (unit: mA) For nodes Extinction ratio of the corresponding optical module (unit: dB) For nodes Optical signal-to-noise ratio (SNR) of the corresponding optical module (unit: dB) For nodes Temperature of the corresponding optical module (unit: °C) Rated extinction ratio (unit: dB) The rated optical signal-to-noise ratio (in dB) is used. The fiber dispersion coefficient is used as the pre-acquired value. (unit: ) and link distance (Unit: km) Structural edge The eigenvectors of are expressed by the following formula:

[0084] ;

[0085] In the formula, For the edge The eigenvectors of , with dimensions 1 For the edge The dispersion coefficient of the corresponding optical fiber (unit: ) Reference dispersion coefficient (unit: ) For the edge Corresponding link distance (unit: km) Use the reference link distance (unit: km) to construct a weighted heterogeneous graph. ,in For a set of nodes, For the edge set. The node feature embedding layer uses a two-layer fully connected network to... Mapping to a 64-dimensional latent space vector, the formula is as follows:

[0086] ;

[0087] In the formula, For nodes The initial 64-dimensional latent space vector, with dimensions 1 This is the weight matrix of the first layer of the node embedding layer, with a dimension of 1. This is the weight matrix of the second layer of the node embedding layer, with a dimension of 1. This is the bias vector for the first layer of the node embedding layer, with a dimension of 1. This is the bias vector of the second layer of the node embedding layer, with a dimension of 1. It is the activation function of the rectified linear unit, that is, taking the input element by element. Operation. Edge features are mapped to a 64-dimensional edge latent space vector via an independent fully connected network, as expressed in the following formula:

[0088] ;

[0089] In the formula, For the edge A 64-dimensional latent space vector with dimensions 1 This is the weight matrix of the first layer of the edge embedding layer, with a dimension of 1. This is the weight matrix of the second layer of the edge embedding layer, with a dimension of 1. This is the bias vector of the first layer of the edge embedding layer, with a dimension of 1. This is the bias vector for the second layer of the edge embedding layer, with a dimension of 1. In the... Physical constraint regularization factors are incorporated into the layered graph attention network. The formula for calculating attention score is as follows:

[0090] ;

[0091] In the formula, For the first Layer nodes For neighboring nodes Attention weights, with a dimension of 1 For the first Layer attention weight matrix, with a dimension of 1 For nodes In the The latent space vector of the layer has a dimension of 1. For nodes In the The latent space vector of the layer has a dimension of 1. For nodes In the The latent space vector of the layer has a dimension of 1. For the edge A 64-dimensional latent space vector with dimensions 1 This indicates vector concatenation, such that... For nodes The set of neighboring nodes. Physical constraint regularization factor. The calculation formula is expressed as follows:

[0092] ;

[0093] In the formula, The physical constraint regularization factor has a dimension of 1. Kerr nonlinear coefficient (unit: (Obtained through pre-calibration of optical fiber material parameters) For nodes Current optical power (unit: mW, obtained in real time by optical module monitoring). Reference optical power (unit: mW) The unit is , The unit is Dividing the two makes Normalized to dimensionless 1 As a physical penalty term superimposed on the attention score, a larger Kerr nonlinearity coefficient or higher optical power corresponds to stronger physical suppression of the attention weights, causing the graph convolution aggregation process to follow the physical laws of nonlinear transmission. The node state update formula is expressed as follows:

[0094] ;

[0095] In the formula, For the first The layer output transformation matrix has a dimension of 1. For the first Layer transformation matrix, dimensionless For nodes In the The latent space vectors of each layer have a dimension of 1. After passing through three layers of graph attention network, the temporal sequence of the node latent space vectors is fed into the Mamba selective state space module. The diagonal elements of its state transition matrix are dynamically generated by linear projection from the current input. This module selectively preserves input dependencies for non-stationary long-range temporal dependencies caused by aging, unlike the fixed state transition mechanism in traditional recurrent neural networks. It is suitable for modeling non-stationary long-range temporal dependencies caused by optical module aging. The node state monitoring module continuously calculates the node input feature vectors at adjacent time steps. and Maharanobis distance The formula is expressed as follows:

[0096] ;

[0097] In the formula, Let Maharanobis distance be the dimensionless distance. For nodes At any moment The input feature vector has a dimension of 1. For nodes At any moment The input feature vector has a dimension of 1. Input the covariance matrix of the feature vectors for historical nodes, with a dimension of 1, to automatically eliminate the influence of differences in the dimensions and variances of various parameters. for The inverse matrix of the given matrix has a dimension of 1, such that... Keep the dimensionless value 1; when Exceeding the mutation detection threshold (When taking no less than 100 actual test jump events) When the 99th quantile of the distribution (dimension 1) is reached, the bypass channel will be triggered. The input-output configuration header bypasses the graph attention network layer. The output configuration header configures the recommendation vector using the output parameters of a two-layer fully connected network. Furthermore, a soft constraint loss term based on physical knowledge is applied, and the total loss function is expressed as follows:

[0098] ;

[0099] In the formula, This is the total loss function, with a dimension of 1. Total number of training samples For the model to nodes Output parameter configuration recommendation vector For nodes The corresponding optimal parameter configuration annotation vector is determined by the expert system using exhaustive search, with units equal to... same Configure the reference normalized vector for the recommended vector of the parameters, with units equal to... Same, make Normalized to dimensionless 1 The weights for the soft constraint loss term based on physical knowledge, with a dimension of 1, are determined using a grid search on the validation set. The constraint loss term, with a dimension of 1, is used to penalize configuration recommendation vectors that violate the steady-state conditions of the laser rate equation. The model is trained using the Adam optimizer with an initial learning rate of 0.001 for a total of no less than 200 epochs. The training dataset consists of no less than 50 optical modules from different batches and with different service durations deployed on a test bench, covering a temperature range of 0–70℃ and no less than 5 batches with different manufacturing process deviations. No less than 100,000 labeled samples are constructed, stratified by temperature range, service duration range, and manufacturing process batch to ensure that the proportions of each stratum are consistent between the training set and the validation set.

[0100] The specific implementation method of step S03 is as follows.

[0101] The discrete parameter configuration states in the optical module parameter space are constructed as a weighted undirected graph. ,in Configure the set of state nodes for the parameters. Configure the state edge set for the parameters, nodes With nodes Between the boundary weights The calculation formula is expressed as follows:

[0102] ;

[0103] In the formula, For weighted undirected graphs Middle node With nodes The edge weights between the two parameters are dimensionless (1). A larger edge weight indicates a higher physical similarity between the two parameter configuration states. For nodes Corresponding fiber dispersion coefficient (unit: ) For nodes Corresponding fiber dispersion coefficient (unit: ) For nodes Corresponding Kerr nonlinear coefficient (unit: ) For nodes Corresponding Kerr nonlinear coefficient (unit: ) Reference normalized values ​​for Kerr nonlinear coefficients (unit: ) To prevent small positive numbers with a denominator of zero, the dimension is 1, and the empirical value is... The weighted undirected graph can be represented as a resistive network, with the graph Laplace matrix... The definition is as follows:

[0104] $L_{p,ij}=\begin{cases}{\displaystyle\sum_{k\neq i}w_{ik},}&i=j\\-w_{ij},&i\neq j,\,(i,j)\in\mathcal{E}_{p}\\0,&\text{Other}\end{cases}$;

[0105] In the formula, For the graph Laplace matrix No. Line 1 Column elements, dimensionless For weighted undirected graphs The total number of nodes; diagonal elements are the sum of all edge weights for the corresponding node, and off-diagonal elements are the negative values ​​of the corresponding edge weights. Configure state nodes with current parameters. With a unit voltage excitation applied and the remaining nodes grounded, establish the graphical Laplace linear equations:

[0106] ;

[0107] In the formula, The potential vectors at each node represent the global parameter perturbation propagation potential field, with dimensions of 1. This is the activation vector, with a dimension of 1, which only activates the source nodes. The corresponding component is 1, and the rest are 0. The sparse conjugate gradient method is used to solve the above system of equations, with a computational complexity of O(n log n). Order of magnitude. Solving for the desired order of magnitude. Then, a greedy search is performed along the potential gradient direction, while simultaneously applying a gradient proportional to the edge weight. The probability is used to perform a random walk for several steps before performing a greedy search. The two are combined to obtain an optimized set of parameter configuration candidates, so that the search process is dynamically balanced between exploring new regions and converging to a high-quality state. The effective resistance value between nodes corresponds to the parameter configuration jump cost, so that the search path tends to jump between parameter configuration states with high physical similarity.

[0108] The specific implementation method of step S04 is as follows.

[0109] Equalizer tap coefficient vector (dimension 1, The number of equalizer taps is updated using a variable step size minimum mean square-affine projection joint algorithm. Frame equalization error The calculation formula is expressed as follows:

[0110] ;

[0111] In the formula, For the first Frame equalization error, with dimensions consistent with signal dimensions. For the first The expected signal in the frame has the same dimensions as the signal. For the first The input signal vector of the frame equalizer has the same dimensions as the signal. This is the inner product of the tap coefficient vector and the input signal vector, with dimensions consistent with the signal dimensions. Error power moving average. The recursive formula is expressed as follows:

[0112] ;

[0113] In the formula, For the first Frame error power moving average, dimensionless. For the first Frame error power moving average, dimensionless. This is the forgetting factor, with a dimension of 1, and is typically taken as 0.95. This is a reference value for signal power, with units equal to... Dimensional consistency, making Normalized to a dimensionless value of 1. Adaptively update step size. The calculation formula is expressed as follows:

[0114] ;

[0115] In the formula, For the first Frame adaptive update step size, with a dimension of 1 This is the upper limit of the step size, with a dimension of 1. This is the lower limit of the step size, with a dimension of 1. To adjust the slope, the dimensionless unit is 1. The reference error power moving average is dimensionless; the above parameters were determined by grid search experiments under no fewer than 20 combinations of optical signal-to-noise ratio and fiber dispersion coefficient; as the error power moving average increases... Increase to accelerate convergence, when the moving average of error power decreases. Reduced to fine-grained adjustments. The affine projection algorithm utilizes the nearest neighbor... Frame input vectors form a matrix ,recent Frame expected signal vector affine projection error vector The calculation formula is expressed as follows:

[0116] ;

[0117] In the formula, For the first The frame affine projection error vector has the same dimensions as the signal. for The transpose, and Multiplying yields the closest result The equalized output vector of the frame has the same dimensions as the signal. The tap coefficient update formula is expressed as follows:

[0118] ;

[0119] In the formula, For the first Frame tap coefficient vector, dimensionless To regularize a small quantity, with a dimension of 1, the empirical value is... for An identity matrix of order 1 with dimensions 1 The mean square norm penalty coefficient for the tap coefficients, with a dimension of 1, is used to suppress tap coefficient oscillations. The optimal value is determined through the above grid search experiment. Normalize the affine projection error vector to a dimensionless value of 1, and update the tap coefficients. Keeping the dimensionless constant, and dimensionless Project the updated direction onto the nearest The affine subspace spanned by the frame input vectors is used to reduce the update variance by utilizing multi-frame error vector information.

[0120] The specific implementation method of step S05 is as follows.

[0121] The online learning update intensity level obtained from S01 (by...) Based on the refined definition, elastic weight consolidation regularization is applied to the physical constraint embedded configuration optimization model for online updates. Let the model parameter vector be... The first in historical task data Diagonal approximation elements of the Fisher information matrix with parameters for:

[0122] ;

[0123] In the formula, For the first The diagonal approximation elements of the Fisher information matrix for each model parameter, with a dimension of 1, reflect the importance of that parameter to the historical task. Total number of historical task samples For the first Label vectors of historical samples For the first Input feature vector of historical samples The optimal model parameter vector after training for the historical task. for The Each component To use the optimal parameters of historical tasks The model output probability is conditional and has a dimension of 1. The elastic weight consolidation regularization online update loss function is expressed as follows:

[0124] ;

[0125] In the formula, The total loss is updated online using elastic weight consolidation regularization, with a dimension of 1. The loss on the current distribution data, with a dimension of 1. The current model parameter vector The Each component For the first The reference normalized value of each parameter, and Having the same dimensions makes Normalized to dimensionless 1 To reinforce the regularization coefficients for elastic weights, the coefficients are dimensionless (1) and are based on... The corresponding range value is taken from the range; the regularization term is applied to... Larger historically significant parameters are subject to stronger secondary penalties to prevent catastrophic amnesia. The aging physics prior model describes the bias current. Compared to extinction ratio The degradation follows a monotonic degradation pattern over time. The degradation pattern parameters are obtained by curve fitting of time-series data collected from no less than 50 optical modules in accelerated aging experiments. These parameters serve as constraints for parameter update direction, making the model parameter updates biased towards the direction that conforms to the physical degradation pattern of aging.

[0126] The specific implementation method of step S06 is as follows.

[0127] The optimized parameter configuration candidate set obtained in S03 and the stable convergent tap coefficient configuration obtained in S04 are distributed to the edge-center two-level collaborative architecture. High-frequency small-amplitude perturbations with bias current changes below 5% of the rated bias current and temperature changes below 3℃ are handled by the optical module's local adaptive rule engine. The cross-module collaborative optimization task is routed to the central server, which calls the physical constraint embedded configuration optimization model and the optical module parameter perturbation propagation algorithm to solve the problem collaboratively, completing the final parameter configuration distribution.

[0128] To better understand and implement this invention, a specific application scenario of the invention is provided below as an example 2: To verify the effectiveness of the invention, technicians built a test environment and deployed an optical module adaptive parameter configuration system in a data center interconnection scenario. The present invention's solution was used to manage the full lifecycle parameter configuration of 32 optical module nodes. The test environment covered a temperature range of 0–70°C, the optical module service life spanned 0–40,000 hours, encompassing five batches with different manufacturing process deviations. The link used both 16QAM (16th-order quadrature amplitude modulation) and PAM4 (4th-order pulse amplitude modulation) modulation formats, and the fiber dispersion coefficient ranged from 16 to 20. The link distance ranges from 40 to 120 km.

[0129] Technicians first followed step S01 to collect real-time data on the bias current, extinction ratio, optical signal-to-noise ratio, temperature, and service life of 32 optical modules, with a sampling period set to 60 seconds. Taking a typical optical module node as an example, at the start of the test, its bias current drift was... The rated bias current is 4.2mA. 80mA, temperature deviation The standard operating temperature is 12℃. 25℃, service life Design life is 28,000 hours. Substituting 100,000 hours into the aging-temperature coupling bias index calculation formula, , , Using typical values ​​of 0.4, 0.35, and 0.25, the aging-temperature coupling bias index was calculated. The value is 0.388, which falls within the medium-level range, and the online learning update intensity level is determined to be medium. Table 1 lists the physical parameter acquisition values ​​of six typical optical module nodes during the test period, along with their corresponding aging-temperature coupling bias index and online learning update intensity level.

[0130] Table 1. Statistics of Physical Parameters and Aging-Temperature Coupling Bias Index of Typical Optical Module Nodes

[0131]

[0132] Following step S02, the normalized bias current, extinction ratio, optical signal-to-noise ratio, and temperature of the 32 optical module nodes are used as graph node features, and the fiber dispersion coefficient and link distance of each link are used as graph edge features, constructing a weighted heterogeneous graph containing 32 nodes. The three-layer graph attention network layer of the physically constrained embedded configuration optimization model incorporates a physical constraint regularization factor, which is composed of the normalized product of the Kerr nonlinear coefficient of each link and the optical power of the node, when aggregating neighbor node features. Figure 3 As shown, the distribution of selective retention weights for the historical states of nodes with different service durations is illustrated after 32 nodes are aggregated in a 3-layer graph attention network. It is evident that the retention weights for the historical states of nodes with long service durations (such as M03 and M06) are significantly higher than those for nodes with short service durations, demonstrating the model's adaptive long-range memory capability in the face of aging non-stationary drift. For node M03, when the service duration reaches 42,000 hours, the Mahalanobis distance between the input feature vectors of adjacent sampling times exceeds the mutation detection threshold, triggering a mode-hopping bypass mechanism. The original node input feature vector bypasses the graph attention network layer and directly inputs and outputs the configuration head, ensuring the accuracy of configuration recommendations under mode-hopping events.

[0133] Following step S03, using the current configuration state of node M01 as the excitation source, the optical module parameter perturbation propagation algorithm is executed on the weighted heterogeneous graph. The optical module parameter space is discretized into 1024 parameter configuration state nodes. The edge weight between two nodes is defined by the reciprocal of the Euclidean distance between the normalized difference of the fiber dispersion coefficient and the Kerr nonlinear coefficient. A sparse graph Laplacian matrix is ​​constructed and solved using the sparse conjugate gradient method to obtain the global parameter perturbation propagation potential field. A greedy search is performed along the potential gradient direction, and a random walk step is performed with a probability proportional to the edge weight. Finally, an optimized parameter configuration candidate set containing 8 candidate parameter configuration states is obtained. The effective resistance value between nodes ensures that the search path is concentrated in the parameter configuration region with high physical similarity.

[0134] Following step S04, using the bias current, extinction ratio, optical signal-to-noise ratio, and temperature of node M01 as inputs, the equalizer tap coefficients are updated using a variable-step-size least mean square-affine projection joint algorithm. The affine subspace consists of the input vector and error vector from the most recent 8 frames. The error power moving average window length is 32 frames, and the tap coefficient regularization constraint coefficients are determined through grid search. Table 2 lists the stable convergence values ​​and corresponding error power moving averages of the equalizer tap coefficients after being updated using the variable-step-size least mean square-affine projection joint algorithm under different optical signal-to-noise ratio conditions.

[0135] Table 2. Statistical table of convergence values ​​and moving average values ​​of equalizer tap coefficients and error power under different optical signal-to-noise ratios.

[0136]

[0137] Following step S05, taking node M01 with an online learning update intensity level as an example, the online update intensity index is... The mean squared error normalization is calculated by combining the recommended vector with the measured optimal parameter configuration vector. The value is 0.41, which is in the middle range. The elastic weight consolidation regularization coefficient and learning rate are both set to the middle range. The important historical parameters are protected by the diagonal approximation of the Fisher information matrix. At the same time, the parameter update direction is constrained by the aging physical prior model. The physical constraint embedded configuration optimization model is gradually adapted to the current distribution to prevent catastrophic forgetting.

[0138] Following step S06, the optimized parameter configuration candidate set and the stable converged tap coefficient configuration are distributed to the edge-center two-level collaborative architecture. For nodes with bias current changes below 5% of the rated bias current and temperature changes below 3℃ (such as M02 and M05), the optical module's local adaptive rule engine directly processes the task, reducing communication overhead. For cross-module collaborative optimization tasks such as nodes M03 and M06, the task is routed to the central server, which calls the physical constraint embedded configuration optimization model and the optical module parameter perturbation propagation algorithm to solve the problem collaboratively, generating the final parameter configuration and distributing it.

[0139] Compared to traditional configuration methods based on rule lookup tables or fixed-step least mean square algorithms, this invention directly embeds physical constraints into the attention weights of the graph attention network layer. This ensures that the configuration recommendation process is constrained by the nonlinear transmission physical laws during the aggregation stage, rather than relying on post-event verification. This fundamentally eliminates the possibility of physically unreasonable configuration vectors propagating in the network. The input-dependent dynamic selection mechanism of the Mamba selective state space module prevents the model's memory of aging non-stationary drift from decaying over time, overcoming the intrinsic limitations of the fixed state transition mechanism of traditional recurrent neural networks in non-stationary scenarios. Elastic weight consolidation regularization delineates historical knowledge protection regions in the parameter space through the Fisher information matrix, mathematically guaranteeing a gradual balance between new and old knowledge in online learning, rather than a simple sliding window forgetting mechanism. The synergistic effect of these mechanisms ensures that the parameter configuration accuracy of this invention remains stable under three-dimensional coupled perturbations.

[0140] It should be noted that the variables involved in this invention are explained in detail in Tables 3, 4, and 5.

[0141] Table 3. Variable Explanation Table (Part 1)

[0142]

[0143] Table 4. Variable Explanation Table (Part Two)

[0144]

[0145] Table 5. Variable Explanation Table (Part 3)

[0146]

[0147] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for adaptive parameter configuration of an optical module, characterized in that, Includes the following steps: The bias current, extinction ratio, optical signal-to-noise ratio, temperature and service life of the optical module are collected in real time. The bias current, extinction ratio, optical signal-to-noise ratio, temperature and service life are input into the aging-temperature coupled bias index calculation formula to obtain the aging-temperature coupled bias index. The online learning update intensity level is determined according to the interval to which the aging-temperature coupled bias index belongs. The collected bias current, extinction ratio, optical signal-to-noise ratio, and temperature are used as graph node features, and the pre-collected fiber dispersion coefficient and link distance are used as graph edge features to construct a weighted heterogeneous graph. The weighted heterogeneous graph is then input into a physical constraint embedded configuration optimization model to obtain an initial parameter configuration recommendation vector. Using the current configuration state corresponding to the collected bias current, extinction ratio, optical signal-to-noise ratio and temperature as the excitation source, the optical module parameter perturbation propagation algorithm is executed on the constructed weighted heterogeneous graph to solve the graph Laplace linear equation system to obtain the global parameter perturbation propagation potential field. Greedy search is performed along the gradient direction of the global parameter perturbation propagation potential field and random walk exploration is fused to obtain the optimized parameter configuration candidate set. Using the collected bias current, extinction ratio, optical signal-to-noise ratio and temperature as inputs, the equalizer tap coefficients are updated using a variable step size minimum mean square-affine projection joint algorithm. The update step size is adaptively adjusted according to the moving average of error power, and tap coefficient regularization constraints are applied to obtain a stable and convergent tap coefficient configuration. Based on the obtained online learning update intensity level, elastic weight consolidation regularization online update is performed on the physical constraint embedded configuration optimization model. The aging physical prior model is used to constrain the parameter update direction, thus completing the gradual adaptation of the physical constraint embedded configuration optimization model to the current distribution. The optimized parameter configuration candidate set and the stable converged tap coefficient configuration are distributed to the edge-center two-level collaborative architecture. High-frequency small-amplitude disturbances with bias current changes less than 5% of the rated bias current and temperature changes less than 3℃ are processed by the optical module's local adaptive rule engine. The cross-module collaborative optimization task is routed to the central server, which calls the physical constraint embedded configuration optimization model and the optical module parameter disturbance propagation algorithm to solve the problem collaboratively, and completes the final parameter configuration distribution.

2. The optical module adaptive parameter configuration method according to claim 1, characterized in that, The optical module parameter perturbation propagation algorithm constructs a weighted undirected graph from the discrete parameter configuration state in the parameter space. The edge weight between two nodes is defined by the reciprocal of the Euclidean distance between the normalized difference of the fiber dispersion coefficient and the normalized difference of the Kerr nonlinear coefficient. The solution of the graph Laplace linear equations adopts the sparse conjugate gradient method.

3. The optical module adaptive parameter configuration method according to claim 2, characterized in that, The variable step size minimum mean square-affine projection joint algorithm uses the error vectors of the most recent frames to form an affine subspace during each update, projects the update direction of the equalizer tap coefficients to the affine subspace, and applies tap coefficient regularization constraints by superimposing a mean square norm penalty term onto the cost function.

4. The optical module adaptive parameter configuration method according to claim 3, characterized in that, The rule for adaptively adjusting the update step size based on the moving average of the error power is as follows: when the moving average of the error power increases, the update step size increases to accelerate convergence; when the moving average of the error power decreases, the update step size decreases to fine-tune the process. The upper and lower limits of the update step size are determined by grid search.

5. The optical module adaptive parameter configuration method according to claim 4, characterized in that, The core of the physical constraint embedded configuration optimization model is a 3-layer graph attention network. The attention weight calculation of each graph attention network layer incorporates a physical constraint regularization factor, which is formed by normalizing the product of the Kerr nonlinear coefficient and the optical power of the current node.

6. The optical module adaptive parameter configuration method according to claim 5, characterized in that, The physical constraint embedded configuration optimization model sets up a Mamba selective state space module after the 3-layer graph attention network layer. It dynamically generates the diagonal elements of the state transition matrix by linear projection of the 64-dimensional latent space vector, and establishes an input-dependent selective historical physical state retention mechanism.

7. The optical module adaptive parameter configuration method according to claim 6, characterized in that, The physical constraint embedded configuration optimization model sets up a mode skipping bypass mechanism. The node state monitoring module continuously calculates the Mahalanobis distance between the input feature vectors of nodes at adjacent time points. When the Mahalanobis distance exceeds the mutation detection threshold, the bypass channel is triggered to bypass the original node input feature vectors and directly input them to the output configuration head, bypassing the 3-layer graph attention network layer.

8. The optical module adaptive parameter configuration method according to claim 7, characterized in that, The aging physics prior model is a physical equation describing the monotonic degradation of the bias current and extinction ratio of optical module devices with service time. The degradation parameters are obtained by curve fitting of time-series data collected from no less than 50 optical modules in accelerated aging experiments.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, which, when executed in a computer, are used to perform the optical module adaptive parameter configuration method according to any one of claims 1-8.

10. An adaptive parameter configuration system for an optical module, characterized in that, The system comprises the computer-readable storage medium of claim 9, wherein the system is a computer, the computer-readable storage medium is disposed within the system, and the system is provided with a microprocessor that executes program instructions stored in the computer-readable storage medium.