A nonlinear equalization method based on regression coupling value of GRU neural network

By using a nonlinear equalization method based on a GRU neural network with regression coupling values, the problems of high complexity and model training cost of learning equalizers in optical fiber communication are solved. This method achieves efficient nonlinear impairment compensation and crosstalk capture in long-distance optical fiber communication systems, improving the robustness and adaptability of the algorithm.

CN117938264BActive Publication Date: 2026-05-15BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2024-01-31
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing learning-based equalizers in fiber optic communication suffer from problems such as susceptibility to data bias, overfitting, high complexity, and divergent equalization results. They are difficult to design with refined unit structures and define accurate model parameters, resulting in high training costs and insufficient nonlinear equalization capabilities.

Method used

A nonlinear equalization method based on regression coupling value of GRU neural network is adopted. The error factor is calculated by reshaping the shape of input samples, gating recursive units and the distributed Fourier method of nonlinear Schrödinger equation, which guides the network parameter update, optimizes the training process, reduces complexity and improves nonlinear damage compensation capability.

Benefits of technology

It effectively reduces model training costs and complexity, improves the compensation capability and stability of nonlinear equalization algorithms, and can flexibly adapt to channel changes in long-distance optical fiber communication systems, reducing training time and improving the capture and compensation effect of signal nonlinear impairments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117938264B_ABST
    Figure CN117938264B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of optical fiber communication, and particularly relates to a GRU neural network nonlinear equalization method based on a regression coupling value, which first reshapes the data structure of a received signal sequence and matrixes it; then in the training stage, a nonlinear damage error factor based on a nonlinear Schrodinger equation is used to guide network updating, and in the application stage, a signal is input into a GRU neural network, and a signal nonlinear damage compensation result is obtained through the propagation algorithm of the network. The method of the present application uses a GRU neural network algorithm based on a regression coupling value to realize the capture and compensation of signal nonlinear damage, solves the problems of high algorithm complexity, multiple iteration numbers and limited nonlinear compensation capacity in digital back propagation and learning type equalization algorithms, further improves the effectiveness and practicability of the nonlinear equalization algorithm, and has an important application prospect in the field of digital signal processing related to optical communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optical fiber communication technology, and in particular to a nonlinear equalization method for GRU neural networks based on regression coupling values. Background Technology

[0002] Against the backdrop of surging internet traffic and services, the development of ultra-high-speed, ultra-large-capacity, ultra-long-distance, and ultra-strong-protection optical fiber communication systems is an inevitable trend. In order to cope with the exponentially increasing capacity demands from next-generation mobile networks and high-bandwidth internet applications, the field of optical fiber communication has begun to intensively research technologies that can fully utilize the capabilities of optical fibers. Regardless of the capacity enhancement method adopted, the main limiting factor for capacity will ultimately be the nonlinear Shannon capacity limit of transmitted information.

[0003] In long-distance, high-bandwidth optical networks, this limitation is mainly attributed to the nonlinearities within and between fiber channels induced by the Kerr effect, including effects such as self-phase modulation, four-wave mixing, optical amplifier noise, spontaneous polarization, and Raman scattering. These effects cause problems such as phase distortion, frequency drift, and power fluctuations in optical signals during transmission, limiting the performance and distance of optical fiber communication systems.

[0004] The digital backpropagation algorithm iteratively propagates the error signal backward, adjusting the signal's phase and amplitude based on nonlinear interactions during transmission. It effectively suppresses nonlinear effects such as self-phase modulation and four-wave mixing, improving the transmission quality of optical signals. In contrast, the learning-based equalization algorithm uses deep learning to build a neural network, enabling it to automatically learn the complex characteristics of nonlinear communication systems without requiring prior knowledge.

[0005] Network models can parameterize single-mode fiber transmission formulas, and neural networks can be trained using forward computation and backpropagation. Learning-based equalization algorithms have greater versatility due to their applicability to various signal-to-noise ratios, signal formats, and channel environments.

[0006] However, learning equalizers have problems such as susceptibility to data bias, overfitting, high complexity, and divergent equalization results. Some learning equalizers have excessively large unit structures, resulting in high complexity and a large requirement for training samples. At the same time, some learning equalizer algorithms deviate from the channel model, causing the network model to be unable to effectively capture nonlinear impairment information, resulting in poor robustness of the equalization results.

[0007] The current challenge lies in how to achieve a more refined unit structure design and more accurate model parameter definition, reduce model training costs, lower algorithm complexity, and further improve the algorithm's nonlinear equilibrium capability. Summary of the Invention

[0008] This invention aims to address the technical challenges in existing technologies, such as how to achieve more refined unit structure design and more accurate model parameter definition, reduce model training costs, and lower algorithm complexity. It provides a nonlinear equilibrium method for GRU neural networks based on regression coupling values.

[0009] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0010] A nonlinear equilibrium method for GRU neural networks based on regression coupling values ​​includes the following steps:

[0011] Reshape the input sample.

[0012] Adjust the data structure of the received signal sequence and matrix it;

[0013] The received QAM samples of length K are normalized and divided into K-2N subsamples using a sliding window;

[0014] The gate structure element, combined with the error factor, guides the network parameter update process, specifically as follows:

[0015] A GRU-based neural network model is established, where the reshaped signal is fed into the input layer, then calculated by gated recurrent units in the hidden layer through activation functions, and finally fully connected and expanded in the output layer; and

[0016] The training process involves iterative steps, namely, updating network parameters using a loss function and backpropagation during a training process to achieve neural network training;

[0017] The loss function and backpropagation update network parameters are calculated using the distributed Fourier method based on the nonlinear Schrödinger equation to obtain the nonlinear damage error factor and guide the network update.

[0018] Specifically, the nonlinear Schrödinger equation is decomposed using the split-step Fourier method to obtain the error factor Δb at each current moment within the fiber span, and the error factor b is added to the loss function and activation function. t Optimize the network update process;

[0019] During application, the output sequence obtained by the output layer is the result of nonlinear compensation of the original signal;

[0020] Finally, in the application stage, the signal is input into the GRU neural network, and the signal nonlinear damage compensation result is obtained through the network's propagation algorithm.

[0021] The signal is simplified as follows:

[0022] s(t)=Re[s(t)·e j2πfc (1);

[0023] Where s(t) is the baseband signal in complex form, Re[·] denotes taking the real part of the complex number, and e j2πfc t It is a carrier wave, f c It is the carrier frequency.

[0024] Specifically, the process of normalizing the received QAM samples of length K and dividing them into K-2N sub-samples through a sliding window is carried out by dividing them into K-2N sub-samples through a sliding window with a width of 2N+1 and a step size of 1.

[0025] Here, the window width represents the size of the input features, so that the new input matrix can be represented in the following form:

[0026]

[0027] The subsamples are reshaped into three-dimensional tensors, and each vector in the features and labels is decomposed into real and imaginary parts so that the data constitutes the input layer samples.

[0028] Specifically, the process of updating network parameters guided by gate update and updating network parameters guided by error factor also includes:

[0029] The steps for defining the network and creating update and reset gates include applying formulas (3) and (4):

[0030] R t =σ(X) t W xr +H t-1 W hr +b r (3);

[0031] Z t =σ(X) t W xz +H t-1 W hz +b z (4);

[0032] Where σ represents the sigmoid function that converts the value to a range between 0 and 1, W xr ∈R m ×n and W xr ∈R m ×n represents the weight matrix involved in the learning process, b r ∈R1×n and b z ∈R1×n represents the bias value;

[0033] The formula for calculating the candidate hidden state t at the current time step is as follows:

[0034]

[0035] Where σ represents the sigmoid function that converts the value to a range between 0 and 1, W xh ∈R m ×n and W hh ∈R m ×n represents the weight matrix involved in the learning process, b z ∈R1×n represents the bias value, and ⊙ represents the Hadamard product;

[0036] The candidate states are combined with the update gate using the following formula:

[0037]

[0038] Obtain the model's memory state or output at the current time t.

[0039] Specifically, during the training phase of the training process, the sequence of the input layer enters the hidden layer, and all GRU units are fully connected to the hidden layer.

[0040] First, calculate the error factor:

[0041] The error factor originates from the channel nonlinearity. For the propagation function, the relationship given by the nonlinear Schrödinger equation is used to discretize the propagation function, and the formula is as follows:

[0042]

[0043] Where α represents attenuation, β2 represents the second-order dispersion coefficient, and γ represents the fiber nonlinearity coefficient;

[0044] Fourier transform is used to convert between the time and frequency domains, decomposing the nonlinear Schrödinger equation into E0... x and E y The differential formulas in both directions handle the nonlinear effects in the frequency domain, and then transform them back to the time domain to obtain the general form of the error components:

[0045]

[0046] The error factor is calculated as follows:

[0047]

[0048] Then, using the calculation factor as the error component, the output at the current time is calculated, as shown in formula (8):

[0049] y t =W Softmax h t +Softmax(b t (8).

[0050] Specifically, during training, an activation function is used to transform the input, generating the neuron's output (out). d The calculation formula is as follows:

[0051] out d =f ReLU (y t (9);

[0052] The network's predicted output and the actual labels are passed to the loss function to calculate the loss value. The loss function is estimated using the mean squared error of cross-validation.

[0053]

[0054] Finally, the difference between the network output and the actual signal is measured using a loss function. The derivative of each parameter is calculated, and the gradient of the loss with respect to the network parameters is calculated through the backpropagation algorithm. The parameters are then updated in the opposite direction of the gradient, thus achieving network training.

[0055] Specifically, in the application phase, the current time output shown in formula (8) is mapped to the original signal strength range, as shown in formula (11):

[0056] Y = y t σ+μ (11);

[0057] Where σ represents the standard deviation of the original data, and μ represents the mean of the original data.

[0058] Among them, formulas (1) to (2) represent the data recombination stage, formulas (3) to (7) represent the damage factor calculation stage, formulas (8) to (10) represent the network update stage, and formula (11) represents the network output in the actual application stage.

[0059] Specifically, the application conditions are: 16-QAM transmission system over long-distance standard single-mode fiber.

[0060] Specifically, it also includes the following steps:

[0061] First, the signal is sampled and ideal dispersion compensation is performed using an ideal frequency domain equalizer;

[0062] At the receiver, an ideal dispersion equalizer requires at least two samples per symbol to provide near-ideal dispersion compensation, with the expectation of focusing only on nonlinear impairments, i.e., higher-order terms, interaction terms, or other nonlinear transformation features of the data.

[0063] Preferably, the transmitting end step is as follows:

[0064] At the transmitting end, 1000 sets of pseudo-random binary sequences with a code length of 128 were first generated;

[0065] Each 4 bits is mapped to a 16QAM symbol, and after being upsampled twice, the baseband is shaped by a root-raised cosine filter.

[0066] The system uses 40 external cavity lasers with a frequency spacing of 100 GHz, one group for odd-numbered paths and one group for even-numbered paths, to output a total of 80 optical carriers with a spacing of 50 GHz, with a wavelength range of 1530-1562 nm.

[0067] The optical power of each wavelength channel is 13dBm;

[0068] This includes: resampling the digital baseband signal and loading it onto an arbitrary waveform generator with a sampling rate of 64GSa / s, converting it into two electrical signals to drive an IQ modulator with a 3dB bandwidth of 29GHz;

[0069] A loop structure with cyclic spans is used to achieve 1000km SSFM transmission, with each fiber span being a cyclic element;

[0070] The optical fiber is divided into 50 uniform steps per span, and an EDFA is added after each span to compensate for transmission loss, with an output power of 23dBm.

[0071] The wavelength selection switch is used to suppress power transfer and gain unevenness under high output power of EDFA. Its insertion loss is 5dB and the minimum adjustable pitch is 50GHz.

[0072] The two output signals are re-multiplexed and enter the next cycle.

[0073] Preferably, at the receiving end, the steps include:

[0074] The coherent optical receiver performs zero-difference detection on the optical signal selected after wave demultiplexing, and the receiving power is controlled at -2dBm;

[0075] The baseband electrical signal was captured by a synchronous 8-channel oscilloscope with a sampling rate of 80 GSa / s and a bandwidth of 33 GHz;

[0076] The receiver DSP process includes dispersion compensation, resampling, and clock recovery in the frequency domain;

[0077] The quadratic timing estimation algorithm is used to preserve four times the signal rate;

[0078] Nonlinear equalization based on regression-coupled value GRU neural network is performed. The decision algorithm adopts auxiliary least mean square algorithm. After decision, 16QAM symbol demapping and final bit error rate calculation are performed.

[0079] The present invention has the following beneficial effects:

[0080] Firstly, this technical solution uses a nonlinear equalization algorithm based on a GRU neural network in a coherent optical communication system to capture and compensate for nonlinear signal impairments. At the same time, it introduces regression coupling values ​​into the network to improve training and update efficiency, resolves the contradiction between complexity and performance in traditional digital backpropagation algorithms and learning-based equalization algorithms, and further enhances the compensation capability and stability of the nonlinear equalization algorithm.

[0081] Secondly, the trained network model is used to capture and compensate for nonlinear signal impairments. Compared with existing learning equalizer schemes, this scheme uses gated recursive units (GRUs) to reduce model complexity and introduces a regression coupling parameter based on the nonlinear Schrödinger equation to optimize the network update calculation process. This parameter can effectively utilize the time-domain information of the sequence and maximize the algorithm's ability to learn nonlinear impairments in the channel.

[0082] Thirdly, the method provided by this invention has important application prospects in high-speed, long-distance optical fiber transmission systems; the solution of this invention can effectively extract and "learn" the channel characteristics caused by nonlinear impairments under specific conditions, thereby effectively compensating for the nonlinear impairments and crosstalk between signals of the original signal.

[0083] When the channel length changes, this method does not require retraining the entire model. Retraining with only a small dataset is sufficient to adapt the network parameters to the new channel length, improving the model's flexibility and adaptability, and effectively reducing training costs and time consumption. Attached Figure Description

[0084] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0085] Figure 1 This is a flowchart of the algorithm of the present invention;

[0086] Figure 2 This is a flowchart of the signal processing at the receiving end of the present invention;

[0087] Figure 3 This invention relates to the multi-level network model structure of the GRU model.

[0088] Figure 4 This is a flowchart of the simulation and experiment of the optical communication system of the present invention;

[0089] The reference numerals in the figure are:

[0090] Input sequence length K, number of GRU layers M, sliding window size N, batch size n;

[0091] Signal error b, current input vector X t Output H at the current time tThe current hidden state Z t ; Detailed Implementation

[0092] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention. It should be noted that, for ease of description, in this application, "left side" is referred to as "first end", "right side" as "second end", "upper side" as "first end", and "lower side" as "second end" in the current view. The purpose of such description is to clearly express the technical solution and should not be construed as an improper limitation of the technical solution of this application.

[0093] This invention aims to address the technical challenges in existing technologies, such as how to achieve more refined unit structure design and more accurate model parameter definition, reduce model training costs, and lower algorithm complexity. It provides a nonlinear equilibrium method for GRU neural networks based on regression coupling values.

[0094] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following description, in conjunction with the accompanying drawings, uses a 32GBd 16-QAM transmission system over a 1000km standard single-mode fiber as an example. This embodiment is implemented based on the design method of this application and designs a nonlinear equalization method for GRU neural networks based on regression coupling values, but is not limited to this transmission system.

[0095] A nonlinear equalization method based on regression coupling values ​​using a GRU neural network is employed to capture and compensate for nonlinear signal impairments. The detailed algorithm flow is as follows: Figure 1 As shown in the diagram. In this method, the data structure of the received signal sequence is first adjusted and matrixed.

[0096] Next, during the training phase, the nonlinear damage error factor is calculated using the distributed Fourier method based on the nonlinear Schrödinger equation and used to guide network updates.

[0097] Finally, in the application stage, the signal is input into the GRU neural network, and the signal nonlinear damage compensation result is obtained through the network's propagation algorithm.

[0098] This invention mainly relates to the digital signal processing at the receiver of a single-polarization long-distance optical communication system, with a focus on compensation for signal nonlinear impairments. The principles of this invention are explained in detail below.

[0099] The receiver signal processing flow used in this method is as follows: Figure 2As shown, the signal is first sampled, and ideal dispersion compensation is performed using an ideal frequency domain equalizer. At the receiver, the ideal dispersion equalizer requires at least two samples per symbol to provide near-ideal dispersion compensation. Because we aim to focus only on nonlinear impairments—i.e., higher-order terms, interaction terms, or other nonlinear transformation characteristics of the data—we assume ideal carrier phase and frequency estimates in this case (without considering laser phase noise). Such signals exhibit significant nonlinear characteristics, and after data reconstruction and dimensional superposition, the network can more effectively capture nonlinear relationships.

[0100] The GRU network model structure used in this study is as follows: Figure 3 As shown;

[0101] By using the given relationships, the propagation formula is discretized, and Fourier transform is used to convert between the time and frequency domains. The nonlinear Schrödinger equation is decomposed into differential formulas in two directions, and the nonlinear effects in the frequency domain are handled. Then, it is converted back to the time domain expression. The general form of the error components can be obtained.

[0102]

[0103] The gated recurrent unit (GRU) network model structure includes a reset gate and an update gate. Through the control of these two gates, the GRU generates a candidate hidden state at each time step, which simultaneously considers the current input and the hidden state from the previous time step. By combining the current candidate state and the hidden state from the previous time step, the GRU calculates the final hidden state, which serves as the output for the current time step.

[0104] The gate update rule guides the GRU neural network to update its parameters in the next iteration through iterative computation. In the current time step t, suppose X... t H is an m-dimensional vector representing the input; t-1 R is an n-dimensional vector representing the hidden state at time t. Reset R t Determine whether the hidden state from the previous time step will be abandoned at time t, and update gate Z. t This will determine whether the state will be updated at time t.

[0105] Figure 4 The present invention provides an experimental model of coherent optical communication, which was tested on a 1000km standard single-mode fiber (SSFM) using a 32GBd 16-QAM transmission system. The setup of the experimental system is shown in the figure, and the specific process is as follows:

[0106] At the transmitter, 1000 sets of pseudo-random binary sequences with a code length of 128 were first generated. Each 4 bits were mapped to a 16QAM symbol, and after doubling the upsampling, baseband shaping was achieved through a root-raised cosine filter. Two sets (one set for odd-numbered paths and one set for even-numbered paths) of 40 external cavity lasers with a frequency spacing of 100 GHz were used to output a total of 80 optical carriers with a 50 GHz spacing, covering a wavelength range of 1530-1562 nm, conforming to the requirements of ITU-T standards C20C59 and H20-H59 channels. The optical power of each wavelength channel was 13 dBm. Subsequent steps included resampling the digital baseband signal and loading it onto an arbitrary waveform generator with a sampling rate of 64 GSa / s, converting it into two electrical signals to drive an IQ modulator with a 3 dB bandwidth of 29 GHz.

[0107] The experiment employed a loop structure with cyclic spans to achieve 1000km SSFM transmission, with each fiber span serving as a cyclic element. The fiber was divided into 50 uniformly spaced steps per span, and an EDFA (Electronic Distribution Array) was added after each span to compensate for transmission loss, with an output power of 23dBm. A wavelength selective switch was used to suppress power transfer and gain unevenness under high EDFA output power; its insertion loss was 5dB, and the minimum adjustable spacing was 50GHz. The two output signals were re-mode multiplexed and entered the next cycle until the fiber transmission distance met the requirements.

[0108] At the receiving end, the coherent optical receiver performs zero-difference detection on the optical signal selected after wave demultiplexing, and the receiving power is controlled at around -2dBm.

[0109] The baseband electrical signal was captured using a synchronous 8-channel oscilloscope with a sampling rate of 80 GSa / s and a bandwidth of 33 GHz. The receiver-side DSP process included dispersion compensation, resampling, and clock recovery in the frequency domain. A squared timing estimation algorithm was employed, preserving four times the signal rate.

[0110] Then, nonlinear equilibrium based on regression coupling value GRU neural network is performed. The decision algorithm adopts auxiliary minimum mean square algorithm. After the decision, 16QAM symbol demapping and final bit error rate calculation are performed.

[0111] To further illustrate, the present invention provides a nonlinear equalization method for GRU neural networks based on regression coupling values, which captures and compensates for nonlinear signal impairments through a trained network model;

[0112] Compared to existing learning equalizer schemes, this scheme uses gated recursive units (GRUs) to reduce model complexity and introduces a regression coupling parameter based on the nonlinear Schrödinger equation to optimize the network update calculation process. This parameter can effectively utilize the time-domain information of the sequence and maximize the algorithm's ability to learn nonlinear impairments in the channel.

[0113] In conclusion, this method has significant application prospects in the field of coherent optical communication.

[0114] The main design parameters in this design method include: attenuation α, second-order dispersion coefficient β², fiber nonlinearity coefficient γ, input sequence length K, number of GRU layers M, sliding window size N, batch size n, signal error b, and the input vector X at the current time step. t Output H at the current time t The current hidden state Z t Quality factor Q, Bit error rate BER.

[0115] First, the shape of the input sample is reshaped, the received QAM sample of length is normalized, and then divided into K-2N subsamples by a sliding window.

[0116] Next, a GRU-based neural network model is established. The reshaped signal is fed into the input layer, and then the gated recurrent units in the hidden layer perform calculations through activation functions. Finally, the output layer is fully connected and expanded.

[0117] During training, the network parameters are updated using a loss function and backpropagation to achieve neural network training.

[0118] During application, the output sequence obtained by the output layer is the result of nonlinear compensation of the original signal.

[0119] Finally, as a further improvement of this invention, this method decomposes the nonlinear Schrödinger equation using the split-step Fourier method to obtain the error factor Δb at each current moment within the fiber span, and adds the error factor b to the loss function and activation function. t Optimize the network update process.

[0120] In this invention, the signal is simplified to:

[0121] s(t)=Re[s(t)·e j2πfc (1);

[0122] Where s(t) is the baseband signal in complex form, Re[·] denotes taking the real part of the complex number, and e j2πfc t It is a carrier wave, f c It is the carrier frequency.

[0123] The main formulas for the nonlinear equilibrium method based on the regression coupling value GRU neural network are as follows.

[0124] First, the received QAM samples of length *k* are normalized and divided into *K-2N* subsamples using a sliding window with a width of 2N+1 and a stride of 1, where the window width represents the size of the input features. The new input matrix can be represented in the following form:

[0125]

[0126] Then, these subsamples are reshaped into a three-dimensional tensor [K-2N-M+1,M,2N+1], and each vector in the features and labels is decomposed into real and imaginary parts. The above data constitutes the input layer samples.

[0127] Next, define the network and create update and reset gates:

[0128] R t =σ(X) t W xr +H t-1 W hr +b r (3);

[0129] Z t =σ(X) t W xz +H t-1 W hz +b z (4);

[0130] Here, W represents the sigmoid function that converts the value to a range between 0 and 1. xr ∈R m ×n and W xr ∈R m ×n represents the weight matrix involved in the learning process, b r ∈R1×n represents the bias value.

[0131] Here, W represents the sigmoid function that converts the value to a range between 0 and 1. xz ∈R m ×n and W hz ∈R m ×n represents the weight matrix involved in the learning process, b z ∈R1×n represents the bias value.

[0132] The formula for calculating the candidate hidden state t at the current time step is as follows:

[0133]

[0134] Where σ represents the sigmoid function that converts the value to a range between 0 and 1, W xh ∈R m ×n and W hh ∈R m ×n represents the weight matrix involved in the learning process, b z ∈R1×n represents the bias value, and ⊙ represents the Hadamard product.

[0135] Finally, the candidate states are combined with the update gate using the following formula:

[0136]

[0137] Obtain the model's memory state or output at the current time t.

[0138] During the training phase, the input sequence enters the hidden layer, and all GRU units are fully connected to the hidden layer. The error factor is calculated first.

[0139] The error factor is calculated as follows:

[0140]

[0141] Next, using the calculation factor as the error component, the output at the current time is calculated, as shown in formula (8):

[0142] y t =W Softmax h t +Softmax(b t (8).

[0143] During training, activation functions are used to transform the input, generating the neuron's output (out). d The calculation formula is as follows:

[0144] The calculation formula is as follows:

[0145] out d =f ReLU (y t (9);

[0146] The network's predicted output and the actual labels are passed to the loss function to calculate the loss value. The loss function is estimated using the mean squared error of cross-validation.

[0147]

[0148] Finally, the difference between the network output and the actual signal is measured using a loss function. The derivative of each parameter is calculated, and the gradient of the loss with respect to the network parameters is calculated through the backpropagation algorithm. The parameters are then updated in the opposite direction of the gradient, thus achieving network training.

[0149] During testing and application, the current output shown in formula (8) is mapped to the original signal strength range, as shown in formula (11):

[0150] Y = y t σ+μ (11);

[0151] Where σ represents the standard deviation of the original data, and μ represents the mean of the original data.

[0152] Among them, formulas (1) to (2) represent the data recombination stage, formulas (3) to (7) represent the damage factor calculation stage, formulas (8) to (10) represent the network update stage, and formula (11) represents the network output in the actual application stage.

[0153] The solution of this invention can effectively extract and "learn" the channel characteristics caused by nonlinear impairment under specific conditions, thereby effectively compensating for the nonlinear impairment and crosstalk between signals of the original signal.

[0154] When the channel length changes, this method does not require retraining the entire model; retraining with only a small dataset is sufficient to adapt the network parameters to the new channel length. This improves the model's flexibility and adaptability, effectively reducing training costs and time. It has significant application prospects in high-speed, long-distance fiber optic transmission systems.

[0155] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art can make other variations or modifications based on the above description.

[0156] It is neither necessary nor possible to exhaustively list all possible implementations here. However, any obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A nonlinear equilibrium method for GRU neural networks based on regression coupling values, characterized in that, Includes the following steps: The process of reshaping input data, Adjust the data structure of the received signal sequence and matrix it; The received QAM samples of length K are normalized and divided into K-2N subsamples using a sliding window; The process of input data propagating forward through the gated loop unit is as follows: A network model is established, the reshaped signal is fed into the input layer, and then the gated recurrent units in the hidden layer perform calculations through activation functions. Finally, the output layer is fully connected and expanded. as well as The training process involves iterative steps, namely, using the loss function and backpropagation to guide the updating of network parameters during a training process to achieve neural network training; In this context, both the activation function and the loss function use the error factor as the offset term b. t The error factor is obtained by calculating the nonlinear damage error term based on the distributed Fourier method of the nonlinear Schrödinger equation. By decomposing the nonlinear Schrödinger equation using the step-by-step Fourier method, the error factor Δb at each current moment within the fiber span can be obtained. Finally, in the application stage, the received signal is input into the GRU neural network, and the output sequence of the output layer is obtained through the forward propagation algorithm of the network. This sequence is the nonlinear compensation result of the original signal. The signal is simplified as follows: Where s(t) is the baseband signal in complex form, Re[·] denotes taking the real part of the complex number, and e j2πfc t It is a carrier wave, f c It is the carrier frequency.

2. The nonlinear equalization method of GRU neural network based on regression coupling value as described in claim 1, characterized in that, The process of normalizing the received QAM samples of length K and dividing them into K-2N sub-samples using a sliding window is described by dividing them into K-2N sub-samples using a sliding window with a width of 2N+1 and a step size of 1. Here, the window width represents the size of the input features, so that the new input matrix can be represented in the following form: Then, each n rows is packaged and integrated into a batch. Thus, the subsamples are reshaped into three-dimensional tensors, and each vector in the features and labels is decomposed into real and imaginary parts, so that the data constitutes the input layer samples.

3. The nonlinear equalization method of GRU neural network based on regression coupling value as described in claim 2, characterized in that, During the training phase of the training process, the hidden layer guides the network parameter updates through an error factor; The channel environment modeling involves discretizing the propagation formula using the relationships given by the nonlinear Schrödinger equation. The formula takes the following form: Where α represents attenuation, β2 represents the second-order dispersion coefficient, and γ represents the fiber nonlinearity coefficient; Fourier transform is used to convert between the time and frequency domains, decomposing the nonlinear Schrödinger equation into E0... x and E y The differential formulas in both directions, which handle nonlinear effects in the frequency domain, and then transformed back to the time domain expression, yield the general form of the error components: The error factor is calculated as follows: Then, using the calculation factor as the error component, the output at the current time is calculated, as shown in formula (8): y t =W Softmax h t +Softmax(b t ) (8)。 4. The nonlinear equilibrium method of GRU neural network based on regression coupling value as described in claim 3, characterized in that, The process of using the gate structure unit to guide network parameter updates in conjunction with error factors also includes: The steps for defining the network and creating update and reset gates include applying formulas (3) and (4): R t =σ(X t W xr +H t-1 W hr +b r ) (3); Z t =σ(X t W xz +H t-1 W hz +b z ) (4); Where σ represents the sigmoid function that converts the value to a range of 0 to 1, and W represents the weight values, which are presented in matrix form. xr ∈R m ×n and W xr ∈R m ×n represents the weight matrix involved in the learning process, b r ∈R1×n and b z ∈R1×n represents the bias value; Based on the GRU network model structure, the gated recursive unit network model structure includes a reset gate and an update gate. Through the control of these two gates, GRU generates a candidate hidden state at each time step, which takes into account both the current input and the hidden state at the previous time step. By combining the current candidate state and the hidden state at the previous time step, GRU calculates the final hidden state, which is used as the output at the current time step. In step S4, the gate update rule guides the GRU neural network to update its parameters for the next time through iterative calculation; In the current time step t, suppose X t It is an m-dimensional vector representing the input; H t-1 It is an n-dimensional vector representing the hidden state at time t. Reset R t Determine whether the hidden state from the previous time step will be abandoned at time t, and update gate Z. t This will determine whether the state will be updated at time t; The formula for calculating the candidate hidden state t at the current time step is as follows: Where σ represents the sigmoid function that converts the value to a range between 0 and 1, W xh ∈R m ×n and W hh ∈R m ×n represents the weight matrix involved in the learning process, b z ∈R1×n represents the bias value, and ⊙ represents the Hadamard product; The candidate states are combined with the update gate using the following formula: Obtain the model's memory state or output at the current time t.

5. The nonlinear equalization method of GRU neural network based on regression coupling value as described in claim 4, characterized in that, During training, activation functions are used to transform the input, generating the neuron's output (out). d The calculation formula is as follows: out d =f ReLU (y t ) (9); The network's predicted output and the actual labels are passed to the loss function to calculate the loss value. The loss function is estimated using the mean squared error of cross-validation. Finally, the difference between the network output and the actual signal is measured using a loss function. The derivative of each parameter is calculated, and the gradient of the loss with respect to the network parameters is calculated through the backpropagation algorithm. The parameters are then updated in the opposite direction of the gradient, thus achieving network training.

6. The nonlinear equalization method of GRU neural network based on regression coupling value as described in claim 5, characterized in that, In the application phase, the current time output shown in formula (8) is mapped to the original signal strength range, as shown in formula (11): Y=y t s+m (11); Where σ represents the standard deviation of the original data, and μ represents the mean of the original data; This process allows for the completion of the balanced sequence output.

7. The nonlinear equalization method of GRU neural network based on regression coupling value as described in claim 6, characterized in that, The application conditions are: long-distance standard single-mode fiber 16-QAM transmission system.

8. The nonlinear equilibrium method for GRU neural networks based on regression coupling values ​​as described in any one of claims 1-7, characterized in that, It also includes an experimental model of coherent optical communication, the specific working process of which includes the transmitting end step and the receiving end step; The transmitting end steps are as follows: At the transmitting end, 1000 sets of pseudo-random binary sequences with a code length of 128 were first generated; Each 4 bits is mapped to a 16QAM symbol, and after being upsampled twice, the baseband is shaped by a root-raised cosine filter. The system uses 40 external cavity lasers with a frequency spacing of 100 GHz, one group for odd-numbered paths and one group for even-numbered paths, to output a total of 80 optical carriers with a spacing of 50 GHz, with a wavelength range of 1530-1562 nm. The optical power of each wavelength channel is 13dBm.

9. The nonlinear equilibrium method of GRU neural network based on regression coupling value as described in claim 8, characterized in that, Also includes: The digital baseband signal is resampled and applied to a waveform generator with a sampling rate of 64GSa / s; It is converted into two electrical signals to drive the IQ modulator, with a 3dB bandwidth of 29GHz; A loop structure with cyclic spans is used to achieve 1000km of standard single-mode fiber transmission, with each span of the fiber being a cyclic element. The optical fiber is divided into 50 uniform steps per span, and an EDFA is added after each span to compensate for transmission loss, with an output power of 23dBm. The wavelength selection switch is used to suppress power transfer and gain unevenness under high output power of EDFA. Its insertion loss is 5dB and the minimum adjustable pitch is 50GHz. The two output signals are re-multiplexed and enter the next cycle.

10. The nonlinear equilibrium method of GRU neural network based on regression coupling value as described in claim 9, characterized in that, The receiving end steps are as follows: The coherent optical receiver performs zero-difference detection on the optical signal selected after wave demultiplexing, and the receiving power is controlled at -2dBm; The baseband electrical signal was captured by a synchronous 8-channel oscilloscope with a sampling rate of 80 GSa / s and a bandwidth of 33 GHz; The receiver DSP process includes dispersion compensation, resampling, and clock recovery in the frequency domain; The quadratic timing estimation algorithm is used to preserve four times the signal rate; Nonlinear equalization based on regression-coupled value GRU neural network is performed. The decision algorithm adopts auxiliary least mean square algorithm. After decision, 16QAM symbol demapping and final bit error rate calculation are performed.