End-to-End Optimization Method and System Based on Differentiable Auxiliary Channels

By adopting an end-to-end optimization method based on a micro-assisted channel in the fiber communication system, using the neural network model and the MSE loss function, the problem of channel gradient and gradient disappearance is solved, and the capacity and efficiency improvement of the fiber communication system is achieved.

CN114117697BActive Publication Date: 2025-06-27SHANGHAI GUANGZHIYU INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111405475.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-24
Publication Date
2025-06-27
Estimated Expiration
2041-11-24

AI Technical Summary

Technical Problem

The optimization of end-to-end deep learning in fiber optic communication systems is limited by the dependence of backpropagation algorithms on channel gradients and the gradient vanishing problem caused by Sigmoid functions in the decoder.

Method used

The end-to-end optimization method based on the micro-assisted channel is adopted. By building the encoder, decoder and micro-assisted channel, the neural network model is used for training, the loss function is adjusted to MSE, and the Sigmoid function on the last layer of the decoder is removed to solve the gradient vanishing problem.

Benefits of technology

End-to-end optimization is achieved without the need for additional channel models, reducing gradient vanishing problems and improving the capacity and efficiency of fiber optic communication systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114117697B_ABST
    Figure CN114117697B_ABST
Patent Text Reader

Abstract

The present invention provides an end-to-end optimization method and system based on a differentiable auxiliary channel, including the following steps: Step S1: Perform end-to-end optimization based on the differentiable auxiliary channel; Step S2: Perform optical fiber transmission according to the optimized end-to-end. The differentiable auxiliary channel and its training method of the present invention solve the problem of blocked backpropagation in end-to-end optimization; adjust the loss function to MSE, the decoder does not require the Sigmoid function, reducing the problem of gradient disappearance; the method does not require a channel model, is applicable to experimental channels, the method is direct, and can be directly used for end-to-end optimization of long-distance transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optical fiber communication systems, and in particular, to an end-to-end optimization method and system based on a differentiable auxiliary channel. Background Art

[0002] In the development of optical fiber communication systems, the transceiver algorithms are all manually designed. However, this kind of block design system optimized by a greedy algorithm-like does not necessarily guarantee the optimal performance. On the other hand, the design theory support at the transmitting end is far less than that of the receiving end algorithm. This makes there still be a lot of room for optimization in the design of the transmitting end. End-to-end deep learning regards the entire communication system as an autoencoder, so that the algorithms of the transmitting and receiving ends can be designed with a global idea, and thus it is a means to explore the algorithm system with global optimality.

[0003] The Chinese patent literature with the publication number CN105490763A discloses an end-to-end broadband mobile MIMO propagation channel model and a modeling method, which are used for communication system optimization research and system performance evaluation. In the broadband mobile MIMO propagation system, the transmitting antennas at the transmitting end and the receiving antennas at the receiving end both adopt mobile polarization antenna arrays. The modeling method includes the following steps: transmitting the signal transmitted by the mobile transmitting end through a narrowband MIMO propagation channel to the mobile receiving end to obtain an end-to-end narrowband mobile MIMO propagation channel model without considering the physical characteristics of the antennas; obtaining the polarization responses of the transmitting antenna and the receiving antenna and the depolarization response of the MIMO polarization propagation channel, and obtaining the polarization factor of the polarization effect through the polarization responses of the transmitting antenna and the receiving antenna and the depolarization response of the MIMO polarization propagation channel; obtaining the coupling coefficient matrix of the transmitting antenna array and the receiving antenna array; combining the polarization factor and the coupling coefficient matrix of the transmitting and receiving antenna arrays to obtain an end-to-end broadband mobile MIMO propagation channel matrix.

[0004] Regarding the above related technologies, the inventor believes that end-to-end deep learning is limited by the backpropagation algorithm, which requires the gradients of the input and output of the channel, and these gradients are often unknown in the actual system. On the other hand, the decoder of the traditional end-to-end deep learning optimization method often uses the Sigmoid function in the last layer, and the loss function is the bit cross-entropy. The Sigmoid function causes the problem of gradient disappearance. Solving the channel gradient and the problem of training gradient disappearance to realize the optimization of the communication system of end-to-end deep learning has high practical significance for improving the capacity of optical fiber communication. Summary of the Invention

[0005] Aiming at the defects in the prior art, the purpose of the present invention is to provide an end-to-end optimization method and system based on a differentiable auxiliary channel.

[0006] An end-to-end optimization method based on a differentiable auxiliary channel provided by the present invention includes the following steps:

[0007] Step S1: Perform end-to-end optimization based on the differentiable auxiliary channel;

[0008] Step S2: Perform optical fiber transmission according to the optimized end-to-end.

[0009] Preferably, the step S1 includes the following steps:

[0010] Construction step: Construct an encoder, a decoder, and a differentiable auxiliary channel;

[0011] Dataset construction step: Construct an end-to-end training dataset based on end-to-end input data and end-to-end output data, and construct a channel dataset based on channel input data and channel output data;

[0012] Training step: Train the differentiable auxiliary channel based on the channel dataset;

[0013] Decoder optimization step: Optimize the decoder based on the end-to-end training dataset;

[0014] Encoder parameter gradient acquisition step: Obtain the gradient of the decoder input through backpropagation of the optimized decoder; Based on the trained differentiable auxiliary channel, the gradient of the decoder input is backpropagated to obtain the encoder parameter gradient;

[0015] Encoder parameter optimization step: Based on the encoder parameter gradient, the encoder performs encoder parameter optimization.

[0016] Preferably, in the construction step, the encoder, the decoder, and the differentiable auxiliary channel are built by a neural network model;

[0017] The encoder output data includes the processed signal and the optimization parameter, and the dimension of the encoder output data is consistent with the input dimensions of the differentiable auxiliary channel and the real channel;

[0018] The differentiable auxiliary channel is composed of the addition of multiple parts: one part represents the linear feature, and the other part represents the non-linear perturbation.

[0019] Preferably, in the dataset construction step, end-to-end propagation is performed to complete the construction of the dataset in batches, and the batch size is required to be greater than a predetermined value; The distance, the input fiber power, and the number of channels of the channels in different batches change;

[0020] Both the input data and the output data of the end-to-end training dataset are bits;

[0021] Both the input data and the output data of the channel dataset are power-normalized, and the normalization rule is:

[0022]

[0023] Among them, S represents the length of data normalization; x i represents the i-th symbol in the data; represents the normalized output of the i-th symbol in the data.

[0024] Preferably, in the training step, the MSE is used as the loss function according to the input and output symbols, and the differentiable auxiliary channel is optimized by using the backpropagation algorithm and gradient descent; the differentiable auxiliary channel is repeatedly trained more than the first predetermined number of times for each batch.

[0025] Preferably, in the decoder optimization step, the decoder is optimized by supervised learning using the MSE loss function based on the end-to-end training dataset; the loss function calculates the loss for each bit; the decoder is repeatedly trained more than the second predetermined number of times for each batch.

[0026] Preferably, the encoder parameter gradient acquisition step includes the following steps:

[0027] Decoder input gradient acquisition step: The decoder is required to obtain the gradient of the loss function with respect to the decoder input through the error backpropagation algorithm;

[0028] Differentiable auxiliary channel output acquisition step: The output data of the encoder is used as the input data of the differentiable auxiliary channel, and the input data of the differentiable input channel is input into the differentiable auxiliary channel to obtain the output data of the differentiable auxiliary channel;

[0029] Encoder loss function step: The output data of the differentiable auxiliary channel is multiplied by the gradient of the decoder input and summed to be used as the encoder loss function;

[0030] Transmitter end gradient step: The encoder parameter gradient is obtained by using the error backpropagation of the encoder loss function.

[0031] An end-to-end optimization system based on a differentiable auxiliary channel provided by the present invention includes the following modules:

[0032] Module M1: Perform end-to-end optimization based on the differentiable auxiliary channel;

[0033] Module M2: Perform optical fiber transmission according to the optimized end-to-end.

[0034] Preferably, the module M1 includes the following modules:

[0035] Construction module: Construct an encoder, a decoder, and a differentiable auxiliary channel;

[0036] Dataset construction module: Construct an end-to-end training dataset based on the end-to-end input data and end-to-end output data, and construct a channel dataset based on the channel input data and channel output data;

[0037] Training module: Training a differentiable auxiliary channel based on a channel data set;

[0038] Decoder optimization module: Optimizing the decoder based on an end-to-end training data set;

[0039] Encoder parameter gradient acquisition module: Obtaining the gradient of the decoder input through backpropagation of the optimized decoder; Based on the trained differentiable auxiliary channel, the gradient of the decoder input is backpropagated to obtain the encoder parameter gradient;

[0040] Encoder parameter optimization module: Based on the encoder parameter gradient, the encoder performs encoder parameter optimization.

[0041] Preferably, in the construction module, the encoder, the decoder, and the differentiable auxiliary channel are built by a neural network model;

[0042] The encoder output data includes the processed signal and the optimization parameters, and the dimension of the encoder output data is consistent with the input dimensions of the differentiable auxiliary channel and the real channel;

[0043] The differentiable auxiliary channel is composed of the addition of multiple parts: one part represents linear features, and the other part represents non-linear perturbations.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] 1. The differentiable auxiliary channel and its training method of the present invention solve the problem of blocked backpropagation in end-to-end optimization;

[0046] 2. The present invention adjusts the loss function to MSE, and the decoder does not require the Sigmoid function, reducing the problem of gradient disappearance;

[0047] 3. The method of the present invention does not require a channel model, is applicable to experimental channels, and the method is direct and can be directly used for end-to-end optimization of long-distance transmission. Description of the Drawings

[0048] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:

[0049] Figure 1 It is a flowchart of the end-to-end optimization method based on a differentiable auxiliary channel of the present invention;

[0050] Figure 2 It is a structural diagram of the differentiable auxiliary channel in the construction step of the end-to-end optimization method based on a differentiable auxiliary channel of the present invention;

[0051] Figure 3 It is a flowchart of the forward processing of the signal in the data set construction step and the training step of the end-to-end optimization method based on a differentiable auxiliary channel of the present invention;

[0052] Figure 4 This is the error backpropagation flowchart for obtaining the encoder parameter gradient in the end-to-end optimization method based on the differentiable auxiliary channel of the present invention. Detailed implementation manners

[0053] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0054] An embodiment of the present invention discloses an end-to-end optimization method based on a differentiable auxiliary channel, as Figure 1 shown, including the following steps: Step S1: Perform end-to-end optimization based on the differentiable auxiliary channel. Step S1 includes the following steps: Construction step: Construct an encoder, a decoder, and a differentiable auxiliary channel. Construct an encoder, a decoder, and a differentiable auxiliary channel based on a non-linear perturbation. The encoder, decoder, and differentiable auxiliary channel are built by a neural network model. The end-to-end input is the input of the encoder, and the end-to-end output is the output of the decoder. The end-to-end includes an encoder, a differentiable auxiliary channel, and a decoder. The encoder input is a bit input, while the decoder input is the output of the channel. The input of the differentiable auxiliary channel is the output of the encoder. The common feature of the three is that the channel parameters (including distance, power, and the number of channels) are additionally input. The bit input and the channel parameter input are the end-to-end input data. The bit output is also the end-to-end output. Therefore, the bits and channel parameters here are used to train the model. The encoder, decoder, and differentiable auxiliary channel all contain channel parameters, so that the encoder, decoder, and differentiable auxiliary channel can adapt to various channel conditions. Both the encoder and the decoder are composed of fully connected neural networks, and the Sigmoid function is not used in the last layer of the decoder.

[0055] The encoder output data includes the processed signal and the optimization parameters. The dimension of the encoder output data is consistent with the input dimensions of the differentiable auxiliary channel and the real channel. The differentiable auxiliary channel is composed of multiple parts added together: one part represents the linear feature, and the other part represents the non-linear perturbation. The encoder output is designed as the processed signal and the optimization parameters, and the signal output by the encoder is consistent with the inputs of the differentiable auxiliary channel and the real channel. The differentiable auxiliary channel does not need to model the noise and is required to be composed of two parts added together: one part represents the linear feature, the overall rotation of dispersion and phase, which is realized by matrix multiplication; the other part represents the non-linear perturbation, which is realized by a neural network.

[0056] As Figure 2As shown, the differentiable auxiliary channel does not need to model noise. The specific steps of the differentiable auxiliary channel modeling include: Linear feature characterization step: Represent linear features according to the input signal, including dispersion and the overall rotation of the non-linear phase, which is achieved by matrix multiplication. Noise perturbation characterization step: Use a neural network to represent noise perturbation according to the input signal, including non-linear perturbation noise, residual linear features, etc., which is achieved by a neural network architecture. Summation step: Sum the linear features and non-linear perturbations. Normalization step: Output after batch normalization of the power, and the formula is

[0057]

[0058] where S represents the length of data normalization, usually the total size of the batch; x i and represent the i-th symbol in the data and the normalized output of the i-th symbol. The average absolute value of the normalized signal remains around 1. Here, the output of the differentiable auxiliary channel needs to be normalized.

[0059] The data output by the encoder is the data obtained after being processed by the encoder. Specifically, the data output here can be flexibly set as the processed signal and the parameters that need to be optimized at the transmitting end. For example, if it is set that the encoder needs to perform optimal filtering on the signal and set the optimal transmission power at the same time, then the encoder output is the filtered signal and the magnitude of the transmission power. The focus of the present invention is on the end-to-end optimization of using optimized transceiver algorithms. The signal output by the encoder can be used as part of the signal processing at the transmitting end, and the optimized parameters can be used to design the optimal transmitting end structure, thereby improving the communication quality. The encoder output is the optimization target, and the present invention is an effective means to achieve the target. In order to achieve differentiation, the dimension of the output here needs to be consistent with the inputs of the differentiable auxiliary channel and the real channel. Processing represents the function that the encoder is set to process, and the processed signal represents the output of the signal after passing through the encoder. For example, if the encoder is set for filtering, then the output of the signal after passing through the encoder is the filtered signal. The optimized parameters represent the parameters that the encoder is set to optimize. For example, the modulation format and transmission power of the transmitting end signal, etc. After the encoder is trained, the best modulation format and the optimal transmission power are given to achieve the improvement of communication quality. According to the perturbation theory, any non-linear function can be represented as linear plus non-linear perturbation. Here, the perturbation theory is exactly used to set the neural network structure as the addition of linear and non-linear perturbations to achieve more accurate modeling of the channel.

[0060] The end-to-end concept refers to the communication from the transmitting end to the receiving end. The end-to-end optimization algorithm of the present invention can optimize both the transmitting and receiving ends simultaneously, and design the communication transmitting and receiving algorithms from a global perspective, so it is called end-to-end optimization. Starting from the actual situation of optical fiber communication, by using the algorithm to optimize both the transmitting and receiving ends of communication, the rate can be improved. The transmitting and receiving ends include many processes and algorithms. The algorithm of the present invention can deploy a network at each of the transmitting and receiving ends. The network deployed at the transmitting end is an encoder, and the one deployed at the receiving end is a decoder. Encoders and decoders are generally designed to solve certain optimization problems at the transmitting and receiving ends. For example, the transmitting end encoder can implement the best modulation format, filter parameters, etc. in communication; the receiving end decoder can implement the best signal recovery and demodulation, etc. Narrowly speaking, the transmitting and receiving ends and the encoder and decoder have a relationship of inclusion and being included. The transmitting end can also be completely implemented by the encoder, and the receiving end can be completely implemented by the decoder.

[0061] Steps for constructing the dataset: Construct an end-to-end training dataset based on the end-to-end input data and end-to-end output data, and construct a channel dataset based on the channel input data and channel output data. Input the original bit stream into the encoder, channel, and decoder to obtain the output. Here, the output is the end-to-end output data. Specifically, it is the output of the decoder, that is, the output after passing through the encoder, channel, and decoder. The end-to-end input and output construct the end-to-end training dataset, and the channel input and output construct the channel dataset. The end-to-end input and output construct the end-to-end training dataset, and the goal is to restore the output to the input. The end-to-end input data includes the original bit stream and may also selectively include channel parameters. The channel is established between the encoder and the decoder. Here, the channel refers to a relatively broad sense of the channel, which is the middle part between the encoder and the decoder. Generally speaking, the output of the encoder can continue to perform some signal processing at the transmitting end, and then pass through the narrow sense channel (that is, the optical fiber). At the receiving end, it is first processed and then sent to the decoder. However, by packing all the intermediate parts between the decoder and the encoder, it can also be regarded as a broad sense channel. Here, the broad sense channel is the part for differentiable auxiliary channel modeling. The original bit stream is obtained at the information source at the transmitting end of optical fiber communication, usually customer information. The communication goal is to perfectly transmit the original bit stream of the information source to the receiving end. For example, in a real scenario, the information bits sent by a computer for Internet access somewhere.

[0062] Implement the forward transmission for end-to-end optimization. The specific process is as Figure 3As shown, the original bitstream is input into the encoder, transmitter signal processing, channel, receiver signal processing, and decoder to obtain the output. The end-to-end input and output construct the end-to-end training dataset. The end-to-end input and output data are the original bitstream, which is the information to be transmitted in communication and is provided by the information source, usually the customer information at a certain end. The channel input and channel output construct the channel dataset. The output of the original bitstream after passing through the encoder is used as the input data of the channel, and the data after passing through the channel is used as the channel output data. Propagate end-to-end to complete the construction of the dataset in batches, and the number of batches is required to be greater than a predetermined value; the distance, input fiber power, and number of channels of the channel vary for different batches of data. The input data and output data of the end-to-end training dataset are both bits. Each epoch, the dataset is constructed in batches, and the number of batches is required to be greater than 1000; for different batches of data, the distance, input fiber power, and number of channels of the channel vary. The input data and output data of the channel dataset are both power-normalized. The channel dataset requires both input and output to be power-normalized, and the normalization rule is:

[0063]

[0064] where S represents the length of data normalization, usually the total size of the batch; x i represents the i-th symbol in the data; represents the normalized output of the i-th symbol in the data. x i and represent the i-th symbol and the normalized output of the i-th symbol in the data. The average absolute value of the normalized signal remains near 1. A set of datasets contains multiple signals. The entire dataset is arranged as a vector and represented by x. Each element of the vector represents one signal. i represents the i-th signal x i in the dataset vector. Here, the dataset of the differentiable auxiliary channel needs to be normalized. The goal is for the differentiable auxiliary channel input to be consistent with the dataset, so the same formula as above is used.

[0065] The end-to-end system performs forward propagation to complete the construction of the dataset in batches, and the number of batches is required to be greater than 1000. The end-to-end training dataset requires both input and output to be bits. It is required that the channel parameters change during the training process (including the transmission length, number of channels, and transmit power of optical fiber communication, etc.) to avoid only training to the result under one channel condition. Random channel parameters mean that during each training process, the above channel conditions are randomly changed, and corresponding channel inputs and outputs will be generated after the change. This pair constructs the training data of the differentiable auxiliary channel, and the channel output is used as the input of the decoder.

[0066] Training steps: Train the differentiable auxiliary channel based on the channel dataset. Train the differentiable auxiliary channel through supervised learning based on the channel dataset. Use the MSE as the loss function according to the input and output symbols, and optimize the differentiable auxiliary channel using the backpropagation algorithm and gradient descent; repeat the training of the differentiable auxiliary channel more than a first predetermined number of times for each batch. The first predetermined number of times includes 5 times. It is required to use the MSE as the loss function according to the input and output symbols, and optimize the differentiable auxiliary channel using the backpropagation algorithm and gradient descent method; repeat the training of the differentiable auxiliary channel more than 5 times for each batch. The end-to-end dataset consists of inputs and outputs, and the goal is to restore the output to the input to complete the decoding of communication. During the training process, a dataset is constructed from the inputs and outputs, the MSE loss is calculated for the output and its corresponding input, and the network model is trained through supervised learning. The forward transmission is included in the dataset construction step and the training step.

[0067] Decoder optimization steps: Optimize the decoder based on the end-to-end training dataset. Optimize the decoder through supervised learning using the MSE loss function based on the end-to-end training dataset; the loss function calculates the loss for each bit; repeat the training of the decoder more than a second predetermined number of times for each batch. The second predetermined number of times includes 5 times. The loss function is required to use the MSE function and calculate the loss for each bit; repeat the training of the decoder more than 5 times for each batch.

[0068] Steps for obtaining the gradient of the encoder parameters: Obtain the gradient of the decoder input through backpropagation of the optimized decoder; based on the trained differentiable auxiliary channel, backpropagate the gradient of the decoder input to obtain the gradient of the encoder parameters. Obtain the gradient of the decoder input through error backpropagation of the decoder error, and based on the differentiable auxiliary channel, backpropagate the gradient of the decoder input to obtain the gradient at the transmitting end (the parameter gradient of the transmitting-end encoder). Obtain the gradient of the decoder input through error backpropagation by the decoder receiving the loss; based on the differentiable auxiliary channel, backpropagate the gradient of the decoder input to obtain the gradient at the transmitting end, and based on the gradient at the transmitting end, the encoder realizes the optimization at the transmitting end.

[0069] As Figure 4 shown, the specific steps of the backpropagation algorithm based on the differentiable auxiliary channel, that is, the steps for obtaining the gradient of the encoder parameters include the following steps: Steps for obtaining the gradient of the decoder input: It is required that the decoder obtains the gradient of the loss function with respect to the decoder input through the error backpropagation algorithm. Its size should be (B, D), where B is the batch size required to be greater than 1000, and D is the dimension of the decoder input, which is consistent with the output dimensions of the channel and the differentiable auxiliary channel.

[0070] Steps for obtaining the output of the differentiable auxiliary channel: The output data of the encoder is used as the input data of the differentiable auxiliary channel. The input data of the differentiable input channel is input into the differentiable auxiliary channel to obtain the output data of the differentiable auxiliary channel. The output of the encoder is used as the input and input into the differentiable auxiliary channel to obtain the output of the differentiable auxiliary channel, and its size should be (B, D).

[0071] Steps for the encoder loss function: The output data of the differentiable auxiliary channel is multiplied element-wise with the gradient of the decoder input and then summed to obtain the encoder loss function. The output of the differentiable auxiliary channel is multiplied element-wise (i.e., corresponding elements are multiplied) with the gradient of the decoder input and then summed to obtain the encoder loss function.

[0072] Steps for the transmitter gradient: The encoder parameter gradient is obtained by backpropagating the error of the encoder loss function. The encoder parameter gradient (the gradient of the transmitter encoder parameters) is obtained by backpropagating the error.

[0073] Steps for optimizing the encoder parameters: Based on the encoder parameter gradient, the encoder performs optimization of the encoder parameters. Based on the transmitter gradient, the encoder achieves transmitter optimization; according to the number of training iterations, the learning rate parameter is optimized. The learning rate starts from the initial value, uses the cosine annealing algorithm, reduces the learning rate to 0, then restarts hot, and follows the cosine annealing from the initial value. The change of the learning rate repeats several cycles. The learning rates of the encoder and decoder decrease gradually to 0 in the form of cosine, and after dropping to 0, they resume to the initial learning rate and continue to decrease to 0 according to the cosine rule, cycling back and forth several times until the specified number of training cycles.

[0074] Steps for achieving optimization: The encoder is optimized using gradient descent through the parameter gradient; the current batch training is ended, and when training for M rounds, the current epoch ends; before the next epoch, every 10 epochs, it is required that the learning rates of the encoder and decoder decrease according to the cosine rule with built-in hot restart, and it is required that the learning rate of the differentiable auxiliary channel decreases according to the cosine rule; when training for N rounds in an epoch, the end-to-end and differentiable auxiliary channel are optimized, that is, the encoder, decoder, and differentiable auxiliary channel are optimized. Every 10 epochs, it is required that the learning rates of the encoder and decoder decrease according to the cosine rule with built-in hot restart, and it is required that the learning rate of the differentiable auxiliary channel decreases according to the cosine rule. An epoch represents a training cycle, which is the number of times to completely train the dataset once. A new epoch represents using the same dataset to retrain a cycle. The present invention has three neural networks, namely an encoder, a decoder, and a differentiable auxiliary channel. The learning rate of the differentiable auxiliary channel uses cosine descent without hot restart, while the encoder and decoder require hot restart.

[0075] The training of the present invention is divided into several loops: First is the outermost loop, which represents that all data are trained for N epochs in total, and i represents the i-th epoch; Secondly, in each epoch, the data are divided into M batches, and one training is performed for each batch, and l represents the l-th batch trained in that epoch; Secondly, during each training process, the decoder and the differentiable auxiliary channel need to be trained well first to ensure accurate gradients. Therefore, during each training, the decoder and the differentiable auxiliary channel are each trained 5 times. j represents the j-th training of the decoder, and k represents the k-th training of the differentiable auxiliary channel. % represents taking the remainder. For example, 40 % 10 = 0, 41 % 10 = 1. i % 10 can represent every 10 times.

[0076] Step S2: Perform optical fiber transmission according to the optimized end-to-end. Design the optimal optical fiber communication algorithm according to the optimized end-to-end for high-speed long-distance optical fiber transmission. According to the optimized end-to-end, design the optical communication transceiver algorithm at the same time, and assist in the parameter design of the DSP from the perspective of global optimization, such as signal coding, constellation shaping, etc. The full English name of DSP is Digital Signal Processing, and the Chinese translation is Digital Signal Processing.

[0077] The present invention can set the signal modulation format at the transmitting end as the object of encoder optimization, construct an end-to-end deep learning optical communication system, design the optimal geometric shaping and coding for the encoder according to the optical channel conditions, and perform the optimal demodulation and decoding for the decoder to maximize the communication mutual information and approach the channel capacity of the communication. End-to-end learning is applied to the design and decoding of the optimal modulation format in optical communication. Improve the mutual information and approach the channel capacity through end-to-end training.

[0078] The embodiment of the present invention also discloses an end-to-end optimization system based on a differentiable auxiliary channel, including the following modules: Module M1: Perform end-to-end optimization based on the differentiable auxiliary channel. Module M1 includes the following modules: Construction module: Construct an encoder, a decoder, and a differentiable auxiliary channel. The encoder, decoder, and differentiable auxiliary channel are built by a neural network model. The data input to the encoder, decoder, and differentiable auxiliary channel include bits and channel parameters. The data output by the encoder includes the processed signal and optimization parameters. The dimension of the data output by the encoder is consistent with the input dimensions of the differentiable auxiliary channel and the real channel. The differentiable auxiliary channel is composed of the addition of multiple parts: one part represents the linear feature, and the other part represents the non-linear perturbation.

[0079] Dataset construction module: Construct an end-to-end training dataset based on end-to-end input data and end-to-end output data, and construct a channel dataset based on channel input data and channel output data. Training module: Train a differentiable auxiliary channel based on the channel dataset. Decoder optimization module: Optimize the decoder based on the end-to-end training dataset. Encoder parameter gradient acquisition module: Obtain the gradient of the decoder input through backpropagation of the optimized decoder; based on the trained differentiable auxiliary channel, the gradient of the decoder input is backpropagated to obtain the encoder parameter gradient. Encoder parameter optimization module, based on the encoder parameter gradient, the encoder performs encoder parameter optimization.

[0080] Module M2: Perform optical fiber transmission according to the optimized end-to-end.

[0081] The technical problem to be solved by the present invention is that the optimization method of end-to-end deep learning needs to rely on the gradient of the known channel or the channel analytical model and the problem of gradient disappearance of the decoder. The end-to-end optimization method based on the differentiable auxiliary channel of the present invention does not require an additional channel model, only needs to obtain the channel input and output data, and realizes a differentiable channel model through a data-driven method; at the same time, a new loss function is used to remove the sigmoid function of the last layer of the decoder to solve the problem of gradient disappearance. The encoder, decoder and differentiable auxiliary channel of the present invention are constructed by neural networks, the loss function is set as the MSE loss, the last layer of the decoder does not require the sigmoid function, the differentiable auxiliary channel and the decoder are trained through supervised learning, the gradient of the decoder input is obtained by backpropagation, and the loss function is constructed with the output of the differentiable auxiliary channel, and the encoder parameter gradient is obtained again using the backpropagation algorithm, so as to realize the optimization of the encoder. This method reduces the problem of gradient disappearance in end-to-end training, does not require knowledge of the channel model, realizes end-to-end communication optimization through a data-driven method, is simple and direct, has high computational efficiency, is applicable to the design and optimization of optical communication systems, and improves the optical fiber transmission capacity. The full English name of MSE is Mean Square Error, and the Chinese translation is mean square error. The sigmoid function represents an S-shaped growth curve.

[0082] Those skilled in the art know that in addition to implementing the system and its various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to make the system and its various devices, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. to achieve the same function. Therefore, the system and its various devices, modules, and units provided by the present invention can be regarded as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structure within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as both software modules for implementing the method and the structure within the hardware component.

[0083] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be arbitrarily combined with each other.

Claims

1. An end-to-end optimization method based on a differentiable auxiliary channel, characterized in that, It includes the following steps: Step S1: Perform end-to-end optimization based on a differentiable auxiliary channel; Step S2: Perform optical fiber transmission according to the optimized end-to-end; The said Step S1 includes the following steps: Construction step: Construct an encoder, a decoder, and a differentiable auxiliary channel; Dataset construction step: Construct an end-to-end training dataset based on end-to-end input data and end-to-end output data, and construct a channel dataset based on channel input data and channel output data; Training step: Train the differentiable auxiliary channel based on the channel dataset; Decoder optimization step: Optimize the decoder based on the end-to-end training dataset; Encoder parameter gradient acquisition step: Obtain the gradient of the decoder input through backpropagation of the optimized decoder; Based on the trained differentiable auxiliary channel, backpropagate the gradient of the decoder input to obtain the encoder parameter gradient; Encoder parameter optimization step: Based on the encoder parameter gradient, the encoder performs encoder parameter optimization; In the construction step, the encoder, the decoder, and the differentiable auxiliary channel are built by a neural network model; The encoder output data includes the processed signal and the optimization parameter, and the dimension of the encoder output data is consistent with the input dimensions of the differentiable auxiliary channel and the real channel; The differentiable auxiliary channel is composed of the addition of multiple parts: one part represents the linear feature, and the other part represents the non-linear perturbation; In the said training step, use MSE as the loss function according to the input and output symbols, and use the backpropagation algorithm and gradient descent to optimize the differentiable auxiliary channel; Repeat training the differentiable auxiliary channel more than the first predetermined number of times for each batch; In the said decoder optimization step, optimize the decoder through supervised learning using the MSE loss function based on the end-to-end training dataset; The loss function calculates the loss for each bit; repeat training the decoder more than the second predetermined number of times for each batch; The said encoder parameter gradient acquisition step includes the following steps: Decoder input gradient acquisition step: Require the decoder to obtain the gradient of the loss function with respect to the decoder input through the error backpropagation algorithm; Differentiable auxiliary channel output acquisition step: The output data of the encoder is used as the input data of the differentiable auxiliary channel, and the input data of the differentiable input channel is input into the differentiable auxiliary channel to obtain the output data of the differentiable auxiliary channel; Encoder loss function step: Perform dot product summation on the output data of the differentiable auxiliary channel and the gradient of the decoder input as the encoder loss function; Transmitter end gradient step: Use the error backpropagation of the encoder loss function to obtain the encoder parameter gradient.

2. The end-to-end optimization method based on a differentiable auxiliary channel according to claim 1, wherein In the said dataset construction step, perform end-to-end propagation to complete the construction of the dataset in batches, and require the batch to be greater than the predetermined value; the distance, input optical power, and number of channels of the channel vary for different batches of data; Both the input data and the output data of the end-to-end training dataset are bits; Both the input data and the output data of the channel dataset are power-normalized, and the normalization rule is: Among them, S represents the length of data normalization; x i represents the i-th symbol in the data; represents the normalized output of the i-th symbol in the data.

3. An end-to-end optimization system based on a differentiable auxiliary channel, characterized in that, It includes the following modules: Module M1: Perform end-to-end optimization based on a differentiable auxiliary channel; Module M2: Perform optical fiber transmission according to the optimized end-to-end; The said Module M1 includes the following modules: Construction module: Construct an encoder, a decoder, and a differentiable auxiliary channel; Data set construction module: Construct an end-to-end training data set based on end-to-end input data and end-to-end output data, and construct a channel data set based on channel input data and channel output data; Training module: Train a differentiable auxiliary channel based on the channel data set; Decoder optimization module: Optimize the decoder based on the end-to-end training data set; Encoder parameter gradient acquisition module: Obtain the gradient of the decoder input through backpropagation of the optimized decoder; Based on the trained differentiable auxiliary channel, the gradient of the decoder input is backpropagated to obtain the encoder parameter gradient; Encoder parameter optimization module: Based on the encoder parameter gradient, the encoder performs encoder parameter optimization; In the construction module, the encoder, decoder, and differentiable auxiliary channel are built by a neural network model; The encoder output data includes the processed signal and optimization parameters, and the dimension of the encoder output data is consistent with the input dimensions of the differentiable auxiliary channel and the real channel; The differentiable auxiliary channel is composed of the sum of multiple parts: one part represents linear features, and the other part represents non-linear perturbations; In the training module, use MSE as the loss function according to the input and output symbols, and use the backpropagation algorithm and gradient descent to optimize the differentiable auxiliary channel; Repeat training the differentiable auxiliary channel more than the first predetermined number of times for each batch; In the decoder optimization module, optimize the decoder through supervised learning using the MSE loss function based on the end-to-end training data set; The loss function calculates the loss for each bit; Repeat training the decoder more than the second predetermined number of times for each batch; The encoder parameter gradient acquisition module includes the following modules: Decoder input gradient acquisition module: Require the decoder to obtain the gradient of the loss function with respect to the decoder input through the error backpropagation algorithm; Differentiable auxiliary channel output acquisition module: The output data of the encoder is used as the input data of the differentiable auxiliary channel, and the input data of the differentiable input channel is input into the differentiable auxiliary channel to obtain the output data of the differentiable auxiliary channel; Encoder loss function module: The output data of the differentiable auxiliary channel and the gradient of the decoder input are multiplied and summed as the encoder loss function; Transmitter end gradient module: Use the error backpropagation of the encoder loss function to obtain the encoder parameter gradient.

Citation Information

Patent Citations

  • End-to-end broadband mobile MIMO (multiple input multiple output) propagating channel model and modeling method

    CN105490763A

  • Deep learning-based joint optimization method of wireless communication physical layer receiving and sending end

    CN111327381A