An artificial intelligence-based optical fiber terahertz communication system optimization method
By employing an AI-based end-to-end optimization method and utilizing independent sideband modulation and neural network training, a globally optimal four-dimensional joint geometric probability constellation shaping scheme is designed. This solves the reliability and architectural complexity issues of fiber optic terahertz communication systems under complex channels, achieving higher communication performance.
Patent Information
- Application Number
- CN202411637977.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-11-15
AI Technical Summary
Existing fiber optic terahertz communication systems struggle to achieve four-dimensional geometric shaping under complex channel conditions, resulting in high system architecture complexity. Furthermore, the independent sideband modulation signals introduce frequency-selective fading due to fiber dispersion, impacting communication reliability.
An AI-based end-to-end optimization method is adopted, which utilizes the additional dimension provided by two independent sidebands. The equalizer and constellation shaper are trained through a pruning neural network and a hybrid density network. The globally optimal four-dimensional joint geometric probability constellation shaping scheme is adaptively designed, and the end-to-end gradient direction propagation is realized by considering the end-to-end damage mechanism.
It reduces the complexity of neural networks, approaches the Shannon limit, improves the reliability and effectiveness of communication systems, simplifies channel modeling, and enhances adaptability to complex channels.
Smart Images

Figure CN119602870B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, and in particular to an optimization method for an artificial intelligence-based fiber optic terahertz communication system. Background Technology
[0002] Fiber optic terahertz communication systems achieve seamless integration of fiber optic networks and terahertz wireless links, offering greater spectrum resources and lower latency. In recent years, artificial intelligence technology has been increasingly applied to other disciplines during its development, achieving performance advantages surpassing traditional algorithms. Simultaneously, the large bandwidth of fiber optic terahertz systems enables the rapid generation of massive amounts of data, making them highly suitable for various data-driven methods. A key reason for the special significance of deep learning methods in the field of communications is their ability to collaboratively design all modules of the entire communication system end-to-end, thereby approximating the global optimum. This differs from traditional communication system design philosophies, which typically consider optimizing different communication modules with different metrics to achieve multiple local optima.
[0003] In recent years, several two-dimensional constellation shaping schemes have been adopted in wireless and fiber optic communication systems. For example, by using an autoencoder with mutual information as the loss function, end-to-end joint geometric probability shaping of additive white Gaussian noise (AWGN) and fading channels has been achieved [Stark, Maximilian, ...]. AitAoudia, and Jakob Hoydis. "Joint learning of geometric and probabilistic constellation shaping." 2019 IEEE Globecom Workshops (GC Wkshps). IEEE, 2019. Further considering bit tags and using generalized mutual information as the loss function, end-to-end joint geometric probabilistic constellation shaping for arbitrary channel models is achieved by uniformly sampling from constellation points to avoid sampling from probabilities. [Aoudia, Ait, and Jakob Hoydis. "Joint learning of probabilistic and geometric shaping for coded modulation systems." GLOBECOM 2020-2020 IEEE Global Communications Conference. IEEE, 2020. However, these methods require prior knowledge of the channel conditional probability distribution and are not suitable for more complex channels, such as fiber optic terahertz channels.
[0004] Compared to two-dimensional geometric shaping, four-dimensional geometric shaping, by sacrificing a small amount of spectral efficiency, can achieve a larger minimum Euclidean distance between constellation points. Furthermore, because four-dimensional space provides a larger search range, constellation diagrams with better nonlinear tolerance can be designed, thus achieving higher coding gain [Chen, Bin, et al. Geometrically-shaped multi-dimensional modulation formats in coherent optical transmission systems. Journal of Lightwave Technology 41.3(2022):897-910.]. However, to achieve four-dimensional geometric shaping, existing optical fiber communication systems typically use two polarizations of light as two additional dimensions, requiring optical polarization multiplexing and polarization diversity, which significantly increases the complexity of the system architecture. In addition, double-sideband signals generated by intensity modulation based on a single-arm Mach-Zehnder modulator will introduce frequency-selective fading under the influence of fiber dispersion. Independent sideband modulation can solve this problem well. Furthermore, independent sideband modulation generates two different sidebands with independent information around the same carrier, which can naturally provide two additional dimensions for the four-dimensional constellation shaping scheme, making it very suitable for constructing four-dimensional constellation shaping. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and provide an artificial intelligence-based optimization method for fiber optic terahertz communication systems. This method utilizes the additional dimensions provided by two independent sidebands, adopts an end-to-end optimization framework, and considers the impact of end-to-end impairment mechanisms (including the effects of fiber optic link loss, dispersion, nonlinear effects, and atmospheric loss and multipath interference of terahertz links on the transmitted signal). It adaptively designs a globally optimal four-dimensional joint geometric probability constellation shaping scheme to further approach the Shannon limit and achieve higher reliability and effectiveness of the communication system.
[0006] To achieve the above-mentioned objectives, this invention provides an artificial intelligence-based optimization method for fiber optic terahertz communication systems, comprising a training phase and a deployment phase, specifically including the following steps:
[0007] Training phase steps:
[0008] Step a1: The transmitter randomly generates a sequence of transmitted symbols, which is then transmitted through a full-link fiber optic terahertz channel to obtain a sequence of received symbols at the receiver.
[0009] An equalizer based on a pruned neural network is trained using transmitted and received symbol sequences as training data.
[0010] Step a2: Train an equalized post-channel model based on a mixture density network (MDN) based on the transmitted symbol sequence and the equalized symbol sequence. This model is then used to connect the constellation shaper (i.e., the four-dimensional joint geometric probability constellation shaper) at the transmitter and the demapper at the receiver to achieve end-to-end gradient direction propagation by taking into account the end-to-end impairment mechanism.
[0011] Among them, the equalized channel model is used to output the conditional probability distribution of the equalized channel. The composition parameters are used to achieve the following: The fit;
[0012] The equalized symbol sequence is obtained by inputting the emitted symbol sequence into a trained equalizer based on a pruned neural network and obtaining the equalized symbol sequence based on its output.
[0013] Step a3: Simultaneously train the constellation shaper at the transmitting end and the demapping mechanism at the receiving end, wherein both the constellation shaper and the demapping mechanism are neural network models.
[0014] The constellation shaper includes a geometric shaper and a probability shaper, which are used to generate a four-dimensional geometric and probabilistic constellation table for shaping the bit sequence at the transmitting end to obtain a four-dimensional frequency domain transmitted symbol sequence x;
[0015] The four-dimensional frequency domain transmitted symbol sequence x is input into the trained equalized channel model to obtain the four-dimensional frequency domain received symbol sequence. Then receive the four-dimensional frequency domain symbol sequence. Input the demapper to obtain the corresponding bit sequence estimate.
[0016] Deployment phase steps:
[0017] Step b1: The trained constellation shaper generates a constellation table, which is used to shape the information bit sequence at the transmitting end to obtain a four-dimensional frequency domain transmitted symbol sequence x.
[0018] Step b2: The four-dimensional frequency domain transmitted symbol sequence x is transmitted through the same full-link fiber terahertz channel as in step a1. The corresponding received symbol sequence is obtained at the receiving end. The received symbol sequence is then sent to a trained equalizer based on a pruned neural network to obtain the equalized received symbol sequence.
[0019] Step b3: After equalization, the received symbol sequence is demapped to obtain the estimated received information bit sequence.
[0020] The implementation process of the artificial intelligence-based optimization method for fiber optic terahertz communication systems provided by this invention is as follows: A four-dimensional constellation shaping is constructed using the additional dimensions provided by two independent sidebands. This end-to-end optimization framework is divided into a training phase and a deployment phase. In the training phase, an equalizer based on a pruned neural network and an equalized post-channel model based on a hybrid density network are first trained by transmitting and receiving random training sequences. Then, the constellation shaper at the transmitting end and the demapper at the receiving end are connected through the equalized post-channel model with fixed parameters (trained). This takes into account the residual distortion and noise of the entire equalized link, and backpropagates the gradient to learn the globally optimal four-dimensional joint geometric probability constellation diagram, further approximating the Shannon limit. In the deployment phase, the transmission of the bit stream is achieved through the four-dimensional geometrically shaped constellation table and probability distribution learned in the training phase, as well as the coordination of the equalizer and demapper at the receiving end.
[0021] Furthermore, in step a1, the transmitting end randomly generates a transmitted symbol sequence, which, after passing through a full-link fiber optic terahertz channel, results in a received symbol sequence at the receiving end, including:
[0022] (a1.1) The transmitter generates a random symbol sequence that follows a uniform probability distribution. After serial-to-parallel conversion, two independent data streams are obtained. These streams are then modulated using Orthogonal Frequency Division Multiplexing (OFDM), pulse-shaped, and up-converted to a frequency f. s This yields two independent double-sideband OFDM signals;
[0023] (a1.2) The two double-sideband signals provide drive inputs for two Mach-Zehnder modulators (MZMs) operating in push-pull mode, to modulate a signal with an operating frequency of f. c The continuous wavelength optical carrier generated by the single-mode laser is used to obtain two double-sideband optical signals. The left and right sidebands are extracted by two optical filters respectively, and the independent sideband (ISB) optical signal is obtained by passing through an optical power combiner.
[0024] (a1.3) The generated ISB optical signal is transmitted through a single-mode optical fiber. At the receiving end of the optical fiber link, it is amplified by an erbium-doped fiber amplifier and then connected to a signal with a working frequency of f. c+t The continuous wavelength optical carrier generated by the single-mode laser is combined. Based on the principle of optical heterodyne beat frequency, the combined signal is converted into a center frequency of f by a single-carrier photodiode (UTC-PD). t ISB terahertz signal;
[0025] (a1.4) The generated ISB terahertz signal is radiated into free space by the terahertz antenna. At the receiving end of the terahertz link, the local radio frequency source is mixed with the received terahertz signal after N frequency multiplications to convert the ISB terahertz signal into an ISB baseband signal. Then, two equal-power ISB baseband signals are obtained through a power divider, and the independent left and right sidebands are separated by two filters.
[0026] (a1.5) The two independent left and right sidebands are respectively down-converted to the center frequency, pulse-shaped and OFDM demodulated to obtain the received symbol sequence.
[0027] Furthermore, the equalizer based on the pruned neural network can adopt a multilayer perceptron model structure (i.e., including multiple perceptron layers), and the number of neurons in the input layer of the pruned neural network equalizer is twice the number of subcarriers in the input transmitted symbol sequence.
[0028] Since intercarrier interference (ICI) mainly occurs on neighboring subcarriers, a pruning technique is used to retain only the connections of neurons corresponding to the subcarriers of the input layer that are adjacent to the neurons of the first hidden layer in the first layer of the model structure (this layer is called the corresponding pruned layer). Subsequent network layers only perform pointwise operations. Based on the above pruning idea, the equalizer of this invention adopts a dual-branch heterogeneous form. One branch is a linear branch that only considers linear distortion, i.e., it does not contain activation functions. The other branch is a nonlinear branch that considers nonlinear distortion, i.e., it contains nonlinear activation functions. That is, the linear branch includes several linear layers, while the nonlinear branch consists of several linear layers with nonlinear activation functions, and the two branches have the same number of layers. At the same time, a self-attention mechanism is introduced, which learns a weight with a value range of [0,1] from the input of the equalizer to weight the output of the linear branch, so that the neural network can better learn the linear and nonlinear distortions themselves. The input data of the equalizer is fed into the self-attention layer to learn the self-attention weights of the linear branch. The output of each subcarrier of the linear branch is weighted and then added to the output of the corresponding subcarrier of the nonlinear branch to obtain the output of the equalizer based on the pruning neural network.
[0029] In this system, each neuron in the first layer of the linear and nonlinear branches is connected only to the neurons in the input layer that represent the subcarrier represented by the current neuron. From the second layer onwards, the linear and nonlinear branches only perform point-by-point operations, meaning that neurons in adjacent layers are only connected to neurons that represent the same subcarrier.
[0030] Furthermore, the weights of each network layer in both linear and nonlinear branches during computation are the weights W of that network layer and the corresponding mask matrix W. mThe dot product is used to achieve pruning of each network layer;
[0031] Wherein, the mask matrix W of the first layer m Based on the number of subcarriers N sub Inter-carrier interference (ICI) length N ICI Determined: 0 indicates the subcarrier at the current position is not interfered with, 1 indicates the subcarrier at the current position is interfered with, and the number of 1s in each row except the first and last rows is 2 × N. ICI +2, the numerical value 2 in this expression is because the real and imaginary parts are separated;
[0032] Linear and nonlinear branches begin from the second layer, and their operation uses a mask matrix W. m It is a block diagonal matrix to achieve the goal of performing only pointwise operations in subsequent network layers.
[0033] Furthermore, the equalizer based on the pruned neural network uses the mean squared error as the loss function during training.
[0034] Furthermore, in the equalized channel model, each neuron in the first hidden layer is connected only to the neurons in the input layer that represent the subcarriers of the current neuron. From the second layer onwards, neurons in adjacent layers are connected only to neurons that represent the same subset of subcarriers.
[0035] Furthermore, the output of the equalized channel model describes the conditional probability distribution of the equalized channel. The set of parameters includes residual linear and nonlinear distortions and noise in the equalized channel. The equalized channel model learns multiple Gaussian distributions and their weighted combinations by transmitting symbols to achieve the conditional probability distribution of the equalized channel. The essence of the channel model output after fitting and equalization is a probability distribution. Therefore, the negative log-likelihood function is selected as its loss function during training.
[0036] Furthermore, in step a3, when simultaneously training the constellation shaper at the transmitting end and the demapping unit at the receiving end, the training algorithm framework of the configured constellation shaper includes a geometric shaper, a probability shaper, a symmetric quadrant constraint layer, a power normalization layer, a Softmax function layer, and a sampler.
[0037] The geometric shaper is used to transform the sequential one-hot symbol table into an unnormalized constellation table in the first quadrant of four-dimensional space. The mapping, where This represents the bit sequence corresponding to the symbol;
[0038] The probability shaper is used to transform an arbitrarily specified constant into a corresponding value. A set of nonstandardized log probabilities Mapping;
[0039] Symmetric quadrant constraint layer is used for and Performing a mirroring operation yields the complete four-dimensional unnormalized constellation table. and corresponding unstandardized log probability
[0040] The Softmax function layer is used to process the output of the symmetric quadrant constraint layer. Performing the Softmax operation yields the probability distribution of the four-dimensional constellation diagram.
[0041] Power normalization layer based on probability distribution The output of the symmetric quadrant constraint layer Perform power normalization to obtain a four-dimensional geometrically shaped constellation table.
[0042] Sampler according to probability distribution The one-hot symbol sequence is obtained by using the method, and then compared with the four-dimensional geometric integer constellation table. Multiplying them yields a four-dimensional frequency domain transmitted symbol sequence x;
[0043] When the preset training convergence conditions are met (such as the number of training iterations reaching a preset limit, or the loss function value converging), based on the currently obtained four-dimensional geometric integer constellation table... Unstandardized log probability The trained constellation shaper is obtained for use in the four-dimensional geometric and probabilistic constellation shaping during the deployment phase.
[0044] Preferably, the sampler can perform sampling based on the Gumbel-Softmax technique.
[0045] Furthermore, in step a3, when simultaneously training the constellation shaper at the transmitting end and the demapper at the receiving end, maximizing generalized mutual information (GMI) is used as the optimization objective for training.
[0046] Furthermore, the loss function used when simultaneously training the constellation shaper at the transmitting end and the demapper at the receiving end is:
[0047]
[0048] Where E(·) is the mathematical expectation, and m is the bit sequence estimate. The sequence length, b i Bit sequence estimation The i-th symbol, H represents the conditional probability distribution corresponding to the demapper.w (X) represents the source entropy, and the subscript w represents the trainable parameters of the constellation shaper and demapper.
[0049] The technical solution provided by this invention brings at least the following beneficial effects:
[0050] (1) The equalization method based on pruned neural networks adopted in this invention naturally takes into account the characteristic that ICI is mainly concentrated on the neighboring subcarriers of the current subcarrier under normal circumstances, reducing the space and time complexity of the neural network equalization algorithm, so that its parameter quantity and computational quantity increase linearly with the increase of the number of subcarriers, rather than exponentially. Furthermore, the dual-branch heterogeneous neural network architecture is adopted, and a self-attention mechanism is introduced to help the network itself better focus its attention on the compensation of linear and nonlinear impairments;
[0051] (2) The equalization-based channel modeling method adopted in this invention fits the conditional probability distribution of the channel through a hybrid density network, rather than learning a deterministic neural network model with fixed parameters, which is more conducive to simulating the stochastic characteristics of the channel. By adopting the idea of equalization before channel modeling, the hybrid density network only needs to learn the residual distortion and noise of the equalized channel, simplifying the difficulty of channel modeling;
[0052] (3) The end-to-end four-dimensional constellation shaping method based on independent sideband modulation adopted in this invention reduces the size of the four-dimensional geometric probability constellation search space by using symmetric quadrant constraints. The constellation shaper at the transmitter and the demapper at the receiver are connected by an equalized channel model to take into account the full-link impairment mechanism. The globally optimal four-dimensional joint geometric probability constellation shaping scheme is adaptively learned, which can further approach the Shannon limit. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is a schematic diagram illustrating the implementation framework of an artificial intelligence-based optimization method for fiber optic terahertz communication systems provided in an embodiment of the present invention.
[0055] Figure 2 This is a diagram of an equalizer model based on a pruned neural network in an embodiment of the present invention.
[0056] Figure 3 This is a diagram of the equalized channel model based on a hybrid density network in an embodiment of the present invention.
[0057] Figure 4 This is a schematic diagram illustrating the implementation process of the end-to-end four-dimensional joint geometric probability constellation shaper in an embodiment of the present invention. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be described in detail and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Generally, the components of the embodiments of the present invention described and shown in the accompanying drawings can be arranged and designed using different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present invention.
[0059] For ease of description, the relevant technical terms appearing in the embodiments of this invention are explained as follows:
[0060] MZM (Mach-Zehnder Modulator): Mach-Zehnder modulator
[0061] OFDM (Orthogonal Frequency Division Multiplexing): Orthogonal frequency division multiplexing;
[0062] ISB (Independent Sideband): Independent Sideband
[0063] MDN (Mixture Density Network): A network with a mixed density.
[0064] ICI (Inter-Carrier Interference): Inter-carrier interference
[0065] GMI (Generalized Mutual Information): Generalized Mutual Information
[0066] As one possible implementation, the artificial intelligence-based optimization method for fiber optic terahertz communication systems provided in this embodiment of the invention is applicable to integrated terahertz fiber optic communication systems and can also be applied to other communication systems. This method includes a training phase and a deployment phase, such as... Figure 1 As shown, the specific implementation steps are as follows:
[0067] I. The training phase includes the following steps:
[0068] S1. The transmitter randomly generates a transmission symbol sequence, which passes through the full-link fiber optic terahertz channel and is then used to obtain a reception symbol sequence at the receiver. The transmission symbol sequence and the reception symbol sequence are used for the subsequent training of the pruned neural network equalizer.
[0069] S1.1 The transmitter generates a random symbol sequence that follows a uniform probability distribution. After serial-to-parallel conversion, two independent data streams are obtained. These streams are then up-converted to the center frequency f after OFDM modulation and pulse shaping. s This yields two independent double-sideband OFDM signals; in this embodiment, the characteristic of a single sideband is set to 10 Gbaud, f s =7.5GHz;
[0070] S1.2, the two double-sideband signals provide drive inputs for two MZMs operating in push-pull mode, to modulate the signal from a single operating frequency f. c A single-mode laser generates a continuous wavelength optical carrier, resulting in two double-sideband optical signals. The left and right sidebands are extracted by two optical filters, and the signals are then combined using an optical power combiner to obtain the ISB optical signal. In this embodiment, f is set... c =193.1 THz;
[0071] S1.3 The generated ISB optical signal is transmitted through a single-mode optical fiber. At the receiving end of the optical fiber link, it is amplified by an erbium-doped fiber amplifier and then connected to a signal with a working frequency of f. c+t The single-mode laser generates continuous wavelength optical carriers, which are combined based on the principle of optical heterodyne beat frequency. The combined signal is then converted by UTC-PD to a center frequency of f. t The ISB terahertz signal; in this embodiment, f is set c+t =193.4THz, f t =300GHz, i.e., f c+t =f c +f t ;
[0072] S1.4 The generated ISB terahertz signal is radiated into free space by the terahertz antenna. At the receiving end of the terahertz link, the local radio frequency source is mixed with the received terahertz signal after N frequency multiplications to convert the ISB terahertz signal into an ISB baseband signal. Then, two equal-power ISB baseband signals are obtained through a power divider, and the independent left and right sidebands are separated by two filters.
[0073] S1.5. The two independent left and right sidebands are respectively down-converted to the center frequency, pulse-shaped and OFDM demodulated to obtain the received symbol sequence;
[0074] S2, the transmitted symbol sequence and the received symbol sequence are divided into training data and validation data to train an equalizer based on a pruned neural network to achieve the mapping from received symbols to transmitted symbols;
[0075] In this embodiment of the invention, the equalizer is implemented using a multilayer perceptron, with mean square error as the loss function, to learn the mapping from received OFDM frequency domain symbols to transmitted OFDM frequency domain symbols. Considering the general case, ICI mainly focuses on the neighboring subcarriers of the current subcarrier. Through pruning techniques, in the first layer of the neural network, only the connections of neurons corresponding to input layer subcarriers adjacent to the neurons corresponding to the first hidden layer are retained; subsequent layers only perform point-by-point operations. The pruned neural network adopts a dual-branch heterogeneous form: one is a linear branch, considering only linear distortion (i.e., without activation functions), and the other is a nonlinear branch, considering nonlinear distortion (i.e., with activation functions). Furthermore, a self-attention mechanism is introduced, learning a 0-1 weight through the input to weight the linear layer, so that the neural network can better learn the linear and nonlinear distortions themselves. In this embodiment, the number of OFDM subcarriers is set to 1024, the number of hidden layers in the pruned neural network is 2, and the ICI length is 1. Since the number of subcarriers is large at this point, it is difficult to represent it graphically. Therefore, a simplified model is used to describe the structure of the pruned neural network. Assuming the number of OFDM subcarriers is 3, the number of hidden layers in the pruned neural network is 2, and the ICI length is 1, the schematic diagram of the pruned neural network structure is as follows: Figure 2 As shown, the circles with numbers represent bundled sets of neurons, connected by fully connected connections. SiLU is used as the activation function, and Attn is the attention layer, which can be represented as two weight parameters w calculated from the input. lin and w nlin Finally, the pruned neural network in this example can be represented by the following expression:
[0076]
[0077] In this context, the subscripts lin and nlin represent linear and nonlinear branches, respectively, and the corresponding numbers are used to distinguish network layers, i.e., the layer numbers of the network layers. For example, lin3 represents the linear branch of the third layer, and w represents the attention weight. p ,b represent the linear layer weights and biases, respectively. The SiLU activation function is used. Note that due to the application of pruning techniques, W... p In reality, it's the original complete linear layer weights W and a mask matrix W m The result of the dot product is that the mask matrix can be easily represented by the number of subcarriers and the ICI length. For example, the mask matrix of the first layer of this simplified pruned neural network model is as follows:
[0078]
[0079] S3. Train an equalized channel model based on MDN using the transmitted symbol sequence and the equalized symbol sequence. This model is then used to connect the transmitter constellation shaper and the receiver demapping unit to achieve end-to-end gradient direction propagation, taking into account the end-to-end impairment mechanism.
[0080] In this embodiment, the channel model is implemented using MDN, with the input being transmitted OFDM frequency domain symbols and the output being a description of the equalized channel conditional probability distribution. A set of parameters, where the latter includes the effects of residual linear and nonlinear distortion and noise in the equalized channel, is used by MDN to learn multiple Gaussian distributions and their weighted combinations through transmitted symbols, thereby achieving a conditional probability distribution of the equalized channel. The fitting is shown below:
[0081]
[0082] Where, π k (x) is the mixing coefficient. The mean is μ k (x), variance is The Gaussian distribution. Note that π... k (x) must satisfy This can be achieved using Softmax, where the variance must be positive, and can be achieved by exponentially activating the output.
[0083] The essence of MDN output is a probability distribution, therefore the negative log-likelihood function is chosen as the loss function:
[0084]
[0085] Where ω is a trainable parameter in MDN.
[0086] In this embodiment, K=3, the number of hidden layers is 1, and the first layer uses pruning techniques to consider only ICIs of length 1. SiLU is selected as the activation function, exponential activation is used to ensure positive variance, and Softmax is used to ensure that the sum of Gaussian weights is 1. Similarly, due to the large number of subcarriers, it is difficult to depict the specific structure of the MDN graphically. Therefore, a simplified model is used here, assuming the number of subcarriers is 3, K=2, and the ICI length is 1. The MDN structure at this time is as follows: Figure 3 As shown.
[0087] S4. The transmitter generates a four-dimensional constellation table through a geometric shaper and a corresponding probability distribution through a probability shaper. During the training of the four-dimensional joint geometric probability constellation shaper (referred to as the constellation shaper, whose training algorithm framework includes a geometric shaper, a probability shaper, a symmetric quadrant constraint layer, a power normalization layer, a softmax function layer, and a sampler) and the demapper, the transmitter generates random data symbols in two independent sidebands. After OFDM modulation, the received symbols are obtained through a trained and fixed-parameter equalized channel model. After OFDM demodulation, the bit sequence is estimated by the demapper. See also... Figure 4 The specific implementation steps include:
[0088] S4.1 The geometric shaper is implemented in the form of a multilayer perceptron, completing the transformation from a sequential one-hot symbol table to an unnormalized constellation table in the first quadrant of four-dimensional space. The mapping, where This represents the bit sequence corresponding to the symbol;
[0089] S4.2 The probability shaper is also implemented in the form of a multilayer perceptron, transforming an arbitrarily specified constant into a corresponding... A set of nonstandardized log probabilities Mapping;
[0090] S4.3, through the symmetric quadrant constraint layer and The mirror operation yields the complete four-dimensional unnormalized constellation table. and corresponding unstandardized log probability
[0091] S4.4, using the Softmax function layer The Softmax operation yields the probability distribution of the four-dimensional constellation diagram. Power normalization layer based on probability distribution right Normalization yields a four-dimensional geometrically shaped constellation table.
[0092] S4.5, Using a sampler based on the Gumbel-Softmax technique, distributed according to probability. Sampling is performed to obtain a one-hot symbol sequence, which is then compared with a four-dimensional geometric integer constellation table. Multiplying these results in a four-dimensional frequency domain transmitted symbol sequence x, which is the output of the constellation shaper.
[0093] S4.6. The equalized four-dimensional frequency domain received symbol sequence is obtained by training the four-dimensional frequency domain transmitted symbol sequence x and equalizing the channel model with fixed parameters.
[0094] S4.7 The demapping unit is implemented in the form of a multilayer perceptron to complete the four-dimensional frequency domain received symbol sequence. To the corresponding bit sequence estimation Mapping;
[0095] S4.8. This end-to-end optimization framework aims to maximize the GMI, which is the achievable speed of the BICM system. Considering the relationship between the bit cross-entropy loss function and GMI, bit sequence estimation is used... The loss function for constructing the end-to-end joint geometric probability constellation shaping is calculated based on the emission symbol distribution as follows:
[0096]
[0097] In this equation, the first term is equivalent to the bit cross-entropy loss of the demapper output. H represents the conditional probability distribution corresponding to the demapper. w (X) represents the source entropy, which needs to be subtracted to avoid the probability shaper learning to an extreme one-hot form. The calculated loss function can be used to train the geometry shaper, probability shaper, and demapper end-to-end;
[0098] II. After the training phase, a pruned neural network equalizer, geometric shaper, probability shaper, and demapper with fixed parameters will be obtained. The specific steps of the deployment phase are as follows:
[0099] S5. A fixed-parameter geometric shaper can be equivalently represented by a fixed constellation table, and a fixed-parameter probability shaper can be equivalently represented by the corresponding probability distribution. First, the transmitted symbol sequence is obtained from the information bit sequence through the constellation table and the corresponding probability distribution. The probability shaping can be achieved by symbol-by-symbol matching or block-by-block matching. The former is similar to lookup tables and dynamic programming, while the latter is similar to CCDM. That is, the four-dimensional constellation table output by the trained four-dimensional joint geometric probability constellation shaper can map the information bit sequence into a four-dimensional frequency domain transmitted symbol sequence x.
[0100] S6. The transmitted symbol sequence is obtained through the full-link fiber optic terahertz channel in step 1 of the same training phase to obtain the corresponding received symbol sequence, and then the received symbol sequence is obtained after being pruned by a fixed-parameter neural network equalizer.
[0101] S7. The received symbol sequence after equalization is demapped to obtain the estimated received information bit sequence.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0103] The above descriptions are merely some embodiments of the present invention. Those skilled in the art can make various modifications and improvements without departing from the inventive concept of the present invention, and these all fall within the scope of protection of the present invention.
Claims
1. An optimization method for an artificial intelligence-based fiber optic terahertz communication system, characterized in that, Includes the following steps: Training phase steps: Step a1: The transmitter randomly generates a sequence of transmitted symbols, which is then transmitted through a full-link fiber optic terahertz channel to obtain a sequence of received symbols at the receiver. An equalizer based on a pruned neural network is trained using transmitted and received symbol sequences as training data. Step a2: Train an equalized post-channel model based on a hybrid density network based on the transmitted symbol sequence and the equalized symbol sequence; Among them, the equalized channel model is used to output the conditional probability distribution of the equalized channel. The composition parameters are used to achieve the following: The fit; The equalized symbol sequence is obtained by inputting the emitted symbol sequence into a trained equalizer based on a pruned neural network and obtaining the equalized symbol sequence based on its output. Step a3: Simultaneously train the constellation shaper at the transmitting end and the demapping mechanism at the receiving end, wherein both the constellation shaper and the demapping mechanism are neural network models. The constellation shaper includes a geometric shaper and a probabilistic shaper, used to generate a four-dimensional geometric and probabilistic constellation table to shape the transmitted bit sequence into a four-dimensional frequency-domain transmitted symbol sequence. ; Transmit symbol sequences in the four-dimensional frequency domain The trained equalization channel model is input to obtain a four-dimensional frequency domain received symbol sequence. Then the four-dimensional frequency domain received symbol sequence Input the demapper to obtain the corresponding bit sequence estimate. ; Deployment phase steps: Step b1: The trained constellation shaper generates a constellation table, which is used to shape the information bit sequence at the transmitting end to obtain a four-dimensional frequency domain transmitted symbol sequence. ; Step b2, four-dimensional frequency domain transmitted symbol sequence Using the same full-link fiber terahertz channel as in step a1, the corresponding received symbol sequence is obtained at the receiving end, and then the received symbol sequence is sent to a trained equalizer based on a pruned neural network to obtain an equalized received symbol sequence. Step b3: After equalization, the received symbol sequence is demapped to obtain the estimated received information bit sequence.
2. The method as described in claim 1, characterized in that, In step a1, the transmitting end randomly generates a transmitted symbol sequence, which, after passing through a full-link fiber optic terahertz channel, results in a received symbol sequence at the receiving end, including: a1.1 The transmitter generates a random symbol sequence following a uniform probability distribution. After serial-to-parallel conversion, two independent data streams are obtained. These streams are then modulated using Orthogonal Frequency Division Multiplexing (OFDM), pulse-shaped, and up-converted to their center frequencies. This yields two independent double-sideband OFDM signals; a1.2 Two independent double-sideband OFDM signals provide drive inputs for two Mach-Zehnder modulators (MZMs) operating in push-pull mode, respectively, to modulate the continuous-wavelength optical carrier generated by a single-mode laser, resulting in two double-sideband optical signals. The operating frequency of the single-mode laser is [insert frequency here]. ; The left and right sidebands of the two double-sideband optical signals are extracted by two optical filters respectively, and the independent sideband ISB optical signals are obtained by optical power combiner. a1.3 The ISB optical signal is transmitted into a single-mode optical fiber. At the receiving end of the optical fiber link, it is amplified by an erbium-doped fiber amplifier and then connected to a signal with a working frequency of [missing information]. The continuous wavelength optical carrier generated by the single-mode laser is combined, and the resulting combined signal is converted into a center frequency of by a single-carrier photodiode. ISB terahertz signal; The 1.4ISB terahertz signal is radiated into free space by the terahertz antenna. At the receiving end of the terahertz link, the local radio frequency source mixes the received terahertz signal after multiple frequency multiplications to convert the ISB terahertz signal into an ISB baseband signal. Then, two equal-power ISB baseband signals are obtained through a power divider, and then the independent left and right sidebands are separated by two filters. a1.5 The two independent left and right sidebands are respectively down-converted to the center frequency, pulse-shaped and OFDM demodulated to obtain the received symbol sequence.
3. The method as described in claim 1, characterized in that, The equalizer based on the pruned neural network consists of multiple perceptual layers. The number of neurons in the input layer of the pruned neural network equalizer is twice the number of subcarriers in the input transmitted symbol sequence.
4. The method as described in claim 3, characterized in that, The equalizer based on the pruned neural network adopts a dual-branch heterogeneous form. One branch is a linear branch with several linear layers, and the other branch is a nonlinear branch with several linear layers with nonlinear activation functions. The number of layers in the two branches is the same. In the linear and nonlinear branches, each neuron in the first hidden layer is connected only to the neurons in the input layer that represent the subcarriers of the current neuron; from the second layer onwards, neurons in the linear and nonlinear branches are connected only to the neurons that represent the same subset of subcarriers. The input data of the equalizer based on the pruning neural network is fed into the self-attention layer to learn the self-attention weights of the linear branch. The output results of each subcarrier of the linear branch are weighted and then added to the output results of the corresponding subcarrier of the nonlinear branch to obtain the output of the equalizer based on the pruning neural network.
5. The method as described in claim 4, characterized in that, The weight of each network layer in linear and nonlinear branches during computation is the weight of that network layer. With the corresponding mask matrix dot product; Among them, the mask matrix of the first layer of linear and nonlinear branches Based on the number of subcarriers Inter-carrier interference (ICI) length Confirmed, 0 indicates that the subcarrier at the current position is not interfered with, 1 indicates that the subcarrier at the current position is interfered with, and the number of 1s in each row is: ; Linear and nonlinear branches, starting from the second layer, use mask matrices during computation. It is a block diagonal matrix.
6. The method according to any one of claims 1 to 5, characterized in that, The equalizer based on the pruned neural network uses the mean squared error as the loss function during training.
7. The method as described in claim 1, characterized in that, In the equalized channel model, each neuron in the first hidden layer is connected only to the neurons in the input layer that represent the subcarriers of the current neuron. From the second layer onwards, neurons in adjacent layers are connected only to neurons that represent the same subset of subcarriers.
8. The method as described in claim 1, characterized in that, The equalized channel model selects the negative log-likelihood function as its loss function during training.
9. The method as described in claim 1, characterized in that, In step a3, when training the constellation shaper at the transmitting end and the demapping unit at the receiving end simultaneously, the training algorithm framework of the constellation shaper includes a geometric shaper, a probability shaper, a symmetric quadrant constraint layer, a power normalization layer, a Softmax function layer, and a sampler. Among them, the geometric shaper is used to transform the sequential one-hot symbol table into an unnormalized constellation table in the first quadrant of four-dimensional space. The mapping, where This represents the bit sequence corresponding to the symbol; The probability shaper is used to transform an arbitrarily specified constant into a corresponding value. A set of nonstandardized log probabilities Mapping; Symmetric quadrant constraint layer is used for and Performing a mirroring operation yields the complete four-dimensional unnormalized constellation table. and corresponding unstandardized log probability ; The Softmax function layer is used to process the output of the symmetric quadrant constraint layer. Performing the Softmax operation yields the probability distribution of the four-dimensional constellation diagram. ; Power normalization layer based on probability distribution The output of the symmetric quadrant constraint layer Perform power normalization to obtain a four-dimensional geometrically shaped constellation table. ; Sampler according to probability distribution Sampling is performed to obtain a one-hot symbol sequence, which is then compared with a four-dimensional geometric integer constellation table. Multiplying yields a four-dimensional frequency domain transmitted symbol sequence ; When the preset training convergence condition is met, based on the currently obtained four-dimensional geometrically shaped constellation table... Unstandardized log probability Obtain a well-trained constellation shaper.
10. The method as described in claim 1 or 9, characterized in that, In step a3, when training the constellation shaper at the transmitting end and the demapping machine at the receiving end simultaneously, maximizing the generalized mutual information is used as the optimization objective of the training.