Optical fiber transmission and reception signal prediction model training method, prediction method and equipment
By training the neural network model and utilizing the hybrid domain feature extraction and peak-aware attention mechanism, the problem of inaccurate received signal prediction in the CS-NFDM system is solved, and efficient fiber optic communication system modeling and performance prediction are achieved.
Patent Information
- Application Number
- CN202511100546.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing technologies cannot effectively predict received signals in continuous spectrum nonlinear frequency division multiplexing (CS-NFDM) systems, resulting in inaccurate modeling of optical fiber communication systems and increased equipment resource consumption.
A neural network model is used for training. Through hybrid domain feature extraction, multi-scale processing and peak perception attention mechanism, the received signal prediction result data is generated, and the neural network model is optimized to improve the prediction accuracy.
It improves the reliability and effectiveness of the optical fiber transmission receiving signal prediction model, shortens the optical fiber communication system modeling cycle, reduces computing resource consumption, provides a more effective modeling basis, and improves the design and performance prediction efficiency of the optical fiber communication system.
Smart Images

Figure CN120601969A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of transmission technology, and in particular to a training method, prediction method, and device for a prediction model of a received signal transmitted through an optical fiber. Background Art
[0002] With the dramatic increase in global data traffic and network carrying capacity demands, high-capacity, high-speed fiber-optic communication technology has become a core pillar of information infrastructure. Nonlinear frequency division multiplexing (NFDM) systems are a key technology for next-generation high-speed optical communications. Continuous spectrum nonlinear frequency division multiplexing (CS-NFDM) is a fiber-optic communication technology based on the nonlinear Fourier transform (NFT). By modulating the signal onto the continuous spectral components of the nonlinear spectrum for transmission, CS-NFDM effectively combats the Kerr nonlinearity effect in optical fibers. Modeling the complex channel characteristics of CS-NFDM has always been a technical challenge.
[0003] At present, although classic waveform simulation methods such as the Split-Step Fourier Transform (SSFM) algorithm and BERT-based polarization-multiplexed fiber channel modeling have been developed in the field of fiber-optic communication system modeling, there is no automatic prediction method specifically suitable for received signals in continuous spectrum nonlinear frequency division multiplexing (CS-NFDM) systems. The accuracy of received signal prediction in CS-NFDM systems cannot be guaranteed, and thus the accuracy of fiber-optic communication system modeling for CS-NFDM systems cannot be guaranteed. This will also affect its modeling efficiency and lead to increased equipment resource consumption. Summary of the Invention
[0004] In view of this, embodiments of the present application provide a method for training a prediction model for optical fiber transmission received signals, a prediction method, and an apparatus to eliminate or improve one or more defects in the prior art.
[0005] One aspect of the present application provides a method for training a prediction model for optical fiber transmission received signals, comprising: Generating a training sample corresponding to a transmission signal pre-generated by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system, and generating a label corresponding to the training sample based on a received signal formed after the transmission signal is pre-transmitted to a receiving end via an optical fiber; A neural network model is trained based on the training samples, so that the neural network model performs mixed domain feature extraction, multi-scale processing, and attention feature extraction based on the peak perception attention mechanism on the training samples to obtain reception signal prediction result data for predicting the reception signal formed after the training samples are transmitted to the receiving end through the optical fiber, and the neural network model is optimized based on the reception signal prediction result data corresponding to the training samples and the labels, so as to train the neural network model as an optical fiber transmission reception signal prediction model for outputting the reception signal prediction result data corresponding to the transmission signal.
[0006] In some embodiments of the present application, the neural network model includes a sample processing module and a peak-aware attention module; The sample processing module is used to obtain the three-dimensional tensor corresponding to the training sample, and perform hybrid domain feature extraction and multi-scale processing on the three-dimensional tensor corresponding to the training sample to obtain multi-scale feature data corresponding to the training sample; The peak-aware attention module is used to perform peak detection on the multi-scale feature data corresponding to the training sample to obtain the peak weight corresponding to the training sample, and perform attention detection on the multi-scale feature data to obtain the attention weight corresponding to the training sample; fuse the peak weight and the attention weight to obtain the corresponding fusion weight, and normalize and feature aggregate the fusion weight to obtain the attention feature data corresponding to the training sample for generating the received signal prediction result data.
[0007] In some embodiments of the present application, the peak-aware attention module includes: a peak detection branch, a standard detection branch, a peak fusion layer, a normalization layer, a feature aggregation layer, and an output convolution layer; the peak detection branch and the standard detection branch are both connected to the peak fusion layer; the peak fusion layer, the normalization layer, the feature aggregation layer, and the output convolution layer are connected in sequence; The peak detection branch is used to perform peak detection on the multi-scale feature data corresponding to the training sample to obtain the peak weight corresponding to the training sample; The standard detection branch is used to perform attention detection on the multi-scale feature data corresponding to the training sample based on the query vector, the key vector and the value vector to obtain the attention weight corresponding to the training sample; The peak fusion layer is used to fuse the peak weight and the attention weight to obtain a corresponding fusion weight; The normalization layer is used to normalize the fusion weight to obtain the normalized fusion weight; The feature aggregation layer is used to multiply the normalized fusion weight and the value vector to obtain corresponding weighted feature data; The output convolution layer is used to perform convolution and residual connection on the weighted feature data to obtain attention feature data corresponding to the training sample.
[0008] In some embodiments of the present application, the neural network model further includes: a peak enhancement module and an output layer; The peak enhancement module is connected to the peak perception attention module, and the peak enhancement module is used to perform channel compression and recovery processing on the attention feature data in sequence based on two convolutional layers to obtain corresponding peak enhancement feature data, and perform weighted residual connection on the peak enhancement feature data and the attention feature data to obtain peak enhancement attention feature data corresponding to the training sample; The output layer is used to generate received signal prediction result data for predicting the received signal formed after the training sample is transmitted to the receiving end through the optical fiber based on the peak enhanced attention feature data; wherein, the received signal prediction result data includes: the real part prediction component and the imaginary part prediction component of the received signal.
[0009] In some embodiments of the present application, the sample processing module includes: an input layer, a mixed domain feature extraction unit, and a multi-scale feature extraction unit; The input layer is used to receive the training sample, wherein the training sample includes: a real component and an imaginary component corresponding to the transmission signal; the input layer is further used to stack the real component and the imaginary component corresponding to the transmission signal in the channel dimension to obtain a three-dimensional tensor corresponding to the training sample; wherein each dimension in the three-dimensional tensor is used to represent the batch size, the number of channels, and the time series length respectively; The mixed domain feature extraction unit is used to perform time domain feature extraction and frequency domain feature extraction on the three-dimensional tensor corresponding to the training sample to obtain time domain feature data and frequency domain feature data corresponding to the training sample, and perform feature fusion on the time domain feature data and the frequency domain feature data to obtain mixed domain feature data corresponding to the training sample; The multi-scale feature extraction unit is used to perform feature extraction of multiple time scales on the mixed domain feature data corresponding to the training sample to obtain multiple time scale feature data corresponding to the training sample, and perform feature fusion processing on each of the time scale feature data to obtain multi-scale feature data corresponding to the training sample.
[0010] In some embodiments of the present application, generating a training sample corresponding to a transmission signal pre-generated by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system, and generating a label corresponding to the training sample based on a received signal formed after the transmission signal is pre-transmitted to a receiving end via an optical fiber, includes: Acquire signal pairs, wherein each signal pair includes a transmission signal pre-generated by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system and a receiving signal formed by transmitting the transmission signal to a receiving end via an optical fiber; performing quantile-based robust normalization processing on the transmitted signal and the received signal in each of the signal pairs; Each of the transmitted signals and each of the received signals after the robust normalization processing is separated into real components and imaginary components to obtain training samples corresponding to each of the signal pairs, and each of the received signals in each of the signal pairs is used as a label corresponding to each of the training samples; wherein each of the training samples contains a real component and an imaginary component of the transmitted signal; and each of the labels contains a real component and an imaginary component of the received signal belonging to the same signal pair as the transmitted signal uniquely corresponding to the label.
[0011] In some embodiments of the present application, optimizing the neural network model based on the received signal prediction result data corresponding to the training sample and the label includes: In the current iteration round, based on the received signal prediction result data corresponding to the training sample and the label, a preset target loss function is used to determine the target loss corresponding to the current iteration round, and the neural network model is optimized based on the target loss; The target loss function is composed of a real part loss function, an imaginary part loss function, an amplitude loss function, and an amplitude loss weight coefficient corresponding to the amplitude loss function; The real part loss function is used to represent the product between a first mean square error and a peak weighted term; wherein the first mean square error includes the mean square error between the real part prediction component corresponding to the received signal prediction result data and the real part component of the received signal corresponding to the label; the peak weighted term includes the sum of 1 and a peak product term; the peak product term includes the product of a preset peak weighting factor and a peak mask; The imaginary part loss function is used to represent the product between the second mean square error and the peak weighted term; wherein the second mean square error includes the mean square error between the imaginary part prediction component corresponding to the received signal prediction result data and the imaginary part component of the received signal corresponding to the label; The amplitude loss function is used to represent the product between the third mean square error and the peak weighted term; wherein the third mean square error includes the mean square error between the signal amplitude corresponding to the pre-acquired received signal prediction result data and the signal amplitude of the received signal corresponding to the pre-acquired label.
[0012] A second aspect of the present application provides a method for predicting a received signal during optical fiber transmission, comprising: The transmission signal currently generated by the transmitting end in the continuous spectrum nonlinear frequency division multiplexing system is input into the optical fiber transmission receiving signal prediction model, so that the optical fiber transmission receiving signal prediction model outputs corresponding receiving signal prediction result data; wherein, the optical fiber transmission receiving signal prediction model is pre-trained based on the optical fiber transmission receiving signal prediction model training method.
[0013] A third aspect of the present application provides a device for training a prediction model for optical fiber transmission reception signals, comprising: A training data construction module is configured to generate training samples corresponding to a transmission signal pre-generated by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system, and to generate labels corresponding to the training samples based on a received signal formed by pre-transmitting the transmission signal to a receiving end via an optical fiber; A model training module is used to train a neural network model based on the training samples, so that the neural network model performs mixed domain feature extraction, multi-scale processing, and attention feature extraction based on the peak perception attention mechanism on the training samples to obtain reception signal prediction result data for predicting the reception signal formed after the training samples are transmitted to the receiving end through the optical fiber, and optimize the neural network model based on the reception signal prediction result data corresponding to the training samples and the labels, so as to train the neural network model as an optical fiber transmission reception signal prediction model for outputting the reception signal prediction result data corresponding to the transmission signal.
[0014] A fourth aspect of the present application provides a device for predicting a received signal during optical fiber transmission, comprising: The model prediction module is used to input the transmission signal currently generated by the transmitting end in the continuous spectrum nonlinear frequency division multiplexing system into the optical fiber transmission receiving signal prediction model, so that the optical fiber transmission receiving signal prediction model outputs corresponding receiving signal prediction result data; wherein, the optical fiber transmission receiving signal prediction model is pre-trained based on the optical fiber transmission receiving signal prediction model training method.
[0015] The fifth aspect of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for training a prediction model for a fiber optic transmission received signal is implemented, and / or the method for predicting a fiber optic transmission received signal is implemented.
[0016] The sixth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it implements the optical fiber transmission received signal prediction model training method and / or implements the optical fiber transmission received signal prediction method.
[0017] The seventh aspect of the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the optical fiber transmission received signal prediction model training method and / or implements the optical fiber transmission received signal prediction method.
[0018] The present application provides a method for training a prediction model for a fiber optic transmission receiving signal, which generates a training sample corresponding to the transmitting signal based on a transmitting signal pre-generated by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system, and generates a label corresponding to the training sample based on a receiving signal formed after the transmitting signal is pre-transmitted to the receiving end via an optical fiber; a neural network model is trained based on the training samples, so that the neural network model performs mixed domain feature extraction, multi-scale processing, and attention feature extraction based on a peak-perceived attention mechanism on the training samples to obtain receiving signal prediction result data for predicting the receiving signal formed after the training sample is transmitted to the receiving end via an optical fiber, and optimizes the neural network model based on the receiving signal prediction result data corresponding to the training sample and the label, so as to train the neural network model to output the receiving signal corresponding to the transmitting signal. The invention discloses an optical fiber transmission receiving signal prediction model based on signal prediction result data, that is, provides a training method for a prediction model of a receiving signal in a continuous spectrum nonlinear frequency division multiplexing system, which can effectively improve the reliability and effectiveness of the optical fiber transmission receiving signal prediction model training process, realize the prediction of the receiving signal in the continuous spectrum nonlinear frequency division multiplexing system, and improve the accuracy of receiving signal prediction using the trained optical fiber transmission receiving signal prediction model, thereby being able to predict the transmission performance of the receiving signal at different environmental distances in advance during the modeling process of the high optical fiber communication system, shortening the R&D cycle of the optical fiber communication system modeling process and reducing the computing resource consumption of the R&D equipment, effectively avoiding the waste of modeling experiment costs, and providing a more effective basis for the design, performance prediction and optimization of the optical fiber communication system, so as to improve the efficiency and transmission performance of the optical fiber communication system constructed according to the modeling results of the optical fiber communication system.
[0019] Additional advantages, purposes, and features of the present application will be described in part in the following description and will become apparent to those skilled in the art upon study of the following or may be learned from practice of the present application. The purposes and other advantages of the present application may be achieved and obtained by the structures specifically pointed out in the specification and drawings.
[0020] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present application are not limited to the above specific description, and the above and other purposes that can be achieved by the present application will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings described herein are intended to provide a further understanding of the present application, constitute a part of the present application, and do not constitute a limitation of the present application. The components in the drawings are not drawn to scale, but are only for the purpose of illustrating the principles of the present application. In order to facilitate the illustration and description of some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, they may become larger than other components in the exemplary device actually manufactured according to the present application. In the drawings: Figure 1 This is a schematic diagram of a first flow chart of a method for training a prediction model for optical fiber transmission received signals in an embodiment of the present application.
[0022] Figure 2 Schematic diagram of the architecture of a neural network model in a method for training a prediction model for optical fiber transmission received signals in one embodiment of the present application.
[0023] Figure 3 Schematic diagram of the architecture of the peak-aware attention module in the neural network model in one embodiment of the present application.
[0024] Figure 4 This is a second flow chart of the optical fiber transmission reception signal prediction model training method in one embodiment of the present application.
[0025] Figure 5 This is a third flow chart of the optical fiber transmission reception signal prediction model training method in one embodiment of the present application.
[0026] Figure 6 This is a schematic diagram illustrating an example of a transmission signal in an application example of the present application.
[0027] Figure 7 This is a schematic diagram illustrating an example of a received signal in an application example of the present application.
[0028] Figure 8 This is a data preprocessing flow chart in an application example of this application.
[0029] Figure 9Schematic diagram of the loss change on the validation set in an application example of this application.
[0030] Figure 10 This is a complete flow chart of a method for training a prediction model for optical fiber transmission reception signals in an application example of this application.
[0031] Figure 11 This is a schematic diagram illustrating an example of a test transmission signal in an application example of the present application.
[0032] Figure 12 This is a schematic diagram illustrating an actual received signal in an application example of the present application.
[0033] Figure 13 This is a schematic diagram illustrating an example of received signal prediction result data in an application example of the present application.
[0034] Figure 14 This is a schematic diagram of the comparison between the actual received signal and the received signal prediction result data in an application example of the present application.
[0035] Figure 15 This is a schematic diagram of the NMSE training trend in an application example of this application.
[0036] Figure 16 Schematic diagram of the comparison of inference time between the Split-Step Fourier Transform algorithm (SSFM) and the peak-aware hybrid-domain NN in an application example of this application. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail in conjunction with the embodiments and drawings. Here, the illustrative embodiments of this application and their descriptions are used to explain this application, but are not intended to limit this application.
[0038] It should also be noted here that in order to avoid obscuring the present application due to unnecessary details, the accompanying drawings only show structures and / or processing steps that are closely related to the scheme according to the present application, while other details that are not closely related to the present application are omitted.
[0039] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.
[0040] It should also be noted that, unless otherwise specified, the term "connection" herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.
[0041] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0042] In the field of optical fiber communication system modeling, the Split-Step Fourier Transform (SSFM) algorithm, a classic waveform simulation method, suffers from high computational overhead due to its inherent segmented iterative mechanism when simulating long-distance transmission. Especially in ultra-high-speed coherent optical communication scenarios, the algorithm's time complexity exhibits a nonlinear growth relationship with transmission distance, becoming a key bottleneck restricting the efficiency of system-level optimization. With the advancement of machine learning, it has gained widespread application in optical communications. Fiber channel modeling based on bidirectional long short-term memory and generative adversarial networks have been proposed for waveform modeling using purely data-driven approaches. These data-driven approaches require large amounts of data to implement modeling, do not rely on theoretical models, and focus solely on time-domain processing. Alternatively, BERT-based polarization-multiplexed fiber channel modeling utilizes a Transformer architecture to process sequential signals and utilizes a multi-head self-attention mechanism to process the input sequence to assist in channel modeling.
[0043] However, the general Transformer architecture is not specifically optimized for the characteristics of optical fiber communication signals. Therefore, in order to achieve automatic prediction of received signals in continuous spectrum nonlinear frequency division multiplexing systems to reduce the R&D cycle of the optical fiber communication system modeling process and the computing resource consumption of R&D equipment, the embodiments of the present application respectively provide a method for training an optical fiber transmission received signal prediction model, an optical fiber transmission received signal prediction model training device for executing the optical fiber transmission received signal prediction model training method, a physical device, a computer-readable storage medium and a computer program product, which can achieve automatic prediction of received signals in continuous spectrum nonlinear frequency division multiplexing systems.
[0044] The details are described in detail through the following examples.
[0045] Based on this, the embodiment of the present application provides a method for training a prediction model of an optical fiber transmission reception signal, which can be implemented by a prediction model training device for an optical fiber transmission reception signal. Figure 1 The optical fiber transmission reception signal prediction model training method specifically includes the following contents: Step 100: Generate training samples corresponding to the transmission signal based on the transmission signal pre-generated by the transmitter in the continuous spectrum nonlinear frequency division multiplexing system, and generate labels corresponding to the training samples based on the received signal formed after the transmission signal is pre-transmitted to the receiver via optical fiber.
[0046] Continuous Spectrum Nonlinear Frequency Division Multiplexing (CS-NFDM) is a fiber-optic communication technology based on the nonlinear Fourier transform. It effectively manages the nonlinear effects of optical fibers by modulating information onto the continuous spectrum portion of the nonlinear spectrum. This technology, based on the Zakharov-Shabat scattering theory, transforms the complex nonlinear transmission problem in optical fibers into a linear evolution problem in the nonlinear frequency domain. The Zakharov-Shabat equation (ZS equation) is a core tool in the inverse scattering theory of the nonlinear Schrödinger equation (NLS), used to solve initial value problems for integrable systems.
[0047] In one or more embodiments of the present application, a transmitter may refer to a client device capable of transmitting signals, a receiver may refer to a client device capable of receiving signals, and the same client device may be both a transmitter and a receiver.
[0048] Step 200: Train a neural network model based on the training samples, so that the neural network model performs hybrid domain feature extraction, multi-scale processing, and attention feature extraction based on the peak perception attention mechanism on the training samples to obtain reception signal prediction result data for predicting the reception signal formed after the training samples are transmitted to the receiving end through the optical fiber, and optimize the neural network model based on the reception signal prediction result data corresponding to the training samples and the labels, so as to train the neural network model as an optical fiber transmission reception signal prediction model for outputting the reception signal prediction result data corresponding to the transmission signal.
[0049] In step 200, at least one iteration round of model training can be performed on the neural network model according to each of the training samples, and a preset training step can be performed in each iteration round: the training step includes: inputting the training sample into the neural network model corresponding to the current iteration round, so that the neural network model performs mixed domain feature extraction, multi-scale processing, and attention feature extraction based on the peak perception attention mechanism on the training sample to obtain reception signal prediction result data for predicting the reception signal formed after the training sample is transmitted to the receiving end through the optical fiber, and calculating the target loss value based on the reception signal prediction result data corresponding to the training sample and the label and optimizing the architecture of the neural network model based on the target loss value, and then judging whether the current iteration round is the preset last iteration round, or whether the neural network model converges, if so, stopping the iteration and using the optimized neural network model as the optical fiber transmission reception signal prediction model for the reception signal prediction result data corresponding to the transmission signal output; if not, using the optimized neural network model as the neural network model corresponding to the next iteration round.
[0050] In step 300, hybrid domain feature extraction refers to a signal processing method that simultaneously utilizes time domain and frequency domain information; multi-scale processing refers to a method of extracting features of different time scales using different convolution kernel sizes; peak perception refers to a mechanism for focusing on and optimizing the peak area of the signal. The self-attention mechanism is an attention mechanism that learns the dependency between positions by calculating the correlation between each position and all positions in the sequence. Correspondingly, the attention feature based on the peak-aware attention mechanism refers to an attention mechanism designed specifically for the peak area of the signal. The standard self-attention is enhanced by generating weights through the peak detection branch to achieve focused attention on important signal areas.
[0051] From the above description, it can be seen that the optical fiber transmission receiving signal prediction model training method provided in the embodiment of the present application provides a training method for the prediction model of the receiving signal in the continuous spectrum nonlinear frequency division multiplexing system, which can effectively improve the reliability and effectiveness of the optical fiber transmission receiving signal prediction model training process, and can realize the prediction of the receiving signal in the continuous spectrum nonlinear frequency division multiplexing system, and can improve the accuracy of the received signal prediction using the trained optical fiber transmission receiving signal prediction model, thereby being able to predict the transmission performance of the receiving signal at different environmental distances in advance during the modeling process of the high optical fiber communication system, shortening the R&D cycle of the optical fiber communication system modeling process and reducing the computing resource consumption of the R&D equipment, effectively avoiding the waste of modeling experiment costs, and providing a more effective basis for the design, performance prediction and optimization of the optical fiber communication system, so as to improve the efficiency and transmission performance of the optical fiber communication system constructed according to the modeling results of the optical fiber communication system.
[0052] In order to further improve the attention paid to the special importance of the peak area of the communication signal during the training process of the optical fiber transmission reception signal prediction model, in the optical fiber transmission reception signal prediction model training method provided in the embodiment of the present application, see Figure 2 , the neural network model specifically includes the following contents: Sample processing module and peak-aware attention module.
[0053] The sample processing module is used to obtain the three-dimensional tensor corresponding to the training sample, and perform hybrid domain feature extraction and multi-scale processing on the three-dimensional tensor corresponding to the training sample to obtain multi-scale feature data corresponding to the training sample.
[0054] The peak-aware attention module is used to perform peak detection on the multi-scale feature data corresponding to the training sample to obtain the peak weight corresponding to the training sample, and perform attention detection on the multi-scale feature data to obtain the attention weight corresponding to the training sample; fuse the peak weight and the attention weight to obtain the corresponding fusion weight, and normalize and feature aggregate the fusion weight to obtain the attention feature data corresponding to the training sample for generating the received signal prediction result data.
[0055] The attention weights can be specifically represented as an attention weight matrix, also known as a similarity matrix. The peak weights can be specifically represented as a peak weight map with a value range between 0 and 1 and a dimension of [batch_size, 1, 1024], where batch_size represents the batch size. The peak weights are used to identify important locations in the transmitted signal that may correspond to peaks in the received signal.
[0056] Based on this, the fusion of the peak weight and the attention weight to obtain the corresponding fusion weight can be: first expand the peak weight map, and then multiply the expanded peak weight map with the attention weight matrix element by element, so that the peak area obtains a higher weight in the attention calculation.
[0057] In order to further improve the effectiveness and reliability of the application of the peak perception attention module in the optical fiber transmission reception signal prediction model, in the optical fiber transmission reception signal prediction model training method provided in the embodiment of the present application, see Figure 3 The peak-aware attention module in the neural network model specifically includes the following contents: A peak detection branch, a standard detection branch, a peak fusion layer, a normalization layer, a feature aggregation layer and an output convolution layer; the peak detection branch and the standard detection branch are both connected to the peak fusion layer; the peak fusion layer, the normalization layer, the feature aggregation layer and the output convolution layer are connected in sequence.
[0058] The peak detection branch is used to perform peak detection on the multi-scale feature data corresponding to the training sample to obtain the peak weight corresponding to the training sample.
[0059] Specifically, the peak-aware attention module is one of the core improvements in this embodiment, specifically designed to process peak characteristics in optical fiber transmission. In CS-NFDM systems, due to nonlinear effects, the received signal exhibits significant peaks compared to the transmitted signal at certain moments. These peaks contain critical transmission information. This module receives multiscale feature data (multiscale_features) as input and consists of two collaborative branches: a peak detection branch and a standard attention branch.
[0060] The peak detection branch automatically identifies peak locations using a two-layer convolutional network. The first convolutional network compresses the multi-scale feature data of the 64 input channels into 16 channels and activates them using the ReLU function. The second convolutional network further compresses the multi-scale feature data of the 16 channels into 1 channel and activates it using the Sigmoid function, outputting a peak weight map (peak_weights) with values ranging from 0 to 1. The ReLU function is a simple and effective activation function that achieves nonlinear mapping by truncating negative inputs to 0 and retaining the original values of positive inputs. The sigmoid function is a common S-shaped function in biology, also known as the S-shaped growth curve.
[0061] The standard detection branch is used to perform attention detection on the multi-scale feature data corresponding to the training sample based on the query vector, the key vector and the value vector to obtain the attention weight corresponding to the training sample.
[0062] Specifically, the standard attention branch implements the self-attention mechanism, which can generate a query vector (Q), a key vector (K), and a value vector (V) through three independent 1×1 convolutional layers. Each vector maps the input 64-channel multi-scale feature data to a new 64-channel representation. The calculation of the attention weight is first obtained by matrix multiplication of Q and K to obtain the similarity matrix, and then divided by the square root of the feature dimension for scaling.
[0063] Among them, the similarity matrix corresponds to the attention weight matrix calculated in the standard attention branch, and the attention weight matrix The calculation formula is as follows: in, Represents a similarity matrix, which represents the similarity relationship between time points. Each element (i, j) in the similarity matrix represents the degree of similarity between the i-th time point and the j-th time point. Before the softmax calculation, the similarity matrix represents the raw similarity scores. Attention represents the attention mechanism. Softmax represents the normalized exponential function.
[0064] The peak fusion layer is used to fuse the peak weight and the attention weight to obtain a corresponding fusion weight.
[0065] Specifically, in the Peak Fusion layer, the peak weights are first expanded from the original dimensions [B, 1, L] to the expanded dimensions [B, L, L], where B (short for Batch or batch_size) represents the batch size, and L (short for Length) represents the time series length. The original peak weights [B, 1, L] represent the peak importance of each time point; the expanded peak weights, with the dimensions [B, L, L], have the same peak weights replicated for each row. This allows element-wise multiplication with the attention matrix of the same dimensions [B, L, L]. This allows the peak weights to modulate the entire attention matrix, ensuring that each time point considers the peak importance of the target time point when calculating attention relative to other time points, resulting in a higher attention weight for the peak region.
[0066] The expanded peak weight is multiplied element-by-element with the similarity matrix, so that the peak area gets a higher weight in the attention calculation. The mathematical expression is: where d k The feature dimension is 64, represents element-wise multiplication, W peak is the expanded peak weight map.
[0067] The normalization layer is used to perform normalization processing on the fusion weight to obtain the normalized fusion weight.
[0068] The feature aggregation layer is used to multiply the normalized fusion weight by the value vector to obtain corresponding weighted feature data.
[0069] The output convolution layer is used to perform convolution and residual connection on the weighted feature data to obtain attention feature data corresponding to the training sample.
[0070] Specifically, the attention weights after softmax normalization are multiplied by V to obtain the attention output, which is then integrated through a 3×3 convolution kernel. Finally, the attention output is added to the original input through a residual connection to maintain feature integrity.
[0071] Among them, in one example of the present application, four layers of peak-aware attention modules connected in sequence can be used to form a deep attention network. Each layer of peak-aware attention modules independently calculates the peak weight and attention weight, and the layers are connected through residual connections. The input of the first-layer peak-aware attention module is the multi-scale feature data (multiscale_features) output by the sample processing module, and the attention feature data output by the first-layer peak-aware attention module serves as the input of the second-layer peak-aware attention module, and so on. This multi-layer structure enables the network to gradually refine its attention to the peak area and better capture the peak evolution characteristics during the transmission process from the transmitted signal to the received signal. The final output attention feature data (attention_features) is still of dimension (batch_size, 64, 1024).
[0072] In order to further improve the reliability and effectiveness of the training process of the optical fiber transmission reception signal prediction model, it is possible to realize the automatic prediction of the reception signal in the continuous spectrum nonlinear frequency division multiplexing system, and to improve the accuracy of the reception signal prediction using the trained optical fiber transmission reception signal prediction model, in a method for training the optical fiber transmission reception signal prediction model provided in an embodiment of the present application, see Figure 2 , the neural network model also specifically includes the following contents: Peak enhancement module and output layer.
[0073] The peak enhancement module is connected to the peak-perceived attention module, and the peak enhancement module is used to perform channel compression and recovery processing on the attention feature data in sequence based on two convolutional layers to obtain corresponding peak enhancement feature data, and perform weighted residual connection on the peak enhancement feature data and the attention feature data to obtain the peak enhancement attention feature data corresponding to the training sample.
[0074] Specifically, the peak enhancement module receives the attention feature data after attention processing and further enhances the expression ability of the peak information through a specialized convolutional network. Considering the importance of peak features in the received signal for accurately modeling the fiber channel, the peak enhancement module adopts a two-layer convolution structure. The first layer compresses the 64 input channels of the attention feature data into 32 channels, and uses a 3×3 convolution kernel and a ReLU activation function to extract peak-related features. The second layer restores the 32 channels of the attention feature data to 64 channels, also using a 3×3 convolution kernel, but adopts the Tanh activation function to output peak-enhanced feature data in the range of -1 to 1. The Tanh function, also known as the hyperbolic tangent activation function, is a variation of the Sigmoid function.
[0075] The peak enhancement feature data is combined with the original attention feature data through a weighted residual connection, with the weight coefficient set to 0.3. This design not only retains the main information of the original feature, but also moderately enhances the expression of the peak feature. The specific calculation is: enhanced_features = attention_features + 0.3 × peak_enhanced Here, enhanced_features represents the peak-enhanced attention feature data; peak_enhanced is the peak-enhanced feature data output by the two convolution layers. The small weight coefficient of 0.3 ensures that the enhancement operation does not excessively alter the original features, but instead serves as a fine-tuning measure, helping the network better reconstruct the peak features in the received signal. The peak-enhanced attention feature data (enhanced_features) output by the peak enhancement module maintains a dimension of (batch_size, 64, 1024) and serves as the input to the output layer.
[0076] The output layer is used to generate received signal prediction result data for predicting the received signal formed after the training sample is transmitted to the receiving end through the optical fiber based on the peak enhanced attention feature data; wherein, the received signal prediction result data includes: the real part prediction component and the imaginary part prediction component of the received signal.
[0077] Specifically, the output layer receives the peak enhanced attention feature data (enhanced_features) and is responsible for generating the final received signal prediction, that is, the predicted received signal prediction result data. This layer uses a two-layer convolutional network structure. The first layer compresses the 64 input channels of the peak enhanced attention feature data into 32 channels, using a 5×5 convolution kernel and the LeakyReLU activation function for preliminary feature mapping. The second layer further compresses the 32 channels of the peak enhanced attention feature data into 2 channels, corresponding to the real and imaginary parts of the received signal, and also uses a 5×5 convolution kernel but no activation function to maintain the linear characteristics of the output. The leaky rectified linear unit (LeakyReLU) is a variant of the ReLU activation function.
[0078] The 2-channel tensor generated by the output layer needs to be channel-separated. Channel 0 corresponds to the real part of the prediction component (pred_real), and channel 1 corresponds to the imaginary part of the prediction component (pred_imag). The dimension of each component is (batch_size, 1024). These two prediction components correspond to the received signal q after optical fiber transmission. rxThe real and imaginary components of (t) have the same format as the real component of the actual received signal (rx_real) and the imaginary component of the received signal (rx_imag) saved in the preprocessing stage, and the loss calculation can be performed directly.
[0079] The data flow throughout the neural network model forms a complete processing chain: the real component of the input transmitted signal (tx_real) and the imaginary component of the transmitted signal (tx_imag) are first organized into input tensors. These components then undergo feature extraction in the time and frequency domains, respectively. The fused features are then processed through modules such as multi-scale processing, peak-aware attention, and peak enhancement. Finally, the output layer generates the real and imaginary predicted components of the received signal, corresponding to the received signal prediction data. This end-to-end design enables the network to learn the complete mapping from transmitted to received signals, achieving accurate modeling of the fiber channel.
[0080] In order to further improve the reliability and effectiveness of the training process of the optical fiber transmission reception signal prediction model, it is possible to realize the automatic prediction of the reception signal in the continuous spectrum nonlinear frequency division multiplexing system, and to improve the accuracy of the reception signal prediction using the trained optical fiber transmission reception signal prediction model, in a method for training the optical fiber transmission reception signal prediction model provided in an embodiment of the present application, see Figure 2 The sample processing module package in the neural network model specifically includes the following contents: Input layer, mixed domain feature extraction unit and multi-scale feature extraction unit.
[0081] The input layer is used to receive the training samples, and the training samples include: real components and imaginary components corresponding to the transmitted signal; the input layer is also used to stack the real components and imaginary components corresponding to the transmitted signal in the channel dimension to obtain a three-dimensional tensor corresponding to the training samples; wherein each dimension in the three-dimensional tensor is used to represent the batch size, the number of channels and the time series length respectively.
[0082] Specifically, the input layer receives the real and imaginary components of the transmitted signal. Each component is a one-dimensional array of length 1024, corresponding to the sample values within a 6-nanosecond time window. The main function of the input layer is to organize these two independent real number sequences into a tensor format suitable for neural network processing. Specifically, the input layer stacks the real and imaginary components of the transmitted signal along the channel dimension, forming a three-dimensional tensor of shape (batch_size, 2, 1024), where the first dimension represents the batch size, the second dimension contains the real and imaginary channels, and the third dimension is the time series length.
[0083] This tensor representation preserves the emission signal q txThe complete information of (t), including the 16-QAM modulation information of its 64 subcarriers and the phase characteristics introduced by continuous spectrum modulation, is obtained. The resulting three-dimensional tensor is fed into both the time-domain feature extraction branch and the frequency-domain feature extraction branch for parallel processing. The goal is to learn the mapping relationship from the transmitted signal to the received signal.
[0084] The mixed domain feature extraction unit is used to perform time domain feature extraction and frequency domain feature extraction on the three-dimensional tensor corresponding to the training sample to obtain time domain feature data and frequency domain feature data corresponding to the training sample, and perform feature fusion on the time domain feature data and frequency domain feature data to obtain mixed domain feature data corresponding to the training sample.
[0085] Specifically, the hybrid domain feature extraction unit includes a time domain feature extraction branch and a frequency domain feature extraction branch, and a feature fusion layer respectively connected to the time domain feature extraction branch and the frequency domain feature extraction branch.
[0086] The time-domain feature extraction branch receives the 3D tensor output from the input layer as input and extracts the signal's temporal features through two layers of one-dimensional convolutional networks. These features are crucial for capturing transient changes in the 3D tensor and pulse broadening due to dispersion. The first convolutional network layer maps the two input channels of the 3D tensor into 64 feature channels, using a convolution kernel of size 5 and padding of 2 to ensure the sequence length remains constant at 1024. The convolutional features are processed using the LeakyReLU activation function with a negative slope of 0.2 to enhance the network's nonlinear representation capabilities. The activated features have a dimension of (batch_size, 64, 1024) and contain the initially extracted time-domain feature information. The second convolutional network layer receives the output of the first layer and refines the features from 64 to 64 channels, also using a convolution kernel of size 5 and the same activation function. This layer further extracts and abstracts the time-domain features to generate a higher-level feature representation. After two layers of convolution, the time-domain branch outputs time-domain feature data (time_features) with dimensions (batch_size, 64, 1024). These time-domain feature data effectively captures the local variation patterns and timing dependencies of the signal in the temporal dimension, particularly the time-domain distortion characteristics during the transmission process from the transmitted signal to the received signal.
[0087] The frequency-domain feature extraction branch processes in parallel with the time-domain branch, also receiving data from a three-dimensional tensor. This branch first reconstructs the real and imaginary components of the three-dimensional tensor into a complex signal, recovering the complex form of the transmitted signal. The reconstructed complex signal has dimensions (batch_size, 1024). A fast Fourier transform (FFT) is then performed on the signals in each batch to obtain a frequency-domain representation. In one or more embodiments of the present application, the real component may be referred to as the real part, and the imaginary component may be referred to as the imaginary part.
[0088] To fully exploit the value of frequency domain information, the system extracts four different frequency domain features from the FFT results. The amplitude spectrum, obtained by calculating the modulus of a complex number, reveals the signal's energy distribution at different frequencies, which is important for understanding frequency-selective fading in optical fiber transmission. The phase spectrum, obtained by calculating the argument of a complex number, reflects the phase relationship between frequency components and is crucial for capturing phase variations caused by dispersion. The real and imaginary spectra, respectively, extract the real and imaginary parts of the FFT results, preserving the complete frequency domain information. These four features are each formed into a tensor of dimension (batch_size, 1, 1024), which is then concatenated along the channel dimension to form a frequency domain feature tensor of dimension (batch_size, 4, 1024). This frequency domain feature tensor is then processed through a two-layer convolutional network. The first layer maps the four input channels into 64 feature channels, while the second layer performs 64-to-64 channel feature extraction. Both layers use a convolution kernel of size 5 and the LeakyReLU activation function. The frequency domain branch finally outputs frequency domain feature data (freq_features) with a dimension of (batch_size, 64, 1024). The frequency domain feature data captures the spectral characteristics and global structure information of the signal and is extremely effective for modeling the frequency domain changes from the transmitted signal to the received signal.
[0089] The feature fusion layer receives the time-domain feature data output by the time-domain branch and the frequency-domain feature data output by the frequency-domain branch. Both are feature tensors of (batch_size, 64, 1024) dimensions. The design of the feature fusion layer takes into account the coupling between time-domain and frequency-domain effects during the transmission process from the transmitted signal to the received signal. The feature fusion layer first concatenates these two features along the channel dimension to form a joint feature representation of (batch_size, 128, 1024) dimensions. This concatenation operation preserves all information in the time and frequency domains, providing a complete input for subsequent feature integration. The concatenated 128-channel features undergo dimensionality reduction through a 1×1 convolution, reducing the number of channels back to 64. This dimensionality reduction not only reduces computational complexity but also enables effective fusion of time-domain and frequency-domain features through a learned weight matrix. The 1×1 convolution is equivalent to performing a linear transformation on the 128-dimensional feature vector at each time point, generating a 64-dimensional fused feature. The dimension of the fused mixed domain feature data (fused_features) is (batch_size, 64, 1024), which contains both the time domain and frequency domain information of the signal, providing a rich and comprehensive input representation for subsequent multi-scale feature extraction.
[0090] The multi-scale feature extraction unit is used to perform feature extraction of multiple time scales on the mixed domain feature data corresponding to the training sample to obtain multiple time scale feature data corresponding to the training sample, and perform feature fusion processing on each of the time scale feature data to obtain multi-scale feature data corresponding to the training sample.
[0091] Specifically, the multi-scale feature extraction unit receives mixed-domain feature data (fused_features) as input and captures signal features at different time scales through a parallel multi-branch convolutional structure. This design is particularly well-suited for handling multi-scale effects in optical fiber transmission: fast nonlinear phase modulation and slow dispersion-induced pulse broadening. The module features four parallel branches, each using a different convolution configuration to extract features at a specific scale. The first branch uses convolution with a kernel size of 3 to capture local details, suitable for detecting rapid changes in the transmitted signal; the second branch uses convolution with a kernel size of 5 to capture medium-range features; the third branch uses convolution with a kernel size of 7 to capture larger-scale features, suitable for modeling dispersion effects; and the fourth branch uses dilated convolution with a dilation rate of 2 and a kernel size of 3 to capture long-range dependencies. Each branch takes 64-channel mixed-domain feature data (fused_features) as input and outputs 16 channels of time-scale feature data. The outputs of the four branches are concatenated along the channel dimension to form 64-channel multi-scale feature data. This design ensures a balanced representation of features at different scales. The concatenated features are fused through a 1×1 convolution and then connected to the original input using a residual connection, which adds the fused features to the input features. Residual connections facilitate the backpropagation of gradients and improve the stability of training.
[0092] Features after residual connection It also needs to go through layer normalization. Layer normalization calculates the mean μ and variance σ² of each sample in the feature dimension, and then normalizes it according to the formula: in, Representation layer normalization; γ and β are learnable parameters, and ε is a numerical stability constant. Layer normalization helps stabilize the training process and accelerate convergence. The multi-scale feature extraction unit uses two consecutive processing layers. The output of the first layer directly serves as the input to the second layer. By stacking multiple layers, the expressive power of features is enhanced. The final output multi-scale feature data (multiscale_features) maintains the dimension (batch_size, 64, 1024) and is passed to the spike-aware attention module.
[0093] In order to further improve the reliability and effectiveness of the training process of the optical fiber transmission reception signal prediction model, it is possible to realize the automatic prediction of the reception signal in the continuous spectrum nonlinear frequency division multiplexing system, and to improve the accuracy of the reception signal prediction using the trained optical fiber transmission reception signal prediction model, in a method for training the optical fiber transmission reception signal prediction model provided in an embodiment of the present application, see Figure 4 Step 100 in the optical fiber transmission reception signal prediction model training method specifically includes the following contents: Step 110: Acquire each signal pair, wherein each signal pair includes a transmission signal pre-generated by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system and a receiving signal formed after the transmission signal is transmitted to a receiving end via an optical fiber.
[0094] Step 120: Perform quantile-based robust normalization processing on the transmitted signal and the received signal in each of the signal pairs.
[0095] Step 130: Separate each of the transmitted signals and each of the received signals after the robust normalization processing into real components and imaginary components to obtain training samples corresponding to each of the signal pairs, and use each of the received signals in each of the signal pairs as a label corresponding to each of the training samples; wherein each of the training samples contains a real component and an imaginary component of the transmitted signal; and each of the labels contains a real component and an imaginary component of the received signal belonging to the same signal pair as the transmitted signal uniquely corresponding to the label.
[0096] In order to further improve the reliability and effectiveness of the training process of the optical fiber transmission reception signal prediction model, it is possible to realize the automatic prediction of the reception signal in the continuous spectrum nonlinear frequency division multiplexing system, and to improve the accuracy of the reception signal prediction using the trained optical fiber transmission reception signal prediction model, in a method for training the optical fiber transmission reception signal prediction model provided in an embodiment of the present application, see Figure 5 Step 200 of the optical fiber transmission reception signal prediction model training method specifically includes the following contents: Step 210: In the current iteration round, the training samples are input into the neural network model so that the neural network model performs mixed domain feature extraction, multi-scale processing, and attention feature extraction based on the peak perception attention mechanism on the training samples to obtain reception signal prediction result data for predicting the reception signal formed after the training samples are transmitted to the receiving end through the optical fiber.
[0097] Step 220: In the current iteration round, based on the received signal prediction result data corresponding to the training sample and the label, a preset target loss function is used to determine the target loss corresponding to the current iteration round, and the neural network model is optimized based on the target loss.
[0098] The target loss function is composed of a real part loss function, an imaginary part loss function, an amplitude loss function, and an amplitude loss weight coefficient corresponding to the amplitude loss function; The real part loss function is used to represent the product between a first mean square error and a peak weighted term; wherein the first mean square error includes the mean square error between the real part prediction component corresponding to the received signal prediction result data and the real part component of the received signal corresponding to the label; the peak weighted term includes the sum of 1 and a peak product term; the peak product term includes the product of a preset peak weighting factor and a peak mask; The imaginary part loss function is used to represent the product between the second mean square error and the peak weighted term; wherein the second mean square error includes the mean square error between the imaginary part prediction component corresponding to the received signal prediction result data and the imaginary part component of the received signal corresponding to the label; The amplitude loss function is used to represent the product between the third mean square error and the peak weighted term; wherein the third mean square error includes the mean square error between the signal amplitude corresponding to the pre-acquired received signal prediction result data and the signal amplitude of the received signal corresponding to the pre-acquired label.
[0099] Specifically, the loss function module receives the real prediction component (pred_real) and imaginary prediction component (pred_imag) of the received signal generated by the network output layer, as well as the real component (rx_real) and imaginary component (rx_imag) of the actual received signal provided by the data preprocessing stage. Considering the received signal q rx (t) compared to the transmitted signal q tx (t) The peak characteristics generated during the transmission process. The first step in loss calculation is to identify the peak area based on the actual received signal. First, calculate the complex amplitude of the actual signal and obtain the signal amplitude value at each time point using the following formula: Here, true_mag represents the amplitude of the actual received signal, and Sqrt represents the square root function. We then search for the maximum value along the time dimension of the batch to obtain the maximum amplitude (max_mag) of each sample.
[0100] Peak regions are identified using a thresholding method, marking the time points where the amplitude exceeds 70% of the maximum value as peak regions. This threshold is chosen based on an understanding of fiber transmission characteristics: in CS-NFDM systems, nonlinear effects can cause significant power concentrations in the signal at certain moments. These high-power regions are crucial for accurate modeling of channel characteristics. The peak mask (M_peak) is defined using the following rule: if the signal amplitude at time t is greater than (maximum amplitude × 0.7), then the peak mask M_peak(t) at that time point is 1; otherwise, M_peak(t) is 0.
[0101] The specific calculation steps are: (1) Calculate the amplitude of the real received signal true_mag = sqrt(rx_real² + rx_imag²); (2) Find the maximum amplitude value max_mag = max(true_mag); max represents the maximum value; (3) Set the peak threshold threshold = 0.7 × max_mag; (4) Generate a binary mask: When true_mag > threshold, the peak mask M_peak = 1 (indicating peak regions), and when true_mag ≤ threshold, the peak mask M_peak = 0 (indicating non-peak regions). This relative threshold design can adapt to the dynamic range of different signals and ensure the robustness of peak detection. The peak mask generated in this way is a binary tensor with the same length as the signal, which is used to identify which regions require additional weight attention in the loss calculation.
[0102] Among them, the target loss function adopts a multi-component design to calculate the real part loss, imaginary part loss and amplitude loss respectively, and comprehensively evaluate the prediction quality from the transmitted signal to the received signal. Before calculating the loss of each component, it is necessary to first calculate the amplitude of the predicted received signal. This amplitude will be used in the calculation of the amplitude loss and also reflects the overall quality of the predicted signal.
[0103] Real part loss function L real It is obtained by calculating the mean squared error between the predicted real part and the true real part, but an innovation lies in the introduction of a peak weighting mechanism. The specific calculation process is: first, the basic MSE error is calculated, and then the error is multiplied by (1 + λ_peak × M_peak), where λ_peak = 3.0, which is the peak weighting factor. This means that the error in the peak region is amplified by 4 times (1 + 3), while the non-peak region retains the original weight. This design ensures that when the network learns the mapping from transmitted to received signals, it pays special attention to the peak regions that contain important transmitted information. Finally, the weighted errors are averaged to obtain the real part loss.
[0104] Real part loss function L real The mathematical expression is: Where MSE stands for Normalized Mean Square Error. represents the real part prediction component corresponding to the received signal prediction result data; r represents the real part of the received signal corresponding to the label; represents the first mean square error; Represents the peak weighting term. mean represents the mean function that can be used to calculate the average value.
[0105] Component imaginary part loss L imag is calculated in exactly the same way as the real loss, except that the real part is replaced by the imaginary part: in, represents the imaginary prediction component corresponding to the received signal prediction result data; i represents the imaginary component of the received signal corresponding to the label.
[0106] Amplitude loss function L imag Compute the weighted mean square error between the predicted and true magnitudes: in, represents the signal amplitude corresponding to the received signal prediction result data; q represents the signal amplitude of the received signal corresponding to the label.
[0107] The three component losses are weighted summed to obtain the target loss L total : Where α = 0.5 is the amplitude loss weight coefficient. This multi-component design ensures that the network optimizes the prediction accuracy of the real part, imaginary part, and amplitude at the same time, while the peak weighting strategy makes the network pay more attention to the accuracy of the peak area during training, which is crucial for accurately modeling the nonlinear mapping relationship from the transmitted signal to the received signal. The calculated target loss L total Is a scalar value that will be used as the starting point for backpropagation to calculate the gradients of the network parameters.
[0108] In order to further illustrate the above embodiment, the present application also provides a specific application example of a method for training a prediction model of an optical fiber transmission reception signal.
[0109] The existing technology has the following problems: (1) The general Transformer architecture is not specifically optimized for the characteristics of optical fiber communication signals; Signal property mismatch: Fiber optic signals have specific amplitude, phase, and frequency characteristics, but general-purpose Transformers cannot effectively model these physical properties. Data preprocessing requirements: There is a lack of specialized preprocessing modules for optical fiber signals, such as amplitude normalization and phase unwrapping; Loss function design: Traditional MSE loss cannot fully reflect the key quality indicators of optical fiber signals (such as bit error rate and signal-to-noise ratio).
[0110] (2) The multi-head self-attention mechanism does not consider the special importance of the peak area of the communication signal; Peak information loss: The standard attention mechanism tends to ignore key peak points in the signal, resulting in the averaging of important information; Unreasonable weight distribution: Multi-head attention cannot adaptively assign higher weights to peak areas, affecting key feature extraction; Insufficient time domain feature modeling: There is a lack of a dedicated modeling mechanism for the signal’s time domain peak pattern.
[0111] (3) Lack of specialized modeling mechanisms for the physical characteristics of signals; Insufficient complex signal processing capabilities: Standard Transformers have difficulty effectively processing the amplitude and phase information of complex signals; Lack of nonlinear distortion modeling: Unable to model nonlinear effects in optical fiber transmission (such as self-phase modulation, cross-phase modulation, etc.); Insufficient utilization of frequency domain features: lack of joint modeling capabilities in frequency and time domains.
[0112] (4) Only processes serialized one-dimensional signal data; (5) Failure to fully utilize the complementary characteristics of the signal in the time domain and frequency domain; (6) Lack of specialized processing for the amplitude and phase of complex signals.
[0113] In order to solve the above technical problems, the optical fiber transmission reception signal prediction model training method provided in the application example of this application specifically includes the following contents: 1. Overview of Continuous Spectrum Nonlinear Frequency Division Multiplexing System In an ideal lossless optical fiber, the transmission evolution of the continuous spectrum follows a certain mathematical law. When the signal propagates a distance L in the optical fiber, the change of its continuous spectrum coefficient is It can be expressed as: Among them, q c represents the continuous spectrum coefficient; exp represents the natural exponential function; i represents the imaginary unit; ξ is the nonlinear frequency variable. This evolution law shows that the continuous spectrum only undergoes changes related to ξ during transmission. 2 The phase rotates proportionally while the amplitude remains unchanged, which provides a theoretical basis for reliable communication in nonlinear channels.
[0114] This application example uses a multi-carrier modulation scheme with 64 subcarriers, supporting the 16-QAM modulation format, to achieve efficient data transmission within a 34GHz bandwidth. The system effectively manages the signal's evolution during transmission by performing phase adjustments at both the transmitter and receiver.
[0115] (2) Transmit Signal Generation 1. Data source processing and 16-QAM constellation mapping First, receive the raw binary data stream {b i}, where i=1,2,...,N bits , respectively representing different bit numbers. According to the requirements of the 16-QAM modulation format, the system groups the continuous original binary data stream into groups of 4 bits. Each group of bits is mapped to a complex modulation symbol, so that each symbol carries 4 bits of information.
[0116] QAM constellation mapping generates complex symbols , where m is the data block index and k is the subcarrier index, The system design uses 64 parallel subcarriers, so the value of k ranges from -32 to +31. The in-phase component I of the constellation diagram k and the quadrature component Q k Each symbol is selected from the normalized amplitude set {−3, −1, +1, +3}, forming a constellation of 16 equally spaced points. This constellation design ensures a good minimum Euclidean distance between symbols, which facilitates reliable detection at the receiver.
[0117] 2. Continuous spectrum modulation processing The continuous spectrum modulation module maps the constellation to the symbol sequence Modulated to the nonlinear frequency domain. The system uses a raised cosine carrier waveform with a roll-off factor of α = 0.5, which not only ensures the orthogonality between subcarriers but also has good spectral characteristics. The center frequency of each subcarrier is set to , where the basic time parameter T0 = 2×10 −9 Seconds determine the carrier interval.
[0118] Multi-carrier modulation is realized by superposition principle, the initial continuous spectrum The construction expression is: Here, h(ξ) represents the raised cosine carrier waveform function. This summation process combines the symbol information carried by the 64 subcarriers into a continuous nonlinear spectral function. To optimize system performance, the modulated signal needs to be power-adjusted by multiplying it by a power control factor A=3: in, Represents the continuous spectrum after power control; the power control parameters are determined based on the fiber's nonlinear threshold and system performance requirements. The system performs 50% phase pre-compensation at the transmitter to optimize the signal's transmission characteristics in the fiber. Pre-compensation is achieved by applying a phase factor: in, =81.3×10 3 This phase pre-compensation is an effective compensation technology for CS-NFDM system transmission.
[0119] 3. Inverse Nonlinear Fourier Transform (INFT) Modulated continuous spectrum It needs to be converted to a time domain signal through an inverse nonlinear Fourier transform. INFT is based on the Zakharov-Shabat inverse scattering theory and recovers the time domain envelope from the nonlinear spectrum by solving the corresponding Riemann-Hilbert problem. The mathematical relationship of this transform is expressed as: in, represents the time-domain envelope recovered from the nonlinear spectrum. The numerical implementation uses the Fast Nonlinear Fourier Transform (FNFT) algorithm, using the fourth-order Runge-Kutta method to solve the associated differential equations. The time-domain signal is discretized using 512 sampling points, covering a time window of (−3 ns, +3 ns) to ensure the complete 6-nanosecond block duration.
[0120] 4. Transmit filtering and signal conditioning The time domain signal output by INFT needs to be filtered to meet the system bandwidth requirements. The transmit filter adopts an ideal low-pass characteristic, and its frequency domain transfer function is for: Where B = 34 GHz is the system design bandwidth, f represents the frequency, and rect() represents the rectangular window function. The filtering process effectively limits the signal bandwidth and prevents interference between adjacent channels. The filtered signal is denormalized to convert the normalized units used in the calculation to physical units: in, Represents the time domain signal at the transmitting end before inputting the optical fiber; represents the filtered time domain signal; the time normalization factor T scale =4×10−10 seconds, power normalization factor P scale Determined according to the system transmission power. The transmitted signal q obtained after processing tx (t) It has all the necessary characteristics for transmission in optical fiber.
[0121] (3) Fiber Optic Transmission Channel In optical fiber transmission channels, signal propagation follows the nonlinear Schrödinger equation (NLSE), which comprehensively describes various physical effects in optical fibers. The loss coefficient is 0.2×10 −3 / m, corresponding to an optical fiber loss of 0.087 dB / km; the group velocity dispersion coefficient is −5.75×10 −27 sec² / m, characterizing the anomalous dispersion characteristics; the nonlinear coefficient is 1.6×10 −3 / (W·m), describes the strength of Kerr nonlinearity.
[0122] Because the NLSE equation contains both linear and nonlinear terms, and these terms do not commutatively satisfy the law, analytical solutions are unavailable and require numerical methods. The system employs the split-step Fourier method (SSFM) to numerically solve the NLSE. The key idea of this method is to rewrite the NLSE equation in operator form and then apply a symmetric split-step approximation within each step.
[0123] The propagation process is decomposed into three steps: the first half of linear propagation, the complete nonlinear propagation and the second half of linear propagation.
[0124] In one example, the 81.3 km transmission distance is divided into 40 computational segments, each with a length of Δz = 2.0325 km. The linear step addresses loss and dispersion effects. Because linear operators have a diagonal form in the frequency domain, performing computations in the frequency domain offers greater efficiency and accuracy. The linear propagation process first converts the time-domain signal to the frequency domain using a fast Fourier transform (FFT), then applies the corresponding transfer function to each frequency component, and finally returns to the time domain using an inverse fast Fourier transform (IFFT). The nonlinear step, performed directly in the time domain, addresses self-phase modulation effects. Because the nonlinear operator depends only on the instantaneous power of the signal and is independent of its time derivative, when the step size is sufficiently small, the power can be assumed to remain constant within the segment, resulting in an analytical solution. Nonlinear propagation causes the signal to undergo a phase change proportional to its instantaneous power. This power-dependent phase modulation is a direct manifestation of the Kerr effect in optical fibers. In the numerical implementation, the signal power is calculated at each time sampling point, and the corresponding nonlinear phase modulation is then applied. The complete step-by-step algorithm follows a symmetrically distributed scheme within each computational segment: a first half of linear propagation handles half the loss and dispersion effects, a full nonlinear propagation handles the self-phase modulation for the entire step length, and a second half of linear propagation handles the remaining loss and dispersion effects. The signal propagates from fiber entry to exit through 40 iterations, completing the entire transmission process.
[0125] After complete optical fiber transmission, the output includes the receiving signal q of various transmission effects rx(t). At this time, the transmitted signal q can be collected tx (t) and the received signal q rx (t) is used as a training dataset.
[0126] (IV) Dataset Construction First, construct a set of 1200 pairs of transmitting signals q tx (t) and the received signal q at the receiving end rx (t) is a training data set of data samples, which are generated by the aforementioned CS-NFDM (Continuous Spectrum Nonlinear Frequency Division Multiplexing) system. Each pair of data samples represents a complete signal transmission process.
[0127] The fiber link parameters include: (1) The loss coefficient is 0.2×10⁻³ m⁻¹; (2) The dispersion coefficient is -5.75×10⁻² 7 s² / m; (3) The nonlinear coefficient is 1.6×10⁻³ (W·m)⁻¹; (4) Fiber span length: 81.3 km; (5) Number of spans: 1 span; (6) Center frequency: 193.1 THz; (7) Time normalization scale: T_scale = 4×10-10 s.
[0128] The dataset is divided into training, validation, and test sets in an 8:2:2 ratio. Specifically, the training set consists of 800 pairs of data samples for network parameter learning, the validation set of 200 pairs of data samples for performance monitoring during training, and the test set of 200 pairs of data samples for final model evaluation. Each data sample contains a complex signal of length 1024, sampling time information, and the corresponding signal amplitude value.
[0129] In one example, the transmission signal is as follows: Figure 6 As described, the received signal is as follows Figure 7 shown.
[0130] (V) Signal preprocessing and normalization During training, the training signal needs to be preprocessed. Figure 8 , the data preprocessing module receives the paired transmission signal q tx (t) and the received signal q rx (t) as input, both are time domain signals in complex form. This application example adopts a quantile-based robust normalization method, first calculating the signal amplitude A = |q(t)|, and then determining the 1st percentile q1 and the 99th percentile q99 As the normalized boundary. The mathematical expression of the normalization process is: Where A is the original signal amplitude, is the signal phase, A norm is the normalized amplitude; represents the normalized complex time-domain signal; e represents the natural exponential function. The normalized amplitude is clipped to the range (0, 1.5). This design allows the peak signal to slightly exceed 1, effectively preserving important peak feature information.
[0131] The processed complex signal is separated into two independent real-valued sequences, the real and imaginary parts, each with a length of 1024. This separation strategy enables the neural network to independently process the two orthogonal components of the complex signal while simplifying subsequent tensor operations. The preprocessing module also saves the normalization parameters (q1, q99, and maximum amplitude), which are used for denormalization of the signal after training. The separated real and imaginary sequences serve as the input and target output of the neural network.
[0132] (6) Training process 1. Training data processing and batch loading The training process uses the 1200 signal pairs constructed above, performing iterative learning in batches. The data loader reads 16 pairs of training samples at a time, ensuring the correct pairing of the sender and receiver files. The loaded data includes the real and imaginary parts of the transmitted and received signals, as well as normalization parameters. After being transferred to the GPU, these tensors maintain the dimension [16, 1024], representing the batch data of 16 signal pairs.
[0133] 2. Forward propagation and loss calculation The forward propagation takes the transmitted signal as input and, through modules such as hybrid domain feature extraction, multi-scale processing, and peak-aware attention, generates a prediction of the received signal. The prediction and the actual received signal are fed into a peak-aware loss function to calculate a weighted loss that includes the real part, imaginary part, and amplitude. This loss function specifically emphasizes peak regions, enabling the network to focus on learning key peak features during the transmission process from the transmitted signal to the received signal.
[0134] 3. Parameter optimization and learning rate scheduling The network uses the AdamW optimizer for parameter updates, with the initial learning rate set to 0.001 and the weight decay coefficient to 1×10^(-4). To prevent gradient explosion in deep networks, the system limits the L2 norm of the gradient to less than 0.5. The learning rate scheduling adopts the learning rate plateau decay (ReduceLROnPlateau) strategy; when the verification loss does not improve for two consecutive epochs, the learning rate is multiplied by 0.7 for decay, and the minimum learning rate is limited to 1×10^(-6). The AdamW optimizer is an improved version of the Adam optimizer. It solves the generalization defects of the traditional Adam by decoupling the weight decay mechanism and has become the default optimizer for current large model training. Parameter updates follow the AdamW algorithm: in, represents the model parameters at the tth training step; represents the updated model parameters at the t+1th training step; represents the first-order moment estimate after bias correction at step t; represents the second-order moment estimate after bias correction at step t; η is the learning rate, and λ is the weight decay coefficient, which ensures that the model maintains good generalization ability while learning the mapping from transmitted signals to received signals; It is a numerical stability constant in the AdamW algorithm to prevent the denominator from being zero and is usually set to 1×10 -8 .
[0135] Among them, convolutional neural networks (CNNs) and peak-aware attention can be replaced by Transformer encoders and improved multi-head attention mechanisms; specific replacement implementations include: position encoding can replace convolutional feature extraction with learnable position encoding; multi-head attention improvement can be achieved by incorporating peak weight modulation into standard multi-head attention; feedforward network adaptation can use MLP instead of convolutional layers for feature transformation.
[0136] Alternatively, replace pure CNN architectures for time series modeling with CNN feature extraction and LSTM / GRU time series modeling. LSTM stands for Long Short-Term Memory, and GRU stands for Gated Recurrent Unit. Specific implementations include: retaining the hybrid-domain CNN feature extraction layer; replacing multi-scale convolution with bidirectional LSTM in the time series modeling layer; and applying peak-aware attention to the LSTM output in the attention layer.
[0137] 4. Early Stopping Mechanism and Model Selection The training process uses an early stopping mechanism to prevent overfitting. The system monitors the loss changes on the validation set, such as Figure 9As shown in the figure, the blue line represents the training loss, and the orange line represents the calibration loss. When the validation loss does not improve for eight consecutive epochs, training is automatically stopped and the model parameters with the best validation performance are restored. This ensures that the final model has the optimal Qtx to Qrx mapping capability. All of the above training steps result in a fast training process.
[0138] Based on this, the complete process of the optical fiber transmission receiving signal prediction model training method provided by the application example is as follows: Figure 10 shown.
[0139] (VII) Performance evaluation indicators Normalized mean square error (NMSE) Normalized Mean Square Error (NMSE) is a core metric for evaluating the accuracy of neural network channel modeling. This metric calculates the mean squared error between the predicted signal and the true signal and normalizes the true signal power, enabling a consistent evaluation of signals of varying power levels. NMSE is particularly suitable for evaluating the end-to-end prediction accuracy from Qtx to Qrx, as it objectively reflects the network's ability to model fiber transmission effects.
[0140] The mathematical definition of normalized mean square error NMSE is: Where: q pred (t) represents the received signal prediction result data predicted by the network; q true (t) represents the actual received signal; E[] represents the statistical expectation operation, which is achieved through time averaging in actual calculations.
[0141] The smaller the NMSE value, the higher the prediction accuracy. The advantage of this metric is that its normalized nature eliminates the impact of signal power variations, allowing direct comparison of modeling performance under different transmission conditions.
[0142] (8) Test results Randomly select a Figure 11 The transmission signal shown in the figure is input into the model for testing. The actual receiving signal corresponding to the transmission signal is as follows Figure 12 As shown, the received signal prediction result data obtained by the optical fiber transmission received signal prediction model is as follows Figure 13 As shown, the comparison between the actual received signal and the received signal prediction result data is shown in Figure 14 As shown, it can be seen in Figure 14 In the example, the actual received signal and the received signal prediction result data are substantially consistent. Figures 11 to 13 The vertical axis represents the signal amplitude, and the horizontal axis represents the time (unit: nanosecond ns). Figure 14 The blue solid line in the figure represents the actual received signal, and the orange dotted line represents the received signal prediction result data.
[0143] The NMSE training trend is as follows Figure 15 shown.
[0144] Also, see Figure 16 The test results also include comparative analysis data on the inference time between the Split-Step Fourier Transform (SSFM) algorithm and the Peak Perception Hybrid Domain NN: Inference time refers to the time required for the model to receive input data and generate output prediction results. In the fiber channel modeling task, inference time specifically refers to the time from the input transmission signal q tx (t) to output received signal q rx The computation time (t) is the time it takes to infer a model. Inference time is a key metric for evaluating a model's real-time performance and directly impacts the system's practical application value. The shorter the inference time, the higher the system's transmission efficiency.
[0145] In other words, the embodiments and application examples of this application provide: (1) Peak-aware attention mechanism, designing a dual-branch parallel architecture: peak detection branch and self-attention branch; (2) Hybrid domain feature extraction architecture: time domain branch, extracting local time features; frequency domain branch: FFT transformation and four-way feature extraction (amplitude spectrum, phase spectrum, frequency domain real part, frequency domain imaginary part); intelligent fusion strategy, learning the optimal weight combination of time and frequency domain features; residual connection to maintain the integrity of the original information.
[0146] (3) Combined with the multi-scale CNN processing module: four-branch parallel design; branch output channels are unified into 16 dimensions, totaling 64-dimensional multi-scale features; 1×1 convolution fusion, residual connection and layer normalization stabilized training strategy.
[0147] In other words, this application example addresses the lack of domain-specific design in transformer solutions: it specifically designs a peak-aware attention mechanism that transforms the physical importance of peak regions in communication systems into network attention weights, achieving a deep fusion of domain knowledge and deep learning. To address the issue of incomplete feature extraction, it designs a hybrid-domain feature extraction architecture. The time-domain branch captures instantaneous characteristics, while the frequency-domain branch extracts four frequency-domain features (amplitude spectrum, phase spectrum, real part, and imaginary part) through FFT, achieving a complete representation of the signal. To address the issue of single-scale modeling, it designs a four-branch parallel multi-scale architecture, using 3×1, 5×1, and 7×1 convolution kernels and dilated convolutions, respectively, to simultaneously capture channel effects at different time scales, such as instantaneous changes, inter-symbol interference, memory effects, and long-range dependencies. To address the issue of high parameter sensitivity, it uses a data-driven learning approach to automatically learn channel characteristics from a large number of input-output signal pairs, eliminating the need for precise physical parameters and significantly improving system robustness.
[0148] The application examples of this application can realize the intelligent modeling of the peak-aware attention mechanism: focusing on and optimizing the signal peak area, and significantly reducing the peak error; making full use of the complementary information of the time and frequency domains: simultaneously extracting and fusing the time domain and frequency domain features to improve the modeling integrity; supporting multi-scale time-dependent modeling: simultaneously capturing the channel effects of different time scales; realizing real-time and efficient modeling: improving the inference speed. The serial characteristics of SSFM limit the acceleration potential. This technical solution is more than 100 times faster than the traditional SSFM method.
[0149] Based on the above-mentioned embodiment and application example of the optical fiber transmission received signal prediction model training method, the present application also provides an embodiment of an optical fiber transmission received signal prediction method, and the optical fiber transmission received signal prediction method specifically includes the following contents: Step 300: Input the transmission signal currently generated by the transmitting end in the continuous spectrum nonlinear frequency division multiplexing system into the optical fiber transmission receiving signal prediction model, so that the optical fiber transmission receiving signal prediction model outputs corresponding receiving signal prediction result data; wherein, the optical fiber transmission receiving signal prediction model is pre-trained based on the optical fiber transmission receiving signal prediction model training method provided in the aforementioned embodiments and / or application examples.
[0150] From a software perspective, the present application also provides a fiber optic transmission received signal prediction model training device for executing all or part of the fiber optic transmission received signal prediction model training method, and the fiber optic transmission received signal prediction model training device specifically includes the following contents: A training data construction module is configured to generate training samples corresponding to a transmission signal pre-generated by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system, and to generate labels corresponding to the training samples based on a received signal formed by pre-transmitting the transmission signal to a receiving end via an optical fiber; A model training module is used to train a neural network model based on the training samples, so that the neural network model performs mixed domain feature extraction, multi-scale processing, and attention feature extraction based on the peak perception attention mechanism on the training samples to obtain reception signal prediction result data for predicting the reception signal formed after the training samples are transmitted to the receiving end through the optical fiber, and optimizes the neural network model based on the reception signal prediction result data corresponding to the training samples and the labels, so as to train the neural network model into an optical fiber transmission reception signal prediction model for outputting the reception signal prediction result data corresponding to the transmission signal. The embodiment of the optical fiber transmission reception signal prediction model training device provided in the present application can be specifically used to execute the processing flow of the embodiment of the optical fiber transmission reception signal prediction model training method in the above-mentioned embodiment. Its function will not be repeated here, and reference can be made to the detailed description of the embodiment of the above-mentioned optical fiber transmission reception signal prediction model training method.
[0151] The portion of the optical fiber transmission received signal prediction model training apparatus that performs optical fiber transmission received signal prediction model training can be performed in a server or client device. The specific selection can be based on the processing capabilities of the client device and the limitations of the user's usage scenario. This application is not limited to this. If all operations are performed in the client device, the client device may further include a processor for specific processing of the optical fiber transmission received signal prediction model training.
[0152] The aforementioned client device may include a communication module (i.e., a communication unit) capable of establishing a communication connection with a remote server to facilitate data transmission with the server. The server may include a server at the task scheduling center or, in other implementation scenarios, a server on an intermediate platform, such as a server on a third-party server platform that is communicatively linked to the task scheduling center server. The server may comprise a single computer device, a server cluster consisting of multiple servers, or a distributed server configuration.
[0153] The server and the client device may communicate using any suitable network protocol, including network protocols that have not yet been developed as of the filing date of this application. Examples of such network protocols include TCP / IP, UDP / IP, HTTP, and HTTPS. Furthermore, examples of such network protocols include RPC (Remote Procedure Call Protocol) and REST (Representational State Transfer) protocols, which are used on top of the aforementioned protocols.
[0154] From the above description, it can be seen that the optical fiber transmission receiving signal prediction model training device provided in the embodiment of the present application provides a training method for the prediction model of the receiving signal in the continuous spectrum nonlinear frequency division multiplexing system, which can effectively improve the reliability and effectiveness of the optical fiber transmission receiving signal prediction model training process, and can realize the prediction of the receiving signal in the continuous spectrum nonlinear frequency division multiplexing system, and can improve the accuracy of the received signal prediction using the trained optical fiber transmission receiving signal prediction model, thereby being able to predict the transmission performance of the receiving signal at different environmental distances in advance during the modeling process of the high-fiber communication system, shortening the R&D cycle of the optical fiber communication system modeling process and reducing the computing resource consumption of the R&D equipment, effectively avoiding the waste of modeling experiment costs, and providing a more effective basis for the design, performance prediction and optimization of the optical fiber communication system, so as to improve the efficiency and transmission performance of the optical fiber communication system constructed according to the modeling results of the optical fiber communication system.
[0155] From a software perspective, the present application further provides an optical fiber transmission received signal prediction device for executing all or part of the optical fiber transmission received signal prediction method. The optical fiber transmission received signal prediction device specifically includes the following contents: A model prediction module is used to input the transmission signal currently generated by the transmitting end in the continuous spectrum nonlinear frequency division multiplexing system into the optical fiber transmission receiving signal prediction model, so that the optical fiber transmission receiving signal prediction model outputs corresponding receiving signal prediction result data; wherein, the optical fiber transmission receiving signal prediction model is pre-trained based on the optical fiber transmission receiving signal prediction model training method provided in the aforementioned embodiment.
[0156] The present application also provides an electronic device that may include a processor, a memory, a receiver, and a transmitter. The processor is configured to execute the optical fiber transmission received signal prediction model training method and / or the optical fiber transmission received signal prediction method described in the above embodiments. The processor and the memory may be connected via a bus or other means, with bus connection being used as an example. The receiver may be connected to the processor and the memory via a wired or wireless means.
[0157] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.
[0158] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the optical fiber transmission received signal prediction model training method and / or the optical fiber transmission received signal prediction method in the embodiments of the present application. The processor executes the non-transitory software programs, instructions, and modules stored in the memory to perform various processor functions and data processing, thereby implementing the optical fiber transmission received signal prediction model training method and / or the optical fiber transmission received signal prediction method in the above-mentioned method embodiments.
[0159] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0160] The one or more modules are stored in the memory, and when executed by the processor, perform the optical fiber transmission reception signal prediction model training method in the embodiment.
[0161] In some embodiments of the present application, the user equipment may include a processor, a memory and a transceiver unit, and the transceiver unit may include a receiver and a transmitter. The processor, memory, receiver and transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.
[0162] As an implementation method, the functions of the receiver and transmitter in this application can be considered to be implemented through a transceiver circuit or a dedicated transceiver chip, and the processor can be considered to be implemented through a dedicated processing chip, a processing circuit or a general-purpose chip.
[0163] As another implementation method, it is possible to use a general-purpose computer to implement the server provided in the embodiments of the present application. That is, the program code for implementing the functions of the processor, receiver, and transmitter is stored in a memory, and the general-purpose processor implements the functions of the processor, receiver, and transmitter by executing the code in the memory.
[0164] The present application also provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the steps of the aforementioned optical fiber transmission reception signal prediction model training method and / or optical fiber transmission reception signal prediction method. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.
[0165] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the aforementioned optical fiber transmission received signal prediction model training method and / or optical fiber transmission received signal prediction method.
[0166] It should be understood by those skilled in the art that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether it is implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted on a transmission medium or communication link via a data signal carried in a carrier.
[0167] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.
[0168] In this application, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.
[0169] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Those skilled in the art will appreciate that various modifications and variations of the present embodiment are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for training a prediction model for optical fiber transmission reception signals, characterized in that: include: Generating a training sample corresponding to a transmission signal pre-generated by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system, and generating a label corresponding to the training sample based on a received signal formed after the transmission signal is pre-transmitted to a receiving end via an optical fiber; A neural network model is trained based on the training samples, so that the neural network model performs mixed domain feature extraction, multi-scale processing, and attention feature extraction based on the peak perception attention mechanism on the training samples to obtain reception signal prediction result data for predicting the reception signal formed after the training samples are transmitted to the receiving end through the optical fiber, and the neural network model is optimized based on the reception signal prediction result data corresponding to the training samples and the labels, so as to train the neural network model as an optical fiber transmission reception signal prediction model for outputting the reception signal prediction result data corresponding to the transmission signal.
2. The optical fiber transmission reception signal prediction model training method according to claim 1, characterized in that: The neural network model includes a sample processing module and a peak perception attention module; The sample processing module is used to obtain the three-dimensional tensor corresponding to the training sample, and perform hybrid domain feature extraction and multi-scale processing on the three-dimensional tensor corresponding to the training sample to obtain multi-scale feature data corresponding to the training sample; The peak-aware attention module is used to perform peak detection on the multi-scale feature data corresponding to the training sample to obtain the peak weight corresponding to the training sample, and perform attention detection on the multi-scale feature data to obtain the attention weight corresponding to the training sample; The peak weight and the attention weight are fused to obtain a corresponding fusion weight, and the fusion weight is normalized and feature-aggregated to obtain attention feature data corresponding to the training sample for generating the received signal prediction result data.
3. The optical fiber transmission reception signal prediction model training method according to claim 2, characterized in that: The peak-aware attention module includes: a peak detection branch, a standard detection branch, a peak fusion layer, a normalization layer, a feature aggregation layer, and an output convolution layer; the peak detection branch and the standard detection branch are both connected to the peak fusion layer; the peak fusion layer, the normalization layer, the feature aggregation layer, and the output convolution layer are connected in sequence; The peak detection branch is used to perform peak detection on the multi-scale feature data corresponding to the training sample to obtain the peak weight corresponding to the training sample; The standard detection branch is used to perform attention detection on the multi-scale feature data corresponding to the training sample based on the query vector, the key vector and the value vector to obtain the attention weight corresponding to the training sample; The peak fusion layer is used to fuse the peak weight and the attention weight to obtain a corresponding fusion weight; The normalization layer is used to normalize the fusion weight to obtain the normalized fusion weight; The feature aggregation layer is used to multiply the normalized fusion weight and the value vector to obtain corresponding weighted feature data; The output convolution layer is used to perform convolution and residual connection on the weighted feature data to obtain attention feature data corresponding to the training sample.
4. The optical fiber transmission reception signal prediction model training method according to claim 2, characterized in that: The neural network model also includes: a peak enhancement module and an output layer; The peak enhancement module is connected to the peak perception attention module, and the peak enhancement module is used to perform channel compression and recovery processing on the attention feature data in sequence based on two convolutional layers to obtain corresponding peak enhancement feature data, and perform weighted residual connection on the peak enhancement feature data and the attention feature data to obtain peak enhancement attention feature data corresponding to the training sample; The output layer is used to generate received signal prediction result data for predicting the received signal formed after the training sample is transmitted to the receiving end through the optical fiber based on the peak enhanced attention feature data; wherein, the received signal prediction result data includes: the real part prediction component and the imaginary part prediction component of the received signal.
5. The optical fiber transmission reception signal prediction model training method according to claim 2, characterized in that: The sample processing module includes: an input layer, a mixed domain feature extraction unit and a multi-scale feature extraction unit; The input layer is used to receive the training sample, wherein the training sample includes: a real component and an imaginary component corresponding to the transmission signal; the input layer is further used to stack the real component and the imaginary component corresponding to the transmission signal in the channel dimension to obtain a three-dimensional tensor corresponding to the training sample; wherein each dimension in the three-dimensional tensor is used to represent the batch size, the number of channels, and the time series length respectively; The mixed domain feature extraction unit is used to perform time domain feature extraction and frequency domain feature extraction on the three-dimensional tensor corresponding to the training sample to obtain time domain feature data and frequency domain feature data corresponding to the training sample, and perform feature fusion on the time domain feature data and the frequency domain feature data to obtain mixed domain feature data corresponding to the training sample; The multi-scale feature extraction unit is used to perform feature extraction of multiple time scales on the mixed domain feature data corresponding to the training sample to obtain multiple time scale feature data corresponding to the training sample, and perform feature fusion processing on each of the time scale feature data to obtain multi-scale feature data corresponding to the training sample.
6. The optical fiber transmission reception signal prediction model training method according to claim 1, characterized in that: Generating a training sample corresponding to a transmission signal pre-generated by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system, and generating a label corresponding to the training sample based on a received signal formed after the transmission signal is pre-transmitted to a receiving end via an optical fiber, includes: Acquire signal pairs, wherein each signal pair includes a transmission signal pre-generated by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system and a receiving signal formed by transmitting the transmission signal to a receiving end via an optical fiber; performing quantile-based robust normalization processing on the transmitted signal and the received signal in each of the signal pairs; Each of the transmitted signals and each of the received signals after the robust normalization processing is separated into real components and imaginary components to obtain training samples corresponding to each of the signal pairs, and each of the received signals in each of the signal pairs is used as a label corresponding to each of the training samples; wherein each of the training samples contains a real component and an imaginary component of the transmitted signal; and each of the labels contains a real component and an imaginary component of the received signal belonging to the same signal pair as the transmitted signal uniquely corresponding to the label.
7. The optical fiber transmission reception signal prediction model training method according to claim 4, characterized in that: The optimizing the neural network model based on the received signal prediction result data corresponding to the training sample and the label includes: In the current iteration round, based on the received signal prediction result data corresponding to the training sample and the label, a preset target loss function is used to determine the target loss corresponding to the current iteration round, and the neural network model is optimized based on the target loss; The target loss function is composed of a real part loss function, an imaginary part loss function, an amplitude loss function, and an amplitude loss weight coefficient corresponding to the amplitude loss function; The real part loss function is used to represent the product between a first mean square error and a peak weighted term; wherein the first mean square error includes the mean square error between the real part prediction component corresponding to the received signal prediction result data and the real part component of the received signal corresponding to the label; the peak weighted term includes the sum of 1 and a peak product term; the peak product term includes the product of a preset peak weighting factor and a peak mask; The imaginary part loss function is used to represent the product between the second mean square error and the peak weighted term; wherein the second mean square error includes the mean square error between the imaginary part prediction component corresponding to the received signal prediction result data and the imaginary part component of the received signal corresponding to the label; The amplitude loss function is used to represent the product between the third mean square error and the peak weighted term; wherein the third mean square error includes the mean square error between the signal amplitude corresponding to the pre-acquired received signal prediction result data and the signal amplitude of the received signal corresponding to the pre-acquired label.
8. A method for predicting optical fiber transmission reception signals, characterized in that: include: The transmission signal currently generated by the transmitting end in the continuous spectrum nonlinear frequency division multiplexing system is input into the optical fiber transmission receiving signal prediction model, so that the optical fiber transmission receiving signal prediction model outputs corresponding receiving signal prediction result data; wherein, the optical fiber transmission receiving signal prediction model is pre-trained based on the optical fiber transmission receiving signal prediction model training method according to any one of claims 1 to 7.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the optical fiber transmission received signal prediction model training method according to any one of claims 1 to 7, and / or implements the optical fiber transmission received signal prediction method according to claim 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the optical fiber transmission reception signal prediction model training method according to any one of claims 1 to 7, and / or implements the optical fiber transmission reception signal prediction method according to claim 8.
Citation Information
Patent Citations
Training method and device for multi-span optical fiber transmission signal prediction system
CN114647976A
Method and device for training optical fiber transmission signal prediction model
CN114647977A
Signal detection method and device, and storage medium
US20250023759A1
Neural network training method, electronic device, and computer storage medium
WO2023077809A1
Cited By
Time sequence prediction method based on multi-scale decomposition and gating fusion
CN120893016A
Optical-infrared target detection method based on coupling nonlinear dynamics
CN122199939A