Optical fiber transmission and reception signal prediction model training method, prediction method and device

By training a neural network model and utilizing hybrid domain feature extraction and peak-sensing attention mechanisms, the accuracy problem of received signal prediction in CS-NFDM systems was solved, improving the efficiency and accuracy of fiber optic communication system modeling and reducing resource consumption.

CN120601969BActive Publication Date: 2025-11-07BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511100546.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-07
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

Existing fiber optic communication system modeling methods cannot accurately predict received signals in continuous spectrum nonlinear frequency division multiplexing (CS-NFDM) systems, resulting in low modeling efficiency and increased equipment resource consumption.

Method used

A neural network model is trained, and the received signal prediction results are generated through hybrid domain feature extraction, multi-scale processing, and peak perception attention mechanism. The neural network model is then optimized to improve prediction accuracy.

Benefits of technology

It improves the accuracy and efficiency of fiber optic communication system modeling, shortens the R&D cycle, reduces computing resource consumption, and provides a more effective basis for design and optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120601969B_ABST
    Figure CN120601969B_ABST
Patent Text Reader

Abstract

The application provides a fiber transmission and reception signal prediction model training method, a prediction method and equipment, relates to the transmission technical field, and the training method comprises the following steps: generating a training sample corresponding to a transmission signal according to a transmission signal generated by a sending end in a continuous spectrum nonlinear frequency division multiplexing system in advance, training a neural network model based on the training sample to perform mixed domain feature extraction, multi-scale processing and attention feature extraction based on a peak value perception attention mechanism on the training sample, obtaining reception signal prediction result data and optimizing the neural network model, so that the neural network model is trained as a fiber transmission and reception signal prediction model. The application can realize the prediction of the reception signal in the continuous spectrum nonlinear frequency division multiplexing system, can improve the accuracy of the reception signal prediction by using the trained fiber transmission and reception signal prediction model, and can reduce the research and development period of the fiber communication system modeling process and the calculation resource consumption of the research and development equipment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the transmission field, and in particular to a fiber transmission and reception signal prediction model training method, a prediction method and equipment. BACKGROUND

[0002] With the rapid rise of global data traffic and network carrying demand, high-capacity high-speed optical fiber communication technology has become the core support of information infrastructure. Nonlinear frequency division multiplexing (NFDM) system is an important technology of next-generation high-speed optical communication, and continuous spectrum nonlinear frequency division multiplexing (CS-NFDM) is an optical fiber communication technology based on nonlinear Fourier transform (NFT), which effectively counteracts the Kerr nonlinear effect in the optical fiber by modulating the signal on the continuous spectrum component of the nonlinear spectrum. The complex channel characteristic modeling has always been a technical challenge.

[0003] At present, although in the field of optical fiber communication system modeling, there are classical waveform simulation methods such as step Fourier algorithm (SSFM) and polarization multiplexing optical fiber channel modeling based on BERT, there is still no automatic prediction method specially applicable to the received signal in the continuous spectrum nonlinear frequency division multiplexing (CS-NFDM) system, and the accuracy of the received signal prediction in the continuous spectrum nonlinear frequency division multiplexing (CS-NFDM) system cannot be guaranteed, and the accuracy of the optical fiber communication system modeling for the continuous spectrum nonlinear frequency division multiplexing (CS-NFDM) system cannot be guaranteed, which will affect the modeling efficiency and increase the consumption of equipment resources. SUMMARY

[0004] In view of this, the embodiments of the present application provide a fiber transmission and reception signal prediction model training method, a prediction method and equipment to eliminate or improve one or more defects in the prior art.

[0005] One aspect of the present application provides a fiber transmission and reception signal prediction model training method, comprising:

[0006] generating a training sample corresponding to the transmission signal according to the transmission signal generated in advance by the sending end in the continuous spectrum nonlinear frequency division multiplexing system, and generating a label corresponding to the training sample according to the received signal formed after the transmission signal is transmitted to the receiving end through the optical fiber in advance;

[0007] training a neural network model based on the training samples, so that the neural network model performs mixed domain feature extraction, multi-scale processing and attention feature extraction based on a peak perception attention mechanism on the training samples respectively to obtain received signal prediction result data for predicting a received signal formed after the training samples are transmitted by an optical fiber to a receiving end, and optimizing the neural network model based on the received signal prediction result data corresponding to the training samples and the labels, so as to train the neural network model into an optical fiber transmission received signal prediction model for outputting corresponding received signal prediction result data according to the transmitted signal.

[0008] In some embodiments of the present application, the neural network model comprises a sample processing module and a peak perception attention module.

[0009] The sample processing module is configured to obtain a three-dimensional tensor corresponding to the training sample, and perform mixed domain feature extraction and multi-scale processing on the three-dimensional tensor corresponding to the training sample to obtain multi-scale feature data corresponding to the training sample.

[0010] The peak perception attention module is configured to perform peak detection on the multi-scale feature data corresponding to the training sample to obtain peak weights corresponding to the training sample, and perform attention detection on the multi-scale feature data to obtain attention weights corresponding to the training sample; fuse the peak weights and the attention weights to obtain corresponding fused weights, and perform normalization and feature aggregation processing on the fused weights to obtain attention feature data corresponding to the training sample for generating the received signal prediction result data.

[0011] In some embodiments of the present application, the peak perception attention module comprises a peak detection branch, a standard detection branch, a peak fusion layer, a normalization layer, a feature aggregation layer and an output convolution layer; the peak detection branch and the standard detection branch are connected to the peak fusion layer; the peak fusion layer, the normalization layer, the feature aggregation layer and the output convolution layer are connected in sequence.

[0012] The peak detection branch is configured to perform peak detection on the multi-scale feature data corresponding to the training sample to obtain peak weights corresponding to the training sample.

[0013] The standard detection branch is configured to perform attention detection on the multi-scale feature data corresponding to the training sample based on a query vector, a key vector and a value vector to obtain attention weights corresponding to the training sample.

[0014] The peak fusion layer is configured to fuse the peak weights and the attention weights to obtain corresponding fused weights.

[0015] The normalization layer is configured to normalize the fusion weight to obtain a normalized fusion weight.

[0016] The feature aggregation layer is configured to multiply the normalized fusion weight and the value vector to obtain corresponding weighted feature data.

[0017] The output convolution layer is configured to perform convolution and residual connection on the weighted feature data to obtain attention feature data corresponding to the training sample.

[0018] In some embodiments of the present application, the neural network model further comprises a peak enhancement module and an output layer.

[0019] The peak enhancement module is connected to the peak-aware attention module, and the peak enhancement module is configured to sequentially perform channel compression and recovery processing on the attention feature data based on two convolution layers to obtain corresponding peak-enhanced feature data, and perform weighted residual connection between the peak-enhanced feature data and the attention feature data to obtain peak-enhanced attention feature data corresponding to the training sample.

[0020] The output layer is configured to generate, according to the peak-enhanced attention feature data, received signal prediction result data for predicting a received signal formed after the training sample is transmitted to a receiving end through an optical fiber; wherein the received signal prediction result data comprises real part prediction components and imaginary part prediction components of the received signal.

[0021] In some embodiments of the present application, the sample processing module comprises an input layer, a mixed domain feature extraction unit and a multi-scale feature extraction unit.

[0022] The input layer is configured to receive the training sample, the training sample comprising real part components and imaginary part components corresponding to the transmitted signal; the input layer is further configured to stack the real part components and the imaginary part components corresponding to the transmitted signal in a channel dimension to obtain a three-dimensional tensor corresponding to the training sample; wherein each dimension in the three-dimensional tensor is used to represent batch size, channel number and time sequence length, respectively.

[0023] The mixed domain feature extraction unit is configured to perform time domain feature extraction and frequency domain feature extraction on the three-dimensional tensor corresponding to the training sample to obtain time domain feature data and frequency domain feature data corresponding to the training sample, and perform feature fusion on the time domain feature data and the frequency domain feature data to obtain mixed domain feature data corresponding to the training sample.

[0024] The multi-scale feature extraction unit is configured to perform feature extraction on the mixed domain feature data corresponding to the training sample in multiple time scales to obtain multiple time scale feature data corresponding to the training sample, and perform feature fusion processing on each time scale feature data to obtain multi-scale feature data corresponding to the training sample.

[0025] In some embodiments of the present application, the training sample corresponding to the transmission signal is generated according to a transmission signal generated in advance by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system, and the label corresponding to the training sample is generated according to a receiving signal formed after the transmission signal is transmitted to a receiving end through an optical fiber in advance.

[0026] Obtain each signal pair, wherein each signal pair contains a transmission signal generated in advance by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system and a receiving signal formed after the transmission signal is transmitted to a receiving end through an optical fiber.

[0027] Perform quantile-based robust normalization processing on the transmission signal and the receiving signal in each signal pair, respectively.

[0028] Separate each transmission signal and each receiving signal after the robust normalization processing into real and imaginary parts to obtain a training sample corresponding to each signal pair, respectively, and take each receiving signal in each signal pair as a label corresponding to each training sample, respectively. Each training sample contains real and imaginary parts of a transmission signal. Each label contains real and imaginary parts of a receiving signal belonging to the same signal pair as the transmission signal corresponding to the label.

[0029] In some embodiments of the present application, the neural network model is optimized based on the receiving signal prediction result data corresponding to the training sample and the label, including:

[0030] In the current iteration round, based on the receiving signal prediction result data corresponding to the training sample and the label, a target loss corresponding to the current iteration round is determined by a preset target loss function, and the neural network model is optimized based on the target loss.

[0031] The target loss function is composed of a real part loss function, an imaginary part loss function, an amplitude loss function, and an amplitude loss weight coefficient corresponding to the amplitude loss function.

[0032] The real part loss function is used to represent the product between a first mean square error and a peak weighting term; wherein the first mean square error comprises a mean square error between the real part prediction component corresponding to the received signal prediction result data and a real part component of the received signal corresponding to the label; the peak weighting term comprises a sum between 1 and a peak product term; the peak product term comprises a product between a preset peak weighting factor and a peak mask;

[0033] The imaginary part loss function is used to represent the product between a second mean square error and the peak weighting term; wherein the second mean square error comprises a mean square error between the imaginary part prediction component corresponding to the received signal prediction result data and an imaginary part component of the received signal corresponding to the label.

[0034] The amplitude loss function is used to represent the product between a third mean square error and the peak weighting term; wherein the third mean square error comprises a mean square error between a pre-acquired signal amplitude corresponding to the received signal prediction result data and a pre-acquired signal amplitude of the received signal corresponding to the label.

[0035] The second aspect of the present application provides a fiber transmission received signal prediction method, comprising:

[0036] Inputting a transmission signal currently generated by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system into a fiber transmission received signal prediction model, so as to make the fiber transmission received signal prediction model output corresponding received signal prediction result data; wherein the fiber transmission received signal prediction model is obtained by pre-training based on the fiber transmission received signal prediction model training method.

[0037] The third aspect of the present application provides a fiber transmission received signal prediction model training device, comprising:

[0038] A training data construction module is configured to generate a training sample corresponding to a transmission signal generated by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system in advance, and generate a label corresponding to the training sample according to a received signal formed after the transmission signal is transmitted to a receiving end through an optical fiber in advance.

[0039] The model training module is configured to train a neural network model based on the training samples, so that the neural network model performs mixed domain feature extraction, multi-scale processing, and attention feature extraction based on a peak value perception attention mechanism on the training samples respectively to obtain received signal prediction result data for predicting a received signal formed after the training samples are transmitted by an optical fiber to a receiving end, and optimizes the neural network model based on the received signal prediction result data corresponding to the training samples and the labels, so as to train the neural network model as an optical fiber transmission received signal prediction model for outputting the corresponding received signal prediction result data according to the transmission signal.

[0040] The fourth aspect of the present application provides an optical fiber transmission received signal prediction device, comprising:

[0041] The model prediction module is configured to input a transmission signal currently generated by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system into an optical fiber transmission received signal prediction model, so that the optical fiber transmission received signal prediction model outputs corresponding received signal prediction result data; wherein the optical fiber transmission received signal prediction model is obtained by pre-training based on the optical fiber transmission received signal prediction model training method.

[0042] The fifth aspect of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the optical fiber transmission received signal prediction model training method, and / or implement the optical fiber transmission received signal prediction method.

[0043] The sixth aspect of the present application provides a computer readable storage medium, which stores a computer program executable by a processor to implement the optical fiber transmission received signal prediction model training method, and / or implement the optical fiber transmission received signal prediction method.

[0044] The seventh aspect of the present application provides a computer program product comprising a computer program executable by a processor to implement the optical fiber transmission received signal prediction model training method, and / or implement the optical fiber transmission received signal prediction method.

[0045] The optical fiber transmission and reception signal prediction model training method provided in the application generates a training sample corresponding to a transmission signal according to a transmission signal generated by a sending end in a continuous spectrum nonlinear frequency division multiplexing system in advance, and generates a label corresponding to the training sample according to a reception signal formed after the transmission signal is transmitted to a receiving end through an optical fiber in advance; a neural network model is trained based on the training sample, so that the neural network model performs mixed domain feature extraction, multi-scale processing, and attention feature extraction based on a peak value perception attention mechanism on the training sample respectively to obtain reception signal prediction result data for predicting a reception signal formed after the training sample is transmitted to the receiving end through the optical fiber, and the neural network model is optimized based on the reception signal prediction result data corresponding to the training sample and the label, so as to train the neural network model into an optical fiber transmission and reception signal prediction model for outputting corresponding reception signal prediction result data according to the transmission signal, that is, a training method for a prediction model of a reception signal in a continuous spectrum nonlinear frequency division multiplexing system is provided, which can effectively improve the reliability and effectiveness of the optical fiber transmission and reception signal prediction model training process, can realize prediction of the reception signal in the continuous spectrum nonlinear frequency division multiplexing system, can improve the accuracy of reception signal prediction by using the trained optical fiber transmission and reception signal prediction model, and can further predict the transmission performance of the reception signal under different environmental distances in the high fiber communication system modeling process, shorten the research and development cycle of the fiber communication system modeling process, reduce the computing resource consumption of the research and development equipment, effectively avoid the waste of modeling experiment cost, provide a more effective basis for fiber communication system design, performance prediction and optimization, and improve the efficiency and transmission performance of the fiber communication system constructed according to the fiber communication system modeling result.

[0046] Additional advantages, objects, and features of the application will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the following or can be learned by practice of the application. The objects and other advantages of the application can be realized and attained by the structure particularly pointed out in the written description and claims hereof as well as the appended drawings.

[0047] It will be understood by those skilled in the art that the objects and advantages of the present application can be realized and attained by the structure particularly pointed out in the written description and claims hereof as well as the appended drawings. BRIEF DESCRIPTION OF DRAWINGS

[0048] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the principles of the application. The components in the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the application. For purposes of clarity and a consistent approach, portions of the drawings may have been exaggerated from the actual implementation of the exemplary embodiments of this application, and are included for purposes of illustration and description. In the drawings:

[0049] Figure 1 The first flowchart of the optical fiber transmission and reception signal prediction model training method in an embodiment of the present application.

[0050] Figure 2 The architecture diagram of the neural network model in the optical fiber transmission and reception signal prediction model training method in an embodiment of the present application.

[0051] Figure 3 The architecture diagram of the peak perception attention module in the neural network model in an embodiment of the present application.

[0052] Figure 4 The second flowchart of the optical fiber transmission and reception signal prediction model training method in an embodiment of the present application.

[0053] Figure 5 The third flowchart of the optical fiber transmission and reception signal prediction model training method in an embodiment of the present application.

[0054] Figure 6 The example diagram of the transmitted signal in an application example of the present application.

[0055] Figure 7 The example diagram of the received signal in an application example of the present application.

[0056] Figure 8 The data preprocessing flowchart in an application example of the present application.

[0057] Figure 9 The loss change diagram on the validation set in an application example of the present application.

[0058] Figure 10 The complete flowchart of the optical fiber transmission and reception signal prediction model training method in an application example of the present application.

[0059] Figure 11 The example diagram of the test transmitted signal in an application example of the present application.

[0060] Figure 12 The example diagram of the real received signal in an application example of the present application.

[0061] Figure 13A comparison diagram between the real received signal and the received signal prediction result data in an application example of the present application.

[0062] Figure 14 A comparison diagram between the real received signal and the received signal prediction result data in an application example of the present application.

[0063] Figure 15 A NMSE training trend diagram in an application example of the present application.

[0064] Figure 16 A comparison diagram between the inference time of the SSFM and the peak perception hybrid domain NN in an application example of the present application. DETAILED DESCRIPTION

[0065] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the embodiments and the accompanying drawings. Herein, the illustrative embodiments of the present application and the descriptions thereof are used to explain the present application, but not to limit the present application.

[0066] It should be noted that, in order to avoid the present application being obscured by unnecessary details, only the structures and / or processing steps closely related to the scheme according to the present application are shown in the accompanying drawings, and other details not closely related to the present application are omitted.

[0067] It should be emphasized that the term "comprise / comprising" is used herein to mean that a feature, element, step or component is present, but not to the exclusion of one or more other features, elements, steps or components.

[0068] It should be noted that, if not specifically stated, the term "connected" herein can not only mean direct connection, but also mean indirect connection with an intermediate.

[0069] Hereinafter, the embodiments of the present application will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference signs represent the same or similar components, or the same or similar steps.

[0070] In the field of optical fiber communication system modeling, the split-step Fourier method (SSFM) is a classical waveform simulation method. Its inherent piecewise iteration mechanism results in high computational resource overhead when simulating long-distance transmission. Especially in the context of ultra-high-speed coherent optical communication, the time complexity of this algorithm increases nonlinearly with transmission distance, becoming a key bottleneck that restricts the efficiency of system-level optimization. With the development of machine learning, it has been widely applied in optical communication. Bidirectional long short-term memory-based optical fiber channel modeling and generative adversarial networks have been proposed for waveform modeling using purely data-driven methods. These data-driven methods require a large amount of data to achieve modeling, do not rely on theoretical models, and only focus on time-domain processing. Or BERT-based polarization multiplexing optical fiber channel modeling, i.e., using the Transformer architecture to process sequence signals, processing input sequences through the multi-head self-attention mechanism to assist in channel modeling.

[0071] However, the general Transformer architecture is not specifically optimized for the characteristics of optical communication signals. Therefore, in order to achieve automatic prediction of received signals in a continuous spectrum nonlinear frequency division multiplexing system to reduce the research and development cycle and computational resource consumption of research and development equipment in the process of optical fiber communication system modeling, the embodiments of the present application respectively provide an optical fiber transmission received signal prediction model training method, an optical fiber transmission received signal prediction model training device for executing the optical fiber transmission received signal prediction model training method, an entity device, a computer readable storage medium and a computer program product, which can realize automatic prediction of received signals in a continuous spectrum nonlinear frequency division multiplexing system.

[0072] The embodiments are specifically described as follows.

[0073] Based on this, the embodiments of the present application provide an optical fiber transmission received signal prediction model training method that can be implemented by an optical fiber transmission received signal prediction model training device, as shown in Figure 1 , the optical fiber transmission received signal prediction model training method specifically includes the following contents:

[0074] Step 100: generating a training sample corresponding to the transmission signal according to the transmission signal generated by the sending end in the continuous spectrum nonlinear frequency division multiplexing system in advance, and generating a label corresponding to the training sample according to the received signal formed after the transmission signal is transmitted to the receiving end through the optical fiber in advance.

[0075] A continuous spectrum nonlinear frequency division multiplexing (CS-NFDM) system is an optical fiber communication technology based on nonlinear Fourier transform, which realizes effective management of the nonlinear effects of optical fiber by modulating information to the continuous spectrum part of the nonlinear spectrum. The technology is based on the Zakharov-Shabat scattering theory, which converts the complex nonlinear transmission problem in the optical fiber into a linear evolution problem in the nonlinear frequency domain. Among them, the Zakharov-Shabat (referred to as ZS equation) is the core tool of the inverse scattering theory of the nonlinear Schrödinger equation (NLS), which is used to solve the initial value problem of the integrable system.

[0076] In one or more embodiments of the present application, the sending end can refer to a client device capable of signal transmission, the receiving end can refer to a client device capable of signal reception, and the same client device can be both the sending end and the receiving end.

[0077] Step 200: training a neural network model based on the training samples, so that the neural network model performs mixed domain feature extraction, multi-scale processing and attention feature extraction based on a peak value perception attention mechanism on the training samples respectively to obtain received signal prediction result data for predicting a received signal formed after the training samples are transmitted to a receiving end through an optical fiber, and optimizing the neural network model based on the received signal prediction result data corresponding to the training samples and the labels, so as to train the neural network model into an optical fiber transmission received signal prediction model for outputting corresponding received signal prediction result data according to the transmitted signal.

[0078] In step 200, at least one iteration round of model training can be performed on the neural network model according to each of the training samples, and a predetermined training step is performed in each iteration round: the training step includes: inputting the training sample into the neural network model corresponding to the current iteration round, so that the neural network model performs mixed domain feature extraction, multi-scale processing and attention feature extraction based on a peak value perception attention mechanism on the training sample respectively to obtain received signal prediction result data for predicting a received signal formed after the training sample is transmitted to a receiving end through an optical fiber, and calculating a target loss value based on the received signal prediction result data corresponding to the training sample and the label and optimizing the architecture of the neural network model based on the target loss value, and then determining whether the current iteration round is a predetermined last iteration round or whether the neural network model converges, if yes, stopping iteration and taking the optimized neural network model as an optical fiber transmission received signal prediction model for outputting corresponding received signal prediction result data according to the transmitted signal; if not, taking the optimized neural network model as the neural network model corresponding to the next iteration round.

[0079] In step 300, the mixed domain feature extraction refers to a signal processing method using time domain and frequency domain information simultaneously; the multi-scale processing refers to a method of extracting features of different time scales using different convolution kernel sizes; the peak perception refers to a mechanism of focusing on and optimizing the peak value area of a signal; and the self-attention mechanism is an attention mechanism that learns the dependency between positions by calculating the correlation between each position and all positions in a sequence. Correspondingly, the attention feature based on the peak perception attention mechanism refers to an attention mechanism specially designed for the peak value area of a signal, which enhances the standard self-attention by generating weights through a peak detection branch, and focuses on important signal areas.

[0080] From the above description, it can be seen that the optical fiber transmission received signal prediction model training method provided by the embodiments of the present application provides a training method for a prediction model of a received signal in a continuous spectrum nonlinear frequency division multiplexing system, which can effectively improve the reliability and effectiveness of the optical fiber transmission received signal prediction model training process, can realize prediction of the received signal in the continuous spectrum nonlinear frequency division multiplexing system, and can improve the accuracy of received signal prediction using the trained optical fiber transmission received signal prediction model, thereby being able to predict the transmission performance of the received signal at different environmental distances in advance in the high fiber communication system modeling process, shorten the research and development cycle of the fiber communication system modeling process and reduce the computing resource consumption of the research and development equipment, effectively avoid the waste of modeling experiment cost, and provide a more effective basis for fiber communication system design, performance prediction and optimization, so as to improve the efficiency and transmission performance of constructing a fiber communication system according to the fiber communication system modeling result.

[0081] In order to further improve the attention to the special importance of the peak value area of the communication signal in the optical fiber transmission received signal prediction model training process, in the optical fiber transmission received signal prediction model training method provided by the embodiments of the present application, referring to Figure 2 , the neural network model specifically includes the following contents:

[0082] The sample processing module and the peak perception attention module.

[0083] The sample processing module is configured to obtain a three-dimensional tensor corresponding to the training sample, and perform mixed domain feature extraction and multi-scale processing on the three-dimensional tensor corresponding to the training sample, to obtain multi-scale feature data corresponding to the training sample.

[0084] The peak perception attention module is configured to perform peak detection on the multi-scale feature data corresponding to the training sample to obtain peak weight corresponding to the training sample, and perform attention detection on the multi-scale feature data to obtain attention weight corresponding to the training sample; the peak weight and the attention weight are fused to obtain corresponding fusion weight, and the fusion weight is normalized and aggregated to obtain attention feature data corresponding to the training sample for generating the prediction result data of the received signal.

[0085] The attention weight can be embodied as an attention weight matrix, i.e., a similarity matrix. The peak weight can be embodied as a peak weight map with a value range of 0 to 1 and a dimension of [batch_size, 1, 1024], where batch_size represents batch size. The peak weight is used to identify important positions in the transmitted signal that can correspond to the peak of the received signal.

[0086] Therefore, the fusion of the peak weight and the attention weight to obtain the corresponding fusion weight can be: first, the peak weight map is expanded, and then the expanded peak weight map and the attention weight matrix are multiplied element by element, so that the peak region obtains higher weight in attention calculation.

[0087] To further improve the application effectiveness and reliability of the peak perception attention module in the optical fiber transmission received signal prediction model, in an embodiment of the optical fiber transmission received signal prediction model training method provided by the present application, referring to Figure 3 The peak perception attention module in the neural network model specifically includes the following contents:

[0088] The peak detection branch, the standard detection branch, the peak fusion layer, the normalization layer, the feature aggregation layer and the output convolution layer; the peak detection branch and the standard detection branch are connected to the peak fusion layer; the peak fusion layer, the normalization layer, the feature aggregation layer and the output convolution layer are connected in sequence.

[0089] The peak detection branch is configured to perform peak detection on the multi-scale feature data corresponding to the training sample to obtain peak weight corresponding to the training sample.

[0090] Specifically, the peak-aware attention module is one of the core improvements of the embodiments of the present application, which is specially designed to process the peak characteristics in fiber transmission. In the CS-NFDM system, due to the effect of nonlinear effect, the received signal will appear significant peak at some time compared with the transmitted signal, and these peaks contain the key transmission information. The module receives multiscale_features as input, which contains two branches working together: peak detection branch and standard attention branch.

[0091] The peak detection branch can realize the automatic identification of peak position through a two-layer convolutional network. The first layer of convolutional network compresses the 64 input channel multiscale_features to 16 channels and uses the ReLU function to activate, and the second layer of convolutional network further compresses the 16 channel multiscale_features to 1 channel and uses the Sigmoid function to activate, outputting the peak_weights with the value range between 0 and 1. The ReLU function is a simple and effective activation function, which realizes nonlinear mapping by truncating negative input to 0 and keeping positive input original value. The sigmoid function is a common S-shaped function in biology, also known as S-shaped growth curve.

[0092] The standard detection branch is used to detect the attention of the multiscale_features corresponding to the training sample based on the query vector, key vector and value vector, to obtain the attention weight corresponding to the training sample.

[0093] Specifically, the standard attention branch realizes the self-attention mechanism, which can generate query vector (Q), key vector (K) and value vector (V) through three independent 1x1 convolutional layers respectively. Each vector maps the input 64-channel multiscale_features to a new 64-channel representation. The calculation of attention weight first obtains the similarity matrix through the matrix multiplication of Q and K, and then scales it by dividing by the square root of the feature dimension.

[0094] Wherein, the similarity matrix corresponds to the attention weight matrix calculated in the standard attention branch, and the attention weight matrix The calculation formula is as follows:

[0095]

[0096] Wherein, The similarity matrix represents the similarity relationship between time points: each element (i, j) of the similarity matrix represents the similarity between the ith time point and the jth time point; before the softmax calculation, the similarity matrix represents the original similarity score. Attention represents the attention mechanism. Softmax represents the normalization exponential function.

[0097] The peak fusion layer is configured to fuse the peak weight and the attention weight to obtain a corresponding fused weight.

[0098] Specifically, in the peak fusion layer, the peak weight is first expanded, and the dimension [B, 1, L] of the original peak weight is expanded to the dimension [B, L, L] of the expanded peak weight, where B (i.e., the abbreviation of Batch or batch_size) represents the batch size, and L (i.e., the abbreviation of Length) represents the length of the time sequence. The original peak weight [B, 1, L] represents the peak importance of each time point; each row of the dimension [B, L, L] of the expanded peak weight replicates the same peak weight, so that the peak weight can be multiplied element by element with the attention matrix of the same dimension [B, L, L], so that the peak weight can modulate the entire attention matrix, and each time point can consider the peak importance of the target time point when calculating the attention with other time points, thereby achieving the effect that the peak region obtains higher attention weight.

[0099] The expanded peak weight is multiplied element by element with the similarity matrix, so that the peak region obtains higher weight in attention calculation. The mathematical expression is as follows:

[0100]

[0101] where d k is the feature dimension 64, represents element-wise multiplication, W peak is the expanded peak weight map.

[0102] The normalization layer is configured to normalize the fused weight to obtain a normalized fused weight.

[0103] The feature aggregation layer is configured to multiply the normalized fused weight and the value vector to obtain corresponding weighted feature data.

[0104] The output convolution layer is configured to convolve and connect the weighted feature data in a residual manner to obtain attention feature data corresponding to the training sample.

[0105] Specifically, the attention weight normalized by the softmax is multiplied by V to obtain attention output, and then a 3x3 convolution kernel is used for feature integration. Finally, the attention output is added to the original input through residual connection to maintain the integrity of the features.

[0106] In an example of the present application, four peak-aware attention modules connected in sequence can be used to form a deep attention network. Each layer of peak-aware attention module independently calculates the peak weight and attention weight, and the layers are connected through residual connection. The input of the first layer of peak-aware attention module is the multiscale_features output by the sample processing module, and the attention feature data output by the first layer of peak-aware attention module is used as the input of the second layer of peak-aware attention module, and so on. This multi-layer structure enables the network to gradually refine the attention on the peak region and better capture the peak evolution characteristics in the transmission process from the emission signal to the receiving signal. The final output is attention_features, still with a dimension of (batch_size, 64, 1024).

[0107] In order to further improve the reliability and effectiveness of the training process of the optical fiber transmission received signal prediction model, automatically predict the received signal in the continuous spectrum nonlinear frequency division multiplexing system, and improve the accuracy of predicting the received signal by using the trained optical fiber transmission received signal prediction model, in the optical fiber transmission received signal prediction model training method provided in the embodiment of the present application, referring to Figure 2 , the neural network model further comprises the following content:

[0108] The peak enhancement module and the output layer.

[0109] The peak enhancement module is connected with the peak-aware attention module, and the peak enhancement module is configured to sequentially perform channel compression and recovery processing on the attention feature data based on two convolutional layers to obtain corresponding peak enhancement feature data, and perform weighted residual connection between the peak enhancement feature data and the attention feature data to obtain peak enhancement attention feature data corresponding to the training sample.

[0110] Specifically, the peak enhancement module receives the attention feature data after attention processing, and further enhances the expression ability of the peak information through a special convolutional network. Considering the importance of the peak feature in the received signal for accurately modeling the optical fiber channel, the peak enhancement module adopts a two-layer convolutional structure. The first layer compresses the 64 input channels of the attention feature data into 32 channels, uses a 3x3 convolutional kernel and a ReLU activation function, and extracts peak-related features. The second layer restores the 32 channels of the attention feature data to 64 channels, also uses a 3x3 convolutional kernel, but uses a Tanh activation function, and outputs peak enhancement feature data with a range of -1 to 1. The Tanh function, also known as the hyperbolic tangent activation function, is a transformation of the Sigmoid function.

[0111] The peak-enhanced feature data and the original attention feature data are combined through a weighted residual connection, and the weight coefficient is set to 0.3. This design not only retains the main information of the original features, but also moderately enhances the expression of the peak features. The specific calculation is as follows:

[0112] enhanced_features = attention_features + 0.3 × peak_enhanced

[0113] Wherein, enhanced_features represents the peak-enhanced attention feature data; peak_enhanced is the peak-enhanced feature data output by the two-layer convolution. The small weight coefficient 0.3 ensures that the enhancement operation will not change the original features too much, but will play a fine adjustment role, helping the network to better reconstruct the peak features in the received signal. The dimension of the peak-enhanced attention feature data (enhanced_features) output by the peak-enhancement module remains (batch_size, 64, 1024), which will be used as the input of the output layer.

[0114] The output layer is used to generate received signal prediction result data for predicting the received signal formed after the training sample is transmitted to the receiving end through the optical fiber; wherein the received signal prediction result data includes: real part prediction component and imaginary part prediction component of the received signal.

[0115] Specifically, the output layer receives the peak-enhanced attention feature data (enhanced_features) and is responsible for generating the final received signal prediction, i.e., the predicted received signal prediction result data. This layer adopts a two-layer convolution network structure, the first layer compresses the 64 input channels of the peak-enhanced attention feature data into 32 channels, uses a 5x5 convolution kernel and a LeakyReLU activation function, and performs preliminary mapping of the features. The second layer further compresses the 32 channels of the peak-enhanced attention feature data into 2 channels, corresponding to the real part and the imaginary part of the received signal, also uses a 5x5 convolution kernel but does not use an activation function, and maintains the linearity of the output. LeakyReLU (Leaky Rectified Linear Unit) is a variant of ReLU activation function.

[0116] The 2-channel tensor generated by the output layer needs to be separated by channel, the 0th channel corresponds to the predicted real part prediction component (pred_real), and the 1st channel corresponds to the predicted imaginary part prediction component (pred_imag), and the dimension of each component is (batch_size, 1024). These two prediction components correspond to the received signal q rxThe real and imaginary parts of (t) have the same format as the real part (rx real) and the imaginary part (rx imag) of the real received signal saved in the preprocessing stage, and the loss calculation can be directly performed.

[0117] The data flow of the entire neural network model forms a complete processing link: the real part (tx real) and the imaginary part (tx imag) of the input transmitted signal are first organized into an input tensor, and then the time domain and frequency domain features are extracted respectively, the fused features pass through the multi-scale processing, peak perception attention, peak enhancement and other modules in turn, and finally the output layer generates the real part prediction component and the imaginary part prediction component of the received signal corresponding to the received signal prediction result data. This end-to-end design enables the network to learn the complete mapping relationship from the transmitted signal to the received signal, and realizes accurate modeling of the optical fiber channel.

[0118] In order to further improve the reliability and effectiveness of the optical fiber transmission received signal prediction model training process, realize automatic prediction of the received signal in the continuous spectrum nonlinear frequency division multiplexing system, and improve the accuracy of received signal prediction using the trained optical fiber transmission received signal prediction model, in the optical fiber transmission received signal prediction model training method provided in the embodiment of the present application, referring to Figure 2 , the sample processing module in the neural network model specifically includes the following contents:

[0119] The input layer, the mixed domain feature extraction unit and the multi-scale feature extraction unit.

[0120] The input layer is used to receive the training sample, and the training sample includes the real part and the imaginary part corresponding to the transmitted signal; the input layer is also used to stack the real part and the imaginary part corresponding to the transmitted signal in the channel dimension to obtain a three-dimensional tensor corresponding to the training sample; wherein each dimension in the three-dimensional tensor is used to represent the batch size, the channel number and the time sequence length.

[0121] Specifically, the input layer receives the real part and the imaginary part of the transmitted signal, and each component is a one-dimensional array with a length of 1024, corresponding to the sampling value in a 6 nanosecond time window. The main function of the input layer is to organize the two independent real number sequences into a tensor format suitable for neural network processing. Specifically, the input layer stacks the real part and the imaginary part of the transmitted signal in the channel dimension to form a three-dimensional tensor with a dimension of (batch_size, 2, 1024), wherein the first dimension represents the batch size, the second dimension contains two channels of real and imaginary parts, and the third dimension is the time sequence length.

[0122] This tensor representation preserves the full information of the transmitted signal q tx (t), including the 16-QAM modulation information of its 64 subcarriers and the phase characteristics introduced by the continuous spectrum modulation. The generated three-dimensional tensor will be simultaneously fed into the time-domain feature extraction branch and the frequency-domain feature extraction branch for parallel processing, with the goal of learning the mapping relationship from the transmitted signal to the received signal.

[0123] The mixed-domain feature extraction unit is configured to perform time-domain feature extraction and frequency-domain feature extraction on the three-dimensional tensor corresponding to the training sample respectively to obtain time-domain feature data and frequency-domain feature data corresponding to the training sample, and perform feature fusion on the time-domain feature data and the frequency-domain feature data to obtain mixed-domain feature data corresponding to the training sample.

[0124] Specifically, the mixed-domain feature extraction unit comprises a time-domain feature extraction branch and a frequency-domain feature extraction branch, and a feature fusion layer connected to the time-domain feature extraction branch and the frequency-domain feature extraction branch respectively.

[0125] The time-domain feature extraction branch receives the three-dimensional tensor output by the input layer as input and extracts the time sequence features of the signal through two one-dimensional convolutional networks. These features are useful for capturing transient changes in the three-dimensional tensor and pulse broadening due to dispersion. The first layer of convolutional network maps 2 input channels of the three-dimensional tensor to 64 feature channels, using a convolution kernel of size 5 and padding of 2, which ensures that the sequence length remains unchanged at 1024. The features after convolution are processed by a LeakyReLU activation function with a negative slope of 0.2, enhancing the non-linear representation ability of the network. The activated feature dimension is (batch_size, 64, 1024), containing the preliminary extracted time-domain feature information. The second layer of convolutional network receives the output of the first layer and refines the features from 64 to 64 channels, using the same convolution kernel size of 5 and activation function settings. The role of this layer is to further extract and abstract the time-domain features and generate higher-level feature representations. After two layers of convolution processing, the time-domain branch outputs time-domain feature data (time_features) with a dimension of (batch_size, 64, 1024). The time-domain feature data effectively captures the local change patterns and time sequence dependencies in the time dimension, especially the time-domain distortion characteristics in the transmission process from the transmitted signal to the received signal.

[0126] The frequency domain feature extraction branch and the time domain branch are parallel, and both receive the three-dimensional tensor. The branch first recombines the real and imaginary parts in the three-dimensional tensor into a complex signal, recovering the complex form of the transmitted signal. The reconstructed complex signal has a dimension of (batch_size, 1024), and then a fast Fourier transform (FFT) is performed on the signal in each batch to obtain a frequency domain representation. In one or more embodiments of the present application, the real part can be referred to as the real part; the imaginary part can be referred to as the imaginary part.

[0127] In order to fully exploit the value of frequency domain information, the system extracts four different frequency domain features from the FFT results. The amplitude spectrum is obtained by calculating the modulus of the complex number, which reveals the energy distribution of the signal at different frequencies, which is important for understanding the frequency selective fading in optical fiber transmission; the phase spectrum is obtained by calculating the argument of the complex number, which reflects the phase relationship of each frequency component, which is crucial for capturing the phase change caused by dispersion; the real part spectrum and the imaginary part spectrum extract the real part and the imaginary part of the FFT result respectively, which preserves the complete frequency domain information. The four features each form a tensor with a dimension of (batch_size, 1, 1024), and then are spliced in the channel dimension to form a frequency domain feature tensor with a dimension of (batch_size, 4, 1024). The frequency domain feature tensor is then processed by a two-layer convolutional network. The first layer maps 4 input channels to 64 feature channels, and the second layer performs 64-to-64 channel feature extraction, both using a convolution kernel with a size of 5 and a LeakyReLU activation function. The frequency domain branch finally outputs frequency domain feature data (freq_features) with a dimension of (batch_size, 64, 1024), which captures the spectral characteristics and global structural information of the signal, and is extremely effective for modeling the frequency domain changes from the transmitted signal to the received signal.

[0128] The feature fusion layer receives the time-domain feature data output by the time-domain branch and the frequency-domain feature data output by the frequency-domain branch, both of which are feature tensors with a dimension of (batch_size, 64, 1024). The design of the feature fusion layer takes into account the coupling characteristics of the time-domain and frequency-domain effects in the transmission process from the transmitting signal to the receiving signal. The feature fusion layer first concatenates the two features in the channel dimension to form a joint feature representation with a dimension of (batch_size, 128, 1024). This concatenation operation preserves all the information of the time-domain and frequency-domain, providing a complete input for subsequent feature integration. The 128-channel feature after concatenation is processed by a 1x1 convolution for dimension reduction, compressing the channel number back to 64. This dimension reduction operation not only reduces the computational complexity, but also achieves effective fusion of time-domain and frequency-domain features through the learning of the weight matrix. The 1x1 convolution is equivalent to a linear transformation of the 128-dimensional feature vector at each time point, generating 64-dimensional fused features. The dimension of the fused mixed-domain feature data (fused_features) is (batch_size, 64, 1024), which contains both time-domain and frequency-domain information of the signal, providing a rich and comprehensive input representation for subsequent multi-scale feature extraction.

[0129] The multi-scale feature extraction unit is configured to perform multi-time-scale feature extraction on the mixed-domain feature data corresponding to the training sample to obtain a plurality of time-scale feature data corresponding to the training sample, and perform feature fusion processing on each of the time-scale feature data to obtain multi-scale feature data corresponding to the training sample.

[0130] Specifically, the multi-scale feature extraction unit receives the fused_features as input, and captures signal features of different time scales through a parallel multi-branch convolution structure. This design is particularly suitable for processing the multi-scale effects in fiber transmission: fast nonlinear phase modulation and slow pulse broadening caused by dispersion. The module designs four parallel branches, each of which uses different convolution configurations to extract features of a specific scale. The first branch uses a convolution with a kernel size of 3 to capture local detail features, suitable for detecting rapid changes in the transmitted signal; the second branch uses a convolution with a kernel size of 5 to capture medium-range features; the third branch uses a convolution with a kernel size of 7 to capture larger-range features, suitable for modeling dispersion effects; the fourth branch uses an expansion convolution with an expansion rate of 2 and a kernel size of 3 to capture long-distance dependencies. The input to each branch is the 64-channel fused_features, and the output is 16-channel time-scale feature data. The outputs of the four branches are concatenated in the channel dimension to form 64-channel multi-scale feature data. This design ensures a balanced representation of features of different scales. The concatenated features are fused through a 1x1 convolution, and then a residual connection is made with the original input, i.e., the fused features are added to the input features. The residual connection helps the gradient to be backpropagated, improving the stability of the training.

[0131] The features after residual connection The features also need to be processed by layer normalization. Layer normalization calculates the mean μ and variance σ2 of each sample in the feature dimension, and then normalizes according to the formula:

[0132]

[0133] wherein, denotes layer normalization; γ and β are learnable parameters, and ε is a numerical stability constant. Layer normalization helps stabilize the training process and accelerate convergence. The multi-scale feature extraction unit uses 2 consecutive processing layers, and the output of the first layer is directly used as the input of the second layer, to enhance the expression ability of the features by stacking multiple layers. The final output of the multi-scale feature data (multiscale_features) maintains a dimension of (batch_size, 64, 1024), which will be passed to the peak perception attention module.

[0134] In order to further improve the reliability and effectiveness of the training process of the optical fiber transmission received signal prediction model, realize automatic prediction of the received signal in the continuous spectrum nonlinear frequency division multiplexing system, and improve the accuracy of received signal prediction using the trained optical fiber transmission received signal prediction model, in the optical fiber transmission received signal prediction model training method provided in the embodiments of the present application, referring to Figure 4The step 100 in the optical fiber transmission and reception signal prediction model training method specifically includes the following contents:

[0135] Step 110: Obtain each signal pair, wherein each signal pair includes a transmission signal generated by a sending end in a continuous spectrum nonlinear frequency division multiplexing system and a reception signal formed after the transmission signal is transmitted to a receiving end through an optical fiber.

[0136] Step 120: Perform quantile-based robust normalization processing on the transmission signal and the reception signal in each signal pair.

[0137] Step 130: Separate each transmission signal and each reception signal after the robust normalization processing into real and imaginary parts to obtain training samples corresponding to each signal pair, and take each reception signal in each signal pair as a label corresponding to each training sample; wherein each training sample includes real and imaginary parts of the transmission signal; and each label includes real and imaginary parts of the reception signal belonging to the same signal pair as the transmission signal corresponding to the label.

[0138] In order to further improve the reliability and effectiveness of the optical fiber transmission and reception signal prediction model training process, realize automatic prediction of the reception signal in the continuous spectrum nonlinear frequency division multiplexing system, and improve the accuracy of the reception signal prediction using the trained optical fiber transmission and reception signal prediction model, in the optical fiber transmission and reception signal prediction model training method provided in the embodiment of the present application, referring to Figure 5 The step 200 in the optical fiber transmission and reception signal prediction model training method specifically includes the following contents:

[0139] Step 210: In the current iteration round, input the training sample into the neural network model, so that the neural network model performs mixed domain feature extraction, multi-scale processing, and attention feature extraction based on the peak value perception attention mechanism on the training sample to obtain reception signal prediction result data for predicting a reception signal formed after the training sample is transmitted to a receiving end through an optical fiber.

[0140] Step 220: In the current iteration round, based on the reception signal prediction result data corresponding to the training sample and the label, determine a target loss corresponding to the current iteration round by using a preset target loss function, and optimize the neural network model based on the target loss.

[0141] The target loss function is composed of a real part loss function, an imaginary part loss function, an amplitude loss function, and an amplitude loss weight coefficient corresponding to the amplitude loss function.

[0142] The real part loss function is used to represent the product between the first mean square error and the peak weighted item; wherein the first mean square error includes the mean square error between the real part prediction component corresponding to the received signal prediction result data and the real part component of the received signal corresponding to the label; the peak weighted item includes the sum between 1 and the peak product item; the peak product item includes the product between the preset peak weighted factor and the peak mask;

[0143] The imaginary part loss function is used to represent the product between the second mean square error and the peak weighted item; wherein the second mean square error includes the mean square error between the imaginary part prediction component corresponding to the received signal prediction result data and the imaginary part component of the received signal corresponding to the label;

[0144] The amplitude loss function is used to represent the product between the third mean square error and the peak weighted item; wherein the third mean square error includes the mean square error between the pre-acquired signal amplitude corresponding to the received signal prediction result data and the pre-acquired signal amplitude of the received signal corresponding to the label.

[0145] Specifically, the loss function module receives the real part prediction component (pred_real) and the imaginary part prediction component (pred_imag) of the received signal generated by the network output layer, and the real part component (rx_real) and the imaginary part component (rx_imag) of the received signal provided in the data preprocessing stage. Considering the peak characteristics generated in the transmission process of the received signal q rx (t) compared with the transmitted signal q tx (t), the first step of loss calculation is to identify the peak area based on the real received signal. First, calculate the complex amplitude of the real signal, and get the signal amplitude value at each time point through the following formula:

[0146]

[0147] Wherein, true_mag represents the amplitude of the real received signal; Sqrt represents the square root function. Then find the maximum value in the batch time dimension to get the maximum amplitude (max_mag) of each sample.

[0148] The identification of peak regions adopts a thresholding method, marking the time points with amplitudes exceeding 70% of the maximum value as peak regions. This threshold selection is based on the understanding of the fiber transmission characteristics: in the CS-NFDM system, nonlinear effects will cause significant power concentration of signals at certain times, and these high-power regions are crucial for accurate modeling of channel characteristics. The definition of the peak mask (M_peak) is achieved by the following rules: if the signal amplitude at time point t is greater than (maximum amplitude x 0.7), the peak mask M_peak(t) = 1 at this time point, otherwise, M_peak(t) = 0.

[0149] In specific calculations, the steps are:

[0150] (1) Calculate the amplitude of the real received signal true_mag = sqrt(rx_real² + rx_imag²);

[0151] (2) Find the maximum amplitude value max_mag = max(true_mag); max represents the maximum value;

[0152] (3) Set the peak threshold threshold = 0.7 x max_mag;

[0153] (4) Generate a binary mask: when true_mag > threshold, the peak mask M_peak = 1 (indicating a peak region), and when true_mag ≤ threshold, the peak mask M_peak = 0 (indicating a non-peak region). This relative threshold design can adapt to the dynamic range of different signals, ensuring the robustness of peak detection. The peak mask generated in this way is a binary tensor with the same length as the signal, which is used to identify which regions need additional weight attention in loss calculation.

[0154] Where the target loss function adopts a multi-component design, respectively calculating the real part loss, the imaginary part loss and the amplitude loss, to comprehensively evaluate the prediction quality from the transmitted signal to the received signal. Before calculating the component loss, first calculate the amplitude of the predicted received signal This amplitude will be used for the calculation of the amplitude loss, and also reflects the overall quality of the predicted signal.

[0155] The real part loss function L realThe mean square error between the predicted real part and the true real part is calculated, but an innovation is introduced in the peak weighting mechanism. The specific calculation process is as follows: first, the basic MSE error is calculated, and then the error is multiplied by (1+ λ_peak × M_peak), where λ_peak = 3.0 is the peak weighting factor. This means that the error in the peak area will be amplified by 4 times (1+3), while the non-peak area remains the original weight. This design ensures that the network pays special attention to the peak area containing important transmission information when learning the mapping from the transmitted signal to the received signal. Finally, the weighted error is averaged to obtain the real part loss.

[0156] The mathematical expression of the real part loss function L real is as follows:

[0157]

[0158] where MSE represents the normalized mean square error (Normalized Mean Square Error); represents the real part prediction component corresponding to the received signal prediction result data; r represents the real part component of the received signal corresponding to the label; represents the first mean square error; represents the peak weighting term. mean represents the mean function that can be used to calculate the average.

[0159] The calculation method of the imaginary part loss L imag is exactly the same as that of the real part loss, except that the real part is replaced by the imaginary part:

[0160]

[0161] where, represents the imaginary part prediction component corresponding to the received signal prediction result data; i represents the imaginary part component of the received signal corresponding to the label.

[0162] The amplitude loss function L imag calculates the weighted mean square error between the predicted amplitude and the true amplitude:

[0163]

[0164] where, represents the signal amplitude corresponding to the received signal prediction result data; q represents the signal amplitude of the received signal corresponding to the label.

[0165] The target loss L total is obtained by weighted summation of the three component losses:

[0166]

[0167] where a = 0.5 is the amplitude loss weight coefficient. This multi-component design ensures that the network optimizes the prediction accuracy of the real part, the imaginary part, and the amplitude simultaneously, while the peak weighting strategy makes the network pay more attention to the accuracy of the peak region during the training process, which is crucial for accurately modeling the nonlinear mapping relationship from the transmitted signal to the received signal. The calculated target loss L total is a scalar value that will be used as the starting point for backpropagation to calculate the gradient of the network parameters.

[0168] To further illustrate the above embodiments, the application also provides a specific application example of a method for training an optical fiber transmission and reception signal prediction model.

[0169] The prior art has the following problems:

[0170] (1) The general Transformer architecture is not specifically optimized for the characteristics of optical fiber communication signals;

[0171] Signal characteristics do not match: optical fiber signals have specific amplitude, phase, and frequency characteristics, but the general Transformer cannot effectively model these physical characteristics;

[0172] Data preprocessing requirements: Lack of specialized preprocessing modules for optical fiber signals, such as amplitude normalization and phase unfolding;

[0173] Loss function design: Traditional MSE loss cannot fully reflect the key quality indicators of optical fiber signals (such as bit error rate and signal-to-noise ratio).

[0174] (2) The multi-head self-attention mechanism does not consider the special importance of the peak region of the communication signal;

[0175] Loss of peak information: Standard attention mechanisms tend to ignore key peak points in the signal, leading to important information being averaged;

[0176] Unreasonable weight allocation: Multi-head attention cannot adaptively allocate higher weights to the peak region, affecting key feature extraction;

[0177] Insufficient modeling of time-domain features: Lack of specialized modeling mechanisms for signal time-domain peak patterns.

[0178] (3) Lack of specialized modeling mechanisms for signal physical characteristics;

[0179] Inadequate complex signal processing capabilities: Standard Transformers have difficulty effectively processing amplitude and phase information of complex signals;

[0180] Nonlinear distortion modeling is missing: unable to model nonlinear effects in optical fiber transmission (such as self-phase modulation and cross-phase modulation);

[0181] Insufficient utilization of frequency domain features: lack of joint modeling capability of frequency domain and time domain.

[0182] (4) Only processing serialized one-dimensional signal data;

[0183] (5) Inadequate use of time domain and frequency domain complementary characteristics of signals;

[0184] (6) Lack of specialized processing for complex signal amplitude and phase.

[0185] To solve the above technical problems, the optical fiber transmission receiving signal prediction model training method provided by the application application contains the following contents:

[0186] (I) Continuous spectrum nonlinear frequency division multiplexing system overview

[0187] In an ideal lossless optical fiber, the transmission evolution of the continuous spectrum follows a certain mathematical law. After the signal propagates in the optical fiber for a distance L, the change of the continuous spectrum coefficient of the signal can be represented as:

[0188]

[0189] Where q c represents the continuous spectrum coefficient; exp represents the natural exponential function; i represents the imaginary unit; ξ is the nonlinear frequency variable. This evolution law shows that the continuous spectrum only undergoes a phase rotation proportional to ξ 2 during transmission, while the amplitude remains unchanged, which provides a theoretical basis for reliable communication in nonlinear channels.

[0190] The application application adopts a multi-carrier modulation scheme of 64 subcarriers, supports 16-QAM modulation format, and realizes efficient data transmission in a bandwidth of 34GHz. The system effectively manages the evolution characteristics of the signal in the transmission process by adjusting the phase at the sending end and the receiving end.

[0191] (II) Generation of transmitted signal

[0192] 1. Data source processing and 16-QAM constellation mapping

[0193] First, the original binary data stream {b i} is received, where i=1,2,...,N bits , respectively representing different bit numbers. According to the requirements of 16-QAM modulation format, the system groups the continuous original binary data stream by every 4 bits as a group, and each group of bits is mapped to a complex modulation symbol, so that each symbol carries 4 bits of information.

[0194] QAM constellation mapping generates complex symbols where m is the data block index and k is the subcarrier index, represents the imaginary unit. The system design adopts 64 parallel subcarriers, so the value of k ranges from -32 to +31. The in-phase component I k and the quadrature component Q k of the constellation are selected from the normalized amplitude set {−3,−1,+1,+3}, forming 16 constellation points with equal interval. This constellation design ensures good minimum Euclidean distance between symbols, which is beneficial for reliable detection at the receiving end.

[0195] 2. Continuous spectrum modulation processing

[0196] The continuous spectrum modulation module modulates the symbol sequence obtained by constellation mapping to the nonlinear frequency domain. The system adopts a raised cosine carrier waveform with a roll-off factor α = 0.5, which ensures the orthogonality between subcarriers and has good spectral characteristics. The center frequency of each subcarrier is set to , where the basic time parameter T0= 2×10 −9 seconds determines the carrier spacing.

[0197] The multi-carrier modulation is realized by the superposition principle, and the initial continuous spectrum is constructed as follows:

[0198]

[0199] where h(ξ) represents the raised cosine carrier waveform function. This summation process combines the symbol information carried on the 64 subcarriers into a continuous nonlinear spectrum function. To optimize system performance, the modulation signal needs to be power adjusted, multiplied by a power control factor A = 3:

[0200]

[0201] where represents the power-controlled continuous spectrum; this power control parameter is determined according to the nonlinear threshold of the optical fiber and the system performance requirements. The system performs a phase 50% pre-compensation at the transmitting end to optimize the transmission characteristics of the signal in the optical fiber. The pre-compensation is realized by applying a phase factor:

[0202]

[0203] where =81.3×10 3 meters is the transmission distance designed for the system. This phase pre-compensation is an effective compensation technique for the transmission of the CS-NFDM system.

[0204] 3. Inverse nonlinear Fourier transform (INFT)

[0205] Modulated continuous spectrum The modulated continuous spectrum needs to be converted into time-domain signal by inverse nonlinear Fourier transform (INFT). INFT is based on Zakharov-Shabat inverse scattering theory, which recovers time-domain envelope from nonlinear spectrum by solving the corresponding Riemann-Hilbert problem. The mathematical relation of the transform is expressed as:

[0206]

[0207] where, represents the time-domain envelope recovered from nonlinear spectrum. Numerical implementation adopts fast nonlinear Fourier transform (FNFT) algorithm, which uses 4th-order Runge-Kutta method to solve the relevant differential equations. The time-domain signal is discretized with 512 sampling points, covering a time window of (−3 ns, +3 ns), ensuring the complete inclusion of the 6-nanosecond block duration.

[0208] 4. Transmission filtering and signal conditioning

[0209] The time-domain signal output by INFT needs to be filtered to meet the system bandwidth requirement. The transmission filter adopts ideal low-pass characteristics, and its frequency-domain transfer function is:

[0210]

[0211] where, B = 34 GHz is the system design bandwidth, f represents frequency, and rect() represents the rectangular window function. The filtering process effectively limits the signal bandwidth, preventing mutual interference between adjacent channels. The filtered signal is subjected to inverse normalization processing, which converts the normalization unit used in the calculation into physical units:

[0212]

[0213] where, represents the time-domain signal at the input end of the optical fiber; represents the filtered time-domain signal; the time normalization factor T scale = 4 × 10−10 seconds, and the power normalization factor P scale is determined according to the system transmission power. The processed transmission signal q tx (t) possesses all the necessary characteristics for transmission in the optical fiber.

[0214] (Three) Optical fiber transmission channel

[0215] In the optical fiber transmission channel, the propagation of the signal follows the nonlinear Schrödinger equation (NLSE), which comprehensively describes various physical effects in the optical fiber, where the loss coefficient is 0.2 × 10 −3 ​ / m, corresponding to a fiber loss of 0.087 dB / km; a group-velocity dispersion coefficient of -5.75 x 10 −27 / m, representing an anomalous dispersion characteristic; a nonlinear coefficient of 1.6 x 10 −3 / (W m), describing the strength of Kerr nonlinearity.

[0216] Since the NLSE equation contains both linear and nonlinear terms, and the two do not satisfy the commutative law, an analytic solution cannot be obtained, and a numerical method must be used to solve it. The system uses the split-step Fourier method to numerically solve the NLSE. The split-step Fourier method (SSFM) provides a numerical solution scheme. The core idea of this method is to rewrite the NLSE equation as an operator form, and then use a symmetric split approximation within each step.

[0217] The propagation process is divided into three steps: linear propagation in the first half, complete nonlinear propagation, and linear propagation in the second half.

[0218] In one example, the transmission distance of 81.3 kilometers is divided into 40 calculation segments, each with a length of Δz = 2.0325 kilometers. The linear step handles the loss and dispersion effects. Since the linear operator has a diagonalized form in the frequency domain, performing the calculation in the frequency domain has higher efficiency and accuracy. The linear propagation process first converts the time-domain signal to the frequency domain through fast Fourier transform (FFT), then applies the corresponding transfer function to each frequency component, and finally returns to the time domain through inverse fast Fourier transform (IFFT). The nonlinear step is calculated directly in the time domain, handling the self-phase modulation effect. Since the nonlinear operator only depends on the instantaneous power of the signal and is independent of the time derivative, when the step size is small enough, the power can be assumed to remain constant within the segment, thus obtaining an analytic solution. Nonlinear propagation causes the signal to acquire a phase change proportional to its instantaneous power, and this power-dependent phase modulation is a direct manifestation of the Kerr effect in optical fiber. In numerical implementation, the signal power is calculated for each time sample point, and then the corresponding nonlinear phase modulation is applied. The complete split-step algorithm performs three steps in each calculation segment according to the symmetric split scheme: the first half of the linear propagation handles half of the loss and dispersion effects, the complete nonlinear propagation handles the self-phase modulation of the entire step, and the second half of the linear propagation handles the remaining loss and dispersion effects. Through 40 iterations of calculation, the signal propagates from the entrance to the exit of the optical fiber, completing the entire transmission process.

[0219] After complete fiber transmission, the received signal q rx (t) containing various transmission effects is obtained. At this time, the transmitted signal q tx (t) and the received signal q rx (t) can be collected as a data set for training.

[0220] (IV) Dataset construction

[0221] First, a training dataset containing 1200 pairs of data samples of the transmitted signal q tx (t) and the received signal q rx (t) at the receiving end is constructed, which are generated by the aforementioned CS-NFDM (Continuous Spectrum Nonlinear Frequency Division Multiplexing) system, and each pair of data samples represents a complete signal transmission process.

[0222] The fiber link parameters include:

[0223] (1) Loss coefficient: 0.2 x 10-3 m-1;

[0224] (2) Dispersion coefficient: -5.75 x 10-2 s2 / m; 7

[0225] (3) Nonlinear coefficient: 1.6 x 10-3 (W-m)-1;

[0226] (4) Fiber span length: 81.3 km;

[0227] (5) Span number: 1 span;

[0228] (6) Center frequency: 193.1 THz;

[0229] (7) Time normalization scale: T_scale= 4 x 10-10 s.

[0230] The dataset is divided into a training set, a validation set, and a test set in a ratio of 8:2:2, and specifically contains: 800 pairs of data samples in the training set for network parameter learning, 200 pairs of data samples in the validation set for performance monitoring during the training process, and 200 pairs of data samples in the test set for final model evaluation. Each data sample contains a complex signal with a length of 1024, sampling time information, and corresponding signal amplitude values.

[0231] In an example, the transmitted signal is as shown in Figure 6 The received signal is as shown in Figure 7

[0232] (V) Signal preprocessing and normalization

[0233] During training, the training signal needs to be preprocessed. Referring to Figure 8 , the data preprocessing module receives the paired transmitted signal q tx (t) and the received signal q rx ​​(t) as input, both in complex form. The application example uses a robust normalization method based on quantiles. First, the signal amplitude A = |q(t)| is calculated, then the 1st percentile q1 and the 99th percentile q99 are determined 99 as the normalization boundary. The mathematical expression of the normalization process is:

[0234]

[0235]

[0236] where A is the original signal amplitude, is the signal phase, A norm is the normalized amplitude; denotes the normalized complex time-domain signal; e denotes the natural exponential function. The normalized amplitude is limited to the range (0, 1.5) by a clipping operation, which allows the peak signal to slightly exceed 1 and effectively preserves important peak feature information.

[0237] The processed complex signal is separated into two independent real sequences, real part and imaginary part, each with a length of 1024. This separation strategy allows the neural network to independently process the two orthogonal components of the complex signal, while simplifying subsequent tensor operations. The preprocessing module also saves the normalization parameters (q1, q99, maximum amplitude value), which will be used for signal denormalization after training is completed. The separated real and imaginary sequences will be used as input and target output of the neural network.

[0238] (VI) Training process

[0239] 1. Training data processing and batch loading

[0240] The training process uses the aforementioned 1200 pairs of signals, which are iteratively learned in batches. The data loader reads 16 pairs of training samples each time, ensuring correct pairing of the sender file and the receiver file. The loaded data includes the real and imaginary parts of the transmitted signal and the real and imaginary parts of the received signal, as well as the normalization parameters. These tensors are transmitted to the GPU device, maintaining a dimension of [16, 1024], representing a batch of 16 pairs of signal data.

[0241] 2. Forward propagation and loss calculation

[0242] Forward propagation takes the transmitted signal as input, passes through the mixed domain feature extraction, multi-scale processing, peak perception attention, etc. modules, and generates the predicted result of the received signal. The predicted result and the true received signal are fed into the peak perception loss function, which calculates the weighted loss including the real part, the imaginary part and the amplitude. The loss function particularly emphasizes the weight of the peak area, so that the network focuses on learning the key peak features in the transmission process from the transmitted signal to the received signal.

[0243] 3. Parameter optimization and learning rate scheduling

[0244] The network uses the AdamW optimizer for parameter updating, with an initial learning rate of 0.001 and a weight decay coefficient of 1x10^(-4). To prevent gradient explosion in deep networks, the system limits the L2 norm of the gradient to within 0.5. The learning rate scheduling uses the learning rate plateau decay (ReduceLROnPlateau) strategy; when the validation loss does not improve for 2 consecutive iteration periods (epochs), the learning rate is multiplied by 0.7 for decay, and the minimum learning rate is limited to 1x10^(-6). The AdamW optimizer is an improved version of the Adam optimizer, which solves the generalization defect of traditional Adam by decoupling the weight decay mechanism and has become the default optimizer for current large model training. Parameter updating follows the AdamW algorithm:

[0245]

[0246] where, represents the model parameters at the t-th training step; represents the updated model parameters at the t+1-th training step; represents the first moment estimate after bias correction at the t-th step; represents the second moment estimate after bias correction at the t-th step; η is the learning rate, and λ is the weight decay coefficient, which ensures that the model maintains good generalization ability while learning the mapping from the transmitted signal to the received signal; is the numerical stability constant in the AdamW algorithm, which prevents the denominator from being zero, and is usually set to 1x10 -8 .

[0247] where, the convolutional neural network (CNN) and peak perception attention can be replaced by the Transformer encoder and improved multi-head attention mechanism; specific alternative implementations include: position encoding replacement can use learnable position encoding to replace convolutional feature extraction; multi-head attention improvement can incorporate peak weight modulation based on standard multi-head attention; feedforward network adaptation can use MLP to replace convolutional layers for feature transformation.

[0248] Or replace the pure CNN architecture with time series modeling with CNN feature extraction and LSTM / GRU time series modeling, where LSTM refers to long short-term memory network and GRU refers to gated recurrent unit. Specific alternative implementations include: the feature extraction layer keeps the mixed domain CNN feature extraction unchanged; the time series modeling layer uses bidirectional LSTM to replace multi-scale convolution; the attention layer applies peak perception attention on the LSTM output.

[0249] 4. Early stopping mechanism and model selection

[0250] The training process employs an early stopping mechanism to prevent overfitting. The system monitors the loss variation on the validation set, as shown in Figure 9 , where the blue line represents the training loss and the orange line represents the calibration loss. Training is automatically stopped when the validation loss does not improve for 8 consecutive epochs, and the model parameters with the best validation performance are restored. This ensures that the final model has the optimal qtx to qrx mapping capability. All the above training procedures enable the entire training process to converge quickly.

[0251] Based on this, the complete process of the optical transmission received signal prediction model training method provided by the application instance is as shown in Figure 10

[0252] (Seven) Performance Evaluation Index

[0253] Normalized Mean Square Error (NMSE)

[0254] Normalized Mean Square Error (NMSE) is the core index for evaluating the accuracy of neural network channel modeling. This index calculates the mean square error between the predicted signal and the real signal, and normalizes the power of the real signal to achieve unified evaluation of signals at different power levels. NMSE is particularly suitable for evaluating the end-to-end prediction accuracy from qtx to qrx, as it objectively reflects the network's ability to model the effects of optical transmission.

[0255] The mathematical definition of normalized mean square error NMSE is:

[0256]

[0257] Where: q pred (t) represents the predicted received signal prediction data predicted by the network; q true (t) represents the real received signal; E[] represents the statistical expectation operation, which is realized by time averaging in actual calculation.

[0258] The smaller the NMSE value, the higher the prediction accuracy. The advantage of this index is that its normalization property eliminates the influence of signal power changes, allowing the modeling performance under different transmission conditions to be directly compared.

[0259] (Eight) Test Results

[0260] A transmission signal input model as shown in Figure 11 is randomly selected from the test set for testing. The real received signal corresponding to the transmission signal is as shown in Figure 12 , and the received signal prediction data predicted by the optical transmission received signal prediction model is as shown in Figure 13 ​As shown, the comparison chart between the real received signal and the prediction result data of the received signal is as shown in FIG. 6. Figure 14 As shown, in FIG. 6, the real received signal and the prediction result data of the received signal are basically coincident. Wherein, Figure 14 the ordinate of FIG. 6 represents the signal amplitude, and the abscissa represents the time (unit: nanosecond ns). Figure 11 to Figure 13 The blue solid line in FIG. 6 represents the real received signal, and the orange dotted line represents the prediction result data of the received signal. Figure 14

[0261] The NMSE training trend in FIG. 6 is as shown in FIG. 7. Figure 15

[0262] In addition, referring to FIG. 8, the test results also include the inference time comparison analysis data between the step-by-step Fourier algorithm (SSFM) and the peak perception hybrid domain NN: the inference time refers to the time required for the model to produce output prediction results from receiving input data. In the optical fiber channel modeling task, the inference time specifically refers to the calculation time from inputting the transmission signal q tx (t) to outputting the received signal q rx (t). The inference time is a key indicator for evaluating the real-time performance of the model, and directly affects the actual application value of the system. The less the inference time consumes, the higher the transmission efficiency of the system. Figure 16 That is, the embodiments and application examples of the present application provide:

[0263] (1) a peak perception attention mechanism, a double-branch parallel architecture is designed: a peak detection branch and a self-attention branch;

[0264] (2) a hybrid domain feature extraction architecture: a time domain branch for extracting local time features; a frequency domain branch: FFT transformation and four-way feature extraction (amplitude spectrum, phase spectrum, frequency domain real part, and frequency domain imaginary part); an intelligent fusion strategy for learning the optimal weight combination of time and frequency domain features; residual connection for maintaining the integrity of the original information.

[0265] (3) combined with a multi-scale CNN processing module: four-branch parallel design; the branch output channels are unified to 16 dimensions, and the total multi-scale features are 64 dimensions; 1x1 convolution fusion, residual connection, and layer normalization for stable training strategy.

[0266]

[0267] ​​​That is, the application instance of the present application aims at the problem that the transformer scheme lacks field-specific design: the application instance of the present application specially designs a peak-aware attention mechanism, converts the physical importance of the peak area in the communication system into the attention weight of the network, and realizes the deep integration of field knowledge and deep learning. In view of the problem that the feature extraction is incomplete: a hybrid domain feature extraction architecture is designed, the time domain branch captures the instantaneous characteristics, the frequency domain branch extracts four-way frequency domain features (amplitude spectrum, phase spectrum, real part, and imaginary part) through FFT, and the complete representation of the signal is realized. In view of the problem that the time scale modeling is single: a four-branch parallel multi-scale architecture is designed, 3x1, 5x1, and 7x1 convolution kernels and dilated convolution are used at the same time, and different time scales of channel effects such as instantaneous change, intersymbol interference, memory effect, and long-range dependence are captured at the same time. In view of the problem that the parameters are sensitive: through a data-driven learning method, the network automatically learns the channel characteristics from a large number of input and output signal pairs, without accurate physical parameters, and the system robustness is significantly improved.

[0268] The application instance of the present application can realize intelligent modeling of the peak-aware attention mechanism: focusing on and optimizing the signal peak area, the peak error is significantly reduced; making full use of the complementary information in the time and frequency domains: simultaneously extracting and fusing time domain and frequency domain features to improve the modeling completeness; supporting multi-scale time-dependent modeling: simultaneously capturing channel effects of different time scales; realizing real-time and efficient modeling: the inference speed is improved, the serial characteristics of SSFM limit the acceleration potential, and the technical scheme improves more than 100 times compared with the traditional SSFM method.

[0269] Based on the above embodiment and application instance of the optical fiber transmission and reception signal prediction model training method, the present application further provides an embodiment of an optical fiber transmission and reception signal prediction method, which specifically includes the following contents:

[0270] Step 300: input the transmission signal currently generated by the transmitting end in the continuous spectrum nonlinear frequency division multiplexing system into the optical fiber transmission and reception signal prediction model, so that the optical fiber transmission and reception signal prediction model outputs the corresponding reception signal prediction result data; wherein the optical fiber transmission and reception signal prediction model is trained in advance based on the optical fiber transmission and reception signal prediction model training method provided in the foregoing embodiment and / or application instance.

[0271] From the software level, the present application further provides an optical fiber transmission and reception signal prediction model training device for executing all or part of the optical fiber transmission and reception signal prediction model training method, which specifically includes the following contents:

[0272] The training data construction module is configured to generate a training sample corresponding to a transmission signal according to the transmission signal generated by a sending end in a continuous spectrum nonlinear frequency division multiplexing system in advance, and generate a label corresponding to the training sample according to a receiving signal formed after the transmission signal is transmitted to a receiving end through an optical fiber in advance;

[0273] The model training module is configured to train a neural network model based on the training sample, so that the neural network model performs mixed domain feature extraction, multi-scale processing, and attention feature extraction based on a peak value perception attention mechanism on the training sample respectively to obtain receiving signal prediction result data for predicting a receiving signal formed after the training sample is transmitted to the receiving end through the optical fiber, and optimize the neural network model based on the receiving signal prediction result data corresponding to the training sample and the label, so as to train the neural network model as an optical fiber transmission receiving signal prediction model for outputting corresponding receiving signal prediction result data according to the transmission signal. The embodiments of the optical fiber transmission receiving signal prediction model training device provided in the present application can be specifically used to execute the processing procedures of the embodiments of the optical fiber transmission receiving signal prediction model training method in the above embodiments, and the functions thereof will not be repeated here. Please refer to the detailed description of the above embodiments of the optical fiber transmission receiving signal prediction model training method.

[0274] The part of the optical fiber transmission receiving signal prediction model training performed by the optical fiber transmission receiving signal prediction model training device can be completed in a server or a client device. Specifically, the processing capacity of the client device and the limitations of the user's use scenario can be selected. The present application does not make any limitation on this. If all operations are completed in the client device, the client device can further include a processor for specific processing of the optical fiber transmission receiving signal prediction model training.

[0275] The above-mentioned client device can have a communication module (i.e. a communication unit) and can be communicatively connected with a remote server to realize data transmission with the server. The server can include a server of a task scheduling center side, and can also include a server of an intermediate platform in other implementation scenarios, such as a server of a third-party server platform communicatively connected with the server of the task scheduling center. The server can include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.

[0276] The server and the client device can communicate using any suitable network protocol, including network protocols that have not yet been developed as of the filing date of this application. The network protocol can include, for example, a TCP / IP protocol, a UDP / IP protocol, an HTTP protocol, an HTTPS protocol, and the like. Of course, the network protocol can also include, for example, a RPC protocol (Remote Procedure Call Protocol), a REST protocol (Representational State Transfer), and the like used on top of the above-mentioned protocols.

[0277] From the above description, it can be known that the optical fiber transmission and reception signal prediction model training device provided by the embodiments of the present application provides a training method for a prediction model of a received signal in a continuous spectrum nonlinear frequency division multiplexing system, which can effectively improve the reliability and effectiveness of the optical fiber transmission and reception signal prediction model training process, can realize prediction of the received signal in the continuous spectrum nonlinear frequency division multiplexing system, and can improve the accuracy of the received signal prediction using the trained optical fiber transmission and reception signal prediction model, thereby being able to predict the transmission performance of the received signal at different environmental distances in advance in the high fiber communication system modeling process, shorten the research and development cycle of the fiber communication system modeling process and reduce the computing resource consumption of the research and development equipment, effectively avoid the waste of modeling experiment cost, and provide a more effective basis for fiber communication system design, performance prediction and optimization, so as to improve the efficiency and transmission performance of the fiber communication system constructed according to the fiber communication system modeling result.

[0278] From the software level, the present application also provides an optical fiber transmission and reception signal prediction device for executing all or part of the optical fiber transmission and reception signal prediction method, which specifically includes the following contents:

[0279] A model prediction module is configured to input a transmission signal currently generated by a sending end in a continuous spectrum nonlinear frequency division multiplexing system into an optical fiber transmission and reception signal prediction model, so that the optical fiber transmission and reception signal prediction model outputs corresponding received signal prediction result data. The optical fiber transmission and reception signal prediction model is trained in advance based on the optical fiber transmission and reception signal prediction model training method provided in the foregoing embodiments.

[0280] The embodiments of the present application also provide an electronic device, which can include a processor, a memory, a receiver and a transmitter. The processor is configured to execute the optical fiber transmission and reception signal prediction model training method and / or the optical fiber transmission and reception signal prediction method mentioned in the foregoing embodiments. The processor and the memory can be connected through a bus or other means. The receiver can be connected with the processor and the memory through a wired or wireless manner.

[0281] The processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, or a combination thereof.

[0282] The memory, as a non-transitory computer readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as program instructions / modules corresponding to the optical fiber transmission received signal prediction model training method and / or the optical fiber transmission received signal prediction method in the embodiments of the present application. The processor executes various function applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory, that is, implements the optical fiber transmission received signal prediction model training method and / or the optical fiber transmission received signal prediction method in the above method embodiments.

[0283] The memory can include a program storage area and a data storage area, wherein the program storage area can store an operating system and application programs required by at least one function; and the data storage area can store data created by the processor and the like. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory disposed remotely with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0284] The one or more modules are stored in the memory, and when executed by the processor, perform the optical fiber transmission received signal prediction model training method in the embodiments.

[0285] In some embodiments of the present application, the user equipment can include a processor, a memory and a transceiver unit which can include a receiver and a transmitter, the processor, the memory, the receiver and the transmitter can be connected through a bus system, the memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to transceive signals.

[0286] As an implementation manner, the functions of the receiver and the transmitter in the present application can be implemented by a transceiver circuit or a transceiver dedicated chip, and the processor can be implemented by a dedicated processing chip, a processing circuit or a general-purpose chip.

[0287] As another implementation manner, the server provided by the embodiments of the present application can be implemented by using a general-purpose computer. That is, the program codes for implementing the functions of the processor, the receiver and the transmitter are stored in the memory, and the general-purpose processor implements the functions of the processor, the receiver and the transmitter by executing the codes in the memory.

[0288] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps of the above-mentioned optical fiber transmission received signal prediction model training method and / or optical fiber transmission received signal prediction method. The computer readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable memory disk, a CD-ROM, or any other form of storage medium known in the art.

[0289] The embodiments of the present application further provide a computer program product, which includes a computer program. The computer program is executed by a processor to implement the steps of the above-mentioned optical fiber transmission received signal prediction model training method and / or optical fiber transmission received signal prediction method.

[0290] Those skilled in the art should understand that the exemplary components, systems and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software or a combination thereof. The actual implementation depends on the specific application and design constraints imposed on the overall system. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are program or code segments used to perform the required tasks. The program or code segments can be stored in a machine readable medium or transmitted by a data signal carried in a carrier wave over a transmission medium or a communication link.

[0291] It is to be expressly understood that the application is not limited to the described and illustrated particular configurations and processes. For the sake of clarity, detailed descriptions of known methods are omitted. In the above described embodiments, several specific steps are described and illustrated as examples. However, the method processes of the application are not limited to the specific steps described and illustrated, and the skilled person can make various changes, modifications and additions, or change the order of the steps, after having understood the spirit of the application.

[0292] In this application, features described and / or illustrated with respect to one embodiment can be used in the same or similar manner in one or more other embodiments and / or combined with or substituted for features of other embodiments.

[0293] The above only describes the preferred embodiments of the application, and is not intended to limit the application. The skilled in the art can make various changes and modifications to the embodiments of the application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application shall be included in the protection scope of the application.

Claims

1. A method for training a prediction model for optical fiber transmission and reception signals, characterized in that, The method comprises the following steps: generating a training sample corresponding to a transmission signal generated in advance by a sending end in a continuous spectrum nonlinear frequency division multiplexing system, and generating a label corresponding to a receiving signal formed after the transmission signal is transmitted to a receiving end through an optical fiber in advance; training a neural network model based on the training sample, so that the neural network model performs mixed domain feature extraction, multi-scale processing, and attention feature extraction based on a peak perception attention mechanism on the training sample to obtain receiving signal prediction result data for predicting the receiving signal formed after the training sample is transmitted to the receiving end through the optical fiber, and optimizing the neural network model based on the receiving signal prediction result data corresponding to the training sample and the label, so as to train the neural network model into an optical fiber transmission receiving signal prediction model for outputting the corresponding receiving signal prediction result data according to the transmission signal; the neural network model comprises a sample processing module and a peak perception attention module; the sample processing module is used for obtaining a three-dimensional tensor corresponding to the training sample, and performing mixed domain feature extraction and multi-scale processing on the three-dimensional tensor corresponding to the training sample to obtain multi-scale feature data corresponding to the training sample; the peak perception attention module is used for performing peak detection on the multi-scale feature data corresponding to the training sample to obtain peak weight corresponding to the training sample, and performing attention detection on the multi-scale feature data to obtain attention weight corresponding to the training sample; the peak weight and the attention weight are fused to obtain corresponding fusion weight, and the fusion weight is normalized and feature aggregated to obtain attention feature data corresponding to the training sample for generating the receiving signal prediction result data. 2.The method of claim 1, wherein, the peak perception attention module comprises a peak detection branch, a standard detection branch, a peak fusion layer, a normalization layer, a feature aggregation layer, and an output convolution layer; the peak detection branch and the standard detection branch are connected to the peak fusion layer; the peak fusion layer, the normalization layer, the feature aggregation layer, and the output convolution layer are connected in sequence; the peak detection branch is used for performing peak detection on the multi-scale feature data corresponding to the training sample to obtain peak weight corresponding to the training sample; the standard detection branch is used for performing attention detection on the multi-scale feature data corresponding to the training sample based on a query vector, a key vector, and a value vector to obtain attention weight corresponding to the training sample; the peak fusion layer is used for fusing the peak weight and the attention weight to obtain corresponding fusion weight; the normalization layer is used for normalizing the fusion weight to obtain normalized fusion weight; the feature aggregation layer is used for multiplying the normalized fusion weight and the value vector to obtain corresponding weighted feature data; the output convolution layer is used for performing convolution and residual connection on the weighted feature data to obtain attention feature data corresponding to the training sample. 3.The method of claim 1, wherein, The neural network model further comprises a peak enhancement module and an output layer; The peak enhancement module is connected with the peak-aware attention module, and is configured to sequentially perform channel compression and recovery processing on the attention feature data based on two convolution layers to obtain corresponding peak enhancement feature data, and perform weighted residual connection on the peak enhancement feature data and the attention feature data to obtain peak enhancement attention feature data corresponding to the training sample; The output layer is configured to generate, according to the peak enhancement attention feature data, a received signal prediction result data for predicting a received signal formed after the training sample is transmitted to a receiving end through an optical fiber; wherein the received signal prediction result data comprises a real part prediction component and an imaginary part prediction component of the received signal. 4.The method of claim 1, wherein, The sample processing module comprises an input layer, a mixed domain feature extraction unit and a multi-scale feature extraction unit; The input layer is configured to receive the training sample, wherein the training sample comprises a real part component and an imaginary part component corresponding to the transmitted signal; and the input layer is further configured to stack the real part component and the imaginary part component corresponding to the transmitted signal in a channel dimension to obtain a three-dimensional tensor corresponding to the training sample; wherein each dimension in the three-dimensional tensor is used to represent batch size, channel number and time sequence length, respectively; The mixed domain feature extraction unit is configured to perform time domain feature extraction and frequency domain feature extraction on the three-dimensional tensor corresponding to the training sample to obtain time domain feature data and frequency domain feature data corresponding to the training sample, and perform feature fusion on the time domain feature data and the frequency domain feature data to obtain mixed domain feature data corresponding to the training sample; The multi-scale feature extraction unit is configured to perform feature extraction on the mixed domain feature data corresponding to the training sample in multiple time scales to obtain multiple time scale feature data corresponding to the training sample, and perform feature fusion processing on each time scale feature data to obtain multi-scale feature data corresponding to the training sample. 5.The method of claim 1, wherein, The training sample corresponding to the transmitted signal is generated according to a transmitted signal generated in advance by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system, and the label corresponding to the training sample is generated according to a received signal formed after the transmitted signal is transmitted to a receiving end through an optical fiber in advance, comprising: Obtaining each signal pair, wherein each signal pair comprises a transmitted signal generated in advance by a transmitting end in a continuous spectrum nonlinear frequency division multiplexing system and a received signal formed after the transmitted signal is transmitted to a receiving end through an optical fiber; Performing quantile-based robust normalization processing on the transmitted signal and the received signal in each signal pair, respectively; The robust normalized each of the transmit signal and each of the receive signal is separated into real and imaginary parts, respectively, to obtain each of the signal corresponding to the training sample, and each of the receive signal in each of the signal is used as the corresponding label of each training sample; wherein each of the training sample contains a real and imaginary part of the transmit signal; each of the label contains a real and imaginary part of the receive signal which belongs to the same signal pair with the transmit signal corresponding to the label. 6.The method of claim 3, wherein, The neural network model is optimized based on the receive signal prediction result data corresponding to the training sample and the label, comprising: In the current iteration, based on the receive signal prediction result data corresponding to the training sample and the label, a target loss corresponding to the current iteration is determined by a preset target loss function, and the neural network model is optimized based on the target loss. The target loss function is composed of a real part loss function, an imaginary part loss function, an amplitude loss function, and an amplitude loss weight coefficient corresponding to the amplitude loss function. The real part loss function is used to represent the product between the first mean square error and the peak weighted item; wherein the first mean square error includes the mean square error between the real part prediction component corresponding to the receive signal prediction result data and the real part of the receive signal corresponding to the label; the peak weighted item includes the sum between 1 and the peak product item; the peak product item includes the product between the preset peak weighted factor and the peak mask; The imaginary part loss function is used to represent the product between the second mean square error and the peak weighted item; wherein the second mean square error includes the mean square error between the imaginary part prediction component corresponding to the receive signal prediction result data and the imaginary part of the receive signal corresponding to the label; The amplitude loss function is used to represent the product between the third mean square error and the peak weighted item; wherein the third mean square error includes the mean square error between the signal amplitude of the receive signal prediction result data pre-acquired and the signal amplitude of the receive signal corresponding to the label pre-acquired.

7. A method of predicting a transmission of a received signal of an optical fiber, characterized by, It comprises: The transmit signal generated by the sending end in the continuous spectrum nonlinear frequency division multiplexing system is input into the optical fiber transmission receive signal prediction model, so that the optical fiber transmission receive signal prediction model outputs the corresponding receive signal prediction result data; wherein the optical fiber transmission receive signal prediction model is pre-trained based on the optical fiber transmission receive signal prediction model training method of any one of claims 1 to 6.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the optical fiber transmission receive signal prediction model training method of any one of claims 1 to 6, and / or implement the optical fiber transmission receive signal prediction method of claim 7.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the optical fiber transmission receive signal prediction model training method of any one of claims 1 to 6, and / or implement the optical fiber transmission receive signal prediction method of claim 7.

Citation Information

Patent Citations

  • Training method and device for multi-span optical fiber transmission signal prediction system

    CN114647976A

  • Method and device for training optical fiber transmission signal prediction model

    CN114647977A