A method for demodulating spatial division multiplexing transmission optical signals based on a neural network master-slave architecture
By employing a spatial multiplexing transmission optical signal demodulation method based on a neural network master-slave architecture, utilizing a Transformer encoder for global modeling and transfer learning, and combining structured pruning, high-precision demodulation and lightweight deployment under complex channels are achieved. This solves the problems of crosstalk suppression and large training data requirements of traditional methods, and is suitable for real-time optical communication systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI UNIV OF TECH
- Filing Date
- 2026-04-14
- Publication Date
- 2026-07-10
AI Technical Summary
Traditional optical signal demodulation methods struggle to achieve high-precision recovery in complex channels, exhibiting problems such as insufficient crosstalk suppression, high training data requirements, difficulties in model deployment, and poor generalization adaptability. These methods fail to meet the high-speed, stable, and low-power real-time demodulation requirements of space-division multiplexing optical communication systems.
A spatial multiplexing transmission optical signal demodulation method based on a neural network master-slave architecture is adopted. By constructing a Transformer encoder for global modeling, combined with transfer learning and structured pruning, lightweight deployment is achieved. A bit-level hybrid demodulation strategy is adopted to dynamically adjust the network depth and computational overhead.
It significantly improves demodulation accuracy and crosstalk suppression capability under complex channels, reduces training data requirements, improves training efficiency and the practicality of model deployment, and is suitable for real-time signal processing hardware platforms.
Smart Images

Figure CN122372101A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of optical communication signal processing technology, and in particular to a method and system for demodulating spatially multiplexed optical signals based on a neural network master-slave architecture. Background Technology
[0002] Space division multiplexing (SDM) transmission technology, through parallel transmission of multiple fiber cores and spatial modes, can significantly improve the transmission capacity of optical fiber communication systems and is one of the core technologies of next-generation high-speed optical communication systems. However, when SDM optical signals are transmitted in optical fibers, they are affected by various linear and nonlinear impairments such as inter-mode crosstalk, dispersion, polarization mode dispersion, and changes in optical signal-to-noise ratio, leading to severe distortion of the received signal. Traditional demodulation methods struggle to achieve high-precision recovery.
[0003] Traditional optical signal demodulation methods primarily rely on channel equalization and log-likelihood ratio calculations. On the one hand, when faced with strong crosstalk and complex channel impairments, it is difficult to globally model the spatiotemporal joint dependencies, limiting demodulation performance. On the other hand, training neural networks independently for multiple spatial modes requires a large amount of labeled data, resulting in high training costs and slow convergence speeds. Furthermore, complete neural network models have a large number of parameters and high computational complexity, making lightweight deployment on embedded and real-time communication hardware difficult. In addition, fixed-structure demodulation networks cannot dynamically adjust computational overhead based on channel quality, making it difficult to achieve an adaptive balance between performance and complexity. Existing technologies generally suffer from insufficient crosstalk suppression capabilities, large training data requirements, difficulties in model deployment, and poor generalization adaptability, failing to meet the high-speed, stable, and low-power real-time demodulation requirements of spatially divided multiplexing optical communication systems. Summary of the Invention
[0004] In order to address the shortcomings of the existing technology, this invention proposes a demodulation method for spatially multiplexed optical signals based on a neural network master-slave architecture. This method aims to improve the demodulation performance, training efficiency, and deployment practicality of spatially multiplexed optical signals in complex channels, thereby meeting the demodulation requirements for high-speed real-time transmission.
[0005] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: The present invention provides a method for demodulating spatially multiplexed optical signals based on a neural network master-slave architecture, characterized by the following steps: Step 1: Obtain the fiber optic cable's... Root core or first Physical parameter vectors in each spatial mode ,in, The first fiber represents the first... Root core or first The first spatial mode One physical parameter, The total number of physical parameters; Obtaining the first fiber Root core or first Transmitted symbol sequence in a spatial mode ,in, The first fiber represents the first... Root core or first The first spatial mode One transmitted symbol, and , express The first in bits, The total number of transmitted symbols ; The number of bits corresponding to each transmitted symbol. , Indicates the total number of fiber cores or spatial modes; The transmitted symbol sequence After transmission through the space-division multiplexing channel, the first fiber optic cable is obtained. Root core or first Received symbol sequences in spatial modes ,in, The first fiber represents the first... Root core or first The first spatial mode One received symbol, and , for Inter-mode crosstalk, for Gaussian white noise; Define the half length of the time window as The length of the time window is Thus constructing the first Time window ; express The central moment, and ; extract In the Time window Received symbol sequence ;in, express The Middle One received symbol; Step 2, from Select any reference mode from the spatial modes. Thus constructing a collection Individual spatial patterns and reference patterns The next spatiotemporal input matrices ;in, Indicates the first Time window The next One received symbol The real part, Indicates the first Time window The next One received symbol The imaginary part, Indicates the first Time window The next One received symbol The real part, Indicates the first Time window The next One received symbol The imaginary part, Indicates reference mode The next One physical parameter; Calculate the reference model using equation (1) Next Central symbol The true log-likelihood ratio is greater than the label The Middle Bit tag value : (1) In equation (1), Indicates reference mode Lower physical parameter vector, Indicates reference mode The next One received symbol; Indicates reference mode The following transmitted symbol sequence The first in 1 bit; Represents probability; Step 3: Construct a Transformer-based spatiotemporal joint neural network Includes: feature projection module, Layer Transformer encoder and output layer, and for The process is performed to obtain the predicted log-likelihood ratio of the label. and with Constructing the mean squared error loss function For spatiotemporal joint neural networks The spatiotemporal joint neural model was obtained after training. Step 4: Obtain the results by performing transfer learning on the trained spatiotemporal joint neural model. The target pattern model trained under each spatial pattern; Step 5, for the first The target pattern model trained under the first spatial pattern is then subjected to structured pruning to obtain the second spatial pattern model. Lightweight target pattern model under a spatial pattern; Step 6, based on , for the Each bit in the spatial mode is assigned a demodulation method to output the first bit. Complete demodulation bit sequence in each spatial mode .
[0006] The method for demodulating spatial multiplexing transmission optical signals based on a neural network master-slave architecture described in this invention is characterized in that step 3 includes the following steps: Step 3.1: The feature projection module uses equation (2) to project the first... spatiotemporal input matrices Mapping to the embedding space, we get the first Embedded sequences ,in, Indicates the model feature dimension: (2) In equation (2), This represents the weight matrix of the feature projection module. This represents the bias vector of the feature projection module; Step 3.2: The feature projection module uses equation (3) to obtain the first... Each input represents : (3) In equation (3), Indicates the location code to be added; Step 3.3, the The layer Transformer encoder utilizes multi-head self-attention and feedforward networks to... The first layer Each hidden state represents Processing yields the first... The first layer Each hidden state represents After iterating layer by layer, the first... The first layer Each hidden state represents ,when season , ; Step 3.4, extract the first The corresponding layer of hidden state representation Time window Spatiotemporal central representation vector at the central moment After projection through the output layer, the predicted log-likelihood ratio label in the reference mode is obtained using equation (4). : (4) In equation (4), Indicates reference mode Next The weight matrix of the output layer at each central time point Indicates reference mode Next The bias vector of the output layer at each center time; Step 3.5: Construct the mean squared error loss function using equation (5) : (5) In equation (5), Indicates the signal source index, used to identify the first... The signal transmitting source node corresponding to each time point express The Middle Predicted label values.
[0007] Furthermore, step 4 includes the following steps: Step 4.1: Construct the first The first spatial mode spatiotemporal input matrices and its corresponding true log-likelihood ratio label ; Step 4.2: Use the trained spatiotemporal joint neural model as the first... Target pattern network under spatial mode and will Input target pattern network Processing is performed to obtain the first... Predicted log-likelihood ratio label in each spatial pattern Therefore, the target pattern network can be constructed using equation (6). Fine-tuning loss function : (6) In equation (6), The regularization coefficient is . The output layer weights of the base network; express The j-th real label value in the data. express The j-th predicted label value in the data. Denotes the L2 norm regularization term. This represents the total number of training samples corresponding to the m-th spatial pattern; Step 4.3: Freeze the pre-trained spatiotemporal joint neural model All parameters of the layer, only the first one is fine-tuned. Layer parameters and output layer weights Thus, the target pattern network Fine-tuning training was performed to obtain the first... The target pattern model trained under each spatial pattern is thus obtained. The target pattern model trained under each spatial pattern.
[0008] Furthermore, step 5 includes the following steps: Step 5.1: Calculate the first step using equation (7). In the target pattern model trained under the spatial pattern, the first... The first layer of the Transformer encoder The importance of attention ,in, For the index of attention head, , Indicates the total number of attention heads: (7) In equation (7), It is the Frobenius norm; like If the value exceeds the set attention threshold, then the first [position] is retained. First, pay attention; otherwise, the first... Each attention head is pruned and removed, thus obtaining the set of attention heads after preliminary pruning and removal. After sorting the attention heads after initial pruning and removal in descending order of importance, they are then pruned according to a preset pruning ratio. Remove attention heads with low importance and retain attention heads with high importance to form the set of retained attention heads in the m-th spatial pattern; Step 5.2: Calculate the first step using equation (8). In the target pattern model trained under the spatial pattern, the first... The first layer of the feedforward network of the Transformer encoder The importance of each neuron : (8) like If the value is greater than the set neuron threshold, then the first neuron is retained. The first neuron, otherwise, the second neuron... One neuron is pruned and removed; thus, a set of neurons after preliminary pruning and removal is obtained. After sorting the neurons after initial pruning and removal in descending order of importance, the order is then determined based on the proportion of neurons pruned. Remove neurons of low importance; retain the set of neurons of high importance to form the retained neurons in the m-th spatial pattern; Based on the set of retained attention heads and the set of retained neurons in the m-th spatial pattern, we obtain the m-th spatial pattern. Target pattern model after pruning in a spatial pattern ; Step 5.3: Input the spatiotemporal matrix Enter the first The target model after pruning in a spatial pattern Process it and output the first... The prediction log-likelihood ratio of the pruning model in the spatial pattern Thus, the first equation (9) is used to construct the second equation. Loss function in each spatial pattern right Training and optimization were performed to finally obtain the first... The target pattern model after lightweighting under the spatial pattern is calculated using equation (9). The j-th bit in the spatial pattern ; (9).
[0009] Furthermore, step 6 includes the following steps: Step 6.1: Define bit-level threshold and with Compare and generate the first Mapping table under each space pattern : like At that time, then let the first Demodulation method of the j-th bit in each spatial mode ; like At that time, then order ;in, This indicates neural network demodulation. This indicates the demodulation algorithm based on the log-likelihood ratio. Step 6.2: Construct the first The first spatial mode One received symbol spatiotemporal input matrix , and enter Processing is performed to obtain the first... The first spatial mode The log-likelihood ratio ; Step 6.3: According to , for the Demodulation method for the j-th bit allocation in each spatial mode: like Then use Calculate the first The first spatial mode The demodulated bit value of the j-th transmitted symbol ; like ,use And calculate the value of the j-th demodulated bit. ; Step 6.4: Place the first All demodulated bit values in each spatial mode After recombination, the output is the first... Complete demodulation bit sequence in each spatial mode ,in, .
[0010] The present invention provides an electronic device, including a memory and a processor, characterized in that the memory is used to store a program supporting the processor in performing the method described therein, and the processor is configured to execute the program stored in the memory.
[0011] The present invention discloses a computer-readable storage medium storing a computer program, characterized in that the computer program is executed by a processor to perform the steps of the method described thereon.
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention adopts a master-slave neural network architecture based on Transformer encoder, utilizes a multi-head self-attention mechanism to globally model the crosstalk and time dependency between spatial multiplexing modes, and integrates the multimodal features of received signals and physical parameters, which significantly improves the demodulation accuracy and crosstalk suppression capability under complex channel impairments.
[0013] 2. This invention adopts a master-slave transfer learning mechanism to train the basic network with the reference mode and quickly adapt to other spatial multiplexing modes. Only a small amount of target mode data is needed to complete model fine-tuning, which greatly reduces the dependence of the multi-mode demodulation system on large-scale labeled data, reduces data acquisition and labeling costs, shortens the model training cycle, and improves the training efficiency and convergence speed of the multi-mode parallel demodulation system, making the demodulation scheme easier to deploy and implement in engineering optical communication receiving equipment.
[0014] 3. This invention achieves structured pruning based on attention heads and neuron importance, intelligently removes redundant network structures, compresses model size and computational complexity, realizes lightweight deployment, and is more suitable for real-time signal processing hardware platforms such as FPGA.
[0015] 4. This invention adopts a bit-level hybrid demodulation strategy, combining the advantages of lightweight neural networks and LogMAP algorithm, and dynamically adjusts the network depth and the number of attention heads according to the channel quality to achieve an adaptive balance between demodulation performance and computational overhead, thereby improving robustness and practicality in different demodulation environments. Attached Figure Description
[0016] Figure 1 This is an overall flowchart of the present invention; Figure 2 This is a diagram of the experimental framework. Figure 3 For Transformer neural network modules; Figure 4 A flowchart for transfer learning; Figure 5 Flowchart for neural network pruning. Detailed Implementation
[0017] In this embodiment, a spatial multiplexing transmission optical signal demodulation method based on a neural network master-slave architecture is proposed. This method utilizes Transformer to globally model spatiotemporal crosstalk and time dependencies, combines transfer learning to reduce the need for multi-mode training data, achieves network lightweighting through structured pruning, and employs bit-level hybrid demodulation and dynamic adaptive inference. This allows it to meet the high-speed, stable, and low-power real-time demodulation requirements of spatial multiplexing optical communication systems. Specifically, as... Figure 1 As shown, the method includes the following steps: Step 1: Obtain the fiber optic cable's... Root core or first Physical parameter vectors in each spatial mode ,in, The first fiber represents the first... Root core or first The first spatial mode One physical parameter, The total number of physical parameters.
[0018] Obtaining the first fiber Root core or first Transmitted symbol sequence in a spatial mode ,in, The first fiber represents the first... Root core or first The first spatial mode One transmitted symbol, and , express The first in bits, The total number of transmitted symbols ; The number of bits corresponding to each transmitted symbol. , This indicates the total number of fiber cores or spatial modes.
[0019] Send symbol sequence After transmission through the space-division multiplexing channel, the first fiber optic cable is obtained. Root core or first Received symbol sequences in spatial modes ,in, The first fiber represents the first... Root core or first The first spatial mode One received symbol, and , for Inter-mode crosstalk, for Gaussian white noise.
[0020] In practice, the training data for the demodulation end comes from, for example... Figure 2 The space division multiplexing transmission equipment shown is used to collect physical parameters of 6 fiber cores / 6th spatial mode, including optical signal-to-noise ratio, dispersion value, polarization mode dispersion coefficient, and transmission distance, forming a P=4-dimensional physical parameter vector. .
[0021] Define the half length of the time window as The length of the time window is Thus constructing the first Time window ; express The central moment, and ; extract In the Time window Received symbol sequence ;in, express The Middle One received symbol.
[0022] Step 2, from Select any reference mode from the spatial modes. Thus constructing a collection Individual spatial patterns and reference patterns The next spatiotemporal input matrices ; Indicates the first Time window The next One received symbol The real part, Indicates the first Time window The next One received symbol The imaginary part, Indicates the first Time window The next One received symbol The real part, Indicates the first Time window The next One received symbol The imaginary part, Indicates reference mode The next One physical parameter.
[0023] Calculate the reference model using equation (1) Next Central symbol The true log-likelihood ratio is greater than the label The Middle Bit tag value : (1) In equation (1), Indicates reference mode Lower physical parameter vector, Indicates reference mode The next One received symbol; Indicates reference mode The following transmitted symbol sequence The first in 1 bit; It represents probability.
[0024] Step 3: Construct a Transformer-based spatiotemporal joint neural network Includes: feature projection module, Layer Transformer encoder and output layer, and for The process is performed to obtain the predicted log-likelihood ratio of the label. and with Constructing the mean squared error loss function For spatiotemporal joint neural networks The trained spatiotemporal joint neural model is obtained; in this embodiment, it includes a feature projection module, a 4-layer Transformer encoder, and an output layer, as follows. Figure 3 As shown.
[0025] Step 3.1: The feature projection module uses equation (2) to project the first... spatiotemporal input matrices Mapping to the embedding space, we get the first Embedded sequences ,in, Indicates the model feature dimension: (2) In equation (2), This represents the weight matrix of the feature projection module. This represents the bias vector of the feature projection module.
[0026] Step 3.2: Add location encoding To retain time and location information: (3) in, The relative position within the time window. For dimensional indexing.
[0027] The feature projection module uses equation (3) to obtain the first... Each input represents : (4) In equation (4), This indicates the added positional encoding; the model feature dimension in this embodiment. =128, weight matrix He normal initialization, bias vector It is initialized as an all-zero vector and is adaptively updated with the loss function during network training. The relative position within the time window; in this embodiment, the window length is 7.
[0028] Step 3.3, the The layer Transformer encoder utilizes multi-head self-attention and feedforward networks to... The first layer Each hidden state represents Processing yields the first... The first layer Each hidden state represents After iterating layer by layer, the first... The first layer Each hidden state represents ,when season , .
[0029] Step 3.4, extract the first The corresponding layer of hidden state representation Time window Spatiotemporal central representation vector at the central moment After projection through the output layer, the predicted log-likelihood ratio label in the reference mode is obtained using equation (4). : (4) In equation (4), Indicates reference mode Next The weight matrix of the output layer at each central time point Indicates reference mode Next The bias vector of the output layer at each center time.
[0030] Step 3.5: Construct the mean squared error loss function using equation (5) : (5) In equation (5), Indicates the signal source index, used to identify the first... The signal transmitting source node corresponding to each time point express The Middle Predicted label values.
[0031] In this example, the Adam optimizer is used with a learning rate of 1×10−4, a batch size of 1024, and a total number of symbols K=1×106. The training continues until the loss converges, resulting in a trained spatiotemporal joint neural network model.
[0032] Step 4: Obtain the results by performing transfer learning on the trained spatiotemporal joint neural model. The target pattern model trained under each spatial pattern; Step 4.1: Construct the first The first spatial mode spatiotemporal input matrices and its corresponding true log-likelihood ratio label .
[0033] Step 4.2: Use the trained spatiotemporal joint neural model as the first... Target pattern network under spatial mode and will Input target pattern network Processing is performed to obtain the first... Predicted log-likelihood ratio label in each spatial pattern Thus, the target pattern network is constructed using equation (8). Fine-tuning loss function : (8) In equation (8), The regularization coefficient is . The output layer weights of the base network; express The j-th real label value in the data. express The j-th predicted label value in the data. Denotes the L2 norm regularization term. This represents the total number of training samples corresponding to the m-th spatial pattern.
[0034] Step 4.3: Freeze the pre-trained spatiotemporal joint neural model All parameters of the layer, only the first one is fine-tuned. Layer parameters and output layer weights Thus, the target pattern network Fine-tuning training was performed to obtain the first... The target pattern model trained under each spatial pattern is thus obtained. The target pattern model trained under each spatial pattern.
[0035] In this embodiment, λ=0.001 is used to freeze the spatiotemporal joint neural model after training. The first three layers of the Transformer encoder contain all parameters; only the first layer is fine-tuned. The layer refers to the parameters of the 4th layer Transformer encoder and the weights of the output layer. The Adam optimizer with a small learning rate is used to optimize the target pattern network. Perform a small number of rounds of fine-tuning training, such as Figure 4 As shown.
[0036] Step 5, for the first The target pattern model trained under the first spatial pattern is then subjected to structured pruning to obtain the second spatial pattern model. The target pattern model after lightweighting under the spatial pattern.
[0037] Step 5.1: Calculate the first step using equation (9). In the target pattern model trained under the spatial pattern, the first... The first layer of the Transformer encoder The importance of attention ,in For attention head index, , Indicates the total number of attention heads: (9) In equation (9), It is the Frobenius norm.
[0038] like If the value exceeds the set attention head threshold, the attention head is retained; otherwise, it is pruned and removed. According to the preset pruning ratio Remove the least important attention head and retain the more important attention heads to form the set of retained attention heads for the m-th spatial pattern. In the subsequent model inference stage, only attention heads within this set are activated, while pruned attention heads are ignored to reduce computational complexity.
[0039] Step 5.2: Calculate the first step using equation (10). In the target pattern model trained under the spatial pattern, the first... The first layer of the feedforward network of the Transformer encoder The importance of each neuron : (10) like If the value exceeds the set neuron threshold, the neuron is retained; otherwise, it is pruned and removed.
[0040] Based on neuron pruning ratio Remove the least important neurons; Based on the retained attention head set obtained in step 5.1 Thus, the first Target pattern model after pruning in a spatial pattern .
[0041] Step 5.3: Input the feature data into the first... The target model after pruning in a spatial pattern The data is processed to output the predicted log-likelihood ratio of the pruning model under this spatial pattern. Therefore, the loss function under this spatial pattern can be constructed using equation (11). right Training and optimization were performed to finally obtain the first... The target mode model after lightweighting under the first spatial mode is calculated, and the target mode model after lightweighting is calculated. The j-th bit in the spatial pattern ; (11) In this embodiment, H=8, and the pruning ratio is set. Remove the 30% of attention heads with the lowest importance and retain the set of attention heads with higher importance. Set the neuron pruning ratio ,like Figure 5 As shown.
[0042] Step 6: Based on the pruned network Assign a demodulation method to the j-th bit to output the j-th bit. Complete demodulation bit sequence in each spatial mode ; Step 6.1: For each pattern Define bit-level threshold ,according to and The comparison results generate the first... Mapping table under each space pattern ; The specific generation rules are as follows: hour, , hour, ;in, Indicates the first The demodulation method of the j-th bit in each spatial mode, and ; This indicates neural network demodulation. This indicates the demodulation algorithm based on the log-likelihood ratio.
[0043] Step 6.2: Construct the first The first spatial mode One received symbol spatiotemporal input matrix And input the pruned pattern model. Processing is performed to obtain the first... The first spatial mode The log-likelihood ratio ; Step 6.3: According to Demodulation method assigned to the j-th bit: like Then use The bit value was calculated. ; like ,use And calculate the bit value ; Step 6.4: Place the first All bit values in each space mode After recombination, the output is the first... Complete demodulation bit sequence in each spatial mode ,in, .
[0044] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.
[0045] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.
Claims
1. A method for demodulating spatially multiplexed optical signals based on a neural network master-slave architecture, characterized in that, Includes the following steps: Step 1: Obtain the fiber optic cable's... Root core or first Physical parameter vectors in each spatial mode ,in, The first fiber represents the first... Root core or first The first spatial mode One physical parameter, The total number of physical parameters; Obtaining the first fiber Root core or first Transmitted symbol sequence in a spatial mode ,in, The first fiber represents the first... Root core or first The first spatial mode One transmitted symbol, and , express The first in bits, The total number of transmitted symbols ; The number of bits corresponding to each transmitted symbol. , Indicates the total number of fiber cores or spatial modes; The transmitted symbol sequence After transmission through the space-division multiplexing channel, the first fiber optic cable is obtained. Root core or first Received symbol sequences in spatial modes ,in, The first fiber represents the first... Root core or first The first spatial mode One received symbol, and , for Inter-mode crosstalk, for Gaussian white noise; Define the half length of the time window as The length of the time window is Thus constructing the first Time window ; express The central moment, and ; extract In the Time window Received symbol sequence ;in, express The Middle One received symbol; Step 2, from Select any reference mode from the spatial modes. Thus constructing a collection Individual spatial patterns and reference patterns The next spatiotemporal input matrices ;in, Indicates the first Time window The next One received symbol The real part, Indicates the first Time window The next One received symbol The imaginary part, Indicates the first Time window The next One received symbol The real part, Indicates the first Time window The next One received symbol The imaginary part, Indicates reference mode The next One physical parameter; Calculate the reference model using equation (1) Next Central symbol The true log-likelihood ratio is greater than the label The Middle Bit tag value : (1) In equation (1), Indicates reference mode Lower physical parameter vector, Indicates reference mode The next One received symbol; Indicates reference mode The following transmitted symbol sequence The first in 1 bit; Represents probability; Step 3: Construct a Transformer-based spatiotemporal joint neural network Includes: feature projection module, Layer Transformer encoder and output layer, and for The process is performed to obtain the predicted log-likelihood ratio of the label. and with Constructing the mean squared error loss function For spatiotemporal joint neural networks The spatiotemporal joint neural model was obtained after training. Step 4: Obtain the results by performing transfer learning on the trained spatiotemporal joint neural model. The target pattern model trained under each spatial pattern; Step 5, for the first The target pattern model trained under the first spatial pattern is then subjected to structured pruning to obtain the second spatial pattern model. Lightweight target pattern model under a spatial pattern; Step 6, based on , for the Each bit in the spatial mode is assigned a demodulation method to output the first bit. Complete demodulation bit sequence in each spatial mode .
2. The method for demodulating spatially multiplexed optical signals based on a neural network master-slave architecture according to claim 1, characterized in that, Step 3 includes the following steps: Step 3.1: The feature projection module uses equation (2) to project the first... spatiotemporal input matrices Mapping to the embedding space, we get the first Embedded sequences ,in, Indicates the model feature dimension: (2) In equation (2), This represents the weight matrix of the feature projection module. This represents the bias vector of the feature projection module; Step 3.2: The feature projection module uses equation (3) to obtain the first... Each input represents : (3) In equation (3), Indicates the location code to be added; Step 3.3, the The layer Transformer encoder utilizes multi-head self-attention and feedforward networks to... The first layer Each hidden state represents Processing yields the first... The first layer Each hidden state represents After iterating layer by layer, the first... The first layer Each hidden state represents ,when season , ; Step 3.4, extract the first The corresponding layer of hidden state representation Time window Spatiotemporal central representation vector at the central moment After projection through the output layer, the predicted log-likelihood ratio label in the reference mode is obtained using equation (4). : (4) In equation (4), Indicates reference mode Next The weight matrix of the output layer at each central time point Indicates reference mode Next The bias vector of the output layer at each center time; Step 3.5: Construct the mean squared error loss function using equation (5) : (5) In equation (5), Indicates the signal source index, used to identify the first... The signal transmitting source node corresponding to each time point express The Middle Predicted label values.
3. The method for demodulating spatially multiplexed optical signals based on a neural network master-slave architecture according to claim 2, characterized in that, Step 4 includes the following steps: Step 4.1: Construct the first The first spatial mode spatiotemporal input matrices and its corresponding true log-likelihood ratio label ; Step 4.2: Use the trained spatiotemporal joint neural model as the first... Target pattern network under spatial mode and will Input target pattern network Processing is performed to obtain the first... Predicted log-likelihood ratio label in each spatial pattern Therefore, the target pattern network can be constructed using equation (6). Fine-tuning loss function : (6) In equation (6), The regularization coefficient is . The output layer weights of the base network; express The j-th real label value in the data. express The j-th predicted label value in the data. Denotes the L2 norm regularization term. Indicates the first The total number of training samples corresponding to each spatial pattern; Step 4.3: Freeze the pre-trained spatiotemporal joint neural model All parameters of the layer, only the first one is fine-tuned. Layer parameters and output layer weights Thus, the target pattern network Fine-tuning training was performed to obtain the first... The target pattern model trained under each spatial pattern is thus obtained. The target pattern model trained under each spatial pattern.
4. The method for demodulating spatially multiplexed optical signals based on a neural network master-slave architecture according to claim 3, characterized in that, Step 5 includes the following steps: Step 5.1: Calculate the first step using equation (7). In the target pattern model trained under the spatial pattern, the first... The first layer of the Transformer encoder The importance of attention ,in, For the index of attention head, , Indicates the total number of attention heads: (7) In equation (7), It is the Frobenius norm; like If the value exceeds the set attention threshold, then the first [position] is retained. First, pay attention; otherwise, the first... Each attention head is pruned and removed, thus obtaining the set of attention heads after preliminary pruning and removal. After sorting the attention heads after initial pruning and removal in descending order of importance, they are then pruned according to a preset pruning ratio. Remove low-importance attention heads and retain high-importance attention heads to form the second... Preservation of attention head set in a spatial pattern; Step 5.2: Calculate the first step using equation (8). In the target pattern model trained under the spatial pattern, the first... The first layer of the feedforward network of the Transformer encoder The importance of each neuron : (8) like If the value is greater than the set neuron threshold, then the first neuron is retained. The first neuron, otherwise, the second neuron... One neuron is pruned and removed; thus, a set of neurons after preliminary pruning and removal is obtained. After sorting the neurons after initial pruning and removal in descending order of importance, the order is then determined based on the proportion of neurons pruned. Remove neurons of low importance; retain the set of neurons of high importance to form the second set. Preserved neurons in a spatial pattern; According to the The set of attention heads and the set of neurons preserved in the spatial pattern are obtained to obtain the first spatial pattern. Target pattern model after pruning in a spatial pattern ; Step 5.3: Input the spatiotemporal matrix Enter the first The target model after pruning in a spatial pattern Process it and output the first... The prediction log-likelihood ratio of the pruning model in the spatial pattern Thus, the first equation (9) is used to construct the second equation. Loss function in each spatial pattern right Training and optimization were performed to finally obtain the first... The target pattern model after lightweighting under the spatial pattern is calculated using equation (9). The j-th bit in the spatial pattern ; (9)。 5. The method for demodulating spatially multiplexed optical signals based on a neural network master-slave architecture according to claim 4, characterized in that, Step 6 includes the following steps: Step 6.1: Define bit-level threshold and with Compare and generate the first Mapping table under each space pattern : like At that time, then let the first Demodulation method of the j-th bit in each spatial mode ; like At that time, then order ;in, This indicates neural network demodulation. This indicates the demodulation algorithm based on the log-likelihood ratio. Step 6.2: Construct the first The first spatial mode One received symbol spatiotemporal input matrix , and enter Processing is performed to obtain the first... The first spatial mode The log-likelihood ratio ; Step 6.3: According to , for the Demodulation method for the j-th bit allocation in each spatial mode: like Then use Calculate the first The first spatial mode The demodulated bit value of the j-th transmitted symbol ; like ,use And calculate the value of the j-th demodulated bit. ; Step 6.4: Place the first All demodulated bit values in each spatial mode After recombination, the output is the first... Complete demodulation bit sequence in each spatial mode ,in, .
6. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports a processor in executing the method of any one of claims 1-5, the processor being configured to execute the program stored in the memory.
7. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program is executed by the processor to perform the steps of the method according to any one of claims 1-5.