An end-to-end self-decoding optimizer based on reinforcement learning and system

CN122533665APending Publication Date: 2026-08-07BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING INST OF TECH
Filing Date
2026-05-11
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

旨在解决传统极化码SCL译码中计算复杂度高、路径管理策略固定僵化、无法适应动态信道条件的问题

Benefits of technology

(1)突破传统模块化优化局限,将发送端、信道和接收端视为统一整体,通过端到端训练实现全局最优译码性能,形成自译码器架构。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122533665A_ABST
    Figure CN122533665A_ABST
Patent Text Reader

Abstract

The application discloses an end-to-end self-decoding optimizer based on reinforcement learning and a system thereof, applied to the technical field of high-speed optical communication, comprising: an IM / DD optical interconnection system experimental platform, which transmits a polar code encoded signal modulated by PAM4, takes a received signal as input, transmits a signal as output, and trains an IM / DD optical fiber channel model based on a bidirectional long short-term memory network; a channel log-likelihood ratio of a model output signal is calculated and input to a self-decoding optimizer based on reinforcement learning assistance for end-to-end training, and finally an end-to-end optimized self-decoding optimizer is formed; wherein the self-decoding optimizer inlays a reinforcement learning intelligent agent as an auxiliary decision mechanism, learns the mapping from a state space to an action space through a Q network, dynamically adjusts the scoring weight and the pruning threshold, and comprehensively considers the decoding correctness and the calculation complexity through a reward function. The application significantly improves the decoding reliability and system robustness while reducing the calculation complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of high-speed optical communication technology, and more specifically to an end-to-end self-decoder optimization method and system based on reinforcement learning. Background Technology

[0002] High-speed intensity-modulated direct-detection (IM / DD) optical interconnect systems have become the cornerstone of cost-effective and energy-efficient connectivity in short-range data centers and access networks. As these systems expand to ultra-high-speed baud rates and higher-order modulation formats, the resulting signal impairments (including attenuation, dispersion, and nonlinear effects) necessitate robust forward error correction (FEC) techniques to ensure reliable data transmission under stringent power and delay constraints.

[0003] Polar codes, possessing both the potential to approximate channel capacity and a hardware-efficient structure, have attracted considerable attention and are widely adopted in modern communication systems. Polar codes approximate channel capacity by leveraging channel polarization phenomena, employing a recursive transformation based on the Kronecker generator matrix during encoding. At the decoding end, the Sequential Elimination (SC) algorithm has low computational complexity, but its error correction performance is limited under finite code lengths. Sequential Elimination List (SCL) decoding significantly improves decoding reliability by maintaining multiple candidate decoding paths, and Cyclic Redundancy Check-assisted SCL (CA-SCL) further enhances path selection accuracy. However, SCL-based decoders have inherent drawbacks: increasing list size leads to a significant increase in computational complexity and memory consumption, and the number of paths grows proportionally with the list size. Existing optimization methods mainly rely on fixed rules or manually designed pruning thresholds, which are often not optimal under dynamically changing channel conditions and hardware constraints, failing to meet the real-time requirements of high-speed IM / DD optical interconnect scenarios.

[0004] In recent years, machine learning techniques have been explored to enhance the performance of polarization decoding, with neural network decoders, including fully connected networks, convolutional neural networks, and recurrent architectures, proposed to replace traditional SC decoding. Hybrid model-driven and data-driven methods embed trainable parameters into belief propagation or SCL decoding frameworks. Reinforcement learning (RL) has also been applied to guide bit-flipping strategies, optimize decoding order, or adjust decoding parameters. End-to-end (E2E) learning strategies treat the transmitter, channel, and receiver as a complete system for joint optimization, promising to overcome the performance bottlenecks of modular designs.

[0005] Specifically, in IM / DD optical interconnect systems, existing technologies have not deeply integrated end-to-end learning frameworks with polar code SCL decoding, resulting in decoders lacking self-learning and adaptive capabilities, and struggling to simultaneously meet low latency and high reliability requirements. There is a need to explore a collaborative mechanism between reinforcement learning intelligent decision-making and end-to-end joint optimization, and to develop self-decoder algorithms for high-speed optical interconnects to enhance the system's online adaptability and decoding robustness under dynamic channel conditions, especially in high-speed communication systems where low latency and high reliability are critical. Therefore, how to provide a reinforcement learning-based end-to-end self-decoder optimization method and system is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] In view of this, the present invention provides an end-to-end self-decoder optimization method and system based on reinforcement learning. It aims to solve the problems of high computational complexity, fixed and rigid path management strategies, and inability to adapt to dynamic channel conditions in traditional polar code SCL decoding. By constructing an end-to-end learning framework to train the transmitter, channel, and receiver of the optical fiber communication system as a whole, and utilizing reinforcement learning agents to adaptively control path expansion and pruning decisions in the SCL decoding process, the invention achieves joint optimization of decoding reliability and computational complexity, ultimately forming an adaptive end-to-end self-decoder.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: An end-to-end self-decoder optimization method based on reinforcement learning includes: Step 1: Based on the IM / DD optical interconnect system experimental platform, the polar code encoded signal modulated by PAM4 is transmitted multiple times, and the PAM4 digital time domain sampled signal is received. The received signal is used as input and the transmitted signal is used as output to train the IM / DD optical fiber channel model based on the bidirectional long short-term memory network. Step 2: Calculate the channel log-likelihood ratio of the model output signal and input it into the reinforcement learning-assisted self-decoder for end-to-end training, ultimately forming an end-to-end optimized self-decoder. The self-decoder embeds a reinforcement learning agent as an auxiliary decision-making mechanism, learns the mapping from the state space to the action space through the Q-network, dynamically adjusts the scoring weights and pruning thresholds, and comprehensively considers decoding correctness and computational complexity through the reward function during training.

[0008] Optionally, in step 1, the IM / DD optical interconnect system experimental platform specifically refers to: Includes: arbitrary waveform generator, electrical amplifier, laser, Mach-Zehnder modulator, single-mode fiber, variable optical attenuator, photodetector and real-time oscilloscope; At the system's transmitting end, a pseudo-random binary sequence is generated as an information bit sequence. After being encoded by polar codes, it is mapped to PAM4 symbols and transmitted through an arbitrary waveform generator. The PAM4 electrical signal is amplified by an electrical amplifier and then modulated onto an optical carrier of a preset wavelength using a Mach-Zehnder modulator. The modulated optical signal is transmitted on a single-mode optical fiber and controlled by a variable optical attenuator to adjust the received optical power. Finally, the waveform is converted by a photodetector and acquired by a real-time oscilloscope at a preset sampling rate for digital signal processing.

[0009] Optionally, at the system's transmitting end, a pseudo-random binary sequence is generated as the information bit sequence, which is then encoded using polar codes and mapped to PAM4 symbols, specifically: Length is Information bit sequence Codewords generated by polar code encoding The length is obtained Encoding sequence ;in, for Second Kronecker power, ; In the encoded sequence In, including The set of most reliable sub-channels transmitting information bits and A set of frozen bits The position and value of the freeze bit are known in both the encoder and decoder.

[0010] Optionally, in step 1, after receiving the PAM4 digital time-domain sampled signal, the following steps are also included: The PAM4 digital time-domain sampled signal is resampled and subjected to decision-guided minimum mean square equalization.

[0011] Optionally, in step 2, the channel log-likelihood ratio of the model output signal is calculated as follows:

[0012] in, Input information bit sequence to the model The corresponding model output signal; Output signal for the model The corresponding channel log-likelihood ratio; This represents the conditional probability.

[0013] Optionally, in step 2, the mapping from the state space to the action space is learned through the Q-network, and the scoring weights and pruning thresholds are dynamically adjusted, specifically as follows:

[0014] in, For the goal value; For the reward function; Discount factor; For state space; For action; State-action value; The current Q-network is updated by minimizing the mean squared error loss function, as follows:

[0015] in, This represents the mean squared error loss value; The number of decoding steps; For the first Each decoding step.

[0016] Optionally, in step 2, the state space is designed as follows: In the decoding steps The agent observes the state vector , It is composed of path metric, normalized LLR, and decoding progress, as follows:

[0017] in, , These are the decoding steps. , No. The path metric and normalized log-likelihood ratio for each path; For normalized decoding progress.

[0018] Optionally, in step 2, the motion space design is as follows: Agent selects action Control path selection behavior, when When the original path metric derived from the SCL decoding process is directly used, it corresponds to a conservative decision; when At the same time, the application of perturbation metrics encourages the exploration of path selection, as follows:

[0019] in, For small random perturbations; For decoding steps , No. The corrected path metric for the path; The path metric weight coefficient; For decoding steps , No. The original path metric for each path.

[0020] Optionally, in step 2, the reward function is designed as follows: The agent evaluates its decisions based on reward signals that reflect the reliability and confidence of the decoding. The instantaneous reward for each step is defined as follows:

[0021] in, For decoding steps Instant rewards; For decoding steps , No. The normalized log-likelihood ratio of the paths; For decoding steps , No. The original path metric for each path.

[0022] This invention also provides a reinforcement learning-based end-to-end self-decoder optimization system utilizing a reinforcement learning-based end-to-end self-decoder optimization method, comprising: IM / DD Fiber Channel Model Training Module: This module is used to train an IM / DD fiber channel model based on a bidirectional long short-term memory network by repeatedly transmitting polar code-encoded signals modulated by PAM4 and receiving PAM4 digital time-domain sampled signals from the experimental platform of the IM / DD optical interconnect system. The received signal is used as the input and the transmitted signal is used as the output. The end-to-end optimized self-decoder training module is used to calculate the channel log-likelihood ratio of the model output signal, which is then input into the reinforcement learning-assisted self-decoder for end-to-end training, ultimately forming an end-to-end optimized self-decoder. The self-decoder embeds a reinforcement learning agent as an auxiliary decision-making mechanism, which learns the mapping from the state space to the action space through a Q-network, dynamically adjusts the scoring weights and pruning thresholds, and comprehensively considers decoding correctness and computational complexity through a reward function during training.

[0023] As can be seen from the above technical solution, compared with the prior art, this invention discloses an end-to-end self-decoder optimization method and system based on reinforcement learning. By modeling the IM / DD optical fiber channel using a Bi-LSTM neural network, the transmitting end, channel, and receiving end of the optical fiber system are jointly constructed into an end-to-end environment. Through continuous interaction between the reinforcement learning agent and the decoding environment, the agent autonomously learns the optimal path management strategy, forming an adaptive end-to-end self-decoder. This invention achieves the following beneficial effects compared with existing polar code SCL decoding methods: (1) Breaking through the limitations of traditional modular optimization, the transmitter, channel and receiver are regarded as a unified whole, and the global optimal decoding performance is achieved through end-to-end training, forming a self-decoder architecture.

[0024] (2) By using reinforcement learning agents to adaptively regulate path expansion and pruning decisions in the SCL decoding process, replacing traditional fixed rules, dynamic adaptation of the decoding strategy is achieved, significantly improving robustness to channel changes.

[0025] (3) Through end-to-end joint optimization, the agent learns a stable and effective path ranking mechanism, suppresses excessive metric fluctuations, reduces path replacement and decoder copying operations, and achieves a favorable balance between reliability and complexity.

[0026] (4) The effects of dispersion and nonlinearity in IM / DD fiber optic systems are considered. Through the coordinated optimization of data-driven channel model and intelligent decoding strategy, adaptive decoding of polar codes in high-speed optical interconnect systems is realized. It is particularly suitable for optical interconnect application scenarios with high throughput, low latency and limited computing resources. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0028] Figure 1 This is a schematic diagram of the method flow provided by the present invention.

[0029] Figure 2 This is a schematic diagram of signal transmission for the experimental platform of the IM / DD optical interconnect system provided by the present invention.

[0030] Figure 3 This is a schematic diagram illustrating the comparison results of bit error rate performance provided by the present invention.

[0031] Figure 4 This is a schematic diagram illustrating the performance comparison results under conditions of equal complexity provided by the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Example 1: Embodiment 1 of this invention discloses an end-to-end self-decoder optimization method based on reinforcement learning, such as... Figure 1 As shown, it includes: Step 1: Based on the IM / DD optical interconnect system experimental platform, the polar code-encoded signal modulated by PAM4 is transmitted multiple times, and the PAM4 digital time-domain sampled signal is received. The received signal is used as input and the transmitted signal is used as output to train the IM / DD optical fiber channel model based on the bidirectional long short-term memory network (Bi-LSTM). Through joint processing of forward and backward hidden states, the attenuation, dispersion and nonlinear impairments in the channel are accurately characterized.

[0034] IM / DD optical interconnect system experimental platform, such as Figure 2 As shown, specifically: Includes: Arbitrary Waveform Generator (AWG), Electrical Amplifier (EA), Laser, Mach-Zehnder Modulator (MZM), Single-Mode Fiber (SMF), Variable Optical Attenuator (VOA), Photodetector (PD), and Real-Time Oscilloscope (OSC). At the system's transmitting end, a pseudo-random binary sequence (PRBS) is generated as an information bit sequence. After being encoded by polar codes, it is mapped to PAM4 symbols and transmitted through an arbitrary waveform generator at a sampling rate of 120 Gsa / s. The PAM4 electrical signal is amplified by an electrical amplifier with a bandwidth of 3dB at 40 GHz and then modulated onto an optical carrier with a preset wavelength (1550 nm in this embodiment) using a Mach-Zehnder modulator. The modulated optical signal is transmitted over a 1-kilometer single-mode optical fiber and controlled by a variable optical attenuator to adjust the received optical power. Finally, the waveform is converted by a 40 GHz photodetector and acquired by a real-time oscilloscope at a preset sampling rate (128 Gsa / s in this embodiment) for digital signal processing.

[0035] At the system's transmitting end, a pseudo-random binary sequence is generated as the information bit sequence, which is then encoded using polar codes and mapped to PAM4 symbols, specifically: Length is Information bit sequence Codewords generated by polar code encoding The length is obtained Encoding sequence ;in, for Second Kronecker power, ; In the encoded sequence In, including The set of most reliable sub-channels transmitting information bits and A set of frozen bits The position and value of the frozen bit are known in both the encoder and decoder (set to 0). In this embodiment of the invention, the Polar (128,96) polar code is selected, where K=96 represents the most reliable subchannel transmission information bit, and NK=32 represents the frozen bit and is fixed at 0.

[0036] After receiving the PAM4 digital time-domain sampled signal, the process also includes: The PAM4 digital time-domain sampled signal is resampled and subjected to decision-guided minimum mean square equalization.

[0037] The Bi-LSTM channel model of this invention is trained using 1.8 million samples, independently validated using 360,000 samples, and optimized using the Adam optimizer. After training the IM / DD fiber optic channel model based on a bidirectional long short-term memory network, it also includes: The channel model is validated using test set data: The output of the Bi-LSTM channel model is compared with the measured data to evaluate the modeling error distribution and ensure the accuracy and generalization ability of the model.

[0038] Step 2: Calculate the channel log-likelihood ratio of the model output signal and input it into the reinforcement learning-assisted self-decoder for end-to-end training, ultimately forming an end-to-end optimized self-decoder. The self-decoder embeds a reinforcement learning agent as an auxiliary decision-making mechanism (the agent is tightly integrated into the SCL decoding loop and regulates the survival state of candidate paths in each decoding step). It learns the mapping from the state space to the action space through the Q network, dynamically adjusts the scoring weights and pruning thresholds, and comprehensively considers decoding correctness and computational complexity through the reward function during training.

[0039] The channel log-likelihood ratio of the model output signal is calculated and used as the input soft information for SCL decoding, as follows:

[0040] in, Input information bit sequence to the model The corresponding model output signal; Output signal for the model The corresponding channel log-likelihood ratio; This represents the conditional probability.

[0041] The Q-network learns the mapping from the state space to the action space, and dynamically adjusts the scoring weights and pruning thresholds, specifically as follows:

[0042] in, For the goal value; For the reward function; Discount factor; For state space; For action; State-action value; The current Q-network is updated by minimizing the mean squared error loss function, as follows:

[0043] in, This represents the mean squared error loss value. The number of decoding steps; For the first Each decoding step.

[0044] The state space is designed as follows: In the decoding steps The agent observes the state vector , It is composed of path metric (PM), normalized LLR, and decoding progress, as follows:

[0045] in, , These are the decoding steps. , No. The path metric and normalized log-likelihood ratio for each path; This represents the normalized decoding progress. This state indicates the location-dependent characteristics of joint capture-cumulative decoding reliability, instantaneous bit confidence, and channel polarization.

[0046] The design of the Action space is as follows: Agent selects action Control path selection behavior, when When the original path metric derived from the SCL decoding process is directly used, it corresponds to a conservative decision; when At the same time, the application of perturbation metrics encourages the exploration of path selection, as follows:

[0047] in, For small random perturbations; For decoding steps , No. The corrected path metric for the path; The path metric weight coefficient; For decoding steps , No. The original path metric for each path. This action design enables the agent to balance the exploration of reliable decoding paths with the exploration of regions of decoding uncertainty.

[0048] The reward function is designed as follows: The agent evaluates its decisions based on reward signals that reflect the reliability and confidence of the decoding. The instantaneous reward for each step is defined as follows:

[0049] in, For decoding steps Instant rewards; For decoding steps , No. The normalized log-likelihood ratio of the paths; For decoding steps , No. The original path metric for each path. This reward encourages the retention of paths with lower cumulative uncertainty and higher bit-level reliability, while penalizing unnecessary path expansion. The reward is set to zero when the decoding process reaches a termination state.

[0050] In polarization decoding, the receiver first calculates the channel LLR, which is propagated through the SC decoding graph and recursively updated according to the polarization code factor graph structure. Hard decision is then performed based on the updated LLR.

[0051]

[0052] To mitigate the inherent error propagation in SC decoding, SCL decoding maintains multiple decoding candidates in parallel. At each information bit position, the current decoding path branch has two candidates ( =0 and =1), each path is associated with a path metric that quantifies its likelihood. Common path metric update rules are:

[0053] After path expansion, only the L most reliable paths (with the lowest PM value) are retained, and the rest are pruned.

[0054] Unlike traditional deterministic metric-driven pruning in SCL decoding, this invention embeds a reinforcement learning agent into the SCL decoding process. The agent participates in the path selection phase, observing the state comprised of the current PM, normalized LLR, and decoding progress, and outputs an action to determine whether to adopt a conventional PM selection or a perturbation path scoring strategy.

[0055] Through continuous interaction, the agent learns to balance reliable paths with uncertain decoding exploration, reducing the risk of discarding the correct path in the early stages and enhancing robustness to LLR distortion and channel uncertainty.

[0056] The end-to-end training process of the self-decoder of this invention is as follows: Information bits are encoded with polar codes and modulated with PAM4, then the received signal is generated using a Bi-LSTM channel model. The received signal is fed into a reinforcement learning-assisted self-decoder after LLR calculation. The agent observes the state, performs actions, and receives rewards at each decoding step. The path management strategy is adaptively adjusted through gradient backpropagation and Q-network updates. During training, the agent dynamically balances exploration and exploitation, gradually learning an adaptive decoding strategy to adapt to channel conditions, ultimately forming an end-to-end optimized self-decoder.

[0057] The trained end-to-end self-decoder was applied to test data to evaluate performance metrics such as bit error rate (BER). Compared with the traditional SCL decoding method, the received optical power (ROP) gain and complexity were calculated and compared at the same BER.

[0058] like Figure 3 As shown, the BER performance of the traditional SCL and the end-to-end self-decoder for Polar (128,96) codes varies with ROP under different list sizes L. For both decoding schemes, increasing the list size improves BER performance. However, for a given list size, the end-to-end self-decoder consistently outperforms the traditional SCL across the entire ROP range. Significantly, at a target BER of 10... -4 At this point, the end-to-end self-decoder achieves a ROP gain of approximately 0.15 dB.

[0059] like Figure 4 As shown, the performance of the two decoding schemes is further compared under the constraint of equal pruning operations. Under the same pruning budget, the end-to-end self-decoder allows for a larger list size, thus achieving better BER performance than the traditional SCL. Specifically, the end-to-end self-decoder with L=8 has similar pruning complexity to the traditional SCL with L=6, but achieves better BER with L=10. -4 It provides an additional ROP gain of approximately 0.2 dB.

[0060] Example 2: Embodiment 2 of the present invention discloses a reinforcement learning-based end-to-end self-decoder optimization system, comprising: IM / DD Fiber Channel Model Training Module: This module is used to train an IM / DD fiber channel model based on a bidirectional long short-term memory network by repeatedly transmitting polar code-encoded signals modulated by PAM4 and receiving PAM4 digital time-domain sampled signals from the experimental platform of the IM / DD optical interconnect system. The received signal is used as the input and the transmitted signal is used as the output. The end-to-end optimized self-decoder training module is used to calculate the channel log-likelihood ratio of the model output signal, which is then input into the reinforcement learning-assisted self-decoder for end-to-end training, ultimately forming an end-to-end optimized self-decoder. The self-decoder embeds a reinforcement learning agent as an auxiliary decision-making mechanism, which learns the mapping from the state space to the action space through a Q-network, dynamically adjusts the scoring weights and pruning thresholds, and comprehensively considers decoding correctness and computational complexity through a reward function during training.

[0061] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0062] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A reinforcement learning-based end-to-end self-decoder optimization method, characterized in that, include: Step 1: Based on the IM / DD optical interconnect system experimental platform, the polar code encoded signal modulated by PAM4 is transmitted multiple times, and the PAM4 digital time domain sampled signal is received. The received signal is used as input and the transmitted signal is used as output to train the IM / DD optical fiber channel model based on the bidirectional long short-term memory network. Step 2: Calculate the channel log-likelihood ratio of the model output signal and input it into the reinforcement learning-assisted self-decoder for end-to-end training, ultimately forming an end-to-end optimized self-decoder; wherein, the self-decoder embeds a reinforcement learning agent as an auxiliary decision-making mechanism, learns the mapping from the state space to the action space through the Q network, dynamically adjusts the scoring weights and pruning thresholds, and comprehensively considers decoding correctness and computational complexity through the reward function during the training process.

2. The end-to-end self-decoder optimization method based on reinforcement learning according to claim 1, characterized in that, In step 1, the IM / DD optical interconnect system experimental platform specifically refers to: Includes: arbitrary waveform generator, electrical amplifier, laser, Mach-Zehnder modulator, single-mode fiber, variable optical attenuator, photodetector and real-time oscilloscope; At the system's transmitting end, a pseudo-random binary sequence is generated as an information bit sequence. After being encoded by polar codes, it is mapped to PAM4 symbols and transmitted through the arbitrary waveform generator. The PAM4 electrical signal is amplified by the electrical amplifier and then modulated onto an optical carrier of a preset wavelength using the Mach-Zehnder modulator. The modulated optical signal is transmitted on the single-mode optical fiber and controlled by the variable optical attenuator to adjust the received optical power. Finally, the waveform is converted by the photodetector and acquired by the real-time oscilloscope at a preset sampling rate for digital signal processing.

3. The end-to-end self-decoder optimization method based on reinforcement learning according to claim 2, characterized in that, At the system's transmitting end, a pseudo-random binary sequence is generated as the information bit sequence, which is then encoded using polar codes and mapped to PAM4 symbols, specifically: Length is Information bit sequence Codewords generated by polar code encoding The length is obtained as Encoding sequence ;in, for Second Kronecker power, ; In the encoded sequence In, including The set of information bits transmitted by the most reliable sub-channels and A set of frozen bits The position and value of the freeze bit are known in both the encoder and decoder.

4. The end-to-end self-decoder optimization method based on reinforcement learning according to claim 1, characterized in that, Step 1, after receiving the PAM4 digital time-domain sampling signal, further includes: The PAM4 digital time-domain sampled signal is resampled and subjected to decision-guided minimum mean square equalization.

5. The end-to-end self-decoder optimization method based on reinforcement learning according to claim 1, characterized in that, In step 2, the channel log-likelihood ratio of the model output signal is calculated as follows: in, Input information bit sequence to the model The corresponding model output signal; Output signal for the model The corresponding channel log-likelihood ratio; This represents the conditional probability.

6. The end-to-end self-decoder optimization method based on reinforcement learning according to claim 1, characterized in that, In step 2, the mapping from the state space to the action space is learned through the Q-network, and the scoring weights and pruning thresholds are dynamically adjusted, specifically as follows: in, For the goal value; For the reward function; Discount factor; For state space; For action; State-action value; The current Q-network is updated by minimizing the mean squared error loss function, as follows: in, This represents the mean squared error loss value. The number of decoding steps; For the first Each decoding step.

7. The end-to-end self-decoder optimization method based on reinforcement learning according to claim 1, characterized in that, In step 2, the design of the state space is as follows: In the decoding steps The agent observes the state vector , It is composed of path metric, normalized LLR, and decoding progress, as follows: in, , These are the decoding steps. , No. The path metric and normalized log-likelihood ratio for each path; For normalized decoding progress.

8. The end-to-end self-decoder optimization method based on reinforcement learning according to claim 1, characterized in that, In step 2, the design of the action space is as follows: Agent selects action Control path selection behavior, when When the original path metric derived from the SCL decoding process is directly used, it corresponds to a conservative decision; when At the same time, the application of perturbation metrics encourages the exploration of path selection, as follows: in, For small random perturbations; For decoding steps , No. The corrected path metric for the path; The path metric weight coefficient; For decoding steps , No. The original path metric for each path.

9. The end-to-end self-decoder optimization method based on reinforcement learning according to claim 1, characterized in that, In step 2, the reward function is designed as follows: The agent evaluates its decisions based on reward signals that reflect the reliability and confidence of the decoding. The instantaneous reward for each step is defined as follows: in, For decoding steps Instant rewards; For decoding steps , No. The normalized log-likelihood ratio of the paths; For decoding steps , No. The original path metric for each path.

10. A reinforcement learning-based end-to-end self-decoder optimization system utilizing the reinforcement learning-based end-to-end self-decoder optimization method according to any one of claims 1-9, characterized in that, include: IM / DD Fiber Channel Model Training Module: This module is used to train an IM / DD fiber channel model based on a bidirectional long short-term memory network by repeatedly transmitting polar code-encoded signals modulated by PAM4 and receiving PAM4 digital time-domain sampled signals from the experimental platform of the IM / DD optical interconnect system. The received signal is used as the input and the transmitted signal is used as the output. End-to-end optimized self-decoder training module: used to calculate the channel log-likelihood ratio of the model output signal, input to the reinforcement learning-assisted self-decoder for end-to-end training, and finally form an end-to-end optimized self-decoder; wherein, the self-decoder embeds a reinforcement learning agent as an auxiliary decision-making mechanism, learns the mapping from the state space to the action space through the Q network, dynamically adjusts the scoring weights and pruning thresholds, and comprehensively considers decoding correctness and computational complexity through the reward function during the training process.