A coherent detection spatial light cooperation transmission method based on local oscillator elastic light splitting
By employing a coherent detection spatial optical cooperative transmission method based on local oscillator elastic beam splitting and utilizing the DDPG algorithm to adaptively allocate local oscillator optical power, the problem of unreasonable optical power resources in traditional allocation schemes is solved, achieving more efficient communication quality and stability.
Patent Information
- Application Number
- CN202510008669.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-01-03
AI Technical Summary
Traditional local oscillator power proportional allocation schemes lead to unreasonable allocation of optical power resources, making it difficult to maximize the mixing gain of the local oscillator and effectively cope with the effects of atmospheric turbulence and path loss in FSO communication.
A coherent detection spatial optical cooperative transmission method based on local oscillator elastic beam splitting is adopted. The DDPG algorithm is used to adaptively allocate local oscillator optical power to each link. The optical power allocation scheme is dynamically adjusted by combining an elastic beam splitter and a voltage controller with a deep deterministic strategy gradient algorithm.
It improves the system's flexibility and transmission efficiency, effectively addresses stability and efficiency under dynamic channel conditions, and enhances the system's communication quality under atmospheric turbulence and path loss.
Smart Images

Figure CN119853815B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of space optical communication counting and relates to a coherent detection space optical cooperative transmission method based on local oscillator elastic beam splitting. Background Technology
[0002] Free-space optical (FSO) communication has become an effective supplementary solution for high-capacity wireless communication due to its advantages such as no spectrum licensing required, large channel capacity, and flexible deployment. Currently, FSO communication technology is mainly divided into two communication systems: intensity modulation / direct detection (IM / DD) and multi-order modulation / coherent detection. Among them, the multi-order modulation / coherent detection system is widely used in existing optical communication terminals due to its higher detection sensitivity, more flexible modulation methods, and better wavelength selectivity, and it is also an important research direction for future space optical communication technology. Despite the many advantages of FSO systems, factors such as atmospheric turbulence and path loss significantly limit the communication quality of FSO links.
[0003] Cooperative communication, as an effective technique for FSO communication to combat the effects of atmospheric turbulence and path loss, has been extensively studied. However, since FSOs are directly exposed to the atmospheric environment, this may lead to differences in the received optical power of different transmission paths in cooperative transmission systems. Traditional proportional allocation schemes for local oscillator optical power result in an unreasonable allocation of optical power resources, making it difficult to maximize the mixing gain of the local oscillator. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a coherent detection space optical cooperative transmission method based on local oscillator elastic beam splitting. In view of the differences in received optical power under different transmission paths caused by the time-varying nature of atmospheric channel conditions, the method uses an elastic beam splitter that can adaptively allocate local oscillator optical power according to channel conditions, and combines a deep deterministic strategy gradient algorithm to determine the output scheme of the voltage controller, thereby determining the power allocation scheme of the elastic beam splitter, thereby improving the flexibility and transmission efficiency of the system.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A coherent detection spatial optical cooperative transmission method based on local oscillator elastic beam splitting, the method comprising:
[0007] At the optical transmitter, the bit sequence is modulated into an optical signal using quadrature amplitude modulation, and then split into multiple uncorrelated sub-signals by a fixed optical splitter for broadcast transmission.
[0008] At the optical receiver, the signal transmitted by the even-numbered link is phase-shifted by a 90° phase shifter to achieve phase orthogonality between the signals of the even-numbered and odd-numbered links; at the same time, the local oscillator power is adaptively allocated to the photodetectors corresponding to each link by an elastic beam splitter; the bit sequence is restored after the output signal of the photodetector is demodulated by a coherent receiver.
[0009] The power distribution of the flexible optical splitter is controlled by a voltage controller, wherein the output voltage of the voltage controller is dynamically determined according to the atmospheric channel state using the DDPG algorithm.
[0010] Furthermore, the DDPG algorithm calculates the output voltage of the voltage controller with the goal of minimizing the bit error rate; in this DDPG algorithm, action a L s(t) represents the power allocation ratio of the local oscillator, s(t) represents the channel state of each link, and the reward function is expressed as:
[0011] r L (t)=|log 10 BER(t)|-r th
[0012] In the formula, r th This represents the initial threshold of the system, referring to the system's |log| when the local oscillator optical power is evenly distributed. 10 BER|, where BER(t) represents the bit error rate; if action a L (t) is applied to the environmental reward function r L When (t)>0, it indicates that the local oscillator power ratio improves the system performance under this state.
[0013] The steps of the DDPG algorithm to solve for the output voltage of the voltage controller include:
[0014] 1) Initialize the experience replay pool D, and initialize the Actor network parameters θ. a and Critic network parameters θ c Set the target network parameter θ c′ ←θ c θ a′ ←θ a Set the learning rate and the initial system threshold r. th and discount factor γ;
[0015] 2) At each time step, the Actor network, based on its learned policy... Select a set of local oscillator power distribution ratio actions a L (t), the Critic network uses the value function Q(s(t),a) L (t)|θ c To evaluate the Actor network μ(s) t ,at |θ a The action taken and the reward received. L (t) and the state at the next time step s(t+1);
[0016] 3) The empirical tuple (s(t), a) L (t),r L (t), s(t+1)) are stored in the experience replay pool D; training begins when the agent has accumulated enough experience tuples, and M experience tuples are randomly selected from the experience replay pool D for training; the parameters θ are updated according to the loss function and the policy gradient function respectively. c and θ a Update the objective function parameters based on soft updates;
[0017] 4) After training, the atmospheric channel state is sensed in real time by the channel state sensor, and the output voltage of the voltage controller is calculated by the DDPG algorithm to adjust the power distribution ratio of the elastic beam splitter in real time.
[0018] The loss function is expressed as follows:
[0019]
[0020] R t =Q(s(t),a L (t)|θ c )
[0021] R′ t =r L (t)+γ·Q(s(t+1),a L (t+1)|θ c )
[0022] In the formula, R t R′ represents the long-term reward of the Critic network. t Indicates the estimated long-term reward;
[0023] The policy gradient function is expressed as follows:
[0024] The beneficial effects of this invention are as follows: This invention utilizes the DDPG algorithm to dynamically estimate the local oscillator splitting ratio based on different atmospheric channel conditions, and sends the decision results to a voltage controller in real time. The voltage controller further controls the dynamic allocation of local oscillator optical power. This invention adaptively determines different link optical power allocation schemes using a local oscillator elastic splitting structure, which can flexibly cope with the impact of different atmospheric channel conditions on transmission performance. It can effectively improve the stability and efficiency of the system under dynamic channel conditions, demonstrating a strong ability to resist channel attenuation and environmental changes.
[0025] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0027] Figure 1 Block diagram of a spatial optical cooperative transmission system with local oscillator elastic beam splitting and IQ modulation / coherent detection;
[0028] Figure 2 This is a schematic diagram of the DDPG algorithm.
[0029] Figure 3 For the bit error rate performance of a system based on a fixed optical splitter;
[0030] Figure 4 For the bit error rate performance of the system based on the elastic beam splitter;
[0031] Figure 5 The system performance under different refractive index constants;
[0032] Figure 6 Performance comparison under different channel conditions;
[0033] Figure 7 This is a performance comparison under similar IQ channel conditions. Detailed Implementation
[0034] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0035] To ensure the rationality of local oscillator power allocation and reduce the bit error rate of system transmission, one embodiment of the present invention provides a coherent detection spatial optical cooperative transmission method based on local oscillator elastic beam splitting. Addressing the differences in received optical power under different transmission paths caused by the time-varying nature of atmospheric channel conditions, this method designs an elastic beam splitter that can adaptively allocate local oscillator power according to channel conditions. The output scheme of the voltage controller is determined through the DDPG algorithm, and the voltage controller then determines the power allocation scheme of the elastic beam splitter.
[0036] Figure 1 The diagram shows a local oscillator elastically splitting IQ modulation / coherent detection space optical cooperative transmission system, which mainly consists of three parts: an optical transmitter, an atmospheric channel, and a low-complexity coherent receiver. The transmitter employs Quadrature Amplitude Modulation (QAM), then splits the modulated signal into multiple channel-independent sub-signals for broadcast transmission. At the receiver, the second and fourth (even-numbered) links are phase-shifted by 90° to intersect with the first and third (odd-numbered) links. At this point, the first and third (odd-numbered) links are I-channel signals, and the second and fourth (even-numbered) links are Q-channel signals orthogonal to the I-channel signals. A simplified coherent structure is then used to demodulate the I-channel and Q-channel signals respectively to achieve symbol recovery.
[0037] At the receiving end of the coherent receiver, an elastic beam splitter is used to allocate local oscillator power to the photodetectors of each channel. The allocation ratio of the local oscillator power is determined by the output scheme of the voltage controller, which is determined by the DDPG algorithm. The specific process of determining the voltage controller output scheme using the DDPG algorithm is as follows:
[0038] 1. First, in the DDPG algorithm, action a... L (t) is set as the local oscillator power ratio, and the state is set as the channel state of each link s(t)=[h1(t-1),h2(t-1),…,h n [(t-1)], h n (t-1) represents the channel state of link n. Action a... L (t) When applied to the environment, an immediate reward r can be obtained. L (t) is used to evaluate action a L (t) represents the performance of s(t) in state s(t). To increase the efficiency of training data utilization, an experience replay mechanism is introduced, which saves historical states, actions, and reward sequences and randomly selects a batch of samples from them for training. The reward function is defined as:
[0039] r L (t)=|log 10 BER(t)|-r th
[0040] In the formula, r th The initial threshold of the system refers to the system's |log| when the local oscillator power is evenly distributed. 10 BER|. When action a L (t) after being applied to the environment r L When (t) > 0, it indicates that the local oscillator power ratio improves system performance in this state. Therefore, the goal of the DDPG algorithm is to maximize long-term reward.
[0041] 2. Initialization: Initialize the experience replay pool D, and initialize the Actor network parameters θ. a and Critic network parameters θ c Set the target network parameter θ c′ ←θ c θ a′ ←θ a Set the learning rate and the initial system threshold r. th Parameters such as discount factor γ.
[0042] 3. Interaction Process: At each moment, the Actor network applies its learned strategy... Select a set of local oscillator optical power ratio actions a L (t), the Critic network uses the value function Q(s(t),a) L (t)|θ c To evaluate the Actor neural network μ(s) t ,a t |θ a The actions taken by the Actor neural network optimize it during online training.
[0043] 4. Training begins when the agent has accumulated enough experience tuples. The Critic network replays these tuples to obtain long-term rewards for state s(t):
[0044] R t =Q(s(t),a L (t)|θ c )
[0045] At the same time, an instant reward r will be given L (t) is stored in the corresponding tuple, and the estimated long-term reward is obtained using the Bellman equation:
[0046] R′ t =r L (t)+γ·Q(s(t+1),a L (t+1)|θ c )
[0047] Since the training of the Critic network aims to learn how to more accurately evaluate the actions of the Actor network, its loss function is defined as:
[0048]
[0049] In the formula, M represents the number of samples taken during training. During training, the parameters θ of the Critic network are updated by minimizing the loss function. c .
[0050] Actor network parameters θ a The policy gradient function is used for updating, and its formula is as follows:
[0051]
[0052] 5. Through continuous interaction, the Actor network gradually improves its action selection in the environment, enhancing policy performance under a given state. After a certain number of training steps, the DDPG algorithm converges, reflected in the stability of the actions generated by the Actor network in the state space and a reduction in the prediction error of the Critic network. Based on the DDPG algorithm, it is possible to adjust the output of the voltage controller in real time under dynamically changing channel conditions, thereby adjusting the splitting ratio of the flexible beam splitter in real time.
[0053] The steps of the DDPG algorithm are as follows:
[0054] 1: Initialize the experience replay pool D;
[0055] 2: Initialize the Actor network θ a And the parameters θ of the Critic network c ;
[0056] 3: Set the target network parameter θ c′ ←-θ c θ a′ ←θ a ;
[0057] 4: Initialize the environment, execute the local oscillator light power distribution action to obtain the initial state s0 and initial...
[0058] threshold r th ;
[0059] 5: for episodes=0,1,…,L do
[0060] 6: Initialize a random process for action exploration;
[0061] 7: for t=0,1,…,T do
[0062] 8: Select the local oscillator power ratio a based on the strategy function and exploration noise.L (t)
[0063] 9: Perform action a L (t) Receive reward r L (t) and the state at the next time step s(t+1);
[0064] 10: The empirical tuple (s(t), a) L (t),r L (t),s(t+1)) are stored in the experience replay pool D;
[0065] 11: Randomly select M samples (s(t), a) from D. L (t),r L (t),s(t+1));
[0066] 12: Update θ based on minimizing the loss function and policy gradient function. a and θ c ;
[0067] 13: Update the objective function parameters based on soft updates;
[0068] 14: end for
[0069] 15: end for
[0070] To analyze the relative advantages of the DDPG algorithm in optimizing optical power allocation, this embodiment compares the power allocation results obtained using the Lagrange method with those obtained using the DDPG algorithm. To analyze the performance advantages of the Elastic Splitter (ES) compared to the Fixed Splitter (FS), a comparative analysis is conducted from two aspects: atmospheric attenuation and atmospheric turbulence. Figure 3 and Figure 4 The performance of coherent probe space optical cooperative transmission systems using FS and ES was compared without the addition of atmospheric turbulence.
[0071] from Figure 3 and Figure 4 It can be seen that the system bit error rate increases with link L I,1 / L Q,2 and L I,3 / L Q,4 The received optical power gradually increases as the received optical power decreases. Simultaneously, when the received optical power of both the I / Q component links falls below -34dBm, the system's transmission performance degrades sharply. When L... I,1 / L Q,2 and L I,3 / L Q,4When the received optical power is unequal, the BER of a coherent probe spatial optical cooperative transmission system using ES is lower than that of a coherent probe spatial optical cooperative transmission system using FS. Furthermore, L... I,1 / L Q,2 and L I,3 / L Q,4 The greater the difference in received optical power, the greater the performance improvement of the low-complexity coherent probe FSO cooperative communication system using ES (e.g. Figure 3 and Figure 4 (As indicated by the circle in the middle), it can improve system performance by up to 1 to 2 orders of magnitude. This is because the local oscillator plays a role in optical amplification during coherent detection, and the flexible beam splitter can flexibly adjust the local oscillator power splitting ratio according to the received optical power of each link to achieve the maximum mixing gain.
[0072] To further understand the effectiveness of ES in improving the turbulence resistance of coherent probe space optical cooperative transmission systems, this embodiment compares the effects of ES and FS coherent probe space optical cooperative transmission systems under different turbulence conditions. The atmospheric loss of the I / Q component link is set to L. I,1 / L Q,2 =6dB / km and L I,3 / L Q,4 =10dB / km, refractive index structure constant is 1×10 -18 m -2 / 3 ~1×10 -14 m -2 / 3 The changes between them. For example... Figure 5 As shown, the system's average BER increases with the increase of the link's refractive index structure constant. This is because the refractive index structure constant... This affects the fluctuation of optical signal intensity within a small range; therefore, using the average BER (Refractive Index Requirement) better reflects the impact of the refractive index structure constant on system performance. The average BER is the value of the refractive index structure constant for each element in the link. The results of 10 experiments. Figure 5 Error bars are used to characterize the statistical distribution of BER values from 10 simulations. When the refractive index structure constant... At that time, the coherent probe spatial optical cooperative transmission system of ES no longer has system gain compared to the coherent probe spatial optical cooperative transmission system of FS. Figure 5 It can be seen that the coherent detection space optical cooperative transmission system of ES is at 1×10 -18 m -2 / 3 ~1×10 -15 m -2 / 3 There is a significant gain in the weak to moderate turbulence range, with the BER improving by up to 1 to 1.5 orders of magnitude.
[0073] Figure 6The performance of ES and FS in a coherent probe spatial optical cooperative transmission system was compared under different channel conditions. The channel parameters under different channel conditions are shown in Table 1. Figure 6 It can be seen that the BER of the ES supported by DDPG applied to the coherent probe FSO cooperative system gradually converges with the increase of the number of iterations. The average BER in the figure is the average BER of 10 training iterations per iteration. The algorithm reaches convergence after 44 iterations, and the BER after convergence is close to 10. -5 Meanwhile, the ES supported by the Lagrange optimization algorithm, when applied to coherent FSO cooperative systems, does not improve system performance and may even worsen it. This is because the local oscillator power ratio obtained by the Lagrange optimization method cannot be applied to different channel conditions. According to... Figure 6 It can be observed that the BER of FS applied to coherent probe space optical cooperative transmission systems is higher than 10. -3 DDPG-supported ES, when applied to coherent detection space optical cooperative transmission systems, can improve system performance by more than two orders of magnitude compared to FS.
[0074] Table 1
[0075]
[0076]
[0077] To compare whether the DDPG algorithm can maintain its advantage when the overall channel states of the I and Q components are similar, this embodiment compares the performance of the Lagrange optimization algorithm and the DDPG algorithm. The system channel parameter settings are shown in Table 2. Figure 7 This paper compares the performance of ES supported by DDPG and ES supported by Lagrange multiplication algorithm in a coherent probe spatial optical cooperative transmission system (SOP) under similar overall channel states for IQ components. The average BER is the average BER of 10 training iterations per iteration. ES supported by DDPG begins to converge in the coherent probe SOP at the 47th iteration, with the optimal BER converging at 4.4 × 10⁻⁶. -5 The Lagrange optimization algorithm, which supports optimal ES (Extreme Stability) in coherent detection space optical cooperative transmission systems, only requires calculating the optimal solution based on the expression, without a convergence process, and the BER (Breakpoint) can be optimized to 2.25 × 10⁻⁶. -5 Compared to the DDPG algorithm, the Lagrange algorithm can achieve optimization more simply and efficiently under these channel conditions.
[0078] Table 2
[0079]
[0080] pass Figure 6 and Figure 7A performance comparison of the Lagrange optimization algorithm and the DDPG algorithm under multiple channel conditions in a low-complexity coherent FSO cooperative communication system based on ES reveals that both algorithms have their own advantages and disadvantages. The DDPG algorithm is more advantageous when dealing with more complex channel conditions, while the Lagrange optimization algorithm is more efficient when the differences between various channel conditions are not significant.
[0081] In summary, this invention proposes a coherent detection spatial optical cooperative transmission method based on local oscillator elastic beam splitting. Addressing the differences in received optical power across different transmission paths caused by the time-varying nature of atmospheric channel conditions, the output scheme of the voltage controller is dynamically determined based on Lagrange and DDPG algorithms according to dynamically changing channel conditions. Furthermore, the voltage controller determines the power allocation scheme of the elastic beam splitter, achieving on-demand allocation of local oscillator power. Depending on different channel conditions, the elastic beam splitting structure designed in this invention can flexibly adjust the output local oscillator power, no longer limited to equal distribution of local oscillator power, but flexibly adjusted according to actual conditions. Simulation results demonstrate the invention's ability to effectively improve the stability and efficiency of the system under dynamic channel conditions, as well as its strong ability to combat channel attenuation and environmental changes.
[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A coherent detection spatial optical cooperative transmission method based on local oscillator elastic spectroscopy, characterized in that, The method includes: At the optical transmitter, the bit sequence is modulated into an optical signal using quadrature amplitude modulation, and then split into multiple uncorrelated sub-signals by a fixed optical splitter for broadcast transmission. At the optical receiver, the signal transmitted by the even-numbered link is phase-shifted by a 90° phase shifter to achieve phase orthogonality between the signals of the even-numbered and odd-numbered links; at the same time, the local oscillator power is adaptively allocated to the photodetectors corresponding to each link through an elastic beam splitter; the output signal of the photodetector is demodulated by a coherent receiver to recover the bit sequence. The power distribution of the flexible optical splitter is controlled by a voltage controller, wherein the output voltage of the voltage controller is dynamically determined according to the atmospheric channel state using the DDPG algorithm. The DDPG algorithm calculates the output voltage of the voltage controller with the goal of minimizing the bit error rate; in this DDPG algorithm, the action... The power distribution ratio of the local oscillator light, state Given the channel state of each link, the reward function is expressed as: In the formula, This represents the initial threshold of the system, referring to the system when the local oscillator optical power is evenly distributed. , Indicates the bit error rate; if the action Reward function applied to the environment If the power ratio of the local oscillator is positive, it indicates that the system performance is improved under this condition; The steps for solving the voltage controller output voltage using the DDPG algorithm include: Initialize the experience replay pool D Initialize Actor network parameters and Critic network parameters Set target network parameters , Set the learning rate and initial system threshold. and discount factor ; At each moment, the Actor network applies its learned policy. Select a set of local oscillator power distribution ratios for operation. The Critic network uses value functions. To evaluate the Actor network The actions taken and the rewards received. and the state at the next moment s ( t+ 1); experience tuples Stored in the experience replay pool D Training begins when the agent has accumulated enough experience tuples, starting from the experience replay pool. D Random selection M Train using empirical tuples; update parameters according to the loss function and policy gradient function respectively. and Update the objective function parameters based on soft updates; After training, the atmospheric channel state is sensed in real time by the channel state sensor, and the output voltage of the voltage controller is calculated by the DDPG algorithm to adjust the power distribution ratio of the elastic beam splitter in real time.
2. The method according to claim 1, characterized in that, The loss function is expressed as: In the formula, This represents the long-term reward of the Critic network. Indicates the estimated long-term reward; The policy gradient function is expressed as follows: .
Citation Information
Patent Citations
Fault-tolerant method for improving underwater robot networking robustness
CN118741573A
Internet of vehicles resource optimization method based on composite priority experience playback sampling
CN118890658A