FSO adaptive modulation method for improving spectrum utilization rate

By using channel state information prediction model and reinforcement learning algorithm for adaptive modulation in the FSO system, the problem of low spectrum utilization and transmission rates in the prior art is solved, and more efficient spectrum utilization and transmission performance is achieved.

CN120150878APending Publication Date: 2025-06-13GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510292041.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing FSO adaptive modulation technology fails to effectively utilize channel state information, resulting in a lower spectrum utilization and transmission rate in complex atmospheric environments, and the calculation complexity and optimization difficulty are high.

Method used

By creating and training a channel state information prediction model, predicting channel state information, and combining different atmospheric channel models, adaptive modulation is converted into a Markov decision-making process, and using reinforcement learning algorithms to jointly schedule the encoding rate, modulation order and transmission power to optimize spectrum utilization and transmission rate.

Benefits of technology

It improves the spectrum utilization and transmission rate of the FSO system in complex atmospheric environments, reduces the computational complexity and optimization difficulty, and enhances the system's adaptability and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120150878A_ABST
    Figure CN120150878A_ABST
Patent Text Reader

Abstract

The invention discloses an FSO adaptive modulation method for improving the spectrum utilization rate, and the method comprises the steps: S1, creating and training a channel state information prediction model, and predicting the channel state information through the trained model; s2, creating different atmospheric channel models according to the channel state information; s3, converting adaptive modulation into a Markov decision process by combining channel state information and different atmospheric channel models and aiming at selecting an optimal action and maximizing a system spectrum utilization rate and a maximum transmission rate, establishing a reinforcement learning algorithm, and performing joint scheduling on parameters including a coding rate, a modulation order and transmission power to obtain an adaptive modulation result; obtaining an optimal reinforcement learning model under the channel state information; and S4, performing adaptive modulation through the reinforcement learning model. According to the invention, adaptive modulation is realized by adopting a deep strong chemical algorithm, the calculation complexity and optimization difficulty can be reduced, and the processing efficiency and performance of the algorithm in coping with a large-scale decision space are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of free space optical networks, and particularly to an FSO adaptive modulation method for improving spectrum utilization rate. Background Art

[0002] With the rapid development of the Internet, the data traffic carried by wireless networks is increasing at an unprecedented speed. People are increasingly pursuing higher bandwidth and higher rates, and traditional radio frequency wireless communication can no longer meet people's needs. Compared with radio frequency technology, free space optical communication technology (FSO) has attracted much attention due to its advantages such as high bandwidth, license-free spectrum, and high rate. However, FSO is vulnerable to the influence of atmospheric turbulence during the transmission process, resulting in an increase in the system bit error rate, a decrease in the frequency band utilization rate, and a reduction in the performance of the entire communication system. Adaptive modulation technology can effectively resist the attenuation caused by atmospheric turbulence. Adaptive modulation technology is based on channel estimation and adaptively adjusts the transmission code mode (coding category, code length, code rate, etc.) according to the channel conditions fed back by the receiving end to match the dynamic changes of the channel, ultimately achieving efficient and reliable signal transmission. In order for the system to adjust the transmission mode in a timely manner, accurate channel estimation and real-time feedback are very important. If the adaptive modulation technology uses outdated channel state information to adjust the transmitter parameters, then the adaptive modulation technology loses its due function.

[0003] In recent years, machine learning techniques have been widely applied in the FSO field. Reference [1] Lionis A, Peppas K, Nistazakis HE, et al. Using machine learning algorithms for accurate received optical power prediction of an FSO link over a maritime environment [C] / / Photonics. MDPI, 2021, 8(6): 212. By modeling the received signal strength index, the accuracy of various machine learning algorithms in predicting the performance of FSO links in a maritime environment is demonstrated. Reference [2] Hameed S M, Mohammed J K, Abdulsatar S M. Performance of adaptive M-PAM modulation for FSO systems based on end-to-end learning [J]. Journal of Optics, 2024: 1-10. Utilizes the performance of adaptive modulation of M-PAM by artificial neural networks in FSO systems. Reference [3] Jaiswal A, Jain V K, Kar S. Adaptive coding and modulation (ACM) technique for performance enhancement of FSO Link [C] / / 2016 IEEE First International Conference on Control, Measurement and Instrumentation (CMI). IEEE, 2016: 53-57. Is evaluated under different turbulence conditions in the scenario of free space laser links. On this basis, the performance of the OOK modulation scheme in FSO links at different data rates and coding rates is studied.

[0004] However, the existing research schemes for FSO adaptive modulation have the following problems:

[0005] 1) At present, FSO adaptive modulation technology is based on information losses such as signal-to-noise ratio, bit error rate, or packet loss rate, and these methods are not linked to actual meteorological parameters. At the same time, the existing technologies mainly focus on channel models with a single turbulence intensity and do not achieve a comprehensive and in-depth study of weak, medium, and strong turbulence intensities.

[0006] 2) The key to adaptive modulation lies in obtaining channel state information in real time, so as to dynamically adjust the transmission parameters at the transmitter to optimize communication performance. However, existing adaptive modulation technologies ignore that the acquisition, processing, and transmission of channel state information all require processing time, and may face the problem of low information timeliness in actual use.

[0007] 3) In FSO adaptive modulation technology, the multi-dimensional channel state information and various modulation and coding strategies increase the decision space, increasing the computational complexity and optimization difficulty of traditional machine learning algorithms. Summary of the Invention

[0008] The object of the present invention is to overcome the deficiencies of the prior art and provide an FSO adaptive modulation method for improving spectrum utilization.

[0009] To achieve the above object, the technical solution provided by the present invention is as follows:

[0010] An FSO adaptive modulation method for improving spectrum utilization, comprising:

[0011] S1. Create and train a channel state information prediction model, and use the trained channel state information prediction model to predict the channel state information;

[0012] S2. Create different atmospheric channel models according to the predicted channel state information;

[0013] S3. Combine the channel state information and different atmospheric channel models, and aim to select the optimal action and maximize the system spectrum utilization and maximum transmission rate. Convert the adaptive modulation into a Markov decision process, build a reinforcement learning algorithm, jointly schedule parameters including coding rate, modulation order, and transmission power, and obtain the optimal reinforcement learning model under the channel state information;

[0014] S4. Perform adaptive modulation through the optimal reinforcement learning model under the channel state information.

[0015] Further, step S1 includes:

[0016] S1-1. Preprocess the data set;

[0017] S1-2. Use the preprocessed data set to train different base learning models;

[0018] S1-3. Select the prediction results of the top N base learning models with the best prediction effect in step S1-2 as feature inputs, and train the channel state information prediction model;

[0019] S1-4. Predict the channel state information through the trained channel state information prediction model.

[0020] Furthermore, preprocess the dataset, including:

[0021] 1) Clean the dataset, find and remove outliers, missing values, and duplicate values in the dataset;

[0022] 2) Standardize the cleaned dataset.

[0023] Furthermore, create different atmospheric channel models according to the predicted channel state information, including:

[0024] h n is the channel state information. When h n <10 -16 , use the lognormal distribution model as the atmospheric channel model, and its probability density function is:

[0025]

[0026] where I a is the unknown, σ s 2 is the log fluctuation variance, and its expression is:

[0027]

[0028] In the formula, k = 2π / λ is the optical wave number, L is the transmission distance of the communication system, D is the diameter of the receiver aperture, λ is the optical wave wavelength, and the parameter σ 0 2 is the Rytov variance, σ 0 2 = 1.23C n 2 k 7 / 6 L 11 / 6

[0029] Step B: When h n > 10 -14 , use the Gamma-Gamma model as the atmospheric channel model, and its probability density function is:

[0030]

[0031] where Γ(·) is the Gamma-Gamma function, K n (·) is the modified Bessel function of the second kind of order n, α is the large-scale turbulence factor, β is the small-scale turbulence factor, and the formulas are as follows:

[0032]

[0033] Further, in step S3, a reinforcement learning algorithm is built to jointly schedule parameters including coding rate, modulation order, and transmission power. During the entire transmission process, it is necessary to ensure that under the constraint that the transmission rate and performance meet the communication requirements, the optimal coding and modulation combination is selected to maximize the spectral efficiency and ensure that the bit error rate is less than the set threshold, specifically including:

[0034] The bit error rate of the nth transmission is e n = f(q n ,P n ,h n ), where q n is the different modulation orders, P n is the different optical powers, h n is the channel state information. Let δ be the maximum bit error rate that the system can accept. If e n is less than δ, then the communication data rate c n is obtained from the coding rate r n and the modulation order M n , otherwise the transmission fails and adaptive modulation is no longer performed;

[0035] And set the minimum rate c * for each transmission. After multiple transmissions, the average transmission rate should not be less than the minimum rate c * .

[0036] Further, the calculation formula of the communication data rate c n is as follows:

[0037]

[0038] where B is the bandwidth, r n is the coding rate, and M n is the modulation order.

[0039] Further, in step S3, a reinforcement learning algorithm is built to jointly schedule parameters including coding rate, modulation order, and transmission power to obtain the optimal reinforcement learning model under the channel state information. The process includes:

[0040] Input the modulation method and the number of transmission rounds;

[0041] A3-1. Initialize the parameters of the Q-value main network and the Q-value delay network. Among them, the policy network parameters and value network parameters of the Q-value main network are θ and ω respectively, and the policy network parameters and value network parameters of the Q-value delay network are θ' and ω' respectively. And initialize the learning rates α, β, exploration rate ε, and discount factor λ;

[0042] A3-2. In each transmission, the state s n = [en-1 , n] In the policy network of the main network that inputs the Q value, the policy network uses the ε-greedy algorithm to select an action a from the action space A = [r, M, P]. n , a n It is expressed as:

[0043]

[0044] Among them, Q(s n , a; θ) is the value expectation of each action in state s n ; r, M, and P respectively represent the coding rate, modulation order, and transmission power;

[0045] A3-4. Randomly draw N s samples (s t , a t , R t , s t+1 ) from the experience pool D as training data. Input s t , a t into the value network of the main network of the Q value to obtain the value function Q(s t of action a in state s t ; ω) and reward R t , a t ; ω) and reward R t . Input s t+1 into the policy network of the Q value delay network to obtain the action set Q t+1 of the Q value delay network; finally, input s t+1 and Q t+1 into the value network of the Q value delay network to obtain the action value function Q(s t+1 , a t+1 ; ω′) and calculate the Target value;

[0046] A3-5. Use (y t , s t , a t ) to calculate the gradient and and update the weights θ, θ′, ω, ω′. The weight update is:

[0047]

[0048] Among them, α 1 , β 1 is the learning rate, and L x (θ) and L Q (ω) are loss functions, defined as:

[0049]

[0050] Among them, Ns The quantity size delay network parameters taken from the experience pool D are updated to:

[0051] ω′ = ρ 1 ω + (1 - ρ 1 )ω′

[0052] θ′ = ρ 2 θ + (1 - ρ 2 )θ′

[0053] where ρ 1 , ρ 2 are update coefficients;

[0054] A3 - 6. Repeat steps A3 - 4 to A3 - 5 until the N s extracted samples are trained to obtain a trained model, that is, the optimal reinforcement learning model under channel state information.

[0055] Furthermore, the reward function is used to select an appropriate modulation method to maximize the average transmission rate after n transmissions and ensure that the average bit error rate is less than a threshold, and the reward R n is calculated by the following formula:

[0056]

[0057] where, is the average rate after n transmissions, the discount factor is a parameter used to ensure that the average bit error rate of the system is less than the threshold, B * = TNc * is the minimum data volume required for transmission, a 1 , a 2 are positive numbers.

[0058] Furthermore, Target is calculated by the following formula:

[0059]

[0060] where, R t is the reward at time t, and λ is the discount factor.

[0061] Furthermore, adaptive modulation is performed through the optimal reinforcement learning model under channel state information, including:

[0062] A4 - 1. Obtain the current state s n = [e n-1 , n], and initialize the time slot T;

[0063] A4 - 2. Substitute s nInput the optimal reinforcement learning model under the channel state information, jointly schedule the parameters including coding rate, modulation order, and transmission power, and calculate the reward R n 、Obtain the state s at the next moment n+1 If T is less than the final time slot, then T = T + 1;

[0064] A4-3. Repeat A4-1 and A4-2 until T reaches the final time slot.

[0065] Compared with the prior art, the principle and advantages of the present technical solution are as follows:

[0066] 1. Using the channel state information prediction model to predict the channel state information can ensure the real-time performance and accuracy of the channel state information, and avoid affecting the communication performance due to the time delay of the channel state information.

[0067] 2. Create different atmospheric channel models according to the predicted channel state information, conduct comprehensive research on various turbulence intensities, greatly improve the adaptability of the FSO system to complex atmospheric environments, and enhance the system performance and reliability.

[0068] 3. Adopt the deep reinforcement learning algorithm to achieve adaptive modulation, which can reduce the computational complexity and optimization difficulty, and improve the processing efficiency and performance of the algorithm when dealing with large-scale decision spaces. Description of the Drawings

[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the services required for the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0070] Figure 1 It is the principle flowchart of an FSO adaptive modulation method for improving spectrum utilization according to the present invention;

[0071] Figure 2 It is the flowchart of predicting the channel state information by the channel state information prediction model completed through training in an FSO adaptive modulation method for improving spectrum utilization according to the present invention. Detailed Embodiments

[0072] The present invention will be further described below in conjunction with specific embodiments:

[0073] As Figure 1 shown, an FSO adaptive modulation method for improving spectrum utilization described in this embodiment includes the following steps:

[0074] S1. Create and train a channel state information prediction model, and use the trained channel state information prediction model to predict the channel state information;

[0075] This step specifically includes:

[0076] S1-1. Preprocess the data set:

[0077] 1) Clean the data set, find and remove outliers, missing values, and duplicate values in the data set;

[0078] 2) Standardize the cleaned data set.

[0079] S1-2. Use the preprocessed data set to train different base learning models (base learners);

[0080] S1-3. Select the prediction results of the top N base learning models with the best prediction effect in step S1-2 as feature inputs, and train a channel state information prediction model (meta-learner);

[0081] S1-4. Predict the channel state information through the trained channel state information prediction model.

[0082] S2. Create different atmospheric channel models according to the predicted channel state information;

[0083] This step specifically includes:

[0084] h n is the channel state information. When h n <10 -16 , use the log-normal distribution model as the atmospheric channel model, and its probability density function is:

[0085]

[0086] where, I a is the unknown, σ s 2 is the log fluctuation variance, and its expression is:

[0087]

[0088] In the formula, k = 2π / λ is the optical wave number, L is the transmission distance of the communication system, D is the diameter of the receiver aperture, λ is the optical wave wavelength, and the parameter σ 0 2 is the Rytov variance, σ 0 2 = 1.23C n 2 k 7 / 6 L 11 / 6

[0089] Step B: When h n > 10 -14 , the Gamma-Gamma model is used as the atmospheric channel model, and its probability density function is as follows:

[0090]

[0091] where Γ(·) is the Gamma-Gamma function, K n (·) is the modified Bessel function of the second kind of order n, α is the large-scale turbulence factor, and β is the small-scale turbulence factor. The formulas are as follows:

[0092]

[0093]

[0094] S3. Combining the channel state information and different atmospheric channel models, aiming at selecting the optimal action and maximizing the system spectral efficiency and maximum transmission rate, the adaptive modulation is converted into a Markov decision process, a reinforcement learning algorithm is built, the parameters including the coding rate, modulation order, and transmission power are jointly scheduled, and the optimal modulation method under the channel state information is selected.

[0095] In this step, a reinforcement learning algorithm is built to train the model to jointly schedule the parameters including the coding rate, modulation order, and transmission power. During the entire transmission process, it is necessary to ensure that under the constraints that the transmission rate and performance meet the communication requirements, the optimal coding and modulation combination is selected to maximize the spectral efficiency and ensure that the bit error rate is less than the set threshold. Specifically, it includes:

[0096] The bit error rate of the nth transmission is e n = f(M n , P n , h n ), where M n is the different modulation orders, P n is the different optical powers, h n is the channel state information. Let δ be the maximum bit error rate that the system can accept. If e n is less than δ, then the communication data rate c n is obtained from the coding rate r n and the modulation order M n , otherwise the transmission fails and the adaptive modulation is no longer performed;

[0097] And set the minimum rate c * for each transmission. After multiple transmissions, the average transmission rate should not be less than the minimum rate c * for each transmission.

[0098] Among the above, the communication data rate c n has the following calculation formula:

[0099]

[0100] where B is the bandwidth, r n is the coding rate, and M n is the modulation order.

[0101] Next, build a reinforcement learning algorithm to jointly schedule parameters including coding rate, modulation order, and transmission power, including:

[0102] Input the modulation method and the number of transmission rounds;

[0103] A3-1. Initialize the parameters of the Q-value main network and the Q-value delay network. Among them, the policy network parameters and value network parameters of the Q-value main network are θ and ω respectively, and the policy network parameters and value network parameters of the Q-value delay network are θ' and ω', and initialize the learning rate α 1 , β 1 , exploration rate ε, discount factor λ;

[0104] A3-2. In each transmission, input the state s n = [e n-1 , n] into the policy network of the Q-value main network. The policy network uses the ε-greedy algorithm to select an action a n from the action space A×[r, M, P], and a n is expressed as:

[0105]

[0106] where Q(s n , a; θ) is the value expectation of each action in state s n ; r, M, and P represent the coding rate, modulation order, and transmission power respectively;

[0107] A3-3. Execute action a n , obtain the reward R n and the next state s n+1 , and store (s n , a n , R n , s n+1 ) in the experience pool D;

[0108] The reward function is used to select an appropriate modulation method to maximize the average transmission rate after n transmissions and ensure that the average bit error rate is less than the threshold, and calculate the reward R n with the following formula:

[0109]

[0110] Among them, is the average rate after n transmissions. The parameter B for ensuring that the system average bit error rate is less than the threshold * = TNc * is the minimum amount of data required for transmission, a 1 , a 2 is a positive number.

[0111] A3-4. Randomly select N s samples (s t , a t , R t , s t+1 ) from the experience pool D as training data. The value function Q(s t , a t is obtained by inputting s t into the value network of the Q-value main network for the action a t in the state s t , a t ; ω) and the reward R t . Input s t+1 into the policy network of the Q-value delay network to obtain the action set Q t+1 of the Q-value delay network; finally, input s t+1 and Q t+1 into the value network of the Q-value delay network to obtain the action value function Q(s t+1 , a t+1 ; ω′) and calculate the Target value;

[0112] The Target calculation is obtained by the following formula:

[0113]

[0114] Among them, R t is the reward at time t, and λ is the discount factor.

[0115] A3-5. Use (y t , s t , a t ) to calculate the gradients and and update the weights ′, θ, ω, ω′. The weight update is:

[0116]

[0117] Among them, α 1 , β 1 is the learning rate, L x (θ) and L QLet \(L(\omega)\) be the loss function, which is defined as:

[0118]

[0119] where \(N\) s is the number taken from the experience pool \(D\), and the delayed network parameter update is:

[0120] \(\omega'=\rho\) 1 \(\omega+(1 - \rho\) 1 )\(\omega'\)

[0121] \(\theta'=\rho\) 2 \(\theta+(1 - \rho\) 2 )\(\theta'\)

[0122] where \(\rho\) 1 , \(\rho\) 2 are the update coefficients;

[0123] A3 - 6. Repeat steps A3 - 4 to A3 - 5 until the \(N\) s extracted samples are trained to obtain a trained model (i.e., the optimal reinforcement learning model under channel state information);

[0124] S4. Perform adaptive modulation through the optimal reinforcement learning model under channel state information. The specific steps are as follows:

[0125] A4 - 1. Obtain the current state \(s\) n =[e n-1 , n], and initialize the time slot \(T\)

[0126] A4 - 2. Input \(s\) n into the model trained in S3, jointly schedule the parameters including coding rate, modulation order, and transmission power, and calculate the reward \(R\) n , obtain the state \(s'\) n+1 at the next moment. If \(T\) is less than the final time slot, then \(T = T + 1\)

[0127] A4 - 3. Repeat A4 - 1 and A4 - 2 until \(T\) reaches the final time slot.

[0128] The above - mentioned embodiments are only the preferred embodiments of the present invention, and do not limit the scope of implementation of the present invention. Therefore, all changes made according to the shape and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A FSO adaptive modulation method for improving spectrum utilization, characterized in that: include: S1. Creating and training a channel state information prediction model, and using the trained channel state information prediction model to predict channel state information; S2, creating different atmospheric channel models according to the predicted channel state information; S3. Combining channel state information and different atmospheric channel models, with the goal of selecting the optimal action and maximizing system spectrum utilization and maximum transmission rate, the adaptive modulation is converted into a Markov decision process, and a reinforcement learning algorithm is built to jointly schedule parameters including coding rate, modulation order and transmission power to obtain the optimal reinforcement learning model under channel state information. S4. Adaptive modulation is performed through the optimal reinforcement learning model under channel state information.

2. The FSO adaptive modulation method for improving spectrum utilization according to claim 1, characterized in that: Step S1 includes: S1-1, preprocess the data set; S1-2, use the preprocessed data set to train different base learning models; S1-3, selecting the prediction results of the top N base learning models with the best prediction effect in step S1-2 as feature input, and training the channel state information prediction model; S1-4. Predict channel state information using the trained channel state information prediction model.

3. The FSO adaptive modulation method for improving spectrum utilization according to claim 2, characterized in that: Preprocess the dataset, including: 1) Clean the data set to find and remove outliers, missing values, and duplicate values ​​in the data set; 2) Standardize the cleaned dataset.

4. The FSO adaptive modulation method for improving spectrum utilization according to claim 1, characterized in that: Different atmospheric channel models are created based on the predicted channel state information, including: h n is the channel state information, when h n <10 -16 When , the log-normal distribution model is used as the atmospheric channel model, and its probability density function is: Among them, I a is an unknown number, σ s 2 is the logarithmic fluctuation variance, and its expression is expressed as: In the formula, k = 2π / λ is the optical wave number, L is the transmission distance of the communication system, D is the diameter of the receiver aperture, λ is the wavelength of the optical wave, and the parameter σ0 2 is the Rytov variance, σ0 2 =1.23C n 2 k 7 / 6 L 11 / 6 Step B: When h n >10 -14 When , the Gamma-Gamma model is used as the atmospheric channel model, and its probability density function is: Among them, Γ(·) is the Gamma-Gamma function, K n (·) is the nth-order second-kind modified Bessel function, α is the large-scale turbulence factor, β is the small-scale turbulence factor, and the formulas are as follows:

5. The FSO adaptive modulation method for improving spectrum utilization according to claim 1, characterized in that: In step S3, a reinforcement learning algorithm is built to jointly schedule parameters including coding rate, modulation order and transmission power. During the entire transmission process, it is necessary to ensure that the transmission rate and performance meet the communication requirements, select the optimal coding and modulation combination to maximize the spectrum efficiency, and ensure that the bit error rate is less than the set threshold. Specifically, it includes: The bit error rate of the nth transmission is e n =f(q n ,P n ,h n ), where q n For different modulation orders, P n For different optical powers, h n is the channel state information, let δ be the maximum bit error rate that the system can accept, if e n is less than δ, then the communication data rate c n The coding rate r n and modulation order M n It is concluded that otherwise the transmission fails and adaptive modulation is no longer performed; And set the minimum rate c for each transmission * After multiple transmissions, the average transmission rate should not be less than the minimum transmission rate c of each transmission. * .

6. The FSO adaptive modulation method for improving spectrum utilization according to claim 5, characterized in that: Communication data rate c n The calculation formula is as follows: Where B is the bandwidth, r n is the coding rate, M n is the modulation order.

7. The FSO adaptive modulation method for improving spectrum utilization according to claim 5 or 6, characterized in that: In step S3, a reinforcement learning algorithm is constructed to jointly schedule parameters including coding rate, modulation order and transmission power to obtain the optimal reinforcement learning model under channel state information. The process includes: Enter the modulation method and the number of transmission rounds; A3-1. Initialize the parameters of the Q-value main network and the parameters of the Q-value delay network, where the policy network parameters and value network parameters of the Q-value main network are θ and ω respectively, and the policy network parameters and value network parameters of the Q-value delay network are θ′ and ω′ respectively, and initialize the learning rate α, β, exploration rate ε, and discount factor λ; A3-2. In each transmission, the state s n =[e n-1 ,n] input Q value into the policy network of the main network, the policy network uses the ε greedy algorithm to select action a from the action space A = [r,M,P] n , a n It is expressed as: Among them, Q(s n ,a;θ) is the state s n The expected value of each action is: r, M, P represent the coding rate, modulation order and transmission power respectively; A3-3. Execute action a n , get reward R n and the next state s n+1 , and (s n ,a n ,R n ,s n+1 ) is stored in the experience pool D; A3-4. Randomly draw N from the experience pool D s Samples (s t ,a t ,R t ,s t+1 ) as training data, s t ,a t Input Q value to the main network and get the value of the network in state s t The next action a t The value function Q(s t ,a t ;ω) and reward R t , will s t+1 Input into the policy network of the Q-value delay network to obtain the action set Q of the Q-value delay network t+1 ; Finally, s t+1 and Q t+1 Input into the value network of the Q value delay network to obtain the action value function Q(s t+1 ,a t+1 ;ω′) and calculate the Target value; A3-5、Use (y t ,s t ,a t ) Calculate the gradient and And update the weights θ, θ′, ω, ω′, and the weights are updated as follows: Among them, α1, β1 are learning rates, L x (θ) and L Q (ω) is the loss function, defined as: Where N s The amount of delay network parameters taken from the experience pool D is updated as: ω′=ρ1ω+(1-ρ1)ω′ θ′=ρ2θ+(1-ρ2)θ′ Among them, ρ1 and ρ2 are update coefficients; A3-6. Repeat steps A3-4 to A3-5 until the extracted N s samples, and obtain the trained model, that is, the optimal reinforcement learning model under channel state information.

8. The FSO adaptive modulation method for improving spectrum utilization according to claim 7, characterized in that: The reward function is used to select the appropriate modulation method to maximize the average transmission rate after n transmissions and ensure that the average bit error rate is less than the threshold. The reward R is calculated. n The formula is as follows: in, is the average rate after n transmissions, and the discount factor Parameter used to ensure that the system average bit error rate is less than the threshold, B * =TNc * is the minimum amount of data required for transmission, and a1 and a2 are positive numbers.

9. The FSO adaptive modulation method for improving spectrum utilization according to claim 7, characterized in that: Target calculation is obtained by the following formula: Among them, R t is the reward at time t, and λ is the discount factor.

10. The FSO adaptive modulation method for improving spectrum utilization according to claim 1, characterized in that: Adaptive modulation is performed through the optimal reinforcement learning model under channel state information, including: A4-1. Get the current status s n =[e n-1 ,n], initialize time slot T; A4-2, s n Input the optimal reinforcement learning model under the channel state information, jointly schedule the parameters including coding rate, modulation order and transmission power, and calculate the reward R n , get the next moment state s n+1 , if T is less than the final time slot, then T = T + 1; A4-3. Repeat A4-1 and A4-2 until T reaches the final time slot.