Dynamic symbol interval underwater acoustic communication time delay-error code joint optimization method based on reinforcement learning
By applying a dynamic symbol interval adjustment method based on reinforcement learning in the water acoustic communication system, the symbol interval is optimized in real time, and the problem of difficulty in optimizing delay and bit error rate at the same time in traditional technologies is solved, and the water acoustic communication performance with low delay and low bit error rate is achieved.
Patent Information
- Application Number
- CN202510273651.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-10
AI Technical Summary
When traditional acoustic communication technology faces complex channel characteristics, it is difficult to optimize the delay and bit error rate simultaneously, and the existing methods lack real-time dynamic adjustment mechanism.
Using a dynamic symbol interval adjustment method based on reinforcement learning, through the DDPG algorithm and Actor-Critic network architecture, the channel state and performance indicators are sensed in real time, and the symbol interval is dynamically optimized to achieve joint optimization of delay-bit error rate.
It realizes the low latency and low bit error rate of the water acoustic communication system, and can adapt to complex channel environments to improve system performance and response speed.
Smart Images

Figure CN120128277A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of underwater acoustic communication technology, and is particularly applicable to a real-time communication system in a time-delay-Doppler double-diffusion channel. Specifically, it is a method for dynamically adjusting symbol intervals based on reinforcement learning, which is used to jointly optimize the time-delay and bit-error rate performance. Background Art
[0002] Nowadays, people are increasingly interested in real-time monitoring of health conditions during underwater activities, and fields such as remote ocean monitors, deep-sea robots, and unmanned submarines are developing rapidly. It is the underwater acoustic communication technology that plays a key supporting role. However, the underwater acoustic channel has complex characteristics such as multipath effects, Doppler frequency shift, and low signal-to-noise ratio. These characteristics will cause an increase in signal transmission delay and an increase in bit-error rate, seriously affecting the performance of the underwater acoustic communication system. Traditional symbol interval design methods usually adopt fixed values and cannot be dynamically adjusted according to the real-time channel state, making it difficult to ensure low delay and bit-error rate simultaneously under different channel conditions. Therefore, a method that can adaptively adjust the symbol interval is needed to improve the performance of underwater acoustic communication. The existing underwater acoustic communication technology methods have the following disadvantages:
[0003] Traditional methods adopt fixed symbol intervals and cannot adapt to the time-varying characteristics of the channel, resulting in redundant guard intervals or inter-symbol interference;
[0004] There is an inherent conflict between time-delay and bit-error rate. Existing joint optimization methods rely on offline modeling and lack a real-time dynamic trade-off mechanism. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for jointly optimizing time-delay and bit-error of underwater acoustic communication with dynamically adjusted symbol intervals based on reinforcement learning, so as to achieve low delay, low bit-error rate, and strong environmental adaptability of the underwater acoustic communication system.
[0006] To achieve the above purpose, the technical solution of the present invention is as follows:
[0007] A method for jointly optimizing time-delay and bit-error of underwater acoustic communication with dynamically adjusted symbol intervals based on reinforcement learning includes the following steps:
[0008] Step 1, based on the maximum channel time-delay spread T_min and the maximum allowable time-delay T_max of the system, set the symbol interval adjustment range as [1.2T_min, 0.8T_max];
[0009] Step 2, initialize the parameters of the Actor network and the Critic network of the DDPG algorithm. The Actor network adopts a fully connected layer structure, and the dimension of its input layer matches the state vector to receive and process state information; the activation function of the output layer is Tanh, which is used to map the output value to the interval [-1, 1], meeting the requirements of action output; the input of the Critic network is the state vector and the action value, and the output is the Q-value estimation, which is used to evaluate the value of the action. At the same time, construct an experience replay buffer to store state transition samples and provide data support for subsequent learning and training;
[0010] Step 3, perceive the channel state information in real time. The channel state information includes multipath delay distribution, Doppler frequency shift amount, and signal-to-noise ratio; calculate the system performance metrics. The performance metrics include the current bit error rate, average transmission delay, and symbol interval historical sequence;
[0011] Step 4, input the channel state information and the performance metrics into the Actor-Critic network architecture, and output the symbol interval adjustment amount through the network architecture to dynamically optimize the symbol interval;
[0012] Step 5, calculate the bit error rate incentive term, delay penalty term, and action smoothing term according to the current bit error rate and delay situation, define a dynamic weighted reward function, and store the transfer samples in the experience replay buffer;
[0013] Step 6, sample samples from the experience replay buffer, calculate the temporal difference error, update the parameters of the Actor and Critic networks, perform soft update of the target network, and the transmitter sends data according to the adjusted symbol interval;
[0014] Step 7, feedback the state information of the receiver for time alignment, set a sliding window to construct a temporal correlation state vector through the LSTM network, adopt an ε-greedy strategy for action exploration, and set an emergency fallback mechanism to achieve low delay and low bit error rate.
[0015] The method for dynamically adjusting the symbol interval in step 4 specifically includes the following steps:
[0016] Step 4.1, normalize the multipath delay distribution τ(t), Doppler frequency shift amount Δf(t), signal-to-noise ratio SNR(t), current bit error rate BER(t), average transmission delay D(t), and symbol interval historical sequence {T(t - 1), T(t - 2),..., T(t - N)} (N≥3) in step 3 to construct a state vector s(t);
[0017] Step 4.2, input s(t) into the Actor network. The activation function of the output layer of the network is Tanh, and the output action a(t) ∈ [-1, 1];
[0018] Step 4.3: Linearly transform a(t) to the symbol interval adjustment ΔT(t) = 0.1·T(t - 1)·a(t), where ΔT(t) ∈ [-0.1T(t - 1), +0.1T(t - 1)];
[0019] Step 4.4: Dynamically optimize the symbol interval. The symbol interval formula is: T(t) = T(t - 1) + ΔT(t), where T(t) ∈ [T_min, T_max];
[0020] Step 4.5: Input a(t) and s(t) into the Critic network, which directly outputs a scalar Q value for evaluating the return of the action.
[0021] Among them, the specific setting and parameter adjustment of the weighted reward function in Step 5 include the following steps:
[0022] Step 5.1: Set the bit error rate incentive term as: Enhance the sensitivity in the low bit error rate region through the logarithmic function to encourage the algorithm to reduce the bit error rate;
[0023] Step 5.2: Set the delay penalty term as: When the delay approaches the upper limit tolerated by the system, the penalty increases sharply to ensure low delay;
[0024] Step 5.3: Set the action smoothing term as: R(t) = ψ·|ΔT(t)| to suppress frequent adjustment of the symbol interval and avoid system oscillation;
[0025] Step 5.4: Set the weight dynamic selection strategy. When BER(t) > 10 -3 increase the weight of ω, and when D(t) > 0.8T_max, increase ξ;
[0026] Step 5.5: Transfer the stored transition samples to the experience replay buffer for subsequent training.
[0027] Among them, the method described in Step 6 specifically includes the following steps:
[0028] Step 6.1: Calculate the absolute value of the temporal difference error |δ i | for each sample i in the experience replay buffer to measure the difference between the estimated value and the actual value;
[0029] Step 6.2: Set the formula for dynamically allocating the sampling probability p i as: η = 10 -5is the smoothing factor, which avoids zero-probability sampling, gives samples with larger errors a higher sampling probability, and improves the training efficiency;
[0030] Step 6.3: According to the samples obtained by sampling, update the parameters of the Actor and Critic networks through the stochastic gradient descent method to optimize the strategy;
[0031] Step 6.4: Perform a soft update of the target network, and its update formula is: θ'←ρθ+(1-ρ)θ'(ρ=0.01) where θ is the parameter of the main network, θ' is the parameter of the target network, and ρ is the coefficient of the soft update;
[0032] Step 6.5: After completing the parameter update, the transmitter sends the target data by adjusting the symbol interval.
[0033] The specific method shown in Step 7 specifically includes the following steps:
[0034] Step 7.1: Align the bit error rate and delay measurement values fed back by the receiver with the current channel state information in time;
[0035] Step 7.2: Set the length of the sliding window to 4, store the state information of the current and the previous 3 moments {s(t), s(t-1),..., s(t-3)}, extract the temporal features through the LSTM network, h(t)=LSTM(s(t), h(t-1)), and output the hidden state h(t)∈R 64 as the input of the policy network to enhance the channel dynamics modeling ability and capture the time series changes of the channel state;
[0036] Step 7.3: Set the initial exploration rate of the ε-greedy policy to 0.3, decay by 0.99ε after every 100 action selections, randomly and uniformly sample in the action space [-0.1T(t-1), +0.1T(t-1)], and when ε<0.05, completely rely on the Actor network to output actions;
[0037] Step 7.4: Set an emergency fallback mechanism. If the BER(t) exceeds 10 -3 the threshold for 3 consecutive times, automatically switch the symbol interval to 1.5T_min, and at the same time freeze the update of the Critic network and only update the Actor network; if the BER(t) is less than 10 -4 for 10 consecutive times, resume the regular training, so as to achieve lower delay and bit error rate. Compared with the prior art, the present invention has the following obvious advantages:
[0038] It can adaptively adjust the symbol interval, effectively cope with the time-varying characteristics of the channel, avoid the redundancy of the guard interval or inter-symbol interference, and improve the performance and response speed of underwater acoustic communication.
[0039] The designed dynamic weighted reward function can autonomously adjust the weight coefficient through online reinforcement learning, optimize the delay-bit error rate conflict, and achieve the joint optimization of low delay and low bit error rate.
[0040] By adopting the LSTM network and the sliding window mechanism, the modeling ability of the dynamic changes of the channel is enhanced. Combining the ε-greedy strategy and the emergency fallback mechanism, the system has stronger robustness and adaptability in complex channel environments. Brief Description of the Drawings
[0041] The present invention will be further described below with reference to the drawings and embodiments.
[0042] Figure 1 It is the overall flowchart of the delay-bit error joint optimization method for underwater acoustic communication with dynamic symbol interval based on reinforcement learning provided by the embodiment of the present invention;
[0043] Figure 2 It is the flowchart of the DDPG algorithm update strategy provided by the embodiment of the present invention;
[0044] Figure 3 It is the structural diagram of the LSTM network provided by the embodiment of the present invention; Detailed Embodiments
[0045] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0046] Referring to "embodiments" in the present invention means that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described in the present invention can be combined with other embodiments.
[0047] As Figure 1 shown, the delay-bit error joint optimization method for underwater acoustic communication with dynamic symbol interval based on reinforcement learning provided by the embodiment of the present invention includes the following steps:
[0048] S1. Based on the maximum channel delay spread \(T_{min}\) and the maximum allowable system delay \(T_{max}\), set the symbol interval adjustment range to \([1.2T_{min}, 0.8T_{max}]\).
[0049] S2. Initialize the parameters of the Actor network and the Critic network of the DDPG algorithm;
[0050] Specifically, the network structure used in this embodiment is as follows: The Actor network of the DDPG algorithm adopts a fully connected layer structure, the dimension of the input layer matches the state vector, and the activation function of the output layer is Tanh; the input of the Critic network is the state vector and the action value, and the output is the Q-value estimate. Create an experience replay buffer for storing state transition samples.
[0051] S3. Collect the channel state and calculate performance metrics, including the multipath delay distribution, Doppler frequency shift amount, signal-to-noise ratio, current bit error rate, average transmission delay, and normalize the historical sequence of symbol intervals to obtain the state vector.
[0052] S4. Determine whether the state of the collected channel and the calculated performance meet the standards.
[0053] S5. Input the state vector into the Actor-Critic network architecture to output action parameters, transform the action parameters into symbol adjustment amounts, so as to automatically adjust to obtain a new symbol interval, input the generated new state into the Critic network, and the network directly outputs a scalar Q-value to implement the update strategy, as Figure 2 shown.
[0054] S6. Define a dynamic weighted reward function, calculate the reward value, and then store the experience replay;
[0055] S61. Set the bit error rate incentive term, delay penalty term, and action smoothing term respectively;
[0056] S62. Set the weight dynamic selection strategy, and adjust different weight coefficients according to the bit error rate and delay;
[0057] S63. Calculate the absolute value of the temporal difference error \(\delta\) for each sample \(i\) in the experience replay buffer i ;
[0058] S64. Set the formula for the dynamically allocated sampling probability \(p\) i : \(\eta = 10\) -5 is the smoothing factor to avoid zero-probability sampling.
[0059] S7. Update the parameters of the Actor and Critic networks by the stochastic gradient descent method.
[0060] S8, perform a soft update on the target network, and its update formula is: θ'←ρθ+(1-ρ)θ'(ρ=0.01) where θ is the parameter of the main network, θ' is the parameter of the target network, and ρ is the coefficient of the soft update.
[0061] S9, the transmitter sends data according to the adjusted symbol interval period.
[0062] S10, after receiving the data, the receiver calculates the bit error rate and the time delay, and feeds back the status information to the transmitter.
[0063] S11, perform time alignment, set a sliding window to construct a time series correlation state vector through the LSTM network, adopt the ε-greedy strategy for action exploration, and set an emergency fallback mechanism;
[0064] S111, align the bit error rate and time delay measurement values fed back by the receiver with the current channel state information in time to ensure that these data are compared and analyzed at the same timestamp to eliminate the impact of time deviation on the optimization process;
[0065] S112, set the length of the sliding window to 4, store the state information of the current and the previous 3 moments to form a state sequence containing 4 time steps, extract the time series features through the LSTM network, and output the hidden state as the input of the policy network. Its algorithm process is as Figure 3 shown;
[0066] S113, set the initial exploration rate of the ε-greedy strategy to 0.3, decay by 0.99ε after every 100 action selections, randomly and uniformly sample in the action space [-0.1T(t-1), +0.1T(t-1)], and when ε<0.05, completely rely on the Actor network to output actions;
[0067] S114, set the emergency fallback mechanism. If when BER(t) exceeds the -3 threshold for 3 consecutive times, automatically switch the symbol interval to 1.5T_min, and at the same time freeze the update of the Critic network and only update the Actor network, which can quickly adopt a strategy when the bit error rate is too high to prevent the serious decline of the communication quality; if when BER(t) is less than the -4 for 10 consecutive times, resume the regular training, continue to optimize the symbol interval, effectively cope with the sudden channel deterioration situation, ensure the stability and reliability of the communication system, and thus achieve a lower time delay and bit error rate.
[0068] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Without departing from the original technical scope of the invention, several improvements and deformations can still be made, and all of these should be included within the protection scope of the present invention.
Claims
1. A dynamic symbol spacing underwater acoustic communication delay-error joint optimization method based on reinforcement learning, characterized in that: The method comprises: Step 1: Based on the maximum delay spread T_min of the channel and the maximum allowable delay T_max of the system, set the symbol interval adjustment range to [1.2T_min, 0.8T_max]; Step 2: Initialize the Actor network and Critic network parameters of the DDPG algorithm and build the experience replay buffer; Step 3, real-time perception of channel state information, the channel state information includes multipath delay distribution τ(t), Doppler frequency shift Δf(t), signal-to-noise ratio SNR(t); and calculation of system performance indicators, the performance indicators include current bit error rate BER(t), average transmission delay D(t) and symbol interval history sequence {T(t-1), T(t-2), ..., T(tN)} (N≥3); Step 4, inputting the channel state information and the performance index into the Actor-Critic network architecture, and dynamically optimizing the symbol interval after the symbol interval adjustment amount is output by the network architecture; Step 5, according to the current bit error rate and delay, calculate the bit error rate incentive term, the delay penalty term and the action smoothing term, define a dynamic weighted reward function, and transfer the storage sample to the experience playback buffer; Step 6, sampling from the experience playback buffer, calculating the timing difference error, updating the Actor and Critic network parameters, performing a soft update of the target network, and the transmitter sends data according to the adjusted symbol interval; In step 7, the state information of the receiving end is fed back for time alignment, a sliding window is set to construct a time-related state vector through the LSTM network, an ε-greedy strategy is used for action exploration, and an emergency fallback mechanism is set to achieve low latency and low bit error rate.
2. According to claim 1, a dynamic symbol spacing underwater acoustic communication delay-error joint optimization method based on reinforcement learning is characterized in that: In step 1, the T_min represents 1.2 times of the maximum delay extension of the channel, and the T_max represents 0.8 times of the maximum allowable delay of the system, which comprehensively considers the channel characteristics and the tolerance of the system to delay, and provides a reasonable range for the dynamic adjustment of the subsequent symbol interval.
3. According to claim 1, a dynamic symbol spacing underwater acoustic communication delay-error joint optimization method based on reinforcement learning is characterized in that: Step 2: The Actor network of the DDPG algorithm adopts a fully connected layer structure. The input layer dimension matches the state vector, and the output layer activation function is Tanh. The Critic network input is the state vector and action value, and the output is the Q value estimate. An experience replay buffer is created to store state transition samples.
4. According to claim 1, a dynamic symbol spacing underwater acoustic communication delay-error joint optimization method based on reinforcement learning is characterized in that: Step 3, normalize the multipath delay distribution τ(t), Doppler frequency shift Δf(t), signal-to-noise ratio SNR(t), current bit error rate BER(t), average transmission delay D(t) and symbol interval historical sequence {T(t-1), T(t-2), ..., T(tN)} (N≥3) to construct a state vector s(t).
5. According to claim 1, a dynamic symbol spacing underwater acoustic communication delay-error joint optimization method based on reinforcement learning is characterized in that: Step 3, the normalization uses the minimum and maximum normalization, and its formula is: Among them, x i (t) represents the value of the above parameters at time t, min(x i ) and max(x i ) represent the minimum and maximum values of the parameter in the entire data set, respectively. The state vector s(t)=[s1(t),s1(t),...,s n (t)].
6. According to claim 1, a dynamic symbol spacing underwater acoustic communication delay-error joint optimization method based on reinforcement learning is characterized in that: Step 4, input the state vector s(t) into the Actor-Critic network architecture, output the action a(t)∈[-1,1], output the symbol interval adjustment value ΔT(t)=0.1·T(t-1)·a(t) through the network architecture, ΔT(t)∈[-0.1T(t-1),+0.1T(t-1)], dynamically optimize the symbol interval, the symbol interval T(t)=T(t-1)+ΔT(t), T(t)∈[T_min,T_max].
7. The method for joint optimization of delay and error in underwater acoustic communication with dynamic symbol spacing based on reinforcement learning according to claim 1 is characterized in that: In step 5, the dynamic weighted reward function is: Among them, the weight coefficients ω, ξ, and ψ are dynamically adjusted according to the QoS requirements, and the initial values are set to ω=0.6, ξ=0.3, and ψ=0.1, and the storage transfer samples are transferred to the experience playback buffer.
8. The method for joint optimization of delay and error in underwater acoustic communication with dynamic symbol spacing based on reinforcement learning according to claim 1 is characterized in that: Step 6: The experience replay training adopts a priority experience replay mechanism to dynamically allocate sampling probabilities, and its function is: where δ i is the timing difference error of sample i, η=10 -5 is a smoothing factor, avoiding zero-probability sampling and accelerating the learning of key experiences.
9. The method for joint optimization of delay and error in underwater acoustic communication with dynamic symbol spacing based on reinforcement learning according to claim 1 is characterized in that: Step 6: Update the Actor and Critic network parameters by stochastic gradient descent method and perform soft update strategy. The update formula is: θ'←ρθ+(1-ρ)θ'(ρ=0.01) Where θ is the parameter of the main network, θ' is the parameter of the target network, and ρ is the coefficient of the soft update; after completing the parameter update, the transmitter sends the target data by adjusting the symbol interval.
10. According to the method for joint optimization of delay-bit error of underwater acoustic communication with dynamic symbol interval based on reinforcement learning in claim 1, step 7, time-aligning the bit error rate and delay measurement value fed back by the receiving end with the current channel state information, setting the sliding window to 4, constructing the time series associated state vector through the LSTM network, driving the DDPG strategy network, using the ε greedy strategy, setting the initial exploration rate to 0.3, decaying by 0.99ε after completing 100 action selections, dynamically adjusting the symbol interval T(t), and setting an emergency fallback mechanism, when BER(t) exceeds 10 for 3 consecutive times -3 When the threshold is reached, it automatically switches to the 1.5T_min conservative interval and starts lightweight retraining to achieve lower latency and bit error rate.
Citation Information
Cited By
Fault detection method for multiple communication channels and self-healing control device
CN120639588A
AUV autonomous collision avoidance planning method based on LSTM-DDPG
CN121724059A