The invention discloses a dynamic symbol interval
underwater acoustic communication time
delay-error code joint optimization method based on
reinforcement learning, and belongs to the technical field of
underwater acoustic communication. The method comprises the following steps: setting a symbol interval adjustment range; a depth deterministic policy gradient (DDPG)
algorithm is adopted, an Actor-Critic
network architecture is combined,
channel state information and
system performance indexes are sensed in real time, and symbol intervals are dynamically adjusted; designing a weighted reward function, fusing an error rate excitation item, a time
delay penalty item and an action
smoothing item, and balancing the conflict between the time
delay and the error rate through a weight dynamic adjustment strategy; the Actor and Critic network parameters are updated, and soft update of the target network is carried out; channel state
time sequence characteristics are extracted by using a sliding window and a
long short term memory (LSTM) network, and the
adaptive capacity to a time-varying channel is enhanced; an action space is explored in combination with an epsilon greedy strategy, and an emergency
rollback mechanism is introduced, so that the model convergence efficiency and the strategy optimization precision are improved. Compared with a traditional fixed symbol
interval method, the method can reduce the
bit error rate and time delay, and is suitable for
underwater sensor networks, unmanned underwater vehicles and other scenes.