The invention discloses a
reinforcement learning routing optimization method considering a dynamic
underwater sound environment and node mobility. The method comprises the following steps: performing parameter initialization on an
underwater sensor network; monitoring and updating
underwater environment parameters and node positions in real time, and dynamically updating a
network topology connection state; during
data transmission, constructing a multi-dimensional
state vector, selecting a routing action by adopting an epsilon-greedy strategy, and generating a
routing decision; if the data
packet transmission fails when the routing action is executed, ending the iteration, otherwise, calculating a multi-target
reward value; updating the Q value table to optimize the
routing decision; and monitoring
network performance indexes in real time, adaptively adjusting
reinforcement learning parameters of the
intelligent agent for next iteration, and outputting an
adaptive routing selection model until an iteration termination condition is met. According to the method, time-varying channel and node mobility challenges can be intelligently dealt with, and multi-objective optimization balance of
energy consumption,
delay and stability is realized while the
data delivery rate is remarkably improved.