Directional MAC protocol design method based on DQN in underwater acoustic communication

By adopting the DQN-based directional MAC protocol in the hydroacoustic communication network, optimizing the channel capacity and interference model, the high collision rate and low throughput problems of traditional MAC protocol in complex environments is solved, and the network throughput is maximized and stability is improved.

CN120390039APending Publication Date: 2025-07-29HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510610898.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the existing water acoustic communication network, the traditional omnidirectional communication MAC protocol is difficult to achieve efficient resource scheduling and node cooperation in complex environments, resulting in high conflict rate, low throughput and unfair resource allocation. The existing water acoustic directed communication MAC protocol fails to effectively improve the utilization rate of communication time slots and idle nodes to improve network capacity.

Method used

The directional MAC protocol design method based on deep Q network (DQN) is adopted, and the mutual interference between nodes is calculated and the deep Q network is combined with the deep Q network to optimize the channel capacity to achieve the improvement of the channel capacity of the surrounding nodes to maximize the throughput of the water acoustic network.

Benefits of technology

It significantly improves the throughput of the water acoustic communication network, reduces the probability of conflict between nodes, and enhances the adaptability and stability of the system in an underwater environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390039A_ABST
    Figure CN120390039A_ABST
Patent Text Reader

Abstract

The invention discloses a directional MAC protocol design method based on a DQN in underwater acoustic communication. The method comprises the following steps: step 1, constructing an underwater acoustic network model; 2, establishing an underwater acoustic communication directional interference model, calculating signal receiving power, and establishing a directional signal interference model; 3, establishing a deep reinforcement learning model; 4, performing DQN updating training by using the data stored in the experience playback buffer area, and applying the trained model to an underwater acoustic network MAC protocol; step 5, storing the new reward of the underwater acoustic communication node and the corresponding state-action pair in an experience playback buffer area, and training based on a dual-network mechanism of a DQN; and step 6, repeating the steps 4-5 until the maximum network throughput is reached. According to the method, mutual interference between nodes is calculated, a deep Q network is combined, and the channel capacity of other surrounding nodes is improved while the large channel capacity of the nodes is considered, so that the throughput of an underwater acoustic network is improved to the maximum extent, and the special requirements of the underwater acoustic network are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of media access control (MAC) protocols for underwater acoustic sensor networks. Specifically, it relates to a method for designing a directional MAC protocol based on DQN in underwater acoustic communication. Background Art

[0002] In recent years, with the country's emphasis on the marine strategy, as an important component of underwater wireless communication, underwater acoustic communication network technology (UWANs) has become one of the key technologies for marine development. Underwater acoustic communication technology not only has a wide range of applications in the civilian field but also plays an increasingly important role in the military and commercial fields. In an underwater acoustic wireless communication network (UWANs), the core responsibility of the media access control (MAC) protocol is to manage the shared use of the common underwater acoustic channel by network nodes to avoid data packet collisions and ensure the fairness of competition among network nodes. Traditional omnidirectional communication MAC protocols, such as TDMA, ALOHA, etc., cannot achieve efficient resource scheduling and node cooperation in such a complex environment, easily leading to high collision rates, low throughput, and unfair resource allocation. Therefore, developing a MAC protocol that can improve the overall data transmission efficiency of the network and reduce the probability of data packet collisions is a key task in the current research field of underwater acoustic wireless communication networks.

[0003] At the same time, the emergence of underwater acoustic directional communication technology has greatly optimized the throughput, collision rate, and resource allocation efficiency of underwater acoustic communication networks. However, the design of its MAC protocol is quite different from the traditional omnidirectional communication method. Although the current underwater acoustic directional communication MAC protocols can all function well on the nodes of underwater acoustic directional communication, there is no underwater acoustic directional communication MAC protocol that can improve the utilization rate of communication time slots and idle nodes, thereby greatly improving the overall network capacity of underwater acoustic networks. Summary of the Invention

[0004] In view of the above problems, the present invention designs a method for designing a directional MAC protocol based on (Deep Q-Network) DQN in underwater acoustic communication. This method calculates the mutual interference between nodes and combines the deep Q-network to improve the channel capacity of the surrounding nodes while taking into account its own relatively large channel capacity, with the aim of maximizing the throughput of the underwater acoustic network and meeting the special requirements of underwater acoustic networks.

[0005] The technical solution of the present invention is as follows:

[0006] A method for designing a directional MAC protocol based on DQN in underwater acoustic communication, comprising the following steps:

[0007] Step 1: Construct an underwater acoustic network model; the underwater acoustic network includes several underwater nodes, which respectively use TDMA, ALOHA, and DIR-MAC protocols and are deployed in the underwater area in a random distribution manner. They are used to collect perception data from the surrounding environment and perform packet reception or transmission operations according to the decision-making mechanism of the MAC protocol;

[0008] Step 2: Establish an underwater acoustic communication directional interference model. Calculate the signal reception power according to the signal diffusion loss, absorption loss, and directional beam directivity transceiver gain, and establish a directional signal interference model;

[0009] 2.1 Construction of the underwater acoustic communication directional interference model, the steps are as follows:

[0010] 2.1.1 Diffusion loss: The path loss PL is the main loss source of the underwater acoustic channel, which includes diffusion loss PL spread and absorption loss PL absorption . The diffusion loss is related to the transmission distance d and the diffusion factor k, and the diffusion loss is expressed as:

[0011] PL spread = 10klog 10 (d)+C (1)

[0012] where C is the reference distance normalization constant;

[0013] 2.1.2 The absorption loss is related to the transmission distance d and the attenuation coefficient, and is expressed as

[0014] PL absorption = α(f)d (2)

[0015] where α(f) is the frequency-dependent attenuation coefficient;

[0016] 2.1.3 Calculate the received power P r

[0017] P r = P t + G t + G r - PL (3)

[0018] where P t , G t , G r , PL are the signal transmission power, directional transmission gain, directional reception gain, and path loss respectively;

[0019] 2.1.4 Assume that N directional nodes transmit data simultaneously, and the co-frequency interference of the signals obtained at the receiving end is modeled as:

[0020]

[0021] where P r,i is the power contribution of the i-th transmitting node at the receiving end, i.e., the received power.

[0022] Step 3: Establish a deep reinforcement learning model. Construct a signal interference model of the channel according to the underwater directional communication nodes. Calculate the reward of the communication node through the reward function to evaluate the effect of the node performing an action in a certain state at a past moment. Store the reward and the corresponding state-action pair in the experience replay buffer for the subsequent training of the agent;

[0023] The settings of the agent, action, state, and reward function are as follows:

[0024] Agent: Each underwater acoustic communication node using the DIR-MAC protocol is an agent;

[0025] Directional beam selection: The action of agent i ∈ {1, 2,..., n} at time slot t is defined as deciding whether to choose to occupy this time slot for data transmission; where n represents constructing a channel with the receiving nodes within the sensing range of the agent as the transmitting node, and 0 represents being the receiving node and preparing to receive data;

[0026] 3.1 Agent node directional beam selection, the steps are as follows:

[0027] 3.1.1 The agent node is composed of N equally sized transmitting beam sectors, and there is no overlap between these transmitting beam sectors. The agent node can send signals in a specific direction through any beam transmitting sector at any time;

[0028] 3.1.2 When making an action decision, the agent node can choose any beam transmitting sector as the transmitting node of the data packet to construct a channel with a certain receiving node, or as the receiving party of the data packet to construct a channel with the transmitting nodes around it that send data packets to it;

[0029] For each agent, classify the data packets transmitted by all transmitting nodes within its sensing range according to the beam direction selected by the transmitting node, for statistically calculating the interference signals that will interfere with the agent, that is, statistically calculating the transmitting nodes within the sensing range that will affect this channel. According to the interference modeling, the signal-to-interference-plus-noise ratio (SINR) under the interference condition is:

[0030]

[0031] where N(f) represents the noise power spectral density in the underwater acoustic network; according to the SINR and the Shannon formula, calculate the directional channel capacity R i formed by the transmitting node S i and the receiving node D i, accumulate to obtain the channel capacity of the entire network;

[0032] State: Calculate the interference caused by all interfering nodes to the receiving node according to the signal co-frequency interference model described in step 2, and obtain the total interference value I suffered by the agent;

[0033] Reward function: If the action is to send, then determine whether the action of the receiving node is n. If not, record the reward as 0; if so, then further determine whether the sending node is the sending node closest to the receiving node (the signal arrives first); if so, calculate the reward according to the following reward calculation method, if not, also record the reward as 0. If the decision result of the action is to receive, calculate the reward according to the reward calculation method if a signal is received, otherwise record the reward as 0;

[0034] Calculate the own channel capacity, that is, the own channel capacity R, through the above SINR and Shannon formula i . To achieve the goal of maximizing the throughput of the underwater acoustic network, we set the reward R of the agent e as the income of all nodes that successfully build channels within its sensing range:

[0035] R e =R i +∑R j (6)

[0036] R j represents the income of other nodes that successfully build channels within the sensing range of the agent, that is, the final reward, and is also calculated using the above SINR and Shannon formula.

[0037] Step 4: Use the data stored in the experience replay buffer to update and train the DQN, and apply the trained model to the underwater acoustic network MAC protocol;

[0038] In DQN, the goal of each agent is to find an optimal policy that maximizes the expected cumulative discounted reward Similarly, the goal of the underwater acoustic sensor network is to find a scheduling that can maximize the throughput. Therefore, we use the deep Q-network (DQN) to handle the channel scheduling problem in the underwater acoustic network, and continuously update the network parameters for training by using the data stored in the experience replay buffer, so as to learn the optimal channel scheduling strategy to achieve the maximization of network throughput.

[0039] In DQN, the max operator is used for both selecting actions and calculating the Q-value of actions; DQN1 is used to find the action with the largest Q-value, and DQN2 is used to estimate the target Q-value; where DQN1 and DQN2 have the same network structure;

[0040] The network structure specifically includes: an input layer, five hidden layers, and an output layer. The number of neurons in the input layer matches the state dimension. Each hidden layer contains 64 neurons, and these hidden layers all use the ReLU activation function and Glorot normal distribution to initialize the weights. The number of neurons in the output layer corresponds to the number of actions in the action space. The network receives the channel state as input and outputs the Q value of each action under state s(t).

[0041] At each time slot, the trained DQN1 makes decisions on the agent's actions. All Q values are generated by DQN1. The agent selects an action according to the ε-greedy strategy and then updates the network state.

[0042] Step 5: Store the new rewards of the underwater acoustic communication nodes and the corresponding state-action pairs in the experience replay buffer, and train based on the double-network mechanism of DQN;

[0043] In the double-network mechanism of DQN, DQN1 updates the target value at each time slot And DQN2 adjusts the target value every once in a while DQN1 updates the θ value of the main network through the loss function:

[0044]

[0045] N E is the number of samples drawn from the experience replay area. After a certain period of time, DQN2 updates the value θ in DQN2 - to the θ value in DQN1.

[0046] Step 6: Repeat Step 4 - Step 5 until the maximum network throughput is reached.

[0047] The beneficial effects of the present invention are as follows:

[0048] By adopting the DQN-based directional MAC protocol design method, the present invention significantly improves the throughput of the underwater acoustic communication network and effectively reduces the probability of node collisions. By introducing a directional transducer and an interference model, the agent can optimize decisions in real time in the underwater environment, taking into account the communication performance of itself and surrounding nodes to maximize the network throughput. The present invention not only effectively improves the overall efficiency of the underwater acoustic communication network, but also significantly reduces the probability of collisions between communication nodes, enhancing the adaptability and stability of the system in the underwater acoustic environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a diagram of the underwater acoustic network model;

[0050] Figure 2 is a schematic diagram of the model;

[0051] Figure 3 It is a framework diagram of a deep reinforcement learning algorithm. Specific implementation manners

[0052] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0053] A design method for a directional MAC protocol based on DQN in underwater acoustic communication includes the following steps:

[0054] Step 1: Construct an underwater acoustic network model;

[0055] As Figure 1 shown, the underwater acoustic network includes several underwater nodes. The underwater nodes respectively use TDMA, ALOHA, and DIR-MAC protocols and are deployed in the underwater area in a random distribution manner to collect sensing data from the surrounding environment and perform packet receiving or sending operations according to the decision-making mechanism of the MAC protocol.

[0056] Step 2: Establish an underwater acoustic communication directional interference model, calculate the signal reception power according to the signal diffusion loss, absorption loss, and directional beam directivity transceiver gain, and establish a directional signal interference model.

[0057] Step 3: Establish a deep reinforcement learning model, construct a signal interference model of the channel according to the underwater directional communication node, calculate the reward of the communication node through the reward function to evaluate the effect of the node performing an action in a certain state at a past moment; store the reward and the corresponding state-action pair in the experience replay buffer for subsequent training of the agent.

[0058] Step 4: As Figure 3 shown, use the data stored in the experience replay buffer to update and train the deep Q network, and apply the trained model to the MAC protocol of the underwater acoustic network.

[0059] Step 5: Store the new reward and the corresponding state-action pair of the underwater acoustic communication node in the experience replay buffer, and perform training based on the double-network mechanism of DQN.

[0060] Step 6: Repeat Step 4 - Step 5 until the maximum network throughput is reached.

[0061] The specific steps of Step 2 are as follows:

[0062] 2.1 Construction of the directional interference model for underwater acoustic communication is as follows:

[0063] 2.1.1 Spreading loss: The path loss PL is the main loss source in the underwater acoustic channel, which mainly includes the spreading loss PL spread and the absorption loss PL absorption . The spreading loss is related to the transmission distance d and the spreading factor k, and the spreading loss can be expressed as:

[0064] PL spread = 10k log 10 (d) + C

[0065] where C is the reference distance normalization constant.

[0066] 2.1.2 The absorption loss is related to the transmission distance d and the attenuation coefficient, and can be expressed as

[0067] PL absorption = α(f)d

[0068] where α(f) is the frequency-dependent attenuation coefficient.

[0069] 2.1.3 Calculate the received power P r .

[0070] P r = P t + G t + G r - PL

[0071] where P t , G t , G r , and PL are the signal transmission power, the directional transmission gain, the directional reception gain, and the path loss respectively.

[0072] 2.1.4 Assume that N directional nodes transmit data simultaneously, and the co-channel interference of the signals obtained at the receiving end is modeled as:

[0073]

[0074] where P r,i is the power contribution of the i-th transmitting node at the receiving end, that is, the received power.

[0075] The specific steps of Step 3 are as follows:

[0076] Settings for the agent, directional beam selection, state, and reward function are as follows:

[0077] Agent: Each underwater acoustic communication node using the DIR-MAC protocol is an agent;

[0078] Directional beam selection: The selected directional beam of agent \(i\in\{1, 2, \ldots, n\}\) at time slot \(t\) is defined as decide whether to select and occupy this time slot for data transmission; where \(n\) represents building a channel with the receiving nodes within the sensing range of the agent as the sending node, and \(0\) represents being the receiving node and preparing for data reception.

[0079] 3.1 Agent node directional beam selection, the steps are as follows:

[0080] 3.1.1 As Figure 2 shown, the agent is composed of \(N\) equal-sized transmitting beam sectors, and there is no overlap between these transmitting beam sectors. The agent can send signals in a specific direction through any beam transmitting sector at any time.

[0081] When making an action decision, the agent node can act as the sending node of the data packet to select any beam transmitting sector to build a channel with a certain receiving node, or act as the receiving party of the data packet to build a channel with the sending nodes around it that send data packets to it.

[0082] For each agent, classify the data packets transmitted by all transmitting nodes within its sensing range according to the beam direction selected by the sending node, which is used to count the interference signals that will cause interference to the agent, that is, count the sending nodes within the sensing range that will affect this channel. According to the interference modeling, the signal-to-interference-plus-noise ratio (SINR) under interference conditions is:

[0083]

[0084] where \(N(f)\) represents the noise power spectral density in the underwater acoustic network; according to SINR and the Shannon formula, calculate the directional channel capacity \(R\) i formed by the sending node \(S\) i and the receiving node \(D\) i , and accumulate to obtain the channel capacity of the entire network.

[0085] State: Calculate the interference caused by all interfering nodes to the receiving node according to the co-channel interference model described in step 2, and obtain the total interference value \(I\) suffered by this agent

[0086] Reward function: If the action is to send, then judge whether the action of the receiving node is \(n\). If not, record the reward as \(0\); if so, then further judge whether the sending node is the sending node closest to this receiving node (the signal arrives first). If so, calculate the reward according to the following reward calculation method. If not, also record the reward as \(0\). If the decision result of the action is to receive, if a signal is received, calculate the reward according to the reward calculation method, otherwise record the reward as \(0\).

[0087] Calculate its own channel capacity, i.e., its own channel capacity R, through the above SINR and Shannon formula i To achieve the goal of maximizing the throughput of the underwater acoustic network, we set the reward R of the agent e as the revenue of all nodes that successfully build channels within its sensing range:

[0088] R e = R i + ∑R j

[0089] R j represents the revenue of other successfully built channels within the agent's sensing range, i.e., the final reward, which is also calculated using the above SINR and Shannon formula

[0090] The specific steps of Step 4 are as follows:

[0091] In DQN, the purpose of each agent is to find an optimal policy that maximizes the expected cumulative discounted reward Similarly, the goal of the underwater acoustic sensor network is to find a scheduling that can maximize the throughput. Therefore, we use the Deep Q-Network (DQN) to handle the channel scheduling problem in the underwater acoustic network. By continuously updating the network parameters using the data stored in the experience replay buffer for training, an optimal channel scheduling policy is learned to achieve the maximization of the network throughput

[0092] In the deep reinforcement learning algorithm DQN, the max operator is used for both selecting actions and calculating the Q-value of actions; DQN1 is used to find the action with the largest Q-value, and DQN2 is used to estimate the target Q-value; where DQN1 and DQN2 have the same network structure

[0093] The network structure specifically includes: one input layer, five hidden layers, and one output layer. The number of neurons in the input layer matches the state dimension. Each hidden layer contains 64 neurons, and these hidden layers all use the ReLU activation function and Glorot normal distribution to initialize the weights. The number of neurons in the output layer corresponds to the number of actions in the action space. The network receives the channel state as input and outputs the Q-value of each action in state s(t)

[0094] At each time slot, the trained DQN1 makes decisions on the agent's actions. All Q-values are generated by DQN1. The agent selects an action according to the ε-greedy policy and then updates the network state

[0095] The specific steps of Step 5 are as follows:

[0096] DQN1 updates the target value at each time slot And DQN2 adjusts the target value at regular intervals. DQN1 updates the value of the main network θ through the loss function:

[0097]

[0098] N E is the number of samples drawn from the experience replay area. After a certain period of time, DQN2 updates the value θ in DQN2 - to the θ value in DQN1.

Claims

1. A design method of a directional MAC protocol based on DQN in underwater acoustic communication, characterized in that Including the following steps: Step 1: Construct an underwater acoustic network model; the underwater acoustic network includes a number of underwater nodes, which respectively use TDMA, ALOHA, and DIR-MAC protocols and are deployed in the underwater area in a random distribution manner, used to collect sensing data from the surrounding environment, and perform packet reception or transmission operations according to the decision-making mechanism of the MAC protocol; Step 2: Establish an underwater acoustic communication directional interference model, calculate the signal reception power according to the signal diffusion loss, absorption loss, and directional beam directivity transceiver gain, and establish a directional signal interference model; Step 3: Establish a deep reinforcement learning model. According to the signal interference model of the channel constructed by the underwater directional communication node, calculate the reward of the communication node through the reward function to evaluate the effect of the node's action in a certain state at a past moment, and store the reward and the corresponding state-action pair in the experience replay buffer for subsequent training of the agent; Step 4: Use the data stored in the experience replay buffer to update and train DQN, and apply the trained model to the underwater acoustic network MAC protocol; Step 5: Store the new reward of the underwater acoustic communication node and the corresponding state-action pair in the experience replay buffer, and perform training based on the double-network mechanism of DQN; Step 6: Repeat Step 4 - Step 5 until the maximum network throughput is reached.

2. The method for designing a directional MAC protocol based on DQN in underwater acoustic communication according to claim 1, wherein The specific steps of Step 2 are as follows: 2.1 Construction of the underwater acoustic communication directional interference model, the steps are as follows: 2.1.1 Spreading loss: The path loss PL is the main loss source in the underwater acoustic channel, including the spreading loss PL spread and the absorption loss PL absorption , the spreading loss is related to the transmission distance d and the spreading factor k, and the spreading loss is expressed as: PL spread = 10klog 10 (d)+C (1) Where C is the reference distance normalization constant; 2.1.2 The absorption loss is related to the transmission distance d and the attenuation coefficient, expressed as PL absorption = α(f)d (2) Where α(f) is the frequency-dependent attenuation coefficient; 2.1.3 Calculate the received power P r P r = P t + G t + G r - PL(3) where P t , G t , G r , and PL are the signal transmission power, the directional transmission gain, the directional reception gain, and the path loss, respectively; 2.1.4 Assume that N directional nodes transmit data simultaneously, and the co-channel interference of the signals obtained at the receiving end is modeled as: Where P r,i is the power contribution of the i-th transmitting node at the receiving end, i.e., the received power.

3. According to the method for designing a directional MAC protocol based on DQN in underwater acoustic communication as described in claim 1, wherein the specific steps of Step 3 are as follows: The settings of the agent, action, state, and reward function are as follows: Agent: Each underwater acoustic communication node using the DIR-MAC protocol is an agent; Directional beam selection: The action of agent \(i\in\{1, 2,\cdots, n\}\) at time slot \(t\) is defined as deciding whether to select and occupy this time slot for data transmission; where \(n\) represents building a channel with the receiving nodes within the sensing range of the agent as the sending node, and \(0\) represents being the receiving node and preparing for data reception. 3.1 Selection of the directional beam of the agent node, the steps are as follows: 3.1.1 The agent node is composed of N equal-sized transmitting beam sectors without overlap. The agent node can send signals in a specific direction through any beam transmitting sector at any time; 3.1.2 When making an action decision, the agent node can act as the sending node of the packet to select any beam transmitting sector to construct a channel with a certain receiving node, or act as the receiving party of the packet to construct a channel with the sending node that sends packets to it in the vicinity; For each agent, classify the packets transmitted by all transmitting nodes within its sensing range according to the beam direction selected by the sending node, used to count the interference signals that will interfere with the agent, that is, count the sending nodes within the sensing range that will affect the channel. According to the interference modeling, the signal-to-interference-plus-noise ratio under interference conditions is obtained as: Among them, N(f) represents the noise power spectral density in the underwater acoustic network; according to the SINR and the Shannon formula, calculate the directional channel capacity R i formed by the sending node S i and the receiving node D i , and accumulate to obtain the channel capacity of the entire network; Status: Calculate the interference caused by all interfering nodes to the receiving node according to the signal co-frequency interference model described in step 2, and obtain the total interference value I suffered by the agent; Reward function: If the action is to send, then determine whether the action of the receiving node is n. If not, record the reward as 0; if so, then further determine whether the sending node is the sending node closest to the receiving node; If so, calculate the reward according to the following reward calculation method. If not, also record the reward as 0. If the decision result of the action is to receive, if a signal is received, calculate the reward according to the reward calculation method, otherwise record the reward as 0; Calculate its own channel capacity, i.e., its own channel capacity R, through the above SINR and Shannon formula i To achieve the goal of maximizing the throughput of the underwater acoustic network, we set the reward R of the agent e as the benefits of all nodes that successfully build channels within its sensing range: R e = R i + ∑R j (6) R j Indicates the reward of other successfully established channels within the perception range of the agent, that is, the final reward, which is also calculated using the above SINR and Shannon formula.

4. The method for designing a directional MAC protocol based on DQN in underwater acoustic communication according to claim 1, wherein The specific steps of step 4 are as follows: In DQN, both the selection of actions and the calculation of the Q-value of actions use the max operator; DQN1 is used to find the action with the largest Q-value, and DQN2 is used to estimate the target Q-value; Among them, DQN1 and DQN2 have the same network structure; The network structure specifically includes: one input layer, five hidden layers, and one output layer. The number of neurons in the input layer matches the state dimension. Each hidden layer contains 64 neurons, and these hidden layers all use the ReLU activation function and Glorot normal distribution to initialize the weights. The number of neurons in the output layer corresponds to the number of actions in the action space. The network receives the channel state as input and outputs the Q-value of each action in state s(t); At each time slot, the trained DQN1 is used to make decisions on the agent's actions. All Q-values will be generated by DQN1. The agent selects an action according to the ε-greedy strategy, and then updates the network state.

5. The design method of the directional MAC protocol based on DQN in underwater acoustic communication according to claim 4, characterized in that The specific steps of step 5 are as follows: In the double-network mechanism of DQN, DQN1 updates the target value in each time slot while DQN2 adjusts the target value every once in a while and DQN1 updates the θ value of the main network through the loss function: N E N is the number of samples drawn from the experience replay buffer. After a certain period of time, DQN2 updates the value θ_ in DQN2 to the θ value in DQN1.