Asynchronous cross-layer-based MAC protocol design method in underwater acoustic sensor network
Through asynchronous cross-layer MAC protocol and multi-agent reinforcement learning, time slot conflicts and energy waste caused by dynamic topological changes in the hydroacoustic sensor network are solved, and network throughput and energy efficiency are improved, and network life cycle is extended.
Patent Information
- Application Number
- CN202510610904.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-29
AI Technical Summary
Traditional MAC protocols are difficult to deal with dynamic topological changes caused by ocean currents, obstacles and node movements in water acoustic sensing networks, resulting in time slot conflicts and energy waste, affecting network throughput performance and life cycle.
A MAC protocol based on asynchronous span layer in a water acoustic sensor network is designed. Through cross-layer collaborative design and multi-agent reinforcement learning, sub-slot allocation, transmission power and relay node selection are dynamically adjusted to realize information sharing and collaborative optimization between the network layer and the MAC layer.
Improve network throughput and energy efficiency, adapt to underwater dynamic topology changes, avoid communication conflicts and resource waste, and extend the network life cycle.
Smart Images

Figure CN120390230A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of media access control (MAC) protocols for underwater acoustic sensor networks. Specifically, it relates to a design method for an asynchronous cross-layer based MAC protocol in underwater acoustic sensor networks. Background Art
[0002] In underwater acoustic sensor networks (UASNs), the MAC protocol is crucial for how nodes effectively access the shared channel and achieve efficient data transmission. However, due to uncertain factors such as ocean current disturbances, obstacle blockages, and environmental noise interference in the actual ocean environment, the network topology exhibits highly dynamic and unstable characteristics. Traditional MAC protocols usually assume that nodes can achieve precise synchronization and a stable topology. If directly applied to underwater scenarios, it is extremely easy to cause slot conflicts, multipath interference, and energy waste, severely limiting the network throughput performance and shortening the network lifetime.
[0003] Therefore, the core challenge in the design of underwater acoustic MAC protocols lies in how to effectively improve the network throughput and ensure communication reliability in a communication environment with high energy consumption and long propagation delays. Since the propagation speed of underwater sound waves is much lower than that of electromagnetic waves, it significantly increases the propagation delay, making it difficult for each node to achieve precise global time synchronization, thus triggering communication conflicts and causing waste of channel resources. In addition, most existing MAC protocols perform slot allocation and channel access control based on static or idealized assumptions, and it is difficult to fully cope with the complex dynamic scenarios brought by ocean currents, obstacles, and the movement of nodes themselves. Summary of the Invention
[0004] In view of the above problems, the present invention designs an asynchronous cross-layer based MAC protocol method for underwater acoustic sensor networks. This method fully considers the uncertainty and dynamicity of the underwater communication environment. Through cross-layer collaborative design, it realizes information sharing between the network layer and the MAC layer, and can effectively improve the network throughput, save energy consumption, and adapt to dynamic topology and channel condition changes in the underwater asynchronous environment.
[0005] The technical solution of the present invention is as follows:
[0006] A design method for an asynchronous cross-layer based MAC protocol in underwater acoustic sensor networks, comprising the following steps:
[0007] Step 1: Establish a network model of the underwater acoustic sensor network; the underwater acoustic sensor network includes several underwater sensor nodes, relay nodes, and a surface Sink node. The underwater sensor nodes use TDMA, slotted ALOHA, and DAC-MAC protocols respectively, and are randomly distributed in the underwater area. They collect the sensed data from the surrounding environment and forward it to the surface Sink node through the relay nodes. The surface Sink node broadcasts its location to each node, and all the other nodes need to calculate and store the distance between themselves and the surface Sink node.
[0008] Step 2: Establish an asynchronous time slot model, and convert the distances between each underwater sensor node and relay node and the surface Sink node into propagation delays in units of sub-time slots δ.
[0009] Step 3: Construct a multi-dimensional utility function that combines three key indicators: the remaining energy of the fusion node, the link quality, and the topological distance gain, and adaptively update the weight coefficients of each indicator of the utility function by real-time monitoring of the network state.
[0010] Step 4: Generate a dynamic candidate forwarding node set by calculating and sorting the utility values of the relay nodes and combining with a threshold screening.
[0011] Step 5: Deploy a multi-agent reinforcement learning algorithm for underwater sensor node deployment delay reward, set action, state, and reward functions, and through an asynchronous cross-layer scheduling mechanism, use multi-agent reinforcement learning at the MAC layer to dynamically adjust sub-time slot allocation, transmission power, and relay node selection strategies, so as to achieve information sharing and collaborative work between the network layer and the MAC layer, thereby meeting the real-time requirements of the network and improving the overall performance.
[0012] Preferably, the specific steps of Step 2 are as follows:
[0013] 2.1 In an asynchronous time slot system, assume that the packet lengths of different underwater sensor nodes and relay nodes are the same, and the returned ACK packets also have the same packet length.
[0014] 2.2 Each time slot length T slot corresponds to the duration T data of packet transmission plus the duration T ack of ACK packet transmission and the guard time t g , and further divide T slot into several smaller sub-time slots δ, satisfying Similarly, the propagation delay D i satisfies c is the underwater acoustic propagation speed, and d i is the distance between the underwater sensor node i and the surface Sink node;
[0015] 2.3 The transmission time of the underwater sensor node within a time slot can be flexibly selected, and the transmission time of each underwater sensor node is a specific delay point T within the time slot r [n,i] (n represents the nth time slot, i is the node number), satisfying:
[0016] T r [n,i] = m·δ, m ∈ {0, 1, 2,..., N - 1} (1)
[0017] The underwater sensor node autonomously selects the optimal transmission delay point T according to its own propagation delay to the receiving node (Sink or relay node) and the real-time network environment r [n,i], so as to maximize the risk of data collision with other nodes
[0018] Preferably, the specific steps of step 3 are as follows:
[0019] 3.1 Design a multi-index utility function, the steps are as follows:
[0020] 3.1.1 Remaining energy ratio; The energy of underwater nodes is limited and difficult to replenish. Energy efficiency is the key to extending the network lifetime. Prioritizing the selection of high-energy nodes to participate in forwarding can effectively balance the energy consumption of network nodes and extend the network lifetime; Therefore, the remaining energy ratio is taken as one of the considered indicators and is defined as follows:
[0021]
[0022] where E res is the remaining energy of the node, and E init is the initial energy of the node;
[0023] 3.1.2 Link quality; The underwater acoustic channel is significantly affected by multipath effects, frequency-selective fading, and environmental noise. The dynamic fluctuation of link quality directly determines the data transmission success rate. Therefore, the link quality LQ is used to represent the link quality between the neighbor node and the current node:
[0024]
[0025] where SNR max represents the maximum value of the signal-to-noise ratio among all neighbor nodes; SNR current represents the signal-to-noise ratio of the current node
[0026] 3.1.3 Topological distance gain; Since the propagation speed of underwater acoustic signals is low and the propagation delay dominates the end-to-end delay, therefore, selecting a node closer to the Sink node can reduce the number of hops and cumulative delay. Define D gain as the distance advantage after the node selects a neighbor node for forwarding:
[0027]
[0028] Among them, d current represents the distance from the current node to the water surface Sink node, and d neighbor represents the distance from the neighbor node to the water surface Sink node.
[0029] 3.1.4 Based on the above three indicators, the utility function of the candidate relay forwarding node j is designed as follows:
[0030]
[0031] Among them, α, β, and γ are weight coefficients, satisfying α + β + γ = 1; d(i, Sink) represents the distance from node i to the water surface Sink node; d(j, Sink); represents the distance from node j to the water surface Sink node; E j represents the remaining energy of node j; SNR ij represents the signal-to-noise ratio between node i and node j.
[0032] 3.2 Design an adaptive weight coefficient adjustment mechanism, and the steps are as follows:
[0033] 3.2.1 Periodically monitor the energy distribution index Among them, N is the total number of candidate nodes in the network, and E ratio,i is the remaining energy ratio of candidate node i, is the average value of the remaining energy ratios of all candidate nodes in the network; once it exceeds the threshold, it indicates that the energy distribution is unbalanced. At this time, trigger the energy balance weight adjustment and increase the weight α;
[0034] 3.2.2 Periodically monitor the link quality attenuation rate represents the average signal-to-noise ratio of candidate node i at the t-th monitoring; once it exceeds the threshold, it indicates that the link stability decreases. At this time, trigger the link reliability weight adjustment and increase the weight β;
[0035] 3.2.3 Periodically monitor the topology stability index is the average value of the distances from all candidate nodes i to the water surface Sink node; once it exceeds the threshold, it indicates that the node position drifts or the link topology is unstable. At this time, trigger the topology stability weight adjustment and increase the weight γ.
[0036] Preferably, the specific steps of step 5 are as follows:
[0037] The settings of the agent, action, state, and reward function are as follows:
[0038] Agent: Each underwater sensor node using the DAC-MAC protocol is an agent;
[0039] Action: The action of agent \(i\in\{1,2,\ldots,n\}\) at time slot \(t\) is defined as where \(m\in{- 1,0,\ldots,M - 1\}\) represents the sub - slot delay, specifically starting to send data at the \(m\delta\) sub - slot of this time slot (\(m=-1\) means not sending in this time slot); \(p\in\{P_0,P_1,P_2\}\) represents the transmit power selection, set to three levels: low, medium, and high; \(c\in\{C_0,C_1,C_2,C_3\}\) represents the candidate forwarding nodes, and the top 4 nodes ranked by the utility value are retained as the final candidate forwarding node set;
[0040] Local state: The local observation of agent \(i\in\{1,2,\ldots,n\}\) consists of three parts. The first part is the transmission indication When agent \(i\) selects the action at time slot \(t\), it will receive an ACK packet returned from the selected relay node \(c\) at time slot \(t + 2D\) i and get which represent successful transmission, collision, and channel idle respectively, and \(\mathbf{h}_{i,c}\) represents the information of the relay node \(c\) selected by agent \(i\), including the remaining energy, signal - to - noise ratio, and the distance to the Sink node, which is used to update the node utility value; The second part represents the number of time slots that agent \(i\) has occupied in total from the start of the time slot to time slot \(t-2D\) i and the number of time slots occupied by agents other than agent \(i\), specifically defined as
[0041]
[0042] The last part is the remaining energy of agent \(i\) Therefore, the action - observation pair of agent \(i\) at time slot \(t\) is denoted as
[0043]
[0044] where are respectively the normalized Specifically By concatenating the local observations, the local state can be obtained
[0045]
[0046] where \(M\) is the length of the historical state;
[0047] Global state: The global observation is defined as Similar to the local state, the global state at time slot \(t\) is \(s\)t = [z t-M+1 , z t-M+2 ,..., z t ;
[0048] Reward function: The reward function is divided into two parts: individual reward and global reward. The purpose of the global reward is to maximize the network throughput, select the optimal relay node and transmit power. The individual reward is used to evaluate the behavior of each agent, balance energy consumption and fairness. At time slot t, the global reward is defined as r t,tot = λ1r t,tot1 + (1 - λ1)r t,tot2
[0049]
[0050] The global reward consists of the relay reward r t,tot1 and the power reward r t,tot2 The relative importance of the two is controlled by the weight λ1 ∈ [0, 1]. By allocating the normalized utility function of the selected relay node c as the reward, the behavior of successful transmission and selecting relay nodes with high utility values is encouraged. By allocating negative rewards, the behaviors of collision and channel idleness are punished. Among them, the utility value U j is updated from the local state The power reward gives different positive rewards for successful transmission according to the transmit power level, encouraging reliable transmission with the minimum power and reducing energy waste and interference; The individual reward needs to consider two factors: fairness and energy consumption. By comparing whether the action of agent i is consistent with the optimal action, the fairness part of the individual reward is defined
[0051]
[0052] The energy consumption part of the individual reward can be directly obtained from the local observation, so that the two can be combined to obtain the complete individual reward function
[0053]
[0054] where λ2 is the weight coefficient used to adjust the trade-off between throughput fairness and energy consumption. Finally, the global reward and the individual reward are combined. The reward at time slot t is expressed as
[0055]
[0056] The beneficial effects of the present invention are as follows:
[0057] The present invention fully considers the uncertainty and dynamics of the underwater communication environment, and realizes information sharing between the network layer and the MAC layer through cross-layer collaborative design. At the network layer, by constructing a dynamic candidate forwarding node set and calculating a utility function that combines residual energy, link quality, and topological distance, the latest routing relay candidate node feedback is provided for the MAC layer. While the MAC layer, with the help of multi-agent reinforcement learning, dynamically adjusts sub-slot allocation, transmit power control, and relay node selection strategies under an asynchronous scheduling mechanism, so as to achieve the coordinated optimization of the overall network throughput and energy consumption efficiency while avoiding communication conflicts and resource waste. The present invention not only fully adapts to the propagation delay and channel fluctuation characteristics of the underwater network, but also provides an efficient and robust scheduling scheme for underwater multi-hop communication, which is of great significance for improving the overall performance of the underwater network. Description of the Drawings
[0058] Figure 1 It is a flowchart of the MAC protocol based on asynchronous cross-layer in the underwater acoustic sensor network;
[0059] Figure 2 It is a network model diagram of the underwater acoustic sensor network;
[0060] Figure 3 It is an asynchronous time slot model diagram. Detailed Embodiment
[0061] In order to enable those skilled in the art of this technology to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0062] As Figure 1 shown, a design method of the MAC protocol based on asynchronous cross-layer in the underwater acoustic sensor network includes the following steps:
[0063] Step 1: Establish a network model of the underwater acoustic sensor network;
[0064] As Figure 1 shown, the underwater acoustic sensor network includes several underwater sensor nodes, relay nodes, and a surface Sink node. The underwater sensor nodes respectively use TDMA, slotted ALOHA, and DAC-MAC protocols, and are randomly distributed in the underwater area. They collect the sensed data from the surrounding environment and forward it to the surface Sink node through the relay nodes. The surface Sink node broadcasts its own location to each node, and all the other nodes need to calculate and store the distance between themselves and the surface Sink node;
[0065] Step 2: Establish an asynchronous time slot model. As shown in Figure 2 , convert the distances between each underwater sensor node and relay node and the surface Sink node into propagation delays in units of sub-time slots δ;
[0066] Step 3: Construct a multi-dimensional utility function that combines three key indicators: the remaining energy of the fusion node, the link quality, and the topological distance gain, and adaptively update the weight coefficients of each indicator of the utility function by real-time monitoring of the network status;
[0067] Step 4: Generate a dynamic candidate forwarding node set by calculating and sorting the utility values of the relay nodes and combining with a threshold for screening;
[0068] Step 5: Deploy an underwater sensor node delay reward multi-agent reinforcement learning algorithm, set the action, state, and reward functions, and through an asynchronous cross-layer scheduling mechanism, use multi-agent reinforcement learning at the MAC layer to dynamically adjust the sub-time slot allocation, transmission power, and relay node selection strategy, realize information sharing and collaborative work between the network layer and the MAC layer, so as to meet the real-time requirements of the network and improve the overall performance.
[0069] Preferably, the specific steps of step 2 are as follows:
[0070] 2.1 In an asynchronous time slot system, assume that the packet lengths of different underwater sensor nodes and relay nodes are the same, and the returned ACK packets also have the same packet length;
[0071] 2.2 Each time slot length T slot corresponds to the duration T of packet transmission data plus the duration T of ACK packet transmission ack and the protection time t g , and further divide T slot into several smaller sub-time slots δ, satisfying Similarly, the propagation delay D i satisfies c is the underwater acoustic propagation speed, and d i is the distance between the underwater sensor node i and the surface Sink node;
[0072] 2.3 The transmission moment of the underwater sensor node within the time slot can be flexibly selected, and the transmission moment of each underwater sensor node is a specific delay point T r [n, i] (n represents the nth time slot, and i is the node number), satisfying:
[0073] T r [n, i] = m·δ, m ∈ {0, 1, 2,..., N - 1} (1)
[0074] The underwater sensor node autonomously selects the optimal transmission delay point T according to its own propagation delay to the receiving node (Sink or relay node) and the real-time network environment r [n,i] to maximize the risk of data collision avoidance with other nodes.
[0075] Preferably, the specific steps of step 3 are as follows:
[0076] 3.1 Design a multi-index utility function, the steps are as follows:
[0077] 3.1.1 Remaining energy ratio; The energy of underwater nodes is limited and difficult to replenish. Energy efficiency is the key to extending the network survival period. Prioritizing the selection of high-energy nodes to participate in forwarding can effectively balance the energy consumption of network nodes and extend the network life cycle; Therefore, the remaining energy ratio is used as one of the indicators to consider, and the definition is as follows:
[0078]
[0079] Among them, E res is the remaining energy of the node, and E init is the initial energy of the node;
[0080] 3.1.2 Link quality; The underwater acoustic channel is significantly affected by multipath effects, frequency-selective attenuation, and environmental noise. The dynamic fluctuation of link quality directly determines the success rate of data transmission. Therefore, the link quality LQ is used to represent the link quality between neighbor nodes and the current node:
[0081]
[0082] Among them, SNR max represents the maximum value of the signal-to-noise ratio among all neighbor nodes.
[0083] 3.1.3 Topological distance gain; Since the propagation speed of underwater acoustic signals is low and the propagation delay dominates the end-to-end delay, therefore, selecting a node closer to the Sink node can reduce the number of hops and cumulative delay. Define D gain as the distance advantage after the node selects a neighbor node for forwarding:
[0084]
[0085] Among them, d current represents the distance from the current node to the surface Sink node, and d neighbor represents the distance from the neighbor node to the surface Sink node.
[0086] 3.1.4 Combining the above three indicators, the utility function of candidate relay forwarding node j is designed as:
[0087]
[0088] Among them, α, β, and γ are weight coefficients, satisfying α + β + γ = 1; d(i, Sink) represents the distance from node i to the water surface Sink node; d(j, Sink); represents the distance from node j to the water surface Sink node; E j represents the remaining energy of node j; SNR ij represents the signal-to-noise ratio between node i and node j.
[0089] 3.2 Design an adaptive weight coefficient adjustment mechanism, and the steps are as follows:
[0090] 3.2.1 Periodically monitor the energy distribution index Among them, N is the total number of candidate nodes in the network, and E ratio,i is the remaining energy ratio of candidate node i, is the average value of the remaining energy ratios of all candidate nodes in the network; once it exceeds the threshold, it indicates that the energy distribution is unbalanced. At this time, trigger the energy balance weight adjustment and increase the weight α;
[0091] 3.2.2 Periodically monitor the link quality attenuation rate represents the average signal-to-noise ratio of candidate node i at the t-th monitoring; once it exceeds the threshold, it indicates that the link stability decreases. At this time, trigger the link reliability weight adjustment and increase the weight β;
[0092] 3.2.3 Periodically monitor the topology stability index is the average value of the distances from all candidate nodes i to the water surface Sink node; once it exceeds the threshold, it indicates that the node position drifts or the link topology is unstable. At this time, trigger the topology stability weight adjustment and increase the weight γ.
[0093] Preferably, the specific steps of step 5 are as follows:
[0094] The settings of the agent, action, state, and reward function are as follows:
[0095] Agent: Each underwater sensor node using the DAC-MAC protocol is an agent;
[0096] Action: The action of agent i ∈ {1, 2,..., n} at time slot t is defined as Among them, m ∈ {-1, 0,..., M - 1} represents the sub-slot delay, specifically starting to send data in the mδ sub-slot of this time slot (m = -1 means not sending in this time slot); p ∈ {P0, P1, P2} represents the transmission power selection, set to three levels: low, medium, and high; c ∈ {C0, C1, C2, C3} represents the candidate forwarding node, and the top 4 nodes are retained as the final candidate forwarding node set after sorting by the utility value;
[0097] Local state: The local observation of agent \(i\in\{1,2,\cdots,n\}\) consists of three parts. The first part is the transmission indication When agent \(i\) selects an action at time slot \(t\) it will receive an ACK packet returned by the selected relay node \(c\) at time slot \(t + 2D\) i and obtain which represent successful transmission, collision, and channel idle respectively. represents the information of the relay node \(c\) selected by agent \(i\), including remaining energy, signal-to-noise ratio, and distance to the Sink node, which is used to update the node utility value; The second part represents the number of time slots occupied by agent \(i\) in total from the start of the time slot to time slot \(t - 2D\) i and the number of time slots occupied by agents other than agent \(i\). The specific definitions are
[0098]
[0099] The last part is the remaining energy of agent \(i\) Therefore, the action-observation pair of agent \(i\) at time slot \(t\) is expressed as
[0100]
[0101] where are respectively the after normalization. By concatenating the local observations, the local state can be obtained
[0102]
[0103] where \(M\) is the length of the historical state;
[0104] Global state: The global observation is defined as Similar to the local state, the global state at time slot \(t\) is \(s\) t =\([z t-M+1 ,z t-M+2 ,\cdots,z t \);
[0105] Reward function: The reward function is divided into two parts: individual reward and global reward. The purpose of the global reward is to maximize the network throughput, select the optimal relay node and transmission power. The individual reward is used to evaluate the behavior of each agent, balance energy consumption and fairness. The global reward at time slot \(t\) is defined as \(r\) t,tot =\(\lambda_1r t,tot1 +(1 - \lambda_1)rt,tot2
[0106]
[0107] The global reward consists of the relay reward r t,tot1 and the power reward r t,tot2 and is composed of two parts. The relative importance of the two is controlled by the weight λ1 ∈ [0, 1]. By assigning the normalized utility function of the selected relay node c as the reward, the behavior of successful transmission and the selection of relay nodes with high utility values are encouraged. By assigning negative rewards, the behaviors of collision and channel idleness are punished. Among them, the utility value U j is updated from the local state The power reward gives different positive rewards for successful transmission according to the transmission power level, encouraging reliable transmission with the minimum power and reducing energy waste and interference; the individual reward needs to consider two factors: fairness and energy consumption. By comparing the action of agent i with the optimal action, the fairness part of the individual reward is defined
[0108]
[0109] The energy consumption part of the individual reward can be directly obtained from the local observation In this way, the two can be combined to obtain the complete individual reward function
[0110]
[0111] Among them, λ2 is the weight coefficient used to adjust the trade-off between throughput fairness and energy consumption. Finally, the global reward and the individual reward are combined, and the reward at time slot t is expressed as
[0112]
[0113] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for designing an asynchronous cross-layer MAC protocol in an underwater acoustic sensor network, characterized in that: The following steps are involved: Step 1: Establish a network model of the underwater acoustic sensor network; the underwater acoustic sensor network includes several underwater sensor nodes, relay nodes and surface sink nodes. The underwater sensor nodes use TDMA, slotted ALOHA and DAC-MAC protocols respectively and are randomly distributed in the underwater area. They collect sensed data from the surrounding environment and forward it to the surface sink node through the relay node. The surface sink node broadcasts its position to each node, and all other nodes need to calculate and store the distance between themselves and the surface sink node. Step 2: Establish an asynchronous time slot model and convert the distance between each underwater sensor node and relay node and the surface sink node into a propagation delay in units of sub-time slot δ; Step 3: Construct a multidimensional utility function based on three key indicators: residual energy of the fusion node, link quality, and topological distance gain. Adaptively update the weight coefficients of each indicator in the utility function by monitoring the network status in real time. Step 4: Generate a dynamic candidate forwarding node set by calculating and ranking the relay node utility values and combining threshold screening; Step 5: Deploy the delayed reward multi-agent reinforcement learning algorithm on the underwater sensor nodes, set the action, state, and reward functions, and use the asynchronous cross-layer scheduling mechanism to dynamically adjust the sub-slot allocation, transmission power, and relay node selection strategy at the MAC layer using multi-agent reinforcement learning to achieve information sharing and collaborative work between the network layer and the MAC layer, thereby meeting the real-time requirements of the network and improving the overall performance.
2. The method for designing a MAC protocol based on asynchronous cross-layer in an underwater acoustic sensor network according to claim 1, wherein The specific steps of step 2 are as follows: 2.1 In the asynchronous time slot system, it is assumed that the data packets of different underwater sensor nodes and relay nodes have the same length, and the returned ACK packets also have the same packet length; 2.2 Length of each time slot T slot Corresponds to the duration T of the data packet transmission data Plus the duration T of ACK packet transmission ack and protection time t g , and further T slot Subdivided into several smaller sub-time slots δ, satisfying T slot =N·δ, Likewise, the propagation delay D i satisfy c is the speed of underwater sound, d i is the distance between the underwater sensor node i and the surface Sink node; 2.3 The transmission time of the underwater sensor node can be flexibly selected within a time slot, and the transmission time of each underwater sensor node is a specific delay point T within the time slot. r [n, i], where n represents the nth time slot and i is the node number, satisfying: T r [n,i] = m·δ, m ∈ {0, 1, 2, ..., N - 1} (1) The underwater sensor node autonomously selects the optimal transmission delay point T according to its own propagation delay to the receiving node and the real-time network environment r [n,i], where the receiving node is a Sink or relay node, to maximize the risk avoidance of data collision with other nodes.
3. The method for designing a MAC protocol based on asynchronous cross-layer in an underwater acoustic sensor network according to claim 1, characterized in that The specific steps of step 3 are as follows: 3.1 Design a multi-index utility function. The steps are as follows: 3.1.1 Residual Energy Ratio: Underwater nodes have limited energy and are difficult to replenish. Energy efficiency is key to extending the network lifecycle. Prioritizing high-energy nodes for forwarding can effectively balance the energy consumption of network nodes and extend the network lifecycle. Therefore, the residual energy ratio is considered as one of the indicators and is defined as follows: in, E res is the remaining energy of the node, E init is the initial energy of the node; 3.1.2 Link Quality: The underwater acoustic channel is significantly affected by multipath effects, frequency selective fading, and environmental noise. The dynamic fluctuation of link quality directly determines the success rate of data transmission. Therefore, the link quality LQ is used to represent the link quality between the neighboring node and the current node: where, SNR max represents the maximum value of the signal-to-noise ratio among all neighbor nodes; SNR current represents the signal-to-noise ratio of the current node; 3.1.3 Topological distance gain: Since underwater acoustic signals have a low propagation speed and propagation delay dominates the end-to-end delay, selecting a node closer to the sink node can reduce the number of hops and cumulative delay. Define D gain The distance advantage after forwarding by neighboring nodes is selected for the node: Among them, d current Indicates the distance from the current node to the water surface Sink node, d neighbor Indicates the distance from the neighbor node to the water surface sink node; 3.1.4 Based on the three indicators described in 3.1.1-3.1.3, the utility function of candidate relay forwarding node j is designed as follows: Among them, α, β, and γ are weight coefficients, satisfying α + β + γ = 1; d(i, Sink) represents the distance from node i to the water surface Sink node; d(j, Sink); represents the distance from node j to the water surface Sink node; E j represents the remaining energy of node j; SNR ij represents the signal-to-noise ratio between node i and node j. 3.2 Design an adaptive weight coefficient adjustment mechanism. The steps are as follows: 3.2.1 Periodic monitoring of energy distribution index Where N is the total number of candidate nodes in the network, E ratio,i is the residual energy ratio of candidate node i, is the average of the remaining energy ratios of all candidate nodes in the network. Once the threshold is exceeded, it indicates that the energy distribution is unbalanced, and the energy balance weight adjustment is triggered, increasing the weight α. 3.2.2 Periodic Monitoring of Link Quality Decay Rate denotes the average signal-to-noise ratio of candidate node i at the t-th monitoring; once it exceeds the threshold, it indicates a decrease in link stability, and at this time, the link reliability weight adjustment is triggered to increase the weight β; 3.2.3 Periodic monitoring of topological stability index is the average distance from all candidate nodes i to the water surface sink node; once it exceeds the threshold, it indicates that the node position drifts or the link topology is unstable, and the topology stability weight adjustment is triggered to increase the weight γ.
4. The method for designing an asynchronous cross-layer MAC protocol in an underwater acoustic sensor network according to claim 1, wherein: The specific steps of step 5 are as follows: The settings of the agent, action, state, and reward function are as follows: Agent: Each underwater sensor node using the DAC-MAC protocol is an agent; Action: The action of agent i∈{1,2,...,n} in time slot t is defined as Where m∈{-1,0,...,M-1} represents the sub-slot delay, specifically starting to send data at the mδ sub-slot of the time slot (m=-1 means no transmission in the time slot); p∈{P0,P1,P2} represents the transmit power selection, which is set to low, medium, and high; c∈{C0,C1,C2,C3} represents the candidate forwarding nodes. After sorting by utility value, the top four nodes are retained as the final candidate forwarding node set; Local state: the local observation of agent \(i\in\{1,2,\cdots,n\}\) consists of three parts. The first part is the transmission indication When agent \(i\) selects an action at time slot \(t\) it will receive an ACK packet returned from the selected relay node \(c\) at time slot \(t + 2D\) i and obtain which respectively represent successful transmission, collision, and channel idle represents the information of the relay node \(c\) selected by agent \(i\), including the remaining energy, signal-to-noise ratio, and the distance to the Sink node, which are used to update the node utility value; the second part represents the number of time slots occupied by agent \(i\) from the start of the time slot to time slot \(t - 2D\) i and the number of time slots occupied by agents other than agent \(i\). The specific definitions are The last part is the remaining energy of agent i Therefore, the action-observation pair of agent i at time slot t is Expressed as Among them, are respectively after normalization Specifically By concatenating local observations, the local state can be obtained Where M is the length of the historical state; Global state: The global observation is defined as Similar to the local state, the global state at time slot t is s t = [z t-M+1 , z t-M+2 ,..., z t ; Reward function: The reward function is divided into two parts: individual reward and global reward. The purpose of the global reward is to maximize the network throughput, select the optimal relay node and transmit power, and the individual reward is used to evaluate the behavior of each agent, balancing energy consumption and fairness. The global reward at time slot t is defined as r t,tot = λ1r t,tot1 +(1 - λ1)r t,tot2 The global reward consists of the relay reward r t,tot1 and the power reward r t,tot2 The two parts are composed of the weight λ1 ∈ [0, 1] to control the relative importance of the two. By allocating the normalized utility function of the selected relay node c as the reward, it encourages the behavior of successful transmission and selecting relay nodes with high utility values. By allocating negative rewards to punish the behaviors of collisions and channel idleness, where the utility value U j is updated from the local state The power reward gives different positive rewards for successful transmission according to the transmission power level, encouraging reliable transmission with the minimum power and reducing energy waste and interference; the individual reward needs to consider two factors: fairness and energy consumption. By comparing the action of agent i to define the fairness part of the individual reward The energy consumption part of the individual reward can be directly obtained from local observations, so that the two can be combined to obtain a complete individual reward function Among them, λ2 is the weight coefficient used to adjust the trade-off between throughput fairness and energy consumption. Finally, the global reward and individual reward are combined together, and the reward at time slot t is expressed as