An Underwater Acoustic Dual-Hop Network Medium Access Control Method Based on Q-Learning
By applying the Q learning algorithm in the water acoustic double-hop network, selecting the parent node and child nodes, and optimizing the media access control, the problems of long channel idle time and weak anti-interference ability are solved, and a water acoustic network with a larger coverage range and high throughput are achieved.
Patent Information
- Application Number
- CN202211186500.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-09-27
AI Technical Summary
The channel in the water-sound dual-hop network has a long idle time, many transmission times and weak anti-interference ability. The existing technology has failed to effectively apply the Q-learning algorithm to optimize the network multiple access order.
In the water-acoustic double-hop network, the Q-learning algorithm is used to select the parent node and the child node, and the time slot allocation and reward mechanism are used to optimize the media access control, reduce unnecessary information exchange, and improve network robustness and coverage.
The coverage of the hydroacoustic network has been expanded, network throughput and robustness has been improved, network complexity has been reduced, adapted to dynamic changes in the marine environment, and reduced energy consumption in information exchange.
Smart Images

Figure CN115843110B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an underwater acoustic network, and in particular to a medium access control method for an underwater acoustic two-hop network based on Q-learning. Background Art
[0002] An underwater acoustic sensor network is an essential and important part of an undersea observatory network; an underwater acoustic sensor network needs to select a central node for data collection and aggregation, and this central node is called a sink. If other nodes in the network can communicate directly with the sink, it is called an underwater acoustic single-hop network. The underwater acoustic medium access control (MAC) protocol is a protocol that determines how nodes share the underwater acoustic channel, and a suitable MAC protocol can achieve conflict-free data collection at the sink. In an underwater acoustic single-hop network, since nodes can communicate directly with the sink, the coordination problem between the transceiver parties is greatly simplified, and most current research on underwater acoustic MAC protocols focuses on underwater acoustic single-hop networks. On the other hand, those nodes that cannot communicate directly with the sink need to forward data to reach the sink. If an underwater acoustic sensor network allows the number of node data forwarding times to be no more than once, it is called an underwater acoustic two-hop network. Research shows that the transmission range of two-hop transmission is more than 1.5 times that of single-hop transmission (Morozs N, et al. Dual-Hop TDA-MAC and Routing for Underwater Acoustic Sensor Networks[J]. IEEE Journal of Oceanic Engineering, 2019, 44(4): 865-880.), and the working range of an underwater acoustic two-hop network is larger than that of an underwater acoustic single-hop network. However, in an underwater acoustic two-hop network, the sink needs to communicate with two-hop nodes through relays, so it is impossible to coordinate the data transmission behaviors of each node in the network by broadcasting signals like an underwater acoustic single-hop network, which increases the design difficulty of the MAC protocol, and at the same time there are problems such as an increase in network idle time, an increase in the number of transmissions, and a weakening of anti-interference ability caused by the inherent characteristics of the underwater acoustic channel.
[0003] Some scholars have tried to apply the Q-learning algorithm in an underwater acoustic single-hop network to optimize the multi-access order problem of the underwater acoustic network (Chen Yougan, Huang Weidi, et al. Medium Access Control Method for Underwater Acoustic Network with Variable Number of Nodes Based on Q-Learning, Chinese Invention Patent No. ZL 202110791390.9), and this method can effectively reduce the information exchange between the transceiver parties and improve the anti-interference ability of the single-hop underwater acoustic network. However, currently, no scholars have applied the Q-learning algorithm to an underwater acoustic two-hop network to study and optimize the MAC protocol problem of the underwater acoustic network. Summary of the Invention
[0004] The object of the present invention is to provide a Q - learning algorithm applicable to an underwater acoustic two - hop network to solve the problems that the slow sound propagation speed leads to a large idle time of the underwater acoustic channel and there is an urgent need to optimize the network multiple access order. This method utilizes the learning ability of the Q - learning algorithm to adapt to the changes of the underwater acoustic channel, reduces unnecessary information exchange in the network, improves the overall network lifetime, and has the advantages of large coverage area, strong robustness, and low complexity.
[0005] The present invention includes the following steps:
[0006] 1) Parameter initialization:
[0007] Consider an underwater acoustic two - hop network, which includes M sensor nodes (hereinafter referred to as "nodes") and 1 sink. The nodes sense information from the ocean environment, and the sink is responsible for collecting the acoustic data sensed by the nodes. Among them, there are A + B nodes within the communication range of the sink and can communicate directly with the sink (hereinafter referred to as "direct communication nodes"); while the other C nodes are outside the communication range of the sink, and there is at least one direct communication node within their communication range. These nodes cannot communicate directly with the sink but can communicate indirectly with the sink through the direct communication nodes (hereinafter referred to as "indirect communication nodes"). The underwater acoustic two - hop network is composed of A + B direct communication nodes and C indirect communication nodes in total (M = A + B + C).
[0008] First, among the A + B direct communication nodes, select A nodes to be responsible for forwarding the information between the indirect communication nodes and the sink, and assist the indirect communication nodes to communicate indirectly with the sink, which are called "parent nodes" F i (i = 1, 2, …, A). The communication ranges of these A parent nodes cover all indirect communication nodes, and the assisted objects do not overlap. That is, any indirect communication node will only be assisted in communication by one parent node F i This indirect communication node is called the child node S i of the parent node F (i,j) (j = 1, 2, …). Among the direct communication nodes, the other B nodes that are not parent nodes are simply referred to as "single - hop nodes" E i (i = 1, 2, …, B).
[0009] Let the time length of a time slot be equal to the data length plus the guard interval. Assume that the data formats and lengths of each node are the same, so the time lengths of the time slots are also the same. Suppose the sink divides the data collection process into N time slots, and receives one data in each time slot. Therefore, in one data collection process, the sink can receive at most N data packets. To ensure that all data of all nodes in the network can be fully received, the number of time slots N can be set equal to the total number of nodes M in the underwater acoustic two-hop network. In Q-learning, the Q matrix applied to medium access control is an M×N matrix. The row m (m = 1, 2, …, M) of the Q matrix represents the node number, and the column n (n = 1, 2, …, N) of the Q matrix represents the time slot number. Therefore, Q(m,n) represents the Q value corresponding to the action that node m selects the nth time slot to send data; the larger the Q value, the greater the priority for node m to select the nth time slot to send data. That is, node m will select the time slot with the highest Q value in the mth row of the Q matrix to send data; if there are multiple identical highest Q values in the mth row, a time slot will be randomly selected from the multiple time slots with the highest Q value in the mth row to send data. To reduce the node operation complexity, each node only needs to store the sub-matrix of the row representing its selected transmission time slot inside, that is, node m only needs to store a sub-matrix Q of size 1×N m , so the Q matrix can be expressed as Q = [Q1; Q2; …; Q m ; …; Q M-1 ; Q M .
[0010] The child nodes cannot directly communicate with the sink. The child nodes need to send the data to the parent node, and the parent node then forwards it to the sink. Suppose a parent node needs to assist D i child nodes to forward data to the sink. In one round of data collection process, the parent node needs to send (1 + D i ) times of data, including 1 time of the parent node's own data and D i times of data sent by the child nodes to it, that is, the sending times of the parent node are (1 + D i ). Therefore, the parent node will select the first (1 + D m ) time slots with the largest Q values in the sub-matrix Q i to send data. While the single-hop nodes and child nodes only need to send 1 piece of their own data in one data collection process.
[0011] Initialize the iteration number k = 0, the maximum iteration number is K, and the initial Q value table is a zero matrix of M×N.
[0012] 2) The control signal is a signal used to coordinate the behaviors of each node in the underwater acoustic network, including the start signal and the feedback signal. Assume that the formats and lengths of the control signals are the same, so the time lengths of the control time slots are also the same; the time length of the control time slot is equal to 2 times the maximum transmission delay plus the control signal duration.
[0013] When data collection starts, the destination broadcasts a start signal to all direct communication nodes in the first control time slot, and then the parent node forwards the start signal to its child nodes in the second control time slot. When node m receives the start signal, it checks the sub-matrix Q according to its own transmission times. m , selects the transmission time slot determined by the sub-matrix Q m , and waits until that time slot to send data.
[0014] 3) Feedback signal design:
[0015] When data collection starts, the destination records the reception situation of each time slot. If the destination successfully receives complete data in a certain time slot, it marks it as a successful time slot and records the time slot number; if the destination fails to successfully receive complete data in a certain time slot (including three cases: data collision in that time slot, unable to successfully receive all data due to poor channel state, no node sending data / idle state of the time slot), the time slot is not marked. After the transmission ends, the destination broadcasts a feedback signal to all direct communication nodes; this feedback signal contains the total number of transmission time slots N and the information of successful time slots. If the time slot selected by node m is not a successful time slot, it means that node m's transmission fails.
[0016] Similarly, when data collection starts, the parent node also records the reception situation of each time slot. If the reception time slot of the parent node is occupied in a certain time slot (including two cases: the parent node selects to transmit data in that time slot, and the data of other single-hop nodes arrives at the parent node in that time slot), the parent node records that time slot as an occupied time slot, and the child node is not allowed to transmit data in the occupied time slot. Only time slots that do not belong to the occupied time slots can be judged as successful time slots or unmarked time slots. In other unoccupied time slots, if the parent node successfully receives complete data in a certain time slot, it is marked as a successful time slot; if the parent node fails to successfully receive complete data in a certain time slot (including three cases: data collision in that time slot, unable to successfully receive all data due to poor channel state, no node sending data / idle state of the time slot), the time slot is not marked. After the transmission ends, the parent node combines the total number of transmission time slots in the feedback signal from the destination and its own recorded reception situation, and broadcasts a feedback signal to the child nodes it is responsible for. This feedback information contains the total number of transmission time slots N, the information of occupied time slots, the information of successful time slots, and the information of unsuccessful time slots.
[0017] 4) Reward mechanism design:
[0018] When a node receives the feedback signal from the destination or the parent node, node m will, according to the action of sending data in the nth time slot it selects, combine the information of unsuccessful time slots in the feedback information, and for the mth row of the Q matrix (that is, the sub-matrix Q of the Q matrix stored internally by node m m ), the mth row R in the reward matrixm (m, :) obtains different values. R m (m, n) represents the reward value obtained by combining the feedback signal after the action that node m selects the nth time slot to send data.
[0019] If node m is a direct communication node, the reward value R(m, n) is set as follows:
[0020] ① If node m sends data in time slot n and time slot n is a successful time slot, it means that the data sent by node m in time slot n is successfully received by the destination. R(m, n) in the reward matrix is a positive value +|ψ| to ensure that the Q(m, n) value increases.
[0021] ② If node m sends data in time slot n and time slot n is not a successful time slot, it means that the data sent by node m in the nth time slot is not successfully received by the destination. R(m, n) in the reward matrix is 0 to ensure that the Q(m, n) value approaches 0.
[0022] ③ If node m does not send data in the nth time slot and time slot n is a successful time slot, it means that the data sent by other nodes in the nth time slot is successfully received by the destination. R(m, n) in the reward matrix is a negative value -|ψ| to ensure that the Q(m, n) value decreases.
[0023] ④ If node m does not send data in the nth time slot and time slot n is not a successful time slot. R(m, n) in the reward matrix is 0 to ensure that the Q(m, n) value approaches 0.
[0024] If node m is a child node, the reward value R(m, n) is set as follows:
[0025] ① If the parent node is occupied in time slot n, R(m, n) in the reward matrix is a negative value -|ψ| to ensure that the Q(m, n) value decreases. Only the time slots that do not belong to the occupied time slots can be judged as successful time slots or unmarked time slots.
[0026] ② If node m sends data in time slot n and time slot n is a successful time slot, it means that the data sent by node m in time slot n is successfully received by the parent node. R(m, n) in the reward matrix is a positive value +|ψ| to ensure that the Q(m, n) value increases.
[0027] ③ If node m sends data in time slot n and time slot n is not a successful time slot, it means that the data sent by node m in the nth time slot is not successfully received by the parent node. R(m, n) in the reward matrix is 0 to ensure that the Q(m, n) value approaches 0.
[0028] ④ If node m does not send data in time slot n, and time slot n is a successful time slot, it means that the data sent by other child nodes in the nth time slot has been successfully received by the parent node. The R(m,n) in the reward matrix is a negative value -|ψ| to ensure that the Q(m,n) value decreases.
[0029] ⑤ If node m does not send data in the nth time slot, and time slot n is not a successful time slot. The R(m,n) in the reward matrix is 0 to ensure that the Q(m,n) value approaches 0.
[0030] 5) Update the Q value table according to the Q-learning formula Q(m,n)←(1-γ)·Q(m,n)+γ·R(m,n), where γ is the learning rate, with a value in the range (0,1], and the reward matrix R is a matrix with the same size as the Q matrix.
[0031] 6) After the Q value is updated, k = k + 1. If the maximum number of iterations K is reached or the Q value table no longer changes, it reaches a stable state; otherwise, repeat steps 2) to 5).
[0032] 7) According to the final Q value table obtained by iteration, let node m select the sub-matrix Q m and send data to the sink or the parent node at the time slot n corresponding to the largest Q value in it. The parent node F i selects the largest (1 + D Fi ) Q values in the sub-matrix Q according to its own transmission times i and sends data at the corresponding time slots. The task of allocating the right to use the underwater acoustic channel medium resources is completed.
[0033] The present invention has the following outstanding advantages:
[0034] 1) Compared with the traditional underwater acoustic single-hop network, the underwater acoustic two-hop network has a larger coverage area, and the sink can collect underwater acoustic data from a larger range and more nodes. In the actual ocean environment, the number of sink deployments can be effectively reduced, and the economic cost can be lowered.
[0035] 2) It is applicable to multiple scenarios where the ocean environment is complex and the underwater acoustic channel is unstable, and thus the underwater acoustic link fails and recovers intermittently. When the direct communication node cannot communicate directly with the sink after the underwater acoustic link fails, Q learning can be used to convert the node identity into a child node and request the parent node to assist in transmission for indirect communication; and after the underwater acoustic link is restored, the indirect communication nodes within the direct communication range can also use Q learning to restore the direct communication node identity, improving the robustness of the underwater acoustic network.
[0036] 3) The child nodes do not need to transfer data to the parent nodes during a separate time period. Instead, they utilize the idle time of the parent nodes during the transmission process to transfer data without adding extra time. This effectively overcomes the drawback that the sink cannot work during the waiting time when the child nodes transfer data to the parent nodes in the transmission process of child node - parent node - sink. At the same time, the Q-learning algorithm is used to solve the problem of coordinating the sink and child nodes in indirect communication, which requires exchanging information to obtain stable transmission and incurs huge time and energy consumption. The proposed method effectively improves the throughput of the underwater acoustic two-hop network and reduces the energy consumption of information exchange. Brief Description of the Drawings
[0037] Figure 1 It is a schematic diagram of an underwater acoustic two-hop network. The figure includes 1 sink, 2 parent nodes, 2 single-hop nodes, and 3 child nodes.
[0038] Figure 2 It is an example of the Q-value change of the underwater acoustic two-hop network medium access control method based on Q-learning of the present invention.
[0039] Figure 3 It is a flowchart of the node transmission process of the underwater acoustic two-hop network medium access control method based on Q-learning of the present invention.
[0040] Figure 4 It is a schematic diagram of the transmission process of the underwater acoustic two-hop network medium access control method based on Q-learning of the present invention.
[0041] Figure 5 It is the changing trend of the network throughput of the underwater acoustic two-hop network medium access control method based on Q-learning of the present invention.
[0042] Figure 6 It is the changing trend of the average transmission energy consumption of the underwater acoustic two-hop network medium access control method based on Q-learning of the present invention.
[0043] Figure 7 It is the changing trend of the packet delivery rate of the underwater acoustic two-hop network medium access control method based on Q-learning of the present invention. Detailed Embodiment
[0044] The following embodiments will describe the present invention in detail with reference to the accompanying drawings.
[0045] In an underwater acoustic two-hop network, according to whether a sensor node can communicate with the sink, it is divided into direct communication nodes that can communicate and indirect communication nodes that require the assistance of other nodes. Multiple nodes are selected from the direct communication nodes as relays connecting the indirect communication nodes and the sink, called parent nodes. The indirect nodes assisted by the parent nodes are called the children nodes of the parent node. The data transmission process of the sink collecting the data sensed by the underwater acoustic sensor nodes is divided into several time slots. The direct communication nodes use the Q-learning algorithm, combined with the feedback signal of the sink, and by reasonably setting the reward mechanism, select appropriate time slots to send data to avoid data collection conflicts at the sink. The children nodes use the Q-learning algorithm to transfer data to the parent nodes during the idle time of the data collection process, without affecting the data collection of the sink, effectively overcoming the disadvantage that the sink cannot work during the waiting time when the children nodes transfer data to the parent nodes in the transmission process of children node - parent node - sink. It has a large network range, high throughput, and strong robustness, can expand the working range of the underwater acoustic network, and effectively improve the network performance.
[0046] The embodiments of the present invention include the following steps:
[0047] 1) Parameter initialization:
[0048] As shown in the schematic diagram of the underwater acoustic two-hop network Figure 1 , consider an underwater acoustic two-hop network, which includes M = 20 sensor nodes (hereinafter referred to as "nodes") and 1 sink. The nodes sense information from the ocean environment, and the sink is responsible for collecting the acoustic data sensed by the nodes. Among them, 11 nodes are within the communication range of the sink and can communicate directly with the sink (hereinafter referred to as "direct communication nodes"); while the other 9 nodes are outside the communication range of the sink, and there is at least one direct communication node within the communication range. These nodes cannot communicate directly with the sink, but can communicate indirectly with the sink through the forwarding of the direct communication nodes (hereinafter referred to as "indirect communication nodes"). The underwater acoustic two-hop network consists of 11 direct communication nodes and 9 indirect communication nodes in total.
[0049] First, among the 11 direct communication nodes, 3 nodes are selected to be responsible for forwarding the information of the indirect communication nodes and the sink, and assist the indirect communication nodes to communicate indirectly with the sink, called "parent nodes" F i (i = 1, 2, 3). The communication ranges of these 3 parent nodes cover all indirect communication nodes, and the assisted objects do not overlap. That is, any indirect communication node will only be assisted in communication by one parent node F i . Suppose each parent node is responsible for assisting 3 indirect communication nodes to send data, and this indirect communication node is called the child node S i of the parent node F (i,j) (j = 1, 2, 3). Among the direct communication nodes, the other 8 nodes that are not parent nodes are simply referred to as "single-hop nodes" E i (i = 1, 2,..., 8).
[0050] Let the time length of a time slot be equal to the data length plus the guard interval. Assume that the data formats and lengths of each node are the same, so the time lengths of the time slots are also the same. Suppose the sink divides the data collection process into N = 20 time slots, and receives one piece of data in each time slot. Therefore, in one data collection process, the sink can receive at most 20 data packets. To ensure that all data of all nodes in the network can be completely received, the number of time slots N can be set equal to the total number M of nodes in the underwater acoustic two-hop network. In Q learning, the Q matrix applied to medium access control is a matrix of size M×N, that is, a matrix of size 20×20. The row m (m = 1, 2, …, 20) of the Q matrix represents the node number, and the column n (n = 1, 2, …, 20) of the Q matrix represents the time slot number. Therefore, Q(m, n) represents the Q value corresponding to the action that node m selects the nth time slot to send data; the larger the Q value, the greater the priority for node m to select the nth time slot to send data. That is, node m will select the time slot with the highest Q value in the mth row of the Q matrix to send data; if there are multiple identical highest Q values in the mth row, a time slot will be randomly selected from the multiple time slots with the highest Q value in the mth row to send data. To reduce the computational complexity of the node, each node only needs to store the sub-matrix of the row representing its selected transmission time slot inside, that is, node m only needs to store a sub-matrix of size 1×20, Q m , so the Q matrix can be expressed as Q = [Q1; Q2; …; Q m ; …; Q 19 ; Q 20 .
[0051] The child nodes cannot directly communicate with the sink. The child nodes need to send the data to the parent node, and the parent node then forwards it to the sink. Suppose a parent node needs to assist 3 child nodes in forwarding data to the sink. In one round of data collection process, the parent node needs to send data (1 + 3) times, including 1 time for its own data and 3 times for the data sent by the child nodes to it, that is, the sending times of the parent node are 4. Therefore, the parent node will select the first 4 time slots with the largest Q values in the sub-matrix Q m to send data. While the single-hop nodes and child nodes only need to send 1 piece of their own data in one data collection process.
[0052] Initialize the iteration number k = 0, the maximum iteration number is K = 20, and the initial Q value table is a 20×20 zero matrix.
[0053] 2) The control signal is a signal used to coordinate the behaviors of each node in the underwater acoustic network, including the start signal and the feedback signal. Assume that the control signals have the same format and length, so the time lengths of the control time slots are also the same; the time length of the control time slot is equal to 2 times the maximum transmission delay plus the control signal duration.
[0054] The node transmission process of the entire underwater acoustic double-hop network medium access control method based on Q-learning is as follows Figure 3 as shown.
[0055] When data collection starts, the destination broadcasts a start signal to all direct communication nodes in the first control time slot, and then the parent node forwards the start signal to the child nodes in the second control time slot. When node m receives the start signal, it checks the sub-matrix Q according to its own transmission times m and selects the transmission time slot determined by the sub-matrix Q m and can send data when it reaches that time slot.
[0056] 3) Feedback signal design:
[0057] When data collection starts, the destination records the reception situation of each time slot. If the destination successfully receives complete data in a certain time slot, it is marked as a successful time slot and the time slot number is recorded; if the destination does not successfully receive complete data in a certain time slot (including three situations: data collision in this time slot, unable to successfully receive all data due to poor channel state, no node sending data / idle state of the time slot), then this time slot is not marked. After the transmission ends, the destination broadcasts a feedback signal to all direct communication nodes; this feedback signal includes the total number of transmission time slots N = 20 and the information of successful time slots. If the time slot selected by node m is not a successful time slot, it means that node m's transmission fails.
[0058] In this embodiment, as Figure 2 shown, it is assumed that the destination receives successfully in all time slots except the 10th and 15th time slots.
[0059] Similarly, when data collection starts, the parent node also records the reception situation of each time slot. If the reception time slot of the parent node is occupied in a certain time slot (including two situations: the parent node selects to transmit data in this time slot, and the data of other single-hop nodes arrives at the parent node in this time slot), the parent node will record this time slot as an occupied time slot, and the child node is not allowed to transmit data in the occupied time slot. Only time slots that do not belong to the occupied time slots can be judged as successful time slots or unmarked time slots. In other unoccupied time slots, if the parent node successfully receives complete data in a certain time slot, it is marked as a successful time slot; if the parent node does not successfully receive complete data in a certain time slot (including three situations: data collision in this time slot, unable to successfully receive all data due to poor channel state, no node sending data / idle state of the time slot), then the time slot is not marked. After the transmission ends, the parent node combines the total number of transmission time slots of the feedback signal from the destination and its own recorded reception situation, and broadcasts a feedback signal to the child nodes it is responsible for. This feedback information includes the total number of transmission time slots N, the information of occupied time slots, the information of successful time slots, and the information of unsuccessful time slots.
[0060] In this embodiment, asFigure 2 As shown, the parent node selects to send data in time slots 4, 5, 9, and 10, while the data of other single-hop nodes arrives at the parent node in time slots 19 and 20. Therefore, the parent node marks time slots 4, 5, 9, 10, 19, and 20 as occupied time slots. In other unoccupied time slots, the parent node successfully receives data from child nodes in time slots 14 and 15.
[0061] 4) Design of reward mechanism:
[0062] After the node receives the feedback signal from the sink or the parent node, node m will, according to the action of sending data in the nth time slot it selects, combined with the information of the unsuccessful time slots in the feedback information, for the mth row of the Q matrix (i.e., the sub-matrix Q of the Q matrix stored internally in node m m ), different values are obtained for the mth row R m (m,:) in the reward matrix. R m (m,n) represents the reward value obtained by node m after the action of sending data in the nth time slot in combination with the feedback signal.
[0063] If node m is a direct communication node, the reward value R(m,n) is set as follows:
[0064] ① If node m sends data in time slot n and time slot n is a successful time slot, it means that the data sent by node m in time slot n is successfully received by the sink, and R(m,n) in the reward matrix is a positive value +|ψ| to ensure that the value of Q(m,n) increases.
[0065] ② If node m sends data in time slot n and time slot n is not a successful time slot, it means that the data sent by node m in the nth time slot is not successfully received by the sink, and R(m,n) in the reward matrix is 0 to ensure that the value of Q(m,n) approaches 0.
[0066] ③ If node m does not send data in the nth time slot and time slot n is a successful time slot, it means that the data sent by other nodes in the nth time slot is successfully received by the sink, and R(m,n) in the reward matrix is a negative value -|ψ| to ensure that the value of Q(m,n) decreases.
[0067] ④ If node m does not send data in the nth time slot and time slot n is not a successful time slot. R(m,n) in the reward matrix is 0 to ensure that the value of Q(m,n) approaches 0.
[0068] In this embodiment, it is assumed that the Q value ranges between -5 and +5, then Ψ = 5. The parent node F1 sends data in time slots 4, 5, 9, and 10, and the single-hop node E1 sends data in time slot 10. The data reception in time slot 10 fails due to data conflict, and there is no data transmission in time slot 15, so the data reception fails. The other time slots are successful reception time slots. The data reception in time slots 4, 5, and 9 selected by the parent node F1 is successful, so the values of R(F1, 4), R(F1, 5), and R(F1, 9) in the reward matrix are +5; the data reception in time slot 10 selected by the parent node F1 and the single-hop node E1 fails, so the values of R(F1, 10) and R(E1, 10) in the reward matrix are 0; the single-hop node E1 sends data in time slot 10 and the data reception in time slot 9 is successful, so the value of R(E1, 9) in the reward matrix is -5. The parent node F1 does not send data in time slot 15 and the data reception in time slot 15 fails, so the value of R(F1, 15) in the reward matrix is 0. The sub-matrices Q F1 and Q E1 change as shown in Figure 2 the following figure.
[0069] If node m is a child node, the reward value R(m, n) is set as follows:
[0070] ① If the time slot n of the parent node is already occupied, R(m, n) in the reward matrix is a negative value -|ψ| to ensure that the Q(m, n) value decreases. Only the time slots that do not belong to the occupied time slots can be judged as successful time slots or unmarked time slots.
[0071] ② If node m sends data in time slot n and time slot n is a successful time slot, it means that the data sent by node m in time slot n is successfully received by the parent node. R(m, n) in the reward matrix is a positive value +|ψ| to ensure that the Q(m, n) value increases.
[0072] ③ If node m sends data in time slot n and time slot n is not a successful time slot, it means that the data sent by node m in the nth time slot is not successfully received by the parent node. R(m, n) in the reward matrix is 0 to ensure that the Q(m, n) value approaches 0.
[0073] ④ If node m does not send data in time slot n and time slot n is a successful time slot, it means that the data sent by other child nodes in the nth time slot is successfully received by the parent node. R(m, n) in the reward matrix is a negative value -|ψ| to ensure that the Q(m, n) value decreases.
[0074] ⑤ If node m does not send data in the nth time slot and time slot n is not a successful time slot. R(m, n) in the reward matrix is 0 to ensure that the Q(m, n) value approaches 0.
[0075] In this embodiment, the parent node is occupied in time slots 4, 5, 9, 10, 19, and 20. Therefore, the values of R(S (1,1) , 4), R(S (1,1) , 5), R(S (1,1) , 9), R(S (1,1) , 10), R(S (1,1) , 19), and R(S (1,1) , 20) in the reward matrix are -5. The child node S (1,1) successfully receives data in the selected time slot 14. Therefore, the value of R(S (1,1) , 14) in the reward matrix is +5. The time slot 13 selected by the child node S (1,2) is an unsuccessful time slot, indicating that the data sent by S (1,2) fails to be received. Therefore, the value of R(S (1,2) , 14) in the reward matrix is 0. The child node S (1,1) does not send data in time slot 15, and data is successfully received in time slot 15. Therefore, the value of R(S (1,1) , 15) in the reward matrix is -5. The child node S (1,1) does not send data in time slot 13, and data is not successfully received in time slot 13. Therefore, the value of R(S (1,1) , 13) in the reward matrix is 0. The sub-matrix Q (1,1) of the child node S S(1,1) changes as shown in Figure 2 .
[0076] 5) Update the Q-value table according to the Q-learning formula Q(m, n) ← (1 - γ)·Q(m, n) + γ·R(m, n), where γ is the learning rate, taking values in (0, 1], and the reward matrix R is a matrix with the same size as the Q matrix.
[0077] 6) After the Q-value update is completed, k = k + 1. If the maximum number of iterations K is reached or the Q-value table no longer changes, the stable state is reached; otherwise, repeat steps 2) to 5).
[0078] 7) According to the final Q-value table obtained by iteration, let the node m select the time slot n corresponding to the maximum Q value in the sub-matrix Q m to send data to the sink or the parent node. The parent node F i selects the time slots corresponding to the largest (1 + D Fi ) Q values in the sub-matrix Q i according to its own number of transmissions to send data. The task of allocating the right to use the underwater acoustic channel medium resources is completed. The entire data collection process is as shown in Figure 4 .
[0079] Next, the feasibility of the method of the present invention is verified by computer simulation.
[0080] Deploy a sink on the sea surface, and distribute 20 nodes underwater within 3000 meters around the sink. Among them, there are 11 direct communication nodes and 9 indirect communication nodes. Select 3 of the 11 direct communication nodes as parent nodes, and each parent node is responsible for assisting 3 child nodes in transmitting data. The simulation parameters are set as follows: the underwater sound speed is 1500 meters per second, the transmission rate between nodes and the sink is 1000 bits per second, the data frame and the feedback signal have the same format and the length is set to 1500 bits. The slot length is the same, set to 2 seconds. Assume that the nodes and the sink consume 3 watts of energy when sending signals, while consuming 1 watt of energy when receiving signals.
[0081] If the data transmitted by the node to the sink fails after 5 transmissions, the data will be discarded and new data will be sent. The simulation time is 1h.
[0082] The following is an analysis of the simulation results of the method described in the present invention.
[0083] 1) Network throughput analysis
[0084] Figure 5 This is the change trend of the network throughput of the underwater acoustic double-hop network medium access control method based on Q learning of the present invention. As Figure 5 can be seen, the network throughput of the underwater acoustic network medium access control method with variable number of nodes based on Q learning of the present invention continuously increases with the number of iterations and finally stabilizes at 700 bits per second. Using Q learning to complete the task of allocating the right to use the underwater acoustic channel medium resources, the network performance is continuously improved with the increase of the number of learning times, effectively reducing the sending and receiving of control frames, shortening the channel idle time, and improving the network throughput.
[0085] 2) Transmission energy consumption analysis
[0086] Figure 6 This is the change trend of the average transmission energy consumption of the underwater acoustic double-hop network medium access control method based on Q learning of the present invention. As Figure 6 can be seen, the energy consumption of the underwater acoustic network medium access control method with variable number of nodes based on Q learning of the present invention gradually stabilizes around 9 watts with the number of iterations. In the early stage of learning, due to uneven channel allocation, data packets from different nodes will arrive at the sink at the same time and cause conflicts, resulting in no data reception in another time slot. Therefore, in the early stage of the method learning, the sink does not receive in each time slot, reducing energy consumption. After learning, the underwater acoustic double-hop network medium access control method based on Q learning completes the task of allocating the right to use the underwater acoustic channel medium resources, can effectively avoid conflicts in the stable stage, and does not require the sending and receiving of control frames, so the transmission energy consumption increases slightly.
[0087] 3) Packet delivery ratio analysis
[0088] Figure 7The packet delivery rate change trend of the underwater acoustic double-hop network medium access control method based on Q-learning of the present invention. From Figure 7 It can be seen that the energy consumption of the medium access control method for the underwater acoustic network with variable number of nodes based on Q-learning of the present invention continuously increases with the number of iterations and finally stabilizes at 100%. By using Q-learning to allocate the medium resources of the underwater acoustic channel, conflicts can be effectively avoided after the learning is completed, the success rate of packet transmission can be improved, and a conflict-free data collection process can be carried out in the underwater acoustic double-hop network, effectively improving the network throughput.
[0089] From the above simulation analysis, it can be seen that the medium access control method for the underwater acoustic double-hop network based on Q-learning can effectively reduce the network energy consumption and achieve high packet delivery rate and high network throughput. On the one hand, this method can effectively complete the task of allocating the right to use the medium resources of the underwater acoustic channel, effectively avoid data conflicts, shorten the idle time, reduce the transmission energy consumption, and improve the network throughput; on the other hand, in the designed Q-learning algorithm, each node only needs to store one row in the Q matrix, with low complexity and simple calculation. Moreover, the entire network can effectively adapt to the changes in the ocean environment, has strong anti-interference ability, and only requires a small amount of learning to achieve a conflict-free data collection process in the changed ocean environment.
[0090] The present invention introduces the reinforcement learning algorithm into the medium access control protocol of the underwater acoustic double-hop network, improves the coverage of the underwater acoustic network, uses Q-learning to complete the task of allocating the right to use the medium resources of the underwater acoustic channel, enables the underwater acoustic double-hop network to effectively adapt to the dynamic changes of the ocean environment, reduces the information exchange in the underwater acoustic network, and improves the overall network lifetime. At present, the combination of underwater acoustic networks and artificial intelligence mostly focuses on the routing optimization design, and only a few studies pay attention to its research in the medium access control protocol, and mainly focus on the application in single-hop underwater acoustic networks. How to apply the Q-learning algorithm in the underwater acoustic double-hop network, adapt to the dynamic changes of the ocean environment, expand the coverage of the underwater acoustic network, maintain high network throughput, and enhance the network robustness is of great significance. By applying the Q-learning algorithm, the present invention enables the nodes in the double-hop network to select the optimal transmission time slot without additional information exchange with other nodes, and has the advantages of fast learning speed, strong anti-interference ability, and being applicable to various network node scales.
Claims
1. A medium access control method for an underwater acoustic double-hop network based on Q-learning, characterized in that: After the node receives the feedback signal from the destination or the parent node, node m will, according to the action of sending data in the nth time slot it selects itself, combine the unsuccessful time slot information in the feedback information, and for the mth row of the Q matrix, that is, the sub-matrix Q of the Q matrix stored inside node m m , the mth row R in the reward matrix m (m, :) obtains different values; R m (m, n) represents the reward value obtained after node m selects the action of sending data in the nth time slot and combines the feedback signal; If node m is a direct communication node, the reward value R(m,n) is set as follows: ① If node m sends data in time slot n and time slot n is a successful time slot, it means that the data sent by node m in time slot n is successfully received by the destination. R(m,n) in the reward matrix is a positive value +|ψ| to ensure an increase in the Q(m,n) value; ② If node m sends data in time slot n and time slot n is not a successful time slot, it means that the data sent by node m in the nth time slot is not successfully received by the destination. R(m,n) in the reward matrix is 0 to ensure that the Q(m,n) value approaches 0; ③ If node m does not send data in the nth time slot and time slot n is a successful time slot, it means that the data sent by other nodes in the nth time slot is successfully received by the destination. R(m,n) in the reward matrix is a negative value -|ψ| to ensure a decrease in the Q(m,n) value; ④ If node m does not send data in the nth time slot and time slot n is not a successful time slot; R(m,n) in the reward matrix is 0 to ensure that the Q(m,n) value approaches 0; If node m is a child node, the reward value R(m,n) is set as follows: ① If the parent node is occupied in time slot n, R(m,n) in the reward matrix is a negative value -|ψ| to ensure a decrease in the Q(m,n) value; Only time slots that are not occupied can be judged as successful time slots or unmarked time slots; ② If node m sends data in time slot n and time slot n is a successful time slot, it means that the data sent by node m in time slot n is successfully received by the parent node. R(m,n) in the reward matrix is a positive value +|ψ| to ensure an increase in the Q(m,n) value; ③ If node m sends data in time slot n and time slot n is not a successful time slot, it means that the data sent by node m in the nth time slot is not successfully received by the parent node. R(m,n) in the reward matrix is 0 to ensure that the Q(m,n) value approaches 0; ④ If node m does not send data in time slot n and time slot n is a successful time slot, it means that the data sent by other child nodes in the nth time slot is successfully received by the parent node. R(m,n) in the reward matrix is a negative value -|ψ| to ensure a decrease in the Q(m,n) value; ⑤ If node m does not send data in the nth time slot and time slot n is not a successful time slot; R(m,n) in the reward matrix is 0 to ensure that the Q(m,n) value approaches 0; Update the Q-value table according to the Q-learning formula Q(m,n) ← (1-γ)·Q(m,n) + γ·R(m,n), where γ is the learning rate, with a value in the range (0, 1], and the reward matrix R is a matrix with the same size as the Q matrix; After the Q-value update is completed, k=k+1. If the maximum iteration number K is reached or the Q-value table no longer changes, the stable state is reached; otherwise, repeat the above method; According to the final Q-value table obtained by iteration, let node m select the sub-matrix Q m and send data to the sink or the parent node at the time slot n corresponding to the largest Q value in Q. The parent node F i selects the sub-matrix Q according to its own transmission times Fi and sends data at the time slots corresponding to the largest (1 + D i ) Q values; Complete the task of allocating the right to use the medium resources of the underwater acoustic channel.
2. The medium access control method for an underwater acoustic double-hop network based on Q-learning according to claim 1, the method further includes: Parameter initialization: Consider an underwater acoustic two-hop network, which includes M nodes and 1 sink. The nodes sense information from the ocean environment, and the sink is responsible for collecting the acoustic data sensed by the nodes. Among them, there are A + B nodes within the communication range of the sink and can communicate directly with the sink, that is, direct communication nodes. And the other C nodes are outside the communication range of the sink, and there is at least one direct communication node within the communication range. These nodes cannot communicate directly with the sink but can be forwarded through the direct communication nodes to communicate indirectly with the sink, that is, indirect communication nodes. The underwater acoustic two-hop network is composed of A + B direct communication nodes and C indirect communication nodes in total, and M = A + B + C. First, within A + B direct communication nodes, A nodes are selected to be responsible for forwarding information between the indirect communication nodes and the destination, assisting the indirect communication nodes in performing indirect communication with the destination, and these are called "parent nodes" F i (i = 1, 2, …, A); the communication ranges of these A parent nodes cover all indirect communication nodes, and there is no overlap in the assisted objects; that is, any indirect communication node will only be assisted by one parent node F i in communication, and this indirect communication node is called the child node S F i of the parent node (i,j) (j = 1, 2, …); among the direct communication nodes, the other B nodes that are not parent nodes are simply called "single-hop nodes" E i (i = 1, 2, …, B); Let the time length of a time slot be equal to the data length plus the guard interval. Assume that the data formats and lengths of each node are the same, so the time lengths of the time slots are also the same. Suppose the sink divides the data collection process into N time slots and receives one data packet in each time slot. Therefore, in one data collection process, the sink can receive at most N data packets. To ensure that all data of all nodes in the network can be completely received, let the number of time slots N be equal to the total number of nodes M in the underwater acoustic two-hop network. In Q-learning, the Q matrix applied to medium access control is an M×N matrix. The row m (m = 1, 2, …, M) of the Q matrix represents the node number, and the column n (n = 1, 2, …, N) of the Q matrix represents the time slot number. Therefore, Q(m, n) represents the Q value corresponding to the action that node m selects the nth time slot to send data. The larger the Q value, the greater the priority for node m to select the nth time slot to send data. That is, node m will select the time slot with the highest Q value in the mth row of the Q matrix to send data. If there are multiple identical highest Q values in the mth row, then a time slot will be randomly selected from the multiple time slots with the highest Q value in the mth row to send data. To reduce the computational complexity of the nodes, each node only needs to store the sub-matrix of the row representing its selected transmission time slot inside, that is, node m only needs to store a sub-matrix Q of size 1×N m , and the Q matrix is represented as Q = [Q1; Q2; …; Q m ; …; Q M-1 ; Q M ; The child node cannot communicate directly with the sink. The child node needs to send data to the parent node, and the parent node then forwards it to the sink. Suppose a parent node needs to assist D i child nodes to forward data to the sink. During a round of data collection, the parent node needs to send (1 + D i ) times of data, including 1 time of its own data and D i times of data sent by the child nodes to it. That is, the number of transmissions of the parent node is (1 + D i ); the parent node will select the first (1 + D m ) time slots with the largest Q values in the sub-matrix Q i to send data; while the single-hop node and the child node only need to send 1 piece of their own data during a data collection process. Initialize the iteration number k = 0, the maximum iteration number is K, and the initial Q-value table is a zero matrix of M×N.
3. The method for medium access control of an underwater acoustic two-hop network based on Q learning according to claim 1, the method further includes: The control signal is a signal used to coordinate the behaviors of each node in the underwater acoustic network, including a start signal and a feedback signal. Assume that the formats and lengths of the control signals are the same, so the time lengths of the control time slots are also the same. The time length of the control time slot is equal to 2 times the maximum transmission delay plus the control signal duration. When starting data collection, the destination broadcasts a start signal to all directly communicating nodes in the first control time slot, and then the parent node forwards the start signal to the child nodes in the second control time slot; when node m receives the start signal, it checks the sub-matrix Q according to its own transmission times. m , selects the transmission time slot determined by the sub-matrix Q m , and can send data when it reaches this time slot.
4. The method for medium access control of an underwater acoustic two-hop network based on Q learning according to claim 1, the method further includes: Feedback signal design: When data collection starts, the sink will record the reception situation of each time slot. If the sink successfully receives the complete data within a certain time slot, it is marked as a successful time slot, and the time slot number is noted down. If the sink does not successfully receive the complete data within a certain time slot, the time slot is not marked. After the transmission ends, the sink broadcasts a feedback signal to all direct communication nodes. The feedback signal includes the total number of transmission time slots N and the information of successful time slots. If the time slot selected by node m is not a successful time slot, it means that node m's transmission fails. Similarly, when data collection starts, the parent node will also record the reception situation of each time slot. If the reception time slot of the parent node is occupied within a certain time slot, the parent node will note down this time slot as an occupied time slot, and the child node is not allowed to transmit data within the occupied time slot. Only the time slots that do not belong to the occupied time slots can be judged as successful time slots or unmarked time slots. In other unoccupied time slots, if the parent node successfully receives the complete data within a certain time slot, it is marked as a successful time slot. If the parent node does not successfully receive the complete data within a certain time slot, the time slot is not marked. After the transmission ends, the parent node combines the total number of transmission time slots of the feedback signal from the sink and its own recorded reception situation, and broadcasts a feedback signal to the child nodes it is responsible for. The feedback signal includes the total number of transmission time slots N, the information of occupied time slots, the information of successful time slots, and the information of unsuccessful time slots. The situation that the destination fails to successfully receive complete data within a certain time slot includes three cases: data collision within the time slot, inability to successfully receive all data due to poor channel state, and no node sending data / idle state of the time slot; the situation that the receiving time slot of the parent node is occupied within a certain time slot includes two cases: the parent node selects to transmit data in the time slot, and the data of other single-hop nodes arrives at the parent node in the time slot; the situation that the parent node fails to successfully receive complete data within a certain time slot includes three cases: data collision within the time slot, inability to successfully receive all data due to poor channel state, and no node sending data / idle state of the time slot.
Citation Information
Patent Citations
A medium access control method for underwater acoustic networks with variable node count based on Q-learning
CN113691391B
Underwater acoustic network MAC protocol switching method based on Q-Learning
CN113301032A
Underwater acoustic network medium access control method based on Q learning and data importance
CN114423083A