Data transmission method, device, equipment and medium for underwater wireless sensor network
By adopting the MAC protocol of deep reinforcement learning in underwater wireless sensor networks, building a time slot model and adding a decision network, the node status and action are optimized, which solves the problem of low underwater transmission efficiency and achieves efficient and energy-saving data transmission.
Patent Information
- Application Number
- CN202411701798.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-11-26
AI Technical Summary
In underwater wireless sensor networks, due to the high latency, low bandwidth, high bit error rate and multipath effect of the underwater acoustic channel, the traditional MAC protocol increases the complexity of channel competition and the probability of packet collision when used underwater, resulting in high energy consumption and low transmission efficiency.
A MAC protocol based on deep reinforcement learning is adopted. By configuring the MAC protocol, building a time slot model, defining the node state calculation formula, node action calculation formula and reward function, and adding a decision network to the time slot model, adaptive optimization of node state and action is achieved, reducing energy consumption and improving transmission efficiency.
It maximizes network throughput in underwater wireless sensor networks, saves energy consumption, adapts to various network topologies, solves the conflict problem caused by the long propagation delay characteristics underwater, and improves data transmission efficiency.
Smart Images

Figure CN119450529B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data transmission, and in particular to a data transmission method, device, equipment and medium for an underwater wireless sensor network. Background Art
[0002] The communication environments of underwater acoustic sensor networks differ significantly from those of terrestrial sensor networks. Underwater acoustic channels are characterized by high latency, low bandwidth, high bit error rates, multipath effects, and Doppler shift. Furthermore, underwater nodes are susceptible to passive movement due to the influence of currents or underwater life. Furthermore, underwater nodes are battery-powered, making battery replacement and maintenance expensive and difficult to implement. These characteristics preclude direct application of terrestrial wireless sensor network protocols underwater, necessitating the development of new protocols tailored to the characteristics of underwater networks. The underwater sensor network protocol stack can be divided into the physical layer, data link layer, network layer, transport layer, and application layer. The data link layer primarily addresses medium access control (MAC) protocols. The long propagation delay underwater poses challenges to the development of underwater MAC protocols, increasing the complexity of channel contention and the probability of collisions, leading to packet collisions. Traditional underwater MAC protocols can generally be categorized as contention-based or fixed-assignment-based. These protocols require handshakes or random retransmissions, resulting in additional communication overhead. Their fixed policy makes them highly dependent on the environment and lacks adaptability.
[0003] From the above, it can be seen that how to comprehensively consider the completeness of node observations and the overall energy consumption of the underwater network, maximize network throughput and save energy of underwater sensor nodes, solve the conflicts caused by the long propagation delay characteristics underwater, and improve the efficiency of data transmission in underwater wireless sensor networks are problems to be solved in this field. Summary of the Invention
[0004] In view of this, the present invention aims to provide a data transmission method, apparatus, device, and medium for an underwater wireless sensor network. These methods comprehensively consider the completeness of node observations and the overall energy consumption of the underwater network, maximize network throughput, conserve energy at underwater sensor nodes, resolve conflicts caused by the long propagation delay characteristics underwater, and improve the efficiency of data transmission in underwater wireless sensor networks. The specific solution is as follows:
[0005] In a first aspect, the present application discloses a data transmission method for an underwater wireless sensor network, which is applied to any sending node in a distributed underwater wireless sensor network, comprising:
[0006] Configure the MAC protocol and basic protocol information, build a time slot model, define a node state calculation formula, a node action calculation formula, and a reward function based on the MAC protocol and basic protocol information, and add a decision network to the time slot model;
[0007] Acquire initial node data, calculate the initial node data using the node state calculation formula and the node action calculation formula to obtain an initial node state and an initial node action, and if the initial node action is to send data, perform data communication with other sending nodes except the local node;
[0008] After obtaining the confirmation character fed back by the other sending node, calculating a reward value based on the confirmation character and using the reward function, and updating the initial node state according to the reward value to obtain an updated node state;
[0009] Updating the decision network in the time slot model using the initial node state, the initial node action, the reward value, and the updated node state to obtain the updated decision network;
[0010] The current node data is obtained and input into the updated decision network to output a node data transmission strategy, and data transmission between local and distributed receiving nodes in the underwater wireless sensor network is performed according to the node data transmission strategy.
[0011] Optionally, the constructing of the time slot model includes:
[0012] Build the initial time slot model and determine the task objectives based on business needs;
[0013] The local data transmission work is used as a training parameter, and the initial time slot model is trained in combination with the task target to obtain the time slot model.
[0014] Optionally, the node status calculation formula is:
[0015] ;
[0016] ;
[0017] in, The action of node j at time slot t observed by node i in the underwater sensor is: is the action selected by node j at time slot t, AP is the receiving node, Whether node i observes that AP returns a confirmation character to node j in time slot t, is the action chosen by node j at time slot t.
[0018] Optionally, the node action calculation formula is:
[0019] ;
[0020] in, is the action selected by node i at time slot t, p is the transmit power level, p = {1, 2, …, p-1, p}.
[0021] Optionally, the reward function is:
[0022] ;
[0023] in, is the reward value obtained by node i in time slot t, p is the transmission power level, p={1,2,…,p-1,p}, Reward and penalty items.
[0024] Optionally, adding a decision network to the time slot model includes:
[0025] A MAAC algorithm based on a policy-value architecture is added to the time slot model; the MAAC algorithm includes a decision network; the decision network includes a policy network and a value network.
[0026] Optionally, updating the decision network in the time slot model using the initial node state, the initial node action, the reward value, and the updated node state includes:
[0027] Generate update information based on the initial node state, the initial node action, the reward value, and the updated node state, and store the update information in a local buffer;
[0028] The decision network in the time slot model is updated using the update information corresponding to each time step in the buffer.
[0029] In a second aspect, the present application discloses a data transmission device for an underwater wireless sensor network, which is applied to any sending node in a distributed underwater wireless sensor network, comprising:
[0030] A construction and definition module is used to configure the MAC protocol and basic protocol information, build a time slot model, define the node state calculation formula, node action calculation formula and reward function based on the MAC protocol and the basic protocol information, and add a decision network to the time slot model;
[0031] a calculation module, configured to obtain initial node data, calculate the initial node data using the node state calculation formula and the node action calculation formula to obtain an initial node state and an initial node action, and if the initial node action is to send data, perform data communication with other sending nodes other than the local node;
[0032] a node state updating module, configured to, upon obtaining a confirmation character fed back by the other sending node, calculate a reward value based on the confirmation character and using the reward function, and update the initial node state according to the reward value to obtain an updated node state;
[0033] a decision network updating module, configured to update the decision network in the time slot model using the initial node state, the initial node action, the reward value, and the updated node state to obtain the updated decision network;
[0034] The data transmission module is used to obtain current node data and input the current node data into the updated decision network to output a node data transmission strategy, and perform data transmission between local and distributed receiving nodes in the underwater wireless sensor network according to the node data transmission strategy.
[0035] In a third aspect, the present application discloses an electronic device, comprising:
[0036] Memory, used to store computer programs;
[0037] The processor is used to execute the computer program to implement the aforementioned data transmission method for the underwater wireless sensor network.
[0038] In a fourth aspect, the present application discloses a computer storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the aforementioned disclosed method for transmitting data in an underwater wireless sensor network are implemented.
[0039] It can be seen that the present application provides a data transmission method for an underwater wireless sensor network, including configuring a MAC protocol and basic protocol information, constructing a time slot model, defining a node state calculation formula, a node action calculation formula and a reward function based on the MAC protocol and the basic protocol information, and adding a decision network to the time slot model; obtaining initial node data, calculating the initial node data using the node state calculation formula and the node action calculation formula to obtain an initial node state and an initial node action, and if the initial node action is to send data, performing data communication with other sending nodes other than the local node; when feedback from the other sending nodes is obtained, After the confirmation character is received, the reward value is calculated based on the confirmation character and using the reward function, and the initial node state is updated according to the reward value to obtain the updated node state; the decision network in the time slot model is updated using the initial node state, the initial node action, the reward value and the updated node state to obtain the updated decision network; the current node data is obtained, and the current node data is input into the updated decision network to output a node data transmission strategy, and data transmission between the local and distributed receiving nodes in the underwater wireless sensor network is performed according to the node data transmission strategy. This application is applied to any sending node in a distributed underwater wireless sensor network, configures the MAC protocol and basic protocol information, builds a time slot model, and maximizes network throughput without the need to cooperate with the traditional MAC protocol. It defines the node state calculation formula, node action calculation formula and reward function, and adds a decision network to the time slot model. It can be applied to various network topologies, takes into account the completeness of node observation and the overall energy consumption of the underwater network, and more effectively maximizes network throughput and saves energy at underwater sensor nodes. It calculates the initial node state and initial node action. If the initial node action is to send data, it communicates with other sending nodes except the local node for data transmission. According to communication, after obtaining the confirmation characters fed back by other sending nodes, the reward value is calculated and the initial node state is updated, so as to update the decision network in the time slot model to obtain an updated decision network, and the current node data is input into the updated decision network to output the node data transmission strategy. According to the node data transmission strategy, data is transmitted between the receiving nodes in the local and distributed underwater wireless sensor network, taking into account the conflict problem caused by the long propagation delay characteristics underwater, solving the conflict caused by the long propagation delay characteristics underwater, and at the same time, without the need to cooperate with the traditional MAC protocol, it has high adaptability and improves the efficiency of data transmission in the underwater wireless sensor network. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0041] Figure 1 This is a flow chart of a data transmission method for an underwater wireless sensor network disclosed in this application;
[0042] Figure 2 A diagram showing the time slot length of a data packet disclosed in this application;
[0043] Figure 3 A flowchart for obtaining an initial node data set and calculating updates disclosed in this application;
[0044] Figure 4 This is a data transmission flow chart of a distributed semi-cooperative underwater acoustic network MAC protocol based on deep reinforcement learning disclosed in this application;
[0045] Figure 5 This is a structural diagram of a data transmission device for an underwater wireless sensor network disclosed in this application;
[0046] Figure 6 This is a structural diagram of an electronic device provided in this application. DETAILED DESCRIPTION
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0048] The communication environments of underwater acoustic sensor networks differ significantly from those of terrestrial sensor networks. Underwater acoustic channels are characterized by high latency, low bandwidth, high bit error rates, multipath effects, and Doppler shift. Furthermore, underwater nodes are susceptible to passive movement due to the influence of currents or underwater life. Furthermore, underwater nodes are battery-powered, making battery replacement and maintenance expensive and difficult to implement. These characteristics preclude direct application of terrestrial wireless sensor network protocols underwater, necessitating the development of new protocols tailored to the characteristics of underwater networks. The underwater sensor network protocol stack can be divided into the physical layer, data link layer, network layer, transport layer, and application layer. The data link layer primarily addresses medium access control (MAC) protocols. The long propagation delay underwater poses challenges to the development of underwater MAC protocols, increasing the complexity of channel contention and the probability of collisions, leading to packet collisions. Traditional underwater MAC protocols can generally be categorized as contention-based or fixed-assignment-based. These protocols require handshake exchanges or random retransmissions, resulting in additional communication overhead. Their fixed policy makes them highly dependent on the environment and lacks adaptability. From the above, it can be seen that how to comprehensively consider the completeness of node observations and the overall energy consumption of the underwater network, maximize network throughput and save energy of underwater sensor nodes, solve the conflicts caused by the long propagation delay characteristics underwater, and improve the efficiency of data transmission in underwater wireless sensor networks are problems to be solved in this field.
[0049] See also Figure 1 As shown, an embodiment of the present invention discloses a data transmission method for an underwater wireless sensor network, which is applied to any sending node in a distributed underwater wireless sensor network and may specifically include:
[0050] Step S11: Configure the MAC protocol and basic protocol information, build a time slot model, define the node state calculation formula, node action calculation formula and reward function based on the MAC protocol and the basic protocol information, and add a decision network to the time slot model.
[0051] In this embodiment, the MAC protocol and basic protocol information are configured, an initial time slot model is constructed, and a task objective is determined according to business needs; local data transmission work is used as a training parameter, and the initial time slot model is trained in combination with the task objective to obtain the time slot model, and a node state calculation formula, a node action calculation formula, and a reward function are defined based on the MAC protocol and the basic protocol information, and a MAAC algorithm based on a strategy-value architecture is added to the time slot model; the MAAC algorithm includes a decision network; the decision network includes a strategy network and a value network.
[0052] Specifically, the present application is applied to any sending node in a distributed underwater wireless sensor network, and a MAC protocol is configured in the sending node. For all sending nodes, the sending nodes in N underwater wireless sensor networks are configured with the same MAC protocol, and data packets are transmitted to the AP central receiving node in a time slot manner, and the basic information of the protocol is configured. Then, an initial time slot model is constructed. The task goal is to train a conflict-free data transmission strategy based on a conventional time slot model to maximize the throughput of the underwater wireless sensor network. Based on this goal, the data transmission power of the underwater wireless sensor node is used as a training parameter of the model to realize the management and allocation of the energy of the underwater sensor node. In the time slot model, it is assumed that the data packets of any underwater wireless sensor data sending node have the same packet length. , and the ACK (Acknowledge character) packet from the AP center receiving node to each data sending node also has the same packet length , the packet time slot length is as follows Figure 2 As shown, the length of each time slot Corresponds to the duration of the node sending the data packet .
[0053] The MAC protocol takes into account the long propagation delays underwater. The central receiving node (AP) can only receive one data packet at a time. If two or more data packets from different underwater wireless sensor nodes arrive at the AP at the same time, a data reception conflict occurs at the receiving node. The data transmission and reception processes of each underwater wireless sensor node are mutually exclusive. This means that if a data packet from another node or an ACK from a receiving node arrives at the current node while the current node is transmitting data at the same time, a receive-transmit conflict occurs. Like the receiving nodes, each underwater wireless sensor node can only receive one data packet at a time.
[0054] In this embodiment, the node status calculation formula is:
[0055] ;
[0056] ;
[0057] in, The action of node j at time slot t observed by node i in the underwater sensor is: is the action selected by node j at time slot t, AP is the receiving node, Whether node i observes that AP returns a confirmation character to node j in time slot t, is the action selected by node j at time slot t;
[0058] The node action calculation formula is:
[0059] ;
[0060] in, is the action selected by node i at time slot t, p is the transmit power level, p = {1, 2, …, p-1, p};
[0061] The reward function is:
[0062] ;
[0063] in, is the reward value obtained by node i in time slot t, p is the transmission power level, p={1,2,…,p-1,p}, Reward and penalty items.
[0064] The specific principles for defining the node state calculation formula, node action calculation formula, and reward function are as follows: (1) Define the node state calculation formula, where the state s can be regarded as the observation information parameter ,The state of each underwater wireless sensor node is divided into two parts, namely whether the node transmits data and the reception of the node feedback information ACK. is the observation information of underwater sensor node i at time slot t, The underwater sensor node i observes the action of node j in time slot t, which is based on the strategy get, Is the sensor node i observing whether the AP returns ACK to node j in time slot t. Ideally, any node i can get In reality, due to the uncertainty of the underwater environment and the collision of data packets transmitted by different nodes in the underwater wireless sensor network, some information may not be obtained correctly, so the above node state calculation formula is defined. In reality, the underwater environment changes dynamically. In order to avoid invalid data caused by a single observation anomaly in a certain time slot, the observation of the actual input model is defined. , which not only includes the observation of underwater sensor node i itself at the current time slot t , we should also consider the state information observed by node i in the past M time slots. ;
[0065] (2) Define the node action calculation formula, set is the action selected by node i at time slot t, which is based on the strategy Calculation shows that when energy consumption is not considered, the action space is only 2, with the two cases of sending or refusing to send a data packet represented by 1 and 0 respectively. If energy consumption is considered, successful communication between underwater wireless sensor nodes depends on sufficient communication signal strength to ensure that the signal can be transmitted reliably within the predetermined distance and overcome the influence of environmental interference and noise. The communication signal strength is described by the SINR (Signal to Interference plus Noise Ratio) at the receiving node:
[0066] ;
[0067] in, is the node receiving power, For interference, The node receiving power is obtained by the node transmitting power after the energy loss due to underwater propagation.
[0068] ;
[0069] ;
[0070] ;
[0071] Among them, spre is the underwater geometric propagation loss, abso is the absorption loss, is Thorp's formula, f is the signal frequency, and is usually measured in Hertz.
[0072] ;
[0073] Let the transmit power level be defined as p={1,2,...,p-1,p}, where p can be adjusted according to the network topology. Each level corresponds to a transmit power level, and a node action calculation formula is defined.
[0074] (3) Define the reward function, let is the case where sensor node i receives the ACK feedback from AP in time slot t, and its definition is as follows:
[0075] ;
[0076] The feedback information is divided into two types, and the node's probability estimate of successful transmission at the current power consumption level is dynamically changed according to different transmission results. The formula defines that when the underwater wireless sensor node does not transmit data or the data packet transmission fails, the feedback value obtained is 0; when the data is successfully transmitted, the feedback value obtained is 1. A successful data transmission process can be described as node i at time slot t using the lowest power level p that meets the packet transmission requirements to send a data packet and the data packet is successfully received, and the receiving node AP feeds back an ACK to node i to confirm , which indicates whether the data sent by node i at time slot t is successfully transmitted to the receiving node.
[0077] definition is the reward obtained by node i in time slot t. According to the goal of the project, it is necessary to minimize the energy consumption of sensor nodes and maximize the throughput while the underwater sensor nodes successfully communicate. Define that if a node does not send a data packet, it will be given ; If the data packet is successfully sent and ACK is received, If the node sending the data packet does not receive the returned ACK, When a node sends a data packet, it will choose a certain gear in p={1,2,…,p-1,p}. Different transmission powers correspond to different node energy consumption. For node energy loss, a reward and penalty term is defined. Each time a data packet is sent, the energy consumption dEng is calculated and the penalty is included in the reward to define the reward function.
[0078] When a time slot ends, the information of the time slot is combined into an experience sample Store it in the training data space for future training. Suppose node i receives the data packet from node j at time slot l, then consider is the sum of the data packet propagation time and the data packet sending time between two nodes, Divide the time slot length and round it up to get the number of interval time slots , then in fact The observation information should be updated. If all observation information is updated before the next step is performed, the idle waiting time cannot be avoided, so the model does not guarantee that all observation information Timely update: the initial small number of data samples are in a state of partial information missing, which can be restored to normal after a certain period of time, and it does not affect model training.
[0079] In this embodiment, the principle of adding a decision network to the time slot model is as follows: the model in the MAC protocol specifically adopts the MAAC algorithm based on the Actor-Critic (policy-value) architecture. The Critic network uses a Value-based approach to evaluate the value of action selection, while the Actor network uses a Policy-based approach to predict the probability distribution of actions through input states, and then generates action selections. Each underwater wireless sensor node is equipped with an Actor network and a Critic network that shares global information. The Actor network inputs the local observation of node i based on the Critic evaluation results and selects the best action, with the goal of maximizing the Q value evaluated by the Critic network. The Critic network uses the temporal difference error to update its network weights, with the goal of minimizing the TD error. The loss function is as follows:
[0080] ;
[0081] in, For reward, is the discount factor, is the entropy coefficient, which controls the balance between exploration and utilization. , which is calculated as follows:
[0082] ;
[0083] in, To contribute to the value of node i itself, the node needs to input the local observation values of all underwater sensor nodes and action sets Calculate the value contribution of other underwater sensor nodes other than node i is the weighted sum of the value of each node:
[0084] .
[0085] Step S12: Obtain initial node data, and use the node state calculation formula and the node action calculation formula to calculate the initial node data to obtain the initial node state and initial node action. If the initial node action is to send data, data communication is performed with other sending nodes except the local node.
[0086] Step S13: After obtaining the confirmation character fed back by the other sending nodes, a reward value is calculated based on the confirmation character and using the reward function, and the initial node state is updated according to the reward value to obtain an updated node state.
[0087] Step S14: updating the decision network in the time slot model using the initial node state, the initial node action, the reward value, and the updated node state to obtain the updated decision network.
[0088] In this embodiment, update information is generated based on the initial node state, initial node action, reward value and updated node state, and the update information is stored in a local buffer; the decision network in the time slot model is updated using the update information corresponding to each time step in the buffer to obtain the updated decision network.
[0089] In this embodiment, the process of obtaining initial node data and calculating updates is as follows: Figure 3 As shown in the figure: Each underwater wireless sensor node initializes its state s and converts it into the state input of the model. Each underwater node selects one of the two actions of waiting or sending data according to its current state through the model. In the initial training phase of the model, the algorithm adopts Greedy strategy, with probability Randomly select actions to explore with probability 1- Select the optimal action under the current model's known strategy, as training progresses Gradually decreases, and the model gradually transitions from the exploration stage to the optimization stage. When each underwater wireless sensor node determines the action and starts to execute, it considers the long propagation delay characteristics of underwater nodes for node communication and observes the situation of the AP receiving the node to obtain the confirmation information ACK. The confirmation information is input into the reward function to obtain the reward value of the current underwater node in the current time slot communication. If the node does not transmit data, the reward value is zero. When the above process is completed, the state of each underwater node in the current time slot is ,action , reward value and the updated status All data are stored in the experience replay buffer and are waiting to be updated. The model will update data based on the experience stored in the experience replay buffer at each time step. Update the policy network, gradually adjust the action selection strategy of the node in different states, and finally obtain the appropriate node sending strategy. Update the model network during the training phase. Initialize the parameters of the Actor and Critic networks of node i at the beginning of the model operation. and After that, each time slot continuously carries out interactive communication between underwater wireless sensor nodes and other nodes in the network in the underwater environment, and the experience of each time step is Store it in the experience replay buffer for later use. When updating the Critic network, sample a batch of data from the experience replay buffer and calculate the target Q value:
[0090] ;
[0091] Then calculate the current Q value , and minimize the TD error To update the parameters of the Critic network , while evaluating the action selected by the Actor network output to calculate the policy gradient, maximize the Q value evaluated by the Critic network to update the parameters of the Actor network , the critic network parameters are periodically adjusted in the retraining phase Copy to target network parameters ,After the training is completed, the state is input into the target network, and the ,corresponding data transmission strategy for underwater wireless sensor nodes can be ,derived.
[0092] Step S15: Acquire current node data and input the current node data into the updated decision network to output a node data transmission strategy, and perform data transmission between local and distributed receiving nodes in the underwater wireless sensor network according to the node data transmission strategy.
[0093] This application proposes a data transmission method of a distributed semi-cooperative MAC protocol based on deep reinforcement learning. The specific process is as follows Figure 4 As shown, first, the MAC protocol and basic protocol information are configured to build a time slot model. Then, the node state calculation formula, node action calculation formula and reward function are defined. Then, a decision network is added to the time slot model. The node state calculation formula and node action calculation formula are used to calculate the initial node data to obtain the initial node state and initial node action. If the initial node action is to send data, data communication is performed with other sending nodes other than the local node. After obtaining the confirmation character fed back by other sending nodes, the reward value is calculated based on the confirmation character and the reward function. The initial node state is updated according to the reward value to obtain the updated node state. The decision network in the time slot model is updated using the initial node state, initial node action, reward value and updated node state to obtain the updated decision network. Finally, the current node data is obtained and input into the updated decision network to output the node data transmission strategy. According to the node data transmission strategy, data transmission is performed between the sending node and the receiving node in the distributed underwater wireless sensor network.
[0094] This application can maximize network throughput without the need to cooperate with traditional MAC protocols, and this application focuses on the energy consumption problem of communication between underwater network nodes, and comprehensively considers the various conflicts in the underwater communication process caused by the long propagation delay characteristics underwater, so that the research work has high fidelity. Traditional MAC protocols require handshake information exchange or random time retransmission, which leads to additional communication overhead, and the fixed strategy makes it highly dependent on the environment and lacks adaptability. Applying reinforcement learning to MAC protocol design enables continuous interaction training between sensor nodes and the environment to derive transmission strategies that can be applied to various network topologies. This application takes into account the completeness of node observations and the overall energy consumption of the underwater network, more effectively achieving maximum network throughput and saving energy for underwater sensor nodes, comprehensively considering the conflict problem caused by the long propagation delay characteristics underwater, improving the detection of conflict situations, and at the same time, being highly adaptable.
[0095] In this embodiment, a MAC protocol and basic protocol information are configured to construct a time slot model. A node state calculation formula, a node action calculation formula and a reward function are defined based on the MAC protocol and the basic protocol information, and a decision network is added to the time slot model. Initial node data is obtained, and the initial node data is calculated using the node state calculation formula and the node action calculation formula to obtain an initial node state and an initial node action. If the initial node action is to send data, data communication is performed with other sending nodes other than the local node. After obtaining the confirmation character fed back by the other sending nodes, a reward value is calculated based on the confirmation character and using the reward function, and the initial node state is updated according to the reward value to obtain an updated node state. The decision network in the time slot model is updated using the initial node state, the initial node action, the reward value and the updated node state to obtain an updated decision network. Current node data is obtained, and the current node data is input into the updated decision network to output a node data transmission strategy. Data transmission between the local and distributed receiving nodes in the underwater wireless sensor network is performed according to the node data transmission strategy. This application is applied to any sending node in a distributed underwater wireless sensor network, configures the MAC protocol and basic protocol information, builds a time slot model, and maximizes network throughput without the need to cooperate with the traditional MAC protocol. It defines the node state calculation formula, node action calculation formula and reward function, and adds a decision network to the time slot model. It can be applied to various network topologies, takes into account the completeness of node observation and the overall energy consumption of the underwater network, and more effectively maximizes network throughput and saves energy at underwater sensor nodes. It calculates the initial node state and initial node action. If the initial node action is to send data, it communicates with other sending nodes except the local node for data transmission. According to communication, after obtaining the confirmation characters fed back by other sending nodes, the reward value is calculated and the initial node state is updated, so as to update the decision network in the time slot model to obtain an updated decision network, and the current node data is input into the updated decision network to output the node data transmission strategy. According to the node data transmission strategy, data is transmitted between the receiving nodes in the local and distributed underwater wireless sensor network, taking into account the conflict problem caused by the long propagation delay characteristics underwater, solving the conflict caused by the long propagation delay characteristics underwater, and at the same time, without the need to cooperate with the traditional MAC protocol, it has high adaptability and improves the efficiency of data transmission in the underwater wireless sensor network.
[0096] See also Figure 5 As shown, an embodiment of the present invention discloses a data transmission device for an underwater wireless sensor network, which is applied to any sending node in a distributed underwater wireless sensor network and may specifically include:
[0097] A construction and definition module 11 is used to configure the MAC protocol and basic protocol information, build a time slot model, define a node state calculation formula, a node action calculation formula, and a reward function based on the MAC protocol and the basic protocol information, and add a decision network to the time slot model;
[0098] a calculation module 12 for obtaining initial node data, calculating the initial node data using the node state calculation formula and the node action calculation formula to obtain an initial node state and an initial node action, and performing data communication with other sending nodes other than the local node if the initial node action is to send data;
[0099] The node state updating module 13 is configured to, upon obtaining the confirmation character fed back by the other sending nodes, calculate a reward value based on the confirmation character and using the reward function, and update the initial node state according to the reward value to obtain an updated node state;
[0100] a decision network updating module 14, configured to update the decision network in the time slot model using the initial node state, the initial node action, the reward value, and the updated node state to obtain the updated decision network;
[0101] The data transmission module 15 is used to obtain current node data and input the current node data into the updated decision network to output a node data transmission strategy, and perform data transmission between local and distributed receiving nodes in the underwater wireless sensor network according to the node data transmission strategy.
[0102] In this embodiment, a MAC protocol and basic protocol information are configured to construct a time slot model. A node state calculation formula, a node action calculation formula and a reward function are defined based on the MAC protocol and the basic protocol information, and a decision network is added to the time slot model. Initial node data is obtained, and the initial node data is calculated using the node state calculation formula and the node action calculation formula to obtain an initial node state and an initial node action. If the initial node action is to send data, data communication is performed with other sending nodes other than the local node. After obtaining the confirmation character fed back by the other sending nodes, a reward value is calculated based on the confirmation character and using the reward function, and the initial node state is updated according to the reward value to obtain an updated node state. The decision network in the time slot model is updated using the initial node state, the initial node action, the reward value and the updated node state to obtain an updated decision network. Current node data is obtained, and the current node data is input into the updated decision network to output a node data transmission strategy. Data transmission between the local and distributed receiving nodes in the underwater wireless sensor network is performed according to the node data transmission strategy. This application is applied to any sending node in a distributed underwater wireless sensor network, configures the MAC protocol and basic protocol information, builds a time slot model, and maximizes network throughput without the need to cooperate with the traditional MAC protocol. It defines the node state calculation formula, node action calculation formula and reward function, and adds a decision network to the time slot model. It can be applied to various network topologies, takes into account the completeness of node observation and the overall energy consumption of the underwater network, and more effectively maximizes network throughput and saves energy at underwater sensor nodes. It calculates the initial node state and initial node action. If the initial node action is to send data, it communicates with other sending nodes except the local node for data transmission. According to communication, after obtaining the confirmation characters fed back by other sending nodes, the reward value is calculated and the initial node state is updated, so as to update the decision network in the time slot model to obtain an updated decision network, and the current node data is input into the updated decision network to output the node data transmission strategy. According to the node data transmission strategy, data is transmitted between the receiving nodes in the local and distributed underwater wireless sensor network, taking into account the conflict problem caused by the long propagation delay characteristics underwater, solving the conflict caused by the long propagation delay characteristics underwater, and at the same time, without the need to cooperate with the traditional MAC protocol, it has high adaptability and improves the efficiency of data transmission in the underwater wireless sensor network.
[0103] In some specific embodiments, the construction and definition module 11 may specifically include:
[0104] The initial model building module is used to build the initial time slot model and determine the task objectives according to business needs;
[0105] The model training module is used to use the local data transmission work as a training parameter and train the initial time slot model in combination with the task target to obtain the time slot model.
[0106] In some specific embodiments, the node status calculation formula is:
[0107] ;
[0108] ;
[0109] in, The action of node j at time slot t observed by node i in the underwater sensor is: is the action selected by node j at time slot t, AP is the receiving node, Whether node i observes that AP returns a confirmation character to node j in time slot t, is the action chosen by node j at time slot t.
[0110] In some specific embodiments, the node action calculation formula is:
[0111] ;
[0112] in, is the action selected by node i at time slot t, p is the transmit power level, p = {1, 2, …, p-1, p}.
[0113] In some specific embodiments, the reward function is:
[0114] ;
[0115] in, is the reward value obtained by node i in time slot t, p is the transmission power level, p={1,2,…,p-1,p}, Reward and penalty items.
[0116] In some specific embodiments, the construction and definition module 11 may specifically include:
[0117] An algorithm adding module is used to add a MAAC algorithm based on a policy-value architecture to the time slot model; the MAAC algorithm includes a decision network; the decision network includes a policy network and a value network.
[0118] In some specific embodiments, the decision network update module 14 may specifically include:
[0119] An update information generation module, configured to generate update information based on the initial node state, the initial node action, the reward value, and the updated node state, and store the update information in a local buffer;
[0120] An updating module is used to update the decision network in the time slot model using the update information corresponding to each time step in the buffer.
[0121] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the underwater wireless sensor network data transmission method performed by the electronic device as disclosed in any of the aforementioned embodiments.
[0122] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0123] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon include an operating system 221, a computer program 222 and data 223, etc. The storage method can be temporary storage or permanent storage.
[0124] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, enabling the processor 21 to operate and process data 223 in the memory 22. It can be run under Windows, Unix, Linux, or other operating systems. In addition to including computer programs capable of implementing the underwater wireless sensor network data transmission method performed by the electronic device 20 as disclosed in any of the aforementioned embodiments, the computer program 222 may also include computer programs capable of performing other specific tasks. Data 223 may include data transmitted from external devices and received by the underwater wireless sensor network's data transmission device, as well as data collected by its own input / output interface 25.
[0125] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0126] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium, in which a computer program is stored. When the computer program is loaded and executed by a processor, the steps of the data transmission method of the underwater wireless sensor network disclosed in any of the aforementioned embodiments are implemented.
[0127] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0128] The above is a detailed introduction to the data transmission method, device, equipment and storage medium of an underwater wireless sensor network provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A data transmission method for an underwater wireless sensor network, characterized in that: Applicable to any sending node in a distributed underwater wireless sensor network, including: Configure the MAC protocol and basic protocol information, build a time slot model, define a node state calculation formula, a node action calculation formula, and a reward function based on the MAC protocol and basic protocol information, and add a decision network to the time slot model; Acquire initial node data, calculate the initial node data using the node state calculation formula and the node action calculation formula to obtain an initial node state and an initial node action, and if the initial node action is to send data, perform data communication with other sending nodes except the local node; After obtaining the confirmation character fed back by the other sending node, calculating a reward value based on the confirmation character and using the reward function, and updating the initial node state according to the reward value to obtain an updated node state; Updating the decision network in the time slot model using the initial node state, the initial node action, the reward value, and the updated node state to obtain the updated decision network; The current node data is obtained and input into the updated decision network to output a node data transmission strategy, and data transmission between local and distributed receiving nodes in the underwater wireless sensor network is performed according to the node data transmission strategy.
2. The data transmission method of underwater wireless sensor network according to claim 1, characterized in that: The constructing of the time slot model includes: Build the initial time slot model and determine the task objectives based on business needs; The local data transmission work is used as a training parameter, and the initial time slot model is trained in combination with the task target to obtain the time slot model.
3. The data transmission method of underwater wireless sensor network according to claim 1, characterized in that: The node status calculation formula is: ; ; in, The action of node j at time slot t observed by node i in the underwater sensor is: is the action selected by node j at time slot t, AP is the receiving node, Whether node i observes that AP returns a confirmation character to node j in time slot t, is the action chosen by node j at time slot t.
4. The data transmission method of underwater wireless sensor network according to claim 1, characterized in that: The node action calculation formula is: ; in, is the action selected by node i at time slot t, p is the transmit power level, p = {1, 2, …, p-1, p}.
5. The data transmission method of underwater wireless sensor network according to claim 1, characterized in that: The reward function is: ; in, is the reward value obtained by node i in time slot t, p is the transmission power level, p={1,2,…,p-1,p}, Reward and penalty items.
6. The data transmission method of underwater wireless sensor network according to claim 1, characterized in that: Adding a decision network to the time slot model includes: A MAAC algorithm based on a policy-value architecture is added to the time slot model; the MAAC algorithm includes a decision network; the decision network includes a policy network and a value network.
7. The data transmission method of an underwater wireless sensor network according to any one of claims 1 to 6, characterized in that: The updating of the decision network in the time slot model using the initial node state, the initial node action, the reward value, and the updated node state includes: Generate update information based on the initial node state, the initial node action, the reward value, and the updated node state, and store the update information in a local buffer; The decision network in the time slot model is updated using the update information corresponding to each time step in the buffer.
8. A data transmission device for an underwater wireless sensor network, characterized in that: Applicable to any sending node in a distributed underwater wireless sensor network, including: A construction and definition module is used to configure the MAC protocol and basic protocol information, build a time slot model, define the node state calculation formula, node action calculation formula and reward function based on the MAC protocol and the basic protocol information, and add a decision network to the time slot model; a calculation module, configured to obtain initial node data, calculate the initial node data using the node state calculation formula and the node action calculation formula to obtain an initial node state and an initial node action, and if the initial node action is to send data, perform data communication with other sending nodes other than the local node; a node state updating module, configured to, upon obtaining a confirmation character fed back by the other sending node, calculate a reward value based on the confirmation character and using the reward function, and update the initial node state according to the reward value to obtain an updated node state; a decision network updating module, configured to update the decision network in the time slot model using the initial node state, the initial node action, the reward value, and the updated node state to obtain the updated decision network; The data transmission module is used to obtain current node data and input the current node data into the updated decision network to output a node data transmission strategy, and perform data transmission between local and distributed receiving nodes in the underwater wireless sensor network according to the node data transmission strategy.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor is configured to execute the computer program to implement the data transmission method for an underwater wireless sensor network according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Used to store a computer program; wherein, when the computer program is executed by a processor, the data transmission method for an underwater wireless sensor network according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Underwater wireless sensor network routing method based on reinforcement learning
CN112954769A
Underwater wireless sensor network routing method based on multi-agent reinforcement learning
CN115843083A