Static UWSNs routing protocol optimization method and system based on convergence guidance
By adopting a Q-Learning-based forwarding decision model and aggregation and guidance mechanism in static UWSNs, the problems of routing detours and routing loops are solved, the packet delivery rate is improved and the delay and energy consumption are reduced.
Patent Information
- Application Number
- CN202510013787.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-06
AI Technical Summary
There are routing detours and routing loop problems in the static topology UWSNs routing protocol, resulting in high-end to-end delays and low delivery rates.
Using a forwarding decision model based on Q-Learning, combined with aggregation guidance mechanism, the optimal forwarding path from the source node to the aggregation node is optimized, and the detours and cycles are reduced.
The routing detour and routing cycle phenomena are effectively optimized, the packet delivery rate is improved, the end-to-end delay is reduced, and the forwarding energy consumption is reduced.
Smart Images

Figure CN120050740A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of optimizing static UWSNs routing protocols, and specifically relates to a method and system for optimizing a static UWSNs routing protocol based on sink guidance. Background Technique
[0002] With the development of the economy and the continuous deepening of people's understanding of the ocean, various scientific and technological innovation demands such as ocean environmental monitoring and ocean resource exploration have emerged. Underwater Wireless Sensor Networks (UWSNs) have excellent advantages in underwater continuous monitoring or underwater target detection. However, the reliable operation of UWSNs depends on the efficiency of data collection of its routing protocol. However, the constraints of underwater channel conditions make the reliability and efficiency of underwater routing relatively low.
[0003] However, the proactive routing protocol has a good energy-saving effect in static UWSNs. However, due to the single forwarding link, there are problems such as routing detours and even routing loops, which pose hazards to the delivery rate and end-to-end delay of UWSNs data collection. Aiming at the problems of high end-to-end delay and low delivery rate caused by routing detours and routing loops in the routing protocol of static topology UWSNs. Summary of the Invention
[0004] The present invention aims to solve the problems of high end-to-end delay and low delivery rate caused by routing detours and routing loops in the routing protocol of static topology UWSNs. To solve the above technical problems, the present invention is realized through the following technical solutions:
[0005] Solution 1: The present invention proposes a method for optimizing a static UWSNs routing protocol based on sink guidance, and the method includes the following steps:
[0006] Step 1: Based on the forwarding decision model of Q-Learning, establish the best forwarding path from the source node to the sink node;
[0007] Step 2: According to whether there are completely shared parameters between nodes in the static UWSNs routing protocol, divide the forwarding decision model in Step 1 into two categories, including a decision model with data packets as agents and a decision model with nodes as agents;
[0008] Determine the decision state of the information of the data packet location node and forwarding candidate nodes in the static UWSNs routing protocol. When this node needs to forward a data packet, calculate the Q values of each neighbor node in turn, and finally select the neighbor node with the largest Q value as the next-hop forwarding target;
[0009] Step 3: Construct a convergence guidance mechanism. The convergence guidance mechanism allows nodes to generate training rewards based on the aggregated location information of each node during the training process of the Q-Learning forwarding decision model, generating a path for data packets to converge towards the central node, and completing the optimization of the static UWSNs routing protocol based on convergence guidance.
[0010] Further, a preferred implementation is provided. In the Q-Learning based forwarding decision model described in Step 1, all nodes share a set of reinforcement learning parameters.
[0011] Further, a preferred implementation is provided. In Step 2, the data packet is realized through a decision model with nodes as agents.
[0012] Further, a preferred implementation is provided. After the static UWSNs routing protocol is deployed, it further includes the steps of SAOL routing protocol design, and the SAOL routing protocol design includes steps of neighbor discovery, cluster formation, and data forwarding.
[0013] Further, a preferred implementation is provided. The convergence guidance mechanism described in Step 3 includes steps of reward function design and fixed V-value mechanism.
[0014] Further, a preferred implementation is provided. The method for designing the reward function is as follows:
[0015] R = R sink + R energy + R drift (4 - 4)
[0016] Where R is the return of the current decision; R sink is the return of the convergence factor; R energy is the return of the energy factor; R drift is the return of the drift factor.
[0017] Further, a preferred implementation is provided. The calculation method of the fixed V-value mechanism in Step 3 is as follows:
[0018]
[0019] In the formula, V π (s) represents the expectation of the total return that can be obtained in state s according to the π policy; Q π (s, a) represents the evaluation of the current decision-making action, and α is the learning rate of the system, 0 < α ≤ 1.
[0020] Solution 2: Optimization method for static UWSNs routing protocol based on convergence guidance. The method includes the following steps:
[0021] The forwarding path determination module is used to establish the best forwarding path from the source node to the sink node based on the Q-Learning-based forwarding decision model;
[0022] The forwarding decision model construction module is used to classify the forwarding decision model described in the forwarding path determination module into two categories based on the static UWSNs routing protocol according to whether the parameters between nodes in the static UWSNs routing protocol are completely shared, including a decision model with data packets as agents and a decision model with nodes as agents;
[0023] Determine the decision state of the information of the data packet location node and the forwarding candidate node in the static UWSNs routing protocol. When the node needs to forward a data packet, calculate the Q value of each neighbor node in turn, and select the neighbor node with the largest Q value as the next-hop forwarding target;
[0024] The static UWSNs routing protocol optimization module is used to construct a convergence guidance mechanism. The convergence guidance mechanism allows nodes to generate training rewards based on the location information of the sink node during the Q-Learning training process, generates a path for data packets to gather towards the central node, and completes the optimization of the static UWSNs routing protocol based on convergence guidance.
[0025] Solution 3: A computer device, including a memory and a processor, characterized in that a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes the method described in any one of Solution 1.
[0026] Solution 4: A computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of Solution 1 are implemented.
[0027] The beneficial effects of the present invention are as follows:
[0028] For the optimization method of the static UWSNs routing protocol based on convergence guidance described in the present invention, a forwarding target selection strategy is designed based on Q-Learning, enabling nodes to gradually learn the forwarding potential of current neighbor nodes during continuous forwarding, and subsequent forwarding can be completed at a lower energy consumption and delay cost. At the same time, in order to accelerate the convergence speed and reduce the probability of detours and loops during training, a convergence guidance mechanism is proposed. In the mechanism, the design of the reward function is affected by the location of the sink node, so that nodes in the training try to forward more in the direction of the sink; at the same time, the V value of the sink node is fixed, so that nodes closer to the sink can converge earlier, thus accelerating the speed of routing training. The simulation experiment results show that the proposed protocol can effectively optimize the routing detour and routing loop phenomena, and shows good performance in improving the delivery rate, reducing the end-to-end delay, and reducing the forwarding energy tax.
[0029] In view of the problems of high delay caused by forwarding detours and low delivery rate caused by forwarding loops in the routing protocol of static topology UWSNs, a routing protocol for static UWSNs based on sink guidance is proposed. An active routing protocol is used to control the forwarding link, decision selection is carried out through Q-Learning training, and a sink guidance mechanism is introduced to accelerate training convergence while optimizing routing detours and loops, so as to optimize the delivery rate and delay performance of the routing protocol for static topology UWSNs.
[0030] The present invention is also applicable to the fields of marine environmental monitoring and marine resource exploration. Description of the Drawings
[0031] Figure 1 Schematic diagram of the drift of the anchor node described in Embodiment XI.
[0032] Figure 2 Schematic diagram of the influence of the fixed V value on forwarding convergence described in Embodiment XI.
[0033] Figure 3 Schematic diagram of the packet structure of the SAQL routing protocol described in Embodiment XI.
[0034] Figure 4 Flowchart of the operation of the SAQL routing protocol described in Embodiment XI.
[0035] Figure 5 Flowchart of the neighbor discovery operation of the SAQL routing protocol described in Embodiment XI.
[0036] Figure 6 Flowchart of the data forwarding operation of the SAQL routing protocol described in Embodiment XI.
[0037] Figure 7 Schematic diagram of the update of the V value of the nodes in a certain forwarding link in the simulation described in Embodiment XI.
[0038] Figure 8 Schematic diagram of the statistical result of the packet delivery rate described in Embodiment XI.
[0039] Figure 9 Schematic diagram of the statistical result of the energy tax described in Embodiment XI.
[0040] Figure 10 Schematic diagram of the statistical result of the end-to-end delay described in Embodiment XI.
[0041] Figure 11 Schematic diagram of the influence of the number of source nodes on the delivery rate described in Embodiment XI.
[0042] Figure 12Schematic diagram of the influence of the number of source nodes on the energy tax described in Embodiment XI.
[0043] Figure 13 Schematic diagram of the influence of the number of source nodes on the end-to-end delay described in Embodiment XI. Specific implementation mode
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application.
[0045] Embodiment 1. This embodiment provides an optimization method for a static UWSNs routing protocol based on aggregation guidance. The method includes the following steps:
[0046] Step 1: Based on the forwarding decision model of Q-Learning, establish the best forwarding path from the source node to the sink node.
[0047] Step 2: According to whether there are completely shared parameters between nodes in the static UWSNs routing protocol, divide the forwarding decision model in Step 1 into two categories, including a decision model with data packets as agents and a decision model with nodes as agents;
[0048] Determine the decision state of the information of the node where the data packet is located and the forwarding candidate nodes in the static UWSNs routing protocol. When the node needs to forward the data packet, calculate the Q values of each neighbor node in turn, and finally select the neighbor with the largest Q value as the next-hop forwarding target.
[0049] Step 3: Construct an aggregation guidance mechanism. The aggregation guidance mechanism allows nodes to generate training rewards based on the location information of each node in the aggregation during the training process of the Q-Learning forwarding decision model, generating a path for data packets to gather towards the central node, and completing the optimization of the static UWSNs routing protocol based on aggregation guidance.
[0050] Embodiment 2. This embodiment further limits the optimization method for the static UWSNs routing protocol based on aggregation guidance described in Embodiment 1. In the forwarding decision model of Q-Learning described in Step 1, all nodes share a set of reinforcement learning parameters.
[0051] Embodiment 3. This embodiment further limits the optimization method for the static UWSNs routing protocol based on aggregation guidance described in Embodiment 1. In Step 2, the data packet is implemented through a decision model with nodes as agents.
[0052] Embodiment 4. This embodiment further limits the optimization method of the static UWSNs routing protocol based on aggregation guidance described in Embodiment 1. After the static UWSNs routing protocol is deployed, it further includes the steps of SAOL routing protocol design, and the SAOL routing protocol design includes the steps of neighbor discovery, cluster formation, and data forwarding.
[0053] Embodiment 5. This embodiment further limits the optimization method of the static UWSNs routing protocol based on aggregation guidance described in Embodiment 1. The aggregation guidance mechanism in step 3 includes the steps of reward function design and fixed V - value mechanism.
[0054] Embodiment 6. This embodiment further limits the optimization method of the static UWSNs routing protocol based on aggregation guidance described in Embodiment 1. The method for designing the reward function is as follows:
[0055] R = R sink +R energy +R drift (4 - 4)
[0056] where R is the return of the current decision; R sink is the return of the aggregation factor; R energy is the return of the energy factor; R drift is the return of the drift factor.
[0057] Embodiment 7. This embodiment further limits the optimization method of the static UWSNs routing protocol based on aggregation guidance described in Embodiment 1. The calculation method of the fixed V - value mechanism in step 3 is as follows:
[0058]
[0059] In the formula, V π (s) represents the expectation of the total return that can be obtained in state s according to the π policy; Q π (s,a) represents the evaluation of the current decision - making action, and α is the learning rate of the system, 0 < α ≤ 1.
[0060] Embodiment 8. This embodiment proposes an optimization method of a static UWSNs routing protocol based on aggregation guidance. The method includes the following steps:
[0061] A forwarding path determination module, which is used to establish the best forwarding path from the source node to the sink node based on the Q - Learning - based forwarding decision model;
[0062] The forwarding decision model construction module is used to classify the forwarding decision models described in the forwarding path determination module into two categories based on the static UWSNs routing protocol and whether the parameters are fully shared among nodes in the static UWSNs routing protocol, including a decision model with data packets as agents and a decision model with nodes as agents;
[0063] Determine the decision state of the information of the node where the data packet is located and the forwarding candidate nodes in the static UWSNs routing protocol. When the node needs to forward the data packet, calculate the Q-values of each neighbor node in turn, and finally select the neighbor with the largest Q-value as the next-hop forwarding target;
[0064] The static UWSNs routing protocol optimization module is used to construct a convergence guidance mechanism. The convergence guidance mechanism allows nodes to generate training rewards based on the location information of the sink node during the Q-Learning training process, generating a path for data packets to gather towards the central node, and completing the optimization of the static UWSNs routing protocol based on convergence guidance.
[0065] Embodiment 8. This embodiment proposes an optimization system for a static UWSNs routing protocol based on convergence guidance. The system includes:
[0066] The forwarding path determination module is used to establish the best forwarding path from the source node to the sink node based on the forwarding decision model of Q-Learning;
[0067] The forwarding decision model construction module is used to classify the forwarding decision models described in the forwarding path determination module into two categories based on the static UWSNs routing protocol and whether the parameters are fully shared among nodes in the static UWSNs routing protocol, including a decision model with data packets as agents and a decision model with nodes as agents;
[0068] Determine the decision state of the information of the node where the data packet is located and the forwarding candidate nodes in the static UWSNs routing protocol. When the node needs to forward the data packet, calculate the Q-values of each neighbor node in turn, and finally select the neighbor with the largest Q-value as the next-hop forwarding target;
[0069] The static UWSNs routing protocol optimization module is used to construct a convergence guidance mechanism. The convergence guidance mechanism allows nodes to generate training rewards based on the location information of the sink node during the Q-Learning training process, generating a path for data packets to gather towards the central node, and completing the optimization of the static UWSNs routing protocol based on convergence guidance.
[0070] Embodiment 9. This embodiment provides a computer device, including a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the method described in any one of Embodiments 1 to 7.
[0071] Embodiment 10. This embodiment provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the method described in any one of Embodiments 1 to 7 are implemented.
[0072] Embodiment 11. This embodiment presents an example for explaining the above embodiments. The example is specifically as follows:
[0073] Refer to Figures 1 to 13 To illustrate this embodiment, this embodiment focuses on the research of the optimization of routing protocols in the scenario of static topology UWSNs. Aiming at the problems of high end-to-end delay and low delivery rate caused by routing detours and routing loops in the static topology UWSNs routing protocol, this embodiment proposes a novel underwater routing algorithm based on sink guidance and Q-Learning. Based on Q-Learning for forwarding decisions, a sink guidance mechanism is introduced to reduce the probability that proactive routing encounters detours or even falls into an infinite loop, so as to improve the delivery rate of multi-hop forwarding and reduce the end-to-end delay.
[0074] Specifically, it includes:
[0075] 1. Forwarding decision-making model based on Q-Learning
[0076] In the design scenario of the static UWSNs routing protocol, according to whether the parameters are completely shared among nodes, the forwarding decision-making model can be divided into two categories: a decision-making model with a data packet as an agent and a decision-making model with a node as an agent. When using the data packet to be forwarded as an agent, the routing decision of the data packet is similar to the "maze walking" model. The node where the data arrives is the current state, and the node to which the data packet will be sent is the action. This model requires all nodes to share a set of reinforcement learning parameters, and the application scenario of this model underwater is very limited. When using the node itself as an agent, each node in UWSNs is an agent. At this time, the local information and neighbor table of the node are the agent states, and the next-hop forwarding target is the decision-making action. This model does not require nodes to share RL parameters before, and can adapt to more underwater working scenarios.
[0077] The routing protocol based on sink attraction and Q-Learning decision-making (SinkAttraction-Q-Learning Routing Protocol, SAQL Routing Protocol) proposed in this embodiment is carried out with nodes as agents. In the protocol, the information of the node where the data packet is located and the forwarding candidate nodes (such as remaining energy, number of neighboring nodes, depth, and location, etc.) are the decision-making states of the agents. When the node needs to forward a data packet, it calculates the Q values of each neighboring node in turn, and finally selects the neighboring node with the largest Q value as the next-hop forwarding target.
[0078] 2. Construction of forwarding candidate set
[0079] To speed up the decision-making process, first exclude the neighbors with relatively poor forwarding performance in the neighbor table. These neighbors are used as the backup forwarding set, and the Q values of the neighbor nodes in this backup set will be calculated only when all forwarding attempts fail.
[0080] The construction of the node forwarding candidate set mainly follows two principles: energy and depth.
[0081]
[0082] Among them, E ne is the remaining energy of the neighboring node; E ave is the average remaining energy of the current neighborhood; k is the energy screening coefficient. The larger k is, the fewer candidate nodes there are, which is more suitable for relatively dense networks. The smaller k is, the more candidate nodes there are, which is more suitable for relatively sparse networks; Dep ne is the depth of the neighboring node; Dep me is the depth of this node. This screening condition can select nodes with higher energy and shallower depth (the sink node is generally deployed on the water surface, and the data forwarding direction is generally towards the water surface) to form the forwarding candidate set, which reduces the action space of the current Q-Learning decision-making, helps to increase the probability of selecting nodes with strong forwarding potential during training, and speeds up the training convergence.
[0083] 3. Q function update
[0084] Q-Learning is a value-oriented reinforcement learning algorithm based on the Markov Decision Process (MDP) framework. It aims to evaluate the value of an agent's decision-making attempts and train the agent's action strategy in the current scenario by rewarding or punishing the agent's actions. The core of Q-Learning is to train the agent to learn the action-value function (Q-function), which assigns a value to the state of the current scenario and each action of the agent, representing the expected return that the agent can obtain by performing that action in the current state. The Bellman equation is the core of the update of the action-value function, which describes the iterative relationship between the current state and the future state of the agent, reveals the key law of estimating the value of one state from the value of another state, reflects the connection between the current decision and future rewards, and enables the agent to continuously adjust the strategy to estimate future rewards to achieve the optimal decision.
[0085] Like most intelligent algorithms, Q-Learning mainly uses the balance between exploration and exploitation to learn the optimal strategy. In Q-Learning, an agent's decision is split into two aspects: the evaluation of the agent's decision-making action is called the Q-function, and the evaluation of the agent's state is called the value function. Their expressions are shown in (4-2):
[0086]
[0087] Equation (4-2) is a representation of the Bellman equation, where V π (s) represents the expected total return that can be obtained in state s according to the π strategy; Q π (s,a) represents the return that can be obtained after performing action a in state s and maintaining this strategy until the final state, that is, the evaluation R(s,a) of the current decision-making action is the immediate return of choosing action a in state s. Based on Equation (4-2), the method for updating the Q value in Q-learning can be obtained:
[0088] Q'(s,a)←(1-α)Q(s,a)+α(R(s,a)+γ×maxQ(s',a))(4-3)
[0089] Among them, Q'(s,a) represents the updated Q value; α is the learning rate of the system, where 0 < α ≤ 1. If α = 1, it indicates that the update of the Q value depends entirely on the latest decision. When α < 1, the calculation of the new Q value will be affected by the past Q value calculations, which helps to smooth the training process. By continuously adjusting the action evaluation method based on the rewards of the current decision and the evaluation of future rewards, the agent can gradually learn which actions to take in a specific state to obtain the best long-term rewards, so as to learn the optimal strategy for solving problems. That is, in the continuous update process, each node gradually learns the current forwarding potential of each neighbor node, which helps to find the best forwarding path.
[0090] 4. Convergence Guidance Mechanism
[0091] In the classic Q-Learning routing algorithm, routing detours or routing loops are likely to occur, which affects the convergence speed of the algorithm. And the training of the algorithm is real-time. The slower the convergence speed, the higher the network delivery rate, end-to-end delay, and even energy tax cost will be. Therefore, it is necessary to guide the algorithm training to increase the probability that the agent selects better nodes and accelerate the network convergence speed, thereby reducing the energy and delay costs in training. The convergence guidance mechanism mainly works in two aspects: reward function design and fixed V value mechanism.
[0092] 5. Reward Function Design
[0093] The reward function defines the immediate reward or punishment that the agent will obtain after taking an action in the current state. It is the direct feedback of the environment to the agent's decision and has a great impact on whether the agent can learn the optimal strategy. To optimize the energy consumption balance, forwarding direction, and link stability of the decision-making, the calculation method of the reward function of the SAQL protocol is shown in Equation (4-4).
[0094] R = R sink + R energy + R drift (4-4)
[0095] Where R is the reward of the current decision; R sink is the reward of the convergence factor; R energy is the reward of the energy factor; R drift is the reward of the drift factor.
[0096] 6. Convergence Factor
[0097] The convergence factor aims to guide the forwarding direction of the algorithm. Specifically, based on the positioning information, this factor calculates the forward forwarding distance that can be achieved by forwarding the data packet to the current candidate node, which affects the node's decision-making. Generally speaking, the larger this forward distance is, the closer the current forwarding is to the sink node, and the closer the formed forwarding link is to the shortest forwarding link, which helps to achieve a smaller end-to-end delay. The calculation of the convergence guidance reward is shown in Equation (4-5).
[0098]
[0099] Among them, (x i , y i , z i ) represents the coordinates of the currently decision-making node; (x next , y next , z next ) represents the coordinates of the currently estimated next-hop forwarding candidate node; (x sink , y sink , z sink ) represents the coordinates of the sink node pre-recorded by the current node (in the case of multiple sink nodes, it is the coordinates of the nearest sink node). Equation (4-5) actually calculates a normalized forward forwarding distance, indicating the degree to which the data packet is closer to the sink node after being forwarded to this candidate node. When the candidate forwarder is farther from the sink node than itself, R sink is calculated as a negative number, which is a penalty for the agent's decision-making and helps to prevent the agent from forwarding along a long way or even falling into a forwarding loop.
[0100] 7. Energy factor
[0101] After underwater nodes are deployed, it is often difficult to charge them in time. Nodes should be avoided from failing prematurely as much as possible. When the energy of the forwarding node is low, try to let the nodes with higher remaining power in the surrounding area undertake this forwarding. In the clustering mode, the energy consumption of the cluster head node itself is slightly higher than that of the member nodes. When the next-hop forwarder happens to be the cluster head, it often causes excessive energy consumption of the forwarding cluster head, which is not conducive to maintaining the current clustering for a long time. Therefore, non-cluster head nodes can also be selected into the candidate forwarding set during forwarding, and the final forwarding is determined according to the remaining energy of the nodes. The calculation method of the energy factor reward function is as shown in (4-6).
[0102]
[0103] Among them, g = -1 is the permanent consumption. Since the node will consume energy whether the forwarding is successful or not, the permanent consumption is introduced to represent that any action attempt requires energy consumption; E next , E i and E initial represent the remaining energy of the candidate forwarding node, the remaining energy of the decision-making node, and the initial energy of the node respectively; EDnext and ED i respectively represent the energy distribution of the candidate forwarding node neighborhood, that is, the average energy within their neighborhoods; α 1 and α 2 respectively represent the weight of the remaining energy impact and the weight of the energy distribution impact, and the two satisfy α 1 +α 2 = 1, α 1 , α 1 ∈[0, 1].
[0104] 8. Drift factor
[0105] UWSNs nodes are affected by changing water currents in water. Even when deployed in an anchored manner, the nodes will drift within a certain range. When the distance between two nodes is far, this drift is likely to cause the forwarding link to be intermittent. This disruption not only affects this transmission but also forces the forwarding node to re-explore new candidate nodes, introducing additional energy consumption. Therefore, a drift factor is introduced to represent the evaluation reward of the residence time of neighbor nodes in the neighborhood.
[0106]
[0107] where d max represents the maximum distance of node communication; d represents the distance between the current candidate forwarding node and this node before, and this distance can be calculated through the Received Signal Strength Indicator (RSSI); h represents the depth difference between the current candidate forwarding node and this node. The depth information of the node itself can be obtained through a pressure sensor, and the depth information of the neighbor node is recorded and shared in the neighbor detection data packet. According to Equation (4-7), when the distance between two nodes is farther and the depth difference is larger, the candidate node is more likely to drift out of the range, so its drift reward is smaller. Such a node is relatively less likely to be preferentially selected as the next-hop node, which helps to optimize the stability of the forwarding path obtained by the decision.
[0108] 9. Fixed V-value mechanism
[0109] The V-value represents the expectation of the cumulative forwarding reward that a certain node can achieve under the current forwarding strategy, and it also has a crucial impact on the Q-Learning decision. Based on the conversion relationship between the V-value and the Q-value expressed by Equation (4-8), the convergence guidance mechanism sets the V-value of all sink nodes to be 0 forever to guide the direction for the agent to forward data packets. By fixing the V-value of the sink node to 0, the forwarding node closest to the sink node will converge the training faster, and then gradually spread to the data packet source node, and a forwarding path can be formed faster.
[0110]
[0111] 10. Packet Design
[0112] After the deployment of UWSNs, due to energy and channel resource constraints, it is difficult for nodes to maintain neighbor information efficiently. Therefore, the SAQL protocol adopts an efficient strategy, that is, combining forwarding and information sharing into one. Each data forwarding carries local information, and any node receiving the packet can update its neighbor information accordingly. The core of this process lies in the designed packet structure, which not only carries data but also the node's status information, such as battery level and location. The detailed design of the packet structure is shown as Figure 3 shown
[0113] When the network load is light and the packet forwarding activity between nodes decreases, there may be a situation where a node urgently needs to forward data while its neighbor information is outdated. To address this situation, the SAQL protocol retains a low-frequency information broadcast mechanism. Even when network activity is low, it periodically broadcasts information packets through the network to update and maintain neighbor information. The destination node address and the next-hop node address of these special broadcast packets are both set to the broadcast address, clearly specifying that the only purpose of these packets is to enforce neighbor maintenance, ensuring that even when the network is less active, the neighbor information of each node can maintain a certain degree of validity, guaranteeing the stability and reliability of the network.
[0114] 11. Protocol Workflow
[0115] The workflow of the SAQL routing protocol mainly consists of three key stages: neighbor discovery, cluster formation, and data forwarding. In the neighbor discovery stage, nodes ensure mutual recognition and perception among nodes through forced local information broadcasting. This stage is particularly crucial during the initial network deployment and needs to be repeated periodically during subsequent operation to maintain the update of the node neighbor table, ensuring that node information remains timely even in the absence of data transmission. The cluster formation stage aims to optimize energy consumption balance through clustering and perform local data integration, reducing the occupancy of channel resources. This not only helps to maintain the coverage area during network operation but also reduces the possibility of packet collisions to a certain extent. The specific content has been elaborated in Chapter 3. The data forwarding stage is the core of the protocol, responsible for realizing the effective forwarding of packets among network nodes to ensure the smooth transmission of information.
[0116] 12. Neighbor Discovery
[0117] In the initial deployment stage of UWSNs, nodes lack information about neighboring nodes in their vicinity, resulting in the inability to construct an effective candidate forwarding set when data transmission is required. Similarly, if a certain area of the network remains inactive for a long time, the neighbor information of nodes cannot be updated, and the established forwarding candidate set will become unreliable. To solve this problem, the solution we proposed adopts a periodic information sharing mechanism. This mechanism allows all nodes to time according to an internal timer. Once the predetermined neighbor information maintenance period is reached, the nodes will suspend the current task, actively send broadcast data packets containing local information, and refresh their neighbor information based on the received data packets. The workflow of this stage is as Figure 5 shown.
[0118] To reduce the energy consumption cost of maintaining neighbor information in a high-load network environment, the SAQL protocol adopts a strategy of sharing information incidentally during data packet forwarding, enabling each node receiving a data packet to directly obtain the data required for maintaining neighbor information from the packet header, maintaining the timeliness of the neighbor information of nodes, and at the same time reducing the frequency of forced neighbor discovery operations to save energy consumption.
[0119] 13. Data Forwarding
[0120] The data packet forwarding stage is the core part of the routing protocol. In this stage, each node first sends the data packet to the cluster head node of its affiliated cluster. After the cluster head node preliminarily reduces the weight of the data, it becomes the only "source node" in the current neighborhood, and then constructs a forwarding candidate set, where each neighbor node is regarded as an action that the SAQL algorithm may take in the current state. By evaluating the Q values of each node, a sorting of the forwarding priorities of the nodes can be obtained, and the nodes forward the data according to this sorting, and use the next forwarding of the data packet by the target node as the transmission confirmation signal.
[0121] When constructing the node forwarding candidate set, two main factors are considered: energy and forwarding history. First, the node searches in the neighbor table according to the target address (i.e., the sink node address) information contained in the data packet. If the target node is found, it will be used as the first choice of the forwarding candidate set and given the highest priority. Then, the node calculates the average remaining power of all nodes in the neighbor table and excludes the neighbors whose remaining power does not meet the forwarding conditions. Finally, the node excludes the forwarding node of the previous hop of the data packet and the nodes that have tried to forward but failed from the candidate set.
[0122] After the candidate set is constructed, the nodes will attempt to forward data in order of priority. Each time a node sends a data packet to the next-hop destination, a local timer is started, and its duration is calculated as shown in (4-9). During this period, the node needs to listen to the forwarding behavior of the destination node. If it detects that the destination node successfully forwards the data packet further, this forwarding is considered successful. Otherwise, it is regarded as a forwarding failure, and the node will then attempt other candidate nodes. If the entire candidate set fails to forward the data packet successfully, the node will make a final attempt, that is, consider the nodes previously excluded due to power reasons as alternatives in the hope of repairing the forwarding link.
[0123]
[0124] where R represents the communication radius of the node; v water represents the acoustic wave speed in the current water area of the node, generally 1500 m / s; L pkt represents the data packet length; S represents the data rate of transmission and reception.
[0125] 14. Simulation Analysis
[0126] To verify the performance of the proposed SAQL routing protocol in terms of delivery ratio, delay, and energy consumption, the protocol was simulated and verified in NS3, and the detailed simulation parameters are shown in Table 4.1.
[0127] Table 4.1 Simulation Parameters of SAQL Algorithm
[0128]
[0129] Figure 7 The following shows the update situation of the V values of each node on a certain forwarding link in the simulation. Among them, node 148 is the source node, node 69 is the first-hop node, node 56 is the second-hop node, and node 64 is the third-hop node. The source node has converged after about two hundred training sessions, and its V value quickly converges to around -2.5. Since the source node continuously generates and sends data, its remaining energy becomes lower and lower, so its V value will still decrease slowly after convergence. For each relay node, its V value is also updated quickly and then decreases slowly as energy is consumed. The closer the node is to the sink, the larger its V value, which guides the selection of the next-hop forwarding target.
[0130] Figure 8Shows the performance of the proposed protocol in terms of packet delivery ratio in NS3 simulation. The simulation results show that the proposed protocol has a leading delivery ratio in both sparse and dense node scenarios. And when the number of nodes is sufficient, the delivery ratio of the proposed algorithm remains above 90% when the number of nodes is more than 70, and when the number of nodes increases to 100 or more, the delivery ratio can always be higher than 95%, which is sufficient to meet the application requirements. Although FVBF has a slightly higher PDR when the number of nodes is larger, its flooding transmission has a large energy cost, details can be seen in Figure 9 .
[0131] Figure 9 Shows the energy tax performance of each algorithm in NS3 simulation tests. The simulation statistical results show that the SAQL protocol has the lowest forwarding energy tax in most cases. This is because the proposed protocol adopts a single forwarding link mechanism, and only one target can undertake the forwarding task for each forwarding, so the energy tax is relatively low. As the number of nodes increases, the number of nodes that can overhear the packet in each forwarding increases, so the forwarding energy tax increases. For QELAR and QLFR, in the case of sparse nodes, their forwarding delivery ratios are relatively low. QELAR will perform multiple forwarding repairs, and QLFR will increase the number of forwarding links, so their energy taxes are very high. For the FVBF and ALRP protocols, both of them are opportunistic forwarding with many redundant links. As the number of nodes increases, the number of nodes participating in forwarding increases, so the energy tax also increases significantly.
[0132] Figure 10 Shows the end-to-end delay performance of each protocol in NS3 simulation tests. According to the statistical results, the proposed protocol has the best end-to-end delay. This is because the convergence guidance mechanism can guide nodes to forward towards the target closer to the sink, which helps to shorten the length of the forwarding link, so it has a better end-to-end delay. Another protocol with end-to-end delay performance close to this algorithm is the QLFR protocol, which adopts a multi-link forwarding mechanism, so there is also a high probability of finding a shorter forwarding path.
[0133] Figure 11It shows the impact of the number of source nodes at the same time on the protocol delivery ratio in the NS3 simulation test. As shown in the figure, due to the differences in various protocol mechanisms and modes, for example, this protocol and the QELAR and QLFR protocols require a certain feedback mechanism to update the decision-making strategy, while several other protocols do not require handshakes. To prevent additional factors from interfering with the protocol, the MAC protocol used in the simulation is the BroadcastMAC protocol, which does not contain any handshake and feedback mechanisms and has the least interference with the routing protocol. However, its collision avoidance effect is also very poor. Therefore, as the number of source nodes increases, the delivery ratios of all protocols show a downward trend. Generally speaking, this protocol and the FVBF protocol still maintain the best delivery ratios. The PDR of this protocol drops more severely than that of the FVBF because this protocol is mainly based on a single communication link. Therefore, the redundant backup during transmission is not as good as the opportunistic protocol FVBF. Nevertheless, this protocol has better PDR performance compared with the similar QELAR and QLFR protocols.
[0134] Figure 12 It shows the impact of the number of source nodes at the same time on the protocol energy tax in the NS3 simulation test. As shown in the figure, as the number of source nodes at the same time increases, there is a certain trend of increasing energy tax for all protocols. However, this protocol still maintains the lowest packet energy tax, demonstrating the good energy-saving performance of this protocol. Among them, the performance changes of most protocols are relatively small, and only the QELAR protocol shows serious deterioration. This is mainly because the collision avoidance effect of the MAC protocol used is relatively poor, and the QELAR itself has not been optimized for the compatibility of multi-link interference. Therefore, as the number of source nodes increases, its delivery ratio drops significantly, leading to serious deterioration of its packet delivery ratio.
[0135] Figure 13 It shows the impact of the number of source nodes at the same time on the protocol end-to-end delay in the NS3 simulation test. As shown in the figure, the increase in the number of source nodes results in two changing trends in the end-to-end delays of several protocols. Among them, this protocol, FVBF, and QELAR show an increasing trend in delay, which may be due to the additional time brought by the protocol for link repair after packet collision. While the HHVBF, QLFR, and ALRP protocols show a basically unchanged trend, which may be because the QLFR itself is a multi-link transmission, and the opportunistic forwarding of the HHVBF and ALRP themselves has a flooding nature. Therefore, as long as their communication links still exist, their delays are relatively stable.
[0136] In summary, the proposed protocol can balance the three metrics of delivery ratio, end-to-end delay, and forwarding energy tax in static UWSNs. Although it is not the best in terms of delivery ratio among all algorithms, it can quickly reach over 90% when the node density is appropriate. Due to the influence of the convergence guidance mechanism, the SAQL protocol is more inclined to select the next-hop node closer to the sink, thus greatly increasing the probability of finding the shortest forwarding link, significantly optimizing the problem of routing detours and even routing loops, and reducing the end-to-end delay. Finally, in static UWSNs, the forwarding link can usually maintain a longer effective time. Therefore, the single-forwarding-link method adopted by this protocol can reduce the number of redundant transmissions, thereby reducing the routing energy consumption and having a lower forwarding energy tax.
[0137] Those skilled in the art can understand that the above description is only the preferred embodiment of the present invention. The features described in each embodiment and / or claim of the present disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. It is not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
[0138] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A static UWSNs routing protocol optimization method based on convergence guidance, characterized in that: The method comprises the following steps: Step 1: Based on the Q-Learning forwarding decision model, establish the best forwarding path from the source node to the sink node; Step 2: According to whether there are fully shared parameters between nodes in the static UWSNs routing protocol, the forwarding decision model described in step 1 is divided into two categories, including a decision model with data packets as intelligent agents and a decision model with nodes as intelligent agents; Determine the decision status of the node where the data packet is located and the forwarding candidate node information in the static UWSNs routing protocol. When the node needs to forward the data packet, calculate the Q value of each neighboring node in turn, and finally select the neighboring node with the largest Q value as the next hop forwarding target; Step 3: Construct a convergence guidance mechanism. The convergence guidance mechanism allows the nodes in the Q-Learning forwarding decision model to generate training rewards based on the location information of each node during the training process, generate a path for data packets to converge to the central node, and complete the static UWSNs routing protocol optimization based on convergence guidance.
2. According to claim 1, the static UWSNs routing protocol optimization method based on convergence guidance is characterized in that: All nodes in the Q-Learning-based forwarding decision model described in step 1 share a set of reinforcement learning parameters.
3. According to claim 1, the static UWSNs routing protocol optimization method based on convergence guidance is characterized in that: In step 2, the data packet is implemented through a decision-making model with nodes as intelligent agents.
4. According to claim 3, the static UWSNs routing protocol optimization method based on convergence guidance is characterized in that: After the static UWSNs routing protocol is deployed, the step of SAOL routing protocol design is also included. The SAOL routing protocol design includes the steps of neighbor discovery, cluster formation and data forwarding.
5. The static UWSNs routing protocol optimization method based on convergence guidance according to claim 1 is characterized in that: The convergence guidance mechanism described in step 3 includes the steps of reward function design and fixed V value mechanism.
6. The static UWSNs routing protocol optimization method based on convergence guidance according to claim 5 is characterized in that: The reward function design method is: R=R sink +R energy +R drift (4-4) Where R is the reward of the current decision; R sink is the pooled factor return; R energy is the energy factor return; R drift is the drift factor return.
7. The static UWSNs routing protocol optimization method based on convergence guidance according to claim 5 is characterized in that: The calculation method for the fixed V value mechanism described in step 3 is: Where V π (s) represents the expected total return that can be obtained according to the π strategy in the s state; Q π (s,a) represents the evaluation of the current decision action, α is the learning rate of the system, 0<α≤1.
8. The static UWSNs routing protocol optimization system based on convergence guidance is characterized by: The system comprises: The forwarding path determination module is used to establish the best forwarding path from the source node to the sink node based on the Q-Learning forwarding decision model; A forwarding decision model building module is used to divide the forwarding decision model described in the forwarding path determination module into two categories, including a decision model with a data packet as an intelligent agent and a decision model with a node as an intelligent agent, based on a static UWSNs routing protocol and whether the parameters are fully shared between nodes in the static UWSNs routing protocol; Determine the decision status of the node where the data packet is located and the forwarding candidate node information in the static UWSNs routing protocol. When the node needs to forward the data packet, calculate the Q value of each neighboring node in turn, and finally select the neighboring node with the largest Q value as the next hop forwarding target; The static UWSNs routing protocol optimization module is used to construct a convergence guidance mechanism. The convergence guidance mechanism allows nodes to generate training rewards based on the location information of the convergence node during the Q-Learning training process, generate a path for data packets to converge to the central node, and complete the static UWSNs routing protocol optimization based on convergence guidance.
9. A computer device comprising a memory and a processor, characterized in that A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-agent reinforcement learning routing algorithm based on geographic position
CN112804726A
Method for designing opportunistic routing protocol of underwater acoustic sensor network
CN118474014A
Distributed collaborative evolution method, UAV and intelligent routing method therefor, and apparatus
WO2024021281A1
Cited By
Underwater sensor network efficient adaptive routing protocol system based on Q-learning and Bayesian optimization
CN121547827A
An underwater sensor network efficient adaptive routing protocol system based on Q-learning and bayesian optimization
CN121547827B