An underwater sensor network efficient adaptive routing protocol system based on Q-learning and bayesian optimization
By employing an adaptive routing protocol optimized by Q-learning and Bayesian methods, cluster heads are dynamically elected, optimizing the energy consumption and path selection of underwater sensor networks. This solves the problems of high energy consumption and uneven cluster head distribution in traditional routing protocols, thereby improving network lifetime and communication reliability.
Patent Information
- Application Number
- CN202610064218.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-19
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2046-01-19
AI Technical Summary
Traditional underwater wireless sensor network routing protocols are energy-intensive, have uneven cluster head distribution, and are not suitable for large-scale networks, resulting in insufficient network lifespan and communication reliability.
An adaptive routing protocol based on Q-learning and Bayesian optimization is adopted. Through intra-cluster and inter-cluster routing design, combined with Q-value tables and Bayesian optimization, cluster heads are dynamically elected to optimize network energy consumption and path selection.
It achieves efficient energy management of underwater sensor networks, improves network lifetime and communication reliability, adapts to different network conditions, and continuously converges to the global optimal performance state.
Smart Images

Figure CN121547827B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of underwater sensor network technology, and in particular to an efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization. Background Technology
[0002] As a critical infrastructure for applications such as marine monitoring and resource exploration, underwater wireless sensor networks rely heavily on energy-efficient routing protocol design for their network lifetime and communication reliability. Traditional low-energy adaptive clustering hierarchical protocols balance energy consumption to some extent by randomly rotating cluster heads, but their randomness leads to uneven distribution of cluster heads, and the single-hop communication mode between cluster heads and sink nodes is not suitable for large-scale networks, resulting in still relatively high energy consumption. Summary of the Invention
[0003] To address at least one of the aforementioned technical problems, the present invention aims to provide an efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization.
[0004] This invention includes an efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization. The efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization includes: At least one aggregation node; Multiple underwater nodes; each underwater node is deployed within a three-dimensional underwater monitoring area, each underwater node maintains its own first Q-value table, constructs a cluster structure based on each first Q-value table, performs intra-cluster routing in the cluster structure, performs Bayesian optimization on the first Q-value table, the convergence node and each underwater node maintain a common second Q-value table, and performs inter-cluster routing based on the second Q-value table in the cluster structure.
[0005] Furthermore, each of the underwater nodes maintains its own first Q-value table, including: For any of the underwater nodes: The underwater node senses its own state characteristics and the average distance to its local neighbors. and mobility stability The state characteristics include the remaining energy ratio. Distance from the aggregation node and normalized node degree ; Each component in the state feature is discretized to obtain the state index. ; Establish the first Q-value table; the first Q-value table includes the discrete level number of each component in the state feature and the corresponding first Q-value; the first Q-value is mapped to the action space, which includes campaign actions and non-campaign actions; According to the formula
[0006]
[0007] Calculate the quality score corresponding to the underwater node. ;in, For topological quality factor, For energy topological weights, The remaining energy is weighted first. The average distance to neighbors is the first weighted factor. The first weight is the node degree; Whenever a successful cluster head election is detected, according to the formula
[0008] Calculate the instant reward corresponding to the underwater node. ; Whenever an event indicating that the cluster structure construction is complete is detected, according to the formula... =2.5+ 1.5- 0.2+ 0.5- 5. If the underwater node participates in the election and is elected as the cluster head node =0.5+0.8 If the underwater node did not participate in the election and successfully joined a cluster =-0.8- 0.3-1.0, if the underwater node did not participate in the election and failed to join any cluster. Calculate the additional reward corresponding to the underwater node. ,in, The initial energy of the underwater node is... The number of members in the cluster to which the underwater node belongs. For the overall score The optimal value is when the underwater node is an isolated node. =1, otherwise =0; according to the formula = +
[0009] Calculate the total reward corresponding to the underwater node. ; Based on the total reward The first Q-value table corresponding to the underwater node is updated.
[0010] Furthermore, the construction of the clustering structure based on each of the first Q-value tables includes: For any of the underwater nodes, the underwater node executes... The -greedy strategy, based on its own state index. With size equal to exploration rate The probability is explored in the action space with a size equal to 1- The probability of selecting the action with the largest current Q value is mapped to the action space. When the selected action is a candidate action, the underwater node becomes a candidate cluster head. When the selected action is a non-candidate action, the underwater node becomes a normal node. A connectivity check is performed on all candidate cluster heads, and isolated nodes are removed from the candidate cluster heads; the isolated nodes are candidate cluster heads that are outside the communication range of all other candidate cluster heads. Based on the set cluster head proportion, the size of each candidate cluster head is controlled; The candidate cluster heads that pass the connectivity check and the size control are elected as cluster head nodes, while the other candidate cluster heads become ordinary nodes. Each cluster head node broadcasts an announcement message; the announcement message includes the cluster head node's ID, location information, and remaining energy. For any of the aforementioned ordinary nodes, the ordinary node listens to each of the aforementioned announcement messages, and for any of the aforementioned announcement messages, according to the formula...
[0011] Calculate the corresponding comprehensive score Join the cluster containing the cluster head node with the highest overall score; where, The remaining energy mentioned in the notification message. The distance score is calculated based on the location information in the notification message. As the load penalty factor, The remaining energy is weighted as the second weight. The average distance to neighbors is the second weighted factor. The remaining energy is the second weight of the node degree. The average distance of the neighbors is the second weight. and the second weight of the node degree According to the remaining energy ratio first weight The average distance of the neighbors is the first weight. and the first weight of the node degree Obtained through normalization.
[0012] Further, performing intra-cluster routing in the clustered structure includes: For the ordinary node and the cluster head node belonging to the same cluster, a single-hop transmission is performed between the ordinary node and the cluster head node.
[0013] Further, the Bayesian optimization of the first Q-value table includes: Determine the combination of hyperparameters to be optimized; the combination of hyperparameters to be optimized includes the learning rate. Discount Factor Exploration rate Energy topology weights Remaining energy ratio of the first weight Average distance to neighbors is the first weighted factor. First weight of node degree ; Perform multiple rounds of iterative optimization until the combination of hyperparameters to be optimized converges to the optimal value, or the maximum number of rounds is reached; each round of iterative optimization includes the following steps: Based on the combination of hyperparameters to be optimized, the protocol flow of the efficient adaptive routing protocol system for the underwater sensor network is simulated. During the operation of the protocol process, simulation performance data is collected; the simulation performance data includes the average packet loss rate and the average energy consumption per packet. Calculate the objective function value for this round based on the simulation performance data; The combination of hyperparameters to be optimized is updated based on the objective function value.
[0014] Furthermore, the convergence node and each of the underwater nodes maintain a common second Q-value table, including: Establish the second Q-value table; the second Q-value table includes the plurality of horizontal and vertical coordinates and the corresponding second Q-values; the horizontal coordinates correspond to the transmitting nodes, and the vertical coordinates correspond to the receiving nodes; Whenever an event is detected in which the cluster head node is elected, a first update process is performed on the second Q-value table; Whenever an event is detected that the sending node has selected the next hop, a second update process is performed on the second Q-value table; The first update process includes: traversing all x-coordinate pairs in the second Q-value table; if both the sending node and the receiving node corresponding to the x-coordinate pair are cluster head nodes, and the second Q-value corresponding to the x-coordinate pair is valid, then the second Q-value is retained; if the second Q-value is negative and close to the marked retention value, then the second Q-value is positive; if the second Q-value corresponding to the x-coordinate pair is invalid, then the second Q-value is set to 0; if at least one of the sending node and the receiving node corresponding to the x-coordinate pair is a normal node, and the second Q-value corresponding to the x-coordinate pair is valid, then the second Q-value is negative and set as the marked retention value; if the second Q-value corresponding to the x-coordinate pair is negative, then the second Q-value is retained; if the sending node corresponding to the x-coordinate pair is a cluster head node and the receiving node is a convergence node, then the second Q-value is retained or set to 0; The second update process includes: when the sending node of the next hop is selected as the node... According to the formula
[0015]
[0016] Calculate the probability of successful transmission and packet loss probability ,in, For nodes The number of successfully received data packets sent to neighboring nodes. For nodes The total number of data packets sent to neighboring nodes; according to the formula
[0017]
[0018]
[0019]
[0020]
[0021]
[0022]
[0023]
[0024]
[0025]
[0026] Calculate the reward for successful forwarding ;in, For fixed penalty items, , and For weight parameters, As an energy reward, Represents a node The cost of remaining energy Represents a node The cost of remaining energy Represents a node The remaining energy is compared with the cluster average energy. Represents a node The remaining energy is compared with the cluster average energy. Represents a node The remaining energy, Represents a node The remaining energy, Represents a node initial energy, Represents a node initial energy, Represents a node The average energy of the cluster it belongs to. Represents a node The average energy of the cluster it belongs to. , , , and For parameters, Represents a node The Euclidean distance between the nearest convergence node and the nearest convergence node. This indicates a link quality reward. Represents a node With nodes All The average link quality of each adjacent node; according to the formula
[0027] Calculate the reward for failed forwarding; among which, For fixed penalty items; according to the formula
[0028] compute nodes The reward function; the node is rewarded according to the reward function. The second Q value corresponding to the next hop is updated.
[0029] Further, the step of performing inter-cluster routing in the clustering structure based on the second Q-value table includes: For any of the cluster head nodes, the cluster head node executes... -Greedy strategy, where size equals exploration rate The probability of exploring another cluster head node as the next hop in the second Q-value table is equal to 1- The cluster head node or sink node corresponding to the largest second Q value at present is selected as the next hop with a certain probability.
[0030] Furthermore, the at least one convergence node includes multiple surface convergence nodes and underwater mobile convergence nodes.
[0031] Furthermore, the underwater mobile convergence node is an autonomous underwater vehicle (AUV), which moves within the three-dimensional underwater monitoring area according to a preset cruise path and broadcasts beacon signals to the outside world.
[0032] Further, the step of performing inter-cluster routing in the clustering structure based on the second Q-value table includes: For any of the cluster head nodes, when the cluster head node detects the beacon signal, it determines the corresponding underwater mobile convergence node as the next hop based on the beacon signal.
[0033] The beneficial effects of this invention are as follows: The efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization in the embodiments realizes the application of Q-learning by executing hierarchical routing consisting of intra-cluster routing and inter-cluster routing, thereby easily finding efficient data forwarding paths and realizing data forwarding from underwater nodes to the sink node; by performing Bayesian optimization, key network control parameters such as hyperparameters can be automatically adjusted, enabling the entire routing system to adapt to different network conditions and continuously converge to the state of optimal global performance. Attached Figure Description
[0034] Figure 1 This is a schematic diagram of the structure of the efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization in the embodiment. Figure 2 This is a schematic diagram of the overall protocol flow executed by the efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization in the embodiment. Figure 3 This is a schematic diagram of the Q-learning network model executed by the efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization in the embodiment. Figure 4This is a schematic diagram illustrating the Bayesian optimization performed by the efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization in the embodiment. Figure 5 This is a schematic diagram of the protocol performance results of the efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization in the embodiment. Part (a) is a schematic diagram of end-to-end latency, part (b) is a schematic diagram of average energy consumption per data packet, and part (c) is a schematic diagram of packet loss rate. Detailed Implementation
[0035] Terminology Explanation: Autonomous Underwater Vehicle (AUV) is an underwater robot that navigates and operates autonomously without cables, relying on its built-in intelligence and energy. It features autonomy, cablelessness, self-powered operation, and the ability to be pre-programmed or make intelligent decisions. AUVs can carry communication equipment to function as aggregation nodes.
[0036] With the development of artificial intelligence technology, reinforcement learning can be introduced into routing design to address the dynamics and uncertainties of underwater environments. For example, adaptive clustering routing protocols based on reinforcement learning can solve the problems of uneven network energy consumption and high transmission latency in underwater communication routing scenarios. These protocols construct a state space by considering features such as node energy, neighbor density, and distance to the sink node, and design corresponding reward functions to dynamically elect cluster heads, thus optimizing the network's lifespan and data transmission paths to some extent. However, the above routing protocols may face the following significant limitations: on the one hand, the frequent exchange of information such as Q-values usually leads to substantial control overhead; on the other hand, if the reward function design is too simplistic and fails to comprehensively consider multi-dimensional factors such as distance, network topology, and node redundancy, it can easily result in slow learning convergence and non-globally optimal path selection, making it difficult to completely solve the problems of energy voids and premature node death. To overcome these challenges, it is possible to consider deeply integrating reinforcement learning with advanced network structure control (such as non-uniform clustering and hierarchical networks) and other intelligent algorithms to achieve more refined energy management. Specifically, the following technological evolutions can be considered: 1. Multidimensional optimization of reward function: By introducing node redundancy, or combining routing vector direction and node density (such as RVC method [4]), the reward function can simultaneously weigh energy, distance and data transmission direction, and guide data to be transmitted along an energy-efficient and directional path.
[0037] 2. Deep integration with hierarchical non-uniform clustering: For example, the network is layered according to the number of hops, the near-sink node layer adopts single-hop clustering, the outer layer adopts non-uniform clustering, and Q-learning is integrated into the cluster head election and multi-hop routing process. The information naturally collected in the clustering process is used to update the Q value, which greatly reduces communication overhead and effectively alleviates the sink node isolation problem.
[0038] 3. Collaborative optimization through the integration of multiple technologies: For example, integrating sine and cosine algorithms for global cluster head search, multi-agent reinforcement learning for adaptive cluster head rotation, Q-learning for energy-aware multi-hop routing, and Markov models for cluster head failure prediction, forming a complete solution that integrates global optimization, distributed learning, and fault tolerance, significantly improving the robustness and lifespan of the network in dynamic underwater environments.
[0039] However, the aforementioned reinforcement learning-based clustering routing schemes still have several key problems: First, their performance is highly dependent on the setting of hyperparameters such as learning rate and discount factor, and parameter tuning relies heavily on experience and manual trial and error, making it difficult to ensure that the network maintains optimal performance in different underwater scenarios. Second, in complex environments such as the deep sea, traditional static sink nodes are prone to causing their surrounding cluster heads to quickly deplete their energy due to excessive load, forming an "energy hole." At the same time, multi-hop relays from long-distance nodes to the sink introduce high communication costs and transmission delays. In addition, existing clustering routing mechanisms do not adequately consider security, lacking lightweight identification and dynamic trust management for malicious nodes (such as black hole attacks), making it difficult to ensure the reliability of data forwarding in complex underwater environments.
[0040] Based on the above principles, this embodiment provides an efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization. The efficient adaptive routing protocol system for underwater sensor networks includes at least one sink node and multiple underwater nodes. Specifically, when only one sink node is set, a conventional surface sink node or an autonomous underwater vehicle (AUV) can be used as the underwater mobile sink node.
[0041] Reference Figure 1 When multiple convergence nodes are set up, multiple conventional surface convergence nodes and multiple autonomous underwater vehicles (AUVs) can be set up as underwater mobile convergence nodes. The AUVs move within the three-dimensional underwater monitoring area according to a preset cruise path (such as random waypoints or scan lines), and periodically update their positions, broadcasting beacon signals to indicate their latest positions.
[0042] Reference Figure 1Each underwater node is deployed within a three-dimensional underwater monitoring area. Each underwater node constructs a clustered structure comprising multiple clusters using Q-learning. Typically, each cluster includes a cluster head node and multiple ordinary nodes. The cluster head node is the underwater node that participates in the election and is elected, while the ordinary nodes are underwater nodes that either did not participate in the election or participated but were not elected.
[0043] In this embodiment, each underwater node and each aggregation node can communicate with each other, thus forming a whole. When initially running the underwater sensor network efficient adaptive routing protocol system, all underwater sensor nodes can be randomly deployed in a designated three-dimensional underwater monitoring area, and the underwater sensor network efficient adaptive routing protocol system can be initially configured with reference to the network parameters shown in Table 1. The network size, the total number of sensor nodes, and the energy consumption model for data packet transmission and reception are shown in Table 1. In the communication parameters, the maximum communication radius of the nodes is set to 2000m, and the speed of sound for acoustic communication is set to 1500m / s. In the AUV system parameters, the number of AUVs is set to 6, and their initial positions are randomly and evenly distributed. After each round of cluster head election, the nodes broadcast the cluster head ID and location information, and the AUVs enter cruise mode, with their target being within the communication range of the nearest cluster head node.
[0044] Table 1 Network Parameter Table
[0045] When the efficient adaptive routing protocol system for underwater sensor networks is running, the following steps can be performed: S1. Each underwater node maintains its own first Q value table, constructs a cluster structure based on each first Q value table, and performs intra-cluster routing in the cluster structure; S2. The aggregation node and each underwater node maintain a common second Q-value table, and perform inter-cluster routing based on the second Q-value table in the cluster structure; S3. Perform Bayesian optimization on the first Q-value table.
[0046] In this embodiment, there is no necessary sequential order between steps S1-S3; for example, steps S1-S3 may be executed simultaneously.
[0047] In this embodiment, the principle of steps S1-S3 is as follows: Figure 2 As shown. (Refer to...) Figure 2When the efficient adaptive routing protocol system for underwater sensor networks is running, it specifically operates a Q-learning network model and a Bayesian optimizer. The Q-learning network model, implemented in steps S1-S2, performs cluster head election and dynamic clustering to form a cluster structure, enabling intra-cluster and inter-cluster routing, and establishing data forwarding paths. This allows each underwater node to send data packets to the sink node according to the forwarding path. The Bayesian optimizer, implemented in step S3, optimizes the hyperparameter combinations used in the Q-learning network model.
[0048] In this embodiment, when performing the step S1 in which each underwater node maintains its own first Q-value table, the following steps can be performed for any underwater node: S101. Underwater nodes perceive their own state characteristics and average distance to local neighbors. and mobility stability ; S102. Discretize each component in the state features to obtain the state index. ; S103. Establish the first Q-value table; S104. According to the formula (1) (2) Calculate the quality score corresponding to the underwater node. ;in, For topological quality factor, For energy topological weights, The remaining energy is weighted first. The average distance to neighbors is the first weighted factor. The first weight is the node degree; S105. Whenever a successful cluster head election event is detected, according to the formula... (3) Calculate the instant reward for the underwater node. ; S106. Whenever an event indicating that the cluster structure construction is complete is detected, according to the formula... ① If an underwater node participates in the election and is elected as the cluster head node: =2.5+ 1.5- 0.2+ 0.5- 5(4) ②If the underwater node did not participate in the election but successfully joined a cluster: =0.5+0.8 (5) ③ If the underwater node did not participate in the election and failed to join any cluster: =-0.8- 0.3-1.0 (6) Calculate the additional rewards corresponding to underwater nodes. ; S107. According to the formula = + (7) Calculate the total reward for the underwater node. ; S108. Based on total reward Update the first Q-value table corresponding to the underwater node.
[0049] In this embodiment, each underwater node executes steps S101-S107 to maintain its own first Q-value table. Taking one of the underwater nodes (e.g., node...) as an example... Taking (e.g.) as an example, steps S101-S107 will be explained.
[0050] In step S101, the underwater node senses its own state characteristics, including the remaining energy ratio. Distance from the aggregation node and normalized node degree Equal components, that is, state characteristics can be represented as continuous values ( , , The meanings of these components are as follows: Remaining energy ratio Underwater node (node) The ratio of the current remaining energy to the initial energy; Distance from the aggregation node Underwater node (node) The Euclidean distance between the nearest sink node and the nearest sink node can be the result after normalization. Normalized node degree Underwater node (node) The ratio of the number of neighbors of an underwater node to the maximum number of neighbors for each individual underwater node reflects the overall value of the underwater node (node). Local connectivity of ).
[0051] In step S101, the underwater node also senses the average distance to its local neighbors. and mobility stability The meanings of these data are as follows: Average distance to local neighbors : Represents an underwater node (node) The average distance of an underwater node (node) to all its neighboring nodes (i.e., other underwater nodes in the same cluster) can reflect the underwater node's (node's) distance to all its neighboring nodes. Intra-cluster communication efficiency, average distance to local neighbors It can be done through the following formula =
[0052] Calculation, where Represents underwater nodes (nodes) The average distance of a node to all its neighboring nodes (i.e., other underwater nodes in the same cluster). Represents underwater nodes (nodes) The communication range of the local neighbors is thus increased, thereby maximizing the average distance to the local neighbors. Normalize to the [0,1] interval; Motion stability In static network scenarios, this can be set to a constant, such as 0.5, to reflect the underwater node (node). Positional stability.
[0053] In step S102, the state features obtained in step S101 ( , , Each component in the ) is discretized to obtain the underwater node (node). ) state index .
[0054] In step S103, an underwater node (node) is established. The first Q-value table. In this embodiment, the dimension of the first Q-value table is... × ,in, Taking the cube to represent the state characteristics ( , , The number of discrete levels for each component in ) =2 indicates that there are two selectable actions in the action space: "candidate action" and "non-candidate action," meaning that the first Q-value table stores the state features ( , , The first Q-value corresponds to each combination of the action space and the action space. Therefore, the first Q-value maps to the action space, meaning that each first Q-value corresponds to a certain state feature ( , , The table lists the first Q values, including "election action" and "non-election action". The initial value for each first Q value in the first Q value table is 0.
[0055] In step S104, the underwater node (node) According to the formula (1) (2) Calculate underwater nodes (nodes) The corresponding quality score . In formula (1)-(2), For topological quality factor, For energy topological weights, The remaining energy is weighted first. The average distance to neighbors is the first weighted factor. It is the first weight for node degree. , , , and The remaining energy ratios obtained in step S101 are respectively Distance from the aggregation node Normalized node degree Average distance to local neighbors and mobility stability .
[0056] In step S104, energy topology weights Remaining energy ratio of the first weight Average distance to neighbors is the first weighted factor. First weight of node degree The equal weight values can be determined through Bayesian optimization.
[0057] In step S105, the underwater node (node) It can monitor the election process of cluster head nodes. Whenever an event is detected in which a cluster head node is successfully elected (the newly elected cluster head node can be a cluster head node or a regular node), the quality score calculated in step S104 is obtained through formula (3). Calculate the underwater nodes (nodes) Instant rewards .
[0058] In step S106, the underwater node (node) It can monitor the construction process of cluster structures. Whenever a cluster structure is detected to be completed, it will be based on the underwater nodes (nodes) In this election, the situation is as follows: underwater nodes (nodes) Whether the node participated in the election, was elected as the cluster head node, or successfully joined the cluster as a regular node, choose one of formulas (4), (5), or (6) to calculate the underwater node (node) status. Additional rewards .
[0059] Specifically, in formula (4), 2.5 is the base reward. For underwater nodes (nodes) The initial energy of ), therefore The 1.5 value represents the energy reward based on the remaining energy ratio; For underwater nodes (nodes) The number of members in the cluster, therefore - The value of 0.2 represents a load penalty, which can prevent underwater nodes (nodes) from being penalized. The cluster head node of the cluster is overloaded. The 0.5 value represents the distance efficiency bonus based on the average distance to local neighbors. The value can be 1 or 0. Specifically, if the underwater node (node) ) is an isolated node [i.e., an underwater node (node)] If a node is the cluster head and its cluster contains no ordinary nodes, then... =1, otherwise =0, therefore - Option 5 can impose additional penalties on isolated nodes.
[0060] In formula (5), For all calculated comprehensive scores The optimal value.
[0061] In step S107, the underwater nodes (nodes) are calculated. Instant rewards and extra rewards The sum of these values yields the underwater nodes (nodes). Total reward .
[0062] In step S108, based on the underwater node (node) Total reward For underwater nodes (nodes) The first Q-value table corresponding to the underwater node (node) is updated. Specifically, this can be done at the underwater node (node). The total reward is added to the first Q value in the corresponding first Q value table. This allows us to obtain the updated first Q-value table.
[0063] In this embodiment, when performing the step S1 of constructing the cluster structure based on each first Q-value table, for any underwater node, the following steps can be specifically performed: S109. Underwater Node Execution -greedy strategy, based on its own state index With size equal to exploration rate The probability is explored in the action space, with a value equal to 1- The probability of selecting the first Q value with the largest current value is mapped to the action in the action space. When the selected action is the election action, the underwater node becomes the candidate cluster head. When the selected action is the non-election action, the underwater node becomes a normal node. S110. Perform connectivity checks on all candidate cluster heads and remove isolated nodes from the candidate cluster heads; S111. Based on the set cluster head proportion, control the size of each candidate cluster head; S112. Candidate cluster heads that pass connectivity checks and size control are elected as cluster head nodes, while other candidate cluster heads become ordinary nodes; S113. Each cluster head node broadcasts a notification message to the outside world; S114. For any ordinary node, the ordinary node listens to each announcement message. For any announcement message, according to the formula... (8) Calculate the corresponding comprehensive score Join the cluster containing the cluster head node with the highest overall score.
[0064] In step S109, underwater nodes (nodes) are used. Taking underwater nodes (nodes) as an example )implement -greedy strategy. Specifically, underwater nodes (nodes) ) with size equal to exploration rate Explore actions in the action space with a probability equal to 1- The probability of selecting the first Q-value with the largest current value is mapped to the action in the action space. Thus, the underwater node (node) The underwater node (node) can explore or select either a "candidate action" or a "non-candidate action" within the action space. If the explored or selected action is a "candidate action," then the underwater node (node) can... If a node participates in the election to become a candidate cluster head, then if the explored or selected action is "not to run for election," then the underwater node (node) will be eliminated. (It does not participate in the election and becomes a regular node.) Among them, the exploration rate... It can be determined through Bayesian optimization.
[0065] Step S109 is executed for each underwater node to determine multiple candidate cluster heads, which together form a cluster head candidate set.
[0066] In step S110, the cluster head candidate set is filtered. Specifically, a connectivity check is performed on all candidate cluster heads in the cluster head candidate set, and candidate cluster heads that belong to isolated nodes (i.e., candidate cluster heads whose other candidate cluster heads are all outside their own communication range) are removed.
[0067] In step S111, the candidate cluster heads selected in step S110 are further filtered. Specifically, a cluster head percentage can be set. (Specifically, it can be 0.3), for all Each candidate cluster head is ranked according to its corresponding quality score. Sort in descending order, keeping only the top few. There are 10 candidate cluster heads. If the total number of candidate cluster heads is less than 100... If so, no processing is required. By executing step S111, the number of candidate cluster heads can be controlled.
[0068] In step S112, the candidate cluster heads that are retained through the connectivity check in step S110 and the size control in step S111 are determined as cluster head nodes, that is, these candidate cluster heads are successfully elected as cluster head nodes, while the candidate cluster heads that are filtered out by the connectivity check in step S110 or the size control in step S111 become ordinary nodes.
[0069] In step S113, each successfully elected cluster head node broadcasts an announcement message. The announcement message contains information such as the cluster head node's ID, location information, and remaining energy.
[0070] Step S114 is the process for ordinary nodes to join the cluster, and each ordinary node executes step S114. Taking any ordinary node as an example, when multiple cluster head nodes broadcast announcement messages, the ordinary node can listen to multiple announcement messages. The ordinary node can extract information such as the location and remaining energy of the cluster head node from each announcement message, and calculate the distance between itself and the cluster head node based on its own location and the location information of the cluster head node, and determine the distance score in a negative correlation. The load penalty factor is determined negatively based on the remaining energy. (This avoids overloading the cluster head node by adding too many ordinary nodes). Next, ordinary nodes are processed using the formula...
[0071]
[0072]
[0073] Remaining energy ratio with first weight Average distance to neighbors is the first weighted factor. First weight of node degree Normalize the three components to obtain the remaining energy ratio of the second weight. Average distance to neighbors (second weight) And node degree second weight Ordinary nodes are determined according to the formula. (8) Calculate the corresponding comprehensive score Since ordinary nodes can listen to multiple announcement messages, a corresponding comprehensive score can be calculated for each announcement message. Therefore, ordinary nodes can choose to join the cluster containing the cluster head node with the highest overall score. Specifically, ordinary nodes can determine all overall scores. The best value in Towards the optimal value The corresponding cluster head node initiates a cluster joining request. The cluster head node that receives the cluster joining request can execute a response, thereby enabling ordinary nodes to join the cluster corresponding to the cluster head node.
[0074] In this embodiment, by executing steps S101-S114, each underwater node can become a cluster head node and a normal node, respectively, and the normal nodes can join the cluster corresponding to the cluster head node, thereby forming... Figure 1 The diagram shows a clustering structure comprising one or more clusters. Therefore, by executing steps S110-S114, reinforcement learning can dynamically and adaptively elect cluster heads and form a clustering structure to optimize network energy efficiency and topology quality.
[0075] In this embodiment, to leverage the advantages of Q-learning for inter-cluster routing, the cluster head percentage is controlled to a relatively high value of 0.3 during the cluster head election process. Therefore, each cluster is relatively small, and single-hop transmission within the cluster is superior to multi-hop transmission. Thus, a strategy of single-hop transmission from ordinary nodes to the cluster head node is adopted when performing intra-cluster routing. Specifically, refer to... Figure 1 If a regular node needs to forward data packets to the outside world, then the regular node selects the cluster head node of its cluster as the next hop, that is, it sends data packets to the cluster head node.
[0076] In this embodiment, a forced clustering mechanism can be implemented: a timeout period is set. For ordinary nodes that have not joined any cluster after the timeout, the system forces them to connect to the nearest cluster head node, even if the connection is not optimal, thereby achieving full network coverage.
[0077] In this embodiment, when performing the step S2 where the convergence node and each underwater node maintain a common second Q-value table, the following steps can be specifically performed: S201. Establish the second Q-value table; S202. Whenever a cluster head node election is successfully detected, a first update process is performed on the second Q-value table. The first update process includes: traversing all x-axis-y-axis pairs in the second Q-value table; if both the sending node and the receiving node corresponding to the x-axis-y-axis pair are cluster head nodes, if the second Q-value corresponding to the x-axis-y-axis pair is valid, then the second Q-value is retained; if the second Q-value is negative and close to the marked retention value, then the second Q-value is positive; if the second Q-value corresponding to the x-axis-y-axis pair is invalid, then the second Q-value is set to 0; if at least one of the sending node and the receiving node corresponding to the x-axis-y-axis pair is a normal node, if the second Q-value corresponding to the x-axis-y-axis pair is valid, then the second Q-value is negative and set as the marked retention value; if the second Q-value corresponding to the x-axis-y-axis pair is negative, then the second Q-value is retained; if the sending node corresponding to the x-axis-y-axis pair is a cluster head node and the receiving node is a sink node, then the second Q-value is retained or set to 0. S203. Whenever an event is detected in which a sending node selects a next hop, a second update process is performed on the second Q-value table. The second update process includes: when the sending node that has selected the next hop is node... According to the formula (9) (10) Calculate the probability of successful transmission and packet loss probability ,in, For nodes The number of successfully received data packets sent to neighboring nodes. For nodes The total number of data packets sent to neighboring nodes; according to the formula (11) (12) (13) (14) (15) (16) (17) (18) (19) (20) Calculate the reward for successful forwarding ;in, For fixed penalty items, , and For weight parameters, As an energy reward, Represents a node The cost of remaining energy Represents a node The cost of remaining energy Represents a node The remaining energy is compared with the cluster average energy. Represents a node The remaining energy is compared with the cluster average energy. Represents a node The remaining energy, Represents a node The remaining energy, Represents a node initial energy, Represents a node initial energy, Represents a node The average energy of the cluster it belongs to. Represents a node The average energy of the cluster it belongs to. , , , and For parameters, Represents a node The Euclidean distance between the nearest convergence node and the nearest convergence node. This indicates a link quality reward. Represents a node With nodes All The average link quality of each adjacent node; according to the formula (twenty one) Calculate the reward for failed forwarding; among which, For fixed penalty items; according to the formula (twenty two) compute nodes The reward function; based on the reward function, the nodes... Update the second Q value corresponding to the next hop.
[0078] In step S201, the established second Q-value table is a × A two-dimensional matrix, where This represents the total number of all nodes (including sink nodes and underwater nodes). The x-axis of the second Q-value table represents the sequence number of the sending node, and the y-axis represents the sequence number of the receiving node. Here, "sending node" and "receiving node" refer to the position of each node (including sink nodes and underwater nodes) in the data transmission process; each node can be either a sending node or a receiving node. Each x-axis-y-axis pair (i.e., a specific x-axis and a specific y-axis) maps to a second Q-value in the second Q-value table.
[0079] For the second Q-value table, all the second Q-values in it can be initialized to an invalid value (specifically, it can be a very large negative value, such as -1e10).
[0080] In step S202, whenever an event of successfully electing a cluster head node is detected, a first update process is performed on the second Q-value table. The first update process includes: (1) Traverse all x-coordinate-y-coordinate pairs in the second Q-value table. Specifically, the following situations may occur: (2) If both the sending node and the receiving node corresponding to the x-axis-y-axis pair are cluster head nodes (i.e., the sending node corresponding to the x-axis is a cluster head node, and the receiving node corresponding to the y-axis is also a cluster head node): ① When the second Q value corresponding to the x-axis-y-axis pair is a valid value (greater than the invalid threshold), then this second Q value is retained (migrated to the updated second Q value table); ② When the second Q value is negative and close to the mark retention value (the difference between it and the mark retention value is less than the threshold), then this second Q value is positive (i.e., the absolute value remains unchanged, only the sign is changed); ③ When the second Q value corresponding to the x-axis-y-axis pair is an invalid value (e.g., equal to -1e10), then this second Q value is set to 0; (3) If at least one of the sending node and receiving node corresponding to the x-axis-y-axis pair is a normal node (for example, the sending node corresponding to the x-axis belongs to the cluster head node, and the receiving node corresponding to the y-axis does not belong to the cluster head node): ① When the second Q value corresponding to the x-axis-y-axis pair is a valid value, the second Q value is negative and set as the reserved value; ② When the second Q value corresponding to the x-axis-y-axis pair is negative, the second Q value is reserved (migrated to the updated second Q value table); (4) If the sending node corresponding to the x-axis-y-axis pair is a cluster head node and the receiving node is a sink node, then the second Q value is retained (if there is a historical value) or set to 0 (if there is no historical value).
[0081] In step S203, whenever an event indicating that the sending node has selected a next hop is detected, a second update process is performed on the second Q-value table. Specifically, in the second update process, the node... Taking the selection of the next-hop sending node as an example, obtain the node. To neighboring nodes (e.g., node) Number of successfully received data packets and nodes Total number of data packets sent to neighboring nodes The nodes are calculated according to formulas (9) and (10). Successful transmission probability and packet loss probability .
[0082] Next, the nodes are calculated according to formulas (11)-(20). Reward for successful forwarding In formulas (11)-(20), the fixed penalty term is... Represents a node Transmission costs, energy rewards Nodes can be encouraged Nodes with high remaining energy and energy higher than the cluster average energy are selected as the next hop, parameters. , , The value can be set to = = =0.5; Delayed rewards can incentivize nodes. Select nodes closest to the sink node as the next hop (the smaller the distance, the higher the reward); link quality reward Nodes can be encouraged Choose nodes with high transmission success rates and good neighbor link quality as the next hop. Represents a node With nodes (node The success rate of data packet transmission between neighboring nodes.
[0083] Next, the nodes are calculated according to formula (21). Reward for failed forwarding In formula (21), the fixed penalty term... With fixed penalty items Their functions are similar. The penalty can be imposed on nodes that are far away. = , Energy reward indicates that only nodes are considered. The remaining energy and the corresponding cluster average energy, = , The link quality reward indicates that only nodes are considered. To the node The success rate of direct transmission.
[0084] Next, based on the calculated nodes Successful transmission probability Packet loss probability Reward for successful forwarding Rewards for failed forwarding The nodes are calculated according to formula (22). reward function .
[0085] Finally, according to the reward function For nodes The second Q value corresponding to the next hop it selects is updated. Specifically, for nodes... The sending node determines the x-coordinate in the second Q-value table, and its selected next hop determines the y-coordinate in the second Q-value table. Based on the selected x-coordinate and y-coordinate, the corresponding second Q-value is determined in the second Q-value table, and a reward function is added to this second Q-value. , and obtain the updated second Q value corresponding to the x and y coordinates.
[0086] In this embodiment, when performing the step S2 of executing inter-cluster routing based on the second Q-value table in the cluster structure, any cluster head node can execute... -Greedy strategy. Specifically, any cluster head node can be accessed with a size equal to the exploration rate. The probability of exploring another cluster head node as the next hop in the second Q-value table is equal to 1- The probability is used to select the node corresponding to the largest second Q value in the current second Q value table (specifically, it can be a cluster head node or a sink node) as the next hop.
[0087] In this embodiment, the underwater mobile convergence node (AUV) can be set as the highest priority next hop. Specifically, for any cluster head node, if the cluster head node detects a beacon signal, it will preferentially select the corresponding underwater mobile convergence node (AUV) as its next hop and directly send the data packets it needs to forward to the underwater mobile convergence node (AUV). In this way, the underwater mobile convergence node (AUV) can collect data packets sent by each cluster head node during its cruise and finally surface or transmit the data packets back to the shore-based center.
[0088] Specifically, since AUVs are mobile, the set of cluster head nodes (accessible set) that can be accessed within their communication range (2000m) is also dynamically changing, thereby realizing dynamic aggregation nodes and improving the speed at which data reaches the aggregation nodes.
[0089] In this embodiment, by executing steps S1-S2, a hierarchical routing consisting of intra-cluster routing and inter-cluster routing can be realized, thereby enabling the application of Q-learning. This makes it easier to find efficient data forwarding paths and realize data forwarding from underwater nodes to the aggregation node. By distributing AUVs within the three-dimensional underwater monitoring area, accessing multiple cluster head nodes and directly receiving data packets, and periodically updating their positions to simulate cruise behavior, an AUV collaborative working mechanism can be realized, which helps to reduce the length of the data forwarding path and improve data forwarding efficiency.
[0090] In this embodiment, it can be as follows: Figure 3 As shown, after each execution of step S1-S2, a round of step S1-S2 is calculated. Performance data for this round of step S1-S2 is collected, such as packet loss rate and data forwarding path length. Then, a new round of step S1-S2 is executed, and performance parameters are collected again. After repeating step S1-S2 for multiple rounds (e.g., 50 rounds), the network configuration of the round of step S1-S2 with the best performance parameters (e.g., the cluster head node election result and the clustering result of ordinary nodes) is selected as the optimal network configuration. This optimal network configuration is saved and used in subsequent runs, thereby achieving reinforcement learning of the routing protocol.
[0091] In this embodiment, when performing step S3, which is the Bayesian optimization of the first Q-value table, the following steps can be performed: S301. Determine the combination of hyperparameters to be optimized; S302. Perform multiple rounds of iterative optimization until the combination of hyperparameters to be optimized converges to the optimal value, or the maximum number of rounds is reached; any round of iterative optimization includes the following steps: S30201. Based on the combination of hyperparameters to be optimized, simulate the protocol flow of the efficient adaptive routing protocol system for underwater sensor networks. S30202. During the operation of the protocol flow, collect simulation performance data; S30203. Calculate the objective function value for this round based on the simulation performance data; S30204. Update the combination of hyperparameters to be optimized based on the objective function value.
[0092] In this embodiment, the principle of step S3 is as follows: Figure 4 As shown.
[0093] Reference Figure 4 In step S301, the learning rate can be... Discount Factor Exploration rate Energy topology weights Remaining energy ratio of the first weight Average distance to neighbors is the first weighted factor. First weight of node degree Equal hyperparameters are used as combinations of hyperparameters to be optimized.
[0094] Next, a multi-round iterative optimization process is executed. Each round of iterative optimization includes steps S30201-S30204. Taking one round of iterative optimization as an example, the process will be explained.
[0095] In step S30201, the protocol flow of the efficient adaptive routing protocol system for underwater sensor networks is simulated based on the hyperparameter combination to be optimized. Specifically, based on the specific values of the hyperparameter combination to be optimized in step S301 (the initial value is used if it is the first round of iterative optimization, and the optimized value is used if it is the subsequent rounds of iterative optimization), the clustering, routing, and data collection processes in steps S1-S2 are simulated, and the forwarding of data packets is simulated during this round of iterative optimization.
[0096] In step S30202, during the execution of step S30201, simulation performance data is collected. In this embodiment, the simulation performance data to be collected includes the average packet loss rate. (The average packet loss rate of each round of iterative optimization process) and the average energy consumption per packet Data such as (average energy consumption per data packet in each round of iterative optimization processes that have been executed).
[0097] In step S30203, the Bayesian optimizer calculates the formula...
[0098] Calculate the objective function value for this round. .
[0099] In step S30204, the Bayesian optimizer calculates the objective function value for this round. A surrogate model of the objective function is constructed, and based on the acquisition function, the next set of hyperparameter values most likely to improve performance is selected, thereby updating the hyperparameter combination to be optimized. Specifically, the Bayesian optimizer can select values that improve the objective function value. The optimization objective is to minimize the value of the combination of hyperparameters to be optimized, thereby enabling the design of routing protocols with higher packet delivery rates and lower energy consumption.
[0100] After executing step S30204, if the combination of hyperparameters to be optimized has converged to the optimal value, then output such a combination of hyperparameters to be optimized. Alternatively, execute an iterative optimization process for the maximum number of rounds (e.g., 100 rounds) and select the optimal combination of hyperparameters to be optimized.
[0101] In this embodiment, by executing step S3, outer Bayesian optimization can be achieved, thereby automatically adjusting the hyperparameters and key network control parameters of the reinforcement learning algorithm and improving the adaptability of the entire routing system.
[0102] In summary, the overall technical principle of the efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization in this embodiment is as follows: 1. By introducing the Q-learning algorithm, nodes can adaptively decide whether to run for cluster head node based on multi-dimensional states such as their remaining energy, distance from moving nodes, and density of neighboring nodes, thereby optimizing the network topology and extending network lifetime. 2. Perform intra-cluster routing and inter-cluster routing separately, and use Q-learning to dynamically select the optimal relay path, effectively reducing data transmission energy consumption and latency; 3. By using autonomous underwater vehicles as mobile aggregation nodes, data can be dynamically collected through their periodic cruising behavior, overcoming the "hot zone" problem caused by fixed aggregation nodes, reducing the number of multi-hop relays of data packets, and significantly reducing the overall network energy consumption. 4. Through the outer Bayesian optimization framework, the hyperparameters and key network control parameters of the reinforcement learning algorithm are automatically adjusted, enabling the entire routing system to adapt to different network conditions and continuously converge to the state of optimal global performance. 5. Through the above measures, the overall performance of underwater wireless sensor networks in terms of energy consumption, latency, reliability, and security will be comprehensively improved.
[0103] In this embodiment, the performance of the underwater sensor network high-efficiency adaptive routing protocol system under different total number of nodes is measured, including end-to-end latency, average energy consumption per data packet, and packet loss rate. The results are as follows: Figure 5 Parts (a), (b), and (c) are shown.
[0104] according to Figure 5 The protocol performance results shown demonstrate that the efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization in this embodiment achieves local intelligent decision-making through the Q-learning algorithm, global parameter tuning through Bayesian optimization, and enhanced system robustness through AUV mobile sinking, ultimately achieving the following comprehensive technical effects: 1. Significantly improves energy efficiency and network lifespan: Intelligent clustering and energy balancing mechanisms avoid energy gaps, and AUV mobile data collection significantly reduces total communication energy consumption; 2. Effectively reduce end-to-end transmission latency: Optimized hierarchical routing and short-distance access provided by AUV reduce the number of multi-hop forwardings and queuing time for data packets; 3. Enhanced network adaptability and robustness: In the face of node failures, topology changes, or different deployment scenarios, the system can automatically adjust its strategies through learning and optimization to maintain high performance; 4. Achieve full system automation and optimization: The introduction of the Bayesian optimization layer frees the burden of manual parameter tuning, enabling the protocol to be "plug and play" and autonomously optimize, making it particularly suitable for complex and ever-changing real-world application environments.
[0105] It should be noted that, unless otherwise specified, when a feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. Furthermore, the descriptions of "upper," "lower," "left," and "right" used in this disclosure are only relative to the relative positional relationships of the components of this disclosure in the accompanying drawings. The singular forms "a" and "the" used in this disclosure are also intended to include the plural forms, unless the context clearly indicates otherwise. Moreover, unless otherwise defined, all technical and scientific terms used in this embodiment have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this embodiment specification is only for describing particular embodiments and is not intended to limit the invention. The term "and / or" as used in this embodiment includes any combination of one or more of the associated listed items.
[0106] It should be understood that although the terms first, second, third, etc., may be used to describe various elements in this disclosure, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, a first element may also be referred to as a second element without departing from the scope of this disclosure, and similarly, a second element may also be referred to as a first element. The use of any and all instances or exemplary language (“e.g.,” “such as,” etc.) provided in this embodiment is intended only to better illustrate embodiments of the invention and, unless otherwise required, does not impose a limitation on the scope of the invention.
[0107] It should be recognized that embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable storage medium. The method can be implemented using standard programming techniques—including a non-transitory computer-readable storage medium configured with a computer program, wherein such a storage medium causes the computer to operate in a specific and predefined manner—according to the methods and drawings described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. Furthermore, for this purpose, the program can run on a programmed application-specific integrated circuit (ASIC).
[0108] Furthermore, the procedures described in this embodiment can be performed in any suitable order unless otherwise indicated by this embodiment or otherwise obviously contradict the context. The procedures (or variations and / or combinations thereof) described in this embodiment can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented by hardware or a combination thereof as code (e.g., executable instructions, one or more computer programs, or one or more applications) that commonly executes on one or more processors. A computer program includes a plurality of instructions executable by one or more processors.
[0109] Furthermore, the method can be implemented in any suitable type of computing platform, including but not limited to personal computers, minicomputers, mainframes, workstations, networked or distributed computing environments, standalone or integrated computer platforms, or in communication with charged particle tools or other imaging devices, etc. Aspects of the invention can be implemented as machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it is readable by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the processes described herein. Furthermore, the machine-readable code, or portions thereof, can be transmitted via wired or wireless networks. The invention of this embodiment includes these and other different types of non-transitory computer-readable storage media when such media comprises instructions or programs that implement the steps above in conjunction with a microprocessor or other data processor. When programmed according to the methods and techniques of the invention, the invention also includes the computer itself.
[0110] A computer program can be applied to input data to perform the functions of this embodiment, thereby transforming the input data to generate output data stored in non-volatile memory. The output information can also be applied to one or more output devices, such as a display. In a preferred embodiment of the invention, the transformed data represents physical and tangible objects, including specific visual depictions of physical and tangible objects generated on the display.
[0111] The above are merely preferred embodiments of the present invention. The present invention is not limited to the above-described embodiments. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention, as long as they achieve the technical effects of the present invention by the same means, should be included within the scope of protection of the present invention. Within the scope of protection of the present invention, the technical solutions and / or implementation methods can have various modifications and variations.
Claims
1. A highly efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization, characterized in that, The underwater sensor network high-efficiency adaptive routing protocol system includes: At least one aggregation node; Multiple underwater nodes; each underwater node is deployed within a three-dimensional underwater monitoring area, each underwater node maintains its own first Q-value table, constructs a cluster structure based on each first Q-value table, performs intra-cluster routing in the cluster structure, performs Bayesian optimization on the first Q-value table, the convergence node and each underwater node maintain a common second Q-value table, and performs inter-cluster routing based on the second Q-value table in the cluster structure; The construction of the clustering structure based on each of the first Q-value tables includes: For any of the underwater nodes, the underwater node executes... -greedy strategy, based on its own state index With size equal to exploration rate The probability is explored in the action space, with a value equal to 1- The probability of selecting the action with the largest current Q value is mapped to the action space. When the selected action is a candidate action, the underwater node becomes a candidate cluster head. When the selected action is a non-candidate action, the underwater node becomes a normal node. A connectivity check is performed on all candidate cluster heads, and isolated nodes are removed from the candidate cluster heads; the isolated nodes are candidate cluster heads that are outside the communication range of all other candidate cluster heads. Based on the set cluster head proportion, the size of each candidate cluster head is controlled; The candidate cluster heads that pass the connectivity check and the size control are elected as cluster head nodes, while the other candidate cluster heads become ordinary nodes. Each cluster head node broadcasts an announcement message; the announcement message includes the cluster head node's ID, location information, and remaining energy. For any of the aforementioned ordinary nodes, the ordinary node listens to each of the aforementioned announcement messages, and for any of the aforementioned announcement messages, according to the formula... Calculate the corresponding comprehensive score Join the cluster containing the cluster head node with the highest overall score; where, The remaining energy mentioned in the notification message. The distance score is calculated based on the location information in the notification message. To normalize the node degree, As the load penalty factor, The remaining energy is weighted as the second weight. The average distance to neighbors is the second weighted factor. The remaining energy is the second weight of the node degree. The average distance of the neighbors is the second weight. and the second weight of the node degree According to the remaining energy ratio first weight The average distance of the neighbors is the first weight. and the first weight of the node degree Obtained through normalization.
2. The efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization according to claim 1, characterized in that, Each of the underwater nodes maintains its own first Q-value table, including: For any of the underwater nodes: The underwater node senses its own state characteristics and the average distance to its local neighbors. and mobility stability The state characteristics include the remaining energy ratio. Distance from the aggregation node and normalized node degree ; Each component in the state feature is discretized to obtain the state index. ; Establish the first Q-value table; the first Q-value table includes the discrete level number of each component in the state feature and the corresponding first Q-value; the first Q-value is mapped to the action space, which includes campaign actions and non-campaign actions; According to the formula Calculate the quality score corresponding to the underwater node. ;in, For topological quality factor, For energy topological weights, The remaining energy is weighted first. The average distance to neighbors is the first weighted factor. The first weight is the node degree; Whenever a successful cluster head election is detected, according to the formula Calculate the instant reward corresponding to the underwater node. ; Whenever an event indicating that the cluster structure construction is complete is detected, according to the formula... =2.5+ 1.5- 0.2+ 0.5- 5. If the underwater node participates in the election and is elected as the cluster head node =0.5+0.8 If the underwater node did not participate in the election and successfully joined a cluster =-0.8- 0.3-1.0, if the underwater node did not participate in the election and failed to join any cluster. Calculate the additional reward corresponding to the underwater node. ,in, The initial energy of the underwater node is... The number of members in the cluster to which the underwater node belongs. For the overall score The optimal value is when the underwater node is an isolated node. =1, otherwise =0; according to the formula = + Calculate the total reward corresponding to the underwater node. ; According to the total reward The first Q-value table corresponding to the underwater node is updated.
3. The efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization according to claim 1, characterized in that, The execution of intra-cluster routing in the clustered structure includes: For the ordinary node and the cluster head node belonging to the same cluster, a single-hop transmission is performed between the ordinary node and the cluster head node.
4. The efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization according to claim 1, characterized in that, The Bayesian optimization of the first Q-value table includes: Determine the combination of hyperparameters to be optimized; the combination of hyperparameters to be optimized includes the learning rate. Discount Factor Exploration rate Energy topology weights Remaining energy ratio of the first weight Average distance to neighbors is the first weighted factor. First weight of node degree ; Perform multiple rounds of iterative optimization until the combination of hyperparameters to be optimized converges to the optimal value, or the maximum number of rounds is reached; each round of iterative optimization includes the following steps: Based on the combination of hyperparameters to be optimized, the protocol flow of the efficient adaptive routing protocol system for the underwater sensor network is simulated. During the operation of the protocol process, simulation performance data is collected; the simulation performance data includes the average packet loss rate and the average energy consumption per packet. Calculate the objective function value for this round based on the simulation performance data; The combination of hyperparameters to be optimized is updated based on the objective function value.
5. The efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization according to claim 1, characterized in that, The convergence node and each of the underwater nodes maintain a common second Q-value table, including: Establish the second Q-value table; the second Q-value table includes multiple horizontal and vertical coordinates and corresponding second Q-values; the horizontal coordinates correspond to the transmitting nodes, and the vertical coordinates correspond to the receiving nodes; Whenever an event is detected in which the cluster head node is elected, a first update process is performed on the second Q-value table; Whenever an event is detected that the sending node has selected the next hop, a second update process is performed on the second Q-value table; The first update process includes: traversing all x-coordinate pairs in the second Q-value table; if both the sending node and the receiving node corresponding to the x-coordinate pair are cluster head nodes, and the second Q-value corresponding to the x-coordinate pair is valid, then the second Q-value is retained; if the second Q-value is negative and close to the marked retention value, then the second Q-value is positive; if the second Q-value corresponding to the x-coordinate pair is invalid, then the second Q-value is set to 0; if at least one of the sending node and the receiving node corresponding to the x-coordinate pair is a normal node, and the second Q-value corresponding to the x-coordinate pair is valid, then the second Q-value is negative and set as the marked retention value; if the second Q-value corresponding to the x-coordinate pair is negative, then the second Q-value is retained; if the sending node corresponding to the x-coordinate pair is a cluster head node and the receiving node is a convergence node, then the second Q-value is retained or set to 0; The second update process includes: when the sending node of the next hop is selected as the node... According to the formula Calculate the probability of successful transmission and packet loss probability ,in, For nodes The number of successfully received data packets sent to neighboring nodes. For nodes The total number of data packets sent to neighboring nodes; according to the formula Calculate the reward for successful forwarding ;in, For fixed penalty items, , and For weight parameters, As an energy reward, Represents a node The cost of remaining energy Represents a node The cost of remaining energy Represents a node The remaining energy is compared with the cluster average energy. Represents a node The remaining energy is compared with the cluster average energy. Represents a node The remaining energy, Represents a node The remaining energy, Represents a node initial energy, Represents a node initial energy, Represents a node The average energy of the cluster it belongs to. Represents a node The average energy of the cluster it belongs to. , , , and For parameters, Represents a node The Euclidean distance between the nearest convergence node and the nearest convergence node. This indicates a link quality reward. Represents a node With nodes All The average link quality of each adjacent node; according to the formula Calculate the reward for failed forwarding; among which, For fixed penalty items; according to the formula compute nodes The reward function; the node is rewarded according to the reward function. The second Q value corresponding to the next hop is updated.
6. The efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization according to claim 5, characterized in that, The step of performing inter-cluster routing in the clustering structure based on the second Q-value table includes: For any of the cluster head nodes, the cluster head node executes... -Greedy strategy, where size equals exploration rate The probability of exploring another cluster head node as the next hop in the second Q-value table is equal to 1- The cluster head node or sink node corresponding to the largest second Q value at present is selected as the next hop with a certain probability.
7. The efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization according to claim 2, characterized in that, The at least one convergence node includes multiple surface convergence nodes and underwater mobile convergence nodes.
8. The efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization according to claim 7, characterized in that, The underwater mobile convergence node is an autonomous underwater vehicle (AUV). The AUV moves within the three-dimensional underwater monitoring area according to a preset cruise path and broadcasts beacon signals to the outside world.
9. The efficient adaptive routing protocol system for underwater sensor networks based on Q-learning and Bayesian optimization according to claim 8, characterized in that, The step of performing inter-cluster routing in the clustering structure based on the second Q-value table includes: For any of the cluster head nodes, when the cluster head node detects the beacon signal, it determines the corresponding underwater mobile convergence node as the next hop based on the beacon signal.
Citation Information
Patent Citations
Static UWSNs routing protocol optimization method and system based on convergence guidance
CN120050740A